Epistemic Biases and Algorithmic Hegemony in Automated Classification Systems
Investigating how automated subject indexing and language models reproduce historical colonial and cultural biases in Dewey Decimal Classification (DDC) and Library of Congress Subject Headings (LCSH).
"To what extent do transformer-based bibliographic cataloging pipelines amplify historical Eurocentric biases present in legacy classification schedules?"
5 Open Access Full-Texts
60% Literature Synthesized
2 High-Impact Niches Identified
Target: 15,000 words (APA 7th)
Recent Synthesized Papers
Classification Ethics in the Age of Automated Language Models: A Critique of Machine-Assisted Subject Indexing
Dr. Elena Rostova, Marcus Vance • Journal of the Association for Information Science and Technology (JASIST) (2024)
Decolonizing Subject Headings: Epistemic Injustice and the Structural Limits of Controlled Vocabularies
Dr. Sarah K. Jenkins, Kweku Mensah • Journal of Documentation (2023)
Scientometric Trajectories of Open Access Mandates: A Ten-Year Longitudinal Evaluation of FAIR Compliance
Dr. Henrik Lindqvist, Claire Dubois • Journal of Informetrics (2023)
Critical Research Gaps
Lack of Standardized Epistemic Auditing Benchmarks for Non-Latin Catalog Records
Current research on automated classification bias (e.g., Rostova & Vance 2024) focuses predominantly on Latin-script Western European national library catalogs. There is no open-source benchmark evaluating transformer subject classification on Arabic, Cyrillic, Indic, or CJK bibliographic corpora.
Poly-hierarchical Linked Data Interoperability with Monolithic DDC/LCSH Trees
While Jenkins & Mensah (2023) emphasize indigenous poly-hierarchical thesauri, existing library management systems (LMS) and BIBFRAME 2.0 endpoints fail to parse non-tree relational graph structures without data loss (Gomez, 2023).