English Language and Linguistics· 2026Q1
Metadata and annotation in historical text corpora: standards, challenges and best practices
- 0citations
- Q1SCImago
- 2026year
Short summary
A new framework for historical text corpora aligns with FAIR principles, using a minimalist encoding strategy (plain text linked to Omeka/Dublin Core metadata) for the MetaLing Corpus (1M tokens, 1500-1700 English metalanguage) to enhance discoverability and sustainability.
AI-generated from the title and abstract; the full text is not read.
Key points
- Historical text corpora require robust metadata curation and annotation for long-term accessibility and interoperability.
- Adherence to FAIR principles (Findable, Accessible, Interoperable, Reusable) is crucial for digital humanities research transparency.
- The MetaLing Corpus (1M tokens, 1500-1700 English metalanguage) adopted a minimalist encoding (plain text + Omeka/Dublin Core) due to TEI/Sketch Engine compatibility issues.
- Key challenges include orthographic variation and the limitations of automated tagging for historical data.
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract The construction of historical text corpora presents distinctive challenges in terms of metadata curation, annotation practices and long-term accessibility. This article explores how standards-based approaches can enhance the discoverability, interoperability and sustainability of historical linguistic data. Emphasis is placed on aligning corpus design with the FAIR principles (Findable, Accessible, Interoperable, Reusable) (Wilkinson et al. 2016), which are increasingly important for research transparency and cross-platform integration in the digital humanities. After outlining the conceptual landscape of metadata and annotation in corpus linguistics, the article examines several existing projects that demonstrate the operationalisation of metadata standards across genres and temporal ranges. These include corpora that integrate CMDI profiles for complex resources (Paquot et al. 2024), and others that have transformed legacy metadata into machine-readable formats to facilitate data exchange and semantic enrichment (Fallucchi & De Luca 2020). Key issues addressed include orthographic variation, the limits of automated tagging tools for historical data (Pettersson & Megyesi 2018) and the need to balance standardisation with the preservation of linguistic idiosyncrasies. The article explores a case study of the MetaLing Corpus , a one-million-token historical corpus of English metalanguage from 1500 to 1700 (Andreani & Russo 2026). Practical constraints and editorial decisions are discussed. A minimalist encoding strategy was adopted following compatibility issues with TEI and Sketch Engine (Kilgarriff et al. 2004; Kilgarriff et al. 2014), resulting in a plain-text corpus linked to metadata managed through Omeka using Dublin Core (Caplan 2003). This study contributes to ongoing discussions on metadata sustainability, annotation design and best practices for corpus construction in underrepresented linguistic domains. It advocates for adaptable, transparent approaches that foster cross-disciplinary collaboration and position corpora as evolving infrastructures within broader digital ecosystems.
The authors' abstract, as published at the source. English Language and Linguistics, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Music
MusicArts and Humanities