English Language and Linguistics· 2026Q1
Tarihi metin derlemlerinde meta veri ve sınıflama: standartlar, zorluklar ve en iyi uygulamalar
Metadata and annotation in historical text corpora: standards, challenges and best practices
- 0atıf
- Q1SCImago
- 2026yıl
Kısa özet
Tarihi metin derlemleri için yeni bir çerçeve, bulunabilirliği ve sürdürülebilirliği artırmak amacıyla, MetaLing Derlemi (1 M belirteç, 1500-1700 İngilizce metalanguage) için minimalist bir kodlama stratejisi (düz metin, Omeka/Dublin Core meta verisine bağlı) kullanarak FAIR prensiplerine uyum sağlamaktadır.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Tarihi metin derlemleri, uzun vadeli erişilebilirlik ve birlikte çalışabilirlik için sağlam meta veri kürasyonu ve sınıflaması gerektirir.
- Dijital beşeri bilimler araştırmalarının şeffaflığı için FAIR prensiplerine (Bulunabilir, Erişilebilir, Birlikte Çalışabilir, Yeniden Kullanılabilir) bağlılık esastır.
- MetaLing Derlemi (1 M belirteç, 1500-1700 İngilizce metalanguage), TEI/Sketch Engine uyumluluk sorunları nedeniyle minimalist bir kodlama (düz metin + Omeka/Dublin Core) benimsemiştir.
- Temel zorluklar arasında ortografik varyasyon ve tarihi veriler için otomatik etiketlemenin sınırlamaları yer almaktadır.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Abstract The construction of historical text corpora presents distinctive challenges in terms of metadata curation, annotation practices and long-term accessibility. This article explores how standards-based approaches can enhance the discoverability, interoperability and sustainability of historical linguistic data. Emphasis is placed on aligning corpus design with the FAIR principles (Findable, Accessible, Interoperable, Reusable) (Wilkinson et al. 2016), which are increasingly important for research transparency and cross-platform integration in the digital humanities. After outlining the conceptual landscape of metadata and annotation in corpus linguistics, the article examines several existing projects that demonstrate the operationalisation of metadata standards across genres and temporal ranges. These include corpora that integrate CMDI profiles for complex resources (Paquot et al. 2024), and others that have transformed legacy metadata into machine-readable formats to facilitate data exchange and semantic enrichment (Fallucchi & De Luca 2020). Key issues addressed include orthographic variation, the limits of automated tagging tools for historical data (Pettersson & Megyesi 2018) and the need to balance standardisation with the preservation of linguistic idiosyncrasies. The article explores a case study of the MetaLing Corpus , a one-million-token historical corpus of English metalanguage from 1500 to 1700 (Andreani & Russo 2026). Practical constraints and editorial decisions are discussed. A minimalist encoding strategy was adopted following compatibility issues with TEI and Sketch Engine (Kilgarriff et al. 2004; Kilgarriff et al. 2014), resulting in a plain-text corpus linked to metadata managed through Omeka using Dublin Core (Caplan 2003). This study contributes to ongoing discussions on metadata sustainability, annotation design and best practices for corpus construction in underrepresented linguistic domains. It advocates for adaptable, transparent approaches that foster cross-disciplinary collaboration and position corpora as evolving infrastructures within broader digital ecosystems.
Yazarların özeti; kaynağından alınmıştır. English Language and Linguistics, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: Müzik
MusicArts and Humanities