International Journal of Bilingualism· 2026Q1
Multilingual Hispanic Speech in California (MuHSiC): A Large-Scale Open-Access Spanish–English Bilingual Corpus
- 0citations
- Q1SCImago
- 2026year
Short summary
The MuHSiC corpus, the largest open-access Spanish-English bilingual speech dataset to date, offers over 700 hours of sociolinguistically stratified recordings with phoneme-level alignment, enabling reproducible research in bilingualism.
AI-generated from the title and abstract; the full text is not read.
Key points
- Introduces MuHSiC, the largest open-access Spanish-English bilingual speech corpus (over 700 hours).
- Corpus is sociolinguistically stratified with detailed speaker metadata.
- Includes 4.7 million time-aligned word tokens and over 16 million phonological segments.
- Designed for reproducible research in phonology, phonetics, morphosyntax, and language contact.
AI-generated from the title and abstract; the full text is not read.
Abstract
Aims and Objectives / Purpose / Research Questions: This article introduces the Multilingual Hispanic Speech in California (MuHSiC) corpus. It addresses two primary questions: (a) how can a large-scale, open-access corpus be designed to capture the linguistic diversity of Spanish–English bilingualism in California, and (b) how can such a resource support reproducible research across linguistic subfields? Design / Methodology / Approach: MuHSiC was constructed as a sociolinguistically stratified spoken corpus based on systematically collected bilingual interviews. The project integrates high-fidelity audio recording, standardized transcription protocols, phoneme-level forced alignment, acoustic extraction, and structured metadata collection to produce a research-ready, publicly distributable dataset. Data and Analysis: The corpus comprises 600 recordings from bilingual speakers representing diverse social and regional backgrounds across California. Each recording is approximately 70 min in length, evenly divided between Spanish and English, yielding over 42,720 min (712 hr) of speech. The dataset contains 4.7 million time-aligned word tokens and over 16 million phonological segments, accompanied by detailed speaker metadata. Findings / Conclusions: MuHSiC demonstrates that large-scale bilingual speech corpora can be constructed with both sociolinguistic breadth and phonetic detail. Its design ensures reliable alignment, replicable acoustic analysis, and compatibility with morphosyntactic and discourse-level research. Originality: To date, MuHSiC is the largest open-access, linguistically annotated corpus of Spanish–English bilingual speech. Its integration of balanced bilingual recordings, phoneme-level segmentation, and large-scale acoustic data distinguishes it from existing corpora. Significance / Implications: MuHSiC provides an unprecedented empirical foundation for research on bilingual phonology, phonetics, morphosyntax, and language contact. It establishes a methodological benchmark for multilingual corpus construction and supports open, collaborative research on linguistic diversity in California and beyond.
The authors' abstract, as published at the source. International Journal of Bilingualism, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Linguistics and Language
Linguistics and LanguageSocial Sciences