PofoliaShared via Pofolia

Journal on Computing and Cultural Heritage· 2026Q1

Automatic Annotation and Analysis of Oral History using LLMs: An Empirical Study of Densho Digital Collection

Komala Subramanyam Cherukuri, Pranav Abishai Moses, Aisa Sakata, Jiangping Chen et al.

Short summary

A new LLM-based framework automatically annotates oral histories for semantic and sentiment, achieving high accuracy (e.g., ChatGPT 88.71% F1 for semantics) and enabling analysis of over 92,000 sentences from the Japanese American Incarceration Oral History collection.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Developed an LLM-based framework for automated semantic and sentiment annotation of oral histories.
  • Evaluated ChatGPT, Llama, and Qwen on a benchmark dataset, with ChatGPT achieving 88.71% F1 for semantic classification and Llama 82.87% for sentiment analysis.
  • Successfully annotated 92,191 sentences from 1,002 interviews in the Japanese American Incarceration Oral History collection.
  • Provides reusable code, annotated data, and prompt templates for applying LLMs to culturally sensitive archival research.

AI-generated from the title and abstract; the full text is not read.

Abstract

Oral histories serve as vital records of lived experience, particularly within communities affected by systemic injustice and historical erasure. Effective and efficient annotation and analysis of oral history archives can promote access and use of the oral histories. However, large scale analysis of these archives remains limited due to their unstructured format, emotional complexity, and the high cost of manual annotation. This paper presents a scalable LLM-based framework to automate semantic and sentiment annotation for oral history archives, with a focus on Japanese American Incarceration Oral History (JAIOH). Using large language models (LLMs), this study seeks to construct a high-quality dataset, systematically evaluate the performance of multiple LLMs, and investigate effective prompt engineering strategies on annotation in historically sensitive contexts. Our multiphase approach combines expert annotation, prompt design, and LLM evaluation using ChatGPT, Llama, and Qwen. We labeled 558 sentences from 15 narrators for sentiment and semantic classification, then developed prompts and evaluated across zero shot, few shot, and retrieval augmented generation (RAG) strategies based on the labeled data. On this benchmark, ChatGPT achieved the highest macro-F1 score for semantic classification (88.71%), followed by Llama (84.99%) and Qwen (83.72%). For sentiment analysis, Llama performed slightly better (82.87%) than Qwen (82.66%) and ChatGPT (82.29%), with all models showing comparable results. Based on the overall evaluation, we selected the best performing prompt configurations for each task and used them to automatically annotate 92,191 sentences from 1,002 interviews in the JAIOH collection. This study develops an LLM-based framework for automatic annotation and analysis of oral history collections. The evaluation results indicate that the proposed LLM-based annotation framework provides promising performance on semantic and sentiment annotation of the JAIOH collection. It contributes a reusable pipeline and practical guidance for applying LLMs in culturally sensitive archival analysis. All code, annotated data, prompt templates, and experimental details are available at the Repository . 1

The authors' abstract, as published at the source. Journal on Computing and Cultural Heritage, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: History

HistoryArts and Humanities