Data· 2026Q2
Comparing LLM and Human Expert Coding in Qualitative Research: A Case Study on Emotional Intelligence
- 0citations
- Q2SCImago
- 2026year
Short summary
LLMs achieved 75.0%-91.7% direct coverage and 83.3%-100.0% broad coverage in recovering a 12-dimension emotional intelligence framework from 34 interviews, with theme-count precision ranging from 50.0% to 69.2% across three models.
AI-generated from the title and abstract; the full text is not read.
Key points
- LLMs achieved 75.0%-91.7% direct coverage and 83.3%-100.0% broad coverage in recovering a 12-dimension EI framework.
- Theme-count precision for LLMs ranged from 50.0% (DeepSeek) to 69.2% (Qwen).
- Eight EI dimensions were directly recovered across all LLM outputs.
- LLM quotations were manually checked against transcripts, with no fabricated content found.
- An independent expert reproduced all 36 human-expert mappings with 100% agreement.
AI-generated from the title and abstract; the full text is not read.
Abstract
Large language models (LLMs) are increasingly used for qualitative coding, but their ability to recover human-defined construct boundaries remains unclear. Existing studies rarely distinguish thematic granularity from dimension-level semantic alignment. We compare a 12-dimension emotional-intelligence (EI) framework developed from 34 Chinese behavioral-event interviews with three zero-shot LLM analyses of the same corpus under an identical prompt. A structured framework-level mapping procedure classified each human dimension as directly recovered, partially or merged, or not recovered. Alignment was summarized using direct, broad, and frequency-weighted coverage, theme-count precision, and pairwise Jaccard similarities for the three mapping types. Two senior experts evaluated 36 dimension-level relationships, and retained LLM quotations were manually checked against the original transcripts; no fabricated quotations were identified. Direct coverage ranged from 75.0% to 91.7%, broad coverage from 83.3% to 100.0%, and frequency-weighted mapping from 80.3% to 94.9%. Theme-count precision was 50.0% for DeepSeek, 69.2% for Qwen, and 55.6% for GPT. Eight dimensions were directly recovered across all outputs, whereas guidance and motivation, big-picture awareness, care and support orientation, and teamwork varied across outputs. An independent expert reproduced all 36 mappings, yielding 100% raw agreement. The findings distinguish thematic granularity from construct-boundary recovery and support LLMs as complementary analytical assistants.
The authors' abstract, as published at the source. Data, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: General Social Sciences
General Social SciencesSocial Sciences