PofoliaShared via Pofolia

PLoS ONE· 2026Q1

AI versus human coding of NIH grant abstracts

Sarah A. Alkhatib, Nadiya Alnoor Jiwa, Dallin Judd, Justin M. Luningham et al.

Short summary

ChatGPT-4.0 generated higher-quality innovation descriptions from NIH grant abstracts than human coders, averaging 4.47/5 for depth and relevance compared to humans' 3.33/5 and 3.24/5, respectively.

AI-generated from the title and abstract; the full text is not read.

Key points

  • ChatGPT-4.0 and human coders analyzed 125 NIH grant abstracts for research innovations.
  • Human evaluators rated ChatGPT's innovation descriptions higher than human-generated ones.
  • ChatGPT outputs averaged 4.47/5 for depth/detail and 4.47/5 for relevance/completeness.
  • Human outputs averaged 3.33/5 for depth/detail and 3.24/5 for relevance/completeness.

AI-generated from the title and abstract; the full text is not read.

Abstract

Large language models (LLMs) are increasingly used for qualitative analysis in substance use research, yet their performance relative to human coders remains underexplored. This study compares ChatGPT-4.0 with human coders in performing qualitative coding tasks using NIH grant abstracts as a test case, focusing on the identification and description of research innovations. Using a sample of NIH HEAL Initiative grant abstracts related to opioid overdose prevention, a total of 125 abstracts were independently coded by ChatGPT and humans to generate innovation descriptions, which were then evaluated by both human raters and ChatGPT for depth/detail and relevance/completeness using 5-point Likert scales. Identical instructions were used across all coding and evaluation stages. ChatGPT-generated descriptions were consistently rated higher than human-generated descriptions on both dimensions. Human evaluators rated ChatGPT outputs at an average of 4.47 for both depth/detail and relevance/completeness, compared to 3.33 and 3.24 for human outputs, respectively (F(1,176)=133.9, p < 0.001). These findings suggest that LLMs, when carefully prompted, can enhance the efficiency and quality of qualitative research evaluation.

The authors' abstract, as published at the source. PLoS ONE, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: General Social Sciences

General Social SciencesSocial Sciences