PofoliaShared via Pofolia

npj Digital Medicine· 2026Q1

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Lan Jiang, Xiangji Ying, Andrew William Brown, Mengfei Lan et al.

Short summary

A new dataset (SPIRIT-CONSORT-ELM) with element-level annotations for 100 RCT protocol-results pairs and an automated pipeline using PubMedBERT and GPT-5 achieve high performance (F1: 0.822) in assessing RCT reporting completeness.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Introduced SPIRIT-CONSORT-ELM, a dataset with element-level annotations for 100 RCT protocol-results pairs.
  • Developed an automated pipeline using PubMedBERT and GPT-5 for assessing RCT reporting completeness.
  • The pipeline achieved high performance with an F1 score of 0.822 and Gwet’s AC1 of 0.796.
  • High inter-annotator agreement (Gwet’s AC1: 0.782) was observed during dataset annotation.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract Randomized controlled trials (RCTs) are central to assessing the benefits and harms of interventions, but incomplete reporting undermines their verifiability and usefulness. Although SPIRIT and CONSORT reporting guidelines promote complete reporting of RCT protocols and results publications, many RCTs remain incompletely reported. Automated manuscript checking could help improve reporting completeness before publication. We previously developed SPIRIT-CONSORT-TM, a corpus of 200 articles (100 protocol-results publication pairs) annotated with 83 checklist items from SPIRIT 2013 and CONSORT 2010, and trained models for item-level assessment. However, checklist items may comprise multiple constituent elements, which prior work did not capture or evaluate. Here, we extend the corpus with element-level annotations (SPIRIT-CONSORT-ELM) and formulate assessment as a machine reading comprehension task operationalized through 119 questions targeting specific reporting elements. Two annotators independently assessed 50 articles (25 pairs), with discrepancies resolved through discussion; one annotator assessed the remaining 150 articles. We then developed an automated pipeline combining PubMedBERT-based evidence retrieval with GPT-5-based question answering. Inter-annotator agreement was high (Gwet’s AC1: 0.782), and the pipeline achieved high performance (F1: 0.822, Gwet’s AC1: 0.796). Component analyses demonstrated the importance of evidence retrieval quality and modest benefits from illustrative in-context examples. SPIRIT-CONSORT-ELM provides a benchmark for fine-grained assessment of RCT reporting completeness, while the automated pipeline establishes a robust baseline and shows potential for supporting authors, reviewers, and editors.

The authors' abstract, as published at the source. npj Digital Medicine, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Statistics, Probability and UncertaintyDecision Sciences