PLoS ONE· 2026Q1
Variability among large language models in assessing CONSORT compliance of published randomized clinical trials
- 0citations
- Q1SCImago
- 2026year
Short summary
ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 assessed CONSORT compliance in 20 randomized trials with significant differences: 80.8%, 64.4%, and 54.7% compliance, respectively.
AI-generated from the title and abstract; the full text is not read.
Key points
- ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 assessed CONSORT compliance in 20 randomized trials.
- Mean compliance scores varied significantly: ChatGPT-4o (80.8%), Gemini 2.5 Flash (64.4%), Claude Sonnet 4.6 (54.7%).
- Only ChatGPT-4o identified trials meeting a 90% compliance threshold (25% of papers).
- Inter-model agreement was only moderate, with no pairwise weighted kappa reaching substantial agreement (0.61).
AI-generated from the title and abstract; the full text is not read.
Abstract
Background Peer review processes may inadequately assess compliance with established reporting guidelines such as the Consolidated Standards of Reporting Trials (CONSORT) criteria. Large language models (LLMs) demonstrate potential for systematic manuscript evaluation; however, how consistently different models assess adherence to CONSORT guidelines in published clinical trials remains unexplored. Methods Twenty randomized controlled trials published in immunology journals between 2015 and 2016 were identified through PubMed. Three LLMs (ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6) independently assessed compliance across 37 CONSORT 2010 subpoints. The primary endpoint was the difference between models in mean CONSORT compliance score. Secondary endpoints included inter-model agreement and the proportion of articles meeting a 90% compliance threshold. Statistical analysis employed repeated measures analysis of variance (ANOVA) with post-hoc pairwise comparisons (α = 0.05). Results Mean CONSORT compliance rates were: ChatGPT-4o 80.8% (95% CI 75.8–85.8%), Gemini 2.5 Flash 64.4% (95% CI 58.0–70.8%), and Claude Sonnet 4.6 54.7% (95% CI 48.0–61.3%). Using a 90% compliance threshold as a quality benchmark, ChatGPT-4o identified 25% of papers (5/20) as meeting this standard, while Gemini 2.5 Flash and Claude Sonnet 4.6 each identified none (0/20) as meeting this standard. Repeated-measures ANOVA demonstrated significant differences between models (F(2,38) = 43.01, p < 0.001, partial η² = 0.694). All pairwise comparisons were statistically significant (ChatGPT-4o versus Gemini 2.5 Flash and ChatGPT-4o versus Claude Sonnet 4.6, both p < 0.001; Gemini 2.5 Flash versus Claude Sonnet 4.6, p = 0.014). Conclusions LLMs varied substantially in their assessment of CONSORT compliance in published randomized trials, with a consistent ordering: ChatGPT-4o scored compliance highest, Gemini 2.5 Flash intermediate, and Claude Sonnet 4.6 lowest. Item-level agreement between models was only moderate, with no pairwise weighted kappa reaching the 0.61 substantial-agreement threshold. This inter-model variability indicates the need for standardized evaluation protocols before LLM-assisted manuscript screening is adopted.
The authors' abstract, as published at the source. PLoS ONE, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Statistics, Probability and UncertaintyDecision Sciences