PofoliaPofolia ile paylaşıldı

PLoS ONE· 2026Q1

Büyük Dil Modelleri, Yayımlanmış Rastgele Klinik Denemelerin CONSORT Uygunluğunu Değerlendirmede Geniş Değişkenlik Gösterdi

Variability among large language models in assessing CONSORT compliance of published randomized clinical trials

Daniel Y. Tsybulnik, Justin J. Gillette, Thomas F Heston

Kısa özet

ChatGPT-4o, Gemini 2.5 Flash ve Claude Sonnet 4.6, 20 rastgele denemede CONSORT uygunluğunu sırasıyla %80,8, %64,4 ve %54,7 oranlarında önemli farklılıklarla değerlendirdi.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • ChatGPT-4o, Gemini 2.5 Flash ve Claude Sonnet 4.6, 20 rastgele denemede CONSORT uygunluğunu değerlendirdi.
  • Ortalama uyumluluk puanları önemli ölçüde farklılık gösterdi: ChatGPT-4o (%80,8), Gemini 2.5 Flash (%64,4), Claude Sonnet 4.6 (%54,7).
  • Yalnızca ChatGPT-4o, %90 uyumluluk eşiğini karşılayan denemeleri belirledi (makalelerin %25'i).
  • Modeller arası uyum yalnızca orta düzeydeydi; hiçbir ikili ağırlıklı kappa değeri önemli uyum eşiği olan 0,61'e ulaşmadı.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Background Peer review processes may inadequately assess compliance with established reporting guidelines such as the Consolidated Standards of Reporting Trials (CONSORT) criteria. Large language models (LLMs) demonstrate potential for systematic manuscript evaluation; however, how consistently different models assess adherence to CONSORT guidelines in published clinical trials remains unexplored. Methods Twenty randomized controlled trials published in immunology journals between 2015 and 2016 were identified through PubMed. Three LLMs (ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6) independently assessed compliance across 37 CONSORT 2010 subpoints. The primary endpoint was the difference between models in mean CONSORT compliance score. Secondary endpoints included inter-model agreement and the proportion of articles meeting a 90% compliance threshold. Statistical analysis employed repeated measures analysis of variance (ANOVA) with post-hoc pairwise comparisons (α = 0.05). Results Mean CONSORT compliance rates were: ChatGPT-4o 80.8% (95% CI 75.8–85.8%), Gemini 2.5 Flash 64.4% (95% CI 58.0–70.8%), and Claude Sonnet 4.6 54.7% (95% CI 48.0–61.3%). Using a 90% compliance threshold as a quality benchmark, ChatGPT-4o identified 25% of papers (5/20) as meeting this standard, while Gemini 2.5 Flash and Claude Sonnet 4.6 each identified none (0/20) as meeting this standard. Repeated-measures ANOVA demonstrated significant differences between models (F(2,38) = 43.01, p < 0.001, partial η² = 0.694). All pairwise comparisons were statistically significant (ChatGPT-4o versus Gemini 2.5 Flash and ChatGPT-4o versus Claude Sonnet 4.6, both p < 0.001; Gemini 2.5 Flash versus Claude Sonnet 4.6, p = 0.014). Conclusions LLMs varied substantially in their assessment of CONSORT compliance in published randomized trials, with a consistent ordering: ChatGPT-4o scored compliance highest, Gemini 2.5 Flash intermediate, and Claude Sonnet 4.6 lowest. Item-level agreement between models was only moderate, with no pairwise weighted kappa reaching the 0.61 substantial-agreement threshold. This inter-model variability indicates the need for standardized evaluation protocols before LLM-assisted manuscript screening is adopted.

Yazarların özeti; kaynağından alınmıştır. PLoS ONE, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Statistics, Probability and UncertaintyDecision Sciences