PofoliaPofolia ile paylaşıldı

npj Digital Medicine· 2026Q1

SPIRIT-CONSORT-ELM: RCT Raporlama Değerlendirmesi İçin Eleman Düzeyinde Açıklamalar ve Büyük Dil Modeli Yaklaşımı

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Lan Jiang, Xiangji Ying, Andrew William Brown, Mengfei Lan ve diğerleri

Kısa özet

100 RKÇ protokol-sonuç çifti için eleman düzeyinde açıklamalar içeren yeni bir veri kümesi (SPIRIT-CONSORT-ELM) ve PubMedBERT ile GPT-5 kullanan otomatik bir işlem hattı, RKÇ raporlama eksiksizliğini değerlendirmede yüksek performans (F1: 0.822) elde etti.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • 100 RKÇ protokol-sonuç çifti için eleman düzeyinde açıklamalar içeren SPIRIT-CONSORT-ELM veri kümesi tanıtıldı.
  • RKÇ raporlama eksiksizliğini değerlendirmek için PubMedBERT ve GPT-5 kullanan otomatik bir işlem hattı geliştirildi.
  • İşlem hattı, F1 skoru 0.822 ve Gwet'in AC1 skoru 0.796 ile yüksek performans gösterdi.
  • Veri kümesi açıklamasında yüksek düzeyde yorumlayıcılar arası uyum (Gwet'in AC1'i: 0.782) gözlemlendi.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Abstract Randomized controlled trials (RCTs) are central to assessing the benefits and harms of interventions, but incomplete reporting undermines their verifiability and usefulness. Although SPIRIT and CONSORT reporting guidelines promote complete reporting of RCT protocols and results publications, many RCTs remain incompletely reported. Automated manuscript checking could help improve reporting completeness before publication. We previously developed SPIRIT-CONSORT-TM, a corpus of 200 articles (100 protocol-results publication pairs) annotated with 83 checklist items from SPIRIT 2013 and CONSORT 2010, and trained models for item-level assessment. However, checklist items may comprise multiple constituent elements, which prior work did not capture or evaluate. Here, we extend the corpus with element-level annotations (SPIRIT-CONSORT-ELM) and formulate assessment as a machine reading comprehension task operationalized through 119 questions targeting specific reporting elements. Two annotators independently assessed 50 articles (25 pairs), with discrepancies resolved through discussion; one annotator assessed the remaining 150 articles. We then developed an automated pipeline combining PubMedBERT-based evidence retrieval with GPT-5-based question answering. Inter-annotator agreement was high (Gwet’s AC1: 0.782), and the pipeline achieved high performance (F1: 0.822, Gwet’s AC1: 0.796). Component analyses demonstrated the importance of evidence retrieval quality and modest benefits from illustrative in-context examples. SPIRIT-CONSORT-ELM provides a benchmark for fine-grained assessment of RCT reporting completeness, while the automated pipeline establishes a robust baseline and shows potential for supporting authors, reviewers, and editors.

Yazarların özeti; kaynağından alınmıştır. npj Digital Medicine, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Statistics, Probability and UncertaintyDecision Sciences