Scientific Reports· 2026Q1
Peyronie hastalığı sorularını yanıtlayan büyük dil modellerinin uzman derecelendirmeli kalitesi ve okunabilirliği karşılaştırıldı.
Comparison of expert-rated quality and readability of four large language models answering patient-oriented questions on Peyronie’s disease
- 0atıf
- Q1SCImago
- 2026yıl
Kısa özet
Dört BDM (GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash, Grok), Peyronie hastalığı sorularına yanıtlarında yanıt uzunluğu ve okunabilirlik açısından anlamlı genel farklılıklar gösterdi, ancak sınırlı değerlendiriciler arası uyum nedeniyle kalitede istatistiksel olarak anlamlı farklılıklar tespit edilemedi.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Peyronie hastalığı hakkında hasta odaklı sorular için dört BDM (GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash, Grok) değerlendirildi.
- Modeller arasında yanıt kelime sayısı ve okunabilirlik skorlarında (FRES, FKGL, GFS, SMOG, Okunabilirlik Derecelendirmesi) anlamlı genel farklılıklar gözlemlendi.
- Uzman ürologlar, BDM yanıtlarının kalitesinde istatistiksel olarak anlamlı farklılıklar bulamadı.
- Gözlemciler arasındaki sınırlı değerlendirici uyumu (ICC = 0.063), kalite bulgularının yorumlanmasını kısıtladı.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Large language models (LLMs) are increasingly used by patients seeking information about medical conditions, including Peyronie’s disease (PD). However, the quality and readability of artificial intelligence (AI)-generated responses to patient-oriented PD questions remain unclear. To compare the responses generated by four widely used AI models for common patient-oriented questions about PD in terms of expert-rated quality, response length, and readability. In this observational study, seven patient-oriented questions on PD were generated based on Google Trends outputs. Identical English-language prompts were submitted once to four publicly available AI systems: ChatGPT powered by GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash (Instant mode), and Grok (Auto mode). The resulting responses were independently evaluated in randomized order by three board-certified urologists using a predefined 4-point expert-rating scale. Response length was assessed using word count, while readability was evaluated using the Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Score (GFS), Simple Measure of Gobbledygook (SMOG), and Readable Rating. Given the small number of paired observations, model comparisons were performed using exact Friedman tests; significant overall tests were followed by exact two-sided Wilcoxon signed-rank tests with Bonferroni correction. Reviewer-specific exact Friedman analyses did not detect statistically significant differences in expert-rated quality among the four AI models (all exact p ≥ 0.111). Inter-rater agreement was limited (single-measure ICC = 0.063, 95% CI, – 0.115 to 0.305), restricting the interpretation of comparative quality findings. Exact Friedman analyses demonstrated significant overall differences in word count, FRES, FKGL, GFS, SMOG, and Readable Rating (all exact p < 0.001). However, no pairwise comparison reached the Bonferroni-corrected significance threshold, and these analyses were severely constrained by the small number of paired observations; their non-significance should therefore not be interpreted as evidence that between-model differences were absent. Commonly used AI models showed significant overall differences in response length and readability when answering patient-oriented questions about PD, whereas reviewer-specific analyses did not detect statistically significant differences in expert-rated quality. However, the limited inter-rater agreement and the absence of external reference-standard verification preclude conclusions regarding equivalent or comparable model quality and warrant cautious interpretation of the expert-rated quality findings.
Yazarların özeti; kaynağından alınmıştır. Scientific Reports, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: Genel Sağlık Meslekleri
General Health ProfessionsHealth Professions