PofoliaShared via Pofolia

Scientific Reports· 2026Q1

Comparison of expert-rated quality and readability of four large language models answering patient-oriented questions on Peyronie’s disease

Anıl Erdik, Hacı İbrahim Çimen, Kemal Demirhan, Deniz Gul

Short summary

Four LLMs (GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash, Grok) showed significant overall differences in response length and readability for Peyronie's disease questions, but no statistically significant differences in expert-rated quality due to limited inter-rater agreement.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Four LLMs (GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash, Grok) were evaluated on patient-oriented questions about Peyronie's disease.
  • Significant overall differences were observed in response word count and readability scores (FRES, FKGL, GFS, SMOG, Readable Rating) across the models.
  • Expert urologists did not find statistically significant differences in the quality of LLM responses.
  • Limited inter-rater agreement (ICC = 0.063) among reviewers restricted the interpretation of quality findings.

AI-generated from the title and abstract; the full text is not read.

Abstract

Large language models (LLMs) are increasingly used by patients seeking information about medical conditions, including Peyronie’s disease (PD). However, the quality and readability of artificial intelligence (AI)-generated responses to patient-oriented PD questions remain unclear. To compare the responses generated by four widely used AI models for common patient-oriented questions about PD in terms of expert-rated quality, response length, and readability. In this observational study, seven patient-oriented questions on PD were generated based on Google Trends outputs. Identical English-language prompts were submitted once to four publicly available AI systems: ChatGPT powered by GPT-5.5 Instant, Gemini 3.5 Flash, DeepSeek-V4-Flash (Instant mode), and Grok (Auto mode). The resulting responses were independently evaluated in randomized order by three board-certified urologists using a predefined 4-point expert-rating scale. Response length was assessed using word count, while readability was evaluated using the Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Score (GFS), Simple Measure of Gobbledygook (SMOG), and Readable Rating. Given the small number of paired observations, model comparisons were performed using exact Friedman tests; significant overall tests were followed by exact two-sided Wilcoxon signed-rank tests with Bonferroni correction. Reviewer-specific exact Friedman analyses did not detect statistically significant differences in expert-rated quality among the four AI models (all exact p ≥ 0.111). Inter-rater agreement was limited (single-measure ICC = 0.063, 95% CI, – 0.115 to 0.305), restricting the interpretation of comparative quality findings. Exact Friedman analyses demonstrated significant overall differences in word count, FRES, FKGL, GFS, SMOG, and Readable Rating (all exact p < 0.001). However, no pairwise comparison reached the Bonferroni-corrected significance threshold, and these analyses were severely constrained by the small number of paired observations; their non-significance should therefore not be interpreted as evidence that between-model differences were absent. Commonly used AI models showed significant overall differences in response length and readability when answering patient-oriented questions about PD, whereas reviewer-specific analyses did not detect statistically significant differences in expert-rated quality. However, the limited inter-rater agreement and the absence of external reference-standard verification preclude conclusions regarding equivalent or comparable model quality and warrant cautious interpretation of the expert-rated quality findings.

The authors' abstract, as published at the source. Scientific Reports, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: General Health Professions

General Health ProfessionsHealth Professions