PofoliaShared via Pofolia

Nutrients· 2026Q1

Nutritional Accuracy and Clinical Applicability of Large Language Model-Generated Therapeutic Diet Plans for Phenylketonuria: A Comparative Simulation Study

Bahar Kulu, Banu Süzen, Nuriye Ece Mintaş, Sibel Burçak Şahin Uyar et al.

Short summary

AI-generated PKU diet plans (ChatGPT-5.2, Claude 4.5, Gemini 3) were 75% likely to provide insufficient phenylalanine, failing to meet clinical prescriptions in 36 of 48 simulated cases.

AI-generated from the title and abstract; the full text is not read.

Key points

  • 75.0% (36/48) of AI-generated PKU diet plans provided less phenylalanine than prescribed.
  • Nutrient calculations from AI models showed substantial variability compared to independent BİAYS recalculations.
  • Dietitians rated plans generated with clinically specified structured prompts higher than those from zero-shot prompts.
  • Inter-rater reliability among dietitians evaluating the plans was modest.

AI-generated from the title and abstract; the full text is not read.

Abstract

Background: Large language models (LLMs) are increasingly used to generate nutrition-related recommendations, but their reliability for therapeutic diets requiring precise nutrient control remains uncertain. This study evaluated AI-generated one-day diet plans for standardized simulated cases of classical phenylketonuria (PKU). Methods: Eight standardized simulated cases representing infancy, adolescence, adulthood, and maternal PKU were evaluated. Between 8 and 10 February 2026, ChatGPT-5.2, Claude 4.5, and Gemini 3 were each used through their standard consumer-facing chat interfaces to generate one diet plan per case under two prompting conditions, yielding 48 plans. Nutrient composition was independently recalculated using BİAYS, and three blinded pediatric metabolic dietitians evaluated the plans using a structured 25-item checklist. Results: Across all 48 plans, phenylalanine provision calculated using BİAYS was below the case-specific prescription in 36 plans (75.0%). After Holm correction, none of the 12 paired AI-reported versus BİAYS-calculated nutrient comparisons remained statistically significant, while plan-level error metrics demonstrated substantial variability for several nutrients. Overall expert scores did not differ among models, whereas clinically specified structured prompts received higher expert ratings than limited-information zero-shot prompts. Inter-rater reliability was modest at the diet-plan level. Conclusions: AI-generated PKU diet plans may incorporate core dietary principles, but numerical plausibility does not ensure agreement with independent nutrient calculations or adherence to individualized clinical prescriptions. Independent nutrient verification and specialist metabolic-dietitian review remain necessary. Because only one output was generated for each case–model–prompt combination, the findings characterize the 48 evaluated outputs rather than within-model reproducibility.

The authors' abstract, as published at the source. Nutrients, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Clinical Biochemistry

Clinical BiochemistryBiochemistry, Genetics and Molecular Biology