npj Digital Medicine· 2026Q1
LEME: open large language models for ophthalmology with advanced reasoning and clinical validation
- 1citations
- Q1SCImago
- 2026year
Short summary
LEME, a suite of open-weight LLMs for ophthalmology, outperforms seven baselines, including GPT-4o (by 3.32% in ROUGE-L), and achieves superior clinician ratings for answering patient queries and higher F1 scores for visual acuity extraction from clinical notes.
AI-generated from the title and abstract; the full text is not read.
Key points
- LEME, an open-weight LLM suite for ophthalmology, was developed using instruction tuning and reinforcement learning.
- LEME outperformed seven baselines, including GPT-4o (by 3.32% in ROUGE-L), on zero-shot question answering and patient-physician consultation benchmarks.
- Clinicians rated LEME highest for answering patient queries, with its completeness exceeding expert-written answers (p=0.015).
- LEME achieved the highest F1 scores for visual acuity extraction from clinical notes, outperforming LLaMA-3 70B by 14.1% and Eye-LLaMA by 59.0%.
AI-generated from the title and abstract; the full text is not read.
Abstract
Large language models (LLMs) offer potential to alleviate the growing public health burden of eye disease; however, few models have been tailored for ophthalmology or evaluated against clinically relevant benchmarks and real-world workflows. Here, we present LEME, a suite of open-weight LLMs for ophthalmology developed via a two-stage framework: instruction tuning on 211,149 examples curated from authoritative sources, and reinforcement learning with 29,747 preference-labeled examples to enhance clinical reasoning and informativeness. LEME was evaluated under zero-shot settings using five curated benchmarks, covering question answering and patient-physician consultation. LEME outperformed all seven baselines (all p < 0.004), notably exceeding GPT-4o by 3.32% in average ROUGE-L. Furthermore, LEME was assessed on three downstream tasks using deidentified patient data. In answering patient queries, LEME received the highest overall ratings from attending clinicians; its completeness rating even surpassed that of the expert-written answers ( p = 0.015). For visual acuity extraction from clinical notes, LEME achieved the highest F1 scores, outperforming LLaMA-3 70B by 14.1% and Eye-LLaMA by 59.0%. Finally, in assessment and treatment plan generation, LEME achieved ratings approaching those of attending clinicians, demonstrating promise for ophthalmic decision support. All models, datasets, and code are made open to support further development and validation.
The authors' abstract, as published at the source. npj Digital Medicine, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Health Informatics
Health InformaticsMedicine