PofoliaShared via Pofolia

BMC Oral Health· 2026Q1

Comparative accuracy of large language models in restorative and endodontic dentistry multiple-choice questions: implications for dental education

İkbal Esra Pehlivan, Sema Kaya, ELİF BAŞTUĞ GÜVEN, Beyza Batmaz

Short summary

Fourteen large language models (LLMs) showed varying accuracy on Turkish dental multiple-choice questions, with Gemini Pro and Gemini Pro (Deep Research) excelling in valid questions (Type 1) and Manus demonstrating 98% accuracy in identifying structurally invalid items (Types 3 & 4).

AI-generated from the title and abstract; the full text is not read.

Key points

  • Fourteen LLMs were tested on 200 Turkish multiple-choice questions in restorative dentistry and endodontics.
  • Gemini Pro and Gemini Pro (Deep Research) achieved the highest accuracy on valid questions (Type 1).
  • Manus demonstrated 98% accuracy in identifying structurally invalid questions (Types 3 and 4).
  • Seven of the fourteen models showed significant discipline-based performance differences, with some favoring endodontics and others restorative dentistry.

AI-generated from the title and abstract; the full text is not read.

Abstract

Large language models are increasingly used as learning and question-answering tools in health professions education. However, their reliability in dental multiple-choice assessments may depend not only on overall accuracy but also on disciplinary context and the structural validity of the questions. This study aimed to compare the performance of 14 large language model-based artificial intelligence systems on multiple-choice questions in restorative dentistry and endodontics and to evaluate their ability to identify structurally invalid items. This cross-sectional experimental study evaluated the performance of 14 large language model (LLM)-based artificial intelligence systems on 200 five-option multiple-choice questions prepared in Turkish. The questions were equally distributed between Restorative Dentistry and Endodontics and categorized into four structural types: Type 1 (valid question, single correct answer), Type 2 (semantically invalid question), Type 3 (valid question, no correct option), and Type 4 (valid question, no incorrect option). The model responses were coded as correct (1) or incorrect (0). Cochran’s Q and McNemar tests were used to compare models, and Wilcoxon signed-rank tests evaluated discipline-based differences ( p < 0.05). Significant inter-model differences were observed for types 1, 2, and 4 ( p < 0.001), but not for type 3 ( p = 0.448). Gemini Pro and Gemini Pro (Deep Research) achieved the highest accuracy in Type 1, while DeepSeek and ChatGPT 5.2 led in Type 2. Manus demonstrated outstanding performance on structurally contradictory questions (Types 3 and 4; 98%). Overall, seven of the fourteen models showed significant discipline-based differences. Endodontics was favored in five models: Gemini Pro, Gemini Pro (Deep Research), ChatGPT 5.2 Web, Perplexity, and SciSpace, whereas DeepSeek and DeepSeek (Deep Search) favored restorative dentistry. When analyzed by question type, differences favoring endodontics were clustered in Type 2, whereas differences favoring Restorative Dentistry predominated in Types 3 and 4. The performance of LLMs in dental education varies across disciplines, question structures, and model systems. Accuracy alone is insufficient for evaluating reliability. The ability to detect structurally invalid items varies markedly across models, with important implications for assessment validity, examination security, and the responsible integration of artificial intelligence into dental education.

The authors' abstract, as published at the source. BMC Oral Health, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: General Dentistry

General DentistryDentistry