PofoliaPofolia ile paylaşıldı

BMC Oral Health· 2026Q1

Büyük dil modellerinin restoratif ve endodontik diş hekimliği çoktan seçmeli sorularındaki karşılaştırmalı doğruluğu: diş hekimliği eğitimi için çıkarımlar

Comparative accuracy of large language models in restorative and endodontic dentistry multiple-choice questions: implications for dental education

İkbal Esra Pehlivan, Sema Kaya, ELİF BAŞTUĞ GÜVEN, Beyza Batmaz

Kısa özet

On dört büyük dil modeli (BDM), Türkçe diş hekimliği çoktan seçmeli sorularında değişken doğruluk gösterdi; Gemini Pro ve Gemini Pro (Deep Research) geçerli sorularda (Tip 1) öne çıkarken, Manus yapısal olarak geçersiz maddeleri belirlemede %98 doğruluk sergiledi.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • On dört BDM, restoratif diş hekimliği ve endodontideki 200 Türkçe çoktan seçmeli soru üzerinde test edildi.
  • Gemini Pro ve Gemini Pro (Deep Research), geçerli sorularda (Tip 1) en yüksek doğruluğu elde etti.
  • Manus, yapısal olarak geçersiz soruları (Tip 3 ve 4) belirlemede %98 doğruluk gösterdi.
  • On dört modelden yedisi, bazıları endodontiyi, bazıları ise restoratif diş hekimliğini destekleyen anlamlı disipline dayalı performans farklılıkları gösterdi.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Large language models are increasingly used as learning and question-answering tools in health professions education. However, their reliability in dental multiple-choice assessments may depend not only on overall accuracy but also on disciplinary context and the structural validity of the questions. This study aimed to compare the performance of 14 large language model-based artificial intelligence systems on multiple-choice questions in restorative dentistry and endodontics and to evaluate their ability to identify structurally invalid items. This cross-sectional experimental study evaluated the performance of 14 large language model (LLM)-based artificial intelligence systems on 200 five-option multiple-choice questions prepared in Turkish. The questions were equally distributed between Restorative Dentistry and Endodontics and categorized into four structural types: Type 1 (valid question, single correct answer), Type 2 (semantically invalid question), Type 3 (valid question, no correct option), and Type 4 (valid question, no incorrect option). The model responses were coded as correct (1) or incorrect (0). Cochran’s Q and McNemar tests were used to compare models, and Wilcoxon signed-rank tests evaluated discipline-based differences ( p < 0.05). Significant inter-model differences were observed for types 1, 2, and 4 ( p < 0.001), but not for type 3 ( p = 0.448). Gemini Pro and Gemini Pro (Deep Research) achieved the highest accuracy in Type 1, while DeepSeek and ChatGPT 5.2 led in Type 2. Manus demonstrated outstanding performance on structurally contradictory questions (Types 3 and 4; 98%). Overall, seven of the fourteen models showed significant discipline-based differences. Endodontics was favored in five models: Gemini Pro, Gemini Pro (Deep Research), ChatGPT 5.2 Web, Perplexity, and SciSpace, whereas DeepSeek and DeepSeek (Deep Search) favored restorative dentistry. When analyzed by question type, differences favoring endodontics were clustered in Type 2, whereas differences favoring Restorative Dentistry predominated in Types 3 and 4. The performance of LLMs in dental education varies across disciplines, question structures, and model systems. Accuracy alone is insufficient for evaluating reliability. The ability to detect structurally invalid items varies markedly across models, with important implications for assessment validity, examination security, and the responsible integration of artificial intelligence into dental education.

Yazarların özeti; kaynağından alınmıştır. BMC Oral Health, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Alan: Genel Diş Hekimliği

General DentistryDentistry