Options
Accuracy of Latest Large Language Models in Answering Multiple Choice Questions in Dentistry: A Comparative Study
Journal
PLOS ONE, Volume 20, Issue 1, e0317423
Date Issued
2025
DOI
https://doi.org/10.1371/journal.pone.0317423
Abstract
This study evaluates the performance of the latest large language models (LLMs) in answering dental multiple-choice questions, including both text-based and image-based formats. A total of 1,490 questions from United States National Board Dental Examination review books were used to assess models such as ChatGPT-4, Gemini, Copilot, Claude, Mistral, and Llama. Statistical analyses revealed significant differences in accuracy across models. Copilot, Claude, and ChatGPT achieved the highest performance, especially on text-based questions, while performance varied more on image-based questions. The findings indicate that multimodal capabilities improve performance and highlight the growing potential of LLMs in dental education and assessment.