Oral medicine is the dental specialty dedicated to the oral health care of medically complex patients and the diagnosis and management of medically related diseases, disorders, and conditions affecting the oral and maxillofacial region. Like other dental and medical specialties, oral medicine patient care is often impacted by challenges such as limited manpower, time, or resources. Artificial intelligence (AI) tools seek to supersede these challenges by automating human tasks and ushering greater efficiency and productivity. For direct patient care in oral medicine, AI has applications in risk prediction modeling, diagnosis establishment, treatment decision-making, and prognosis and outcomes prediction modeling.
Key points
-
•
Artificial intelligence applications in oral medicine related to direct patient care include risk prediction modeling, diagnosis establishment, treatment decision-making, and prognosis and outcomes prediction modeling.
-
•
Artificial intelligence has shown technical feasibility for use as an adjunctive tool for oral medicine specialists across many direct patient care domains.
-
•
Limitations and challenges for artificial intelligence use in oral medicine include lack of generalizable results, questionable explainability, and privacy and ethical concerns.
-
•
Further research evaluating performance, workflow considerations, and net benefit analysis is needed to establish justification for artificial intelligence use in oral medicine patient care.
Abbreviations
| AI | artificial intelligence |
| ANN | artificial neural networks |
| BRONJ | bisphosphonate-related osteonecrosis of the jaw |
| LLMs | large language models |
| MRONJ | medication-related osteonecrosis of the jaw |
| NEJM | New England Journal of Medicine |
| OLP | oral lichen planus |
| OPMDs | oral potentially malignant disorders |
| OSCC | oral cavity squamous cell carcinoma |
| TMD | temporomandibular disorder |
| TRIPOD | Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis |
Introduction
Oral medicine is one of the newest dental specialties recognized by the American Dental Association. The American Academy of Oral Medicine defines oral medicine as the specialty of dentistry responsible for the oral health care of medically complex patients and the diagnosis and management of medically related diseases, disorders, and conditions affecting the oral and maxillofacial region. Oral medicine encompasses oral mucosal diseases, chemosensory disorders, salivary gland disorders, oral manifestations of systemic diseases, and orofacial pain conditions.
While oral medicine specialists perform a multitude of management modalities, including surgical procedures, oral medicine is more diagnostics-driven than procedural compared to many other dental specialties. Patient care in oral medicine often requires the consolidation of diverse streams of data by the clinician—such as data elicited by history and performance of a physical examination, imaging, chairside diagnostic testing, or laboratory results—to stratify patients’ risks for disease, render diagnoses and treatment decisions, and predict outcomes of either disease or its therapy.
A multitude of biomedical and psychosocial factors are intertwined in the risk of developing different oral and maxillofacial diseases. The consolidation of individual patient risk factors and navigating their interrelationships in disease pathogenesis is a critical step in patient care, yet this can often be time-consuming for practitioners. The broad spectrum of oral and maxillofacial diseases and the diversity in the human expression of these diseases frequently give rise to diagnostic challenges in oral medicine, with the differential diagnosis of each clinician being informed by and often limited by his or her own experiential knowledge base. Patients with the same diagnosis will have unique prognoses and responses to treatment, requiring precise and personalized patient management plans. However, clinicians may be prone to imprecision, uncertainty, or limited standardization, which can impact treatment decision-making and outcomes prediction.
Despite their unique challenges, oral medicine disease risk prediction, diagnostics, treatment decision-making, and prognosis and outcomes prediction are united by the fact that they are heavily data-driven. Many of the shortcomings present in these fields of direct patient care are due to human fallibilities and limitations—such as knowledge gaps, finite manpower, and time constraints. In this context, oral medicine direct patient care stands out as a particularly ripe domain for artificial intelligence (AI) application. AI, which refers to machine execution of human cognitive functions and tasks, possesses distinct advantages over human intelligence. These include the ability to process larger volumes of data, greater speed in data analysis and decision-making, and indefatigability. The rationale behind the application of AI in oral medicine is the potential to facilitate and enhance the activities within the scope of oral medicine practice and to realize tangible gains for practitioners and patients alike. Fig. 1 demonstrates the various domains of possible AI application in oral medicine with respect to direct patient care.
Domains of artificial intelligence applications in oral medicine related to direct patient care.
Like human intelligence, AI is imperfect, and the limitations and implications of AI must be considered critically when navigating its potential in oral medicine. This article reviews the applications of AI toward direct patient care in oral medicine and discusses the current considerations of AI implementation in each domain. Readers interested in AI applications in oral medicine related to indirect patient care and academia are encouraged to refer to article, “Artificial Intelligence and Its Applications in Oral Medicine– Part 2 . ”
Disease risk prediction modeling
Disease risk prediction modeling refers to the process of identifying disease-predisposing factors and forecasting patients’ risks for disease development. Risk prediction modeling can facilitate the selection of appropriate patients for relevant screening, advanced testing, and monitoring programs, leading to early detection of critical diseases. By enabling personalized preventive strategies, it can also reduce the probability of disease development at the individual level and mitigate disease incidence at the population level.
Importantly, predictive modeling can be performed based on data obtained either with or without clinical examination of a patient. The former approach can render a preliminary risk stratification for the identification of patients for screening. The latter approach incorporates clinical examination findings. In the context of oral cancer risk prediction, the data incorporated into the pre-examination predictive model may be a composite of demographics, medical history, and risk factor history. The data incorporated into a postexamination predictive model may also account for lesion characteristics (eg, size, focality, color, texture, and symptoms).
AI is poised to revolutionize multivariable prediction modeling due to its unmatched ability to analyze extensive data points across large patient cohorts. Several AI algorithms have been adapted to predictive modeling of disease risk, including machine learning models such as logistic regressions, decision trees, K-nearest neighbors, support vector machines, and random forests and deep learning models such as artificial neural networks (ANN).
Disease risk prediction modeling studies for diseases in the scope of oral medicine have largely been retrospective in nature and have investigated the risk of medication-related osteonecrosis of the jaw (MRONJ) in patients receiving antiresorptive therapy, , the risk for oral cancer, , the risk for temporomandibular disorders, the risk for xerostomia in elderly patients, and the risk for oral mucositis. These studies have largely utilized commonly collected data, such as demographics, medical and risk factor history, ± clinical data. However, several studies have specifically evaluated risks pertaining to the inclusion of various biomarkers that are not routinely used in oral medicine practice. ,,
Kim and colleagues performed a retrospective study on 41 cases of bisphosphonate-related osteonecrosis of the jaw (BRONJ) after dental extraction and 84 controls to evaluate the risk of BRONJ development induced by dental extraction. The authors surveyed 9 demographic and clinical variables. Out of numerous algorithms tested, the random forest model was the superior performer with an area under the curve (AUC) of 0.973, sensitivity of 1, and specificity of 0.833. The authors established that drug holiday, serum C-terminal cross-linking telopeptide (CTX), and patient age had among the highest importance in predictive modeling.
Warin and colleagues performed a retrospective study on 5224 cases of MRONJ and 81 controls to evaluate the risk of MRONJ development. The authors surveyed extensive demographic and clinical variables as possible risk factors. The risk factors found to be significantly associated with an increased risk of MRONJ were multiple myeloma, hypertension, dyslipidemia, metabolic syndrome, thyroid disease, intravenous pamidronate, nonspecific antiresorptive agents, and concomitant intravenous corticosteroids. The authors tested several algorithms for both prediction of MRONJ development and time-to-event prediction for development within or beyond 24 months. Random forest was the superior model for time-specific risk of development prediction, with an area under the receiver operating characteristic curve (AUROC) of 0.90, sensitivity of 0.91, specificity of 0.75, and accuracy of 0.88.
Adeoye and colleagues performed a retrospective study on 1467 patients, 24 of whom had oral cavity squamous cell carcinoma (OSCC), with the intent of creating a decision support system for frontline workers to identify high-risk candidates for oral cancer screening. The authors evaluated several demographic and clinical variables and constructed a classifier based on 4 algorithms, namely K-nearest neighbors, random forest, AdaBoost, and ExtraTrees. The model had a sensitivity, specificity, and AUROC of 0.83, 0.86, and 0.85, respectively. The authors established that increasing age, lifetime alcohol consumption, lower educational attainment, reduced dental visitation, type of familial cancer, and smoking pack years were the most critical risk factors in their meta-classifier.
Alhazmi and colleagues performed a retrospective study on 51 patients with OSCC and 22 oral potentially malignant disorders (OPMDs) or benign OSCC mimics. The authors evaluated 29 clinical and lifestyle variables and constructed an ANN predictive model. The ANN demonstrated an accuracy of 0.789, sensitivity of 0.857, and specificity of 0.6. The authors did not comment on the relative importance of variables studied and acknowledged the need for further studies with larger cohorts.
Lee and colleagues performed a retrospective study on 4744 patients, 101 of whom had a self-reported temporomandibular disorder (TMD), and 68 of whom had a doctor-diagnosed TMD, to evaluate TMD risk factors. The authors surveyed extensive demographic, socioeconomic, mental/environmental, biological, and medical comorbidity variables and tested multiple predictive algorithms. They established ANN as the superior performer for self-reported TMD prediction, with accuracy of 0.9789 and AUC of 0.68, and logistic regression as the superior performer for doctor-diagnosed TMD, with accuracy of 0.984 and AUC of 0.68. Body mass index, monthly household income, age, daily sleep, and subjective obesity were the top 5 most relevant variables for both self-reported and doctor-diagnosed TMD.
Kuark-Fontes and colleagues performed a prospective clinical study evaluating the risk for oral mucositis with head and neck radiation, enrolling 157 patients with oral cavity or oropharyngeal SCC who had undergone curative protocols of radiotherapy alone or radiotherapy–chemotherapy. The authors investigated numerous demographic, clinical, and treatment-related data and employed supervised and unsupervised machine learning algorithms to develop a risk prediction model for oral mucositis. The authors found better results when restricting their algorithms to the variables found to have the highest association with oral mucositis, namely age, smoking status, head and neck surgery, radiotherapy dose, cancer treatment modality, histopathological differentiation, presence of oral cancer lesion, and tumor location. The k-nearest neighbors (KNN) algorithm outperformed the decision tree, random forest, and XGBoost models, achieving an accuracy of 65%, sensitivity of 58%, specificity of 68%, F1-score of 63%, and AUC of 0.63.
AI-based disease risk prediction models have shown promise for various diseases under the scope of oral medicine. These prediction models may be applicable for both oral medicine specialists and primary care providers and can be meaningful with data that almost every clinician can easily capture. In the primary care setting, AI-based risk prediction models may be utilized to facilitate the identification of at-risk patients by frontline workers for further screening, such as with the oral cancer risk prediction model described by Adeoye and colleagues. This would facilitate an arguably more effective use of time and resources in comparison to sole reliance on specialists. On the other hand, oral medicine specialists may benefit from more prediction models that can influence complex decision-making such as with extraction-related MRONJ risk as studied by Kim and colleagues.
Despite promising results, certain considerations must be recognized before determining suitability of AI-based disease risk predictive modeling for clinical practice. First, clinical risk prediction models must be tested with diverse datasets that represent the geographic breadth and clinical spectrum of the disease in question. Disease risk factors may vary across populations, and predictive models can only be plausible for populations that they have been validated for statistically. Moreover, prospective cohort trials are necessary to ensure that predictive models generated from retrospective datasets remain currently valid. External validation with new cohorts is also critical to establish generalizability of their results. Ultimately, the value of AI-based disease risk prediction modeling must undergo a net benefit analysis and be justifiable in light of actionable steps and improvements in patient outcomes and health care costs.
Diagnosis establishment
Diagnostics in oral medicine can range from a clinical diagnosis, which is one based on clinical examination, to a differential diagnosis, which represents a prioritized list of probable diagnoses, to a definitive diagnosis, which is one based on histopathological examination or other gold standard criteria. While histopathological examination has been the gold standard for diagnosis establishment for most oral mucosal diseases, the invasive nature of a biopsy has motivated researchers to investigate technology that can facilitate clinical diagnosis establishment with similar accuracy to that derived from histopathological examination. AI has been studied across a spectrum of oral mucosal diseases as well as other oral and maxillofacial diseases such as orofacial pain conditions. While robust AI research in oral and maxillofacial pathology and oral radiology has evaluated histology and medical imaging, respectively, AI models focusing on clinical parameters such as demographic data and clinical photographs remain the most applicable to oral medicine practice. AI diagnostics based on clinical photographs falls under computer vision, the AI discipline dedicated to image and video analysis. Chairside diagnosis of oral mucosal diseases with computer vision techniques can possibly reduce the need for other invasive, time-consuming, and/or costly diagnostic tests as well as allow the potential for oral medicine diagnostics to be subsumed, in part, by primary practitioners, resolving many common issues pertaining to resource limitations. Machine learning and deep learning, and more recently large language models (LLMs) and their larger scale multimodal counterparts, and foundational models have been studied for diagnostic purposes in oral medicine. Fig. 2 demonstrates an example framework for the application of either deep learning or LLMs to the diagnostic process.
Framework for the use of deep learning algorithms or large language models (eg, ChatGPT) for diagnosis establishment of proliferative verrucous leukoplakia from multimodal patient data.
Machine Learning and Deep Learning
The richest research landscape of AI diagnostics in oral medicine is that of OSCC and OPMDs. The high stakes associated with diagnosis of OSCC, the complexity of the clinical detection of early OSCCs, and their differentiation from OPMDs and benign look-alike lesions have rendered this field conducive to exploration with AI and computer vision.
Numerous systematic reviews have evaluated AI in OSCC and OPMD diagnosis from clinical photographs and other imaging modalities. ,,,,,,,,, Most of this literature has either explicitly aimed to assist primary care providers with the decision to refer or not refer to a specialist any lesions deemed suspicious for OSCC or OPMD or focused on broad tasks such as the differentiation of OSCC from OMPDs, benign lesions, or normal mucosa. No study to date has sought to render precision diagnoses across the full spectrum of diseases (OSCC/OMPDs vs benign look-alike diseases), which would have greater utility for oral medicine specialists.
Machine learning and deep learning have constituted the standard thus far in computer vision studies for OSCC and OPMD diagnosis. Several articles have comprehensively reviewed the clinical and engineering workflow concepts, considerations, and limitations in this domain. ,,, AI-based models for the diagnosis of OSCC and OPMD demonstrate promise with high statistical performance over an ever-increasing number of datasets and published studies. However, further growth depends critically on validation with larger and more representative datasets and improved explainability and uncertainty quantification of models.
AI research beyond OSCC and OPMDs has investigated the diagnosis of elementary oral lesions, anemia, oral lichen planus (OLP), , recurrent aphthous stomatitis, and oral ulcers of varying etiology. Several studies have evaluated AI discrimination among a broader range of oral lesions or disorders. ,,,,,,
AI-based diagnosis of systemic disease from oral images, given easy accessibility and lack of invasiveness, has been considered by Donmez and colleagues, who carried out a prospective study to detect anemia from lip vermilion photographs with CNN-based models. The authors tested several CNN models and demonstrated precision, recall, and F1 score values above 0.96. However, the sample size was small (with only 138 patients total and 29 patients with anemia), and no external validation was conducted.
The diagnosis of OLP, one of the most common oral mucosal diseases, has been studied by Acharit and colleagues among others. Acharit and colleagues performed a classification study on 609 OLP images and 480 non-OLP images (comprising diagnoses that would be included in the differential diagnosis of OLP such as hyperkeratosis, epithelial dysplasia, carcinoma in situ, recurrent aphthous stomatitis, traumatic ulcer, pemphigus vulgaris, mucus membrane pemphigoid, lupus erythematosus, and erythematous candidiasis). The authors tested several models and demonstrated accuracy of 0.88, sensitivity of 0.93, specificity of 0.84, and F1 score of 0.89 for their best-performing model.
Zhou and colleagues studied the photograph-based classification and object detection of recurrent aphthous ulcers using an image dataset consisting of 251 recurrent aphthous ulcers, 263 other oral mucosal diseases, and 271 of normal oral mucosa. Several models were tested for classification, with the best-performing one demonstrating precision of 0.93, sensitivity of 0.92, and specificity of 9.96, F1 score of 0.92, and AUC of 0.99. The best-performing object detection model demonstrated a precision of 0.99, recall of 0.79, F1 score of 0.88, and AUC of 0.91.
The differential diagnosis of ulcers frequently presents a challenge given multiple etiologies—traumatic, infectious, immune-mediated, or malignant. Jain and colleagues studied the photograph-based differentiation of oral ulcers of traumatic, infectious, and malignant etiology with 89 images. Their model demonstrated accuracies of 0.86 for binary classification of traumatic versus infectious, 0.88 for binary classification of infectious versus malignant, and 0.91 for binary classification of malignant versus traumatic, respectively.
Machine and deep learning have demonstrated significant growth in diagnostic potential in oral medicine, particularly for OSCC and OPMDs. The studies discussed in this review are restricted primarily to those involving clinical photographs, and specific limitations and considerations exist for this field. First, dataset factors such as the generally limited size and diversity of images, subsequently questionable generalizability of said datasets, and unknown impact of image quality and data augmentation need to be addressed. Second, no consensus remains regarding the threshold for clinical importance of performance metrics. Methodological aspects such as validation techniques, explainability techniques, and uncertainty quantification techniques are also among the most important elements necessitating refinement and elaboration, as otherwise clinical applicability remains problematic. Prospective clinical trials that elucidate AI performance in comparison to humans in addition to net benefit analysis are also critical to bridge the gap to clinical applicability.
Large Language Models
LLMs, which are now becoming multimodal and called multimodal large language models (MLLMs), have also been studied for the diagnosis of various oral lesions from textual case descriptions ,, as well as from clinical photographs alone or in conjunction with textual case descriptions. ,
Albagieh and colleagues evaluated the performance of ChatGPT 3.5 and 2 other LLMs (Stablediffusion and popAI) to that of third-year and fourth-year oral medicine and oral and maxillofacial pathology residents in answering 20 text-based oral medicine case questions in multiple-choice format. The authors found that ChatGPT 3.5 performance surpassed that of the residents as well as that of the 2 other LLMs. There was no statistically significant difference between the LLMs and the residents, however. The performance results of the 3 LLMs were not statistically compared to one another.
Tomo and colleagues evaluated the performance of ChatGPT 3.5 and 4.0 in providing differential diagnoses for 37 oral medicine cases in comparison to oral medicine and/or oral pathology specialists, general dentists, and final-year dental students. ChatGPT was only provided textual case descriptions, while human participants were also provided any applicable clinical and radiographic images. Differential diagnoses were rated on a 4 point Likert scale (correct, plausible, nonplausible, or absent). The oral medicine and/or oral pathology specialists surpassed ChatGPT and other human groups in providing a correct primary diagnosis, plausible primary diagnosis, correct or plausible alternative diagnosis, and correct diagnosis considering any hypothesis. However, ChatGPT performance was not significantly different from that of the specialists. The general dentists and final-year dental students, on the other hand, had significantly lower performance compared to ChatGPT.
Diniz-Freitas and colleagues evaluated the performance of ChatGPT 4V to that of humans in providing the diagnosis for oral diseases from 36 cases that included an image and brief clinical description, presented in a multiple-choice format. These images were derived from the New England Journal of Medicine (NEJM). NEJM’s regular human readers had achieved the correct diagnosis on these cases far less frequently than ChatGPT (49% vs 80.5%). The authors also compared ChatGPT’s performance based on the image alone, text alone, and the image and text together. They found that ChatGPT performed similarly when presented with text alone as when presented with both the image and text. It also performed significantly better in these 2 contexts than when presented with the image alone (80.5%, 80.5%, and 33.3%, respectively).
Yu and colleagues performed a recognition accuracy test of OLP on 3 LLMs, namely ChatGPT-4o, Chat-Diagrams, and Claude Opus using 128 OLP images. The untrained accuracies of the 3 LLMs were 59%, 68%, and 15%, respectively. Training the LLMs by providing a textual description of the clinical presentations and imaging characteristics of OLP at various oral sites led to statistically significant improvement in performance in all 3 groups, with accuracies of 77%, 80%, and 50%, respectively.
Vueghs and colleagues conducted a ChatGPT-4 diagnostics study for orofacial pain conditions. The authors anonymized 100 patient case descriptions and presented them to ChatGPT-4, prompting it to generate primary and differential diagnoses for each case using the International Classification of Orofacial Pain criteria. These cases included dentoalveolar pain, myofascial pain, temporomandibular joint pain, pain attributed to lesions or diseases of the cranial nerves, orofacial pain resembling primary headaches, and idiopathic orofacial pain. The authors also performed a comparative analysis with 2 orofacial pain experts, 2 final-year medical students, and 2 general dentists for 24 of the cases. ChatGPT-4 demonstrated an accuracy of 38% in determining the exact diagnosis for each case but provided accurate differential diagnoses for 80% of cases. ChatGPT-4 had a comparable performance to the medical students and general practitioners but performed below the level of clinician experts.
Although the published literature on LLMs in oral medicine diagnostics is still limited, LLMs appear to have demonstrated promise across a number of studies compared to nonexpert clinicians. Unlike other machine and deep learning models that have differed vastly in design, parameters, and training across studies, which has limited the wide adoption of any single model, the same major LLMs, chief among them ChatGPT, have dominated this new literature. This uniformity provides a distinct advantage when it comes to research growth, as technical advances made in well-funded, ubiquitous LLMs such as ChatGPT automatically become available to future researchers. Moreover, LLMs are capable of processing multimodal media, that is, images, videos, text, and speech, in larger quantities and with reduced need for supervision compared to machine and deep learning models. It is important to note, nonetheless, that LLMs have historically processed text more successfully than other media, and as noted earlier, image-based diagnostic potential is weak in comparison. Further improvements in image processing are duly expected with time, and future research can gauge this growth and better delineate the scope of multimodal LLM diagnostics. Well-documented limitations of LLMs in diagnostics must still be acknowledged, including data biases and questionable generalizability, hallucinations, lack of consistency, sensitivity to minor alterations in prompt language, limited contextual understanding, and inadequate explainability.
Stay updated, free dental videos. Join our Telegram channel
VIDEdental - Online dental courses