Skip Navigation
Skip to contents

Clin Endosc : Clinical Endoscopy

OPEN ACCESS

Articles

Page Path
HOME > Clin Endosc > Ahead-of print articles > Article
Review Translating artificial intelligence into clinical practice for gastrointestinal endoscopy: current applications and future perspectives
Hyeong Ho Jo1,*orcid, Jin Ho Choi2,*orcid, Joo Seong Kim3orcid, Seung-Joo Nam4orcid, Do Hoon Kim5orcid, Woo Hyun Paik2orcid, Jung Ho Bae6orcid, Seung Wook Hong5orcid, Chang Seok Bang7orcid, Da Hyun Jung8orcid, Seong Ji Choi9orcid, Hyunsoo Chung10orcid, The Research Group for Artificial Intelligence, Korean Society for Gastrointestinal Endoscopy

DOI: https://doi.org/10.5946/ce.2025.419
Published online: July 30, 2026

1Department of Internal Medicine, Daegu Catholic University School of Medicine, Daegu, Korea

2Department of Internal Medicine, Seoul National University Hospital, Seoul, Korea

3Department of Internal Medicine, Seoul Metropolitan Government–Seoul National University Boramae Medical Center, Seoul, Korea

4Department of Internal Medicine, Kangwon National University School of Medicine, Chuncheon, Korea

5Department of Gastroenterology, Asan Medical Center, University of Ulsan College of Medicine, Seoul, Korea

6Department of Internal Medicine and Healthcare Research Institute, Healthcare System Gangnam Center, Seoul National University Hospital, Seoul, Korea

7Department of Internal Medicine, Hallym University Chuncheon Sacred Heart Hospital, Hallym University College of Medicine, Chuncheon, Korea

8Department of Internal Medicine, Yonsei University College of Medicine, Seoul, Korea

9Department of Internal Medicine, Korea University Guro Hospital, Korea University College of Medicine, Seoul, Korea

10Department of Internal Medicine, Seoul National University College of Medicine, Seoul, Korea

Correspondence: Hyunsoo Chung Department of Internal Medicine, Seoul National University Hospital, Seoul National University College of Medicine, 101 Daehak-ro, Jongno-gu, Seoul 03080, Korea E-mail: h.chung@snu.ac.kr
*Hyeong Ho Jo and Jin Ho Choi contributed equally to this work.
• Received: November 13, 2025   • Revised: February 12, 2026   • Accepted: February 28, 2026

© 2026 Korean Society of Gastrointestinal Endoscopy

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 1,100 Views
  • 79 Download
  • Artificial intelligence (AI) has emerged as a transformative tool in gastrointestinal (GI) endoscopy, addressing challenges in detection, diagnosis, and decision-making. In upper GI endoscopy, AI supports blind spot monitoring, Helicobacter pylori diagnosis, and the identification of premalignant and malignant lesions, with high accuracy and reduced miss rates. In lower GI endoscopy, computer-aided detection improves adenoma detection, whereas computer-aided diagnosis supports “resect-and-discard” and “diagnose-and-leave” strategies. However, real-world benefits remain modest, with concerns regarding overdetection and variable performance across lesion types and colon segments. In inflammatory bowel disease, AI standardizes endoscopic and histologic scoring, reduces interobserver variability, and accelerates capsule endoscopy interpretation, including high diagnostic accuracy for Crohn’s disease. Pancreatobiliary applications, including endoscopic ultrasound, endoscopic retrograde cholangiopancreatography, and cholangioscopy, demonstrate strong performance in differentiating pancreatic masses and biliary strictures and in predicting postprocedural complications. Despite expert-level performance across multiple domains, most studies remain single-center or retrospective, and explainability, workflow integration, medicolegal responsibility, and cost-effectiveness continue to limit adoption. Emerging solutions, including explainable AI and AI-generated common data model-compatible reports, may bridge these gaps. With rigorous multicenter validation and real-world implementation, AI can evolve from an experimental adjunct into a core component of routine endoscopic practice.
Artificial intelligence (AI) is rapidly reshaping gastrointestinal (GI) endoscopy techniques by improving detection, characterization, and clinical decision-making. In upper GI endoscopy, AI has been used for blind spot detection, Helicobacter pylori (H. pylori) diagnosis, and early detection of gastric and esophageal neoplasia, often achieving accuracies above 90%. Similarly, in colonoscopy, computer-aided detection (CADe) systems increase adenoma detection rates (ADR), whereas computer-aided diagnosis (CADx) supports real-time histological prediction. Although trial data are encouraging, real-world benefits remain modest, with concerns regarding overdetection and variable performance. In inflammatory bowel disease (IBD), AI reduces interobserver variability in endoscopic scoring, enhances histological assessment, and accelerates capsule endoscopy interpretation. Pancreatobiliary (PB) applications, including endoscopic ultrasound (EUS), endoscopic retrograde cholangiopancreatography (ERCP), and cholangioscopy, have shown strong diagnostic performance, particularly for pancreatic masses and biliary strictures. Advances in explainable AI (XAI) and standardized report generation have further supported integration into clinical workflows. Despite these advances, validation, workflow adaptation, interpretability, and medicolegal challenges remain critical for translation into routine practice. In this review, we discuss these applications in greater detail and consider future directions for AI applications in GI endoscopy.
AI in upper GI endoscopy
AI has been increasingly applied to various aspects of upper GI endoscopy. Except for studies on Barrett’s esophagus, most studies in this field have been conducted in Asian countries, particularly China and Japan. Although the field is still evolving, AI-assisted upper GI endoscopy can be broadly classified into four main domains: (1) blind spot detection and quality improvement; (2) diagnosis of H. pylori infection and gastritis; (3) diagnosis of gastric premalignant and malignant lesions; and (4) diagnosis of esophageal premalignant and malignant lesions. AI has also been explored for other applications, including grading of gastroesophageal reflux disease, subepithelial tumor characterization, esophageal varix detection, automated standardization of reports using large language models (LLMs), and training support for novice endoscopists.

1) Blind spot detection and quality improvement

Studies on quality indicators in upper GI endoscopy remain relatively limited compared with those in colonoscopy. Most investigations have focused on image-based analyses, particularly evaluating the completeness of photo documentation, which lends itself to deep-learning applications. A research group led by Prof. Hong-Gang Yu at Wuhan University developed the WISENSE algorithm for blind spot monitoring, demonstrating a significant reduction in blind spot rates from 22.46% to 5.86% among endoscopists with 1–3 years of experience.1 The same group subsequently refined this approach, introducing ENDOANGEL-LD, which in a randomized controlled tandem trial significantly reduced the miss rate of gastric neoplasms from 25.6% to 6.4%.2 They later developed AI-EARS, an automated system capable of capturing endoscopic images, identifying lesion location and characteristics, and generating endoscopy reports without manual image acquisition or documentation by the endoscopist.3 Most other studies have focused on algorithmic accuracy using still images or recorded videos rather than real-time clinical utility. However, recent investigations in Hong Kong and Korea have reported encouraging results regarding the feasibility and effectiveness of AI-based blind spot detection systems in real or simulated endoscopic settings.4,5 Such systems are expected to enhance procedural completeness and may serve as objective quality indicators in daily clinical practice.

2) Diagnosis of H. pylori infection and gastritis

An increasing number of studies have explored AI models for the diagnosis of H. pylori infection using upper GI endoscopic images. Most systems employ convolutional neural networks (CNNs), whereas earlier models adopted traditional machine learning methods, such as support vector machines. A recent meta-analysis reported pooled sensitivity and specificity of approximately 90%–95%.6 Given this high diagnostic performance, AI-based image interpretation may reduce the need for unnecessary biopsies and lower associated costs. However, as most studies have been conducted in Asian populations, mainly in Japan, China, and Korea, where H. pylori prevalence is high, further validation in non-Asian populations with diverse demographic and microbiological characteristics is required.
In addition to H. pylori detection, AI has been applied to automated histopathologic grading of gastritis, atrophy, and intestinal metaplasia according to the Updated Sydney system and to enable risk stratification using the Operative Link on Gastritis Assessment (OLGA) and the Operative Link on Gastric Intestinal Metaplasia Assessment (OLGIM) systems.7
Multiple CNN-based models have achieved diagnostic accuracies comparable to or exceeding those of expert endoscopists, with pooled accuracies approaching 95%.8 Despite these promising results, most existing studies have been retrospective and single-center in design, which limits generalizability and impedes large-scale clinical adoption.

3) Diagnosis of gastric premalignant and malignant lesions

The detection and characterization of neoplastic lesions, particularly gastric cancer, represent one of the most active areas of AI research in upper GI endoscopy. Numerous studies have shown that AI systems can achieve sensitivity, specificity, and overall accuracy exceeding 90% using both white-light endoscopy (WLE) and narrow-band imaging (NBI), often matching or even surpassing the performance of expert endoscopists.9-11 Importantly, AI implementation does not uniformly require image-enhanced platforms. Several CNN models trained exclusively on WLE images have demonstrated sensitivities above 90% and diagnostic accuracies comparable to expert endoscopists.12 In one randomized controlled trial (RCT), AI-assisted standard WLE significantly reduced the miss rate of gastric neoplasms, underscoring that clinically meaningful benefits can be achieved within routine WLE workflows.2 However, RCTs remain limited, with only two completed studies and one ongoing study conducted in China.2,13,14 Accordingly, further multicenter prospective validation is required to establish the clinical utility of AI-assisted CADe and CADx systems. The recent position statement of the Japanese Gastroenterological Endoscopy Society has also emphasized the necessity of robust clinical evidence before routine implementation.15 Future work should focus on real-time AI integration during endoscopic resection and on developing multimodal frameworks that combine endoscopic, radiologic, and histopathologic data to support comprehensive clinical decision-making.

4) Diagnosis of premalignant and malignant esophageal lesions

AI has also demonstrated significant potential for the detection and characterization of esophageal malignancies. Owing to regional epidemiologic differences, studies on esophageal squamous cell carcinoma (ESCC) have primarily been conducted in Asia, whereas investigations of esophageal adenocarcinoma (EAC) have predominated in Western countries. Early-stage esophageal cancers often present with subtle mucosal changes, making AI particularly valuable for early detection. Recent meta-analyses have reported pooled sensitivity and specificity of 91.2% and 80.0% for ESCC and 93.1% and 86.9% for EAC, respectively.16 However, RCTs are scarce, with only a single study reported to date, conducted in China.17 Future research should aim to validate these findings across diverse populations and imaging platforms, ensuring reproducibility and global applicability. Representative studies highlighting these developments across the four key domains of upper GI endoscopy are summarized in Table 1.1-9,13,16,17

5) Current limitations and translational challenges

Although many studies have demonstrated that the performance of AI is comparable to that of expert endoscopists, several critical challenges must be addressed before AI can be seamlessly integrated into clinical practice. First, the dynamics of human–AI interaction remain inadequately understood, particularly regarding whether AI can function reliably in real time without disrupting the workflow and whether clinicians can intuitively interpret its outputs. Second, legal and ethical considerations, including responsibility for diagnostic errors and data privacy, are essential. Third, variations in disease prevalence and institutional practices (e.g., photo documentation protocols or pathology reporting standards) may result in inconsistent AI performance across clinical environments. A recent study revealed that the same AI system exhibited different levels of clinical impact depending on the implementation setting.18 These findings underscore the importance of multicenter RCTs and the accumulation of real-world data across diverse regions and institutions.
Furthermore, even when pooled sensitivity and specificity appear robust across studies, regional differences in baseline disease prevalence, such as the predominance of ESCC in Asia and EAC in Western countries, may influence the predictive value and the perceived clinical utility of AI-assisted detection. Variations in case mix, imaging platforms, and endoscopic protocols can further introduce spectrum effects and domain shifts, potentially altering the diagnostic performance during external validation. Therefore, multicenter validation across diverse prevalence settings and careful calibration of the operating thresholds are essential before global implementation.
Moreover, although highly specialized AI models have been developed for discrete tasks (e.g., detection of Borrmann type IV gastric cancer, undifferentiated-type gastric carcinoma, or ESCC), routine endoscopy requires a comprehensive AI system capable of handling multiple lesion types simultaneously. Needing to run numerous narrow-purpose algorithms during a single procedure is impractical. Whether such a comprehensive model should be constructed by incrementally expanding the base model or by integrating several high-performing task-specific models remains an open question. Furthermore, the addition of new functions may degrade the performance of previously optimized modules, which requires careful design and rigorous validation.
Ultimately, progress in the application of AI to upper GI interventions will depend on multicenter collaboration, external validation, incorporation of XAI frameworks, and the development of workflow-compatible systems that can safely and efficiently transition from research to routine clinical use.
AI in lower GI endoscopy
AI has emerged as a powerful tool for enhancing diagnostic precision, procedural quality, and clinical decision-making in lower GI endoscopy. Applications have centered primarily on two domains: (1) detection and characterization of colorectal neoplasia and (2) assessment and monitoring of IBD. Colonoscopy remains the gold standard for colorectal cancer (CRC) prevention; however, its performance is highly operator-dependent and subject to variations in experience, fatigue, and concentration, which lead to inconsistent ADR and polyp detection rates. Similarly, in IBD, conventional endoscopic and histological indices suffer from interobserver variability and incomplete reproducibility. Therefore, AI-based tools have been developed to standardize diagnostic quality, reduce variability, and improve efficiency for both neoplastic and inflammatory diseases of the lower GI tract.

1) AI in colorectal neoplasia

CADe systems utilize deep-learning algorithms to analyze real-time colonoscopy videos and provide instant visual alerts for potential polyps. RCTs and meta-analyses have consistently shown that CADe increases ADR and reduces adenoma miss rates without significantly prolonging the procedure time or increasing adverse events. A comprehensive meta-analysis of 33 RCTs reported a relative 24% increase in ADR with CADe use.19 Nevertheless, real-world results have been modest. In a prospective observational study of 502 colonoscopies, ADR increased from 30.5% to 34.7% with CADe, but this difference was not statistically significant.20 Multivariable analysis suggested that effects observed in controlled trials may overestimate benefits in routine practice.21 Moreover, CADe performance can vary by colon segment, lesion morphology, withdrawal time, and endoscopist behavior.22 To date, there is no conclusive evidence that CADe use translates into reduced interval CRC incidence, underscoring the need for long-term, outcome-based evaluation.
Importantly, the vast majority of CADe systems evaluated in both RCTs and real-world implementation studies have been developed and deployed using standard high-definition white-light colonoscopy without requiring image-enhanced modalities such as NBI or virtual chromoendoscopy.23,24 This suggests that clinically meaningful AI-assisted detection can be integrated into routine colonoscopy workflows without additional imaging platforms.
CADx systems are designed to provide real-time histologic prediction of detected polyps, supporting two clinical strategies: “diagnose-and-leave” and “resect-and-discard.” The European Society of GI Endoscopy (ESGE) recommends a sensitivity ≥90%, specificity ≥80%, and negative predictive value (NPV) ≥90% for implementation of the diagnose-and-leave strategy. A meta-analysis of nine studies (3,237 polyps) reported sensitivity 87.3%, specificity 88.9%, and NPV 93.6%, meeting NPV and specificity thresholds but not sensitivity.25 For the resect-and-discard approach, ESGE requires both sensitivity and specificity ≥80%. A pooled analysis of 11 studies (7,400 polyps) demonstrated 87% sensitivity but only 75% specificity, indicating that CADx accuracy remains insufficient for broad clinical adoption.26 Real-time implementation studies have likewise shown limited impact on management decisions compared with conventional practice.27
In terms of imaging platforms, many CADx models for optical diagnosis have incorporated image-enhanced modalities, particularly NBI or blue-light imaging, to better capture the surface and vascular patterns. However, recent in vivo studies suggest that white-light-based CADx systems can also achieve expert-level diagnostic accuracy, indicating that image-enhanced endoscopy may be beneficial but is not uniformly required for AI-assisted optical diagnosis in routine practice.28
Overall, although CADe reliably improves lesion detection in controlled settings and CADx demonstrates promising diagnostic performance, both technologies require robust prospective validation in diverse real-world environments before full integration into CRC prevention workflows. Furthermore, although sensitivity and specificity are theoretically prevalence-independent, regional variations in baseline adenoma prevalence, such as differences among screening programs and population risk profiles, may substantially influence predictive values and perceived clinical utility through classic spectrum effects. Accordingly, multicenter RCTs and multivariable analyses across heterogeneous populations are essential to confirm generalizability and ensure a consistent clinical impact.

2) AI in IBD

IBD, including ulcerative colitis (UC) and Crohn’s disease (CD), requires continuous endoscopic evaluation for diagnosis, activity assessment, and treatment monitoring. Conventional indices, including the Mayo Endoscopic Subscore, UC Endoscopic Index of Severity, and CD Endoscopic Index of Severity, are widely used but are limited by subjectivity and interobserver variability. AI systems have been developed to automate these scoring processes and to provide more objective and reproducible evaluations. Recent CNN models trained on colonoscopy videos have predicted disease severity with area under the curve (AUC) values exceeding 0.90, achieving segment-level accuracy above 85%.29-31 One automated video frame analysis system demonstrated high diagnostic accuracy under experimental conditions but decreased performance in real-world multicenter validation.32 A meta-analysis reported pooled sensitivity and specificity of 87% and 92%, respectively, for AI in identifying endoscopic remission, with an overall AUC of 0.96.33 These findings suggest that AI can standardize endoscopic activity scoring and reduce observer variability, although generalizability across institutions remains a concern.
AI has also been applied to the histological evaluation of IBD. Computer vision algorithms analyzing digitized biopsy slides achieved an AUC of 0.95, with a sensitivity of 96% and specificity of 80% for detecting histologic inflammation in UC.29 Other models have accurately classified histologic remission with concordance rates exceeding 85%.31 Integrating endoscopic and histologic AI outputs could facilitate a unified “treat-to-target” approach and mitigate inter-pathologist variability that complicates both clinical care and research endpoints.
AI applications in capsule endoscopy and imaging have revolutionized the assessment of small-bowel and pan-intestinal inflammation in CD. Deep-learning analysis of pan-intestinal capsule endoscopy has achieved a sensitivity of 96%–97% and specificity of 90%–93% while reducing reading time by more than 90%.34 Another AI-assisted capsule system outperformed human gastroenterologists in detecting ulcers and grading disease activity.30 These tools promise to enhance diagnostic efficiency and standardization, although most remain in the pre-implementation phase and require multicenter validation. In addition to imaging, natural language processing (NLP) has been used to extract IBD-related data from electronic health records. NLP algorithms have identified extraintestinal manifestations with 94% accuracy (κ=0.76), outperforming administrative coding methods.35 This capability underscores the value of AI in large-scale population surveillance, real-world evidence generation, and automated outcome tracking.
Despite substantial progress, the real-world adoption of AI in IBD remains limited. Most studies have been retrospective, single-center, or simulation-based, raising concerns about external validity and reproducibility.32,36 The lack of explainability further hinders clinician trust, as black-box predictions without interpretable visualization (e.g., heat maps that highlight mucosal inflammation) are difficult to integrate into high-stakes decision-making. Moreover, while AI-assisted endoscopy has demonstrated the potential for dysplasia detection, no evidence has shown a reduction in the incidence of CRC among patients with IBD. Finally, cost-effectiveness, interoperability with hospital information systems, and regulatory approvals remain underexplored. Overall, AI in IBD shows strong potential for reducing interobserver variability, standardizing endoscopic and histologic scoring, enhancing capsule endoscopy efficiency, and enabling population-level data analytics. However, clinical translation requires large-scale, prospective, multicenter validation, incorporation of XAI frameworks, and demonstration of tangible outcome benefits before routine implementation can be achieved. The representative studies on AI applications in lower GI endoscopy, including colorectal neoplasia detection, optical diagnosis, and IBD assessment, are summarized in Table 2.19-21,25,26,29,31,33-35

3) AI in video capsule endoscopy

Video capsule endoscopy (VCE), particularly small-bowel capsule endoscopy (SBCE), has become a first-line, minimally invasive modality for evaluating obscure GI bleeding and small-bowel inflammation, including CD, owing to its ability to provide comprehensive visualization of otherwise inaccessible intestinal segments.37,38 Each SBCE examination generates tens of thousands of frames, rendering interpretation labor-intensive, reader-dependent, and susceptible to interobserver variability and missed lesions, particularly in high-volume settings.37,39 These inherent characteristics make VCE one of the most technically compatible platforms for AI implementation, as large annotated datasets and repetitive image patterns are well suited to deep-learning architectures.40,41
Early AI applications in SBCE focused on automated detection of clinically relevant lesions, including erosions, ulcers, angiodysplasia, vascular abnormalities, protruding lesions, and active bleeding, primarily using CNNs trained on large image datasets.40,42 In a study including 53,555 capsule endoscopy images, a multi-class CNN achieved an overall accuracy of 99%, with a sensitivity of 88%, specificity of 99%, and high discriminative performance across lesion categories while also stratifying hemorrhagic potential according to Saurin’s classification.42 Such risk-oriented classification moves beyond simple lesion detection toward clinically meaningful triage and prioritization in patients with suspected small-bowel bleeding.42
Subsequent development has shifted from static image-based models toward workflow-integrated systems capable of analyzing full-length capsule videos and simultaneously performing GI localization and multi-lesion detection.37,39 A recently developed CNN-based model trained on over 87,000 images achieved localization AUC values exceeding 0.99 and lesion detection AUCs between 0.98 and 0.99 in internal validation, whereas external validation demonstrated a reduction in mean reading time from approximately 54–57 minutes to under 10 minutes without compromising detection rates.37,39 By incorporating two-step detection-classification pipelines, temporal smoothing, and ensemble learning, these systems more closely replicate the human reading process and address frame redundancy inherent to capsule imaging.37,39 Such integrated “localization+detection” architectures represent a significant step toward clinically deployable AI-assisted SBCE interpretation and may enhance workflow efficiency in both specialized and resource-limited settings.37
In IBD, particularly CD, VCE and panenteric capsule endoscopy enable comprehensive assessment of small- and large-bowel involvement within a single examination, facilitating disease staging and monitoring.38 AI-assisted capsule interpretation in CD has demonstrated high diagnostic accuracy, with a recent meta-analysis of 8 studies and 11 AI models reporting pooled sensitivity of 94%, specificity of 97%, and an overall AUC of 0.99 for CD detection.43 Deep-learning systems applied to panenteric capsule videos have shown promising performance in detecting CD-related ulcers and erosions while substantially shortening review time compared with conventional reading, suggesting potential integration into treat-to-target strategies and longitudinal disease monitoring.38,43 However, most available evidence remains retrospective, and robust data linking AI-assisted capsule analysis to improved long-term clinical outcomes, such as reduced flares or hospitalization rates, are lacking.38,43
Beyond investigational models, commercially available capsule platforms have incorporated AI-enhanced or machine learning–based reading modes that prioritize informative frames and accelerate review, functioning primarily as triage tools rather than fully autonomous diagnostic systems.44 Clinical evaluations indicate substantial reductions in reading time across commercial platforms, yet detection performance varies according to device and algorithm, and a proportion of lesions may remain undetected.44
Despite encouraging diagnostic metrics and substantial efficiency gains, several challenges currently limit routine clinical integration of AI-assisted capsule endoscopy.40,44 Most published models rely on retrospectively collected, single-platform datasets, often derived from limited geographic populations, raising concerns regarding spectrum bias and cross-platform generalizability.37,39,43 External validation cohorts are frequently modest in size, and lesion categories in many models remain restricted to common bleeding-related abnormalities, excluding tumors, strictures, and other clinically relevant entities.37,39 Moreover, the predominance of black-box deep-learning architectures necessitates the development of XAI frameworks to enhance transparency, clinician trust, and regulatory acceptance.40,41
Future research priorities include large-scale prospective multicenter trials, cross-platform training and validation, expansion of lesion taxonomies, and integration of AI-assisted capsule interpretation with clinical endpoints, such as rebleeding, need for device-assisted enteroscopy, CD activity control, and healthcare utilization.38,43,44 Collaborative efforts to establish standardized, annotated, multi-institutional capsule datasets may further accelerate benchmarking and facilitate regulatory harmonization across devices.37,44
In summary, AI in VCE has demonstrated high diagnostic accuracy for small-bowel bleeding and CD detection, along with some of the most pronounced workflow efficiency gains observed in GI AI applications, particularly in terms of reading time reduction.37,39,42,43 Nevertheless, definitive evidence demonstrating improved long-term clinical outcomes, cost-effectiveness, and cross-platform robustness continues to evolve, and careful validation is essential before AI-assisted capsule endoscopy can be fully integrated into routine GI practice.38,43,44
AI in PB endoscopy
AI has been increasingly applied in PB endoscopy, including EUS, ERCP, and cholangioscopy. These applications span a wide spectrum, from lesion detection and disease classification to procedural guidance and postprocedural risk prediction. Although early studies have reported excellent diagnostic accuracy,45 current research has shifted toward real-world validation, external generalizability, and workflow integration to support clinical decision-making in complex PB diseases.

1) AI in EUS

AI-assisted EUS has shown promising accuracy in differentiating pancreatic ductal adenocarcinoma (PDAC) from benign pancreatic conditions, often outperforming conventional diagnostic criteria and guideline-based approaches.46,47 Several CNN-based models have demonstrated diagnostic accuracies exceeding 90% for distinguishing solid pancreatic masses with sensitivity and specificity comparable to those of expert endosonographers. Importantly, AI assistance has been shown to improve the diagnostic performance of novice operators in RCTs,48 highlighting its potential as an educational and quality-enhancing tool.
AI has also been applied to needle-based confocal laser endomicroscopy (nCLE), enabling high-resolution, cellular-level imaging of pancreatic cystic lesions. One AI system has achieved superior accuracy compared with expert reviewers in the detection of high-grade dysplasia within intraductal papillary mucinous neoplasms.49 Although nCLE is not yet widely available, this illustrates the potential of AI to standardize the interpretation of advanced imaging modalities.
Beyond image analysis, multimodal decision-support frameworks, such as CompCyst and EBM-CFMM, have integrated clinical, radiologic, and molecular data to stratify the risk of malignancy and reduce unnecessary surgery.50,51 These models exemplify a shift from image-based detection to comprehensive, data-driven diagnostic assistance. However, most available studies remain retrospective and involve a single center, with limited external validation. Accordingly, AI in EUS is evolving from proof-of-concept detection models to integrated clinical decision-support systems that combine multiomic and imaging data to guide management.
In EUS applications, the modality requirements are model-specific rather than universally mandated. Most early AI models were developed using conventional B-mode EUS images alone, demonstrating high diagnostic accuracy for differentiating PDAC from benign lesions.52 However, subsequent studies have incorporated contrast-enhanced EUS, elastography, or Doppler imaging to improve lesion characterization and vascular assessment, with some reports suggesting incremental performance gains in selected clinical scenarios.53,54 Importantly, current evidence does not support a universal requirement for advanced imaging modalities, as robust diagnostic performance has also been achieved using B-mode–based models alone.52 Prospective head-to-head comparisons and real-world validation studies are needed to clarify the incremental value of multimodal integration over standard B-mode–based approaches.53

2) AI in ERCP

Applications of AI in ERCP have primarily focused on procedural guidance, anatomical recognition, and the prediction of complication risk. AI algorithms have been trained to automatically identify the papilla of Vater, estimate cannulation difficulty, and assist in navigation during complex procedures.55 Machine learning models, particularly ensemble-based approaches, have achieved robust predictive performance for post-ERCP pancreatitis and post-ERCP cholecystitis, both of which represent clinically significant adverse events.56-58 These results indicate that ensemble methods can deliver clinically feasible risk stratification and guide tailored prophylaxis. For example, Zhang et al.58 developed a random forest model using data from 1,117 patients with common bile duct stones and intact gallbladders, achieving an AUC of 0.89 and accuracy of 85.5% for predicting post-ERCP cholecystitis. External validation yielded comparable results (AUC, 0.889), and a publicly available online platform was created to facilitate clinical application. Such tools may guide individualized prophylactic strategies such as early antibiotic administration or closer post-procedure monitoring in high-risk patients.
AI has also shown the potential for improving patient selection for ERCP. One machine learning model outperformed guideline-based scoring systems in predicting choledocholithiasis,59 suggesting its potential to reduce unnecessary invasive procedures. Collectively, these findings highlight the potential of AI to optimize the safety and efficacy of ERCP. Nonetheless, large-scale prospective multicenter studies are required to confirm these results and ensure their generalizability across diverse clinical settings.

3) AI in cholangioscopy

In cholangioscopy, AI has been applied primarily for the differentiation between benign and malignant biliary strictures, which is a long-standing diagnostic challenge in PB endoscopy. Recent studies have reported diagnostic accuracies exceeding 90% and area under the receiver operating characteristic values >0.95 for distinguishing malignant from benign strictures.45,60-62 Multicenter validation has confirmed reproducibility across institutions,45 and several independent investigations have demonstrated similar performances.63,64 AI-enhanced cholangioscopy systems have also shown the potential to support real-time biopsy targeting, improve sampling accuracy, and reduce the rate of indeterminate strictures. These advances may shorten diagnostic delays, reduce unnecessary surgical exploration, and enhance patient outcomes. However, despite their high algorithmic accuracy, most existing studies are retrospective and single-center, underscoring the need for multicenter prospective validation and real-time integration with endoscopic hardware systems. Representative studies on AI applications in PB endoscopy, including EUS, ERCP, and cholangioscopy, are summarized in Table 3.46,48-50,55-60,64

4) Current limitations and translational challenges

Despite the substantial progress, AI applications in PB endoscopy face several barriers to clinical translation. First, many studies to date have been retrospective and based on limited single-center datasets, raising concerns about overfitting and lack of generalizability. Second, workflow integration remains technically challenging, particularly for real-time AI assistance during EUS and ERCP procedures, where dynamic anatomical variations and motion artifacts complicate image interpretation. Third, explainability and accountability are crucial because black-box predictions in high-stakes interventional settings may pose a medicolegal risk. Furthermore, the relatively low prevalence and referral-based case selection of many PB conditions may limit generalizability across institutions with different case volumes and disease spectra. Finally, regulatory approval and cost-effectiveness analyses are lacking, limiting widespread implementation. Future research should focus on multicenter prospective validation with standardized reporting frameworks, integration of XAI into real-time procedural guidance, and cross-platform compatibility with endoscopic systems from different manufacturers. Ultimately, the successful translation of AI into PB endoscopy depends on balancing algorithmic sophistication with clinical practicality, ensuring safety, interpretability, and tangible improvements in patient outcomes.
Legal liability and regulatory considerations
The integration of AI into GI endoscopy raises complex medicolegal questions regarding liability allocation and regulatory functions. Currently, most AI systems in endoscopy function as assistive decision-support tools rather than autonomous diagnostic agents.65 Accordingly, the ultimate responsibility for clinical decision-making remains with the endoscopist.66 In cases of diagnostic error, liability may be distributed among three parties: the physician (clinical judgment), the manufacturer (algorithmic or design defects), and the healthcare institution (implementation and oversight).67
A central legal debate concerns whether AI should be regarded as an adjunct tool or a semi-autonomous decision-maker. As AI systems evolve toward greater autonomy, the traditional framework, where physicians bear primary responsibility, may become increasingly challenged.66 Furthermore, the use of black-box deep-learning models complicates malpractice assessment because limited explainability may hinder the judicial evaluation of negligence or standard-of-care deviations.68
The need for informed consent is another unresolved issue. Although no universal regulatory requirement currently mandates the explicit disclosure of AI assistance during endoscopy, ethical arguments increasingly support transparency when AI meaningfully influences diagnostic decisions.66
From a regulatory perspective, most endoscopic AI systems are classified as software as a medical device by agencies such as the U.S. Food and Drug Administration, the European CE marking framework, and the Korean Ministry of Food and Drug Safety.69-71 However, regulatory pathways for continuous learning or adaptive AI models remain under development.
Although large-scale malpractice litigation involving endoscopic AI has not yet emerged, delayed cancer detection or unnecessary interventions attributable to AI errors may generate future legal challenges.66 Accordingly, integration of XAI, robust validation, and clear documentation protocols will be essential to ensure accountability and maintain clinician trust.
Implementation considerations: online vs. offline AI systems
AI tools in GI endoscopy differ substantially in their deployment modes and hardware integration. Most CADe systems for polyp detection or blind-spot monitoring are designed for real-time (online) applications and are integrated into standard high-definition endoscopy processors. Although compatibility with specific processor platforms may be necessary, these systems do not generally require dedicated endoscopic models.
In contrast, certain CADx systems for optical characterization may operate in real time, but can incorporate image-enhanced modalities, such as NBI, depending on the algorithm design. However, advanced imaging is not universally required for AI implementation.
In VCE, AI algorithms are predominantly deployed in an offline post-acquisition reading mode within proprietary software platforms, functioning primarily as triage or frame prioritization tools rather than real-time diagnostic systems.
In EUS, many AI models have been developed using stored B-mode images and applied offline during image analysis, although real-time integration is an area of ongoing development. Similarly, most risk prediction models for ERCP are currently implemented offline using clinical datasets rather than live procedural imaging.
Overall, most currently available AI systems do not require special endoscopic hardware. However, successful implementation may depend on compatible processors, software ecosystems, or specific imaging modalities according to the intended clinical application.
XAI in CADx for colorectal polyps
The adoption of CADx systems in colonoscopy holds strong potential for optimizing the management of colorectal polyps, particularly through the “resect-and-discard” and “diagnose-and-leave” strategies.72 These approaches aim to improve procedural efficiency, reduce pathology-related costs, and maintain diagnostic safety. Despite remarkable advances in AI model performance, recent large-scale investigations have shown that human–AI collaboration in the optical diagnosis of colorectal polyps has not produced substantial accuracy gains compared with expert endoscopists working independently.73 This suggests that the next stage of progress depends not merely on algorithmic accuracy but also on the development of explainability and trust in human–AI interactions.74
The “black box” nature of deep learning–based CADx systems remains a key barrier to clinician confidence. Opaque decision-making processes hinder the ability of the physician to interpret, verify, and take responsibility for AI-assisted decision-making. In response, professional organizations, such as the American Society for GI Endoscopy AI Task Force, have emphasized that AI in clinical endoscopy must be explainable, transparent, and aligned with human reasoning.75 Accordingly, XAI has become essential for ensuring transparency, interpretability, and regulatory compliance. Multiple XAI techniques are gaining traction, including Local Interpretable Model-Agnostic Explanations, Shapley Additive Explanations (SHAP), and Gradient-weighted Class Activation Mapping (Grad-CAM), which help to visualize model attention and interpret decision features.74,76 Other approaches, such as Case-Based Reasoning or hybrid deep-learning systems, incorporate clinically meaningful endoscopic criteria and further enhance interpretability.
For example, the “niceAI” system links deep feature extraction to the NBI International Colorectal Endoscopic Classification, translating AI features into human-recognizable radiomic parameters.75 This mapping enables endoscopists to audit AI-derived decisions, identify potential errors, and ensure alignment with established diagnostic logic. Such traceability fulfills ethical and legal requirements while reinforcing user trust and accountability in clinical practice.
However, the clinical integration of explainable CADx systems remains limited. Most supporting evidence originates from retrospective validations rather than from prospective real-time clinical trials. To bridge this gap, the Korean Society of GI Endoscopy AI Research Group is conducting multicenter prospective studies to evaluate the niceAI system against conventional black-box models across endoscopists with varying experience levels. In parallel, the development of a standardized optical diagnostic workflow and implementation of digital literacy training programs for clinicians will be critical for achieving effective, safe, and explainable CADx adoption in everyday endoscopic practice.
AI-assisted ERCP: ampulla detection and cannulation difficulty
AI is poised to redefine GI endoscopy as both a diagnostic and therapeutic technology, particularly for the intricate management of ampulla of Vater (AoV) lesions. These lesions pose significant clinical challenges owing to their morphological variability and rarity, which often hinder accurate diagnosis and procedural efficiency. Recent developments have demonstrated that AI can help overcome these limitations by enhancing real-time visualization, optimizing cannulation strategies, and improving procedural success.
AI algorithms have been applied to automatically detect the precise location of the AoV during ERCP and to predict the difficulty of selective cannulation.53,55 In a study by Kim et al.,55 an AI-assisted system accurately identified the AoV with performance comparable to that of human experts, despite variations in shape, size, and surface texture. For predicting cannulation difficulty, the model achieved a recall of 72% for easy cases and 61% for difficult cases, suggesting the potential for preprocedural planning and early recognition of challenging cases. Such predictive insights may improve ERCP quality by guiding the timely use of advanced techniques or providing secondary operator support.
In addition, AI-powered hierarchical classification frameworks that integrate white-light imaging and NBI have shown promising results in differentiating ampullary lesions ranging from normal mucosa to adenoma and carcinoma. By refining lesion classification and optimizing indications for endoscopic papillectomy, these models can enhance curative outcomes while minimizing unnecessary or high-risk resections. This application exemplifies the expanding role of AI not only as a diagnostic adjunct but also as a procedural intelligence layer capable of informing endoscopic therapy decisions.
Despite encouraging preliminary data, real-world implementation remains limited by dataset size, image heterogeneity, and institutional variability. Large-scale multicenter validation, integration with endoscopic navigation systems, and real-time deployment within ERCP suites will be essential for clinical translation.
AI-generated Common Data Model-compatible endoscopy reports
In parallel with algorithmic advances, AI-driven standardization of endoscopy reporting is emerging as a foundational step toward scalable, data-driven gastroenterology.
The Common Data Model (CDM), particularly the OMOP-CDM framework developed by the Observational Health Data Sciences and Informatics consortium, enables the harmonization of heterogeneous data from hospitals, health systems, and countries into a unified analytical structure.77 Although CDMs have traditionally focused on structured data such as diagnoses, medications, and procedures, interest is increasing in extending their scope to unstructured data, including endoscopic images and narrative reports.78
Endoscopy reports, which are rich in clinical content, are typically composed of free text and thus remain poorly structured for large-scale analysis. Recent advances in NLP and LLMs, such as GPT and Med-PaLM, allow the automated extraction and codification of clinical information with high contextual fidelity. These models can identify key findings (e.g., a 2 cm sessile polyp in the ascending colon resected with a snare) and convert them into standardized CDM-compatible codes (e.g., Systematized Nomenclature of Medicine–Clinical Terms and Logical Observation Identifiers Names and Codes).79,80 This capability lays the foundation for AI-generated, CDM-ready endoscopy reports.
In principle, a deep-learning system can integrate image recognition and NLP pipelines to analyze endoscopic videos, detect procedural events (e.g., biopsy, polypectomy), and automatically populate CDM tables such as procedure_occurrence or observation. Although full end-to-end automation has not yet been clinically implemented, the early components are under validation. For example, NLP algorithms have demonstrated high accuracy in extracting findings from colonoscopy reports,81 and prototype systems can map this information to Observational Medical Outcomes Partnership (OMOP)-CDM fields.82 The FUJIFILM AR-C1 system further illustrates this trajectory: it automatically detects device usage (e.g., snare, biopsy forceps), captures procedural timestamps, and generates structured summaries for clinician verification.83 Although current iterations do not directly output CDM-formatted data, they demonstrate the technical feasibility of real-time generation of structured reports.84
The convergence of AI-based automation and CDM-compatible structuring will not only streamline endoscopic documentation but also enable large-scale data analysis, multicenter collaboration, and real-world evidence generation. Future LLMs fine-tuned to extensive corpora of medical narratives and endoscopy reports could serve as the backbone of real-time CDM-compatible reporting systems. Such models would be capable of both transforming existing free-text documentation into standardized structured formats and assisting endoscopists during procedures by automatically suggesting report elements as each step is completed. In the future, endoscopy reports may no longer be confined to routine clinical documentation but could evolve into clinical research-ready, interoperable data assets for outcome tracking and population health studies through the seamless integration of deep learning and standardized data modeling.
Impact of AI on endoscopic training and education
In addition to its influence on diagnostic performance, AI is poised to reshape the paradigm of endoscopic training. Traditional apprenticeship-based models largely depend on procedural volume and subjective faculty assessments. Although case number thresholds offer practical benchmarks, they do not necessarily reflect true cognitive or technical competence and are susceptible to interobserver variability.85
AI-enabled systems provide objective, data-driven performance metrics during real-time procedures.86,87 Tools capable of monitoring blind-spot coverage, mucosal inspection quality, withdrawal time, lesion detection rate, photo-documentation completeness, and ADR enable structured and reproducible feedback. Such capabilities support a shift from volume-based certification to competence-based training, in which advancement to independent practice is guided by validated performance thresholds rather than procedure counts alone.88
The integration of AI with endoscopic simulation platforms further enhances educational potential. AI-assisted simulators can analyze inspection patterns, lesion recognition accuracy, and missed lesion profiles, thereby delivering personalized feedback tailored to individual trainees. This adaptive learning model may accelerate skill acquisition and standardize evaluation across training centers.88 Importantly, the observed improvement was most pronounced among novice endoscopists, suggesting that AI may help bridge the expertise gap during early training. By reinforcing rapid pattern recognition within a defined decision window, CADx systems may accelerate learning curves and promote a more standardized optical diagnosis performance across varying levels of experience.89
However, increasing reliance on CADe systems raises concerns regarding the assessment of intrinsic diagnostic competence.90 Continuous AI assistance may obscure whether lesion detection reflects independent cognitive skills or technological augmentation.90 Future training frameworks may therefore require structured “AI-withdrawal” or AI-blinded evaluation phases to ensure the preservation of core perceptual and decision-making abilities.
Ultimately, AI will not replace endoscopists but will redefine how competence is measured, validated, and maintained throughout professional development. As AI becomes embedded into routine practice, training systems must evolve to ensure that technological augmentation strengthens, rather than substitutes, fundamental clinical expertise.
AI is poised to transform GI endoscopy by enhancing diagnostic precision, reducing interobserver variability, and supporting data-driven clinical decision-making across a wide spectrum of diseases. From upper and lower GI neoplasia to IBD and PB disorders, AI systems have consistently demonstrated diagnostic performance comparable to, and in some domains exceeding, that of expert endoscopists. However, broad clinical adoption remains contingent on addressing key challenges, most notably the need for rigorous multicenter validation, explainable and transparent model design, seamless workflow integration, and clearly defined medicolegal accountability. Future research should emphasize prospective, real-world trials, incorporation of XAI frameworks, and the development of standardized, interoperable data infrastructures to support sustainable implementation. With continued refinement and responsible integration, AI is expected to evolve from a promising experimental adjunct to an indispensable and integral component of routine endoscopic practice.
Table 1.
Representative studies on artificial intelligence in upper gastrointestinal endoscopy
Study Year Country Dataset type Sample size (n) Aim Main results
Blind spot detection/quality
 Wu et al.1 2019 China Video (RT) 5,438 AI-assisted EGD for blind spot monitoring (WISENSE) Blind spot rate decreased from 22.5% to 5.9%
 Wu et al.2 2021 China Video (RCT) 1,012 AI to reduce missed gastric neoplasms (ENDOANGEL-LD) Miss rate decreased from 25.6% to 6.4%
 Ahn et al.4 2025 Korea Video (RT) 1,000 Evaluate real-world blind spot detection Improved completeness of endoscopic photo-documentation
 Chan et al.5 2025 Hong Kong Simulation/video - Training assistance for novice endoscopists Improved procedural quality compared with control
Helicobacter pylori/gastritis
 Parkash et al.6 2024 Multi (Asia) Image (still) Meta-analysis (8) AI diagnosis of H. pylori infection Pooled sensitivity and specificity approximately 90%–95%
 Turtoi et al.8 2024 Europe Image (still) Meta-analysis (10) Automated gastritis grading using CNN Diagnostic accuracy approximately 95%
Gastric premalignant/malignant lesions
 Arribas et al.9 2020 Global Image/short video Pooled (18) Standalone AI for gastric neoplasia Sensitivity 94%, specificity 87%
 Zhou et al.13 2024 China Video (RT) 2,000 Real-time AI detection of early gastric cancer Detection accuracy improved by 19%
Esophageal premalignant/malignant lesions
 Guidozzi et al.16 2023 Multi Image (still) Pooled (15) AI diagnosis of EAC and ESCC For ESCC: sensitivity, 91.2%, specificity, 80%; for EAC: sensitivity, 93.1%; specificity, 86.9%
 Yuan et al.17 2024 China Video (RT) 1,290 AI detection of superficial ESCC Sensitivity, 94.8%; specificity, 84.3%

AI, artificial intelligence; RT, real-time; EGD, esophagogastroduodenoscopy; RCT, randomized controlled trial; CNN, convolutional neural network; EAC, esophageal adenocarcinoma; ESCC, esophageal squamous cell carcinoma; -, not applicable.

Table 2.
Representative studies on artificial intelligence in lower gastrointestinal endoscopy
Study Year Country Dataset type Sample size (n) Aim Main results
Colorectal neoplasia (CADe)
 Lou et al.19 2023 Multi Video (RT) 33 RCTs Evaluate CADe effect on adenoma detection rate Adenoma detection rate increased by 24% (relative improvement)
 Rønborg et al.20 2024 Denmark Video (real-world) 502 Assess CADe in daily colonoscopy Adenoma detection rate 34.7% vs 30.5% (not significant)
 Makar et al.21 2025 Global Video (RT) Meta-analysis (35) Summarize CADe performance Improved adenoma detection and reduced miss rate
Colorectal neoplasia (CADx)
 Hassan et al.25 2024 Global Image (optical) 3,237 Polyps Validate CADx 'diagnose-and-leave' strategy Sensitivity 87.3%, specificity 88.9%, negative predictive value 93.6%
 Hassan et al.26 2024 Global Image (optical) 7,400 Polyps Evaluate CADx 'resect-and-discard' strategy Sensitivity 87%, specificity 75%
IBD
 Bossuyt et al.29 2020 Belgium Endoscopy+pathology 40 AI for ulcerative colitis inflammation scoring Area under the ROC curve 0.95; agreement comparable to experts
 Byrne et al.31 2023 USA/Canada Video (colonoscopy) 341 AI prediction of ulcerative colitis severity Accuracy 87%, area under the ROC curve 0.94
 Lv et al.33 2023 China Video (RT) Meta-analysis (8) AI detection of ulcerative colitis remission Pooled sensitivity 87%, specificity 92%, area under the ROC curve 0.96
 Brodersen et al.34 2024 Europe Video (capsule) 1,062 AI-assisted Crohn’s activity detection Sensitivity 96–97%, specificity 90–93%, reading time reduced by 90%
 Stidham et al.35 2023 USA Clinical data (NLP) 18,000 NLP for extraintestinal IBD features Accuracy 94%, κ=0.76

RT, real-time; RCT, randomized controlled trial; CADe, computer-aided detection; CADx, computer-aided diagnosis; AI, artificial intelligence; IBD, inflammatory bowel disease; ROC, receiver operating characteristic curve; NLP, natural language processing.

Table 3.
Representative studies on artificial intelligence in pancreatobiliary endoscopy
Study Year Country Dataset type Sample size (n) Aim Main results
EUS
 Saraiva et al.46 2024 Portugal/USA Image (EUS) 378 Differentiate PDAC from benign lesions Accuracy greater than 90%
 Cui et al.48 2024 China Video (EUS) 68 Trainees Enhance EUS interpretation accuracy among novice endosonographers Accuracy improved from 69% to 90%
 Krishna et al.49 2025 USA Video (nCLE) 64 Detect high-grade dysplasia in intraductal papillary mucinous neoplasms Sensitivity and specificity approximately 78%, outperforming experts
 Springer et al.50 2019 USA Multimodal 426 Stratify pancreatic cyst malignancy risk using CompCyst model Reduced unnecessary surgery by approximately 60%
ERCP
 Kim et al.55 2021 Korea Video (ERCP) 300 Detect ampulla and predict cannulation difficulty Recall rate 72% for easy cases and 61% for difficult cases
 Archibugi et al.56 2023 Italy Clinical data 1,150 Predict post-ERCP pancreatitis using machine learning Area under the ROC curve 0.67 (internal validation)
 Takahashi et al.57 2024 Japan Clinical data - Externally validate post-ERCP pancreatitis prediction Area under the ROC curve 0.82
 Zhang et al.58 2022 China Clinical data 1,117 Predict post-ERCP cholecystitis using random forest model Area under the ROC curve 0.89; accuracy 85.5%
 Jovanovic et al.59 2014 Serbia Clinical data 380 To predict the need for therapeutic ERCP using AI AI model outperformed conventional guidelines
Cholangioscopy
 Marya et al.60 2023 USA Video (cholangioscopy) 2.3 Million frames Differentiate benign and malignant biliary strictures using CNN Area under the ROC curve 0.94; accuracy 90.6%
 Robles-Medranda et al.64 2023 Ecuador / EU Video (cholangioscopy) 1,200 Validate CNN model for biliary neoplasia detection Sensitivity 91.7%, specificity 94.4%, area under the ROC curve 0.95

EUS, endoscopic ultrasound; PDAC, pancreatic ductal adenocarcinoma; nCLE, needle-based confocal laser endomicroscopy; ERCP, endoscopic retrograde cholangiopancreatography; AI, artificial intelligence; CNN, convolutional neural network; ROC, receiver operating characteristic; -, not applicable.

  • 1. Wu L, Zhang J, Zhou W, et al. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy. Gut 2019;68:2161–2169.ArticlePubMedPMC
  • 2. Wu L, Shang R, Sharma P, et al. Effect of a deep learning-based system on the miss rate of gastric neoplasms during upper gastrointestinal endoscopy: a single-centre, tandem, randomised controlled trial. Lancet Gastroenterol Hepatol 2021;6:700–708.ArticlePubMed
  • 3. Zhang L, Lu Z, Yao L, et al. Effect of a deep learning-based automatic upper GI endoscopic reporting system: a randomized crossover study (with video). Gastrointest Endosc 2023;98:181–190.ArticlePubMed
  • 4. Ahn BY, Lee J, Seol J, et al. Evaluation of an artificial intelligence-based system for real-time high-quality photodocumentation during esophagogastroduodenoscopy. Sci Rep 2025;15:4693.ArticlePubMedPMCPDF
  • 5. Chan SM, Chan D, Yip HC, et al. Artificial intelligence-assisted esophagogastroduodenoscopy improves procedure quality for endoscopists in early stages of training. Endosc Int Open 2025;13:a25476645.Article
  • 6. Parkash O, Lal A, Subash T, et al. Use of artificial intelligence for the detection of Helicobacter pylori infection from upper gastrointestinal endoscopy images: an updated systematic review and meta-analysis. Ann Gastroenterol 2024;37:665–673.Article
  • 7. Fang S, Liu Z, Qiu Q, et al. Diagnosing and grading gastric atrophy and intestinal metaplasia using semi-supervised deep learning on pathological images: development and validation study. Gastric Cancer 2024;27:343–354.ArticlePubMedPMCPDF
  • 8. Turtoi DC, Brata VD, Incze V, et al. Artificial intelligence for the automatic diagnosis of gastritis: a systematic review. J Clin Med 2024;13:4818.ArticlePubMedPMC
  • 9. Arribas J, Antonelli G, Frazzoni L, et al. Standalone performance of artificial intelligence for upper GI neoplasia: a meta-analysis. Gut 2021;70:1458–1468.ArticlePubMed
  • 10. Tokat M, van Tilburg L, Koch AD, et al. Artificial Intelligence in upper gastrointestinal endoscopy. Dig Dis 2022;40:395–408.ArticlePubMedPDF
  • 11. Kikuchi R, Okamoto K, Ozawa T, et al. Endoscopic artificial intelligence for image analysis in gastrointestinal neoplasms. Digestion 2024;105:419–435.ArticlePubMedPDF
  • 12. Yuan XL, Zhou Y, Liu W, et al. Artificial intelligence for diagnosing gastric lesions under white-light endoscopy. Surg Endosc 2022;36:9444–9453.ArticlePubMedPDF
  • 13. Zhou R, Liu J, Zhang C, et al. Efficacy of a real-time intelligent quality-control system for the detection of early upper gastrointestinal neoplasms: a multicentre, single-blinded, randomised controlled trial. EClinicalMedicine 2024;75:102803.ArticlePubMedPMC
  • 14. Dong Z, Zhu Y, Du H, et al. The effectiveness of a computer-aided system in improving the detection rate of gastric neoplasm and early gastric cancer: study protocol for a multi-centre, randomized controlled trial. Trials 2023;24:323.ArticlePubMedPMCPDF
  • 15. Mori Y, Ishihara R, Ogata H, et al. Artificial intelligence in gastrointestinal endoscopy: the japan gastroenterological endoscopy society position statements. Dig Endosc 2025;37:1116–1122.ArticlePubMed
  • 16. Guidozzi N, Menon N, Chidambaram S, et al. The role of artificial intelligence in the endoscopic diagnosis of esophageal cancer: a systematic review and meta-analysis. Dis Esophagus 2023;36:doad048.ArticlePubMedPMCPDF
  • 17. Yuan XL, Liu W, Lin YX, et al. Effect of an artificial intelligence-assisted system on endoscopic diagnosis of superficial oesophageal squamous cell carcinoma and precancerous lesions: a multicentre, tandem, double-blind, randomised controlled trial. Lancet Gastroenterol Hepatol 2024;9:34–44.ArticlePubMed
  • 18. Huang L, Xu M, Li Y, et al. Gastric neoplasm detection of computer-aided detection-assisted esophagogastroduodenoscopy changes with implement scenarios: a real-world study. J Gastroenterol Hepatol 2024;39:2787–2795.ArticlePubMedPDF
  • 19. Lou S, Du F, Song W, et al. Artificial intelligence for colorectal neoplasia detection during colonoscopy: a systematic review and meta-analysis of randomized clinical trials. EClinicalMedicine 2023;66:102341.ArticlePubMedPMC
  • 20. Rønborg SN, Ujjal S, Kroijer R, et al. Assessing the potential of artificial intelligence to enhance colonoscopy adenoma detection in clinical practice: a prospective observational trial. Clin Endosc 2024;57:783–789.ArticlePubMedPMCPDF
  • 21. Makar J, Abdelmalak J, Con D, et al. Use of artificial intelligence improves colonoscopy performance in adenoma detection: a systematic review and meta-analysis. Gastrointest Endosc 2025;101:68–81.ArticlePubMed
  • 22. Djinbachian R, Taghiakbari M, Calce SI, et al. Withdrawal time, CADe and adenoma detection: a prospective study. Gut 2025 Jun 30 [Epub]. https://doi.org/10.1136/gutjnl-2025-335380ArticlePubMed
  • 23. Bencardino S, Lodola I, Centanni L, et al. Artificial intelligence in advanced endoscopic imaging: transforming optical diagnosis in gastroenterology. Front Med (Lausanne) 2025;12:1719145.ArticlePubMedPMC
  • 24. Ali H, Muzammil MA, Dahiya DS, et al. Artificial intelligence in gastrointestinal endoscopy: a comprehensive review. Ann Gastroenterol 2024;37:133–141.ArticlePubMedPMC
  • 25. Hassan C, Misawa M, Rizkala T, et al. Computer-aided diagnosis for leaving colorectal polyps in situ: a systematic review and meta-analysis. Ann Intern Med 2024;177:919–928.ArticlePubMedPDF
  • 26. Hassan C, Rizkala T, Mori Y, et al. Computer-aided diagnosis for the resect-and-discard strategy for colorectal polyps: a systematic review and meta-analysis. Lancet Gastroenterol Hepatol 2024;9:1010–1019.ArticlePubMed
  • 27. Rizkala T, Menini M, Massimi D, et al. Role of artificial intelligence for colon polyp detection and diagnosis and colon cancer. Gastrointest Endosc Clin N Am 2025;35:389–400.Article
  • 28. García-Rodríguez A, Tudela Y, Córdova H, et al. In vivo computer-aided diagnosis of colorectal polyps using white light endoscopy. Endosc Int Open 2022;10:E1201–E1207.ArticlePubMedPMC
  • 29. Bossuyt P, Nakase H, Vermeire S, et al. Automatic, computer-aided determination of endoscopic and histological inflammation in patients with mild to moderate ulcerative colitis based on red density. Gut 2020;69:1778–1786.ArticlePubMed
  • 30. Ruan G, Qi J, Cheng Y, et al. Development and validation of a deep neural network for accurate identification of endoscopic images from patients with ulcerative colitis and crohn's disease. Front Med (Lausanne) 2022;9:854677.ArticlePubMedPMC
  • 31. Byrne MF, Panaccione R, East JE, et al. Application of deep learning models to improve ulcerative colitis endoscopic disease activity scoring under multiple scoring systems. J Crohns Colitis 2023;17:463–471.ArticlePubMedPDF
  • 32. Yao H, Najarian K, Gryak J, et al. Fully automated endoscopic disease activity assessment in ulcerative colitis. Gastrointest Endosc 2021;93:728–736.ArticlePubMed
  • 33. Lv B, Ma L, Shi Y, et al. A systematic review and meta-analysis of artificial intelligence-diagnosed endoscopic remission in ulcerative colitis. iScience 2023;26:108120.Article
  • 34. Brodersen JB, Jensen MD, Leenhardt R, et al. Artificial intelligence-assisted analysis of pan-enteric capsule endoscopy in patients with suspected Crohn's disease: a study on diagnostic performance. J Crohns Colitis 2024;18:75–81.ArticlePubMedPDF
  • 35. Stidham RW, Yu D, Zhao X, et al. Identifying the presence, activity, and status of extraintestinal manifestations of inflammatory bowel disease using natural language processing of clinical notes. Inflamm Bowel Dis 2023;29:503–510.ArticlePubMedPDF
  • 36. George AT, Rubin DT. Artificial intelligence in inflammatory bowel disease. Gastrointest Endosc Clin N Am 2025;35:367–387.Article
  • 37. Liu HR. Deep learning meets small-bowel capsule endoscopy: a step toward faster and more consistent diagnosis of obscure gastrointestinal bleeding. World J Gastrointest Endosc 2025;17:113184.ArticlePubMedPMC
  • 38. Ukashi O, Soffer S, Klang E, et al. Capsule endoscopy in inflammatory bowel disease: panenteric capsule endoscopy and application of artificial intelligence. Gut Liver 2023;17:516–528.ArticlePubMedPMC
  • 39. Kwon YS, Park TY, Kim SE, et al. Deep learning-based localization and lesion detection in capsule endoscopy for patients with suspected small-bowel bleeding. World J Gastroenterol 2025;31:106819.ArticlePubMedPMC
  • 40. George AA, Tan JL, Kovoor JG, et al. Artificial intelligence in capsule endoscopy: development status and future expectations. Mini Invasive Surg 2024;8:4.Article
  • 41. Lodola I, D'Amico F, Danese S, et al. Artificial intelligence in inflammatory bowel disease endoscopy: a review of current evidence and a critical perspective on future challenges. Ther Adv Gastroenterol 2025;18:17562848251350896.ArticlePubMedPMCPDF
  • 42. Mascarenhas Saraiva MJ, Afonso J, Ribeiro T, et al. Deep learning and capsule endoscopy: automatic identification and differentiation of small bowel lesions with distinct haemorrhagic potential using a convolutional neural network. BMJ Open Gastroenterol 2021;8:e000753.ArticlePubMedPMC
  • 43. Bin Y, Peng R, Lee Y, et al. Artificial intelligence-assisted capsule endoscopy for detecting lesions in Crohn's disease: a systematic review and meta-analysis. Front Artif Intell 2025;8:1531362.ArticlePubMedPMC
  • 44. Giordano A, Romero-Mascarell C, González-Suárez B, et al. Integration of artificial intelligence-enhanced capsule endoscopy in clinical practice: a review of market-available tools for clinical practice. Dig Dis Sci 2025;70:2966–2976.ArticlePubMedPMCPDF
  • 45. Mascarenhas M, Almeida MJ, González-Haba M, et al. Artificial intelligence for automatic diagnosis and pleomorphic morphological characterization of malignant biliary strictures using digital cholangioscopy. Sci Rep 2025;15:5447.ArticlePDF
  • 46. Saraiva MM, González-Haba M, Widmer J, et al. Deep learning and automatic differentiation of pancreatic lesions in endoscopic ultrasound: a transatlantic study. Clin Transl Gastroenterol 2024;15:e00771.ArticlePubMedPMC
  • 47. Lee D, Jesry F, Maliekkal JJ, et al. Application of artificial intelligence in pancreatic cyst management: a systematic review. Cancers (Basel) 2025;17:2558.ArticlePubMedPMC
  • 48. Cui H, Zhao Y, Xiong S, et al. Diagnosing solid lesions in the pancreas with multimodal artificial intelligence: a randomized crossover trial. JAMA Netw Open 2024;7:e2422454.ArticlePubMedPMC
  • 49. Krishna SG, Abdelbaki A, Li Z, et al. Towards automating risk stratification of intraductal papillary mucinous Neoplasms: artificial intelligence advances beyond human expertise with confocal laser endomicroscopy. Pancreatology 2025;25:658–666.ArticlePubMedPMC
  • 50. Springer S, Masica DL, Dal Molin M, et al. A multimodality test to guide the management of patients with a pancreatic cyst. Sci Transl Med 2019;11:eaav4772.ArticlePubMedPMC
  • 51. Lavista Ferres JM, Oviedo F, Robinson C, et al. Performance of explainable artificial intelligence in guiding the management of patients with a pancreatic cyst. Pancreatology 2024;24:1182–1191.ArticlePubMed
  • 52. Dumitrescu EA, Ungureanu BS, Cazacu IM, et al. Diagnostic value of artificial intelligence-assisted endoscopic ultrasound for pancreatic cancer: a systematic review and meta-analysis. Diagnostics (Basel) 2022;12:309.ArticlePubMedPMC
  • 53. Araújo CC, Frias J, Mendes F, et al. Unlocking the potential of AI in EUS and ERCP: a narrative review for pancreaticobiliary disease. Cancers (Basel) 2025;17:1132.ArticlePubMedPMC
  • 54. Bharwad AV, Ahuja R, Jain P, et al. Artificial intelligence in pancreatobiliary endoscopy: current advances, opportunities, and challenges. J Clin Med 2025;14:7519.ArticlePubMedPMC
  • 55. Kim T, Kim J, Choi HS, et al. Artificial intelligence-assisted analysis of endoscopic retrograde cholangiopancreatography image for identifying ampulla and difficulty of selective cannulation. Sci Rep 2021;11:8381.ArticlePubMedPMCPDF
  • 56. Archibugi L, Ciarfaglia G, Cárdenas-Jaén K, et al. Machine learning for the prediction of post-ERCP pancreatitis risk: a proof-of-concept study. Dig Liver Dis 2023;55:387–393.ArticlePubMed
  • 57. Takahashi H, Ohno E, Furukawa T, et al. Artificial intelligence in a prediction model for postendoscopic retrograde cholangiopancreatography pancreatitis. Dig Endosc 2024;36:463–472.Article
  • 58. Zhang X, Yue P, Zhang J, et al. A novel machine learning model and a public online prediction platform for prediction of post-ERCP-cholecystitis (PEC). EClinicalMedicine 2022;48:101431.ArticlePubMedPMC
  • 59. Jovanovic P, Salkic NN, Zerem E. Artificial neural network predicts the need for therapeutic ERCP in patients with suspected choledocholithiasis. Gastrointest Endosc 2014;80:260–268.ArticlePubMed
  • 60. Marya NB, Powers PD, Petersen BT, et al. Identification of patients with malignant biliary strictures using a cholangioscopy-based deep learning artificial intelligence (with video). Gastrointest Endosc 2023;97:268–278.ArticlePubMed
  • 61. Tanisaka Y, Hawes R. Peroral cholangioscopy: past, present and future. Clin Endosc 2025;58:360–369.ArticlePubMedPMCPDF
  • 62. Nakai Y, Hakuta R, Shimamatsu Y, et al. Endoscopic approach to indeterminate biliary strictures. Clin Endosc 2026;59:40–48.ArticlePubMedPMCPDF
  • 63. Saraiva MM, Ribeiro T, Ferreira JPS, et al. Artificial intelligence for automatic diagnosis of biliary stricture malignancy status in single-operator cholangioscopy: a pilot study. Gastrointest Endosc 2022;95:339–348.ArticlePubMed
  • 64. Robles-Medranda C, Baquerizo-Burgos J, Alcivar-Vasquez J, et al. Artificial intelligence for diagnosing neoplasia on digital cholangioscopy: development and multicenter validation of a convolutional neural network model. Endoscopy 2023;55:719–727.ArticlePubMedPMC
  • 65. Ahmad OF, Mori Y, Bretthauer M, et al. The legal and ethical framework for artificial intelligence in gastrointestinal endoscopy: a world endoscopy organization international consensus statement. Ann Intern Med 2026;179:270–275.ArticlePubMedPDF
  • 66. Elamin S, Duffourc M, Berzin TM, et al. Artificial Intelligence and medical liability in gastrointestinal endoscopy. Clin Gastroenterol Hepatol 2024;22:1165–1169.ArticlePubMed
  • 67. Jovanovic I. AI in endoscopy and medicolegal issues: the computer is guilty in case of missed cancer? Endosc Int Open 2020;8:E1385–E1386.ArticlePubMedPMC
  • 68. Morley J, Machado CCV, Burr C, et al. The ethics of AI in health care: a mapping review. Soc Sci Med 2020;260:113172.ArticlePubMed
  • 69. Abulibdeh R, Celi LA, Sejdić E. The illusion of safety: a report to the FDA on AI healthcare product approvals. PLOS Digit Health 2025;4:e0000866.ArticlePubMedPMC
  • 70. Ebad SA, Alhashmi A, Amara M, et al. Artificial intelligence-based software as a medical device (AI-SaMD): a systematic review. Healthcare (Basel) 2025;13:817.ArticlePubMedPMC
  • 71. Park SH, Dean G, Ortiz EM, et al. Overview of South Korean guidelines for approval of large language or multimodal models as medical devices: key features and areas for improvement. Korean J Radiol 2025;26:519–523.ArticlePubMedPMCPDF
  • 72. Hassan C, Balsamo G, Lorenzetti R, et al. Artificial intelligence allows leaving-in-situ colorectal polyps. Clin Gastroenterol Hepatol 2022;20:2505–2513.ArticlePubMed
  • 73. Mori Y, Hassan C. Computer-aided diagnosis of colorectal polyps: assisted or autonomous? Clin Endosc 2025;58:514–517.ArticlePDF
  • 74. Mascarenhas M, Mendes F, Martins M, et al. Explainable AI in digestive healthcare and gastrointestinal endoscopy. J Clin Med 2025;14:549.Article
  • 75. Parasa S, Repici A, Berzin T, et al. Framework and metrics for the clinical use and implementation of artificial intelligence algorithms into endoscopy practice: recommendations from the American Society for Gastrointestinal Endoscopy Artificial Intelligence Task Force. Gastrointest Endosc 2023;97:815–824.Article
  • 76. Shin Y, Bae JH, Kim J, et al. A novel approach to overcome black box of AI for optical diagnosis in colonoscopy. Sci Rep 2025;15:21220.ArticlePDF
  • 77. Hripcsak G, Duke JD, Shah NH, et al. Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform 2015;216:574–578.PubMedPMC
  • 78. Hripcsak G, Shang N, Peissig PL, et al. Facilitating phenotype transfer using a common data model. J Biomed Inform 2019;96:103253.ArticlePubMedPMC
  • 79. Wu S, Roberts K, Datta S, et al. Deep learning in clinical natural language processing: a methodical review. J Am Med Inform Assoc 2020;27:457–470.ArticlePubMedPMCPDF
  • 80. Sundgren M, Mistry R, Maeztu G. Harnessing unstructured data and hospital interoperability. Appl Clin Trials 2024;33.
  • 81. Ryu B, Yoon E, Kim S, et al. Transformation of pathology reports into the common data model with oncology module: use case for colon cancer. J Med Internet Res 2020;22:e18526.ArticlePubMedPMC
  • 82. Keloth VK, Banda JM, Gurley M, et al. Representing and utilizing clinical textual data for real world studies: an OHDSI approach. J Biomed Inform 2023;142:104343.ArticlePubMedPMC
  • 83. Naito T, Nosaka T, Tanaka T, et al. Usefulness of an artificial intelligence-based colonoscopy report generation support system. Clin Endosc 2025;58:327–330.ArticlePubMedPMCPDF
  • 84. Qu JY, Li Z, Su JR, et al. Development and validation of an automatic image-recognition endoscopic report generation system: a multicenter study. Clin Transl Gastroenterol 2020;12:e00282.ArticlePubMedPMC
  • 85. Yang D, Hayat M, Draganov PV. Training in advanced endoscopy: current methods, challenges, and emerging innovations. Tech Innov Gastrointest Endosc 2026;28:250958.Article
  • 86. Zhang Z, Chen BS, Du L, et al. Expert-AI collaborative training for novice endoscopists: a path to enhanced efficiency. Bioengineering (Basel) 2025;12:582.ArticlePubMedPMC
  • 87. Lu Z, Zhang L, Yao L, et al. Assessment of the role of artificial intelligence in the association between time of day and colonoscopy quality. JAMA Netw Open 2023;6:e2253840.Article
  • 88. Ho JCL, Qian Z, Lau LHS, et al. Artificial intelligence in digestive endoscopy training-the past, present, and future. Dig Endosc 2026;38:e70047.ArticlePubMedPMCPDF
  • 89. Kang HY, Kang S, Chung GE, et al. Structured integration of an artificial intelligence-based system for the optical diagnosis of colorectal polyps. Gut Liver 2026;20:86–96.ArticlePubMedPMC
  • 90. Budzyń K, Romańczyk M, Kitala D, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol 2025;10:896–903.ArticlePubMed

Figure & Data

REFERENCES

    Citations

    Citations to this article as recorded by  

      • Cite
        CITE
        export Copy Download
        Close
        Download Citation
        Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

        Format:
        • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
        • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
        Include:
        • Citation for the content below
        Translating artificial intelligence into clinical practice for gastrointestinal endoscopy: current applications and future perspectives
        Close
      • XML DownloadXML Download
      Related articles
      Translating artificial intelligence into clinical practice for gastrointestinal endoscopy: current applications and future perspectives
      Translating artificial intelligence into clinical practice for gastrointestinal endoscopy: current applications and future perspectives
      Study Year Country Dataset type Sample size (n) Aim Main results
      Blind spot detection/quality
       Wu et al.1 2019 China Video (RT) 5,438 AI-assisted EGD for blind spot monitoring (WISENSE) Blind spot rate decreased from 22.5% to 5.9%
       Wu et al.2 2021 China Video (RCT) 1,012 AI to reduce missed gastric neoplasms (ENDOANGEL-LD) Miss rate decreased from 25.6% to 6.4%
       Ahn et al.4 2025 Korea Video (RT) 1,000 Evaluate real-world blind spot detection Improved completeness of endoscopic photo-documentation
       Chan et al.5 2025 Hong Kong Simulation/video - Training assistance for novice endoscopists Improved procedural quality compared with control
      Helicobacter pylori/gastritis
       Parkash et al.6 2024 Multi (Asia) Image (still) Meta-analysis (8) AI diagnosis of H. pylori infection Pooled sensitivity and specificity approximately 90%–95%
       Turtoi et al.8 2024 Europe Image (still) Meta-analysis (10) Automated gastritis grading using CNN Diagnostic accuracy approximately 95%
      Gastric premalignant/malignant lesions
       Arribas et al.9 2020 Global Image/short video Pooled (18) Standalone AI for gastric neoplasia Sensitivity 94%, specificity 87%
       Zhou et al.13 2024 China Video (RT) 2,000 Real-time AI detection of early gastric cancer Detection accuracy improved by 19%
      Esophageal premalignant/malignant lesions
       Guidozzi et al.16 2023 Multi Image (still) Pooled (15) AI diagnosis of EAC and ESCC For ESCC: sensitivity, 91.2%, specificity, 80%; for EAC: sensitivity, 93.1%; specificity, 86.9%
       Yuan et al.17 2024 China Video (RT) 1,290 AI detection of superficial ESCC Sensitivity, 94.8%; specificity, 84.3%
      Study Year Country Dataset type Sample size (n) Aim Main results
      Colorectal neoplasia (CADe)
       Lou et al.19 2023 Multi Video (RT) 33 RCTs Evaluate CADe effect on adenoma detection rate Adenoma detection rate increased by 24% (relative improvement)
       Rønborg et al.20 2024 Denmark Video (real-world) 502 Assess CADe in daily colonoscopy Adenoma detection rate 34.7% vs 30.5% (not significant)
       Makar et al.21 2025 Global Video (RT) Meta-analysis (35) Summarize CADe performance Improved adenoma detection and reduced miss rate
      Colorectal neoplasia (CADx)
       Hassan et al.25 2024 Global Image (optical) 3,237 Polyps Validate CADx 'diagnose-and-leave' strategy Sensitivity 87.3%, specificity 88.9%, negative predictive value 93.6%
       Hassan et al.26 2024 Global Image (optical) 7,400 Polyps Evaluate CADx 'resect-and-discard' strategy Sensitivity 87%, specificity 75%
      IBD
       Bossuyt et al.29 2020 Belgium Endoscopy+pathology 40 AI for ulcerative colitis inflammation scoring Area under the ROC curve 0.95; agreement comparable to experts
       Byrne et al.31 2023 USA/Canada Video (colonoscopy) 341 AI prediction of ulcerative colitis severity Accuracy 87%, area under the ROC curve 0.94
       Lv et al.33 2023 China Video (RT) Meta-analysis (8) AI detection of ulcerative colitis remission Pooled sensitivity 87%, specificity 92%, area under the ROC curve 0.96
       Brodersen et al.34 2024 Europe Video (capsule) 1,062 AI-assisted Crohn’s activity detection Sensitivity 96–97%, specificity 90–93%, reading time reduced by 90%
       Stidham et al.35 2023 USA Clinical data (NLP) 18,000 NLP for extraintestinal IBD features Accuracy 94%, κ=0.76
      Study Year Country Dataset type Sample size (n) Aim Main results
      EUS
       Saraiva et al.46 2024 Portugal/USA Image (EUS) 378 Differentiate PDAC from benign lesions Accuracy greater than 90%
       Cui et al.48 2024 China Video (EUS) 68 Trainees Enhance EUS interpretation accuracy among novice endosonographers Accuracy improved from 69% to 90%
       Krishna et al.49 2025 USA Video (nCLE) 64 Detect high-grade dysplasia in intraductal papillary mucinous neoplasms Sensitivity and specificity approximately 78%, outperforming experts
       Springer et al.50 2019 USA Multimodal 426 Stratify pancreatic cyst malignancy risk using CompCyst model Reduced unnecessary surgery by approximately 60%
      ERCP
       Kim et al.55 2021 Korea Video (ERCP) 300 Detect ampulla and predict cannulation difficulty Recall rate 72% for easy cases and 61% for difficult cases
       Archibugi et al.56 2023 Italy Clinical data 1,150 Predict post-ERCP pancreatitis using machine learning Area under the ROC curve 0.67 (internal validation)
       Takahashi et al.57 2024 Japan Clinical data - Externally validate post-ERCP pancreatitis prediction Area under the ROC curve 0.82
       Zhang et al.58 2022 China Clinical data 1,117 Predict post-ERCP cholecystitis using random forest model Area under the ROC curve 0.89; accuracy 85.5%
       Jovanovic et al.59 2014 Serbia Clinical data 380 To predict the need for therapeutic ERCP using AI AI model outperformed conventional guidelines
      Cholangioscopy
       Marya et al.60 2023 USA Video (cholangioscopy) 2.3 Million frames Differentiate benign and malignant biliary strictures using CNN Area under the ROC curve 0.94; accuracy 90.6%
       Robles-Medranda et al.64 2023 Ecuador / EU Video (cholangioscopy) 1,200 Validate CNN model for biliary neoplasia detection Sensitivity 91.7%, specificity 94.4%, area under the ROC curve 0.95
      Table 1. Representative studies on artificial intelligence in upper gastrointestinal endoscopy

      AI, artificial intelligence; RT, real-time; EGD, esophagogastroduodenoscopy; RCT, randomized controlled trial; CNN, convolutional neural network; EAC, esophageal adenocarcinoma; ESCC, esophageal squamous cell carcinoma; -, not applicable.

      Table 2. Representative studies on artificial intelligence in lower gastrointestinal endoscopy

      RT, real-time; RCT, randomized controlled trial; CADe, computer-aided detection; CADx, computer-aided diagnosis; AI, artificial intelligence; IBD, inflammatory bowel disease; ROC, receiver operating characteristic curve; NLP, natural language processing.

      Table 3. Representative studies on artificial intelligence in pancreatobiliary endoscopy

      EUS, endoscopic ultrasound; PDAC, pancreatic ductal adenocarcinoma; nCLE, needle-based confocal laser endomicroscopy; ERCP, endoscopic retrograde cholangiopancreatography; AI, artificial intelligence; CNN, convolutional neural network; ROC, receiver operating characteristic; -, not applicable.


      Clin Endosc : Clinical Endoscopy Twitter Facebook
      Close layer
      TOP