Abstract
- Since its introduction in 2000, capsule endoscopy (CE) has transformed gastrointestinal (GI) diagnostics by enabling noninvasive visualization of the entire GI tract using a swallowable capsule. However, CE still has several limitations, including long reading times, inter-reader variability, and missed lesions due to poor image quality or incomplete examinations. Recent advances in artificial intelligence (AI) have significantly improved CE interpretation. Deep-learning models, particularly convolutional neural networks, can detect small-bowel lesions with accuracy comparable to that of expert endoscopists, while greatly reducing reading time. AI algorithms can also provide objective assessments of small-bowel cleanliness. Transformer-based models can further enhance video-level analysis by recognizing global patterns and sequential relationships among CE images. In addition, foundation models demonstrate high adaptability and robust performance across different CE systems and a wide range of lesion types. Future AI-assisted CE reading is expected to integrate real-time image analysis, autonomous capsule movement, and multimodal sensing technologies to create an intelligent diagnostic platform. Ultimately, AI is transforming CE into an efficient, reliable, and data-driven diagnostic tool suitable for diverse clinical settings. Furthermore, AI-assisted CE is extending its clinical utility beyond small-bowel lesion detection to the comprehensive evaluation of the stomach and colon.
-
Keywords: Artificial intelligence; Capsule endoscopy; Deep learning; Gastrointestinal tract
INTRODUCTION
Capsule endoscopy (CE) was introduced in 2000 as a breakthrough technology that allows noninvasive visualization of the entire gastrointestinal (GI) tract using a swallowable capsule. CE has since become an essential diagnostic tool for evaluating small-bowel bleeding, inflammation, and other mucosal lesions.1,2 Despite its advantages, CE still has practical limitations. Reading a complete CE video typically takes 30 to 120 minutes, and interpretation depends heavily on the experience and concentration of the endoscopist. Missed lesions may occur due to prolonged reading time, fatigue, and poor image quality. In addition, suboptimal bowel preparation and residual debris can impair visualization quality, and there is still no standardized method for objective assessment. These factors limit the clinical efficiency and widespread adoption of CE.3
In recent years, artificial intelligence (AI) has been increasingly applied to medical imaging analysis. Deep learning–based approaches, particularly convolutional neural network (CNN) models, have demonstrated substantial potential to assist CE reading by automatically detecting lesions, classifying findings, and grading bowel cleanliness. More recently, transformer and foundation models have introduced new possibilities for video-level analysis and cross-device generalization (Fig. 1).4,5 This review summarizes current progress in AI applications for CE, focusing on clinically relevant advances, including lesion detection, image-quality assessment, and the emerging roles of transformer and foundation models. To provide a clear comparison of these evolving technologies, their key clinical implications and characteristics are summarized in Table 1. The review also discusses future directions and remaining challenges for safe and effective clinical implementation.
CNN MODEL-BASED ANALYSIS IN CE IMAGES
Concept
The technical concept of CNN models in CE centers on a deep learning architecture designed to automate the analysis of the thousands of images generated during a single procedure. CNNs function by hierarchically learning spatial patterns, such as color, edges, and shapes, through multiple layers of convolutional kernels. These kernels extract features that evolve from low-level textures to high-level lesion morphologies, enabling the model to recognize complex visual information. By training on large datasets, CNNs can effectively differentiate normal mucosa from abnormalities. This automated process not only reduces subjective variability among endoscopists but also minimizes the risk of human error during the labor-intensive reading process.3-5
Clinical applications
1) Automated lesion detection
Early studies applying CNN models to CE images were first reported around 2019. These models aimed to automatically detect small-bowel vascular, inflammatory, and tumorous lesions with high accuracy.6-8 In one of the earliest studies, a CNN-based Single Shot Multibox Detector achieved an accuracy of 90.8% for identifying erosions and ulcers, suggesting the feasibility of automated lesion detection. However, false-positive findings were frequent due to artifacts such as bubbles or debris.7 A milestone multicenter study collected over 100 million CE images from approximately 7,000 patients across 77 institutions. Using this dataset, a CNN model classified normal mucosa, inflammation, ulcers, polyps, lymphangiectasia, and bleeding lesions with more than 99% sensitivity. This model demonstrated diagnostic accuracy superior to that of expert endoscopists (99% vs. 75%) and significantly reduced mean reading time (5.9 vs. 96.6 minutes).9
Following these landmark results, subsequent research focused on binary classification models that determine the presence or absence of significant lesions in CE images. Using this approach, a CNN model was developed based on more than 400,000 CE images to perform binary classification between clinically significant lesions (inflammatory lesions and vascular lesions or bleeding) and normal findings. This model demonstrated a diagnostic accuracy of 98%.10 In another study using an Inception-ResNet-V2 model, the lesion detection rate increased from 29.5% to 63.1% (p=0.01), and the mean reading time decreased from 1,343 to 667 minutes (p=0.03).11 CNN-based binary classification was particularly effective in improving diagnostic performance and reducing reading time for less-experienced readers.
After the development of single-lesion detection and binary classification, further studies applied CNN models to the detection of multiple lesion types in CE images. Using the RetinaNet model, one study demonstrated excellent recognition of inflammatory, vascular, and tumorous lesions (area under the curve [AUC], 0.996, 0.950, and 0.950, respectively).12 Similarly, a CNN model based on ResNet50 and the Single Shot Multibox Detector improved the per-patient lesion detection rate compared with the QuickView mode (99% vs. 89%, p<0.001). In this system, detection rates for inflammatory lesions, angioectasia, protruding lesions, and blood content were 100%, 97%, 99%, and 100%, respectively.13 The SmartScan model, trained on 17 standardized CE structured terminology lesion types, demonstrated a higher detection rate than conventional reading (95.9% vs. 76.1%) and reduced reading time by 89%.14 A recent multicenter prospective study conducted across four countries evaluated CE videos acquired from three different CE systems. CNN-assisted reading showed a markedly higher lesion detection rate (96.1% vs. 76.3%) and sensitivity (97.5% vs. 78.2%) compared with conventional reading. The mean reading time was only 203 seconds per case, demonstrating that accurate interpretation can be achieved within a remarkably short time.15 Notably, reanalysis of previously negative CE videos using a CNN model revealed clinically meaningful lesions in 61% of cases (63/103).16
Overall, CNN-assisted CE reading provides higher diagnostic accuracy, improved efficiency, and reduced human error. It supports endoscopists by highlighting suspicious images and enabling targeted review rather than exhaustive image-by-image inspection.
2) Bowel cleanliness evaluation
Accurate assessment of bowel cleanliness is essential for high-quality CE interpretation. Traditionally, several subjective scales, such as 3- or 5-point grading systems, have been proposed; however, they are not standardized and require manual scoring by the reader.17-20 To overcome these limitations, CNN models have been developed for the automated evaluation of bowel cleanliness.
In one study, a CNN model trained on 600 normal CE images classified images as “adequate” or “inadequate.” When applied to 156 CE videos, the model achieved 90% sensitivity, 83% specificity, and 89% accuracy using a threshold of 79% adequate images.21 Another CNN model introduced a 5-grade cleansing score system (1=visualized mucosa <25% to 5=visualized mucosa >90%) to assess bowel cleanliness. This model applied a cutoff score of 3.25 to define adequate bowel preparation and achieved excellent discrimination between adequate and inadequate cleansing levels (AUC, 0.977).22 When the same CNN model was applied to CE performed immediately after colonoscopy, bowel cleanliness scores correlated with colon preparation quality (poor: 3.39 vs. excellent: 4.42, p<0.05).23 Using the same CNN model, automatic exclusion of low-quality images (poorly visualized mucosa, score 1–2) maintained high diagnostic agreement with conventional reading (k=0.954).24 These findings suggest that CNN-based cleanliness scoring can replace subjective human assessment, providing a standardized and efficient quality-control tool. Automated image reduction by excluding poor-quality images may also reduce review time and improve lesion detection.
3) Structure and localization recognition
Because CE frequently captures images of the esophagus, stomach, and colon before or after small-bowel transit, automated anatomical recognition is valuable for efficient reading. Moreover, magnetically controlled CE is increasingly used for gastric examination, highlighting the need for accurate structural classification.25,26
A study using more than 80,000 CE images with a DenseNet-161 model evaluated the ability to identify anatomic landmarks. The localization rates of the first small-bowel image (96.9% vs. 87.5%, p=0.18) and the first colonic image (81.2% vs. 62.5%, p=0.08) were comparable between AI-assisted and conventional readings, suggesting that AI support can achieve similar localization accuracy with greater efficiency.27 In another study combining CNN and long short-term memory networks, the model classified the stomach, small bowel, and colon with over 95% accuracy. Importantly, this hybrid model estimated the time of small-bowel entrance with a mean difference of only 258 seconds compared with expert assessment, demonstrating strong agreement with human interpretation.28 These results suggest that AI can shorten capsule reading time by automatically identifying anatomic landmarks and transition points during capsule transit, thereby improving the overall efficiency of CE interpretation.
Performance summary and current limitations
In summary, CNN-assisted reading has demonstrated high per-patient lesion detection rates of approximately 99% and diagnostic accuracies of 98% in binary classification tasks. CNN-assisted reading has also reduced reading times by up to 89% compared with conventional reading. CNN models have enabled objective and accurate automated assessment of bowel cleanliness, achieving up to 89% accuracy in cleanliness classification and an AUC of 0.977 using a 5-grade cleansing score system. Furthermore, CNN-based anatomical recognition has shown over 95% accuracy in differentiating the stomach, small bowel, and colon.7,10,14,21,22,28
Despite these performance gains, practical challenges remain, including frequent false-positive findings caused by bubbles, debris, or poor image quality. In addition, because CNNs primarily focus on single-image analysis, they may overlook the temporal continuity essential for complex video interpretation. This limitation has prompted the development of more advanced architectures, such as transformers, which are better suited for sequential data analysis.29
TRANSFORMER MODEL-BASED ANALYSIS IN CE
Concept
The transformer model, originally developed for natural language processing, uses a self-attention mechanism to capture long-range dependencies and global contextual relationships within data. In medical imaging, vision transformers (ViTs) have recently been applied to endoscopic image analysis, offering advantages over CNNs by recognizing patterns across sequential frames rather than within a single isolated image. This approach improves the detection of subtle or transient abnormalities that may be missed by image-by-image analysis.30-32
Clinical applications
1) Video-level lesion detection
In CE videos, thousands of sequential frames must be interpreted within a temporal context, as capsule movement, luminal folds, and residual fluid can obscure small-bowel lesions. Transformer models are well-suited for this purpose because they can integrate information from multiple consecutive frames.
In a study using a transformer model named VWCE-Net, video-level analysis of small-bowel CE was performed and compared with CNN models such as YOLOv4 and XceptionNet. VWCE-Net achieved higher lesion detection performance, with a sensitivity of 95.1% and a specificity of 83.4%. The best results were obtained when the model analyzed groups of four consecutive frames, demonstrating the importance of temporal context in CE video reading.33 This model’s ability to analyze multiple consecutive frames suggests its potential for real-time triage of high-risk segments during CE review. Clinically, this approach indicates that transformers can minimize false negatives by incorporating spatiotemporal continuity, which is particularly useful for detecting subtle mucosal changes or transient bleeding that may not be evident in a single image.
2) Structure and localization recognition
A transformer model applied to upper GI structure recognition classified 15 anatomic regions, including the esophagogastric junction, fundus, body, antrum, and pylorus, with an accuracy of 99.6%, sensitivity of 96.4%, and specificity of 99.8%, demonstrating excellent agreement with endoscopists.34 Another study using a Swin Transformer estimated capsule position and achieved classification accuracies of 93.5% for the stomach, 97.3% for the small bowel, and 98.7% for the colon.35 Errors in anatomic landmark identification and time offsets were within clinically acceptable ranges, indicating potential applicability for real-time localization and automated reporting.
Performance summary and current limitations
Overall, transformer-based analysis enhances diagnostic consistency in CE by leveraging spatiotemporal continuity across sequential frames, achieving nearly 100% accuracy in anatomical recognition and approximately 95% sensitivity in video-level lesion detection, with fewer false negatives than conventional CNN-based single-frame analysis.32-34
However, although transformers are well-suited for video-level analysis, they require substantially greater computational power and larger datasets than standard CNN models. Due to these high demands, their use in routine clinical practice for real-time reading remains challenging, mainly because of slow processing speed and hardware limitations. To address these issues, current research is increasingly focused on foundation models, which are designed to operate more robustly and efficiently across different capsule systems and clinical environments.30,36
FOUNDATION MODEL–BASED ANALYSIS IN CE
Concept
Foundation models represent a new generation of large-scale, pretrained deep learning systems trained on diverse and multimodal datasets. Unlike conventional CNNs or transformer models, which are typically designed for a specific task or dataset, foundation models can adapt to new tasks through fine-tuning or few-shot learning. In medical imaging, these models have demonstrated strong generalizability across modalities and institutions.37,38 CE produces highly variable images depending on device type, bowel cleanliness, transit speed, and underlying small-bowel disease. Foundation models therefore have the potential to provide more reliable interpretation than task-specific CNNs or transformer models.
Clinical applications
The introduction of foundation models into GI endoscopy has opened a new era of universal and expandable AI analysis. The unique characteristics of CE, such as long video sequences, frequent artifacts, and inconsistent preparation quality, require algorithms with strong adaptability. To address this need, a foundation model called EndoFM-LV was trained on 6,469 long endoscopic video samples. This model outperformed previous CNN-based models, EndoFM and EndoSSL, in multitask settings, including classification, segmentation, detection, and workflow recognition. By effectively learning both spatial and temporal context within extended sequences, EndoFM-LV demonstrated high data efficiency and clinical relevance for CE video analysis.39 Building on these findings, multicenter initiatives have focused on creating large, diverse datasets for foundation model development. The GastroNet-5M project compiled more than five million GI endoscopic images from multiple centers to support the training and adaptation of foundation models for various tasks, including CE. Based on this dataset, a foundation model specifically optimized for endoscopic imaging, GastroNet-5M, was developed using self-supervised learning frameworks such as SimCLRv2, MoCov2, and DINO. Through endoscopy-specific self-supervised pretraining, this model achieved 1.6% to 4.6% higher accuracy in angiodysplasia detection and polyp segmentation compared with conventional models such as ResNet50 and ViTs pretrained on natural-image datasets. Moreover, the endoscopy-based foundation model demonstrated improved stability under variable image quality conditions (noise, blur, and illumination changes) and maintained excellent generalization across different endoscopic device manufacturers, image qualities, and patient populations.40 Such large-scale datasets provide the basis for universal AI models that can maintain stable performance regardless of device manufacturer, image resolution, or bowel cleanliness level. They also enable domain adaptation and few-shot learning for rare small-bowel lesions, which is critical for clinical translation.41,42 Representative studies on transformer and foundation models in GI endoscopy are summarized in Table 2.33-35,39,40
Performance summary and current limitations
In CE, foundation model–based analysis could enable comprehensive interpretation that integrates lesion detection, anatomical localization, and quality assessment within a single framework. By pretraining on millions of images, these models achieved up to 4.6% higher accuracy in lesion detection and polyp segmentation compared with conventional CNNs and ViTs. Furthermore, they maintained stable performance even under poor image quality conditions, such as noise, blur, or illumination changes.40
However, validated performance data for small-bowel lesion detection using these large-scale models remain limited. Most current studies focus on anatomical localization or rely on general endoscopic datasets rather than small-bowel–specific images. Therefore, further clinical studies are required to confirm their reliability in the small bowel before routine clinical implementation. With continued refinement and validation, these models may serve as the backbone for next-generation CE platforms capable of real-time, cross-device interpretation.
FUTURE DIRECTIONS OF AI-ASSISTED CE
The next generation of AI-assisted CE is expected to evolve toward autonomous capsule systems capable of self-navigation and real-time image interpretation. Future capsules will focus on integrating multiple analytical tasks, such as lesion detection, anatomical localization, and image-quality assessment, into a single real-time platform (Fig. 2).43,44 As recent advances in transformer and foundation models have enabled highly accurate anatomical localization, the clinical scope of CE is expected to expand beyond the small bowel to include lesion detection in the stomach and colon. In addition, the development of foundation models is expected to facilitate multimodal fusion of imaging, textual, and clinical data, further enhancing diagnostic precision. Future capsules are anticipated to incorporate sensors to monitor biochemical and physiological parameters, such as pH, temperature, microbiome activity, inflammatory markers, and tumor-associated molecules, which will be analyzed by integrated AI systems. These comprehensive models could transform CE from a passive imaging test into an intelligent, decision-supporting diagnostic tool.
CONCLUSIONS
The application of AI in CE is both inevitable and transformative, and its integration into CE has already begun to redefine clinical practice. Recent advances have demonstrated remarkable progress in automated image classification, lesion detection, and bowel quality assessment. Foundation models trained on large and diverse datasets now show strong potential for universal, cross-platform CE interpretation. Global initiatives in Asia and Western countries are converging toward intelligent CE systems capable of evaluating the entire GI tract within a single capsule. In future, a “single AI-empowered capsule” may visualize and autonomously interpret the entire GI tract with high accuracy. To realize these benefits in daily practice, prospective multicenter validation, transparent reporting, and interdisciplinary collaboration among clinicians, engineers, and industry partners are essential to achieve safe, efficient, and patient-centered application of AI-assisted CE.
Conflicts of Interest
The authors have no potential conflicts of interest.
Funding
This work was supported by the Dongguk University, College of Medicine Research Fund of 2025
Author Contributions
Conceptualization: all authors; Investigation: DJO; Project administration: YJL; Supervision: YJL; Visualization: DJO; Writing–original draft: DJO; Writing–review & editing: YJL.
Fig. 1.Evolution and hierarchical relationships of artificial intelligence technologies applied to capsule endoscopy. Artificial intelligence has progressed from traditional machine learning to deep learning architectures such as convolutional neural networks (CNNs) and transformers, ultimately leading to large-scale foundation models. Each stage represents a milestone in the development of capsule endoscopy, from frame-based lesion detection to temporal analysis and multimodal interpretation. CNNs and transformers represent different deep learning architectures, whereas foundation models denote a training paradigm based on large-scale pretraining that can be applied across various architectures.
Fig. 2.Proposed framework for integrated artificial intelligence in capsule endoscopy interpretation. This comprehensive system illustrates how artificial intelligence can support capsule endoscopy by combining multiple functions, including detection, localization, quality control, and clinical decision support, into a single end-to-end workflow for efficient and accurate interpretation.
Table 1.Clinical implications and comparative characteristics of artificial intelligence approaches
|
Category |
Convolutional neural network |
Transformer |
Foundation model |
|
Concept |
Image-based deep learning |
Video-based sequential analysis |
Large-scale pretraining paradigm |
|
Mechanism |
Extracts visual features from single images |
Analyzes temporal context across frames |
Learns generalizable representations from massive datasets |
|
Applications |
Lesion detection; cleanliness grading |
Video-level analysis; anatomical localization |
Cross-platform interpretation |
|
Clinical strengths |
High accuracy in static lesion detection |
Reduced false-positives through temporal modeling |
Multimodal data fusion |
|
Limitations |
Limited contextual understanding |
High resources and cost for real-time use |
Limited clinical evidence for lesion detection |
|
Generalizability |
Low (sensitive to devices and image quality) |
Intermediate (improved consistency) |
High (stable across devices and image quality) |
|
Rare lesions recognition |
Vulnerable (requires large task-specific datasets) |
Limited (requires task-specific training data) |
Strong (adaptable via few-shot learning) |
Table 2.Summary of transformer and foundation models in gastrointestinal endoscopy
|
Study |
Year |
Model |
Dataset |
Outcomes |
Performance |
|
Oh et al.33
|
2023 |
VWCE-Net Transformer |
260 CE cases |
Temporal lesion detection (vs. XceptionNet & YOLOV4) |
Sensitivity, 95.1%; specificity, 83.4% (highest results) |
|
Li et al.34
|
2024 |
Transformer |
3,343 CE cases |
Gastric structure recognition |
Accuracy, 99.6%; specificity, 99.8% |
|
Zhang et al.35
|
2025 |
Swin Transformer |
196 CE cases |
CE position localization (stomach, small bowel, and colon) |
Accuracy (93.5%, 97.3%, 98.7%) |
|
Wang et al.39
|
2025 |
EndoFM-LV |
6,469 CE cases |
Multitask (classification, segmentation, detection, workflow recognition) |
AUC 0.972 |
|
Boers et al.40
|
2025 |
GastroNet-5M |
5 Million endoscopic images (including CE) |
Angiodysplasia detection and polyp segmentation |
1.6%–4.6% higher mean accuracy than ResNet50 or ViT |
REFERENCES
- 1. Iddan G, Meron G, Glukhovsky A, et al. Wireless capsule endoscopy. Nature 2000;405:417.ArticlePDF
- 2. Wang A, Banerjee S, Barth BA, et al. Wireless capsule endoscopy. Gastrointest Endosc 2013;78:805–815.ArticlePubMed
- 3. Oh DJ, Kim KS, Lim YJ. A new active locomotion capsule endoscopy under magnetic control and automated reading program. Clin Endosc 2020;53:395–401.ArticlePubMedPMCPDF
- 4. Hwang Y, Park J, Lim YJ, et al. Application of artificial intelligence in capsule endoscopy: where are we now? Clin Endosc 2018;51:547–551.ArticlePubMedPMCPDF
- 5. Kim M, Jang HJ. Recent technological advances in video capsule endoscopy: a comprehensive review. Clin Endosc 2025 Sep 29 [Epub]. https://doi.org/10.5946/ce.2025.135Article
- 6. Leenhardt R, Vasseur P, Li C, et al. A neural network algorithm for detection of GI angiectasia during small-bowel capsule endoscopy. Gastrointest Endosc 2019;89:189–194.ArticlePubMed
- 7. Aoki T, Yamada A, Aoyama K, et al. Automatic detection of erosions and ulcerations in wireless capsule endoscopy images based on a deep convolutional neural network. Gastrointest Endosc 2019;89:357–363.ArticlePubMed
- 8. Saito H, Aoki T, Aoyama K, et al. Automatic detection and classification of protruding lesions in wireless capsule endoscopy images based on a deep convolutional neural network. Gastrointest Endosc 2020;92:144–151.ArticlePubMed
- 9. Ding Z, Shi H, Zhang H, et al. Gastroenterologist-level identification of small-bowel diseases and normal variants by capsule endoscopy using a deep-learning model. Gastroenterology 2019;157:1044–1054.ArticlePubMed
- 10. Kim SH, Hwang Y, Oh DJ, et al. Efficacy of a comprehensive binary classification model using a deep convolutional neural network for wireless capsule endoscopy. Sci Rep 2021;11:17479.ArticlePubMedPMCPDF
- 11. Park J, Hwang Y, Nam JH, et al. Artificial intelligence that determines the clinical significance of capsule endoscopy images can increase the efficiency of reading. PLoS One 2020;15:e0241474.ArticlePubMedPMC
- 12. Otani K, Nakada A, Kurose Y, et al. Automatic detection of different types of small-bowel lesions on capsule endoscopy images using a newly developed deep convolutional neural network. Endoscopy 2020;52:786–791.ArticlePubMed
- 13. Aoki T, Yamada A, Kato Y, et al. Automatic detection of various abnormalities in capsule endoscopy videos by a deep learning-based system: a multicenter study. Gastrointest Endosc 2021;93:165–173.ArticlePubMed
- 14. Xie X, Xiao YF, Zhao XY, et al. Development and validation of an artificial intelligence model for small bowel capsule endoscopy video review. JAMA Netw Open 2022;5:e2221992.ArticlePubMedPMC
- 15. Mascarenhas Saraiva M, Ferreira J, Afonso J, et al. Real-Life Clinical Validation of Artificial Intelligence-Assisted Detection and Differentiation of Pleomorphic Lesions in Capsule Endoscopy. Am J Gastroenterol 2025 Aug 28 [Epub]. http://doi.org/10.14309/ajg.0000000000003756Article
- 16. Choi KS, Park D, Kim JS, et al. Deep learning in negative small-bowel capsule endoscopy improves small-bowel lesion detection and diagnostic yield. Dig Endosc 2024;36:437–445.ArticlePubMed
- 17. Leighton JA, Brock AS, Semrad CE, et al. Quality indicators for capsule endoscopy and deep enteroscopy. Gastrointest Endosc 2022;96:693–711.ArticlePubMed
- 18. Sidhu R, Shiha MG, Carretero C, et al. Performance measures for small-bowel endoscopy: a European Society of Gastrointestinal Endoscopy (ESGE) Quality Improvement Initiative - Update 2025. Endoscopy 2025;57:366–389.ArticlePubMed
- 19. Park SC, Keum B, Hyun JJ, et al. A novel cleansing score system for capsule endoscopy. World J Gastroenterol 2010;16:875–880.ArticlePubMedPMC
- 20. Brotz C, Nandi N, Conn M, et al. A validation study of 3 grading systems to evaluate small-bowel cleansing for wireless capsule endoscopy: a quantitative index, a qualitative evaluation, and an overall adequacy assessment. Gastrointest Endosc 2009;69:262–270.ArticlePubMed
- 21. Leenhardt R, Souchaud M, Houist G, et al. A neural network-based algorithm for assessing the cleanliness of small bowel during capsule endoscopy. Endoscopy 2021;53:932–936.ArticlePubMed
- 22. Nam JH, Hwang Y, Oh DJ, et al. Development of a deep learning-based software for calculating cleansing score in small bowel capsule endoscopy. Sci Rep 2021;11:4417.ArticlePubMedPMCPDF
- 23. Oh DJ, Hwang Y, Nam JH, et al. Small bowel cleanliness in capsule endoscopy: a case-control study using validated artificial intelligence algorithm. Sci Rep 2022;12:18265.ArticlePubMedPMCPDF
- 24. Oh DJ, Hwang Y, Kim SH, et al. Reading of small bowel capsule endoscopy after frame reduction using an artificial intelligence algorithm. BMC Gastroenterol 2024;24:80.ArticlePubMedPMCPDF
- 25. Xiao YF, Wu ZX, He S, et al. Fully automated magnetically controlled capsule endoscopy for examination of the stomach and small bowel: a prospective, feasibility, two-centre study. Lancet Gastroenterol Hepatol 2021;6:914–921.ArticlePubMed
- 26. Oh DJ, Lee YJ, Kim SH, et al. Efficacy and safety of three-dimensional magnetically assisted capsule endoscopy for upper gastrointestinal and small bowel examination. PLoS One 2024;19:e0295774.ArticlePubMedPMC
- 27. Kwon YS, Park TY, Kim SE, et al. Deep learning-based localization and lesion detection in capsule endoscopy for patients with suspected small-bowel bleeding. World J Gastroenterol 2025;31:106819.ArticlePubMedPMC
- 28. Nam SJ, Moon G, Park JH, et al. Deep learning-based real-time organ localization and transit time estimation in wireless capsule endoscopy. Biomedicines 2024;12:1704.ArticlePubMedPMC
- 29. Dray X, Iakovidis D, Houdeville C, et al. Artificial intelligence in small bowel capsule endoscopy - current status, challenges and future promise. J Gastroenterol Hepatol 2021;36:12–19.ArticlePDF
- 30. Li J, Chen J, Tang Y, et al. Transforming medical imaging with Transformers? A comparative review of key properties, current progresses, and future perspectives. Med Image Anal 2023;85:102762.ArticlePubMedPMC
- 31. Azad R, Kazerouni A, Heidari M, et al. Advances in medical image analysis with vision Transformers: a comprehensive review. Med Image Anal 2024;91:103000.ArticlePubMed
- 32. Ebert N, Stricker D, Wasenmüller O. PLG-ViT: Vision Transformer with Parallel Local and Global Self-Attention. Sensors (Basel) 2023;23:3447.ArticlePubMedPMC
- 33. Oh S, Oh D, Kim D, et al. Video analysis of small bowel capsule endoscopy using a transformer network. Diagnostics (Basel) 2023;13:3133.ArticlePubMedPMC
- 34. Li Q, Xie W, Wang Y, et al. A deep learning application of capsule endoscopic gastric structure recognition based on a transformer model. J Clin Gastroenterol 2024;58:937–943.ArticlePubMed
- 35. Zhang R, Peng B, Liu Y, et al. Localization of capsule endoscope in alimentary tract by computer-aided analysis of endoscopic images. Sensors (Basel) 2025;25:746.ArticlePubMedPMC
- 36. Shamshad F, Khan S, Zamir SW, et al. Transformers in medical imaging: a survey. Med Image Anal 2023;88:102802.ArticlePubMed
- 37. Zhang S, Metaxas D. On the challenges and perspectives of foundation models for medical image analysis. Med Image Anal 2024;91:102996.ArticlePubMed
- 38. Awais M, Naseer M, Khan S, et al. Foundation models defining a new era in vision: a survey and outlook. IEEE Trans Pattern Anal Mach Intell 2025;47:2245–2264.ArticlePubMed
- 39. Wang Z, Liu C, Zhu L, et al. Improving foundation model for endoscopy video analysis via representation learning on long sequences. IEEE J Biomed Health Inform 2025;29:3526–3536.ArticlePubMed
- 40. Boers TGW, Fockens KN, van der Putten JA, et al. Foundation models in gastrointestinal endoscopic AI: Impact of architecture, pre-training approach and data efficiency. Med Image Anal 2024;98:103298.ArticlePubMed
- 41. Wang D, Wang X, Wang L, et al. A real-world dataset and benchmark for foundation model adaptation in medical image classification. Sci Data 2023;10:574.ArticlePubMedPMCPDF
- 42. Jong MR, Boers TGW, Fockens KN, et al. GastroNet-5M: a multicenter dataset for developing foundation models in gastrointestinal endoscopy. Gastroenterology 2026;170:174–187.ArticlePubMed
- 43. Oh DJ, Hwang Y, Lim YJ. A current and newly proposed artificial intelligence algorithm for reading small bowel capsule endoscopy. Diagnostics (Basel) 2021;11:1183.ArticlePubMedPMC
- 44. Giordano A, Romero-Mascarell C, González-Suárez B, et al. Integration of artificial intelligence-enhanced capsule endoscopy in clinical practice: a review of market-available tools for clinical practice. Dig Dis Sci 2025;70:2966–2976.ArticlePubMedPMCPDF
Citations
Citations to this article as recorded by
