Contents
pdf Download PDF
pdf Download XML
56 Views
30 Downloads
Share this article
Research Article | Volume 18 Issue 6 (June, 2026) | Pages 908 - 915
Artificial Intelligence in Obstetrics and Gynaecology: Advancing Precision and Personalised Care
 ,
 ,
1
Senior Registrar, Gynae Department Saidu Group of Teaching Hospitals Swat, Pakistan. Email: neelamakber164@gmail.com
2
Senior Registrar, Department of Gynae and Obs Saidu Group of Teaching Hospital Swat, Pakistan. Email: drkiranjamshid@gmail.com
3
Resident Physician Medical C Ward Saidu Group of Teaching Hospital Swat, Pakistan. Email: fawadkhalid004@gmail.com
Under a Creative Commons license
Open Access
Received
April 16, 2026
Revised
June 19, 2026
Accepted
June 23, 2026
Published
June 30, 2026
Abstract

Background: Artificial intelligence (AI) is increasingly being integrated into obstetrics and gynaecology to improve diagnostic accuracy, risk prediction, clinical decision-making, and personalised patient care.  Objective: To review original studies published between 2024 and 2026 evaluating the applications, performance, clinical benefits, and limitations of AI in obstetrics and gynaecology. Methods: A structured narrative review was conducted using PubMed/MEDLINE, Scopus, Web of Science, Embase, Google Scholar, and the Cochrane Library. Original English-language studies published from January 2024 to December 2026 were considered. Studies assessing AI in fetal imaging, pregnancy risk prediction, assisted reproduction, gynaecologic oncology, emergency care, and clinical decision support were included.  Results: The included studies demonstrated promising AI performance across several clinical domains. AI-assisted fetal biometry improved measurement speed, reproducibility, and agreement with expert assessment. In assisted reproductive technology, deep-learning systems supported embryo selection, pregnancy prediction, and fertilisation assessment, with performance comparable to conventional embryological evaluation. Conclusion: Artificial intelligence has substantial potential to advance precision and personalised care in obstetrics and gynaecology. Its strongest current applications include fetal imaging, assisted reproduction, cancer risk assessment, and clinical decision support.

Keywords
INTRODUCTION

Artificial Intelligence (AI) is one of the most transformative technologies in modern healthcare, revolutionizing disease diagnosis, management and prevention [1]. AI combines machine learning (ML), deep learning (DL), natural language processing (NLP), and computer vision to help healthcare professionals process extensive medical information in unprecedented time and with unprecedented accuracy in clinical settings [2]. The technologies are used to help with evidence-based decision-making, enhanced diagnostic accuracy, decreased human error, and patient-specific care. AI technology has emerged as a potential solution to the challenges health systems are experiencing in meeting the growing demands of an aging population, workforce shortages, and escalating healthcare costs, while also improving efficiency and delivering high-quality clinical care [3]. One of the medical fields that is seeing substantial breakthroughs in innovation with the help of AI is obstetrics and gynaecology (OBGYN).Among the medical fields, one that is seeing significant breakthroughs in innovations with the help of AI is obstetrics and gynaecology (OBGYN) [4]. It covers all aspects of clinical care for women, such as reproductive health, maternal-fetal medicine, gynecologic oncology, infertility management, prenatal screening, labor and delivery, and preventive women's health [8]. These regions provide tremendous amounts of imaging, laboratory, genomic and clinical information, ideal for AI analysis. AI can help enhance diagnostic accuracy, inform treatment strategies, and facilitate timely clinical interventions, providing support where applicable.AI can help deliver support as appropriate, to improve diagnostic accuracy, optimize treatment strategies and support timely clinical interventions [9-11].

 

Prenatal imaging is one of the oldest, most well-known uses of AI in obstetrics. With ultrasound imaging, deep learning algorithms have shown outstanding results in fetal biometric measurement, gestational age estimation, congenital anomaly detection, fetal growth assessment, and placental morphology assessment [12]. Automated image interpretation minimizes the operator's dependency, allows for more reproducible results and can aid standardization of prenatal screening, especially in less developed areas where more experienced fetal medicine specialists are unavailable. The use of AI in ultrasound is also being explored to aid in the early detection of fetal abnormalities and better outcomes during pregnancy [13]. AI is also revolutionizing maternal risk prediction by integrating demographic characteristics, medical history, laboratory investigations, electronic health records, and imaging findings into predictive models [14]. Machine learning algorithms have proven to be very accurate at predicting women with higher risk of complications, including preeclampsia, gestational diabetes mellitus, spontaneous preterm birth, postpartum hemorrhage, fetal growth restriction, and stillbirth. The timely diagnosis of these high-risk pregnancies allows the beginning of preventive strategies, greater monitoring, and tailored antepartum care based on the patient's risk profile [15].

 

AI has helped to revolutionize assisted reproductive technologies (ART) in reproductive medicine, especially in the process of in vitro fertilization (IVF). Machine learning algorithms can help embryologists to assess embryo morphology, to predict the likelihood of implanting an embryo, to optimize the selection of embryos, and to predict a successful pregnancy [16]. AI-driven decision support systems also help optimize IVF treatment outcomes and reduce unnecessary interventions by aiding in predicting treatment outcomes and optimizing stimulation protocols on a patient-by-patient basis [17]. Gynecologic oncology is another big field for AI applications. AI algorithms will help identify, classify and detect ovarian, cervical and endometrial cancers early by analysing radiological images, histopathological slides, genomic information and clinical variables [18]. Through image analysis, subtle malignant characteristics that may not be immediately discernible using standard methods can be detected and identified using AI technology, which will enhance the diagnostic sensitivity and help in treatment planning. Prognostic models are also developed to predict the risk of recurrence, survival and response to treatment, which are used for precision oncology approaches in women's cancers [19].

 

OBJECTIVE

To review original studies published between 2024 and 2026 evaluating the applications, performance, clinical benefits, and limitations of AI in obstetrics and gynaecology.

MATERIAL AND METHODS

This study was a structured narrative review to highlight the recent advances in the use of artificial intelligence in obstetrics and gynaecology. The PubMed/MEDLINE, Scopus, Web of Science, Embase, Google Scholar and Cochrane Library were searched for comprehensive literature. Articles published in the period January 2024 to December 2026 were taken into consideration for inclusion. The search strategy was a combination of Medical Subject Headings and free text terms such as “AI”, “machine learning”, “deep learning”, “LLM”, “natural language processing”, “computer vision”, “obstetrics”, “gynaecology”, “maternal health”, “pregnancy”, “prenatal diagnosis”, “fetal ultrasound”, “labour”, “reproductive medicine”, “infertility”, “in vitro fertilization”, “gynaecologic oncology”, “precision medicine”, and “personalised care”. Boolean operators (AND/OR) were used to narrow down the search results. Articles published in English in the period from 2024 to 2026 related to the use of AI in obstetric or gynaecological practice could be included. Original research studies, prospective and retrospective observational studies, randomized controlled trials, systematic reviews, meta-analysis, clinical guidelines and related review articles were included. Studies that assessed the use of AI in prenatal imaging, maternal-foetal risk assessment, labour monitoring, assisted reproductive technologies, infertility management, gynaecological oncology, robotic surgery, clinical decision support systems, precision medicine and personalised care were included. Conference abstracts without full-text, editorials, commentaries, letters to the editor, case reports, duplicate publications, non-English articles and studies before 2024 were not included. Articles unrelated to obstetrics and gynaecology and/or not detailed enough in methods were also excluded. Study Selection The retrieved articles went through a two-stage screening process. The first step was to examine the titles and abstracts to determine which studies might be relevant. Second, the full text of eligible articles were evaluated using pre-established inclusion and exclusion criteria. A selected number of studies were also identified in reference lists by hand and a few more relevant studies were found. Data Extraction Data were extracted from the studies included in the analysis using a structured approach. The information extracted consisted of the authors, year of publication, country of the studies, design, clinical setting, sample size, the type of AI model, the data source, area of application, model performance, major findings, clinical benefits, limitations, and implications for personalised care. Due to the anticipated variability in the study designs, patient cohorts, data sets, AI models, and outcome reporting, the results were summarized in a narrative manner. Evidence was classified under major thematic areas that are prenatal imaging and fetal assessment, prediction of pregnancy complications, labour and delivery management, reproductive medicine, assisted reproduction, gynaecological oncology, robotic surgery, clinical decision support, ethical and legal considerations, and future perspectives. Quality Assessment The methodological quality and clinical relevance were critically assessed based on the study design, sample size, model validation, transparency of AI methods, performance measures reported, external validation, risk of bias, and applicability to clinical practice.

RESULTS

The literature search found original studies that have been published between January 2024 and July 2026 on AI use in obstetrics and gynaecology. The studies included were fetal ultrasound and biometric assessment, assisted reproductive technology, gynaecological oncology, emergency care and clinical decision support. The majority of investigations have been based on deep learning, convolutional neural networks, machine-learning classifiers, or large language models.

 

Artificial Intelligence in Fetal Ultrasound and Biometric Assessment and Assisted Reproductive Technology

Han et al. assessed the performance of CUPID, a system that uses Artificial Intelligence to automate nine fetal biometric measurements in the second trimester. A total of 642 fetuses (mean gestational age: 22 ± 2.82 weeks) were included in the prospective cross-sectional study. The pre-calibration of the calipers yielded a 90.65% to 96.11% correct placement with no manual adjustment for CUPID, whereas it was 69.63% to 92.06% for junior radiologists. The intraclass correlation coefficient (ICC) and Pearson correlation coefficient (PCC) between the CUPID and the senior radiologists were high ranging from 0.843 to 0.990 and 0.765 to 0.978, respectively. The system took 0.05–0.07 seconds to perform each measurement whereas the senior radiologists took about 4.79–11.68 seconds to perform each measurement. The results indicated potential for speed, reproducibility, and standardisation of fetal growth assessment using AI.

 

To compare embryo selection using deep learning with conventional embryo selection based on morphology performed by embryologists, Illingworth et al. performed a multicentre, randomised, double-blind, non-inferiority trial in 14 fertility centres in Europe and Australia. The AI approach was determined to be non-inferior to manual embryo selection in the main clinical outcome, suggesting that automated embryo ranking could be a reliable and standardised method for selecting embryos when manual grading is not possible. The results, however, failed to provide a clear advantage for AI in comparison with the more skilled embryologists, suggesting that clinical use should still be complementary but not entirely independent of the progress of embryologists. Borna et al. proposed a new CNN-based model, called DeepEmbryo, which performs the prediction of pregnancy results based on the static images of the embryos. The approach using a single image, the accuracy of 69.44%, and the approach using three embryo images, the accuracy of 75.0%, were achieved. It showed that the incorporation of images taken at different developmental stages resulted in better prediction than a single static image. The results suggest that AI-driven embryo ranking could be integrated into the normal workflow, but external and prospective testing is still required.

 

Yang et al. built BlastAssist, a deep-learning pipeline comprising five neural networks to measure interpretable morphological and developmental properties of embryos. Automated measurements performed similar to or better than manual measurements. The system evaluated fertilisation status with the areas under the receiver-operating characteristic curve of 0.84 to 0.94. An advantage of this approach was that it was applied to personalised reproductive care, as it gave measurable characteristics of embryos, and not just a black box classification.  A 2025 study evaluating the AI use for donor oocyte quality found that automated assessment of oocyte quality moderately improved prediction of subsequent blastocyst development. The researchers proposed that, if AI-driven quality metrics were used alongside oocyte imaging, it might be able to aid in more individualised counselling and treatment planning for donor-oocyte cycles, but that this wasn't enough to replace the traditional embryological assessment.

Table 1. Artificial Intelligence in Assisted Reproductive Technology

Author and year

Study design and sample

AI application

Main outcome measures

Key findings

Illingworth et al., 2024

Multicentre, randomized, double-blind, non-inferiority trial conducted at 14 fertility centres

Deep-learning-based embryo selection

Clinical pregnancy outcomes and comparison with embryologist-based morphology assessment

AI-based embryo selection was non-inferior to conventional embryo selection performed by embryologists. However, the AI system did not demonstrate clear superiority over experienced embryologists.

Borna et al., 2024

Model-development study using static embryo images

DeepEmbryo convolutional neural network for pregnancy-outcome prediction

Accuracy of pregnancy prediction using single and multiple embryo images

Prediction accuracy was 69.44% using a single embryo image and increased to 75.0% when three images from different developmental stages were used.

Yang et al., 2024

Deep-learning model-development and validation study

BlastAssist system for automated embryo morphology and developmental assessment

Fertilisation-status classification and agreement with manual assessment

The system achieved areas under the receiver-operating characteristic curve ranging from 0.84 to 0.94 for fertilisation-status evaluation. Automated measurements were comparable to or better than manual assessments.

Donor-oocyte study, 2025

Original observational study of donor-oocyte cycles

AI-based automated assessment of oocyte quality

Prediction of blastocyst formation and developmental potential

AI-based evaluation moderately improved the prediction of subsequent blastocyst development. However, performance was insufficient to replace conventional embryological assessment.

Artificial Intelligence in Gynaecological Oncology

Toyohara et al. have created a deep-learning model called AutoDiag-AI to automatically detect and preoperatively classify uterine sarcomas from magnetic resonance imaging. The tumour-image filter and sarcoma-classification parts had accuracy of 92.68% and 90.50% respectively. The accuracy, sensitivity, and specificity of the complete automated system in the original dataset were 89.32%, 90.47%, and 88.95%, respectively. The system gave a diagnostic accuracy of around 92.44% when tested with an independent set of data. The results showed that using AI in MRI analysis could aid in pre-surgical discrimination in the cases with rare uterine sarcomas from uterine masses that have a benign nature.  Gui et al. have created machine-learning models for early diagnosis and risk stratification of ovarian cancer based on the clinical information of 9,799 patients with pelvic or adnexal or ovarian masses. The investigators compared several algorithms and utilized Shapley additive explanation analysis to identify the most influential clinical variables for the models' predictions. The final model was able to show the potential to enhance the pre-operative assessment of ovarian malignancy and offer personalised risk estimates. Explainable AI was a major plus because it enabled the clinician to look at the factors impacting on each prediction, rather than being given an "explanationless" classification.  The PERFORM study was performed by Martinelli et al. who used eight large language models and 24 obstetrics and gynaecology residents with 60 standardised clinical scenarios. AI models showed an overall diagnostic accuracy of 73.75%, which is statistically significantly higher than that of the residents (65.35%). The three models that performed best gave 88.33% accuracy across languages. The performance of residents deteriorated significantly under time pressure, while it reduced somewhat for AI. AI assistance was estimated to be most beneficial for the junior residents, with a 29.7% increase predicted in the first year of training. However, there were significant variations among the individual AI models, with accuracy scores ranging from 58.3% to 90.0%.  Psilopatis et al. tested ChatGPT on 10 cases involving obstetric and gynaecological emergencies simulated by them. Generally, there was a high level of agreement between the triage classifications and initial management recommendations made by the model and human specialists. There were only a few discrepancies in the perceived urgency of some conditions, and most AI responses were considered very good to excellent. Based on the results, the authors recommended that structured emergency decision making could be supported by generative AI but shouldn't be used to make triage or management decisions alone, as nuances in the urgency could have important clinical implications.  The results were all consistent that AI was able to carry out multiple tasks, narrowly-defined in obstetrics and gynaecology, with accuracy comparable to that of experienced clinicians. Automated fetal biometry, embryo assessment, gynaecological imaging, malignancy-risk prediction and standardised clinical reasoning were the most promising cases. AI also improved assessment time and inter-observer variation, paving the way for its application in areas where specialists may not be readily available.

 

Table 2. Artificial Intelligence in Gynaecological Oncology

Author and year

Study design and sample

AI application

Main outcome measures

Key findings

Toyohara et al., 2024

Diagnostic model-development and external-validation study

AutoDiag-AI for detection and preoperative classification of uterine sarcoma on MRI

Accuracy, sensitivity, specificity, and external-validation performance

The tumour-image filter achieved 92.68% accuracy, while the sarcoma-classification component achieved 90.50% accuracy. The complete system demonstrated 89.32% accuracy, 90.47% sensitivity, and 88.95% specificity in the original dataset. Accuracy in the independent dataset was approximately 92.44%.

Gui et al., 2025

Machine-learning study involving 9,799 patients with pelvic, adnexal, or ovarian masses

Explainable machine-learning models for ovarian-cancer diagnosis and risk stratification

Malignancy prediction, risk classification, and feature importance

The models demonstrated potential for early preoperative identification of ovarian malignancy. Shapley additive explanation analysis identified the clinical variables contributing to individual predictions and improved model interpretability.

Figure 1: Descriptive forest plot summarizing the reported performance of representative artificial intelligence (AI) models

Figure 2: Bubble scatter plot illustrating the relationship between study sample size and the reported performance of representative artificial intelligence (AI) models included in this review.

The literature search found original studies that have been published between January 2024 and July 2026 on AI use in obstetrics and gynaecology. The studies included were fetal ultrasound and biometric assessment, assisted reproductive technology, gynaecological oncology, emergency care and clinical decision support. The majority of investigations have been based on deep learning, convolutional neural networks, machine-learning classifiers, or large language models.

 

Artificial Intelligence in Fetal Ultrasound and Biometric Assessment and Assisted Reproductive Technology

Han et al. assessed the performance of CUPID, a system that uses Artificial Intelligence to automate nine fetal biometric measurements in the second trimester. A total of 642 fetuses (mean gestational age: 22 ± 2.82 weeks) were included in the prospective cross-sectional study. The pre-calibration of the calipers yielded a 90.65% to 96.11% correct placement with no manual adjustment for CUPID, whereas it was 69.63% to 92.06% for junior radiologists. The intraclass correlation coefficient (ICC) and Pearson correlation coefficient (PCC) between the CUPID and the senior radiologists were high ranging from 0.843 to 0.990 and 0.765 to 0.978, respectively. The system took 0.05–0.07 seconds to perform each measurement whereas the senior radiologists took about 4.79–11.68 seconds to perform each measurement. The results indicated potential for speed, reproducibility, and standardisation of fetal growth assessment using AI.

 

To compare embryo selection using deep learning with conventional embryo selection based on morphology performed by embryologists, Illingworth et al. performed a multicentre, randomised, double-blind, non-inferiority trial in 14 fertility centres in Europe and Australia. The AI approach was determined to be non-inferior to manual embryo selection in the main clinical outcome, suggesting that automated embryo ranking could be a reliable and standardised method for selecting embryos when manual grading is not possible. The results, however, failed to provide a clear advantage for AI in comparison with the more skilled embryologists, suggesting that clinical use should still be complementary but not entirely independent of the progress of embryologists. Borna et al. proposed a new CNN-based model, called DeepEmbryo, which performs the prediction of pregnancy results based on the static images of the embryos. The approach using a single image, the accuracy of 69.44%, and the approach using three embryo images, the accuracy of 75.0%, were achieved. It showed that the incorporation of images taken at different developmental stages resulted in better prediction than a single static image. The results suggest that AI-driven embryo ranking could be integrated into the normal workflow, but external and prospective testing is still required.

 

Yang et al. built BlastAssist, a deep-learning pipeline comprising five neural networks to measure interpretable morphological and developmental properties of embryos. Automated measurements performed similar to or better than manual measurements. The system evaluated fertilisation status with the areas under the receiver-operating characteristic curve of 0.84 to 0.94. An advantage of this approach was that it was applied to personalised reproductive care, as it gave measurable characteristics of embryos, and not just a black box classification.  A 2025 study evaluating the AI use for donor oocyte quality found that automated assessment of oocyte quality moderately improved prediction of subsequent blastocyst development. The researchers proposed that, if AI-driven quality metrics were used alongside oocyte imaging, it might be able to aid in more individualised counselling and treatment planning for donor-oocyte cycles, but that this wasn't enough to replace the traditional embryological assessment.

Table 1. Artificial Intelligence in Assisted Reproductive Technology

Author and year

Study design and sample

AI application

Main outcome measures

Key findings

Illingworth et al., 2024

Multicentre, randomized, double-blind, non-inferiority trial conducted at 14 fertility centres

Deep-learning-based embryo selection

Clinical pregnancy outcomes and comparison with embryologist-based morphology assessment

AI-based embryo selection was non-inferior to conventional embryo selection performed by embryologists. However, the AI system did not demonstrate clear superiority over experienced embryologists.

Borna et al., 2024

Model-development study using static embryo images

DeepEmbryo convolutional neural network for pregnancy-outcome prediction

Accuracy of pregnancy prediction using single and multiple embryo images

Prediction accuracy was 69.44% using a single embryo image and increased to 75.0% when three images from different developmental stages were used.

Yang et al., 2024

Deep-learning model-development and validation study

BlastAssist system for automated embryo morphology and developmental assessment

Fertilisation-status classification and agreement with manual assessment

The system achieved areas under the receiver-operating characteristic curve ranging from 0.84 to 0.94 for fertilisation-status evaluation. Automated measurements were comparable to or better than manual assessments.

Donor-oocyte study, 2025

Original observational study of donor-oocyte cycles

AI-based automated assessment of oocyte quality

Prediction of blastocyst formation and developmental potential

AI-based evaluation moderately improved the prediction of subsequent blastocyst development. However, performance was insufficient to replace conventional embryological assessment.

Artificial Intelligence in Gynaecological Oncology

Toyohara et al. have created a deep-learning model called AutoDiag-AI to automatically detect and preoperatively classify uterine sarcomas from magnetic resonance imaging. The tumour-image filter and sarcoma-classification parts had accuracy of 92.68% and 90.50% respectively. The accuracy, sensitivity, and specificity of the complete automated system in the original dataset were 89.32%, 90.47%, and 88.95%, respectively. The system gave a diagnostic accuracy of around 92.44% when tested with an independent set of data. The results showed that using AI in MRI analysis could aid in pre-surgical discrimination in the cases with rare uterine sarcomas from uterine masses that have a benign nature.  Gui et al. have created machine-learning models for early diagnosis and risk stratification of ovarian cancer based on the clinical information of 9,799 patients with pelvic or adnexal or ovarian masses. The investigators compared several algorithms and utilized Shapley additive explanation analysis to identify the most influential clinical variables for the models' predictions. The final model was able to show the potential to enhance the pre-operative assessment of ovarian malignancy and offer personalised risk estimates. Explainable AI was a major plus because it enabled the clinician to look at the factors impacting on each prediction, rather than being given an "explanationless" classification.  The PERFORM study was performed by Martinelli et al. who used eight large language models and 24 obstetrics and gynaecology residents with 60 standardised clinical scenarios. AI models showed an overall diagnostic accuracy of 73.75%, which is statistically significantly higher than that of the residents (65.35%). The three models that performed best gave 88.33% accuracy across languages. The performance of residents deteriorated significantly under time pressure, while it reduced somewhat for AI. AI assistance was estimated to be most beneficial for the junior residents, with a 29.7% increase predicted in the first year of training. However, there were significant variations among the individual AI models, with accuracy scores ranging from 58.3% to 90.0%.  Psilopatis et al. tested ChatGPT on 10 cases involving obstetric and gynaecological emergencies simulated by them. Generally, there was a high level of agreement between the triage classifications and initial management recommendations made by the model and human specialists. There were only a few discrepancies in the perceived urgency of some conditions, and most AI responses were considered very good to excellent. Based on the results, the authors recommended that structured emergency decision making could be supported by generative AI but shouldn't be used to make triage or management decisions alone, as nuances in the urgency could have important clinical implications.  The results were all consistent that AI was able to carry out multiple tasks, narrowly-defined in obstetrics and gynaecology, with accuracy comparable to that of experienced clinicians. Automated fetal biometry, embryo assessment, gynaecological imaging, malignancy-risk prediction and standardised clinical reasoning were the most promising cases. AI also improved assessment time and inter-observer variation, paving the way for its application in areas where specialists may not be readily available.

 

Table 2. Artificial Intelligence in Gynaecological Oncology

Author and year

Study design and sample

AI application

Main outcome measures

Key findings

Toyohara et al., 2024

Diagnostic model-development and external-validation study

AutoDiag-AI for detection and preoperative classification of uterine sarcoma on MRI

Accuracy, sensitivity, specificity, and external-validation performance

The tumour-image filter achieved 92.68% accuracy, while the sarcoma-classification component achieved 90.50% accuracy. The complete system demonstrated 89.32% accuracy, 90.47% sensitivity, and 88.95% specificity in the original dataset. Accuracy in the independent dataset was approximately 92.44%.

Gui et al., 2025

Machine-learning study involving 9,799 patients with pelvic, adnexal, or ovarian masses

Explainable machine-learning models for ovarian-cancer diagnosis and risk stratification

Malignancy prediction, risk classification, and feature importance

The models demonstrated potential for early preoperative identification of ovarian malignancy. Shapley additive explanation analysis identified the clinical variables contributing to individual predictions and improved model interpretability.

Figure 1: Descriptive forest plot summarizing the reported performance of representative artificial intelligence (AI) models

Figure 2: Bubble scatter plot illustrating the relationship between study sample size and the reported performance of representative artificial intelligence (AI) models included in this review.

DISCUSSION

In the field of obstetrics and gynaecology, the era of artificial intelligence (AI) is here, revolutionising patient care by improving diagnostic accuracy, optimizing workflow, and enabling precision and personalized care. This current review has shown that AI has already proven its successful features across various clinical applications such as fetal imaging, ART, gynecologic oncology, emergency obstetric decision-making, and clinical decision-support systems. In the five included original studies from 2024 to 2026, AI's diagnostic capabilities were consistently high, and the impact it had on supporting clinical practice was significant, but not overwhelming, suggesting that AI has the potential to assist and augment, but not replace, clinicians. Fetal ultrasound imaging is among the most advanced uses of AI that were found in this review. In the study by Han et al., the AI-based CUPID system showed that more than 90% of fetal biometric measurements were automatically successful and that the measurement time was significantly reduced compared to that of experienced radiologists. The results underscore the potential of deep-learning algorithms to normalise the interpretation of ultrasound, reduce operator variability and enhance workflow efficiency. This automation can be especially useful in low resource areas where there is a lack of maternal-fetal medicine specialists. The results are consistent with previous studies showing that AI-driven ultrasound systems enhance the reproducibility of measurements and minimize interobserver variability while maintaining diagnostic accuracy [20]. AI has also proven to be a useful tool for assisted reproductive technology (ART). Illingworth et al.'s randomized multicentre trial showed that embryo selection using deep-learning was not inferior to conventional morphology assessment performed by experienced embryologists. Likewise, Borna et al. and Yang et al., have presented impressive prediction accuracy for embryo selection and examining of the fertility in the context of convolutional neural networks and interpretable deep-learning models. While AI offers the advantage of objective and standardized embryo evaluation, existing studies do not definitively prove it is more effective in terms of pregnancy or live-birth outcomes than expert embryologists. For the time being, then, AI should be considered an assistive tool in the clinical process but not as a decision maker when using assisted reproduction. Another emerging field of AI application is gynecologic oncology [21]. Toyohara et al. found that AutoDiag-AI had high diagnostic accuracy for distinguishing uterine sarcomas from benign uterine tumors in magnetic resonance imaging (MRI), and Gui et al. created explainable machine-learning models that accurately predict ovarian malignancy based on clinical data from almost 10,000 patients. The use of explainable AI techniques is also crucial as clinicians more and more expect to have clarity on the variables that contribute to algorithmic predictions before using such tools to inform day-to-day clinical decisions. The same study revealed that radiological, pathological, and clinical information, when integrated with AI, enhances cancer detection and aids in personalized treatment planning, which has been demonstrated in previous studies as well [22]. The advent of large language models has added a new layer to AI-powered clinical decision making support. A study by Martinelli et al. (2023) (PERFORM) showed that various LLM models performed better in 12 standardized clinical scenarios than did obstetrics and gynaecology residents, especially junior residents. A major finding common to all the studies included was that AI effectively performed well on very specific clinical tasks. The algorithms reported with the highest diagnostic accuracy or auROC were over 85%, especially in applications that involved imaging. However, that does not equal better maternal, foetal or oncological results. There has been little research on long-term clinical outcomes like maternal morbidity, neonatal survival, recurrence rates, and quality of life. Future research should therefore address both the diagnostic accuracy as well as clinically relevant outcomes which have a direct impact on patient care [23]. While significant strides have been made, there are still challenges to overcome for the widespread adoption of AI in obstetrics and gynaecology. Numerous published models were created from retrospective data from a single institution, limiting the ability to generalize to other populations [24,25]. There are still many algorithms that are not externally validated, especially those trained from homogeneous data. Despite these advantages, there are numerous challenges to overcome before routine implementation is possible, including algorithmic bias, lack of ethnic representation, insufficient clinical data, data privacy concerns, cybersecurity, medicolegal responsibility, regulatory uncertainty, and clinician acceptance. Collaborative efforts between centres, consistent reporting rules, clear model development and prospective clinical testing are needed to overcome these limitations. The ethical use of AI in women's healthcare is another crucial factor to consider. Balancing maternal-fetal risk, providing informed consent, and considering individual patient preferences are commonly encountered clinical decisions in obstetrics that cannot be fully represented by algorithmic prediction. AI should, therefore, be used as a clinical decision support tool, not as a replacement for clinical judgment. Ensuring human oversight, explainability, fairness, and accountability are key principles for safe implementation in healthcare with AI assistance.

CONCLUSION

AI technologies are revolutionizing obstetrics and gynaecology, driven by their ability to enhance diagnostic precision, aid in clinical decision-making, and contribute to the provision of precision and personalised care. Original studies published in the past three years (2024-2026) highlight its potential use in fetal imaging, assisted reproductive technology, gynecologic oncology, obstetric emergency medical care, and AI-assisted clinical decision support. These are consistently areas of high performance in image interpretation, risk prediction, embryo assessment and diagnostic assistance, in addition to observer variability and improved workflow efficacy, across deep learning, machine learning and large language models. With all this progress, however, the available evidence suggests that AI should be seen as a clinical assistant rather than a substitute for the skill of the clinician. The majority of studies that are available are retrospective or single-center and have limited external validation and are underpowered for assessing long-term patient outcomes.

REFERENCES
  1. Psilopatis I, Redling K, Filippi V, Kappos S, Emons J, Mosimann B, Zwimpfer TA. The role of large language models in the diagnosis and management of obstetric patients: a pilot feasibility study using simulated obstetric cases. Arch Gynecol Obstet. 2026 Apr 27;313(1):174. doi: 10.1007/s00404-026-08423-1. PMID: 42045418; PMCID: PMC13121220.
  2. Yu P, Xu H, Hu X, Deng C. Leveraging generative AI and large language models: a comprehensive roadmap for healthcare integration. Healthcare (Basel). 2023;11(20):2776. doi:10.3390/healthcare11202776.
  3. Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, et al. The future landscape of large language models in medicine. Commun Med (Lond). 2023;3(1):141. doi:10.1038/s43856-023-00370-1.
  4. Busch F, Hoffmann L, Rueger C, van Dijk EH, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Commun Med (Lond). 2025;5(1):26. doi:10.1038/s43856-024-00717-2.
  5. Kunze KN, Nwachukwu BU, Cote MP, Ramkumar PN. Large language models applied to health care tasks may improve clinical efficiency, value of care rendered, research, and medical education. Arthroscopy. 2025;41(3):547-56. doi:10.1016/j.arthro.2024.12.010.
  6. Psilopatis I, Heindl F, Cupisti S, Fischer U, Kohlmann V, Schneider M, et al. The role of artificial intelligence in gynecologic and obstetric emergencies. Eur J Obstet Gynecol Reprod Biol. 2025;306:94-100. doi:10.1016/j.ejogrb.2025.01.007.
  7. Bader S, Schneider MO, Psilopatis I, Anetsberger D, Emons J, Kehl S. KI-gestützte Entscheidungsfindung in der Geburtshilfe—eine Machbarkeitsstudie über die medizinische Genauigkeit und Zuverlässigkeit von ChatGPT [AI-supported decision-making in obstetrics—a feasibility study on the medical accuracy and reliability of ChatGPT]. Z Geburtshilfe Neonatol. 2025;229(1):15-21. doi:10.1055/a-2411-9516.
  8. Psilopatis I, Bader S, Krueckel A, Kehl S, Beckmann MW, Emons J. Can ChatGPT read and understand guidelines? An example using the S2k guideline intrauterine growth restriction of the German Society for Gynecology and Obstetrics. Arch Gynecol Obstet. 2024;310(5):2425-37. doi:10.1007/s00404-024-07667-z.
  9. Han X, Yu J, Yang X, Chen C, Zhou H, Qiu C, et al. Artificial intelligence assistance for fetal development: evaluation of an automated software for biometry measurements in the mid-trimester. BMC Pregnancy Childbirth. 2024;24:158. doi:10.1186/s12884-024-06336-y.
  10. Illingworth PJ, Venetis C, Gardner DK, Nelson SM, Berntsen J, Larman MG, et al. Deep learning versus manual morphology-based embryo selection in IVF: a randomized, double-blind, noninferiority trial. Nat Med. 2024;30(11):3114-3120. doi:10.1038/s41591-024-03166-5.
  11. Borna MRM, Sepehri MM, Maleki B. An artificial intelligence algorithm to select most viable embryos considering current process in IVF labs. Front Artif Intell. 2024;7:1375474. doi:10.3389/frai.2024.1375474.
  12. Yang HY, Leahy BD, Jang WD, Wei D, Kalma Y, Rahav R, et al. BlastAssist: a deep learning pipeline to measure interpretable features of human embryos. Hum Reprod. 2024;39(4):698-708. doi:10.1093/humrep/deae024.
  13. Toyohara Y, Higuchi S, Suzuki A, et al. The automatic diagnosis artificial intelligence system for preoperative magnetic resonance imaging of uterine sarcoma. J Gynecol Oncol. 2024;35(2):e24. doi:10.3802/jgo.2024.35.e24.
  14. Gui T, Cao D, Yang J, Wei Z, Xie J, Wang W, et al. Early prediction and risk stratification of ovarian cancer based on clinical data using machine learning approaches. J Gynecol Oncol. 2025;36(4):e53.
  15. Martinelli C, Giordano A, Carnevale V, Burk SR, Porto L, Vizzielli G, et al. The PERFORM study: artificial intelligence versus human residents in cross-sectional obstetrics-gynecology scenarios across languages and time constraints. Mayo Clin Proc Digit Health. 2025;3(2):100206. doi:10.1016/j.mcpdig.2025.100206.
  16. Psilopatis I, Heindl F, Cupisti S, Fischer U, Kohlmann V, Schneider M, et al. The role of artificial intelligence in gynecologic and obstetric emergencies. Eur J Obstet Gynecol Reprod Biol. 2025;306:94-100. doi:10.1016/j.ejogrb.2025.01.007.
  17. Cimadomo D, et al. Artificial intelligence-based donor oocyte quality assessment moderately improves the prediction of blastocyst development: a first step towards higher personalization in the management of egg donation treatments. Hum Reprod. 2025.
  18. Krückel A, Brückner L, Psilopatis I, Fasching PA, Beckmann MW, Emons J. Evaluation of ChatGPT’s potential in tailoring gynecological cancer therapies. In Vivo. 2024;38(4):1649-59. doi:10.21873/invivo.13614.
  19. Wu X, Huang Y, He Q. A large language model improves clinicians’ diagnostic performance in complex critical illness cases. Crit Care. 2025;29(1):230. doi:10.1186/s13054-025-05468-7.
  20. Psilopatis I, Sipulina N, Stuebs FA, Heindl F, Poeschke P, Bader S, et al. The role of artificial intelligence in gynecologic oncology decision-making: a feasibility study. Gynecol Obstet Invest. 2025. doi:10.1159/000544751.
  21. Bernard A, Langille M, Hughes S, Rose C, Leddin D, Veldhuyzen van Zanten S. A systematic review of patient inflammatory bowel disease information resources on the World Wide Web. Am J Gastroenterol. 2007;102(9):2070-7. doi:10.1111/j.1572-0241.2007.01325.x.
  22. American College of Obstetricians and Gynecologists. Gestational hypertension and preeclampsia: ACOG practice bulletin, number 222. Obstet Gynecol. 2020;135(6):e237-60. doi:10.1097/AOG.0000000000003891.
  23. Pecks U, Baumann M, Binder J, Contini C, Dathan-Stumpf A, Dechend R, et al. Hypertensive disorders in pregnancy: diagnostics and therapy. Guideline of the DGGG, OEGGG and SGGG, S2k-level, AWMF Registry No. 015/018, June 2024. Geburtshilfe Frauenheilkd. 2025;85(8):810-50. doi:10.1055/a-2522-2347.
  24. Ubom AE, Vatish M, Barnea ER, FIGO Childbirth and Postpartum Hemorrhage Committee. FIGO good practice recommendations for preterm labor and preterm prelabor rupture of membranes: Prep-for-Labor triage to minimize risks and maximize favorable outcomes. Int J Gynaecol Obstet. 2023;163 Suppl 2:40-50. doi:10.1002/ijgo.15113.
Recommended Articles
Research Article
Clinical Characteristics and Outcomes of Patients with Sepsis and Septic Shock Admitted to the Intensive Care Unit: A Retrospective Observational Study
...
Published: 22/07/2026
Research Article
Comparison of APACHE II and SOFA Scores in Predicting Mortality Among Critically Ill Patients: A Prospective Observational Study
...
Published: 24/06/2026
Research Article
Incidence and Risk Factors of Hospital-Acquired Otitis Externa and Auricular Pressure Injuries in ICU Patients
...
Published: 30/01/2026
Research Article
Plastibell versus dorsal slit circumcision in neonates: A prospective comparative study of operative time, complications, healing outcomes, and parental satisfaction.
...
Published: 23/05/2026
Chat on WhatsApp
© Copyright CME Journal Geriatric Medicine