AI Chatbots Struggle with Surgical Accuracy: New Study

The rapid integration of **Artificial Intelligence (AI)** into clinical decision support systems faces a significant hurdle, according to a recent analysis of **ChatGPT-5** performance. Researchers investigating the platform’s diagnostic and prognostic capabilities found that the Large Language Model (LLM) consistently overestimates the success rates of **macular hole surgery**.

A **macular hole** is a small break in the **macula**, the specialized area in the center of the **retina** responsible for sharp, detailed vision. While standard surgical interventions, such as **pars plana vitrectomy**, are highly effective, patient outcomes are inherently influenced by variables like hole size, duration of symptoms, and the patient’s overall ocular health.

The study highlights a critical gap between generalized medical information and precision clinical guidance. When prompted to predict surgical outcomes, the AI model provided overly optimistic success projections compared to established, evidence-based clinical benchmarks. This discrepancy raises urgent concerns regarding the use of generative AI as a primary tool for patient counseling or preoperative planning.

Healthcare experts warn that relying on such models for complex **ophthalmological** procedures can lead to misinformed expectations among patients. Because these models operate on probabilistic text generation rather than a true understanding of **biostatistical** data, they lack the nuance required to evaluate individual surgical risks.

The findings underscore the persistent risks of “hallucinations” or inaccuracies within LLMs, even as they advance through newer iterations. For **ophthalmologists** and surgeons, these results serve as a cautionary tale: AI should not be utilized as a substitute for professional clinical judgment or peer-reviewed literature.

As AI models continue to be trained on massive datasets, the researchers emphasize the need for specialized, medically-validated training loops. Without dedicated **fine-tuning** on high-quality clinical trial data, LLMs may inadvertently propagate misinformation that compromises the surgeon-patient relationship.

Moving forward, the medical community must prioritize the development of “human-in-the-loop” AI architectures. Until these tools achieve a higher threshold of accuracy regarding surgical prognosis, providers should caution patients against using general-purpose chatbots for specific healthcare consultations. Relying on verified clinical data remains the gold standard for navigating delicate retinal procedures.