LARGE LANGUAGE MODELS IN HEPATOLOGY: A SYSTEMATIC REVIEW

Thanathip SUENGHATAIPHORN, Narisara TRIBUDDHARAT, Pojsakorn DANPANICHKUL, Narathorn KULTHAMRONGSRI

Hepatology Forum - 2026;7(3):227-235

Department of Internal Medicine, Griffin Hospital, Derby, CT, United States

 

Background and Aim: The rapid advancement of generative artificial intelligence (AI), particularly large language models (LLMs), has opened new frontiers in healthcare, with emerging implications for hepatology. This systematic review synthesizes the current state of research on the application of LLMs in hepatology, focusing on their capabilities in real-world clinical settings, limitations, and future directions. Materials and Methods: Electronic databases, including MEDLINE, EMBASE, and OVID as a search platform, were used to identify eligible studies from inception to January 2025. Eligible studies investigated the clinical utility and performance of LLMs in hepatology, with a clear comparison to a defined ground truth. Key findings were extracted and synthesized narratively. The ROBINS-I tool was used to assess the risk of bias in each study. Results: Twenty-one studies were included in this review. Our analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease. LLMs effectively assisted with radiological image interpretation, provided clinical decision support, and generated patient education materials. However, the accuracy of these models was highly variable, depending on the specific task and the complexity of the clinical scenario. Limitations, such as the generation of inaccurate or misleading information ("hallucinations"), dependence on training data quality, and ethical considerations, were identified across multiple studies. Conclusion: Generative AI demonstrates feasibility across various hepatology applications, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety. Future integration necessitates further research into training methods, data quality, ethical considerations, and real-world validation against standardized benchmarks.