user Admin_Adham
27th Mar, 2026 12:00 AM
Test

Can a Machine Learning Model Predict Liver Cancer?

A machine learning model can accurately predict an individual’s risk of developing hepatocellular carcinoma (HCC) using routine clinical data, according to a new study.

The findings point to a potential way to better identify patients for liver cancer screening, particularly those who fall outside current guidelines.

“Current surveillance approaches are largely based on cirrhosis, but this misses HCC cases [because] chronic liver disease — and especially cirrhosis — is often underdiagnosed,” said senior author Carolin V. Schneider, MD, of University Hospital RWTH Aachen in Aachen, Germany.

Beyond that, the percentage of HCC cases diagnosed in patients without cirrhosis has been steadily growing, highlighting a need for efficient, low-cost strategies to identify more people at increased risk.

Schneider told Medscape Medical News that her team’s model “introduces a prescreening approach” — based on patient demographics, lifestyle factors, diagnoses, and routine blood work — that may identify patients who need further evaluation.

SUGGESTED FOR YOU

For their study, published in Cancer Discovery, the researchers developed and tested machine learning models using data from more than 500,000 participants in the UK Biobank, including 538 diagnosed with HCC.

Of note, 69% of the HCC cases occurred in individuals without a prior diagnosis of cirrhosis, viral hepatitis, or other chronic liver disease.

The models were trained using multiple types of clinical data, including demographics, electronic health records (EHRs), blood tests, genomics, and metabolomics. Performance was internally validated and then tested externally in the All of Us Research Program, which included more than 400,000 US participants and 445 HCC cases.

Schneider’s team found that the best-performing model incorporated demographics, EHRs, and blood tests, achieving an area under the receiver operating characteristic curve of 0.88. Adding genomic or metabolomic data did not meaningfully improve predictive accuracy.

Currently, there are several noninvasive risk scores, like the fibrosis-4 (FIB-4) index, that can be used as surrogates for liver fibrosis and predictors of liver-related mortality. But those tools were developed for high-risk patients and perform poorly in the general population.

Schneider’s team found that their model outperformed FIB-4 and other existing risk scores, even when they used the simplest version, which incorporated just 15 routinely collected clinical features.

“Our study shows that interpretable machine learning algorithms can accurately stratify individual risk of developing HCC on a population scale,” the investigators wrote.

However, prospective validation is still needed.

“We have therefore made the score and full pipeline openly available, with the explicit aim of enabling independent testing and external validation across many health systems,” Schneider said.

Ultimately, though, any real-world implementation may depend less on model performance and more on data integration.

“Routine care data are heterogeneous, including differences in units, coding practices, and missingness,” Schneider pointed out.

In addition, she said, baseline risk differs by population, which means that calibration and thresholding cannot be assumed to “transport unchanged.”

Still, Schneider suggested that these obstacles can be overcome with standardized data and local recalibration of the model.

“This could pave the way for personalized, molecular prevention and risk assessment,” the investigators wrote, “with highest implications for earlier intervention and potentially curative treatment.”

The study was supported by multiple public and institutional funding sources, including the German Cancer Aid and the German Federal Ministry of Research, Technology and Space. Two co-authors disclosed having financial relationships with Johnson & Johnson, Abbott Laboratories, CSL Behring, and others.


Share This Article

Comments

Leave a comment