Evidence map›Paper›PMID 42404572›Full record

ArticleFrontiers in medicine2026

Interpretable machine learning for severity classification of thyroid eye disease using orbital anatomical features.

Ruixin Shi, Leiming Gao, Shengzhi Jiao, Liuzi Wang, Jianing Li, Bei Wang

Abstract read
In one paragraph

Article in Frontiers in medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Ruixin ShiSchool of Nursing, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.
Leiming GaoSchool of Nursing, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.
Shengzhi JiaoSchool of Nursing, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.
Liuzi WangSchool of Nursing, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.
Jianing LiDepartment of Nursing, Affiliated Hospital of Traditional Chinese and Western Medicine, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.
Bei WangDepartment of Nursing, Affiliated Hospital of Traditional Chinese and Western Medicine, Nanjing University of Chinese Medicine, Nanjing, Jiangsu, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Thyroid eye disease (TED) severity assessment using the EUGOGO severity classification is partly subjective and prone to interobserver variability. While MRI-derived anatomical measurements offer objective features, such as ocular protrusion and extraocular muscle thickness, these are underutilized in machine learning (ML) models that often rely on non-interpretable radiomic features. Moreover, the inclusion of longitudinal scans from the same patient may artificially inflate model performance due to unaccounted intra-patient correlations and temporal redundancy. Purpose: To develop an interpretable machine learning framework for objective TED severity stratification (mild, moderate-to-severe, sight-threatening) by quantitatively integrating established orbital anatomical parameters aligned with clinical assessment criteria, and to evaluate how data handling strategies influence model generalizability. Methods: We retrospectively analyzed 443 patients with TED, yielding 1,054 orbital MRI units from all available examinations. Two datasets were constructed: Dataset A included all available scans, while Dataset B retained only the first-visit orbital MRI units per patient (886 orbital units) to reduce temporal and repeat-visit bias. As bilateral orbits from the same patient were included, Dataset B reduces visit-related confounding but retains inherent inter-eye correlations. Six ML models-Logistic Regression (LR), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Random Forest (RF), XGBoost, and LightGBM-were trained and evaluated using cross-validation. Model performance was compared in terms of area under the receiver operating characteristic curve (AUC), F1-score, and recall. Feature importance was assessed using RF, XGBoost, and LightGBM. Results: Under class-imbalance strategies, Random Forest with class weighting achieved the highest AUC (0.811). Random Forest with SMOTE achieved the highest recall (0.669), F1-score (0.648), and specificity (0.815). Performance on Dataset A demonstrated how unaccounted longitudinal scan correlations can inflate metrics, reinforcing the necessity of temporal deduplication and cautious handling of orbital-level correlations. Feature importance analysis consistently ranked ocular protrusion as the top predictor, followed by rectus muscle thicknesses and orbital geometric parameters. Conclusion: Controlling for longitudinal redundancy and intra-patient correlations significantly impacts model evaluation and generalizability. Random Forest with class weighting demonstrated the best discriminative performance in our internal validation on temporally deduplicated first-visit scans. Rather than relying on isolated diagnostic thresholds, the framework integrates measurable anatomical parameters to generate predictions that complement categorical clinical grading. This approach emphasizes standardized quantification and workflow reproducibility, highlighting the need for rigorous data structuring and transparent feature alignment in medical AI.

Indexed as

disease severityfeature importancemachine learningmagnetic resonance imagingorbital anatomythyroid eye disease

Identifiers

PMID42404572
PMCPMC13327989

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.