Evidence map›Paper›PMID 41455821›Full record

ArticleCommunications medicine2025

Compact vision language models enable efficient and interpretable optical coherence tomography through layer-specific multimodal learning.

Tania Haghighi, Sina Gholami, Jared Todd Sokol, Aayush Biswas, Jennifer I Lim, Theodore Leng, Atalie C Thompson, Hamed Tabkhi, Minhaj Nur Alam

Abstract read
In one paragraph

Article in Communications medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

9 authors.

Tania HaghighiDepartment of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA.
Sina GholamiDepartment of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA.
Jared Todd SokolByers Eye Institute at Stanford, Stanford University School of Medicine, Stanford, CA, USA.
Aayush BiswasDepartment of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA.
Jennifer I LimUniversity of Illinois Chicago, Chicago, IL, USA.
Theodore LengByers Eye Institute at Stanford, Stanford University School of Medicine, Stanford, CA, USA.ORCID http://orcid.org/0000-0002-8461-3562
Atalie C ThompsonSurgical Ophthalmology, Atrium Health Wake Forest Baptist, Winston-Salem, NC, USA.
Hamed TabkhiDepartment of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA.
Minhaj Nur AlamDepartment of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA. minhaj.alam@charlotte.edu.ORCID http://orcid.org/0000-0003-3095-2232

Funding

Distributed approaches to train machine learning models in diabetic retinopathyR15EY035804 · NEI · UNIVERSITY OF NORTH CAROLINA CHARLOTTE · PI ALAM, MINHAJ NUR · 2024 to 2024
$460k
Domain-adaptive federated learning to develop machine learning models for predicting incident and progression of geographic atrophyR21EY035271 · NEI · UNIVERSITY OF NORTH CAROLINA CHARLOTTE · PI ALAM, MINHAJ NUR · 2024 to 2024
$379k
NEI NIH HHS R15 EY035804NEI NIH HHS R21 EY035271UNC | University of North Carolina at Charlotte (UNC Charlotte) AI4Health seed garntU.S. Department of Health & Human Services | NIH | National Eye Institute (NEI) R15EY035804U.S. Department of Health & Human Services | NIH | National Eye Institute (NEI) R21EY035271
6 · The paper itself

Abstract

backgroundTranslating the intricate anatomical signatures of retinal disease from optical coherence tomography (OCT) B-scans into clear, accurate clinical narratives demands algorithms that seamlessly fuse visual features with domain expertise.

methodsWe curated a multimodal dataset of 40,000 OCT B-scans from public repositories and private clinical cohorts, each paired with expert-validated summaries spanning six conditions: diabetic macular edema, diabetic retinopathy, geographic atrophy, drusen, choroidal neovascularization, and healthy retina. We introduce LO-VLM, a compact (247M parameter) vision-language model (VLM) that infuses anatomical guidance into both encoder and decoder for free-form summary generation and multiclass disease classification. Benchmarking against state-of-the-art RetinaVLM, LLaVA-Med, and a ViT vision only model demonstrates superior performance.

resultsIn a blinded evaluation by three board certified retina specialists, LO-VLM narratives achieves a mean = 8.5 (standard deviation = 1.15) out of 10, compared to a mean = 5.5 (standard 32 deviation = 1.13) for RetinaVLM (p < 0.0001). In quantitative evaluations, LO-VLM achieves an SBERT similarity of 80.3% and a BERTScore F1 of 71.5%, representing improvements of 8.2% and 28.8% over specialized VLM baselines. For disease classification, LO-VLM reaches 96% accuracy (F1 = 96%), outperforming ViT by 13% and exceeding medical VLM benchmarks by over 62% (p < 0.05).

conclusionsBy reconciling interpretability with computational efficiency, LO-VLM establishes a paradigm for efficient AI models in OCT interpretation.

Identifiers

PMID41455821
PMCPMC12816039

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.