Evidence map›Paper›PMID 40950025›Full record

ArticlebioRxiv : the preprint server for biology2025

PlantCAD2: A Long-Context DNA Language Model for Cross-Species Functional Annotation in Angiosperms.

Jingjing Zhai, Aaron Gokaslan, Sheng-Kai Hsu, Szu-Ping Chen, Zong-Yan Liu, Edgar Marroquin, Eric Czech, Betsy Cannon, Ana Berthel, M Cinta Romay and 3 more

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

13 authors.

Jingjing ZhaiInstitute for Genomic Diversity, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0002-1535-3103
Aaron GokaslanDepartment of Computer Science, Cornell University, Ithaca, NY, USA 14853.ORCID 0000-0002-3575-2961
Sheng-Kai HsuInstitute for Genomic Diversity, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0002-6942-7163
Szu-Ping ChenSection of Plant Breeding and Genetics, Cornell University, Ithaca, NY USA 14853.ORCID 0009-0002-1821-4181
Zong-Yan LiuSection of Plant Breeding and Genetics, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0002-9039-7843
Edgar MarroquinDepartment of Computer Science, Cornell University, Ithaca, NY, USA 14853.ORCID 0009-0009-0265-1214
Eric CzechOpen Athena AI Foundation, New York, NY, USA 10001.ORCID 0000-0002-4254-4255
Betsy CannonOpen Athena AI Foundation, New York, NY, USA 10001.ORCID 0009-0000-0581-3703
Ana BerthelInstitute for Genomic Diversity, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0001-6849-982X
M Cinta RomayInstitute for Genomic Diversity, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0001-9309-1586
Matt PennellDepartment of Computational Biology, Cornell University, Ithaca, NY, USA 14853.ORCID 0000-0002-2886-3970
Volodymyr KuleshovDepartment of Computer Science, Cornell University, Ithaca, NY, USA 14853.ORCID 0000-0002-5150-3308
Edward S BucklerInstitute for Genomic Diversity, Cornell University, Ithaca, NY USA 14853.ORCID 0000-0002-3100-371X

Funding

Next-Generation Algorithms in Statistical Genetics Based on Modern Machine LearningR35GM151243 · NIGMS · CORNELL UNIVERSITY · PI Volodymyr Kuleshov · 2023 to 2026
$1.6M
Leveraging phylogenetic approaches to investigate the evolution of geneexpressionR35GM151348 · NIGMS · UNIVERSITY OF SOUTHERN CALIFORNIA · PI Matthew Wesley Pennell · 2023 to 2026
$1.6M
NIGMS NIH HHS R35 GM151243NIGMS NIH HHS R35 GM151348
6 · The paper itself

Abstract

Understanding how DNA sequence encodes biological function remains a fundamental challenge in biology. Flowering plants (angiosperms), the dominant terrestrial clade, exhibit maximal biochemical complexity, extraordinary species diversity (over 100,000 species), relatively recent origins (~160 million years), ~200-fold variation in genome size and relative compact coding regions compared with other eukaryotes. These features present both a unique challenge and opportunity for pre-training DNA language models to understand plant-specific evolutionary conservation, regulatory architectures and genomic functions. Here, we introduce PlantCAD2, a long-context, plant-specific DNA language model with single-nucleotide resolution, pre-trained on 65 angiosperm genomes, together with a series of public benchmarks for evaluation. Comprehensive zero-shot testing shows that PlantCAD2 (676 million parameters) efficiently captures evolutionary conservation, surpassing the 7-billion-parameter Evo2 model in 10 of 12 tasks. With parameter-efficient fine-tuning, PlantCAD2 also outperforms the 1-billion-parameter AgroNT across seven cross-species tasks including chromatin accessible region, gene expression and protein translation. Moreover, its 8,192bp context window substantially improves accessible chromatin prediction in large genomes such as maize (AUPRC increasing from 0.587 to 0.711), underscoring the importance of long-range context for modeling distal regulation. Together, these results establish PlantCAD2 as a powerful, efficient, and versatile foundation model for plant genomics, enabling accurate genome annotation across diverse species.

Identifiers

PMID40950025
PMCPMC12425018

What Socratic holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.