Evidence map›Paper›PMID 40654677›Full record

ArticlebioRxiv : the preprint server for biology2025

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

Conner J Langeberg, Taehan Kim, Roma Nagle, Charlotte Meredith, Dimple Amitha Garuadapuri, Jennifer A Doudna, Jamie H D Cate

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

7 authors.

Conner J LangebergInnovative Genomics Institute; University of California, Berkeley, CA, USA.ORCID 0000-0002-5609-3758
Taehan KimDepartment of Computer Science, University of California, Berkeley; Berkeley, CA, USA.ORCID 0009-0000-4431-4640
Roma NagleDepartment of Molecular and Cell Biology, University of California, Berkeley; Berkeley, CA, USA.ORCID 0009-0008-8550-3542
Charlotte MeredithDepartment of Molecular and Cell Biology, University of California, Berkeley; Berkeley, CA, USA.ORCID 0009-0002-4597-3239
Dimple Amitha GaruadapuriDepartment of Bioengineering, University of California, Berkeley; Berkeley, CA, USA.ORCID 0009-0009-3486-2999
Jennifer A DoudnaInnovative Genomics Institute; University of California, Berkeley, CA, USA.ORCID 0000-0001-9161-999X
Jamie H D CateInnovative Genomics Institute; University of California, Berkeley, CA, USA.ORCID 0000-0001-5965-7902

Funding

Targeting Viroporins and Coronavirus M ProteinU19AI171110 · NIAID · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI Nevan J Krogan · 2022 to 2026
$103.4M
Project 3U54AI170792 · NIAID · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI Nevan J Krogan · 2022 to 2026
$35.7M
RESEARCH PROJECT 2U19AI135990 · NIAID · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI Dexter Pratt · 2018 to 2026
$21.4M
Resource Core II: In Vivo CoreU19NS132303 · NINDS · UNIVERSITY OF CALIFORNIA BERKELEY · PI NIREN MURTHY · 2023 to 2026
$19.7M
Mechanisms of Translation Control in HumansR35GM148352 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI JAMIE H CATE · 2023 to 2026
$2.3M
Expanding CRISPR-Cas editing technology through exploration of novel Cas proteins and DNA repair systemsU01AI142817 · NIAID · UNIVERSITY OF CALIFORNIA BERKELEY · PI BANFIELD, JILLIAN, DOUDNA, JENNIFER A · 2018 to 2022
$2.0M
Cas9 RNP delivery to immune cells in vivo via molecular targetingUH3AI150552 · NIAID · UNIVERSITY OF CALIFORNIA BERKELEY · PI DOUDNA, JENNIFER A, WILSON, ROSS C · 2022 to 2022
$1.3M
NIAID NIH HHS U01 AI142817NIAID NIH HHS U19 AI135990NIAID NIH HHS U19 AI171110NIAID NIH HHS U54 AI170792NIAID NIH HHS UH3 AI150552NIGMS NIH HHS R35 GM148352NINDS NIH HHS U19 NS132303
6 · The paper itself

Abstract

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Identifiers

PMID40654677
PMCPMC12247784

What Socratic holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.