Evidence map›Paper›PMID 41648417›Full record

ArticlebioRxiv : the preprint server for biology2026

Deconvolving Phylogenetic Distance Mixtures.

Shayesteh Arasti, Ali Osman Berk Şapcı, Eleonora Rachtman, Mohammed El-Kebir, Siavash Mirarab

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Shayesteh ArastiDepartment of Computer Science and Engineering, UC San Diego, CA 92093, USA.ORCID 0009-0005-1607-5300
Ali Osman Berk ŞapcıBioinformatics and Systems Biology Graduate Program, UC San Diego, CA 92093, USA.ORCID 0000-0003-4396-817X
Eleonora RachtmanDepartment of Electrical and Computer Engineering, UC San Diego, CA 92093, USA.ORCID 0000-0002-6104-5750
Mohammed El-KebirSiebel School of Computing and Data Science, University of Illinois Urbana-Champaign, Urbana, IL 61801, USA.ORCID 0000-0002-1468-2407
Siavash MirarabDepartment of Electrical and Computer Engineering, UC San Diego, CA 92093, USA.ORCID 0000-0001-5410-1518

Funding

Biology-aware machine learning methods for characterizing microbiome genotype and phenotypeR35GM142725 · NIGMS · UNIVERSITY OF CALIFORNIA, SAN DIEGO · PI MIR ARABBAYGI, SIAVASH · 2021 to 2025
$1.9M
NIGMS NIH HHS R35 GM142725
6 · The paper itself

Abstract

Mixtures of multiple constituent organisms are sequenced in several widely used applications, including metagenomics and metabarcoding. Characterizing the elements of the sequence mixture and their abundance with respect to a reference set of known organisms has been the subject of intense research across several domains, including microbiome analyses, and methods must overcome two key challenges. First, the mixture constituents are related to each other through an evolutionary history, and hence, should not be considered independent entities. Second, sequence data is noisy, with each short read providing a limited signal. While existing approaches attempt to address these challenges, addressing both challenges simultaneously has proved challenging. For evolutionary dependencies, methods either define hierarchical clusters (e.g., taxonomies or operational taxonomic/genomic units) or use phylogenetic trees. For the second challenge, they either assemble reads into contigs, use statistical priors to summarize read placements, or attempt to analyze all reads jointly using k-mers. Despite this rich literature, a natural approach to simultaneously address both challenges has been underexplored: compute a distance from the mixture to all references, deconvolve those distances, and place the sample on multiple branches of a reference phylogeny with associated abundances. This multi-placement approach is a natural extension of the single-read phylogenetic placement used in practice. We argue that by placing the entire sample on multiple branches instead of placing reads individually, we can obtain a less noisy profile of the mixture. We formalize this approach as the phylogenetic distance deconvolution (PDD) problem, show some limits on the identifiability of PDDs, propose a slow exact algorithm, and an efficient heuristic greedy algorithm with local refinements. Benchmarking shows that these heuristics are effective and that our implementation of the PDD approach (called DecoDiPhy) can accurately deconvolve phylogenetic mixture distances while scaling quadratically. Applied to metagenomics, DecoDiPhy consolidates reads mapped to a large number of branches on a reference tree to a much smaller number of placements. The consolidated placements improve the accuracy of downstream tasks, such as sample differentiation and detection of differentially abundant taxa.

Indexed as

MetagenomicsMixture deconvolutionPhylogenetic distancesPhylogenetic mixture analysis

Identifiers

PMID41648417
PMCPMC12871782

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.