Open research questions in Genomics and Phylogenetic Studies
55 unresolved questions extracted from the limitations and future-work sections of 729 Genomics and Phylogenetic Studies papers in our library. Each links back to the study that raised it.
What the literature leaves open
Mitonuclear co-introgression, a process whereby alleles at N-mt genes move across species boundaries in concert with mitochondrial genomes, has been suggested as a mechanism whereby species could capture heterospecific mitochondria while avoiding mitonuclear incompatibilities, but evidence for this phenomenon is sparse.
Mitochondrial introgression in North American red-backed voles is facilitated by co-introgression at nuclear-encoded mitochondrial genes · 2026 · DOIHowever, there are more than 400 oak species today, divided into eight phylogenetic sections and distributed over four continents, and the extent to which introgression has affected the whole oak phylogeny remains unknown.
ImportanceCandidate Phyla Radiation (CPR) microorganisms represent a major fraction of Earths microbial diversity, yet their biology remains poorly understood.
Premise: Genome skimming (GS) is a cost-effective approach for plant phylogenomics, but its ability to recover informative datasets from different genomic compartments, particularly genome-wide SNPs, remains poorly explored in Solanum.
Making the most out of it: shallow genome-skimming possibilities for the systematics of prickly lineages of Solanum (Solanaceae) · 2026 · DOILateral gene transfer (LGT) has contributed to the genetic makeup of various eukaryotic lineages, yet its prevalence and long-term significance remain poorly understood, particularly for transfers between eukaryotes.
Furthermore, a major barrier to the clinical and epidemiological integration of WGS is the current lack of standardized, accessible identification tools capable of translating complex genomic similarity into harmonized taxonomic assignments.
Supplementary datasets of "MAC-Explorer: bridging genome-based taxonomy and an identification tool for the Mycobacterium avium complex" article · 2026 · DOIHowever, many existing simulators were developed for earlier versions of ONT sequencing or use generic long-read assumptions, and their realism for contemporary ONT data is unclear.
Yet, even if it is difficult and impractical to examine each candidate DNM in multiple pedigree studies with numerous individuals, a subset of candidate DNMs could be examined to estimate FPR.
How precise are mutation rate estimates? Comparison of different approaches to estimate de novo mutation rates · 2026 · DOIOur genome-based reevaluation of Deinococcus taxonomy revealed a clearer picture of evolutionary relationships within this extremophilic genus. Genomic analyses, including digital DNA-DNA hybridization (dDDH), average nucleotide identity (ANI), average amino acid identity (AAI), and phylogenomic reconstructions, identified three key clusters of closely related species. For D. ficus and D. enclensis, despite a low 16S rRNA similarity (97.6%) and conventional DDH (23.9%), high genomic coherence (dDDH 79.5%, ANI 97.2%, AAI 97.6%) supports their classification as a single species, with D. enclensis proposed as a later heterotypic synonym of D. ficus. Similarly, D. seoulensis and D. knuensis share sufficient genomic similarity (dDDH 77.8%, ANI 97.6%, AAI 97.6%) to be conspecific, yet their divergence, likely driven by niche-specific accessory gene variations, warrants subspecies status, with the designation of D. seoulensis subsp. seoulensis subsp. nov. and D. seoulensis subsp. knuensis subsp. nov. In the third cluster, D. arenae and D. kurensis exhibit close genomic profiles (dDDH 81.0%, ANI 97.9%, AAI 98.0%), justifying their unification under D. arenae, with D. kurensis as a synonym, while D. actinosclerus remains distinct (dDDH 64.5–64.7%, ANI 96.2–96.4%, AAI 96.2–96.4%) due to low dDDH values as well as phenotypic and ecological differences. These findings highlight the limitations of traditional markers like 16S rRNA and conventional DDH, emphasizing the necessity of whole-genome approaches to resolve the true evolutionary relationships among bacteria and to establish classifications that accurately reflect bacterial diversity. The refined framework enhances our understanding of Deinococcus diversity and sets a foundation for exploring their ecological and biotechnological significance. 137 Page 10 of 15 Table 2 Phenotypic characteristics of D. enclensis NIO-1023T and D. ficus CC-FR2—10T (Lai et al. 2006; Thorat et al.
Comprehensive phylogenomic analyses support the reclassification of several species belonging to the Deinococcus genus · 2026 · DOILoRTIA Plus occupies an intermediate position between conservative reference-concordance approaches and permissive novel-discovery approaches; comparative studies directly measuring the trade-off between reference agreement and novel isoform recovery across different annotation pipelines in specific application domains (e.g., cancer transcriptomics, viral transcriptomics) are needed.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIThe benchmark framework was intentionally restricted to reference features recovered by at least one annotator within the defined tolerance, which biases evaluation toward the accessible reference space rather than the complete reference set; establishing methods to identify and characterize truly biologically active but currently undetected transcript isoforms and boundaries would improve false-positive assessment.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOILoRTIA Plus shows strongest advantages in the novel TES space for ONT dRNA data using poly(A)-motif-based filtering; the generalizability of poly(A)-motif-based endpoint detection should be evaluated across different tissue types and biological conditions to determine if motif patterns vary sufficiently to require condition-specific filter parameterization.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIFurther chemistry-aware parameter optimization—particularly of support thresholds, artifact filters, and endpoint tolerances—could improve individual annotation pipelines; systematic hyperparameter tuning studies specific to ONT cDNA, ONT dRNA, and PacBio IsoSeq chemistries should be conducted to establish optimal parameter ranges for each long-read platform.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIThe human TSS and TES truth sets rely on CAGE-, PolyASite-, and GENCODE/SQANTI3-based resources, each with different detection limits and positional uncertainty; a systematic cross-validation study quantifying the coordinate discrepancies and detection blind spots among these three reference sources would clarify which transcript endpoints are most challenging to validate.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIThe KSHV reference set is necessarily incomplete with low-abundance isoforms and rare boundary events likely remaining unannotated; a comprehensive characterization of which low-abundance isoforms and rare splice-site junction patterns are systematically missed by current long-read annotation pipelines would improve ground-truth benchmarking.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIONT dRNA protocols currently suffer from loss of extreme 5' terminal nucleotides and basecalling uncertainty near the poly(A) tail at the 3' end; future dRNA protocols that improve 5' cap preservation or recover missing 5' information through adapter ligation should be systematically evaluated to assess how adapter-aware boundary detection in LoRTIA Plus performs with these improved endpoint recovery methods.
LoRTIA Plus: a chemistry-agnostic, feature-first software package for long-read transcriptome annotation · 2026 · DOIHow the random walk on tree space using OLA encoding compares to other random walk strategies (e.g., directly on NNI or SPR neighborhoods) in terms of convergence and computational cost remains unexplored.
The paper proves theoretical bounds on tree distances (NNI, SPR) but does not explore whether these bounds are tight or how often they are achieved in practice.
No comparison is provided between the OLA-based random walk and other established methods for traversing tree space in terms of mixing time or practical efficiency.
Despite this correlation between TR length and phenotype, TRs have been understudied owing to the difficulty in developing accurate, high-throughput, genomewide assays13.
2 One limitation of GENA-LMs arises from the granular- ity imposed by the use of BPE tokenization, which confines predictions to specific tokens. Sparse GENA-LMs, on the other hand, are limited to the lengths on which they were trained.
INFERRING PHYLOGENETIC RELATIONSHIPS BETWEEN CLOSELY RELATED TAXA CAN BE HINDERED BY THREE FACTORS: (1) the lack of informative molecular variation at short evolutionary timescale; (2) the lack of established markers in poorly studied taxa; and (3) the potential phylogenetic conflicts among different genomic regions due to incomplete lineage sorting or introgression.
Is <scp>RAD</scp>‐seq suitable for phylogenetic inference? An in silico assessment and optimization · 2013 · DOIWhat is lacking and needed now is a concerted effort, comparable to the Human Genome Project (HGP), to complete a global biodiversity survey — pole to pole, whales to bacteria, and in a reasonably short period of time.
IMPORTANCE Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined.
Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample · 2026 · DOIConifers, which comprise nearly two-thirds of extant gymnosperm species, are ecologically and economically important but remain genomically understudied because of their exceptionally large, repeat-rich genomes.
Chromosome-scale assembly of the Cupressus sempervirens genome unravels new insights into the evolutionary history of conifers · 2026 · DOI
Most-cited papers in Genomics and Phylogenetic Studies
- The Sequence of the Human Genome · Science · 2001 · 10,184 citations
- Interactive Tree of Life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool · Nucleic Acids Research · 2024 · 3,940 citations
- A Genomic Perspective on Protein Families · Science · 1997 · 2,948 citations
- Towards complete and error-free genome assemblies of all vertebrate species · Nature · 2021 · 2,803 citations
- KEGG: biological systems database as a model of the real world · Nucleic Acids Research · 2024 · 2,207 citations
- InterPro: the protein sequence classification resource in 2025 · Nucleic Acids Research · 2024 · 942 citations
- A DNA barcoding framework for taxonomic verification in the Darwin Tree of Life Project · Wellcome Open Research · 2024 · 662 citations
- Genomes on a Tree (GoaT): A versatile, scalable search engine for genomic and sequencing project metadata across the eukaryotic tree of life · Wellcome Open Research · 2023 · 630 citations
- Ensembl 2025 · Nucleic Acids Research · 2024 · 612 citations
- The UCSC Genome Browser database: 2025 update · Nucleic Acids Research · 2024 · 604 citations
Most recent work
- WitChi: Efficient Detection and Pruning of Compositional Bias in Phylogenomic Alignments Using Empirical Chi-Squared Testing · bioRxiv · 2026
- Biodiversity Genomics Europe (BGE) Project – Abridged Grant Proposal · Research Ideas and Outcomes · 2026
- Comparative phylotranscriptomics of four sympatric tetrigids provides implications for convergent evolution and morphological discordance · BMC Genomics · 2026
- GeneCAD: Plant Genome Annotation with a DNA Foundation Model · bioRxiv · 2026
- Quartet-based species tree methods enable fast and consistent tree of blobs reconstruction under network multispecies coalescent · bioRxiv · 2026
- Rapid Speciation Characterized by Incomplete Lineage Sorting in the Globally Distributed Bacterium Sulfitobacter · bioRxiv · 2026
- Rapid and consistent clustering of millions of genomes highlights the diversity of prokaryotic life · bioRxiv · 2026
- CrossFilt: a cross-species filtering tool that eliminates alignment bias in comparative genomics studies of primates · Genome Biology · 2026
- New lineages provide insights into the convergent evolution of extreme salt adaptation within symbiotic Archaea · Molecular Biology and Evolution · 2026
- Purifying selection purges harmful variants in the rarest pine · bioRxiv · 2026
Find a gap in your own Genomics and Phylogenetic Studies sub-topic
This page shows what the Genomics and Phylogenetic Studies literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →