Speaker
Description
Contemporary de novo assemblers often produce near-perfect telomere-to-telomere assemblies of chromosomes from long-read data, alongside hundreds of shorter accessory contigs with diverse origins: under-collapsed heterozygous sequence, organellar genome isoforms, rDNA, and cobionts. We developed a software tool called dnadis (de novo assembly disambiguator) to automate common analyses following assembly and produce machine-readable tabular reports, rich HTML summaries, and classified sequence files to classify contigs, diagnose assembly errors, produce useful biological insights, and prepare for downstream annotation.
Beginning with whole genome alignment to one or more reference genomes, dnadis orchestrates chromosome-length contig assignment, naming, reorientation, scaffolding. When genomic read sets are available, dnadis automates coverage analysis to surface aneuploidy and distinguish rearrangements from misassemblies. Hybrid and polyploid chromosome sets are disentangled with an implementation of Gaussian mixture model segmentation and/or simultaneous mapping to multiple reference genomes. Classification modules using community-standard external tools and novel algorithms perform organelle identification, rDNA array annotation, cobiont and contaminant screening, and quality assessment. If several closely related query assemblies are provided, dnadis produces additional comparative reports and plots which can be especially helpful in identifying conserved changes to chromosome architecture within a clade.
dnadis is open-source software and supports local and distributed execution for scalable use with dozens of assemblies.
Keywords
Comparative genomics, Curation, Validation, Visualization
| Corresponding author email | martiens@cshl.edu |
|---|---|
| Scientific Session | Genes, Genomes and Evolution |