HomeCategoriesGenomics
134 tools

Genomics Tools

Discover our collection of 134 research tools and applications for genomics.

Related Categories

Sequence analysis52
Bioinformatics25
Transcriptomics16
Genome assembly14
Genome annotation11
Sequence assembly10
+92 more

Tools in Genomics

Found 135 of 135 tools

BEDTools is an extensive suite of utilities for comparing genomic features in BED format.

This R package provide functions that are used in the BREW3R workflow. This mainly contains a function that extend a gtf as GRanges using information from another gtf (also as GRanges). The process allows to extend gene annotation without increasing the overlap between gene ids.

Provides measures for quantitative assessment of genome assembly, gene set, and transcriptome completeness based on evolutionarily informed expectations of gene content from near-universal single-copy orthologs.

Predict the location of ribosomal RNA genes in genomes. It supports bacteria (5S,23S,16S), archaea (5S,5.8S,23S,16S), mitochondria (12S,16S) and eukaryotes (5S,5.8S,28S,18S).

Bismark is a tool to map bisulfite treated sequencing reads and perform methylation calling in a quick and easy-to-use fashion.

Detect blocks of overlapping reads using a gaussian-distribution approach

De-novo assembly tool for long read chemistry like Nanopore data and PacBio data.

Cactus is a reference-free whole-genome multiple alignment program.

F-Seq2 is a Python-based peak caller for high-throughput sequencing data (ChIP-seq, DNase-seq, ATAC-seq) that uses kernel density estimation combined with a local Poisson statistical framework to identify biologically meaningful genomic regions.

Collection of command line tools for Short-Reads FASTA/FASTQ files preprocessing.

Filtlong is a tool for filtering long reads by quality. It can take a set of long reads and produce a smaller, better subset. It uses both read length (longer is better) and read identity (higher is better) when choosing which reads pass the filter.

Bayesian genetic variant detector designed to find small polymorphisms, specifically SNPs, indels, multi-nucleotide polymorphisms, and complex events (composite insertion and substitution events) smaller than the length of a short-read sequencing alignment.

GCTA (Genome-wide Complex Trait Analysis) is a software package initially developed to estimate the proportion of phenotypic variance explained by all genome-wide SNPs for a complex trait but has been greatly extended for many other analyses of data from genome-wide association studies (GWASs).

Test for association in genome-wide association studies (GWAS) using a standard linear mixed model to account for population stratification and sample structure. It calculates exact Wald or likelihood ratio test statistics and p-values, and is computationally efficient for large GWAS.

GFAffix identifies walk-preserving shared affixes in variation graphs and collapses them into a non-redundant graph structure.

Genomic Mapping and Alignment Program for mRNA and EST Sequences.

A high-speed next-gen sequencing structural variation caller. It calls variants based on alignment-guided positional de Bruijn graph breakpoint assembly, split read, and read pair evidence.

A comprehensive package for performing gene set enrichment analysis in Python.

A software package for analyzing various features of gene models.

GangSTR is a tool for genome-wide profiling tandem repeats from short reads. A key advantage of GangSTR over existing genome-wide TR tools is that it can handle repeats that are longer than the read length. GangSTR takes aligned reads (BAM) and a set of repeats in the reference genome as input and outputs a VCF file containing genotypes for each locus.

GapFiller is a seed-and-extend local assembler to fill the gap within paired reads. It can be used for both DNA and RNA and it has been tested on Illumina data. GapFiller can be used whenever a sequence is to be assembled starting from reads lying on its ends, provided a loose estimate of sequence length.

GeMoMa

Gene Model Mapper is a homology-based gene prediction program. GeMoMa uses the annotation of protein-coding genes in a reference genome to infer the annotation of protein-coding genes in a target genome. Thereby, it utilizes amino acid sequence and intron position conservation. In addition, it allows to incorporate RNA-seq evidence for splice site prediction.

GenMap is a fast and exact tool for computing genome mappability, calculating the uniqueness of k-mers across genomic positions while allowing for a specified number of mismatches, helping identify unique and repetitive genomic regions.

GenomeScope2 is a reference-free tool that uses k-mer frequency analysis to estimate genome size, heterozygosity, ploidy, and repeat content from raw sequencing data, supporting both diploid and polyploid genomes.

Genomepy is designed to provide a simple and straightforward way to download and use genomic data. This includes (1) searching available data, (2) showing the available metadata, (3) automatically downloading, preprocessing and matching data and (4) generating optional aligner indexes. All with sensible, yet controllable defaults. Currently, genomepy supports Ensembl, UCSC, NCBI and GENCODE.

GenomicConsensus

The GenomicConsensus package provides the variantCaller tool, which allows you to apply the Quiver or Arrow algorithm to mapped PacBio reads to get consensus and variant calls.

Genrich is a peak-caller for genomic enrichment assays (e.g. ChIP-seq, ATAC-seq). It analyzes alignment files generated following the assay and produces a file detailing peaks of significant enrichment.

A new gene finder based on a Generalized Hidden Markov Model. Although the gene finder conforms to the overall mathematical framework of a GHMM, additionally it incorporates splice site models adapted from the GeneSplicer program and a decision tree adapted from GlimmerM. It also utilizes Interpolated Markov Models for the coding and noncoding models . Currently, GlimmerHMM's GHMM structure includes introns of each phase, intergenic regions, and four types of exons.

HELEN (Homopolymer Encoded Long-read Error-corrector for Nanopore) is a highly optimized genome polishing pipeline that uses deep learning to correct errors in Nanopore long-read assemblies, working in conjunction with MarginPolish and supporting GPU acceleration for scalability.

HMMER is used for searching sequence databases for sequence homologs, and for making sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs).

The main purpose of HTSlib is to provide access to genomic information files, both alignment data (SAM, BAM, and CRAM formats) and variant data (VCF and BCF formats). The library also provides interfaces to access and index genome reference data in FASTA format and tab-delimited files with genomic coordinates. It is utilized and incorporated into both SAMtools and BCFtools.

HTStream is a quality control and processing pipeline for High Throughput Sequencing data. The difference between HTStream and other tools is that HTStream uses a tab delimited fastq format that allows for streaming from application to application. This streaming creates some awesome efficiencies when processing HTS data and makes it fully interoperable with other standard Linux tools.

Hail is a scalable, cloud-native genomic analysis tool designed for large datasets. It provides a query language for genomic data and supports batch computing for efficient variant calling and other analyses.

Hapo-G is a tool that aims to improve the quality of genome assemblies by polishing the consensus with accurate reads. It capable of incorporating phasing information from high-quality reads (short or long-reads) to polish genome assemblies and in particular assemblies of diploid and heterozygous genomes.

This tool was designed to process Hi-C data, from raw fastq files (paired-end Illumina data) to the normalized contact maps. Since version 2.7.0, it can analyze data from digestion protocols as well as data from protocols that do not require restriction enzyme such as DNase Hi-C. The pipeline is flexible, scalable and optimized. It can operate either on a single laptop or on a computational cluster using the PBS-Torque scheduler.

A web server for reproducible Hi-C, capture Hi-C and single-cell Hi-C data analysis, quality control and visualization. HiCExplorer — HiCExplorer 3.6 documentation. scHiCExplorer — scHiCExplorer 7 documentation. Free document hosting provided by Read the Docs.

Homopolish is a method for the removal of systematic errors in nanopore sequencing by homologous polishing.

HyPo, a Hybrid Polisher, utilizes short as well as long reads within a single run to polish a long reads assembly of small and large genomes.

High-performance visualization tool for interactive exploration of large, integrated datasets. It supports a wide variety of data types and format, including short-read alignments in the SAM/BAM format. Data can be viewed from local files or over the web via http.

IMPUTE2

IMPUTE2 is a genotype imputation and haplotype phasing tool that uses a multi-population reference panel to impute missing genotypes in genome-wide association studies (GWAS), enabling researchers to increase the density of genetic markers and improve statistical power for association analyses.

Automated identification of insertion sequence elements in prokaryotic genomes.

ISMapper searches for IS positions in sequence data using paired end Illumina short reads, an IS query/queries of interest and a reference genome. ISMapper reports the IS positions it has found in each isolate, relative to the provided reference genome.

A database which integrates together predictive information about proteins' function from a number of partner resources, giving an overview of the families that a protein belongs to and the domains and sites it contains. Users who have novel nucleotide or protein sequences that they wish to functionally characterise can use the software package InterProScan to run the scanning algorithms from the InterPro database in an integrated way. Sequences are submitted in FASTA format. Matches are then calculated against all of the required member database's signatures and the results are then output in a variety of formats.

IsoQuant is a tool for the genome-based analysis of long RNA reads, such as PacBio or Oxford Nanopores.

IsoSeq v3 contains the newest tools to identify transcripts in PacBio single-molecule sequencing data. Starting in SMRT Link v6.0.0, those tools power the IsoSeq GUI-based analysis application. A composable workflow of existing tools and algorithms, combined with a new clustering technique.

KMA is mapping a method designed to map raw reads directly against redundant databases, in an ultra-fast manner using seed and extend.

KMC is a utility designed for counting k-mers (sequences of consecutive k symbols) in a set of reads from genome sequencing projects.

KentUtils is a collection of UCSC command-line bioinformatic utilities for genome browser data processing, including tools for format conversion, genome annotation, sequence analysis, and manipulation of genomic data formats such as BED, BigWig, BigBed, and BAM files.

KisSplice

KisSplice is a software that enables to analyse RNA-seq data with or without a reference genome.

KmerGenie

KmerGenie estimates the best k-mer length for genome de novo assembly. Given a set of reads, KmerGenie first computes the k-mer abundance histogram for many values of k. Then, for each value of k, it predicts the number of distinct genomic k-mers in the dataset, and returns the k-mer length which maximizes this number. Experiments show that KmerGenie's choices lead to assemblies that are close to the best possible over all k-mer lengths.

Kover is an out-of-core implementation of rule-based machine learning algorithms that has been tailored for genomic biomarker discovery. It produces highly interpretable models, based on k-mers, that explicitly highlight genotype-to-phenotype associations.

LAST is a sequence alignment tool that finds and aligns related regions between large biological sequences, supporting DNA-DNA, DNA-protein, and protein-protein comparisons with sensitivity comparable to BLAST but with greater speed and flexibility for large datasets.

A tool for (1) aligning two DNA sequences, and (2) inferring appropriate scoring parameters automatically.

The Long Read Aligner for Sequences and Contigs. LRA, the long read aligner for sequences and assembly contigs LRA is a sequence alignment program that aligns long reads from single-molecule sequencing (SMS) instruments, or megabase-scale contigs from SMS assemblies. LRA implements seed chaining sparse dynamic programming with a convex gap function to read and assembly alignment, which is also extended to allow for inversion cases. Through the Truvari analysis of LRA, Minimap2 and NGM-LR alignments. LRA achieves higher f1 score over HG002 HiFi, CLR and ONT datasets. Home: https://github.com/ChaissonLab/LRA. Long read aligner for sequences and contigs.

LTR_Finder (Long Terminal Repeat Finder) is an efficient program for finding full-length LTR retrotransposons in genome sequences.

LTRpred is an R package for de novo annotation and prediction of LTR retrotransposons in genome sequences, using structural features and sequence homology to identify and classify LTR retrotransposon families.

LYVE version of the Snp Extraction Tool (SET), a method of using hqSNPs to create a phylogeny.

Lambda is a local aligner optimized for many query sequences and searches in protein space. It is compatible to BLAST, but much faster than BLAST and many other comparable tools.

An accurate gene annotation mapping tool.

LiftoffTools is a toolkit for comparing gene annotations mapped between genome assemblies, enabling the detection and analysis of gene sequence variants, synteny, and gene copy number changes. It leverages Liftoff for annotation transfer and offers modules for analyzing protein-coding genes, gene synteny, and gene copy number.

LongPhase is an ultra-fast program for simultaneously co-phasing SNPs, small indels, large SVs, and (5mC) modifications for Nanopore and PacBio platforms. It can produce nearly chromosome-scale haplotype blocks by using Nanpore ultra-long reads without the need for additional trios, chromosome conformation, and strand-seq data. LongPhase can phase a 30x human genome in ~1 minute

LongQC is a tool for the data quality control of the PacBio and ONT long reads, and it has two functionalities: sample qc and platform qc.

Computational tool to identify important genes from the recent genome-scale CRISPR-Cas9 knockout screens technology.

MUMmer is a modular system for the rapid whole genome alignment of finished or draft sequence. Basically it is a ultra-fast alignment of large-scale DNA and protein sequences

Antimicrobial peptide screening in genomes and metagenomes. Macrel: (Meta)genomic AMP Classification and Retrieval. Pipeline to mine antimicrobial peptides (AMPs) from (meta)genomes.

Portable and easily configurable genome annotation pipeline. It’s purpose is to allow smaller eukaryotic and prokaryotic genome projects to independently annotate their genomes and to create genome databases.

MinCED is a program to find Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs) in full genomes or environmental datasets such as assembled contigs from metagenomes.

Nextclade is an open-source project for viral genome alignment, mutation calling, clade assignment, quality checks and phylogenetic placement.

Optimized dynamic genome graph implementation: a toolkit for understanding pangenome graphs

Software tool to annotate bacterial, archaeal and viral genomes quickly and produce standards-compliant output files.

A high speed stand alone pan genome pipeline, which takes annotated assemblies in GFF3 format (produced by Prokka (Seemann, 2014)) and calculates the pan genome.

De novo assembly from Oxford Nanopore reads.

TransDecoder identifies candidate coding regions within transcript sequences, such as those generated by de novo RNA-Seq transcript assembly using Trinity, or constructed based on RNA-Seq alignments to the genome using Tophat and Cufflinks.

API and command line utilities for the manipulation of VCF files.

ABACAS is intended to rapidly contiguate (align, order, orientate) , visualize and design primers to close gaps on shotgun assembled contigs based on a reference sequence. It uses MUMmer to find alignment positions and identify syntenies of assembly contigs against the reference. The output is then processed to generate a pseudomolecule taking overlaping contigs and gaps in to account. MUMmer's alignment generating programs, Nucmer and Promer are used followed by the 'delta-filter' utility function. Users could also run tblastx on contigs that are not used to generate the pseudomolecule.

Abismal is a mapper of FASTQ bisulfite-converted short reads (between 50 and 1000 bases) to a FASTA reference genome.

abPOA: an SIMD-based C library for fast partial order alignment using adaptive band. abPOA can perform multiple sequence alignment (MSA) on a set of input sequences and generate a consensus sequence by applying the heaviest bundling algorithm to the final alignment graph.

Mass screening of contigs for antimicrobial resistance or virulence genes.

ACTC (Align subreads to CCS reads) is developed by Pacific Biosciences and provides a one-click solution for aligning individual subreads to the corresponding circular consensus (CCS) reads — useful in workflows involving HiFi/CCS read analysis from PacBio sequencing.

AdapterRemoval searches for and removes adapter sequences from High-Throughput Sequencing (HTS) data and (optionally) trims low quality bases from the 3' end of reads following adapter removal. AdapterRemoval can analyze both single end and paired end data, and can be used to merge overlapping paired-ended reads into (longer) consensus sequences. Additionally, AdapterRemoval can construct a consensus adapter sequence for paired-ended reads, if which this information is not available.

adVNTR is a tool for genotyping Variable Number Tandem Repeats (VNTR) from sequence data. It works with both NGS short reads (Illumina HiSeq) and SMRT reads (PacBio) and finds diploid repeating counts for VNTRs and identifies possible mutations in the VNTR sequences.

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data AfterQC can simply go through all fastq files in a folder and then output three folders: good, bad and QC folders, which contains good reads, bad reads and the QC results of each fastq file/pair.

Another Gff Analysis Toolkit (AGAT) Suite of tools to handle gene annotations in any GTF/GFF format.

AGFusion (pronounced 'A G Fusion') is a python package for annotating gene fusions from the human or mouse genomes.

Alien_hunter is an application for the prediction of putative Horizontal Gene Transfer (HGT) events with the implementation of Interpolated Variable Order Motifs (IVOMs).

AlignStats produces various alignment, whole genome coverage, and capture coverage metrics for sequence alignment files in SAM, BAM, and CRAM format. This program is designed to serve reporting and quality control purposes in sequencing analysis pipelines at the Baylor College of Medicine Human Genome Sequencing Center (BCM-HGSC).

AMPtk is a series of scripts to process NGS amplicon data using USEARCH and VSEARCH, it can also be used to process any NGS amplicon data and includes databases setup for analysis of fungal ITS, fungal LSU, bacterial 16S, and insect COI amplicons. It can handle Ion Torrent, MiSeq, and 454 data.

AnchorWave (Anchored Wavefront Alignment) identifies collinear regions via conserved anchors (full-length CDS and full-length exon have been implemented currently) and breaks collinear regions into shorter fragments, i.e., anchor and inter-anchor intervals. By performing sensitive sequence alignment for each shorter interval via a 2-piece affine gap cost strategy and merging them together, AnchorWave generates a whole-genome alignment for each collinear block. AnchorWave implements commands to guide collinear block identification with or without chromosomal rearrangements and provides options to use known polyploidy levels or whole-genome duplications to inform alignment.

Convert various sequence formats to FASTA

Scaffolding genome assemblies using 10X Genomics Chromium data or stLFR linked reads

GUI program that allows users to interact with the assembly graphs made by de novo assemblers such as Velvet, SPAdes, MEGAHIT and others. It visualises assembly graphs, with connections, using graph layout algorithms.

Tools for early stage NGS alignment file processing including fast sorting and duplicate marking.

Software for mapping Single Molecule Sequencing (SMS) reads that are thousands of bases long, with divergence between the read and genome dominated by insertion and deletion error.

Visualisation, quality control and taxonomic partitioning of genome datasets.

Bowtie is an ultrafast, memory-efficient short read aligner.

BRAKER3 is a pipeline for fully automated prediction of protein coding gene structures with GeneMark-ES/ET and AUGUSTUS in novel eukaryotic genomes. BRAKER3 is the latest pipeline in the BRAKER suite. It enables the usage of RNA-seq and protein data in a fully automated pipeline to train and predict highly reliable genes with GeneMark-ETP and AUGUSTUS. The result of the pipeline is the combined gene set of both gene prediction tools, which only contains genes with very high support from extrinsic evidence.

Fast and accurate alignment of BS-Seq reads using bwa-mem and a 3-letter genome

CheckM provides a set of tools for assessing the quality of genomes recovered from isolates, single cells, or metagenomes.

Rust implementation of NanoFilt+NanoLyse, both originally written in Python. This tool, intended for long read sequencing such as PacBio or ONT, filters and trims a fastq file.

Chromap is an ultrafast method for aligning and preprocessing high throughput chromatin profiles. Typical use cases include: (1) trimming sequencing adapters, mapping bulk ATAC-seq or ChIP-seq genomic reads to the human genome and removing duplicates; (2) trimming sequencing adapters, mapping single cell ATAC-seq genomic reads to the human genome, correcting barcodes, removing duplicates and performing Tn5 shift; (3) split alignment of Hi-C reads against a reference genome. In all these three cases, Chromap is 10-20 times faster while being accurate.

A tool to circularize genome assemblies. Circlator will attempt to identify each circular sequence and output a linearised version of it. It does this by assembling all reads that map to contig ends and comparing the resulting contigs with the input assembly.

CIRIquant is a comprehensive analysis pipeline for circRNA detection and quantification in RNA-Seq data

CIRIquant is a comprehensive analysis pipeline for circRNA detection and quantification in RNA-Seq data

Suite of tools to discover structural variations such as (larger) insertions and deletions in genomes from paired-end sequencing reads.

A tool for CNV discovery and genotyping from depth-of-coverage by mapped reads

A tool for quick quality assessment of cram and bam files, intended for long read sequencing.

An efficient tool for converting genome coordinates between assemblies. CrossMap supports most of the commonly used file formats, including BAM, sequence alignment map, Wiggle, BigWig, browser extensible data, general feature format, gene transfer format and variant call format.

Find and remove adapter sequences, primers, poly-A tails and other types of unwanted sequence from your high-throughput sequencing reads.

A genome assembler that reduces the computational time of human genome assembly from 400,000 CPU hours to 2,000 CPU hours, utilizing long erroneous 3GS sequencing reads and short accurate NGS sequencing reads.

DeepConsensus uses gap-aware sequence transformers to correct errors in Pacific Biosciences (PacBio) Circular Consensus Sequencing (CCS) data. This results in greater yield of high-quality reads.

deepTools addresses the challenge of handling the large amounts of data that are now routinely generated from DNA sequencing centers. deepTools contains useful modules to process the mapped reads data for multiple quality checks, creating normalized coverage files in standard bedGraph and bigWig file formats, that allow comparison between different files (for example, treatment and control). Finally, using such normalized and standardized files, deepTools can create many publication-ready visualizations to identify enrichments and for functional annotations of the genome.

Delly is an integrated structural variant (SV) prediction method that can discover, genotype and visualize deletions, tandem duplications, inversions and translocations at single-nucleotide resolution in short-read and long-read massively parallel sequencing data. It uses paired-ends, split-reads and read-depth to sensitively and accurately delineate genomic rearrangements throughout the genome.

Assemble bacterial isolate genomes from Nanopore reads. Dragonflye is a pipeline that aims to make assembling Oxford Nanopore reads quick and easy.

Fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication.

For fast functional annotation of novel sequences. It uses precomputed orthologous groups and phylogenies from the eggNOG database to transfer functional information from fine-grained orthologs only. Its common uses include the annotation of novel genomes, transcriptomes or even metagenomic gene catalogs. The use of orthology predictions for functional annotation permits a higher precision than traditional homology searches, as it avoids transferring annotations from close paralogs.

Tool for predicting effects of variants for any genome in Ensembl or with genome annotation (via GFF). This includes vertebrates and also plants, fungi, protists, metazoa and bacteria.

fastStructure is an algorithm for inferring population structure from large SNP genotype data. It is based on a variational Bayesian framework for posterior inference and is written in Python2.x.

A tool to download FASTQs associated with Study, Experiment, or Run accessions.

fastq-scan reads a FASTQ from STDIN and outputs summary statistics (read lengths, per-read qualities, per-base qualities) in JSON format.

Rewrite paired end fastq files to make sure that all reads have a mate and to separate out singletons. This code does one thing: it takes two fastq files, and generates four fastq files. That's right, for free it doubles the number of fastq files that you have!!

Set of Linux utilities to validate and manipulate fastq files. It also includes a set of programs to preprocess barcodes (namely UMIs, cells and samples), add the barcodes as tags in BAM files and count UMIs.

Command line utility for manipulating FASTQ files

funannotate is a pipeline for genome annotation (built specifically for fungi, but will also work with higher eukaryotes).

fwdpy11 is a Python package for forward-time population genetic simulation, using a C++ back-end (fwdpp) for efficiency. It supports flexible modelling of selection, demography, and multiple populations, with custom temporal samplers for analyzing populations during simulation.

gfatools is a set of tools for manipulating sequence graphs in the GFA or the rGFA format. It has implemented parsing, subgraph and conversion to FASTA/BED.

Program for comparing, annotating, merging and tracking transcripts in GFF files.

Program for filtering, converting and manipulating GFF files

Python package for working with GFF and GTF files. It allows operations which would be complicated or time-consuming using a text-file-only approach.

iVar is a computational package that contains functions broadly useful for viral amplicon-based sequencing.

Interactive assembly and analysis of RADseq datasets. ipyrad: interactive assembly and analysis of RAD-seq data sets. Welcome to ipyrad, an interactive toolkit for assembly and analysis of restriction-site associated genomic data sets (e.g., RAD, ddRAD, GBS) for population genetic and phylogenetic studies. Welcome to ipyrad — ipyrad documentation.

khmer is a set of command-line tools for working with DNA shotgun sequencing data from genomes, transcriptomes, metagenomes, and single cells. khmer can make de novo assemblies faster, and sometimes better. khmer can also identify (and fix) problems with shotgun data.

A command-line algorithm for counting k-mers in DNA sequence.

KofamScan is a gene function annotation tool based on KEGG Orthology and hidden Markov model. You need KOfam database to use this tool.

lima is the standard tool to identify barcode and primer sequences in PacBio single-molecule sequencing data. It powers the Demultiplex Barcodes GUI-based analysis applications.

LoFreq* (i.e. LoFreq version 2) is a fast and sensitive variant-caller for inferring SNVs and indels from next-generation sequencing data. It makes full use of base-call qualities and other sources of errors inherent in sequencing (e.g. mapping or base/indel alignment uncertainty), which are usually ignored by other methods or only used for filtering.