HomeCategoriesSequence analysis
91 tools

Sequence analysis Tools

Discover our collection of 91 research tools and applications for sequence analysis.

Related Categories

Genomics52
Genome assembly15
Bioinformatics12
Metagenomics12
Comparative genomics11
Transcriptomics9
+69 more

Tools in Sequence analysis

Found 92 of 92 tools

Multi-objective sequence design by inverse folding.

A tool that finds regions of similarity between biological sequences. The program compares nucleotide or protein sequences to sequence databases and calculates the statistical significance.

Provides measures for quantitative assessment of genome assembly, gene set, and transcriptome completeness based on evolutionarily informed expectations of gene content from near-universal single-copy orthologs.

Bowtie 2 is an ultrafast and memory-efficient tool for aligning sequencing reads to long reference sequences. It is particularly good at aligning reads of about 50 up to 100s or 1,000s of characters, and particularly good at aligning to relatively long (e.g. mammalian) genomes. Bowtie 2 indexes the genome with an FM Index to keep its memory footprint small: for the human genome, its memory footprint is typically around 3.2 GB. Bowtie 2 supports gapped, local, and paired-end alignment modes.

Cactus is a reference-free whole-genome multiple alignment program.

F-Seq2 is a Python-based peak caller for high-throughput sequencing data (ChIP-seq, DNase-seq, ATAC-seq) that uses kernel density estimation combined with a local Poisson statistical framework to identify biologically meaningful genomic regions.

Collection of command line tools for Short-Reads FASTA/FASTQ files preprocessing.

Infers approximately-maximum-likelihood phylogenetic trees from alignments of nucleotide or protein sequences.

Filtlong is a tool for filtering long reads by quality. It can take a set of long reads and produce a smaller, better subset. It uses both read length (longer is better) and read identity (higher is better) when choosing which reads pass the filter.

FragGeneScan is an application for finding (fragmented) genes in short reads. It can also be applied to predict prokaryotic genes in incomplete assemblies or complete genomes.

Freyja is a tool to recover relative lineage abundances from mixed SARS-CoV-2 samples from a sequencing dataset (BAM aligned to the Hu-1 reference). The method uses lineage-determining mutational "barcodes" derived from the UShER global phylogenetic tree as a basis set to solve the constrained (unit sum, non-negative) de-mixing problem.

The Genome Analysis Toolkit (GATK) is a set of bioinformatic tools for analyzing high-throughput sequencing (HTS) and variant call format (VCF) data. The toolkit is well established for germline short variant discovery from whole genome and exome sequencing data. GATK4 expands functionality into copy number and somatic analyses and offers pipeline scripts for workflows.

Software aimed at pairwise sequence comparison generating high quality results (equivalent to MUMmer) with controlled memory consumption and comparable or faster execution times particularly with long sequences.

GFAffix identifies walk-preserving shared affixes in variation graphs and collapses them into a non-redundant graph structure.

Genomic Mapping and Alignment Program for mRNA and EST Sequences.

A high-speed next-gen sequencing structural variation caller. It calls variants based on alignment-guided positional de Bruijn graph breakpoint assembly, split read, and read pair evidence.

GangSTR is a tool for genome-wide profiling tandem repeats from short reads. A key advantage of GangSTR over existing genome-wide TR tools is that it can handle repeats that are longer than the read length. GangSTR takes aligned reads (BAM) and a set of repeats in the reference genome as input and outputs a VCF file containing genotypes for each locus.

GenMap is a fast and exact tool for computing genome mappability, calculating the uniqueness of k-mers across genomic positions while allowing for a specified number of mismatches, helping identify unique and repetitive genomic regions.

GenomeScope2 is a reference-free tool that uses k-mer frequency analysis to estimate genome size, heterozygosity, ploidy, and repeat content from raw sequencing data, supporting both diploid and polyploid genomes.

Genomepy is designed to provide a simple and straightforward way to download and use genomic data. This includes (1) searching available data, (2) showing the available metadata, (3) automatically downloading, preprocessing and matching data and (4) generating optional aligner indexes. All with sensible, yet controllable defaults. Currently, genomepy supports Ensembl, UCSC, NCBI and GENCODE.

GenomicConsensus

The GenomicConsensus package provides the variantCaller tool, which allows you to apply the Quiver or Arrow algorithm to mapped PacBio reads to get consensus and variant calls.

HELEN (Homopolymer Encoded Long-read Error-corrector for Nanopore) is a highly optimized genome polishing pipeline that uses deep learning to correct errors in Nanopore long-read assemblies, working in conjunction with MarginPolish and supporting GPU acceleration for scalability.

HMMER is used for searching sequence databases for sequence homologs, and for making sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs).

The main purpose of HTSlib is to provide access to genomic information files, both alignment data (SAM, BAM, and CRAM formats) and variant data (VCF and BCF formats). The library also provides interfaces to access and index genome reference data in FASTA format and tab-delimited files with genomic coordinates. It is utilized and incorporated into both SAMtools and BCFtools.

HTStream is a quality control and processing pipeline for High Throughput Sequencing data. The difference between HTStream and other tools is that HTStream uses a tab delimited fastq format that allows for streaming from application to application. This streaming creates some awesome efficiencies when processing HTS data and makes it fully interoperable with other standard Linux tools.

Homopolish is a method for the removal of systematic errors in nanopore sequencing by homologous polishing.

High-performance visualization tool for interactive exploration of large, integrated datasets. It supports a wide variety of data types and format, including short-read alignments in the SAM/BAM format. Data can be viewed from local files or over the web via http.

A fast and effective stochastic algorithm to infer phylogenetic trees by maximum likelihood. IQ-TREE compares favorably to RAxML and PhyML in terms of likelihoods with similar computing time

ISMapper searches for IS positions in sequence data using paired end Illumina short reads, an IS query/queries of interest and a reference genome. ISMapper reports the IS positions it has found in each isolate, relative to the provided reference genome.

sRNA target prediction.

KMA is mapping a method designed to map raw reads directly against redundant databases, in an ultra-fast manner using seed and extend.

KMC is a utility designed for counting k-mers (sequences of consecutive k symbols) in a set of reads from genome sequencing projects.

KaKs_Calculator2.0

Adopts model selection and model averaging to calculate nonsynonymous (Ka) and synonymous (Ks) substitution rates, attempting to include as many features as needed for accurately capturing evolutionary information in protein-coding sequences. In addition, several existing methods for calculating Ka and Ks are also incorporated into KaKs_Calculator.

Program for the taxonomic assignment of high-throughput sequencing reads, e.g., Illumina or Roche/454, from whole-genome sequencing of metagenomic DNA. Reads are directly assigned to taxa using the NCBI taxonomy and a reference database of protein sequences from Bacteria, Archaea, Fungi, microbial eukaryotes and viruses.

KentUtils is a collection of UCSC command-line bioinformatic utilities for genome browser data processing, including tools for format conversion, genome annotation, sequence analysis, and manipulation of genomic data formats such as BED, BigWig, BigBed, and BAM files.

KisSplice

KisSplice is a software that enables to analyse RNA-seq data with or without a reference genome.

KmerGenie

KmerGenie estimates the best k-mer length for genome de novo assembly. Given a set of reads, KmerGenie first computes the k-mer abundance histogram for many values of k. Then, for each value of k, it predicts the number of distinct genomic k-mers in the dataset, and returns the k-mer length which maximizes this number. Experiments show that KmerGenie's choices lead to assemblies that are close to the best possible over all k-mer lengths.

KneadData is a tool designed to perform quality control on metagenomic and metatranscriptomic sequencing data, especially data from microbiome experiments. In these experiments, samples are typically taken from a host in hopes of learning something about the microbial community on the host. However, sequencing data from such experiments will often contain a high ratio of host to bacterial reads. This tool aims to perform principled in silicoseparation of bacterial reads from these "contaminant" reads, be they from the host, from bacterial 16S sequences, or other user-defined sources.

Kraken 2 is the newest version of Kraken, a taxonomic classification system using exact k-mer matches to achieve high accuracy and fast classification speeds. This classifier matches each k-mer within a query sequence to the lowest common ancestor (LCA) of all genomes containing the given k-mer. The k-mer assignments inform the classification algorithm. Any assumption that Kraken’s raw read assignments can be directly translated into species or strain-level abundance estimates is flawed. Bracken (Bayesian Reestimation of Abundance after Classification with KrakEN), estimates species abundances in metagenomics samples by probabilistically re-distributing reads in the taxonomic tree. (Lu, Jennifer et al. “Bracken: estimating species abundance in metagenomics data.”)

KrakenTools is a suite of scripts to be used alongside the Kraken, KrakenUniq, Kraken 2, or Bracken programs. These scripts are designed to help Kraken users with downstream analysis of Kraken results.

LAST is a sequence alignment tool that finds and aligns related regions between large biological sequences, supporting DNA-DNA, DNA-protein, and protein-protein comparisons with sensitivity comparable to BLAST but with greater speed and flexibility for large datasets.

The Long Read Aligner for Sequences and Contigs. LRA, the long read aligner for sequences and assembly contigs LRA is a sequence alignment program that aligns long reads from single-molecule sequencing (SMS) instruments, or megabase-scale contigs from SMS assemblies. LRA implements seed chaining sparse dynamic programming with a convex gap function to read and assembly alignment, which is also extended to allow for inversion cases. Through the Truvari analysis of LRA, Minimap2 and NGM-LR alignments. LRA achieves higher f1 score over HG002 HiFi, CLR and ONT datasets. Home: https://github.com/ChaissonLab/LRA. Long read aligner for sequences and contigs.

Lambda is a local aligner optimized for many query sequences and searches in protein space. It is compatible to BLAST, but much faster than BLAST and many other comparable tools.

LiftoffTools is a toolkit for comparing gene annotations mapped between genome assemblies, enabling the detection and analysis of gene sequence variants, synteny, and gene copy number changes. It leverages Liftoff for annotation transfer and offers modules for analyzing protein-coding genes, gene synteny, and gene copy number.

LongPhase is an ultra-fast program for simultaneously co-phasing SNPs, small indels, large SVs, and (5mC) modifications for Nanopore and PacBio platforms. It can produce nearly chromosome-scale haplotype blocks by using Nanpore ultra-long reads without the need for additional trios, chromosome conformation, and strand-seq data. LongPhase can phase a 30x human genome in ~1 minute

LongQC is a tool for the data quality control of the PacBio and ONT long reads, and it has two functionalities: sample qc and platform qc.

MAFFT (Multiple Alignment using Fast Fourier Transform) is a high speed multiple sequence alignment program.

This tool performs multiple sequence alignments of nucleotide or amino acid sequences.

Portable and easily configurable genome annotation pipeline. It’s purpose is to allow smaller eukaryotic and prokaryotic genome projects to independently annotate their genomes and to create genome databases.

Nextclade is an open-source project for viral genome alignment, mutation calling, clade assignment, quality checks and phylogenetic placement.

PACU is a workflow for whole genome sequencing based phylogeny of Illumina and ONT R9/R10 data. PACU stands for the Prokaryotic Awesome variant Calling Utility and is named after an omnivorous fish (that eats both Illumina and ONT reads).

A tool for Phylogenetic Analysis and Post-Analysis of Large Phylogenies.

A program that screens DNA sequences for interspersed repeats and low complexity DNA sequences. The output of the program is a detailed annotation of the repeats that are present in the query sequence as well as a modified version of the query sequence in which all the annotated repeats have been masked (default: replaced by Ns).

SAMtools are widely used for processing and analysing high-throughput sequencing data. They include tools for file format conversion and manipulation, sorting, querying, statistics, variant calling, and effect analysis amongst other methods.

A multiple sequence alignment package that can be used for DNA, RNA and protein sequences. It can be used to align sequences or to combine the output of other alignment methods (Clustal, Mafft, Probcons, Muscle...) into one unique alignment.

Variant tool set that discovers short variants from Next Generation Sequencing data.

abPOA: an SIMD-based C library for fast partial order alignment using adaptive band. abPOA can perform multiple sequence alignment (MSA) on a set of input sequences and generate a consensus sequence by applying the heaviest bundling algorithm to the final alignment graph.

adVNTR is a tool for genotyping Variable Number Tandem Repeats (VNTR) from sequence data. It works with both NGS short reads (Illumina HiSeq) and SMRT reads (PacBio) and finds diploid repeating counts for VNTRs and identifies possible mutations in the VNTR sequences.

This is a tool to plot allele frequencies in VCF files.

BAM Statistics, Feature Counting and Annotation

Get assembly statistics from FASTA and FASTQ files.

Markov chain Monte Carlo software for simultaneous Bayesian estimation of alignment and phylogeny (and other parameters). It handles generic Bayesian modeling via probabilistic programming.

BamTools provides a fast, flexible C++ API & toolkit for reading, writing, and managing BAM files.

Bioawk is an extension to Brian Kernighan's awk, adding the support of several common biological data formats, including optionally gzip'ed BED, GFF, SAM, VCF, FASTA/Q and TAB-delimited formats with column names. It also adds a few built-in functions and an command line option to use TAB as the input/output delimiter. When the new functionality is not used, bioawk is intended to behave exactly the same as the original BWK awk.

breseq is a computational pipeline for finding mutations relative to a reference sequence using high-throughput DNA resequencing data. It is intended for haploid microbial genomes (<20 Mb). breseq is a command line tool implemented in C++ and R.

Chromap is an ultrafast method for aligning and preprocessing high throughput chromatin profiles. Typical use cases include: (1) trimming sequencing adapters, mapping bulk ATAC-seq or ChIP-seq genomic reads to the human genome and removing duplicates; (2) trimming sequencing adapters, mapping single cell ATAC-seq genomic reads to the human genome, correcting barcodes, removing duplicates and performing Tn5 shift; (3) split alignment of Hi-C reads against a reference genome. In all these three cases, Chromap is 10-20 times faster while being accurate.

Suite of tools to discover structural variations such as (larger) insertions and deletions in genomes from paired-end sequencing reads.

ClonalFrameML is a maximum likelihood implementation of the Bayesian software ClonalFrame which was previously described by Didelot and Falush (2007). The recombination model underpinning ClonalFrameML is exactly the same as for ClonalFrame, but this new implementation is a lot faster, is able to deal with much larger genomic dataset, and does not suffer from MCMC convergence issues

CoverM aims to be a configurable, easy to use and fast DNA read coverage and relative abundance calculator focused on metagenomics applications.

Library and software package for fast parsing and querying of VCF and BCF files and illustrate its speed, simplicity and utility.

A genome assembler that reduces the computational time of human genome assembly from 400,000 CPU hours to 2,000 CPU hours, utilizing long erroneous 3GS sequencing reads and short accurate NGS sequencing reads.

DeepConsensus uses gap-aware sequence transformers to correct errors in Pacific Biosciences (PacBio) Circular Consensus Sequencing (CCS) data. This results in greater yield of high-quality reads.

Delly is an integrated structural variant (SV) prediction method that can discover, genotype and visualize deletions, tandem duplications, inversions and translocations at single-nucleotide resolution in short-read and long-read massively parallel sequencing data. It uses paired-ends, split-reads and read-depth to sensitively and accurately delineate genomic rearrangements throughout the genome.

Sequence aligner for protein and translated DNA searches and functions as a drop-in replacement for the NCBI BLAST software tools. It is suitable for protein-protein search as well as DNA-protein search on short reads and longer sequences including contigs and assemblies, providing a speedup of BLAST ranging up to x20,000.

Dnaapler is a simple tool that reorients complete circular microbial genomes.

Assemble bacterial isolate genomes from Nanopore reads. Dragonflye is a pipeline that aims to make assembling Oxford Nanopore reads quick and easy.

Fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication.

Diverse suite of tools for sequence analysis; many programs analagous to GCG; context-sensitive help for each tool.

A tool for pairwise sequence alignment. It enables alignment for DNA-DNA and DNA-protein pairs and also gapped and ungapped alignment.

A tool designed to provide fast all-in-one preprocessing for FastQ files. This tool is developed in C++ with multithreading supported to afford high performance.

Command line utility for manipulating FASTQ files

gfatools is a set of tools for manipulating sequence graphs in the GFA or the rGFA format. It has implemented parsing, subgraph and conversion to FASTA/BED.

Program for comparing, annotating, merging and tracking transcripts in GFF files.

Program for filtering, converting and manipulating GFF files

Python package for working with GFF and GTF files. It allows operations which would be complicated or time-consuming using a text-file-only approach.

iVar is a computational package that contains functions broadly useful for viral amplicon-based sequencing.

Interactive assembly and analysis of RADseq datasets. ipyrad: interactive assembly and analysis of RAD-seq data sets. Welcome to ipyrad, an interactive toolkit for assembly and analysis of restriction-site associated genomic data sets (e.g., RAD, ddRAD, GBS) for population genetic and phylogenetic studies. Welcome to ipyrad — ipyrad documentation.

A program for quantifying abundances of transcripts from RNA-Seq data, or more generally of target sequences using high-throughput sequencing reads. It is based on the novel idea of pseudoalignment for rapidly determining the compatibility of reads with targets, without the need for alignment.

khmer is a set of command-line tools for working with DNA shotgun sequencing data from genomes, transcriptomes, metagenomes, and single cells. khmer can make de novo assemblies faster, and sometimes better. khmer can also identify (and fix) problems with shotgun data.

A command-line algorithm for counting k-mers in DNA sequence.

lima is the standard tool to identify barcode and primer sequences in PacBio single-molecule sequencing data. It powers the Demultiplex Barcodes GUI-based analysis applications.

Tool for the automated removal of spurious sequences or poorly aligned regions from a multiple sequence alignment.

    Sequence analysis Tools - bundlecore