HomeCategoriesBioinformatics
48 tools

Bioinformatics Tools

Discover our collection of 48 research tools and applications for bioinformatics.

Related Categories

Genomics25
Sequence analysis12
Computational Genomics9
Computational Biology8
Transcriptomics5
Genome assembly5
+54 more

Tools in Bioinformatics

Found 50 of 50 tools

A tool that finds regions of similarity between biological sequences. The program compares nucleotide or protein sequences to sequence databases and calculates the statistical significance.

COSG is a cosine similarity-based method for more accurate and scalable marker gene identification.

GOATools is a Python library for parsing and analyzing Gene Ontology (GO) annotations, enabling statistical analysis of GO term enrichment, visualization of GO hierarchies, and propagation of GO annotations across biological datasets.

A comprehensive package for performing gene set enrichment analysis in Python.

The main purpose of HTSlib is to provide access to genomic information files, both alignment data (SAM, BAM, and CRAM formats) and variant data (VCF and BCF formats). The library also provides interfaces to access and index genome reference data in FASTA format and tab-delimited files with genomic coordinates. It is utilized and incorporated into both SAMtools and BCFtools.

ABACAS is intended to rapidly contiguate (align, order, orientate) , visualize and design primers to close gaps on shotgun assembled contigs based on a reference sequence. It uses MUMmer to find alignment positions and identify syntenies of assembly contigs against the reference. The output is then processed to generate a pseudomolecule taking overlaping contigs and gaps in to account. MUMmer's alignment generating programs, Nucmer and Promer are used followed by the 'delta-filter' utility function. Users could also run tblastx on contigs that are not used to generate the pseudomolecule.

Abismal is a mapper of FASTQ bisulfite-converted short reads (between 50 and 1000 bases) to a FASTA reference genome.

abPOA: an SIMD-based C library for fast partial order alignment using adaptive band. abPOA can perform multiple sequence alignment (MSA) on a set of input sequences and generate a consensus sequence by applying the heaviest bundling algorithm to the final alignment graph.

ACTC (Align subreads to CCS reads) is developed by Pacific Biosciences and provides a one-click solution for aligning individual subreads to the corresponding circular consensus (CCS) reads — useful in workflows involving HiFi/CCS read analysis from PacBio sequencing.

AdapterRemoval searches for and removes adapter sequences from High-Throughput Sequencing (HTS) data and (optionally) trims low quality bases from the 3' end of reads following adapter removal. AdapterRemoval can analyze both single end and paired end data, and can be used to merge overlapping paired-ended reads into (longer) consensus sequences. Additionally, AdapterRemoval can construct a consensus adapter sequence for paired-ended reads, if which this information is not available.

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data AfterQC can simply go through all fastq files in a folder and then output three folders: good, bad and QC folders, which contains good reads, bad reads and the QC results of each fastq file/pair.

Another Gff Analysis Toolkit (AGAT) Suite of tools to handle gene annotations in any GTF/GFF format.

BAM Statistics, Feature Counting and Annotation

Alien_hunter is an application for the prediction of putative Horizontal Gene Transfer (HGT) events with the implementation of Interpolated Variable Order Motifs (IVOMs).

AlignStats produces various alignment, whole genome coverage, and capture coverage metrics for sequence alignment files in SAM, BAM, and CRAM format. This program is designed to serve reporting and quality control purposes in sequencing analysis pipelines at the Baylor College of Medicine Human Genome Sequencing Center (BCM-HGSC).

AMPtk is a series of scripts to process NGS amplicon data using USEARCH and VSEARCH, it can also be used to process any NGS amplicon data and includes databases setup for analysis of fungal ITS, fungal LSU, bacterial 16S, and insect COI amplicons. It can handle Ion Torrent, MiSeq, and 454 data.

AnchorWave (Anchored Wavefront Alignment) identifies collinear regions via conserved anchors (full-length CDS and full-length exon have been implemented currently) and breaks collinear regions into shorter fragments, i.e., anchor and inter-anchor intervals. By performing sensitive sequence alignment for each shorter interval via a 2-piece affine gap cost strategy and merging them together, AnchorWave generates a whole-genome alignment for each collinear block. AnchorWave implements commands to guide collinear block identification with or without chromosomal rearrangements and provides options to use known polyploidy levels or whole-genome duplications to inform alignment.

ANNOgesic is the swiss army knife for RNA-Seq based annotation of bacterial/archaeal genomes. It is a modular, command-line tool that can integrate different types of RNA-Seq data based on dRNA-Seq (differential RNA-Seq) or RNA-Seq protocols that inclusde transcript fragmentation to generate high quality genome annotations. It can detect genes, CDSs/tRNAs/rRNAs, transcription starting sites (TSS) and processing sites, transcripts, terminators, untranslated regions (UTR) as well as small RNAs (sRNA), small open reading frames (sORF), circular RNAs, CRISPR related RNAs, riboswitches and RNA-thermometers. It can also perform RNA-RNA and protein-protein interactions prediction.

Antimicrobial Resistance Identification By Assembly

Somatic copy number analysis using WGS paired end wholegenome sequencing

ASGAL (Alternative Splicing Graph ALigner) is a tool for detecting the alternative splicing events expressed in a RNA-Seq sample with respect to a gene annotation. The main idea behind ASGAL is the following one: the alternative splicing events can be detected by aligning the RNA-Seq reads against the splicing graph of the gene.

Get assembly statistics from FASTA and FASTQ files.

Atropos is tool for specific, sensitive, and speedy trimming of NGS reads.

bcl2fastq is a Linux-based command-line tool from Illumina that converts raw base call (BCL) files from Illumina sequencers into FASTQ format, while simultaneously demultiplexing data based on sample indexes. It is crucial for analyzing sequencing data, requiring a sample sheet and producing FASTQ files, statistics, and reports.

Bioawk is an extension to Brian Kernighan's awk, adding the support of several common biological data formats, including optionally gzip'ed BED, GFF, SAM, VCF, FASTA/Q and TAB-delimited formats with column names. It also adds a few built-in functions and an command line option to use TAB as the input/output delimiter. When the new functionality is not used, bioawk is intended to behave exactly the same as the original BWK awk.

Tools for early stage NGS alignment file processing including fast sorting and duplicate marking.

Breakpoints via assembly - Identifies breaks and attempts to assemble rearrangements in whole genome sequencing data.

Cell Ranger is a set of analysis pipelines for processing Chromium single cell data. It performs barcode processing and UMI counting to quantify gene expression (from 3', 5', and Flex assays), assembles V(D)J immune receptor sequences, and analyzes Feature Barcode data for applications such as cell surface protein analysis and sample multiplexing.

cellranger-atac

Cell Ranger ATAC is a set of automated analysis pipelines developed by 10x Genomics to process and analyze Chromium Single Cell ATAC (Assay for Transposase-Accessible Chromatin) data. It transforms raw sequencing data (FASTQ files) into insights about chromatin accessibility at the single-cell level.

Chromap is an ultrafast method for aligning and preprocessing high throughput chromatin profiles. Typical use cases include: (1) trimming sequencing adapters, mapping bulk ATAC-seq or ChIP-seq genomic reads to the human genome and removing duplicates; (2) trimming sequencing adapters, mapping single cell ATAC-seq genomic reads to the human genome, correcting barcodes, removing duplicates and performing Tn5 shift; (3) split alignment of Hi-C reads against a reference genome. In all these three cases, Chromap is 10-20 times faster while being accurate.

A tool to circularize genome assemblies. Circlator will attempt to identify each circular sequence and output a linearised version of it. It does this by assembling all reads that map to contig ends and comparing the resulting contigs with the input assembly.

CirComPara2 is a computational pipeline to detect, quantify, and correlate expression of linear and circular RNAs from RNA-seq data that combines multiple circRNA-detection methods.

CIRIquant is a comprehensive analysis pipeline for circRNA detection and quantification in RNA-Seq data

Clairvoyante: a multi-task convolutional deep neural network for variant calling in Single Molecule Sequencing.

Control-FREEC is a tool for detection of copy-number changes and allelic imbalances (including LOH) using deep-sequencing data originally developed by the Bioinformatics Laboratory of Institut Curie (Paris). Since 2016, the project has moved to Insitut Cochin, INSERM U1016 (Paris).

A tool for quick quality assessment of cram and bam files, intended for long read sequencing.

Analysis of deep sequencing data for rapid and intuitive interpretation of genome editing experiments. CRISPResso2 can be used to analyze genome editing outcomes using cleaving nucleases (e.g. Cas9 or Cpf1) or noncleaving nucleases (e.g. base editors).

An efficient tool for converting genome coordinates between assemblies. CrossMap supports most of the commonly used file formats, including BAM, sequence alignment map, Wiggle, BigWig, browser extensible data, general feature format, gene transfer format and variant call format.

A cross-platform, efficient and practical CSV/TSV toolkit in Golang. Similar to FASTA/Q format in field of Bioinformatics, CSV/TSV formats are basic and ubiquitous file formats in both Bioinformatics and data science. csvtk is convenient for rapid data investigation and also easy to integrate into analysis pipelines. It could save you lots of time in (not) writing Python/R scripts.

GNU datamash is a command-line program which performs basic numeric, textual and statistical operations on input textual data files.

A genome assembler that reduces the computational time of human genome assembly from 400,000 CPU hours to 2,000 CPU hours, utilizing long erroneous 3GS sequencing reads and short accurate NGS sequencing reads.

DeepConsensus uses gap-aware sequence transformers to correct errors in Pacific Biosciences (PacBio) Circular Consensus Sequencing (CCS) data. This results in greater yield of high-quality reads.

Delly is an integrated structural variant (SV) prediction method that can discover, genotype and visualize deletions, tandem duplications, inversions and translocations at single-nucleotide resolution in short-read and long-read massively parallel sequencing data. It uses paired-ends, split-reads and read-depth to sensitively and accurately delineate genomic rearrangements throughout the genome.

Dnaapler is a simple tool that reorients complete circular microbial genomes.

dnaio is a Python 3 library for very efficient parsing and writing of FASTQ and also FASTA files.

Assemble bacterial isolate genomes from Nanopore reads. Dragonflye is a pipeline that aims to make assembling Oxford Nanopore reads quick and easy.

fastq-scan reads a FASTQ from STDIN and outputs summary statistics (read lengths, per-read qualities, per-base qualities) in JSON format.

Rewrite paired end fastq files to make sure that all reads have a mate and to separate out singletons. This code does one thing: it takes two fastq files, and generates four fastq files. That's right, for free it doubles the number of fastq files that you have!!

Parse multiple Antimicrobial Resistance Analysis Reports into a common data structure

Sinto is a toolkit for processing aligned single-cell data.