scispace - formally typeset
Open AccessJournal ArticleDOI

HTSeq—a Python framework to work with high-throughput sequencing data

Simon Anders, +2 more
- 15 Jan 2015 - 
- Vol. 31, Iss: 2, pp 166-169
TLDR
This work presents HTSeq, a Python library to facilitate the rapid development of custom scripts for high-throughput sequencing data analysis, and presents htseq-count, a tool developed with HTSequ that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes.
Abstract
Motivation: A large choice of tools exists for many standard tasks in the analysis of high-throughput sequencing (HTS) data. However, once a project deviates from standard workflows, custom scripts are needed. Results: We present HTSeq, a Python library to facilitate the rapid development of such scripts. HTSeq offers parsers for many common data formats in HTS projects, as well as classes to represent data, such as genomic coordinates, sequences, sequencing reads, alignments, gene model information and variant calls, and provides data structures that allow for querying via genomic coordinates. We also present htseq-count, a tool developed with HTSeq that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes. Availability and implementation: HTSeq is released as an opensource software under the GNU General Public Licence and available from http://www-huber.embl.de/HTSeq or from the Python Package Index at https://pypi.python.org/pypi/HTSeq. Contact: sanders@fs.tum.de

read more

Content maybe subject to copyright    Report

Citations
More filters
Journal ArticleDOI

A Tunable Mechanism Determines the Duration of the Transgenerational Small RNA Inheritance in C. elegans

TL;DR: RNA-sequencing analysis reveals that, aside from silencing of genes with complementary sequences, dsRNA-induced RNAi affects the production of heritable endogenous small RNAs, which regulate the expression of RNAi factors.
Journal ArticleDOI

Lineage-specific rediploidization is a mechanism to explain time-lags between genome duplication and evolutionary diversification

TL;DR: Under LORe, which is predicted following many WGD events, the functional outcomes of WGD need not appear ‘explosively’, but can arise gradually over tens of millions of years, promoting lineage-specific diversification regimes under prevailing ecological pressures.
Journal ArticleDOI

The draft genome of tropical fruit durian ( Durio zibethinus )

TL;DR: The durian genome provides a resource for tropical fruit biology and agronomy and a potential association between ethylene biosynthesis and methionine regeneration via the Yang cycle is suggested.
Journal ArticleDOI

Single-cell RNA-Sequencing uncovers transcriptional states and fate decisions in haematopoiesis

TL;DR: A marker-free approach is used to computationally reconstruct the blood lineage tree in zebrafish and order cells along their differentiation trajectory, based on their global transcriptional differences, and a refined model of developmental progression of haematopoietic cells is proposed.
Journal ArticleDOI

Tracing the expression of circular RNAs in human pre-implantation embryos

TL;DR: This study reports the first analysis of the whole transcriptome comprising both polyA+ mRNAs and polyA– RNAs including circ RNAs during human pre-implantation development, providing an invaluable resource for analyzing the unique function and complex regulatory mechanisms of circRNAs during this process.
References
More filters
Journal ArticleDOI

Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2

TL;DR: This work presents DESeq2, a method for differential analysis of count data, using shrinkage estimation for dispersions and fold changes to improve stability and interpretability of estimates, which enables a more quantitative analysis focused on the strength rather than the mere presence of differential expression.
Journal ArticleDOI

The Sequence Alignment/Map format and SAMtools

TL;DR: SAMtools as discussed by the authors implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments.
Journal ArticleDOI

Trimmomatic: a flexible trimmer for Illumina sequence data

TL;DR: Timmomatic is developed as a more flexible and efficient preprocessing tool, which could correctly handle paired-end data and is shown to produce output that is at least competitive with, and in many cases superior to, that produced by other tools, in all scenarios tested.
Journal ArticleDOI

edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.

TL;DR: EdgeR as mentioned in this paper is a Bioconductor software package for examining differential expression of replicated count data, which uses an overdispersed Poisson model to account for both biological and technical variability and empirical Bayes methods are used to moderate the degree of overdispersion across transcripts, improving the reliability of inference.
Journal ArticleDOI

BEDTools: a flexible suite of utilities for comparing genomic features

TL;DR: A new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format, which allows the user to compare large datasets (e.g. next-generation sequencing data) with both public and custom genome annotation tracks.
Related Papers (5)