Home
/
Authors
/
Felix Schlesinger

Author

Felix Schlesinger

Other affiliations: Cold Spring Harbor Laboratory, Watson School of Biological Sciences

Bio: Felix Schlesinger is an academic researcher from Illumina. The author has contributed to research in topics: Human genome & Gene. The author has an hindex of 15, co-authored 18 publications receiving 30715 citations. Previous affiliations of Felix Schlesinger include Cold Spring Harbor Laboratory & Watson School of Biological Sciences.

Topics: Human genome, Gene, Genome, Gene expression profiling, Paragraph ...read more

Papers

PDF

Open Access

More filters

Journal Article•DOI•

STAR: ultrafast universal RNA-seq aligner

[...]

Alexander Dobin¹, Carrie A. Davis¹, Felix Schlesinger¹, Jorg Drenkow¹, Chris Zaleski¹, Sonali Jha¹, Philippe Batut¹, Mark Chaisson¹, Thomas R. Gingeras¹ - Show less +5 more•Institutions (1)

Cold Spring Harbor Laboratory¹

01 Jan 2013-Bioinformatics

TL;DR: The Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure outperforms other aligners by a factor of >50 in mapping speed.

...read moreread less

Abstract: Motivation Accurate alignment of high-throughput RNA-seq data is a challenging and yet unsolved problem because of the non-contiguous transcript structure, relatively short read lengths and constantly increasing throughput of the sequencing technologies. Currently available RNA-seq aligners suffer from high mapping error rates, low mapping speed, read length limitation and mapping biases. Results To align our large (>80 billon reads) ENCODE Transcriptome RNA-seq dataset, we developed the Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure. STAR outperforms other aligners by a factor of >50 in mapping speed, aligning to the human genome 550 million 2 × 76 bp paired-end reads per hour on a modest 12-core server, while at the same time improving alignment sensitivity and precision. In addition to unbiased de novo detection of canonical junctions, STAR can discover non-canonical splices and chimeric (fusion) transcripts, and is also capable of mapping full-length RNA sequences. Using Roche 454 sequencing of reverse transcription polymerase chain reaction amplicons, we experimentally validated 1960 novel intergenic splice junctions with an 80-90% success rate, corroborating the high precision of the STAR mapping strategy. Availability and implementation STAR is implemented as a standalone C++ code. STAR is free open source software distributed under GPLv3 license and can be downloaded from http://code.google.com/p/rna-star/.

...read moreread less

30,684 citations

Journal Article•DOI•

Landscape of transcription in human cells

[...]

Sarah Djebali, Carrie A. Davis¹, Angelika Merkel, Alexander Dobin¹, Timo Lassmann, Ali Mortazavi², Ali Mortazavi³, Andrea Tanzer, Julien Lagarde, Wei Lin¹, Felix Schlesinger¹, Chenghai Xue¹, Georgi K. Marinov³, Jainab Khatun⁴, Brian A. Williams³, Chris Zaleski¹, Joel Rozowsky⁵, Marion S. Röder, Felix Kokocinski⁶, Rehab F. Abdelhamid, Tyler Alioto, Igor Antoshechkin³, Michael T. Baer¹, Nadav Bar⁷, Philippe Batut¹, Kimberly Bell¹, Ian Bell⁸, Sudipto K. Chakrabortty¹, Xian Chen⁹, Jacqueline Chrast¹⁰, Joao Curado, Thomas Derrien, Jorg Drenkow¹, Erica Dumais⁸, Jacqueline Dumais⁸, Radha Duttagupta⁸, Emilie Falconnet¹¹, Meagan Fastuca¹, Kata Fejes-Toth¹, Pedro G. Ferreira, Sylvain Foissac⁸, Melissa J. Fullwood¹², Hui Gao⁸, David Gonzalez, Assaf Gordon¹, Harsha P. Gunawardena⁹, Cédric Howald¹⁰, Sonali Jha¹, Rory Johnson, Philipp Kapranov⁸, Brandon King³, Colin Kingswood, Oscar Junhong Luo¹², Eddie Park², Kimberly Persaud¹, Jonathan B. Preall¹, Paolo Ribeca, Brian A. Risk⁴, Daniel Robyr¹¹, Michael Sammeth, Lorian Schaffer³, Lei-Hoon See¹, Atif Shahab¹², Jørgen Skancke⁷, Ana Maria Suzuki, Hazuki Takahashi, Hagen Tilgner¹³, Diane Trout³, Nathalie Walters¹⁰, Huaien Wang¹, John A. Wrobel⁴, Yanbao Yu⁹, Xiaoan Ruan¹², Yoshihide Hayashizaki, Jennifer Harrow⁶, Mark Gerstein⁵, Tim Hubbard⁶, Alexandre Reymond¹⁰, Stylianos E. Antonarakis¹¹, Gregory J. Hannon¹, Morgan C. Giddings⁴, Morgan C. Giddings⁹, Yijun Ruan¹², Barbara J. Wold³, Piero Carninci, Roderic Guigó¹⁴, Thomas R. Gingeras⁸, Thomas R. Gingeras¹ - Show less +84 more•Institutions (14)

Cold Spring Harbor Laboratory¹, University of California, Irvine², California Institute of Technology³, Florida State University College of Arts and Sciences⁴, Yale University⁵, Wellcome Trust Sanger Institute⁶, Norwegian University of Science and Technology⁷, Affymetrix⁸, University of North Carolina at Chapel Hill⁹, University of Lausanne¹⁰, University of Geneva¹¹, Genome Institute of Singapore¹², Stanford University¹³, Pompeu Fabra University¹⁴

06 Sep 2012-Nature

TL;DR: Evidence that three-quarters of the human genome is capable of being transcribed is reported, as well as observations about the range and levels of expression, localization, processing fates, regulatory regions and modifications of almost all currently annotated and thousands of previously unannotated RNAs that prompt a redefinition of the concept of a gene.

...read moreread less

Abstract: Eukaryotic cells make many types of primary and processed RNAs that are found either in specific subcellular compartments or throughout the cells. A complete catalogue of these RNAs is not yet available and their characteristic subcellular localizations are also poorly understood. Because RNA represents the direct output of the genetic information encoded by genomes and a significant proportion of a cell's regulatory capabilities are focused on its synthesis, processing, transport, modification and translation, the generation of such a catalogue is crucial for understanding genome function. Here we report evidence that three-quarters of the human genome is capable of being transcribed, as well as observations about the range and levels of expression, localization, processing fates, regulatory regions and modifications of almost all currently annotated and thousands of previously unannotated RNAs. These observations, taken together, prompt a redefinition of the concept of a gene.

...read moreread less

4,450 citations

An integrated encyclopedia of DNA elements in the human genome

[...]

Ian Dunham, Anshul Kundaje, Shelley Force Aldred, Patrick J. Collins +439 more

01 Sep 2012

TL;DR: The Encyclopedia of DNA Elements project provides new insights into the organization and regulation of the authors' genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

2,767 citations

Journal Article•DOI•

A User's Guide to the Encyclopedia of DNA Elements (ENCODE)

[...]

Richard M. Myers, John A. Stamatoyannopoulos¹, Michael Snyder², Ian Dunham +325 more•Institutions (31)

01 Apr 2011-PLOS Biology

TL;DR: An overview of the project and the resources it is generating and the application of ENCODE data to interpret the human genome are provided.

...read moreread less

Abstract: The mission of the Encyclopedia of DNA Elements (ENCODE) Project is to enable the scientific and medical communities to interpret the human genome sequence and apply it to understand human biology and improve health. The ENCODE Consortium is integrating multiple technologies and approaches in a collective effort to discover and define the functional elements encoded in the human genome, including genes, transcripts, and transcriptional regulatory regions, together with their attendant chromatin states and DNA methylation patterns. In the process, standards to ensure high-quality data have been implemented, and novel algorithms have been developed to facilitate analysis. Data and derived results are made available through a freely accessible database. Here we provide an overview of the project and the resources it is generating and illustrate the application of ENCODE data to interpret the human genome.

...read moreread less

1,446 citations

Journal Article•DOI•

Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications

[...]

Xiaoyu Chen¹, Ole Schulz-Trieglaff¹, Richard Shaw¹, Bret Barnes¹, Felix Schlesinger¹, Morten Källberg¹, Anthony J. Cox¹, Semyon Kruglyak¹, Christopher T. Saunders¹ - Show less +5 more•Institutions (1)

Illumina¹

15 Apr 2016-Bioinformatics

TL;DR: Manta is optimized for rapid germline and somatic analysis, calling structural variants, medium-sized indels and large insertions on standard compute hardware in less than a tenth of the time that comparable methods require to identify only subsets of these variant types.

...read moreread less

Abstract: UNLABELLED : We describe Manta, a method to discover structural variants and indels from next generation sequencing data. Manta is optimized for rapid germline and somatic analysis, calling structural variants, medium-sized indels and large insertions on standard compute hardware in less than a tenth of the time that comparable methods require to identify only subsets of these variant types: for example NA12878 at 50× genomic coverage is analyzed in less than 20 min. Manta can discover and score variants based on supporting paired and split-read evidence, with scoring models optimized for germline analysis of diploid individuals and somatic analysis of tumor-normal sample pairs. Call quality is similar to or better than comparable methods, as determined by pedigree consistency of germline calls and comparison of somatic calls to COSMIC database variants. Manta consistently assembles a higher fraction of its calls to base-pair resolution, allowing for improved downstream annotation and analysis of clinical significance. We provide Manta as a community resource to facilitate practical and routine structural variant analysis in clinical and research sequencing scenarios. AVAILABILITY AND IMPLEMENTATION Manta is released under the open-source GPLv3 license. Source code, documentation and Linux binaries are available from https://github.com/Illumina/manta. CONTACT csaunders@illumina.com SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.

...read moreread less

1,224 citations

1
2
3
4
…

Cited by

PDF

Open Access

More filters

Journal Article•DOI•

STAR: ultrafast universal RNA-seq aligner

[...]

Cold Spring Harbor Laboratory¹

01 Jan 2013-Bioinformatics

...read moreread less

30,684 citations

疟原虫var基因转换速率变化导致抗原变异[英]／Paul H, Robert P, Christodoulou Z, et al//Proc Natl Acad Sci U S A

[...]

宁北芳, 朱淮民

28 Jul 2005

TL;DR: PfPMP1）与感染红细胞、树突状组胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作�ly.

...read moreread less

Abstract: 抗原变异可使得多种致病微生物易于逃避宿主免疫应答。表达在感染红细胞表面的恶性疟原虫红细胞表面蛋白1（PfPMP1）与感染红细胞、内皮细胞、树突状细胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作用。每个单倍体基因组var基因家族编码约60种成员，通过启动转录不同的var基因变异体为抗原变异提供了分子基础。

...read moreread less

18,940 citations

Journal Article•DOI•

An integrated encyclopedia of DNA elements in the human genome

[...]

Principal investigators¹, Nhgri groups², Data production leads³, Lead analysts³•Institutions (3)

Wellcome Trust¹, University of Washington², Pennsylvania State University³

06 Sep 2012-Nature

...read moreread less

Abstract: The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

13,548 citations

Journal Article•DOI•

HISAT: a fast spliced aligner with low memory requirements

[...]

Daehwan Kim¹, Ben Langmead¹, Steven L. Salzberg¹•Institutions (1)

Johns Hopkins University School of Medicine¹

01 Apr 2015-Nature Methods

TL;DR: Tests showed that HISAT is the fastest system currently available, with equal or better accuracy than any other method, and requires only 4.3 gigabytes of memory.

...read moreread less

Abstract: HISAT (hierarchical indexing for spliced alignment of transcripts) is a highly efficient system for aligning reads from RNA sequencing experiments. HISAT uses an indexing scheme based on the Burrows-Wheeler transform and the Ferragina-Manzini (FM) index, employing two types of indexes for alignment: a whole-genome FM index to anchor each alignment and numerous local FM indexes for very rapid extensions of these alignments. HISAT's hierarchical index for the human genome contains 48,000 local FM indexes, each representing a genomic region of ∼64,000 bp. Tests on real and simulated data sets showed that HISAT is the fastest system currently available, with equal or better accuracy than any other method. Despite its large number of indexes, HISAT requires only 4.3 gigabytes of memory. HISAT supports genomes of any size, including those larger than 4 billion bases.

...read moreread less

13,192 citations

Journal Article•DOI•

TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions

[...]

Daehwan Kim¹, Daehwan Kim², Geo Pertea³, Cole Trapnell⁴, Cole Trapnell⁵, Harold Pimentel⁶, Kelley Ryan Matthew⁷, Steven L. Salzberg³, Steven L. Salzberg² - Show less +5 more•Institutions (7)

University of Maryland, College Park¹, Johns Hopkins University School of Medicine², Johns Hopkins University³, Harvard University⁴, Broad Institute⁵, University of California, Berkeley⁶, Illumina⁷

25 Apr 2013-Genome Biology

TL;DR: TopHat2 is described, which incorporates many significant enhancements to TopHat, and combines the ability to identify novel splice sites with direct mapping to known transcripts, producing sensitive and accurate alignments, even for highly repetitive genomes or in the presence of pseudogenes.

...read moreread less

Abstract: TopHat is a popular spliced aligner for RNA-sequence (RNA-seq) experiments. In this paper, we describe TopHat2, which incorporates many significant enhancements to TopHat. TopHat2 can align reads of various lengths produced by the latest sequencing technologies, while allowing for variable-length indels with respect to the reference genome. In addition to de novo spliced alignment, TopHat2 can align reads across fusion breaks, which can occur after genomic translocations. TopHat2 combines the ability to identify novel splice sites with direct mapping to known transcripts, producing sensitive and accurate alignments, even for highly repetitive genomes or in the presence of pseudogenes. TopHat2 is available at http://ccb.jhu.edu/software/tophat.

...read moreread less

11,380 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse