Home
/
Authors
/
Debasish Raha

Author

Debasish Raha

Bio: Debasish Raha is an academic researcher from Yale University. The author has contributed to research in topics: Chromatin immunoprecipitation & Genome. The author has an hindex of 21, co-authored 27 publications receiving 12471 citations. Previous affiliations of Debasish Raha include Stanford University.

Topics: Chromatin immunoprecipitation, Genome, Promoter, Human genome, RNA polymerase II ...read more

Papers

PDF

Open Access

More filters

An integrated encyclopedia of DNA elements in the human genome

[...]

Ian Dunham, Anshul Kundaje, Shelley Force Aldred, Patrick J. Collins +439 more

01 Sep 2012

TL;DR: The Encyclopedia of DNA Elements project provides new insights into the organization and regulation of the authors' genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

2,767 citations

Journal Article•DOI•

The Transcriptional Landscape of the Yeast Genome Defined by RNA Sequencing

[...]

Ugrappa Nagalakshmi¹, Zhong Wang¹, Karl Waern¹, C. Shou¹, Debasish Raha¹, Mark Gerstein¹, Michael Snyder¹ - Show less +3 more•Institutions (1)

Yale University¹

06 Jun 2008-Science

TL;DR: A quantitative sequencing-based method is developed for mapping transcribed regions, in which complementary DNA fragments are subjected to high-throughput sequencing and mapped to the genome, and it is demonstrated that most (74.5%) of the nonrepetitive sequence of the yeast genome is transcribed.

...read moreread less

Abstract: The identification of untranslated regions, introns, and coding regions within an organism remains challenging. We developed a quantitative sequencing-based method called RNA-Seq for mapping transcribed regions, in which complementary DNA fragments are subjected to high-throughput sequencing and mapped to the genome. We applied RNA-Seq to generate a high-resolution transcriptome map of the yeast genome and demonstrated that most (74.5%) of the nonrepetitive sequence of the yeast genome is transcribed. We confirmed many known and predicted introns and demonstrated that others are not actively used. Alternative initiation codons and upstream open reading frames also were identified for many yeast genes. We also found unexpected 3'-end heterogeneity and the presence of many overlapping genes. These results indicate that the yeast transcriptome is more complex than previously appreciated.

...read moreread less

2,506 citations

Journal Article•DOI•

ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia

[...]

Stephen G. Landt¹, Georgi K. Marinov², Anshul Kundaje¹, Pouya Kheradpour³, Florencia Pauli, Serafim Batzoglou¹, Bradley E. Bernstein⁴, Peter J. Bickel⁵, James B. Brown⁵, Philip Cayting¹, Yiwen Chen⁶, Gilberto DeSalvo², Charles B. Epstein⁴, Katherine I. Fisher-Aylor², Ghia Euskirchen¹, Mark Gerstein⁷, Jason Gertz, Alexander J. Hartemink⁸, Michael M. Hoffman⁹, Vishwanath R. Iyer¹⁰, Youngsook L. Jung⁶, Subhradip Karmakar¹¹, Manolis Kellis³, Peter V. Kharchenko¹⁰, Qiang Li¹², Tao Liu⁶, X. Shirley Liu⁶, Lijia Ma¹¹, Aleksandar Milosavljevic¹³, Richard M. Myers, Peter J. Park⁶, Michael J. Pazin¹⁴, Marc D. Perry¹⁵, Debasish Raha⁷, Timothy E. Reddy⁸, Joel Rozowsky⁷, Noam Shoresh⁴, Arend Sidow¹, Matthew Slattery¹¹, John A. Stamatoyannopoulos⁹, Michael Y. Tolstorukov⁶, Kevin P. White¹¹, Simon Xi¹⁶, Peggy J. Farnham¹⁷, Jason D. Lieb¹⁸, Barbara J. Wold², Michael Snyder¹ - Show less +43 more•Institutions (18)

Stanford University¹, California Institute of Technology², Massachusetts Institute of Technology³, Broad Institute⁴, University of California, Berkeley⁵, Harvard University⁶, Yale University⁷, Duke University⁸, University of Washington⁹, University of Texas at Austin¹⁰, University of Chicago¹¹, Pennsylvania State University¹², Baylor College of Medicine¹³, National Institutes of Health¹⁴, Ontario Institute for Cancer Research¹⁵, University of Massachusetts Medical School¹⁶, University of Southern California¹⁷, University of North Carolina at Chapel Hill¹⁸

01 Sep 2012-Genome Research

TL;DR: This work discusses how ChIP quality, assessed in these ways, affects different uses of ChIP-seq data and develops a set of working standards and guidelines for ChIP experiments that are updated routinely.

...read moreread less

Abstract: Chromatin immunoprecipitation (ChIP) followed by high-throughput DNA sequencing (ChIP-seq) has become a valuable and widely used approach for mapping the genomic location of transcription-factor binding and histone modifications in living cells. Despite its widespread use, there are considerable differences in how these experiments are conducted, how the results are scored and evaluated for quality, and how the data and metadata are archived for public use. These practices affect the quality and utility of any global ChIP experiment. Through our experience in performing ChIP-seq experiments, the ENCODE and modENCODE consortia have developed a set of working standards and guidelines for ChIP experiments that are updated routinely. The current guidelines address antibody validation, experimental replication, sequencing depth, data and metadata reporting, and data quality assessment. We discuss how ChIP quality, assessed in these ways, affects different uses of ChIP-seq data. All data sets used in the analysis have been deposited for public viewing and downloading at the ENCODE (http://encodeproject.org/ENCODE/) and modENCODE (http://www.modencode.org/) portals.

...read moreread less

1,801 citations

Journal Article•DOI•

Architecture of the human regulatory network derived from ENCODE data

[...]

Mark Gerstein¹, Anshul Kundaje², Manoj Hariharan², Stephen G. Landt², Koon-Kiu Yan¹, Chao Cheng¹, Xinmeng Jasmine Mu¹, Ekta Khurana¹, Joel Rozowsky¹, Roger P. Alexander¹, Renqiang Min³, Renqiang Min¹, P. Alves¹, Alexej Abyzov¹, Nicholas Addleman², Nitin Bhardwaj¹, Alan P. Boyle², Philip Cayting², Alexandra Charos¹, David Z. Chen¹, Yong Cheng², Declan Clarke¹, Catharine L. Eastman², Ghia Euskirchen², Seth Frietze⁴, Yao Fu¹, Jason Gertz⁵, Fabian Grubert², Arif Harmanci¹, Preti Jain⁵, Maya Kasowski², Phil Lacroute², Jing Jane Leng¹, Jin Lian¹, Hannah Monahan¹, Henriette O'Geen⁶, Zhengqing Ouyang², E. Christopher Partridge⁵, Dorrelyn Patacsil², Florencia Pauli⁵, Debasish Raha¹, Lucía Ramírez², Timothy E. Reddy⁵, Brian Reed¹, Minyi Shi², Teri Slifer², Jing Wang¹, Linfeng Wu², Xinqiong Yang², Kevin Y. Yip¹, Kevin Y. Yip⁷, Gili Zilberman-Schapira¹, Serafim Batzoglou², Arend Sidow², Peggy J. Farnham⁴, Richard M. Myers⁵, Sherman M. Weissman¹, Michael Snyder² - Show less +54 more•Institutions (7)

Yale University¹, Stanford University², Princeton University³, University of Southern California⁴, Joint Genome Institute⁵, University of California, Davis⁶, The Chinese University of Hong Kong⁷

06 Sep 2012-Nature

TL;DR: The combinatorial, co-association of transcription factors is found to be highly context specific: distinct combinations of factors bind at specific genomic locations.

...read moreread less

Abstract: Transcription factors bind in a combinatorial fashion to specify the on-and-off states of genes; the ensemble of these binding events forms a regulatory network, constituting the wiring diagram for a cell. To examine the principles of the human transcriptional regulatory network, we determined the genomic binding information of 119 transcription-related factors in over 450 distinct experiments. We found the combinatorial, co-association of transcription factors to be highly context specific: distinct combinations of factors bind at specific genomic locations. In particular, there are significant differences in the binding proximal and distal to genes. We organized all the transcription factor binding into a hierarchy and integrated it with other genomic information (for example, microRNA regulation), forming a dense meta-network. Factors at different levels have different properties; for instance, top-level transcription factors more strongly influence expression and middle-level ones co-regulate targets to mitigate information-flow bottlenecks. Moreover, these co-regulations give rise to many enriched network motifs (for example, noise-buffering feed-forward loops). Finally, more connected network components are under stronger selection and exhibit a greater degree of allele-specific activity (that is, differential binding to the two parental alleles). The regulatory information obtained in this study will be crucial for interpreting personal genome sequences and understanding basic principles of human biology and disease.

...read moreread less

1,449 citations

Journal Article•DOI•

A User's Guide to the Encyclopedia of DNA Elements (ENCODE)

[...]

Richard M. Myers, John A. Stamatoyannopoulos¹, Michael Snyder², Ian Dunham +325 more•Institutions (31)

01 Apr 2011-PLOS Biology

TL;DR: An overview of the project and the resources it is generating and the application of ENCODE data to interpret the human genome are provided.

...read moreread less

Abstract: The mission of the Encyclopedia of DNA Elements (ENCODE) Project is to enable the scientific and medical communities to interpret the human genome sequence and apply it to understand human biology and improve health. The ENCODE Consortium is integrating multiple technologies and approaches in a collective effort to discover and define the functional elements encoded in the human genome, including genes, transcripts, and transcriptional regulatory regions, together with their attendant chromatin states and DNA methylation patterns. In the process, standards to ensure high-quality data have been implemented, and novel algorithms have been developed to facilitate analysis. Data and derived results are made available through a freely accessible database. Here we provide an overview of the project and the resources it is generating and illustrate the application of ENCODE data to interpret the human genome.

...read moreread less

1,446 citations

1
2
3
4
…
5
6

Collapse

Cited by

PDF

Open Access

More filters

Journal Article•DOI•

STAR: ultrafast universal RNA-seq aligner

[...]

Alexander Dobin¹, Carrie A. Davis¹, Felix Schlesinger¹, Jorg Drenkow¹, Chris Zaleski¹, Sonali Jha¹, Philippe Batut¹, Mark Chaisson¹, Thomas R. Gingeras¹ - Show less +5 more•Institutions (1)

Cold Spring Harbor Laboratory¹

01 Jan 2013-Bioinformatics

TL;DR: The Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure outperforms other aligners by a factor of >50 in mapping speed.

...read moreread less

Abstract: Motivation Accurate alignment of high-throughput RNA-seq data is a challenging and yet unsolved problem because of the non-contiguous transcript structure, relatively short read lengths and constantly increasing throughput of the sequencing technologies. Currently available RNA-seq aligners suffer from high mapping error rates, low mapping speed, read length limitation and mapping biases. Results To align our large (>80 billon reads) ENCODE Transcriptome RNA-seq dataset, we developed the Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure. STAR outperforms other aligners by a factor of >50 in mapping speed, aligning to the human genome 550 million 2 × 76 bp paired-end reads per hour on a modest 12-core server, while at the same time improving alignment sensitivity and precision. In addition to unbiased de novo detection of canonical junctions, STAR can discover non-canonical splices and chimeric (fusion) transcripts, and is also capable of mapping full-length RNA sequences. Using Roche 454 sequencing of reverse transcription polymerase chain reaction amplicons, we experimentally validated 1960 novel intergenic splice junctions with an 80-90% success rate, corroborating the high precision of the STAR mapping strategy. Availability and implementation STAR is implemented as a standalone C++ code. STAR is free open source software distributed under GPLv3 license and can be downloaded from http://code.google.com/p/rna-star/.

...read moreread less

30,684 citations

Journal Article•DOI•

Ultrafast and memory-efficient alignment of short DNA sequences to the human genome

[...]

Ben Langmead¹, Cole Trapnell¹, Mihai Pop¹, Steven L. Salzberg¹•Institutions (1)

University of Maryland, College Park¹

04 Mar 2009-Genome Biology

TL;DR: Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches and can be used simultaneously to achieve even greater alignment speeds.

...read moreread less

Abstract: Bowtie is an ultrafast, memory-efficient alignment program for aligning short DNA sequence reads to large genomes. For the human genome, Burrows-Wheeler indexing allows Bowtie to align more than 25 million reads per CPU hour with a memory footprint of approximately 1.3 gigabytes. Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches. Multiple processor cores can be used simultaneously to achieve even greater alignment speeds. Bowtie is open source http://bowtie.cbcb.umd.edu.

...read moreread less

20,335 citations

疟原虫var基因转换速率变化导致抗原变异[英]／Paul H, Robert P, Christodoulou Z, et al//Proc Natl Acad Sci U S A

[...]

宁北芳, 朱淮民

28 Jul 2005

TL;DR: PfPMP1）与感染红细胞、树突状组胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作�ly.

...read moreread less

Abstract: 抗原变异可使得多种致病微生物易于逃避宿主免疫应答。表达在感染红细胞表面的恶性疟原虫红细胞表面蛋白1（PfPMP1）与感染红细胞、内皮细胞、树突状细胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作用。每个单倍体基因组var基因家族编码约60种成员，通过启动转录不同的var基因变异体为抗原变异提供了分子基础。

...read moreread less

18,940 citations

Journal Article•DOI•

RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome

[...]

Bo Li¹, Colin N. Dewey¹•Institutions (1)

University of Wisconsin-Madison¹

04 Aug 2011-BMC Bioinformatics

TL;DR: It is shown that accurate gene-level abundance estimates are best obtained with large numbers of short single-end reads, and estimates of the relative frequencies of isoforms within single genes may be improved through the use of paired- end reads, depending on the number of possible splice forms for each gene.

...read moreread less

Abstract: RNA-Seq is revolutionizing the way transcript abundances are measured. A key challenge in transcript quantification from RNA-Seq data is the handling of reads that map to multiple genes or isoforms. This issue is particularly important for quantification with de novo transcriptome assemblies in the absence of sequenced genomes, as it is difficult to determine which transcripts are isoforms of the same gene. A second significant issue is the design of RNA-Seq experiments, in terms of the number of reads, read length, and whether reads come from one or both ends of cDNA fragments. We present RSEM, an user-friendly software package for quantifying gene and isoform abundances from single-end or paired-end RNA-Seq data. RSEM outputs abundance estimates, 95% credibility intervals, and visualization files and can also simulate RNA-Seq data. In contrast to other existing tools, the software does not require a reference genome. Thus, in combination with a de novo transcriptome assembler, RSEM enables accurate transcript quantification for species without sequenced genomes. On simulated and real data sets, RSEM has superior or comparable performance to quantification methods that rely on a reference genome. Taking advantage of RSEM's ability to effectively use ambiguously-mapping reads, we show that accurate gene-level abundance estimates are best obtained with large numbers of short single-end reads. On the other hand, estimates of the relative frequencies of isoforms within single genes may be improved through the use of paired-end reads, depending on the number of possible splice forms for each gene. RSEM is an accurate and user-friendly software tool for quantifying transcript abundances from RNA-Seq data. As it does not rely on the existence of a reference genome, it is particularly useful for quantification with de novo transcriptome assemblies. In addition, RSEM has enabled valuable guidance for cost-efficient design of quantification experiments with RNA-Seq, which is currently relatively expensive.

...read moreread less

14,524 citations

Journal Article•DOI•

An integrated encyclopedia of DNA elements in the human genome

[...]

Principal investigators¹, Nhgri groups², Data production leads³, Lead analysts³•Institutions (3)

Wellcome Trust¹, University of Washington², Pennsylvania State University³

06 Sep 2012-Nature

...read moreread less

Abstract: The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

13,548 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse