Home
/
Authors
/
Mark Gerstein

Author

Mark Gerstein

Other affiliations: Rutgers University, Structural Genomics Consortium, University of Antwerp ...read more

Bio: Mark Gerstein is an academic researcher from Yale University. The author has contributed to research in topics: Genome & Gene. The author has an hindex of 168, co-authored 751 publications receiving 149578 citations. Previous affiliations of Mark Gerstein include Rutgers University & Structural Genomics Consortium.

Topics: Genome, Gene, Human genome, Genomics, Pseudogene ...read more

Papers published on a yearly basis

2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
2000
1999
1998
1997
1996
1995
1994
1993
1992
1991

Papers

PDF

Open Access

More filters

Journal Article•DOI•

Transcriptional landscape of the prenatal human brain

[...]

Jeremy A. Miller¹, Songlin Ding¹, Susan M. Sunkin¹, Kimberly A. Smith¹, Lydia Ng¹, Aaron Szafer¹, Amanda Ebbert¹, Zackery L. Riley¹, Joshua J. Royall¹, Kaylynn Aiona¹, James M. Arnold¹, Crissa Bennet¹, Darren Bertagnolli¹, Krissy Brouner¹, Stephanie Butler¹, Shiella Caldejon¹, Anita Carey¹, Christine Cuhaciyan¹, Rachel A. Dalley¹, Nick Dee¹, Tim A. Dolbeare¹, Benjamin A.C. Facer¹, David Feng¹, Tim P. Fliss¹, Garrett Gee¹, Jeff Goldy¹, Lindsey Gourley¹, Benjamin W. Gregor¹, Guangyu Gu¹, Robert Howard¹, Jayson M. Jochim¹, Chihchau L. Kuan¹, Christopher Lau¹, Changkyu Lee¹, Felix Lee¹, Tracy Lemon¹, Phil Lesnar¹, Bergen McMurray¹, Naveed Mastan¹, Nerick Mosqueda¹, Theresa Naluai-Cecchini², Nhan Kiet Ngo¹, Julie Nyhus¹, Aaron Oldre¹, Eric Olson¹, Jody Parente¹, Patrick D. Parker¹, Sheana Parry¹, Allison Stevens³, Mihovil Pletikos⁴, Melissa Reding¹, Kate Roll¹, David Sandman¹, Melaine Sarreal¹, Sheila Shapouri¹, Nadiya V. Shapovalova¹, Elaine H. Shen¹, Nathan Sjoquist¹, Clifford R. Slaughterbeck¹, Michael W. Smith¹, Andy J. Sodt¹, Derric Williams¹, Lilla Zöllei³, Bruce Fischl⁵, Mark Gerstein⁴, Daniel H. Geschwind⁶, Ian A. Glass², Michael Hawrylycz¹, Robert F. Hevner², Hao Huang⁷, Allan R. Jones¹, James A. Knowles⁸, Pat Levitt⁸, John W. Phillips¹, Nenad Sestan⁴, Paul Wohnoutka¹, Chinh Dang¹, Amy Bernard¹, John G. Hohmann¹, Ed S. Lein¹ - Show less +76 more•Institutions (8)

Allen Institute for Brain Science¹, University of Washington², Harvard University³, Yale University⁴, Massachusetts Institute of Technology⁵, University of California, Los Angeles⁶, University of Texas Southwestern Medical Center⁷, University of Southern California⁸

10 Apr 2014-Nature

TL;DR: An anatomically comprehensive atlas of the mid-gestational human brain is described, including de novo reference atlases, in situ hybridization, ultra-high-resolution magnetic resonance imaging (MRI) and microarray analysis on highly discrete laser-microdissected brain regions.

...read moreread less

Abstract: The anatomical and functional architecture of the human brain is mainly determined by prenatal transcriptional processes. We describe an anatomically comprehensive atlas of the mid-gestational human brain, including de novo reference atlases, in situ hybridization, ultra-high-resolution magnetic resonance imaging (MRI) and microarray analysis on highly discrete laser-microdissected brain regions. In developing cerebral cortex, transcriptional differences are found between different proliferative and post-mitotic layers, wherein laminar signatures reflect cellular composition and developmental processes. Cytoarchitectural differences between human and mouse have molecular correlates, including species differences in gene expression in subplate, although surprisingly we find minimal differences between the inner and outer subventricular zones even though the outer zone is expanded in humans. Both germinal and post-mitotic cortical layers exhibit fronto-temporal gradients, with particular enrichment in the frontal lobe. Finally, many neurodevelopmental disorder and human-evolution-related genes show patterned expression, potentially underlying unique features of human cortical formation. These data provide a rich, freely-accessible resource for understanding human brain development.

...read moreread less

1,114 citations

Journal Article•DOI•

Mapping copy number variation by population-scale genome sequencing

[...]

Ryan E. Mills¹, Klaudia Walter², Chip Stewart³, Robert E. Handsaker⁴ +371 more•Institutions (21)

03 Feb 2011-Nature

TL;DR: A map of unbalanced SVs is constructed based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations, and serves as a resource for sequencing-based association studies.

...read moreread less

Abstract: Genomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.

...read moreread less

1,085 citations

Journal Article•DOI•

Global Identification of Human Transcribed Sequences with Genome Tiling Arrays

[...]

Paul Bertone¹, Viktor Stolc¹, Viktor Stolc², Thomas Royce¹, Joel Rozowsky¹, Alexander E. Urban¹, Xiaowei Zhu¹, John L. Rinn¹, Waraporn Tongprasit, Manoj P. Samanta², Sherman M. Weissman¹, Mark Gerstein¹, Michael Snyder¹ - Show less +9 more•Institutions (2)

Yale University¹, Ames Research Center²

24 Dec 2004-Science

TL;DR: This work constructed a series of high-density oligonucleotide tiling arrays representing sense and antisense strands of the entire nonrepetitive sequence of the human genome and found 10,595 transcribed sequences not detected by other methods.

...read moreread less

Abstract: Elucidating the transcribed regions of the genome constitutes a fundamental aspect of human biology, yet this remains an outstanding problem. To comprehensively identify coding sequences, we constructed a series of high-density oligonucleotide tiling arrays representing sense and antisense strands of the entire nonrepetitive sequence of the human genome. Transcribed sequences were located across the genome via hybridization to complementary DNA samples, reverse-transcribed from polyadenylated RNA obtained from human liver tissue. In addition to identifying many known and predicted genes, we found 10,595 transcribed sequences not detected by other methods. A large fraction of these are located in intergenic regions distal from previously annotated genes and exhibit significant homology to other mammalian proteins.

...read moreread less

1,073 citations

Journal Article•DOI•

Genomic analysis of regulatory network dynamics reveals large topological changes

[...]

Nicholas M. Luscombe¹, M. Madan Babu², Haiyuan Yu¹, Michael Snyder¹, Sarah A. Teichmann², Mark Gerstein¹ - Show less +2 more•Institutions (2)

Yale University¹, Laboratory of Molecular Biology²

16 Sep 2004-Nature

TL;DR: The dynamics of a biological network on a genomic scale is presented, by integrating transcriptional regulatory information and gene-expression data for multiple conditions in Saccharomyces cerevisiae, using an approach for the statistical analysis of network dynamics, called SANDY, combining well-known global topological measures, local motifs and newly derived statistics.

...read moreread less

Abstract: Network analysis has been applied widely, providing a unifying language to describe disparate systems ranging from social interactions to power grids. It has recently been used in molecular biology, but so far the resulting networks have only been analysed statically. Here we present the dynamics of a biological network on a genomic scale, by integrating transcriptional regulatory information and gene-expression data for multiple conditions in Saccharomyces cerevisiae. We develop an approach for the statistical analysis of network dynamics, called SANDY, combining well-known global topological measures, local motifs and newly derived statistics. We uncover large changes in underlying network architecture that are unexpected given current viewpoints and random simulations. In response to diverse stimuli, transcription factors alter their interactions to varying degrees, thereby rewiring the network. A few transcription factors serve as permanent hubs, but most act transiently only during certain conditions. By studying sub-network structures, we show that environmental responses facilitate fast signal propagation (for example, with short regulatory cascades), whereas the cell cycle and sporulation direct temporal progression through multiple stages (for example, with highly inter-connected transcription factors). Indeed, to drive the latter processes forward, phase-specific transcription factors inter-regulate serially, and ubiquitously active transcription factors layer above them in a two-tiered hierarchy. We anticipate that many of the concepts presented here--particularly the large-scale topological changes and hub transience--will apply to other biological networks, including complex sub-systems in higher eukaryotes.

...read moreread less

1,007 citations

Journal Article•DOI•

Expanded encyclopaedias of DNA elements in the human and mouse genomes

[...]

Jill Moore¹, Michael J. Purcaro¹, Henry Pratt¹, Charles B. Epstein², Noam Shoresh², Jessika Adrian³, Trupti Kawli³, Carrie A. Davis⁴, Alexander Dobin⁴, Rajinder Kaul⁵, Jessica Halow, Eric L. Van Nostrand⁶, Peter Freese⁷, David U. Gorkin⁸, David U. Gorkin⁶, Yin Shen⁸, Yin Shen⁹, Yupeng He¹⁰, Mark Mackiewicz, Florencia Pauli-Behn, Brian A. Williams¹¹, Ali Mortazavi¹², Cheryl A. Keller¹³, Xiao-Ou Zhang¹, Shaimae I. Elhajjajy¹, Jack Huey¹, Diane E. Dickel¹⁴, Valentina Snetkova¹⁴, Xintao Wei¹⁵, Xiaofeng Wang¹⁶, Xiaofeng Wang¹⁷, Juan Carlos Rivera-Mulia¹⁸, Juan Carlos Rivera-Mulia¹⁹, Joel Rozowsky²⁰, Jing Zhang²⁰, Surya B. Chhetri²¹, Jialing Zhang²⁰, Alec Victorsen²², Kevin P. White, Axel Visel¹⁴, Axel Visel²³, Gene W. Yeo⁶, Christopher B. Burge⁷, Eric Lécuyer¹⁷, Eric Lécuyer¹⁶, David M. Gilbert¹⁸, Job Dekker¹, John L. Rinn²⁴, Eric M. Mendenhall²¹, Joseph R. Ecker¹⁰, Manolis Kellis⁷, Manolis Kellis², Robert J. Klein²⁵, William Stafford Noble⁵, Anshul Kundaje³, Roderic Guigó²⁶, Peggy J. Farnham²⁷, J. Michael Cherry³, Richard M. Myers, Bing Ren⁶, Bing Ren⁸, Brenton R. Graveley¹⁵, Mark Gerstein²⁰, Len A. Pennacchio²⁸, Len A. Pennacchio¹⁴, Michael Snyder³, Bradley E. Bernstein²⁹, Barbara J. Wold¹¹, Ross C. Hardison¹³, Thomas R. Gingeras⁴, John A. Stamatoyannopoulos⁵, Zhiping Weng³⁰, Zhiping Weng³¹, Zhiping Weng¹ - Show less +70 more•Institutions (31)

University of Massachusetts Medical School¹, Broad Institute², Stanford University³, Cold Spring Harbor Laboratory⁴, University of Washington⁵, University of California, San Diego⁶, Massachusetts Institute of Technology⁷, Ludwig Institute for Cancer Research⁸, University of California, San Francisco⁹, Salk Institute for Biological Studies¹⁰, California Institute of Technology¹¹, University of California, Irvine¹², Pennsylvania State University¹³, Lawrence Berkeley National Laboratory¹⁴, University of Connecticut Health Center¹⁵, McGill University¹⁶, Université de Montréal¹⁷, Florida State University¹⁸, University of Minnesota¹⁹, Yale University²⁰, University of Alabama in Huntsville²¹, University of Chicago²², University of California, Merced²³, University of Colorado Boulder²⁴, Icahn School of Medicine at Mount Sinai²⁵, Pompeu Fabra University²⁶, University of Southern California²⁷, University of California, Berkeley²⁸, Harvard University²⁹, Tongji University³⁰, Boston University³¹

29 Jul 2020-Nature

TL;DR: The authors summarize the data produced by phase III of the Encyclopedia of DNA Elements (ENCODE) project, a resource for better understanding of the human and mouse genomes, which have produced 5,992 new experimental datasets, including systematic determinations across mouse fetal development.

...read moreread less

Abstract: The human and mouse genomes contain instructions that specify RNAs and proteins and govern the timing, magnitude, and cellular context of their production. To better delineate these elements, phase III of the Encyclopedia of DNA Elements (ENCODE) Project has expanded analysis of the cell and tissue repertoires of RNA transcription, chromatin structure and modification, DNA methylation, chromatin looping, and occupancy by transcription factors and RNA-binding proteins. Here we summarize these efforts, which have produced 5,992 new experimental datasets, including systematic determinations across mouse fetal development. All data are available through the ENCODE data portal (https://www.encodeproject.org), including phase II ENCODE1 and Roadmap Epigenomics2 data. We have developed a registry of 926,535 human and 339,815 mouse candidate cis-regulatory elements, covering 7.9 and 3.4% of their respective genomes, by integrating selected datatypes associated with gene regulation, and constructed a web-based server (SCREEN; http://screen.encodeproject.org) to provide flexible, user-defined access to this resource. Collectively, the ENCODE data and registry provide an expansive resource for the scientific community to build a better understanding of the organization and function of the human and mouse genomes.

...read moreread less

999 citations

1
2
3
…
4
5
6
7
8
9
10
…
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161

Collapse

Cited by

PDF

Open Access

More filters

Journal Article•DOI•

Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.

[...]

Stephen F. Altschul¹, Thomas L. Madden, Alejandro A. Schäffer¹, Jinghui Zhang, Zheng Zhang², Webb Miller², David J. Lipman - Show less +3 more•Institutions (2)

National Institutes of Health¹, Pennsylvania State University²

01 Sep 1997-Nucleic Acids Research

TL;DR: A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original.

...read moreread less

Abstract: The BLAST programs are widely used tools for searching protein and DNA databases for sequence similarities. For protein comparisons, a variety of definitional, algorithmic and statistical refinements described here permits the execution time of the BLAST programs to be decreased substantially while enhancing their sensitivity to weak similarities. A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original. In addition, a method is introduced for automatically combining statistically significant alignments produced by BLAST into a position-specific score matrix, and searching the database using this matrix. The resulting Position-Specific Iterated BLAST (PSIBLAST) program runs at approximately the same speed per iteration as gapped BLAST, but in many cases is much more sensitive to weak but biologically relevant sequence similarities. PSI-BLAST is used to uncover several new and interesting members of the BRCT superfamily.

...read moreread less

70,111 citations

Journal Article•DOI•

The Protein Data Bank

[...]

Helen M. Berman¹, John D. Westbrook, Zukang Feng, Gary L. Gilliland, Talapady N. Bhat, Helge Weissig, Ilya N. Shindyalov, Philip E. Bourne - Show less +4 more•Institutions (1)

Rutgers University¹

01 Jan 2000-Nucleic Acids Research

TL;DR: The goals of the PDB are described, the systems in place for data deposition and access, how to obtain further information and plans for the future development of the resource are described.

...read moreread less

Abstract: The Protein Data Bank (PDB; http://www.rcsb.org/pdb/ ) is the single worldwide archive of structural data of biological macromolecules. This paper describes the goals of the PDB, the systems in place for data deposition and access, how to obtain further information, and near-term plans for the future development of the resource.

...read moreread less

34,239 citations

Journal Article•DOI•

STAR: ultrafast universal RNA-seq aligner

[...]

Alexander Dobin¹, Carrie A. Davis¹, Felix Schlesinger¹, Jorg Drenkow¹, Chris Zaleski¹, Sonali Jha¹, Philippe Batut¹, Mark Chaisson¹, Thomas R. Gingeras¹ - Show less +5 more•Institutions (1)

Cold Spring Harbor Laboratory¹

01 Jan 2013-Bioinformatics

TL;DR: The Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure outperforms other aligners by a factor of >50 in mapping speed.

...read moreread less

Abstract: Motivation Accurate alignment of high-throughput RNA-seq data is a challenging and yet unsolved problem because of the non-contiguous transcript structure, relatively short read lengths and constantly increasing throughput of the sequencing technologies. Currently available RNA-seq aligners suffer from high mapping error rates, low mapping speed, read length limitation and mapping biases. Results To align our large (>80 billon reads) ENCODE Transcriptome RNA-seq dataset, we developed the Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure. STAR outperforms other aligners by a factor of >50 in mapping speed, aligning to the human genome 550 million 2 × 76 bp paired-end reads per hour on a modest 12-core server, while at the same time improving alignment sensitivity and precision. In addition to unbiased de novo detection of canonical junctions, STAR can discover non-canonical splices and chimeric (fusion) transcripts, and is also capable of mapping full-length RNA sequences. Using Roche 454 sequencing of reverse transcription polymerase chain reaction amplicons, we experimentally validated 1960 novel intergenic splice junctions with an 80-90% success rate, corroborating the high precision of the STAR mapping strategy. Availability and implementation STAR is implemented as a standalone C++ code. STAR is free open source software distributed under GPLv3 license and can be downloaded from http://code.google.com/p/rna-star/.

...read moreread less

30,684 citations

Journal Article•DOI•

Ultrafast and memory-efficient alignment of short DNA sequences to the human genome

[...]

Ben Langmead¹, Cole Trapnell¹, Mihai Pop¹, Steven L. Salzberg¹•Institutions (1)

University of Maryland, College Park¹

04 Mar 2009-Genome Biology

TL;DR: Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches and can be used simultaneously to achieve even greater alignment speeds.

...read moreread less

Abstract: Bowtie is an ultrafast, memory-efficient alignment program for aligning short DNA sequence reads to large genomes. For the human genome, Burrows-Wheeler indexing allows Bowtie to align more than 25 million reads per CPU hour with a memory footprint of approximately 1.3 gigabytes. Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches. Multiple processor cores can be used simultaneously to achieve even greater alignment speeds. Bowtie is open source http://bowtie.cbcb.umd.edu.

...read moreread less

20,335 citations

疟原虫var基因转换速率变化导致抗原变异[英]／Paul H, Robert P, Christodoulou Z, et al//Proc Natl Acad Sci U S A

[...]

宁北芳, 朱淮民

28 Jul 2005

TL;DR: PfPMP1）与感染红细胞、树突状组胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作�ly.

...read moreread less

Abstract: 抗原变异可使得多种致病微生物易于逃避宿主免疫应答。表达在感染红细胞表面的恶性疟原虫红细胞表面蛋白1（PfPMP1）与感染红细胞、内皮细胞、树突状细胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作用。每个单倍体基因组var基因家族编码约60种成员，通过启动转录不同的var基因变异体为抗原变异提供了分子基础。

...read moreread less

18,940 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse