Home
/
Authors
/
Jonathan A. Eisen

Author

Jonathan A. Eisen

Other affiliations: Purdue University, Joint Genome Institute, University of Washington ...read more

Bio: Jonathan A. Eisen is an academic researcher from University of California, Davis. The author has contributed to research in topics: Genome & Whole genome sequencing. The author has an hindex of 96, co-authored 493 publications receiving 59611 citations. Previous affiliations of Jonathan A. Eisen include Purdue University & Joint Genome Institute.

Topics: Genome, Whole genome sequencing, Gene, Metagenomics, Genomics ...read more

Papers published on a yearly basis

2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
2000
1999
1998
1997
1995
1994
1993
1992

Papers

PDF

Open Access

More filters

Journal Article•DOI•

Genome sequence of the human malaria parasite Plasmodium falciparum

[...]

Malcolm J. Gardner¹, Neil Hall¹, Eula Fung¹, Owen White¹, Matthew Berriman¹, Richard W. Hyman¹, Jane M. Carlton¹, Arnab Pain¹, Karen E. Nelson¹, Sharen Bowman¹, Ian T. Paulsen¹, Keith D. James¹, Jonathan A. Eisen¹, Kim Rutherford¹, Steven L. Salzberg¹, Alister Craig¹, Sue Kyes¹, Man Suen Chan¹, Vishvanath Nene¹, Shamira J. Shallom¹, Bernard B. Suh¹, Jeremy Peterson¹, Samuel V. Angiuoli¹, Mihaela Pertea¹, Jonathan E. Allen¹, Jeremy D. Selengut¹, Daniel H. Haft¹, Michael W. Mather¹, Akhil B. Vaidya¹, David M. A. Martin¹, Alan H. Fairlamb¹, Martin Fraunholz¹, David S. Roos¹, Stuart A. Ralph¹, Geoffrey I. McFadden¹, Leda M. Cummings¹, G. Mani Subramanian¹, Christopher J. Mungall¹, J. Craig Venter¹, Daniel J. Carucci¹, Stephen L. Hoffman¹, Chris I. Newbold¹, Ronald W. Davis¹, Claire M. Fraser¹, Bart Barrell¹ - Show less +41 more•Institutions (1)

J. Craig Venter Institute¹

03 Oct 2002-Nature

TL;DR: The genome sequence of P. falciparum clone 3D7 is reported, which is the most (A + T)-rich genome sequenced to date and is being exploited in the search for new drugs and vaccines to fight malaria.

...read moreread less

Abstract: The parasite Plasmodium falciparum is responsible for hundreds of millions of cases of malaria, and kills more than one million African children annually. Here we report an analysis of the genome sequence of P. falciparum clone 3D7. The 23-megabase nuclear genome consists of 14 chromosomes, encodes about 5,300 genes, and is the most (A + T)-rich genome sequenced to date. Genes involved in antigenic variation are concentrated in the subtelomeric regions of the chromosomes. Compared to the genomes of free-living eukaryotic microbes, the genome of this intracellular parasite encodes fewer enzymes and transporters, but a large proportion of genes are devoted to immune evasion and host-parasite interactions. Many nuclear-encoded proteins are targeted to the apicoplast, an organelle involved in fatty-acid and isoprenoid metabolism. The genome sequence provides the foundation for future studies of this organism, and is being exploited in the search for new drugs and vaccines to fight malaria.

...read moreread less

4,312 citations

Journal Article•DOI•

Environmental Genome Shotgun Sequencing of the Sargasso Sea

[...]

J. Craig Venter¹, Karin A. Remington¹, John F. Heidelberg², Aaron L. Halpern, Doug Rusch, Jonathan A. Eisen², Dongying Wu², Ian T. Paulsen², Karen E. Nelson², William C. Nelson², Derrick E. Fouts², Samuel Levy, Anthony H. Knap³, Michael W. Lomas³, Kenneth H. Nealson⁴, Owen White², Jeremy Peterson², Jeff Hoffman¹, Rachel Parsons³, Holly Baden-Tillson¹, Cynthia Pfannkoch¹, Yu-Hui Rogers², Hamilton O. Smith¹ - Show less +19 more•Institutions (4)

Alternatives¹, J. Craig Venter Institute², Bermuda Biological Station for Research³, University of Southern California⁴

02 Apr 2004-Science

TL;DR: Over 1.2 million previously unknown genes represented in these samples, including more than 782 new rhodopsin-like photoreceptors are identified, suggesting substantial oceanic microbial diversity.

...read moreread less

Abstract: We have applied “whole-genome shotgun sequencing” to microbial populations collected en masse on tangential flow and impact filters from seawater samples collected from the Sargasso Sea near Bermuda. A total of 1.045 billion base pairs of nonredundant sequence was generated, annotated, and analyzed to elucidate the gene content, diversity, and relative abundance of the organisms within these environmental samples. These data are estimated to derive from at least 1800 genomic species based on sequence relatedness, including 148 previously unknown bacterial phylotypes. We have identified over 1.2 million previously unknown genes represented in these samples, including more than 782 new rhodopsin-like photoreceptors. Variation in species present and stoichiometry suggests substantial oceanic microbial diversity. Microorganisms are responsible for most of the biogeochemical cycles that shape the environment of Earth and its oceans. Yet, these organisms are the least well understood on Earth, as the ability to study and understand the metabolic potential of microorganisms has been hampered by the inability to generate pure cultures. Recent studies have begun to explore environ

...read moreread less

4,210 citations

Journal Article•DOI•

The Sorcerer II Global Ocean Sampling Expedition: Northwest Atlantic through Eastern Tropical Pacific

[...]

Douglas B. Rusch¹, Aaron L. Halpern¹, Granger G. Sutton¹, Karla B. Heidelberg², Karla B. Heidelberg¹, Shannon J. Williamson¹, Shibu Yooseph¹, Dongying Wu¹, Dongying Wu³, Jonathan A. Eisen³, Jonathan A. Eisen¹, Jeff Hoffman¹, Karin A. Remington¹, Karen Beeson¹, Bao Duc Tran¹, Hamilton O. Smith¹, Holly Baden-Tillson¹, Clare Stewart¹, Joyce Thorpe¹, Jason Freeman¹, Cynthia Andrews-Pfannkoch¹, Joseph E. Venter¹, Kelvin Li¹, Saul A. Kravitz¹, John F. Heidelberg¹, John F. Heidelberg², T. Utterback¹, Yu-Hui Rogers¹, Luisa I. Falcón⁴, Valeria Souza⁴, Germán Bonilla-Rosso⁴, Luis E. Eguiarte⁴, David M. Karl⁵, Shubha Sathyendranath⁶, Trevor Platt⁶, Eldredge Bermingham⁷, Victor A. Gallardo⁸, Giselle Tamayo-Castillo⁹, Michael Ferrari¹⁰, Robert L. Strausberg¹, Kenneth H. Nealson², Kenneth H. Nealson¹, Robert Friedman¹, Marvin Frazier¹, J. Craig Venter¹ - Show less +41 more•Institutions (10)

J. Craig Venter Institute¹, University of Southern California², University of California, Davis³, National Autonomous University of Mexico⁴, University of Hawaii⁵, Bedford Institute of Oceanography⁶, Smithsonian Tropical Research Institute⁷, University of Concepción⁸, University of Costa Rica⁹, Rutgers University¹⁰

13 Mar 2007-PLOS Biology

TL;DR: A metagenomic study of the marine planktonic microbiota in which surface (mostly marine) water samples were analyzed as part of the Sorcerer II Global Ocean Sampling expedition, which yielded an extensive dataset consisting of 7.7 million sequencing reads.

...read moreread less

Abstract: The world's oceans contain a complex mixture of micro-organisms that are for the most part, uncharacterized both genetically and biochemically. We report here a metagenomic study of the marine planktonic microbiota in which surface (mostly marine) water samples were analyzed as part of the Sorcerer II Global Ocean Sampling expedition. These samples, collected across a several-thousand km transect from the North Atlantic through the Panama Canal and ending in the South Pacific yielded an extensive dataset consisting of 7.7 million sequencing reads (6.3 billion bp). Though a few major microbial clades dominate the planktonic marine niche, the dataset contains great diversity with 85% of the assembled sequence and 57% of the unassembled data being unique at a 98% sequence identity cutoff. Using the metadata associated with each sample and sequencing library, we developed new comparative genomic and assembly methods. One comparative genomic method, termed "fragment recruitment," addressed questions of genome structure, evolution, and taxonomic or phylogenetic diversity, as well as the biochemical diversity of genes and gene families. A second method, termed "extreme assembly," made possible the assembly and reconstruction of large segments of abundant but clearly nonclonal organisms. Within all abundant populations analyzed, we found extensive intra-ribotype diversity in several forms: (1) extensive sequence variation within orthologous regions throughout a given genome; despite coverage of individual ribotypes approaching 500-fold, most individual sequencing reads are unique; (2) numerous changes in gene content some with direct adaptive implications; and (3) hypervariable genomic islands that are too variable to assemble. The intra-ribotype diversity is organized into genetically isolated populations that have overlapping but independent distributions, implying distinct environmental preference. We present novel methods for measuring the genomic similarity between metagenomic samples and show how they may be grouped into several community types. Specific functional adaptations can be identified both within individual ribotypes and across the entire community, including proteorhodopsin spectral tuning and the presence or absence of the phosphate-binding gene PstS.

...read moreread less

1,982 citations

Journal Article•DOI•

Insights into the phylogeny and coding potential of microbial dark matter

[...]

Christian Rinke¹, Patrick Schwientek¹, Alexander Sczyrba¹, Alexander Sczyrba², Natalia Ivanova¹, Iain Anderson¹, Jan Fang Cheng¹, Aaron E. Darling³, Aaron E. Darling⁴, Stephanie Malfatti¹, Brandon K. Swan⁵, Esther A. Gies⁶, Jeremy A. Dodsworth⁷, Brian P. Hedlund⁷, Georgios Tsiamis⁸, Stefan M. Sievert⁹, Wen Tso Liu¹⁰, Jonathan A. Eisen⁴, Steven J. Hallam⁶, Nikos C. Kyrpides¹, Ramunas Stepanauskas⁵, Edward M. Rubin¹, Philip Hugenholtz¹¹, Tanja Woyke¹ - Show less +20 more•Institutions (11)

Joint Genome Institute¹, Bielefeld University², University of Technology, Sydney³, University of California, Davis⁴, Bigelow Laboratory For Ocean Sciences⁵, University of British Columbia⁶, University of Nevada, Las Vegas⁷, University of Patras⁸, Woods Hole Oceanographic Institution⁹, University of Illinois at Urbana–Champaign¹⁰, University of Queensland¹¹

14 Jul 2013-Nature

TL;DR: This study applies single-cell genomics to target and sequence 201 archaeal and bacterial cells from nine diverse habitats belonging to 29 major mostly uncharted branches of the tree of life and provides a systematic step towards a better understanding of biological evolution on the authors' planet.

...read moreread less

Abstract: Genome sequencing enhances our understanding of the biological world by providing blueprints for the evolutionary and functional diversity that shapes the biosphere. However, microbial genomes that are currently available are of limited phylogenetic breadth, owing to our historical inability to cultivate most microorganisms in the laboratory. We apply single-cell genomics to target and sequence 201 uncultivated archaeal and bacterial cells from nine diverse habitats belonging to 29 major mostly uncharted branches of the tree of life, so-called 'microbial dark matter'. With this additional genomic information, we are able to resolve many intra- and inter-phylum-level relationships and to propose two new superphyla. We uncover unexpected metabolic features that extend our understanding of biology and challenge established boundaries between the three domains of life. These include a novel amino acid use for the opal stop codon, an archaeal-type purine synthesis in Bacteria and complete sigma factors in Archaea similar to those in Bacteria. The single-cell genomes also served to phylogenetically anchor up to 20% of metagenomic reads in some habitats, facilitating organism-level interpretation of ecosystem function. This study greatly expands the genomic representation of the tree of life and provides a systematic step towards a better understanding of biological evolution on our planet.

...read moreread less

1,856 citations

Journal Article•DOI•

DNA sequence of both chromosomes of the cholera pathogen Vibrio cholerae

[...]

John F. Heidelberg, Jonathan A. Eisen, William C. Nelson, Rebecca A. Clayton, Michelle L. Gwinn, Robert J. Dodson, Daniel H. Haft, Erin Hickey, Jeremy Peterson, Lowell Umayam, Steven R. Gill, Karen E. Nelson, Timothy D. Read, Hervé Tettelin, Delwood Richardson, Maria D. Ermolaeva, Jessica Vamathevan, Steven Bass, Haiying Qin, Ioana Dragoi, Patrick Sellers, Lisa McDonald, Teresa Utterback, Robert D. Fleishmann, William C. Nierman, Owen White, Steven L. Salzberg, Hamilton O. Smith¹, Rita R. Colwell², Rita R. Colwell³, John J. Mekalanos⁴, J. Craig Venter¹, Claire M. Fraser - Show less +29 more•Institutions (4)

Celera Corporation¹, University of Maryland Biotechnology Institute², University of Maryland, College Park³, Harvard University⁴

03 Aug 2000-Nature

TL;DR: The V. cholerae genomic sequence provides a starting point for understanding how a free-living, environmental organism emerged to become a significant human bacterial pathogen.

...read moreread less

Abstract: Here we determine the complete genomic sequence of the Gram negative, g-Proteobacterium Vibrio cholerae El Tor N16961 to be 4,033,460 base pairs (bp). The genome consists of two circular chromosomes of 2,961,146 bp and 1,072,314 bp that together encode 3,885 open reading frames. The vast majority of recognizable genes for essential cell functions (such as DNA replication, transcription, translation and cell-wall biosynthesis) and pathogenicity (for example, toxins, surface antigens and adhesins) are located on the large chromosome. In contrast, the small chromosome contains a larger fraction (59%) of hypothetical genes compared with the large chromosome (42%), and also contains many more genes that appear to have origins other than the g-Proteobacteria. The small chromosome also carries a gene capture system (the integron island) and host ‘addiction’ genes that are typically found on plasmids; thus, the small chromosome may have originally been a megaplasmid that was captured by an ancestral Vibrio species. The V. cholerae genomic sequence provides a starting point for understanding how a free-living, environmental organism emerged to become a significant human bacterial pathogen.

...read moreread less

1,785 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102

Collapse

Cited by

PDF

Open Access

More filters

疟原虫var基因转换速率变化导致抗原变异[英]／Paul H, Robert P, Christodoulou Z, et al//Proc Natl Acad Sci U S A

[...]

宁北芳, 朱淮民

28 Jul 2005

TL;DR: PfPMP1）与感染红细胞、树突状组胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作�ly.

...read moreread less

Abstract: 抗原变异可使得多种致病微生物易于逃避宿主免疫应答。表达在感染红细胞表面的恶性疟原虫红细胞表面蛋白1（PfPMP1）与感染红细胞、内皮细胞、树突状细胞以及胎盘的单个或多个受体作用，在黏附及免疫逃避中起关键的作用。每个单倍体基因组var基因家族编码约60种成员，通过启动转录不同的var基因变异体为抗原变异提供了分子基础。

...read moreread less

18,940 citations

Journal Article•DOI•

The Pfam protein families database

[...]

Marco Punta¹, Penny Coggill¹, Ruth Y. Eberhardt¹, Jaina Mistry¹, John Tate¹, Chris Boursnell¹, Ningze Pang¹, Kristoffer Forslund¹, Goran Ceric¹, Jody Clements¹, Andreas Heger¹, Liisa Holm¹, Erik L. L. Sonnhammer¹, Sean R. Eddy¹, Alex Bateman¹, Robert D. Finn¹ - Show less +12 more•Institutions (1)

Wellcome Trust Sanger Institute¹

01 Jan 2000-Nucleic Acids Research

TL;DR: The definition and use of family-specific, manually curated gathering thresholds are explained and some of the features of domains of unknown function (also known as DUFs) are discussed, which constitute a rapidly growing class of families within Pfam.

...read moreread less

Abstract: Pfam is a widely used database of protein families and domains. This article describes a set of major updates that we have implemented in the latest release (version 24.0). The most important change is that we now use HMMER3, the latest version of the popular profile hidden Markov model package. This software is approximately 100 times faster than HMMER2 and is more sensitive due to the routine use of the forward algorithm. The move to HMMER3 has necessitated numerous changes to Pfam that are described in detail. Pfam release 24.0 contains 11,912 families, of which a large number have been significantly updated during the past two years. Pfam is available via servers in the UK (http://pfam.sanger.ac.uk/), the USA (http://pfam.janelia.org/) and Sweden (http://pfam.sbc.su.se/).

...read moreread less

14,075 citations

Journal Article•DOI•

The sequence of the human genome.

[...]

J. Craig Venter¹, Mark Raymond Adams¹, Eugene W. Myers¹, Peter W. Li¹ +269 more•Institutions (12)

16 Feb 2001-Science

TL;DR: Comparative genomic analysis indicates vertebrate expansions of genes associated with neuronal function, with tissue-specific developmental regulation, and with the hemostasis and immune systems are indicated.

...read moreread less

Abstract: A 2.91-billion base pair (bp) consensus sequence of the euchromatic portion of the human genome was generated by the whole-genome shotgun sequencing method. The 14.8-billion bp DNA sequence was generated over 9 months from 27,271,853 high-quality sequence reads (5.11-fold coverage of the genome) from both ends of plasmid clones made from the DNA of five individuals. Two assembly strategies-a whole-genome assembly and a regional chromosome assembly-were used, each combining sequence data from Celera and the publicly funded genome effort. The public data were shredded into 550-bp segments to create a 2.9-fold coverage of those genome regions that had been sequenced, without including biases inherent in the cloning and assembly procedure used by the publicly funded group. This brought the effective coverage in the assemblies to eightfold, reducing the number and size of gaps in the final assembly over what would be obtained with 5.11-fold coverage. The two assembly strategies yielded very similar results that largely agree with independent mapping data. The assemblies effectively cover the euchromatic regions of the human chromosomes. More than 90% of the genome is in scaffold assemblies of 100,000 bp or more, and 25% of the genome is in scaffolds of 10 million bp or larger. Analysis of the genome sequence revealed 26,588 protein-encoding transcripts for which there was strong corroborating evidence and an additional approximately 12,000 computationally derived genes with mouse matches or other weak supporting evidence. Although gene-dense clusters are obvious, almost half the genes are dispersed in low G+C sequence separated by large tracts of apparently noncoding sequence. Only 1.1% of the genome is spanned by exons, whereas 24% is in introns, with 75% of the genome being intergenic DNA. Duplications of segmental blocks, ranging in size up to chromosomal lengths, are abundant throughout the genome and reveal a complex evolutionary history. Comparative genomic analysis indicates vertebrate expansions of genes associated with neuronal function, with tissue-specific developmental regulation, and with the hemostasis and immune systems. DNA sequence comparisons between the consensus sequence and publicly funded genome data provided locations of 2.1 million single-nucleotide polymorphisms (SNPs). A random pair of human haploid genomes differed at a rate of 1 bp per 1250 on average, but there was marked heterogeneity in the level of polymorphism across the genome. Less than 1% of all SNPs resulted in variation in proteins, but the task of determining which SNPs have functional consequences remains an open challenge.

...read moreread less

12,098 citations

Journal Article•DOI•

phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data.

[...]

Paul J. McMurdie¹, Susan Holmes¹•Institutions (1)

Stanford University¹

22 Apr 2013-PLOS ONE

TL;DR: The phyloseq project for R is a new open-source software package dedicated to the object-oriented representation and analysis of microbiome census data in R, which supports importing data from a variety of common formats, as well as many analysis techniques.

...read moreread less

Abstract: Background The analysis of microbial communities through DNA sequencing brings many challenges: the integration of different types of data with methods from ecology, genetics, phylogenetics, multivariate statistics, visualization and testing. With the increased breadth of experimental designs now being pursued, project-specific statistical analyses are often needed, and these analyses are often difficult (or impossible) for peer researchers to independently reproduce. The vast majority of the requisite tools for performing these analyses reproducibly are already implemented in R and its extensions (packages), but with limited support for high throughput microbiome census data. Results Here we describe a software project, phyloseq, dedicated to the object-oriented representation and analysis of microbiome census data in R. It supports importing data from a variety of common formats, as well as many analysis techniques. These include calibration, filtering, subsetting, agglomeration, multi-table comparisons, diversity analysis, parallelized Fast UniFrac, ordination methods, and production of publication-quality graphics; all in a manner that is easy to document, share, and modify. We show how to apply functions from other R packages to phyloseq-represented data, illustrating the availability of a large number of open source analysis techniques. We discuss the use of phyloseq with tools for reproducible research, a practice common in other fields but still rare in the analysis of highly parallel microbiome census data. We have made available all of the materials necessary to completely reproduce the analysis and figures included in this article, an example of best practices for reproducible research. Conclusions The phyloseq project for R is a new open-source software package, freely available on the web from both GitHub and Bioconductor.

...read moreread less

11,272 citations

SPAdes, a new genome assembly algorithm and its applications to single-cell sequencing ( 7th Annual SFAF Meeting, 2012)

[...]

Glenn Tesler

01 Jun 2012

TL;DR: SPAdes as mentioned in this paper is a new assembler for both single-cell and standard (multicell) assembly, and demonstrate that it improves on the recently released E+V-SC assembler and on popular assemblers Velvet and SoapDeNovo (for multicell data).

...read moreread less

Abstract: The lion's share of bacteria in various environments cannot be cloned in the laboratory and thus cannot be sequenced using existing technologies. A major goal of single-cell genomics is to complement gene-centric metagenomic data with whole-genome assemblies of uncultivated organisms. Assembly of single-cell data is challenging because of highly non-uniform read coverage as well as elevated levels of sequencing errors and chimeric reads. We describe SPAdes, a new assembler for both single-cell and standard (multicell) assembly, and demonstrate that it improves on the recently released E+V-SC assembler (specialized for single-cell data) and on popular assemblers Velvet and SoapDeNovo (for multicell data). SPAdes generates single-cell assemblies, providing information about genomes of uncultivatable bacteria that vastly exceeds what may be obtained via traditional metagenomics studies. SPAdes is available online ( http://bioinf.spbau.ru/spades ). It is distributed as open source software.

...read moreread less

10,124 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse