A global reference for human genetic variation.

doi:10.1038/NATURE15393

Home
/
Papers
/
A global reference for human genetic variation.

Journal Article•DOI•

A global reference for human genetic variation.

Adam Auton¹, Gonçalo R. Abecasis², David Altshuler³, Richard Durbin⁴ +514 more•Institutions (90)

01 Oct 2015-Nature (Nature Publishing Group)-Vol. 526, Iss: 7571, pp 68-74

TL;DR: The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations, and has reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-generation sequencing, deep exome sequencing, and dense microarray genotyping.

read less

Abstract: The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes >99% of SNP variants with a frequency of >1% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies.

...read moreread less

Citations

PDF

Open Access

More filters

Journal Article•DOI•

Physiological and Genetic Adaptations to Diving in Sea Nomads

[...]

Melissa Ilardo¹, Ida Moltke¹, Thorfinn Sand Korneliussen², Thorfinn Sand Korneliussen¹, Jade Cheng³, Aaron J. Stern³, Fernando Racimo¹, Peter de Barros Damgaard¹, Martin Sikora¹, Andaine Seguin-Orlando¹, Simon Rasmussen⁴, Inge C.L. van den Munckhof⁵, Rob ter Horst⁵, Leo A. B. Joosten⁵, Mihai G. Netea⁵, Mihai G. Netea⁶, Suhartini Salingkat, Rasmus Nielsen¹, Rasmus Nielsen³, Eske Willerslev¹, Eske Willerslev⁷, Eske Willerslev² - Show less +18 more•Institutions (7)

University of Copenhagen¹, University of Cambridge², University of California, Berkeley³, Technical University of Denmark⁴, Radboud University Nijmegen⁵, University of Bonn⁶, Wellcome Trust Sanger Institute⁷

19 Apr 2018-Cell

TL;DR: It is shown that natural selection on genetic variants in the PDE10A gene have increased spleen size in the Bajau, providing them with a larger reservoir of oxygenated red blood cells and evidence of strong selection specific to the Bjau on BDKRB2, a gene affecting the human diving reflex.

...read moreread less

116 citations

Cites methods from "A global reference for human geneti..."

...We therefore merged our sequencing data from the Bajau and Saluan with Han Chinese genomes from the 1000 Genomes Project (Auton et al., 2015) and performed a genome-wide selection scan using a new method for detecting local selection, akin to the PBS statistic (Yi et al....
[...]
...We therefore merged our sequencing data from the Bajau and Saluan with Han Chinese genomes from the 1000 Genomes Project (Auton et al., 2015) and performed a genome-wide selection scan using a new method for detecting local selection, akin to the PBS statistic (Yi et al., 2010) but based on an explicit likelihood model and adjusted to account for admixture and differing ancestral components (Cheng et al., 2016)....
[...]
...We therefore merged our sequencing data from the Bajau and Saluan with Han Chinese genomes from the 1000 Genomes Project (Auton et al., 2015) and performed a genome-wide selection scan using a new method for detecting local selection, akin to the PBS statistic (Yi et al., 2010) but based on an…...
[...]

Journal Article•DOI•

Burden Testing of Rare Variants Identified through Exome Sequencing via Publicly Available Control Data

[...]

Michael H. Guo¹, Michael H. Guo², Michael H. Guo³, Lacey Plummer², Yee-Ming Chan³, Joel N. Hirschhorn³, Joel N. Hirschhorn², Joel N. Hirschhorn¹, Margaret F. Lippincott² - Show less +5 more•Institutions (3)

Broad Institute¹, Harvard University², Boston Children's Hospital³

04 Oct 2018-American Journal of Human Genetics

TL;DR: The approach "re-discovered" genes previously implicated in IHH and introduced an approach for highly adaptable variant quality filtering that leads to well-calibrated results, and developed a user-friendly software package for performing gene-based burden testing against public databases.

...read moreread less

Abstract: The genetic causes of many Mendelian disorders remain undefined. Factors such as lack of large multiplex families, locus heterogeneity, and incomplete penetrance hamper these efforts for many disorders. Previous work suggests that gene-based burden testing—where the aggregate burden of rare, protein-altering variants in each gene is compared between case and control subjects—might overcome some of these limitations. The increasing availability of large-scale public sequencing databases such as Genome Aggregation Database (gnomAD) can enable burden testing using these databases as controls, obviating the need for additional control sequencing for each study. However, there exist various challenges with using public databases as controls, including lack of individual-level data, differences in ancestry, and differences in sequencing platforms and data processing. To illustrate the approach of using public data as controls, we analyzed whole-exome sequencing data from 393 individuals with idiopathic hypogonadotropic hypogonadism (IHH), a rare disorder with significant locus heterogeneity and incomplete penetrance against control subjects from gnomAD (n = 123,136). We leveraged presumably benign synonymous variants to calibrate our approach. Through iterative analyses, we systematically addressed and overcame various sources of artifact that can arise when using public control data. In particular, we introduce an approach for highly adaptable variant quality filtering that leads to well-calibrated results. Our approach “re-discovered” genes previously implicated in IHH (FGFR1, TACR3, GNRHR). Furthermore, we identified a significant burden in TYRO3, a gene implicated in hypogonadotropic hypogonadism in mice. Finally, we developed a user-friendly software package TRAPD (Test Rare vAriants with Public Data) for performing gene-based burden testing against public databases.

...read moreread less

116 citations

Journal Article•DOI•

Complete genomic and epigenetic maps of human centromeres

[...]

01 Apr 2022-Science

TL;DR: In this paper , a complete, telomere-to-telomere human genome assembly (T2T-CHM13) has enabled the comprehensively characterize pericentromeric and centromeric repeats, which constitute 6.2% of the genome.

...read moreread less

Abstract: Existing human genome assemblies have almost entirely excluded repetitive sequences within and near centromeres, limiting our understanding of their organization, evolution, and functions, which include facilitating proper chromosome segregation. Now, a complete, telomere-to-telomere human genome assembly (T2T-CHM13) has enabled us to comprehensively characterize pericentromeric and centromeric repeats, which constitute 6.2% of the genome (189.9 megabases). Detailed maps of these regions revealed multimegabase structural rearrangements, including in active centromeric repeat arrays. Analysis of centromere-associated sequences uncovered a strong relationship between the position of the centromere and the evolution of the surrounding DNA through layered repeat expansions. Furthermore, comparisons of chromosome X centromeres across a diverse panel of individuals illuminated high degrees of structural, epigenetic, and sequence variation in these complex and rapidly evolving regions.

...read moreread less

116 citations

Posted Content•DOI•

Genomic Dissection of Bipolar Disorder and Schizophrenia, Including 28 Subphenotypes

[...]

Jo Knight¹, Jo Knight²•Institutions (2)

University of Toronto¹, Centre for Addiction and Mental Health²

14 Jun 2018-bioRxiv

TL;DR: For the first time, specific loci pointing to a potential role of 4 genes (DARS2, ARFGEF2, DCAKD and GATAD2A) that distinguish between BD and SCZ are identified, providing an opportunity to understand the biology contributing to clinical differences of these disorders.

...read moreread less

Abstract: Schizophrenia and bipolar disorder are two distinct diagnoses that share symptomology. Understanding the genetic factors contributing to the shared and disorder-specific symptoms will be crucial for improving diagnosis and treatment. In genetic data consisting of 53,555 cases (20,129 bipolar disorder [BD], 33,426 schizophrenia [SCZ]) and 54,065 controls, we identified 114 genome-wide significant loci implicating synaptic and neuronal pathways shared between disorders. Comparing SCZ to BD (23,585 SCZ, 15,270 BD) identified four genomic regions including one with disorder-independent causal variants and potassium ion response genes as contributing to differences in biology between the disorders. Polygenic risk score (PRS) analyses identified several significant correlations within case-only phenotypes including SCZ PRS with psychotic features and age of onset in BD. For the first time, we discover specific loci that distinguish between BD and SCZ and identify polygenic components underlying multiple symptom dimensions. These results point to the utility of genetics to inform symptomology and potential treatment.

...read moreread less

116 citations

Posted Content•DOI•

Trans-ancestry genetic study of type 2 diabetes highlights the power of diverse populations for discovery and translation

[...]

Anubha Mahajan¹, Cassandra N. Spracklen², Weihua Zhang³, Maggie C.Y. Ng⁴ +242 more•Institutions (95)

23 Sep 2020-medRxiv

TL;DR: Improved fine-mapping enabled systematic assessment of candidate causal genes and molecular mechanisms through which T2D associations are mediated, laying foundations for functional investigations.

...read moreread less

Abstract: We assembled an ancestrally diverse collection of genome-wide association studies of type 2 diabetes (T2D) in 180,834 cases and 1,159,055 controls (48.9% non-European descent). We identified 277 loci at genome-wide significance (p 50% posterior probability. This improved fine-mapping enabled systematic assessment of candidate causal genes and molecular mechanisms through which T2D associations are mediated, laying foundations for functional investigations. Trans-ancestry genetic risk scores enhanced transferability across diverse populations, providing a step towards more effective clinical translation to improve global health.

...read moreread less

115 citations

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
…
113
114
115
116
117
118
119
…
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse

References

PDF

Open Access

More filters

Journal Article•DOI•

Basic Local Alignment Search Tool

[...]

Stephen F. Altschul¹, Warren Gish¹, Webb Miller², Eugene W. Myers³, David J. Lipman¹ - Show less +1 more•Institutions (3)

National Institutes of Health¹, Pennsylvania State University², University of Arizona³

01 Oct 1990-Journal of Molecular Biology

TL;DR: A new approach to rapid sequence comparison, basic local alignment search tool (BLAST), directly approximates alignments that optimize a measure of local similarity, the maximal segment pair (MSP) score.

...read moreread less

88,255 citations

Journal Article•DOI•

The Sequence Alignment/Map format and SAMtools

[...]

Heng Li¹, Bob Handsaker², Alec Wysoker², T. J. Fennell², Jue Ruan³, Nils Homer², Gabor T. Marth⁴, Gonçalo R. Abecasis², Richard Durbin¹ - Show less +5 more•Institutions (4)

Wellcome Trust Sanger Institute¹, University of California, Los Angeles², Chinese Academy of Sciences³, Boston College⁴

01 Aug 2009-Bioinformatics

TL;DR: SAMtools as discussed by the authors implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments.

...read moreread less

Abstract: Summary: The Sequence Alignment/Map (SAM) format is a generic alignment format for storing read alignments against reference sequences, supporting short and long reads (up to 128 Mbp) produced by different sequencing platforms. It is flexible in style, compact in size, efficient in random access and is the format in which alignments from the 1000 Genomes Project are released. SAMtools implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments. Availability: http://samtools.sourceforge.net Contact: [email protected]

...read moreread less

45,957 citations

Journal Article•DOI•

BEDTools: a flexible suite of utilities for comparing genomic features

[...]

Aaron R. Quinlan¹, Ira M. Hall¹•Institutions (1)

University of Virginia¹

15 Mar 2010-Bioinformatics

TL;DR: A new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format, which allows the user to compare large datasets (e.g. next-generation sequencing data) with both public and custom genome annotation tracks.

...read moreread less

Abstract: Motivation: Testing for correlations between different sets of genomic features is a fundamental task in genomics research. However, searching for overlaps between features with existing webbased methods is complicated by the massive datasets that are routinely produced with current sequencing technologies. Fast and flexible tools are therefore required to ask complex questions of these data in an efficient manner. Results: This article introduces a new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format. BEDTools also supports the comparison of sequence alignments in BAM format to both BED and GFF features. The tools are extremely efficient and allow the user to compare large datasets (e.g. next-generation sequencing data) with both public and custom genome annotation tracks. BEDTools can be combined with one another as well as with standard UNIX commands, thus facilitating routine genomics tasks as well as pipelines that can quickly answer intricate questions of large genomic datasets. Availability and implementation: BEDTools was written in C++. Source code and a comprehensive user manual are freely available at http://code.google.com/p/bedtools

...read moreread less

18,858 citations

Journal Article•DOI•

An integrated encyclopedia of DNA elements in the human genome

[...]

Principal investigators¹, Nhgri groups², Data production leads³, Lead analysts³•Institutions (3)

Wellcome Trust¹, University of Washington², Pennsylvania State University³

06 Sep 2012-Nature

TL;DR: The Encyclopedia of DNA Elements project provides new insights into the organization and regulation of the authors' genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

Abstract: The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research.

...read moreread less

13,548 citations

Journal Article•DOI•

The variant call format and VCFtools

[...]

Petr Danecek¹, Adam Auton², Gonçalo R. Abecasis³, Cornelis A. Albers¹, Eric Banks⁴, Mark A. DePristo⁴, Robert E. Handsaker⁴, Gerton Lunter², Gabor T. Marth⁵, Stephen T. Sherry⁶, Gilean McVean², Richard Durbin¹ - Show less +8 more•Institutions (6)

Wellcome Trust¹, University of Oxford², University of Michigan³, Broad Institute⁴, Boston College⁵, National Institutes of Health⁶

01 Aug 2011-Bioinformatics

TL;DR: VCFtools is a software suite that implements various utilities for processing VCF files, including validation, merging, comparing and also provides a general Perl API.

...read moreread less

Abstract: Summary: The variant call format (VCF) is a generic format for storing DNA polymorphism data such as SNPs, insertions, deletions and structural variants, together with rich annotations. VCF is usually stored in a compressed manner and can be indexed for fast data retrieval of variants from a range of positions on the reference genome. The format was developed for the 1000 Genomes Project, and has also been adopted by other projects such as UK10K, dbSNP and the NHLBI Exome Project. VCFtools is a software suite that implements various utilities for processing VCF files, including validation, merging, comparing and also provides a general Perl API. Availability: http://vcftools.sourceforge.net Contact: [email protected]

...read moreread less

10,164 citations