scispace - formally typeset
Search or ask a question
Author

Nilesh J. Samani

Bio: Nilesh J. Samani is an academic researcher from University of Leicester. The author has contributed to research in topics: Genome-wide association study & Population. The author has an hindex of 149, co-authored 779 publications receiving 113545 citations. Previous affiliations of Nilesh J. Samani include University Hospitals of Leicester NHS Trust & Glenfield Hospital.


Papers
More filters
Journal ArticleDOI
TL;DR: Several genetic loci that, individually and in aggregate, substantially affect the risk of development of coronary artery disease are identified.
Abstract: A b s t r ac t Background Modern genotyping platforms permit a systematic search for inherited components of complex diseases. We performed a joint analysis of two genomewide association studies of coronary artery disease. Methods We first identified chromosomal loci that were strongly associated with coronary ar- tery disease in the Wellcome Trust Case Control Consortium (WTCCC) study (which involved 1926 case subjects with coronary artery disease and 2938 controls) and looked for replication in the German MI (Myocardial Infarction) Family Study (which involved 875 case subjects with myocardial infarction and 1644 controls). Data on other single- nucleotide polymorphisms (SNPs) that were significantly associated with coronary artery disease in either study (P 80%) of a true association: chromosomes 1p13.3 (rs599839), 1q41 (rs17465637), 10q11.21 (rs501120), and 15q22.33 (rs17228212). Conclusions We identified several genetic loci that, individually and in aggregate, substantially affect the risk of development of coronary artery disease.

2,000 citations

01 Jan 2010
TL;DR: 18 new loci associated with body mass index are identified, one of which includes a copy number variant near GPRC5B, and genes in other newly associated loci may provide new insights into human body weight regulation.
Abstract: Obesity is globally prevalent and highly heritable, but its underlying genetic factors remain largely elusive. To identify genetic loci for obesity susceptibility, we examined associations between body mass index and approximately 2.8 million SNPs in up to 123,865 individuals with targeted follow up of 42 SNPs in up to 125,931 additional individuals. We confirmed 14 known obesity susceptibility loci and identified 18 new loci associated with body mass index (P < 5 x 10(-)(8)), one of which includes a copy number variant near GPRC5B. Some loci (at MC4R, POMC, SH2B1 and BDNF) map near key hypothalamic regulators of energy balance, and one of these loci is near GIPR, an incretin receptor. Furthermore, genes in other newly associated loci may provide new insights into human body weight regulation.

1,953 citations

Journal ArticleDOI
Benjamin F. Voight1, Benjamin F. Voight2, Benjamin F. Voight3, Gina M. Peloso4, Gina M. Peloso5, Marju Orho-Melander6, Ruth Frikke-Schmidt7, Maja Barbalić8, Majken K. Jensen2, George Hindy6, Hilma Holm9, Eric L. Ding2, Toby Johnson10, Heribert Schunkert11, Nilesh J. Samani12, Nilesh J. Samani13, Robert Clarke14, Jemma C. Hopewell14, John F. Thompson13, Mingyao Li1, Gudmar Thorleifsson9, Christopher Newton-Cheh, Kiran Musunuru3, Kiran Musunuru2, James P. Pirruccello2, James P. Pirruccello3, Danish Saleheen15, Li Chen16, Alexandre F.R. Stewart16, Arne Schillert11, Unnur Thorsteinsdottir9, Unnur Thorsteinsdottir17, Gudmundur Thorgeirsson17, Sonia S. Anand18, James C. Engert19, Thomas M. Morgan20, John A. Spertus21, Monika Stoll22, Klaus Berger22, Nicola Martinelli23, Domenico Girelli23, Pascal P. McKeown24, Christopher Patterson24, Stephen E. Epstein25, Joseph M. Devaney25, Mary Susan Burnett25, Vincent Mooser26, Samuli Ripatti27, Ida Surakka27, Markku S. Nieminen27, Juha Sinisalo27, Marja-Liisa Lokki27, Markus Perola4, Aki S. Havulinna4, Ulf de Faire28, Bruna Gigante28, Erik Ingelsson28, Tanja Zeller29, Philipp S. Wild29, Paul I.W. de Bakker, Olaf H. Klungel30, Anke-Hilse Maitland-van der Zee30, Bas J M Peters30, Anthonius de Boer30, Diederick E. Grobbee30, Pieter Willem Kamphuisen31, Vera H.M. Deneer, Clara C. Elbers30, N. Charlotte Onland-Moret30, Marten H. Hofker31, Cisca Wijmenga31, W. M. Monique Verschuren, Jolanda M. A. Boer, Yvonne T. van der Schouw30, Asif Rasheed, Philippe M. Frossard, Serkalem Demissie5, Serkalem Demissie4, Cristen J. Willer32, Ron Do2, Jose M. Ordovas33, Jose M. Ordovas34, Gonçalo R. Abecasis32, Michael Boehnke32, Karen L. Mohlke35, Mark J. Daly3, Mark J. Daly2, Candace Guiducci3, Noël P. Burtt3, Aarti Surti3, Elena Gonzalez3, Shaun Purcell2, Shaun Purcell3, Stacey Gabriel3, Jaume Marrugat, John F. Peden14, Jeanette Erdmann11, Patrick Diemert11, Christina Willenborg11, Inke R. König11, Marcus Fischer36, Christian Hengstenberg36, Andreas Ziegler11, Ian Buysschaert37, Diether Lambrechts37, Frans Van de Werf37, Keith A.A. Fox38, Nour Eddine El Mokhtari39, Diana Rubin, Jürgen Schrezenmeir, Stefan Schreiber39, Arne Schäfer39, John Danesh15, Stefan Blankenberg29, Robert Roberts16, Ruth McPherson16, Hugh Watkins14, Alistair S. Hall40, Kim Overvad41, Eric B. Rimm2, Eric Boerwinkle8, Anne Tybjærg-Hansen7, L. Adrienne Cupples5, L. Adrienne Cupples4, Muredach P. Reilly1, Olle Melander6, Pier Mannuccio Mannucci42, Diego Ardissino, David S. Siscovick43, Roberto Elosua, Kari Stefansson9, Kari Stefansson17, Christopher J. O'Donnell2, Christopher J. O'Donnell4, Veikko Salomaa4, Daniel J. Rader1, Leena Peltonen44, Leena Peltonen27, Stephen M. Schwartz43, David Altshuler, Sekar Kathiresan 
11 Aug 2012
TL;DR: In this paper, a Mendelian randomisation analysis was performed to compare the effect of HDL cholesterol, LDL cholesterol, and genetic score on risk of myocardial infarction.
Abstract: Methods We performed two mendelian randomisation analyses. First, we used as an instrument a single nucleotide polymorphism (SNP) in the endothelial lipase gene (LIPG Asn396Ser) and tested this SNP in 20 studies (20 913 myocardial infarction cases, 95 407 controls). Second, we used as an instrument a genetic score consisting of 14 common SNPs that exclusively associate with HDL cholesterol and tested this score in up to 12 482 cases of myocardial infarction and 41 331 controls. As a positive control, we also tested a genetic score of 13 common SNPs exclusively associated with LDL cholesterol. – ¹³) but similar levels of other lipid and non-lipid risk factors for myocardial infarction compared with noncarriers. This diff erence in HDL cholesterol is expected to decrease risk of myocardial infarction by 13% (odds ratio [OR] 0·87, 95% CI 0·84–0·91). However, we noted that the 396Ser allele was not associated with risk of myocardial infarction (OR 0·99, 95% CI 0·88–1·11, p=0·85). From observational epidemiology, an increase of 1 SD in HDL cholesterol was associated with reduced risk of myocardial infarction (OR 0·62, 95% CI 0·58–0·66). However, a 1 SD increase in HDL cholesterol due to genetic score was not associated with risk of myocardial infarction (OR 0·93, 95% CI 0·68–1·26, p=0·63). For LDL cholesterol, the estimate from observational epidemiology (a 1 SD increase in LDL cholesterol associated with OR 1·54, 95% CI 1·45–1·63) was concordant with that from genetic score (OR 2·13, 95% CI 1·69–2·69, p=2×10

1,878 citations

Journal ArticleDOI
Andrew R. Wood1, Tõnu Esko2, Jian Yang3, Sailaja Vedantam4  +441 moreInstitutions (132)
TL;DR: This article identified 697 variants at genome-wide significance that together explained one-fifth of the heritability for adult height, and all common variants together captured 60% of heritability.
Abstract: Using genome-wide data from 253,288 individuals, we identified 697 variants at genome-wide significance that together explained one-fifth of the heritability for adult height. By testing different numbers of variants in independent studies, we show that the most strongly associated ∼2,000, ∼3,700 and ∼9,500 SNPs explained ∼21%, ∼24% and ∼29% of phenotypic variance. Furthermore, all common variants together captured 60% of heritability. The 697 variants clustered in 423 loci were enriched for genes, pathways and tissue types known to be involved in growth and together implicated genes and pathways not highlighted in earlier efforts, such as signaling by fibroblast growth factors, WNT/β-catenin and chondroitin sulfate-related genes. We identified several genes and pathways not previously connected with human skeletal growth, including mTOR, osteoglycin and binding of hyaluronic acid. Our results indicate a genetic architecture for human height that is characterized by a very large but finite number (thousands) of causal variants.

1,872 citations

Journal ArticleDOI
Majid Nikpay1, Anuj Goel2, Won H-H.3, Leanne M. Hall4  +164 moreInstitutions (60)
TL;DR: This article conducted a meta-analysis of coronary artery disease (CAD) cases and controls, interrogating 6.7 million common (minor allele frequency (MAF) > 0.05) and 2.7 millions low-frequency (0.005 < MAF < 0.5) variants.
Abstract: Existing knowledge of genetic variants affecting risk of coronary artery disease (CAD) is largely based on genome-wide association study (GWAS) analysis of common SNPs. Leveraging phased haplotypes from the 1000 Genomes Project, we report a GWAS meta-analysis of ∼185,000 CAD cases and controls, interrogating 6.7 million common (minor allele frequency (MAF) > 0.05) and 2.7 million low-frequency (0.005 < MAF < 0.05) variants. In addition to confirming most known CAD-associated loci, we identified ten new loci (eight additive and two recessive) that contain candidate causal genes newly implicating biological processes in vessel walls. We observed intralocus allelic heterogeneity but little evidence of low-frequency variants with larger effects and no evidence of synthetic association. Our analysis provides a comprehensive survey of the fine genetic architecture of CAD, showing that genetic susceptibility to this common disease is largely determined by common SNPs of small effect size.

1,839 citations


Cited by
More filters
28 Jul 2005
TL;DR: PfPMP1)与感染红细胞、树突状组胞以及胎盘的单个或多个受体作用,在黏附及免疫逃避中起关键的作�ly.
Abstract: 抗原变异可使得多种致病微生物易于逃避宿主免疫应答。表达在感染红细胞表面的恶性疟原虫红细胞表面蛋白1(PfPMP1)与感染红细胞、内皮细胞、树突状细胞以及胎盘的单个或多个受体作用,在黏附及免疫逃避中起关键的作用。每个单倍体基因组var基因家族编码约60种成员,通过启动转录不同的var基因变异体为抗原变异提供了分子基础。

18,940 citations

Journal ArticleDOI
Giuseppe Mancia1, Robert Fagard, Krzysztof Narkiewicz, Josep Redon, Alberto Zanchetti, Michael Böhm, Thierry Christiaens, Renata Cifkova, Guy De Backer, Anna F. Dominiczak, Maurizio Galderisi, Diederick E. Grobbee, Tiny Jaarsma, Paulus Kirchhof, Sverre E. Kjeldsen, Stéphane Laurent, Athanasios J. Manolis, Peter M. Nilsson, Luis M. Ruilope, Roland E. Schmieder, Per Anton Sirnes, Peter Sleight, Margus Viigimaa, Bernard Waeber, Faiez Zannad, Michel Burnier, Ettore Ambrosioni, Mark Caufield, Antonio Coca, Michael H. Olsen, Costas Tsioufis, Philippe van de Borne, José Luis Zamorano, Stephan Achenbach, Helmut Baumgartner, Jeroen J. Bax, Héctor Bueno, Veronica Dean, Christi Deaton, Çetin Erol, Roberto Ferrari, David Hasdai, Arno W. Hoes, Juhani Knuuti, Philippe Kolh2, Patrizio Lancellotti, Aleš Linhart, Petros Nihoyannopoulos, Massimo F Piepoli, Piotr Ponikowski, Juan Tamargo, Michal Tendera, Adam Torbicki, William Wijns, Stephan Windecker, Denis Clement, Thierry C. Gillebert, Enrico Agabiti Rosei, Stefan D. Anker, Johann Bauersachs, Jana Brguljan Hitij, Mark J. Caulfield, Marc De Buyzere, Sabina De Geest, Geneviève Derumeaux, Serap Erdine, Csaba Farsang, Christian Funck-Brentano, Vjekoslav Gerc, Giuseppe Germanò, Stephan Gielen, Herman Haller, Jens Jordan, Thomas Kahan, Michel Komajda, Dragan Lovic, Heiko Mahrholdt, Jan Östergren, Gianfranco Parati, Joep Perk, Jorge Polónia, Bogdan A. Popescu, Zeljko Reiner, Lars Rydén, Yuriy Sirenko, Alice Stanton, Harry A.J. Struijker-Boudier, Charalambos Vlachopoulos, Massimo Volpe, David A. Wood 
TL;DR: In this article, a randomized controlled trial of Aliskiren in the Prevention of Major Cardiovascular Events in Elderly people was presented. But the authors did not discuss the effect of the combination therapy in patients living with systolic hypertension.
Abstract: ABCD : Appropriate Blood pressure Control in Diabetes ABI : ankle–brachial index ABPM : ambulatory blood pressure monitoring ACCESS : Acute Candesartan Cilexetil Therapy in Stroke Survival ACCOMPLISH : Avoiding Cardiovascular Events in Combination Therapy in Patients Living with Systolic Hypertension ACCORD : Action to Control Cardiovascular Risk in Diabetes ACE : angiotensin-converting enzyme ACTIVE I : Atrial Fibrillation Clopidogrel Trial with Irbesartan for Prevention of Vascular Events ADVANCE : Action in Diabetes and Vascular Disease: Preterax and Diamicron-MR Controlled Evaluation AHEAD : Action for HEAlth in Diabetes ALLHAT : Antihypertensive and Lipid-Lowering Treatment to Prevent Heart ATtack ALTITUDE : ALiskiren Trial In Type 2 Diabetes Using Cardio-renal Endpoints ANTIPAF : ANgioTensin II Antagonist In Paroxysmal Atrial Fibrillation APOLLO : A Randomized Controlled Trial of Aliskiren in the Prevention of Major Cardiovascular Events in Elderly People ARB : angiotensin receptor blocker ARIC : Atherosclerosis Risk In Communities ARR : aldosterone renin ratio ASCOT : Anglo-Scandinavian Cardiac Outcomes Trial ASCOT-LLA : Anglo-Scandinavian Cardiac Outcomes Trial—Lipid Lowering Arm ASTRAL : Angioplasty and STenting for Renal Artery Lesions A-V : atrioventricular BB : beta-blocker BMI : body mass index BP : blood pressure BSA : body surface area CA : calcium antagonist CABG : coronary artery bypass graft CAPPP : CAPtopril Prevention Project CAPRAF : CAndesartan in the Prevention of Relapsing Atrial Fibrillation CHD : coronary heart disease CHHIPS : Controlling Hypertension and Hypertension Immediately Post-Stroke CKD : chronic kidney disease CKD-EPI : Chronic Kidney Disease—EPIdemiology collaboration CONVINCE : Controlled ONset Verapamil INvestigation of CV Endpoints CT : computed tomography CV : cardiovascular CVD : cardiovascular disease D : diuretic DASH : Dietary Approaches to Stop Hypertension DBP : diastolic blood pressure DCCT : Diabetes Control and Complications Study DIRECT : DIabetic REtinopathy Candesartan Trials DM : diabetes mellitus DPP-4 : dipeptidyl peptidase 4 EAS : European Atherosclerosis Society EASD : European Association for the Study of Diabetes ECG : electrocardiogram EF : ejection fraction eGFR : estimated glomerular filtration rate ELSA : European Lacidipine Study on Atherosclerosis ESC : European Society of Cardiology ESH : European Society of Hypertension ESRD : end-stage renal disease EXPLOR : Amlodipine–Valsartan Combination Decreases Central Systolic Blood Pressure more Effectively than the Amlodipine–Atenolol Combination FDA : U.S. Food and Drug Administration FEVER : Felodipine EVent Reduction study GISSI-AF : Gruppo Italiano per lo Studio della Sopravvivenza nell'Infarto Miocardico-Atrial Fibrillation HbA1c : glycated haemoglobin HBPM : home blood pressure monitoring HOPE : Heart Outcomes Prevention Evaluation HOT : Hypertension Optimal Treatment HRT : hormone replacement therapy HT : hypertension HYVET : HYpertension in the Very Elderly Trial IMT : intima-media thickness I-PRESERVE : Irbesartan in Heart Failure with Preserved Systolic Function INTERHEART : Effect of Potentially Modifiable Risk Factors associated with Myocardial Infarction in 52 Countries INVEST : INternational VErapamil SR/T Trandolapril ISH : Isolated systolic hypertension JNC : Joint National Committee JUPITER : Justification for the Use of Statins in Primary Prevention: an Intervention Trial Evaluating Rosuvastatin LAVi : left atrial volume index LIFE : Losartan Intervention For Endpoint Reduction in Hypertensives LV : left ventricle/left ventricular LVH : left ventricular hypertrophy LVM : left ventricular mass MDRD : Modification of Diet in Renal Disease MRFIT : Multiple Risk Factor Intervention Trial MRI : magnetic resonance imaging NORDIL : The Nordic Diltiazem Intervention study OC : oral contraceptive OD : organ damage ONTARGET : ONgoing Telmisartan Alone and in Combination with Ramipril Global Endpoint Trial PAD : peripheral artery disease PATHS : Prevention And Treatment of Hypertension Study PCI : percutaneous coronary intervention PPAR : peroxisome proliferator-activated receptor PREVEND : Prevention of REnal and Vascular ENdstage Disease PROFESS : Prevention Regimen for Effectively Avoiding Secondary Strokes PROGRESS : Perindopril Protection Against Recurrent Stroke Study PWV : pulse wave velocity QALY : Quality adjusted life years RAA : renin-angiotensin-aldosterone RAS : renin-angiotensin system RCT : randomized controlled trials RF : risk factor ROADMAP : Randomized Olmesartan And Diabetes MicroAlbuminuria Prevention SBP : systolic blood pressure SCAST : Angiotensin-Receptor Blocker Candesartan for Treatment of Acute STroke SCOPE : Study on COgnition and Prognosis in the Elderly SCORE : Systematic COronary Risk Evaluation SHEP : Systolic Hypertension in the Elderly Program STOP : Swedish Trials in Old Patients with Hypertension STOP-2 : The second Swedish Trial in Old Patients with Hypertension SYSTCHINA : SYSTolic Hypertension in the Elderly: Chinese trial SYSTEUR : SYSTolic Hypertension in Europe TIA : transient ischaemic attack TOHP : Trials Of Hypertension Prevention TRANSCEND : Telmisartan Randomised AssessmeNt Study in ACE iNtolerant subjects with cardiovascular Disease UKPDS : United Kingdom Prospective Diabetes Study VADT : Veterans' Affairs Diabetes Trial VALUE : Valsartan Antihypertensive Long-term Use Evaluation WHO : World Health Organization ### 1.1 Principles The 2013 guidelines on hypertension of the European Society of Hypertension (ESH) and the European Society of Cardiology …

14,173 citations

Journal ArticleDOI
Adam Auton1, Gonçalo R. Abecasis2, David Altshuler3, Richard Durbin4  +514 moreInstitutions (90)
01 Oct 2015-Nature
TL;DR: The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations, and has reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-generation sequencing, deep exome sequencing, and dense microarray genotyping.
Abstract: The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes >99% of SNP variants with a frequency of >1% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies.

12,661 citations

Journal ArticleDOI
Paul Burton1, David Clayton2, Lon R. Cardon, Nicholas John Craddock3  +192 moreInstitutions (4)
07 Jun 2007-Nature
TL;DR: This study has demonstrated that careful use of a shared control group represents a safe and effective approach to GWA analyses of multiple disease phenotypes; generated a genome-wide genotype database for future studies of common diseases in the British population; and shown that, provided individuals with non-European ancestry are excluded, the extent of population stratification in theBritish population is generally modest.
Abstract: There is increasing evidence that genome-wide association ( GWA) studies represent a powerful approach to the identification of genes involved in common human diseases. We describe a joint GWA study ( using the Affymetrix GeneChip 500K Mapping Array Set) undertaken in the British population, which has examined similar to 2,000 individuals for each of 7 major diseases and a shared set of similar to 3,000 controls. Case-control comparisons identified 24 independent association signals at P < 5 X 10(-7): 1 in bipolar disorder, 1 in coronary artery disease, 9 in Crohn's disease, 3 in rheumatoid arthritis, 7 in type 1 diabetes and 3 in type 2 diabetes. On the basis of prior findings and replication studies thus-far completed, almost all of these signals reflect genuine susceptibility effects. We observed association at many previously identified loci, and found compelling evidence that some loci confer risk for more than one of the diseases studied. Across all diseases, we identified a large number of further signals ( including 58 loci with single-point P values between 10(-5) and 5 X 10(-7)) likely to yield additional susceptibility loci. The importance of appropriately large samples was confirmed by the modest effect sizes observed at most loci identified. This study thus represents a thorough validation of the GWA approach. It has also demonstrated that careful use of a shared control group represents a safe and effective approach to GWA analyses of multiple disease phenotypes; has generated a genome-wide genotype database for future studies of common diseases in the British population; and shown that, provided individuals with non-European ancestry are excluded, the extent of population stratification in the British population is generally modest. Our findings offer new avenues for exploring the pathophysiology of these important disorders. We anticipate that our data, results and software, which will be widely available to other investigators, will provide a powerful resource for human genetics research.

9,244 citations

Journal ArticleDOI
Monkol Lek, Konrad J. Karczewski1, Konrad J. Karczewski2, Eric Vallabh Minikel2, Eric Vallabh Minikel1, Kaitlin E. Samocha, Eric Banks1, Timothy Fennell1, Anne H. O’Donnell-Luria2, Anne H. O’Donnell-Luria3, Anne H. O’Donnell-Luria1, James S. Ware, Andrew J. Hill4, Andrew J. Hill2, Andrew J. Hill1, Beryl B. Cummings1, Beryl B. Cummings2, Taru Tukiainen2, Taru Tukiainen1, Daniel P. Birnbaum1, Jack A. Kosmicki, Laramie E. Duncan1, Laramie E. Duncan2, Karol Estrada2, Karol Estrada1, Fengmei Zhao2, Fengmei Zhao1, James Zou1, Emma Pierce-Hoffman2, Emma Pierce-Hoffman1, Joanne Berghout5, David Neil Cooper6, Nicole A. Deflaux7, Mark A. DePristo1, Ron Do, Jason Flannick2, Jason Flannick1, Menachem Fromer, Laura D. Gauthier1, Jackie Goldstein2, Jackie Goldstein1, Namrata Gupta1, Daniel P. Howrigan2, Daniel P. Howrigan1, Adam Kiezun1, Mitja I. Kurki1, Mitja I. Kurki2, Ami Levy Moonshine1, Pradeep Natarajan, Lorena Orozco, Gina M. Peloso1, Gina M. Peloso2, Ryan Poplin1, Manuel A. Rivas1, Valentin Ruano-Rubio1, Samuel A. Rose1, Douglas M. Ruderfer8, Khalid Shakir1, Peter D. Stenson6, Christine Stevens1, Brett Thomas1, Brett Thomas2, Grace Tiao1, María Teresa Tusié-Luna, Ben Weisburd1, Hong-Hee Won9, Dongmei Yu, David Altshuler1, David Altshuler10, Diego Ardissino, Michael Boehnke11, John Danesh12, Stacey Donnelly1, Roberto Elosua, Jose C. Florez2, Jose C. Florez1, Stacey Gabriel1, Gad Getz1, Gad Getz2, Stephen J. Glatt13, Christina M. Hultman14, Sekar Kathiresan, Markku Laakso15, Steven A. McCarroll1, Steven A. McCarroll2, Mark I. McCarthy16, Mark I. McCarthy17, Dermot P.B. McGovern18, Ruth McPherson19, Benjamin M. Neale1, Benjamin M. Neale2, Aarno Palotie, Shaun Purcell8, Danish Saleheen20, Jeremiah M. Scharf, Pamela Sklar, Patrick F. Sullivan21, Patrick F. Sullivan14, Jaakko Tuomilehto22, Ming T. Tsuang23, Hugh Watkins16, Hugh Watkins17, James G. Wilson24, Mark J. Daly1, Mark J. Daly2, Daniel G. MacArthur2, Daniel G. MacArthur1 
18 Aug 2016-Nature
TL;DR: The aggregation and analysis of high-quality exome (protein-coding region) DNA sequence data for 60,706 individuals of diverse ancestries generated as part of the Exome Aggregation Consortium (ExAC) provides direct evidence for the presence of widespread mutational recurrence.
Abstract: Large-scale reference data sets of human genetic variation are critical for the medical and functional interpretation of DNA sequence changes. Here we describe the aggregation and analysis of high-quality exome (protein-coding region) DNA sequence data for 60,706 individuals of diverse ancestries generated as part of the Exome Aggregation Consortium (ExAC). This catalogue of human genetic diversity contains an average of one variant every eight bases of the exome, and provides direct evidence for the presence of widespread mutational recurrence. We have used this catalogue to calculate objective metrics of pathogenicity for sequence variants, and to identify genes subject to strong selection against various classes of mutation; identifying 3,230 genes with near-complete depletion of predicted protein-truncating variants, with 72% of these genes having no currently established human disease phenotype. Finally, we demonstrate that these data can be used for the efficient filtering of candidate disease-causing variants, and for the discovery of human 'knockout' variants in protein-coding genes.

8,758 citations