scispace - formally typeset
Search or ask a question
Author

David Altshuler

Bio: David Altshuler is an academic researcher from University of Michigan. The author has contributed to research in topics: Genome-wide association study & Population. The author has an hindex of 162, co-authored 345 publications receiving 201782 citations. Previous affiliations of David Altshuler include Vertex Pharmaceuticals & Massachusetts Institute of Technology.


Papers
More filters
Journal ArticleDOI
TL;DR: The failure of standard methods to detect stratification in case-control association studies indicates that new methods may be required, and a SNP in the gene LCT that varies widely in frequency across Europe was strongly associated with height.
Abstract: Population stratification occurs in case-control association studies when allele frequencies differ between cases and controls because of ancestry. Stratification may lead to false positive associations, although this issue remains controversial. Empirical studies have found little evidence of stratification in European-derived populations, but potentially significant levels of stratification could not be ruled out. We studied a European American panel discordant for height, a heritable trait that varies widely across Europe. Genotyping 178 SNPs and applying standard analytical methods yielded no evidence of stratification. But a SNP in the gene LCT that varies widely in frequency across Europe was strongly associated with height (P < 10(-6)). This apparent association was largely or completely due to stratification; rematching individuals on the basis of European ancestry greatly reduced the apparent association, and no association was observed in Polish or Scandinavian individuals. The failure of standard methods to detect this stratification indicates that new methods may be required.

459 citations

Journal ArticleDOI
Zari Dastani1, Hivert M-F.2, Hivert M-F.3, N J Timpson4  +615 moreInstitutions (128)
TL;DR: A meta-analysis of genome-wide association studies in 39,883 individuals of European ancestry to identify genes associated with metabolic disease identifies novel genetic determinants of adiponectin levels, which, taken together, influence risk of T2D and markers of insulin resistance.
Abstract: Circulating levels of adiponectin, a hormone produced predominantly by adipocytes, are highly heritable and are inversely associated with type 2 diabetes mellitus (T2D) and other metabolic traits. We conducted a meta-analysis of genome-wide association studies in 39,883 individuals of European ancestry to identify genes associated with metabolic disease. We identified 8 novel loci associated with adiponectin levels and confirmed 2 previously reported loci (P = 4.5×10(-8)-1.2×10(-43)). Using a novel method to combine data across ethnicities (N = 4,232 African Americans, N = 1,776 Asians, and N = 29,347 Europeans), we identified two additional novel loci. Expression analyses of 436 human adipocyte samples revealed that mRNA levels of 18 genes at candidate regions were associated with adiponectin concentrations after accounting for multiple testing (p<3×10(-4)). We next developed a multi-SNP genotypic risk score to test the association of adiponectin decreasing risk alleles on metabolic traits and diseases using consortia-level meta-analytic data. This risk score was associated with increased risk of T2D (p = 4.3×10(-3), n = 22,044), increased triglycerides (p = 2.6×10(-14), n = 93,440), increased waist-to-hip ratio (p = 1.8×10(-5), n = 77,167), increased glucose two hours post oral glucose tolerance testing (p = 4.4×10(-3), n = 15,234), increased fasting insulin (p = 0.015, n = 48,238), but with lower in HDL-cholesterol concentrations (p = 4.5×10(-13), n = 96,748) and decreased BMI (p = 1.4×10(-4), n = 121,335). These findings identify novel genetic determinants of adiponectin levels, which, taken together, influence risk of T2D and markers of insulin resistance.

456 citations

Journal ArticleDOI
TL;DR: A statistical method that takes a list of disease regions and automatically assesses the degree of relatedness of implicated genes using 250,000 PubMed abstracts, and offers a statistically robust approach to identifying functionally related genes from across multiple disease regions—that likely represent key disease pathways.
Abstract: Translating a set of disease regions into insight about pathogenic mechanisms requires not only the ability to identify the key disease genes within them, but also the biological relationships among those key genes. Here we describe a statistical method, Gene Relationships Among Implicated Loci (GRAIL), that takes a list of disease regions and automatically assesses the degree of relatedness of implicated genes using 250,000 PubMed abstracts. We first evaluated GRAIL by assessing its ability to identify subsets of highly related genes in common pathways from validated lipid and height SNP associations from recent genome-wide studies. We then tested GRAIL, by assessing its ability to separate true disease regions from many false positive disease regions in two separate practical applications in human genetics. First, we took 74 nominally associated Crohn's disease SNPs and applied GRAIL to identify a subset of 13 SNPs with highly related genes. Of these, ten convincingly validated in follow-up genotyping; genotyping results for the remaining three were inconclusive. Next, we applied GRAIL to 165 rare deletion events seen in schizophrenia cases (less than one-third of which are contributing to disease risk). We demonstrate that GRAIL is able to identify a subset of 16 deletions containing highly related genes; many of these genes are expressed in the central nervous system and play a role in neuronal synapses. GRAIL offers a statistically robust approach to identifying functionally related genes from across multiple disease regions—that likely represent key disease pathways. An online version of this method is available for public use (http://www.broad.mit.edu/mpg/grail/).

455 citations

Journal ArticleDOI
TL;DR: Evidence is found for three functional alleles of IRF5: the previously described exon 1B splice site variant, a 30-bp in-frame insertion/deletion variant of exon 6 that alters a proline-, glutamic acid-, serine- and threonine-rich domain region, and a variant in a conserved polyA+ signal sequence that alters the length of the 3′ UTR and stability of IRf5 mRNAs.
Abstract: Systematic genome-wide studies to map genomic regions associated with human diseases are becoming more practical. Increasingly, efforts will be focused on the identification of the specific functional variants responsible for the disease. The challenges of identifying causal variants include the need for complete ascertainment of genetic variants and the need to consider the possibility of multiple causal alleles. We recently reported that risk of systemic lupus erythematosus (SLE) is strongly associated with a common SNP in IFN regulatory factor 5 (IRF5), and that this variant altered spicing in a way that might provide a functional explanation for the reproducible association to SLE risk. Here, by resequencing and genotyping in patients with SLE, we find evidence for three functional alleles of IRF5: the previously described exon 1B splice site variant, a 30-bp in-frame insertion/deletion variant of exon 6 that alters a proline-, glutamic acid-, serine- and threonine-rich domain region, and a variant in a conserved polyA+ signal sequence that alters the length of the 3' UTR and stability of IRF5 mRNAs. Haplotypes of these three variants define at least three distinct levels of risk to SLE. Understanding how combinations of variants influence IRF5 function may offer etiological and therapeutic insights in SLE; more generally, IRF5 and SLE illustrates how multiple common variants of the same gene can together influence risk of common disease.

441 citations

Journal ArticleDOI
TL;DR: The data suggest that TXNIP might play a key role in defective glucose homeostasis preceding overt T2DM, as it regulates both insulin-dependent and insulin-independent pathways of glucose uptake in human skeletal muscle.
Abstract: Background Type 2 diabetes mellitus (T2DM) is characterized by defects in insulin secretion and action. Impaired glucose uptake in skeletal muscle is believed to be one of the earliest features in the natural history of T2DM, although underlying mechanisms remain obscure. Methods and Findings We combined human insulin/glucose clamp physiological studies with genome-wide expression profiling to identify thioredoxin interacting protein (TXNIP) as a gene whose expression is powerfully suppressed by insulin yet stimulated by glucose. In healthy individuals, its expression was inversely correlated to total body measures of glucose uptake. Forced expression of TXNIP in cultured adipocytes significantly reduced glucose uptake, while silencing with RNA interference in adipocytes and in skeletal muscle enhanced glucose uptake, confirming that the gene product is also a regulator of glucose uptake. TXNIP expression is consistently elevated in the muscle of prediabetics and diabetics, although in a panel of 4,450 Scandinavian individuals, we found no evidence for association between common genetic variation in the TXNIP gene and T2DM.

437 citations


Cited by
More filters
Journal ArticleDOI
TL;DR: Bowtie 2 combines the strengths of the full-text minute index with the flexibility and speed of hardware-accelerated dynamic programming algorithms to achieve a combination of high speed, sensitivity and accuracy.
Abstract: As the rate of sequencing increases, greater throughput is demanded from read aligners. The full-text minute index is often used to make alignment very fast and memory-efficient, but the approach is ill-suited to finding longer, gapped alignments. Bowtie 2 combines the strengths of the full-text minute index with the flexibility and speed of hardware-accelerated dynamic programming algorithms to achieve a combination of high speed, sensitivity and accuracy.

37,898 citations

Journal ArticleDOI
TL;DR: The Gene Set Enrichment Analysis (GSEA) method as discussed by the authors focuses on gene sets, that is, groups of genes that share common biological function, chromosomal location, or regulation.
Abstract: Although genomewide RNA expression analysis has become a routine tool in biomedical research, extracting biological insight from such information remains a major challenge. Here, we describe a powerful analytical method called Gene Set Enrichment Analysis (GSEA) for interpreting gene expression data. The method derives its power by focusing on gene sets, that is, groups of genes that share common biological function, chromosomal location, or regulation. We demonstrate how GSEA yields insights into several cancer-related data sets, including leukemia and lung cancer. Notably, where single-gene analysis finds little similarity between two independent studies of patient survival in lung cancer, GSEA reveals many biological pathways in common. The GSEA method is embodied in a freely available software package, together with an initial database of 1,325 biologically defined gene sets.

34,830 citations

Journal ArticleDOI
TL;DR: This work introduces PLINK, an open-source C/C++ WGAS tool set, and describes the five main domains of function: data management, summary statistics, population stratification, association analysis, and identity-by-descent estimation, which focuses on the estimation and use of identity- by-state and identity/descent information in the context of population-based whole-genome studies.
Abstract: Whole-genome association studies (WGAS) bring new computational, as well as analytic, challenges to researchers. Many existing genetic-analysis tools are not designed to handle such large data sets in a convenient manner and do not necessarily exploit the new opportunities that whole-genome data bring. To address these issues, we developed PLINK, an open-source C/C++ WGAS tool set. With PLINK, large data sets comprising hundreds of thousands of markers genotyped for thousands of individuals can be rapidly manipulated and analyzed in their entirety. As well as providing tools to make the basic analytic steps computationally efficient, PLINK also supports some novel approaches to whole-genome data that take advantage of whole-genome coverage. We introduce PLINK and describe the five main domains of function: data management, summary statistics, population stratification, association analysis, and identity-by-descent estimation. In particular, we focus on the estimation and use of identity-by-state and identity-by-descent information in the context of population-based whole-genome studies. This information can be used to detect and correct for population stratification and to identify extended chromosomal segments that are shared identical by descent between very distantly related individuals. Analysis of the patterns of segmental sharing has the potential to map disease loci that contain multiple rare variants in a population-based linkage analysis.

26,280 citations

Journal ArticleDOI
Eric S. Lander1, Lauren Linton1, Bruce W. Birren1, Chad Nusbaum1  +245 moreInstitutions (29)
15 Feb 2001-Nature
TL;DR: The results of an international collaboration to produce and make freely available a draft sequence of the human genome are reported and an initial analysis is presented, describing some of the insights that can be gleaned from the sequence.
Abstract: The human genome holds an extraordinary trove of information about human development, physiology, medicine and evolution. Here we report the results of an international collaboration to produce and make freely available a draft sequence of the human genome. We also present an initial analysis of the data, describing some of the insights that can be gleaned from the sequence.

22,269 citations

Journal ArticleDOI
TL;DR: The philosophy and design of the limma package is reviewed, summarizing both new and historical features, with an emphasis on recent enhancements and features that have not been previously described.
Abstract: limma is an R/Bioconductor software package that provides an integrated solution for analysing data from gene expression experiments. It contains rich features for handling complex experimental designs and for information borrowing to overcome the problem of small sample sizes. Over the past decade, limma has been a popular choice for gene discovery through differential expression analyses of microarray and high-throughput PCR data. The package contains particularly strong facilities for reading, normalizing and exploring such data. Recently, the capabilities of limma have been significantly expanded in two important directions. First, the package can now perform both differential expression and differential splicing analyses of RNA sequencing (RNA-seq) data. All the downstream analysis tools previously restricted to microarray data are now available for RNA-seq as well. These capabilities allow users to analyse both RNA-seq and microarray data with very similar pipelines. Second, the package is now able to go past the traditional gene-wise expression analyses in a variety of ways, analysing expression profiles in terms of co-regulated sets of genes or in terms of higher-order expression signatures. This provides enhanced possibilities for biological interpretation of gene expression differences. This article reviews the philosophy and design of the limma package, summarizing both new and historical features, with an emphasis on recent enhancements and features that have not been previously described.

22,147 citations