scispace - formally typeset
Search or ask a question
Author

Ning Li

Bio: Ning Li is an academic researcher from University of Minnesota. The author has contributed to research in topics: Gene & Population. The author has an hindex of 51, co-authored 449 publications receiving 14228 citations. Previous affiliations of Ning Li include Beijing Institute of Genomics & University of Science and Technology of China.


Papers
More filters
Journal ArticleDOI
Erich D. Jarvis1, Siavash Mirarab2, Andre J. Aberer3, Bo Li4, Bo Li5, Bo Li6, Peter Houde7, Cai Li5, Cai Li4, Simon Y. W. Ho8, Brant C. Faircloth9, Benoit Nabholz, Jason T. Howard1, Alexander Suh10, Claudia C. Weber10, Rute R. da Fonseca11, Jianwen Li, Fang Zhang Zhang, Hui Li, Long Zhou, Nitish Narula7, Nitish Narula12, Liang Liu13, Ganesh Ganapathy1, Bastien Boussau, Shamsuzzoha Bayzid2, Volodymyr Zavidovych1, Sankar Subramanian14, Toni Gabaldón15, Salvador Capella-Gutierrez, Jaime Huerta-Cepas, Bhanu Rekepalli16, Bhanu Rekepalli17, Kasper Munch18, Mikkel H. Schierup18, Bent E. K. Lindow11, Wesley C. Warren19, David A. Ray, Richard E. Green20, Michael William Bruford21, Xiangjiang Zhan21, Xiangjiang Zhan22, Andrew Dixon, Shengbin Li6, Ning Li23, Yinhua Huang23, Elizabeth P. Derryberry24, Elizabeth P. Derryberry25, Mads F. Bertelsen26, Frederick H. Sheldon25, Robb T. Brumfield25, Claudio V. Mello27, Claudio V. Mello28, Peter V. Lovell27, Morgan Wirthlin27, Maria Paula Cruz Schneider28, Francisco Prosdocimi28, José Alfredo Samaniego11, Amhed Missael Vargas Velazquez11, Alonzo Alfaro-Núñez11, Paula F. Campos11, Bent O. Petersen29, Thomas Sicheritz-Pontén29, An Pas, Thomas L. Bailey, R. Paul Scofield30, Michael Bunce31, David M. Lambert14, Qi Zhou, Polina L. Perelman32, Amy C. Driskell33, Beth Shapiro20, Zijun Xiong, Yongli Zeng, Shiping Liu, Zhenyu Li, Binghang Liu, Kui Wu, Jin Xiao, Xiong Yinqi, Quiemei Zheng, Yong Zhang, Huanming Yang, Jian Wang, Linnéa Smeds10, Frank E. Rheindt34, Michael J. Braun35, Jon Fjeldså11, Ludovic Orlando11, F. Keith Barker4, Knud A. Jønsson4, Warren E. Johnson33, Klaus-Peter Koepfli33, Stephen J. O'Brien36, David Haussler, Oliver A. Ryder, Carsten Rahbek4, Eske Willerslev11, Gary R. Graves33, Gary R. Graves4, Travis C. Glenn13, John E. McCormack37, Dave Burt38, Hans Ellegren10, Per Alström, Scott V. Edwards39, Alexandros Stamatakis3, David P. Mindell40, Joel Cracraft4, Edward L. Braun41, Tandy Warnow42, Tandy Warnow2, Wang Jun, M. Thomas P. Gilbert31, M. Thomas P. Gilbert4, Guojie Zhang11, Guojie Zhang5 
12 Dec 2014-Science
TL;DR: A genome-scale phylogenetic analysis of 48 species representing all orders of Neoaves recovered a highly resolved tree that confirms previously controversial sister or close relationships and identifies the first divergence in Neoaves, two groups the authors named Passerea and Columbea.
Abstract: To better determine the history of modern birds, we performed a genome-scale phylogenetic analysis of 48 species representing all orders of Neoaves using phylogenomic methods created to handle genome-scale data. We recovered a highly resolved tree that confirms previously controversial sister or close relationships. We identified the first divergence in Neoaves, two groups we named Passerea and Columbea, representing independent lineages of diverse and convergently evolved land and water bird species. Among Passerea, we infer the common ancestor of core landbirds to have been an apex predator and confirm independent gains of vocal learning. Among Columbea, we identify pigeons and flamingoes as belonging to sister clades. Even with whole genomes, some of the earliest branches in Neoaves proved challenging to resolve, which was best explained by massive protein-coding sequence convergence and high levels of incomplete lineage sorting that occurred during a rapid radiation after the Cretaceous-Paleogene mass extinction event about 66 million years ago.

1,624 citations

Journal ArticleDOI
06 Nov 2008-Nature
TL;DR: Genotyping analysis showed that SNP identification had high accuracy and consistency, indicating the high sequence quality of this assembly, and the potential usefulness of next-generation sequencing technologies for personal genomics.
Abstract: Here we present the first diploid genome sequence of an Asian individual. The genome was sequenced to 36-fold average coverage using massively parallel sequencing technology. We aligned the short reads onto the NCBI human reference genome to 99.97% coverage, and guided by the reference genome, we used uniquely mapped reads to assemble a high-quality consensus sequence for 92% of the Asian individual's genome. We identified approximately 3 million single-nucleotide polymorphisms (SNPs) inside this region, of which 13.6% were not in the dbSNP database. Genotyping analysis showed that SNP identification had high accuracy and consistency, indicating the high sequence quality of this assembly. We also carried out heterozygote phasing and haplotype prediction against HapMap CHB and JPT haplotypes (Chinese and Japanese, respectively), sequence comparison with the two available individual genomes (J. D. Watson and J. C. Venter), and structural variation identification. These variations were considered for their potential biological impact. Our sequence data and analyses demonstrate the potential usefulness of next-generation sequencing technologies for personal genomics.

963 citations

Journal ArticleDOI
Guojie Zhang1, Guojie Zhang2, Cai Li2, Qiye Li2, Bo Li2, Denis M. Larkin3, Chul Hee Lee4, Jay F. Storz5, Agostinho Antunes6, Matthew J. Greenwold7, Robert W. Meredith8, Anders Ödeen9, Jie Cui10, Qi Zhou11, Luohao Xu2, Hailin Pan2, Zongji Wang12, Lijun Jin2, Pei Zhang2, Haofu Hu2, Wei Yang2, Jiang Hu2, Jin Xiao2, Zhikai Yang2, Yang Liu2, Qiaolin Xie2, Hao Yu2, Jinmin Lian2, Ping Wen2, Fang Zhang2, Hui Li2, Yongli Zeng2, Zijun Xiong2, Shiping Liu12, Long Zhou2, Zhiyong Huang2, Na An2, Jie Wang13, Qiumei Zheng2, Yingqi Xiong2, Guangbiao Wang2, Bo Wang2, Jingjing Wang2, Yu Fan14, Rute R. da Fonseca1, Alonzo Alfaro-Núñez1, Mikkel Schubert1, Ludovic Orlando1, Tobias Mourier1, Jason T. Howard15, Ganeshkumar Ganapathy15, Andreas R. Pfenning15, Osceola Whitney15, Miriam V. Rivas15, Erina Hara15, Julia Smith15, Marta Farré3, Jitendra Narayan16, Gancho T. Slavov16, Michael N Romanov17, Rui Borges6, João Paulo Machado6, Imran Khan6, Mark S. Springer18, John Gatesy18, Federico G. Hoffmann19, Juan C. Opazo20, Olle Håstad21, Roger H. Sawyer7, Heebal Kim4, Kyu-Won Kim4, Hyeon Jeong Kim4, Seoae Cho4, Ning Li22, Yinhua Huang22, Michael William Bruford23, Xiangjiang Zhan13, Andrew Dixon, Mads F. Bertelsen24, Elizabeth P. Derryberry25, Wesley C. Warren26, Richard K. Wilson26, Shengbin Li27, David A. Ray19, Richard E. Green28, Stephen J. O'Brien29, Darren K. Griffin17, Warren E. Johnson30, David Haussler28, Oliver A. Ryder, Eske Willerslev1, Gary R. Graves31, Per Alström21, Jon Fjeldså32, David P. Mindell33, Scott V. Edwards34, Edward L. Braun35, Carsten Rahbek32, David W. Burt36, Peter Houde37, Yong Zhang2, Huanming Yang38, Jian Wang2, Erich D. Jarvis15, M. Thomas P. Gilbert1, M. Thomas P. Gilbert39, Jun Wang 
12 Dec 2014-Science
TL;DR: This work explored bird macroevolution using full genomes from 48 avian species representing all major extant clades to reveal that pan-avian genomic diversity covaries with adaptations to different lifestyles and convergent evolution of traits.
Abstract: Birds are the most species-rich class of tetrapod vertebrates and have wide relevance across many research fields. We explored bird macroevolution using full genomes from 48 avian species representing all major extant clades. The avian genome is principally characterized by its constrained size, which predominantly arose because of lineage-specific erosion of repetitive elements, large segmental deletions, and gene loss. Avian genomes furthermore show a remarkably high degree of evolutionary stasis at the levels of nucleotide sequence, gene synteny, and chromosomal structure. Despite this pattern of conservation, we detected many non-neutral evolutionary changes in protein-coding genes and noncoding regions. These analyses reveal that pan-avian genomic diversity covaries with adaptations to different lifestyles and convergent evolution of traits.

872 citations

Journal ArticleDOI
TL;DR: A draft genome anchored onto nine chromosomes and annotated 38,801 genes was produced and key chromosome reshuffling events were detected through collinearity identification between foxtail millet, rice and sorghum.
Abstract: Completion of genome sequences for the diploid Setaria italica reveals features of C4 photosynthesis that could enable improvement of the polyploid biofuel crop switchgrass (Panicum virgatum). The genetic basis of biotechnologically relevant traits, including drought tolerance, photosynthetic efficiency and flowering control, is also highlighted. Foxtail millet (Setaria italica), a member of the Poaceae grass family, is an important food and fodder crop in arid regions and has potential for use as a C4 biofuel. It is a model system for other biofuel grasses, including switchgrass and pearl millet. We produced a draft genome (∼423 Mb) anchored onto nine chromosomes and annotated 38,801 genes. Key chromosome reshuffling events were detected through collinearity identification between foxtail millet, rice and sorghum including two reshuffling events fusing rice chromosomes 7 and 9, 3 and 10 to foxtail millet chromosomes 2 and 9, respectively, that occurred after the divergence of foxtail millet and rice, and a single reshuffling event fusing rice chromosome 5 and 12 to foxtail millet chromosome 3 that occurred after the divergence of millet and sorghum. Rearrangements in the C4 photosynthesis pathway were also identified.

553 citations

Journal ArticleDOI
TL;DR: Comparing the genome of Tibetan wild boars with those of neighboring Chinese domestic pigs further showed the impact of thousands of years of artificial selection and different signatures of selection in wild boar and domestic pig.
Abstract: We report the sequencing at 131× coverage, de novo assembly and analyses of the genome of a female Tibetan wild boar. We also resequenced the whole genomes of 30 Tibetan wild boars from six major distributed locations and 18 geographically related pigs in China. We characterized genetic diversity, population structure and patterns of evolution. We searched for genomic regions under selection, which includes genes that are involved in hypoxia, olfaction, energy metabolism and drug response. Comparing the genome of Tibetan wild boar with those of neighboring Chinese domestic pigs further showed the impact of thousands of years of artificial selection and different signatures of selection in wild boar and domestic pig. We also report genetic adaptations in Tibetan wild boar that are associated with high altitudes and characterize the genetic basis of increased salivation in domestic pig.

412 citations


Cited by
More filters
Journal ArticleDOI
TL;DR: The GATK programming framework enables developers and analysts to quickly and easily write efficient and robust NGS tools, many of which have already been incorporated into large-scale sequencing projects like the 1000 Genomes Project and The Cancer Genome Atlas.
Abstract: Next-generation DNA sequencing (NGS) projects, such as the 1000 Genomes Project, are already revolutionizing our understanding of genetic variation among individuals. However, the massive data sets generated by NGS—the 1000 Genome pilot alone includes nearly five terabases—make writing feature-rich, efficient, and robust analysis tools difficult for even computationally sophisticated individuals. Indeed, many professionals are limited in the scope and the ease with which they can answer scientific questions by the complexity of accessing and manipulating the data produced by these machines. Here, we discuss our Genome Analysis Toolkit (GATK), a structured programming framework designed to ease the development of efficient and robust analysis tools for next-generation DNA sequencers using the functional programming philosophy of MapReduce. The GATK provides a small but rich set of data access patterns that encompass the majority of analysis tool needs. Separating specific analysis calculations from common data management infrastructure enables us to optimize the GATK framework for correctness, stability, and CPU and memory efficiency and to enable distributed and shared memory parallelization. We highlight the capabilities of the GATK by describing the implementation and application of robust, scale-tolerant tools like coverage calculators and single nucleotide polymorphism (SNP) calling. We conclude that the GATK programming framework enables developers and analysts to quickly and easily write efficient and robust NGS tools, many of which have already been incorporated into large-scale sequencing projects like the 1000 Genomes Project and The Cancer Genome Atlas.

20,557 citations

Journal ArticleDOI
TL;DR: Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches and can be used simultaneously to achieve even greater alignment speeds.
Abstract: Bowtie is an ultrafast, memory-efficient alignment program for aligning short DNA sequence reads to large genomes. For the human genome, Burrows-Wheeler indexing allows Bowtie to align more than 25 million reads per CPU hour with a memory footprint of approximately 1.3 gigabytes. Bowtie extends previous Burrows-Wheeler techniques with a novel quality-aware backtracking algorithm that permits mismatches. Multiple processor cores can be used simultaneously to achieve even greater alignment speeds. Bowtie is open source http://bowtie.cbcb.umd.edu.

20,335 citations

Journal Article
Fumio Tajima1
30 Oct 1989-Genomics
TL;DR: It is suggested that the natural selection against large insertion/deletion is so weak that a large amount of variation is maintained in a population.

11,521 citations

Journal ArticleDOI
28 Oct 2010-Nature
TL;DR: The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype as mentioned in this paper, and the results of the pilot phase of the project, designed to develop and compare different strategies for genomewide sequencing with high-throughput platforms.
Abstract: The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.

7,538 citations