scispace - formally typeset
Search or ask a question
Author

Sébastien Aubourg

Bio: Sébastien Aubourg is an academic researcher from Institut national de la recherche agronomique. The author has contributed to research in topics: Gene & Genome. The author has an hindex of 21, co-authored 34 publications receiving 6956 citations. Previous affiliations of Sébastien Aubourg include University of Paris-Sud & University of Paris.

Papers
More filters
Journal ArticleDOI
26 Aug 2007-Nature
TL;DR: A high-quality draft of the genome sequence of grapevine is obtained from a highly homozygous genotype, revealing the contribution of three ancestral genomes to the grapevine haploid content and explaining the chronology of previously described whole-genome duplication events in the evolution of flowering plants.
Abstract: The analysis of the first plant genomes provided unexpected evidence for genome duplication events in species that had previously been considered as true diploids on the basis of their genetics. These polyploidization events may have had important consequences in plant evolution, in particular for species radiation and adaptation and for the modulation of functional capacities. Here we report a high-quality draft of the genome sequence of grapevine (Vitis vinifera) obtained from a highly homozygous genotype. The draft sequence of the grapevine genome is the fourth one produced so far for flowering plants, the second for a woody species and the first for a fruit crop (cultivated for both fruit and beverage). Grapevine was selected because of its important place in the cultural heritage of humanity beginning during the Neolithic period. Several large expansions of gene families with roles in aromatic features are observed. The grapevine genome has not undergone recent genome duplication, thus enabling the discovery of ancestral traits and features of the genetic organization of flowering plants. This analysis reveals the contribution of three ancestral genomes to the grapevine haploid content. This ancestral arrangement is common to many dicotyledonous plants but is absent from the genome of rice, which is a monocotyledon. Furthermore, we explain the chronology of previously described whole-genome duplication events in the evolution of flowering plants.

3,311 citations

Journal ArticleDOI
TL;DR: A detailed bioinformatic analysis of 441 members of the Arabidopsis PPR family plus genomic and genetic data on the expression, localization, and general function of many family members confirm, but massively extend, the very sparse observations previously obtained from detailed characterization of individual mutants in other organisms.
Abstract: The complete sequence of the Arabidopsis thaliana genome revealed thousands of previously unsuspected genes, many of which cannot be ascribed even putative functions. One of the largest and most enigmatic gene families discovered in this way is characterized by tandem arrays of pentatricopeptide repeats (PPRs). We describe a detailed bioinformatic analysis of 441 members of the Arabidopsis PPR family plus genomic and genetic data on the expression (microarray data), localization (green fluorescent protein and red fluorescent protein fusions), and general function (insertion mutants and RNA binding assays) of many family members. The basic picture that arises from these studies is that PPR proteins play constitutive, often essential roles in mitochondria and chloroplasts, probably via binding to organellar transcripts. These results confirm, but massively extend, the very sparse observations previously obtained from detailed characterization of individual mutants in other organisms.

1,207 citations

Journal ArticleDOI
TL;DR: High-quality de novo assembly of the apple genome is produced and genome-wide DNA methylation data suggest that epigenetic marks may contribute to agronomically relevant aspects, such as apple fruit development.
Abstract: Using the latest sequencing and optical mapping technologies, we have produced a high-quality de novo assembly of the apple (Malus domestica Borkh.) genome. Repeat sequences, which represented over half of the assembly, provided an unprecedented opportunity to investigate the uncharacterized regions of a tree genome; we identified a new hyper-repetitive retrotransposon sequence that was over-represented in heterochromatic regions and estimated that a major burst of different transposable elements (TEs) occurred 21 million years ago. Notably, the timing of this TE burst coincided with the uplift of the Tian Shan mountains, which is thought to be the center of the location where the apple originated, suggesting that TEs and associated processes may have contributed to the diversification of the apple ancestor and possibly to its divergence from pear. Finally, genome-wide DNA methylation data suggest that epigenetic marks may contribute to agronomically relevant aspects, such as apple fruit development.

588 citations

Journal ArticleDOI
TL;DR: The data suggest that illegitimate DNA recombination, leading to various genomic rearrangements, constitutes one of the major evolutionary mechanisms in wheat species.
Abstract: The Hardness (Ha) locus controls grain hardness in hexaploid wheat (Triticum aestivum) and its relatives (Triticum and Aegilops species) and represents a classical example of a trait whose variation arose from gene loss after polyploidization. In this study, we investigated the molecular basis of the evolutionary events observed at this locus by comparing corresponding sequences of diploid, tertraploid, and hexaploid wheat species (Triticum and Aegilops). Genomic rearrangements, such as transposable element insertions, genomic deletions, duplications, and inversions, were shown to constitute the major differences when the same genomes (i.e., the A, B, or D genomes) were compared between species of different ploidy levels. The comparative analysis allowed us to determine the extent and sequences of the rearranged regions as well as rearrangement breakpoints and sequence motifs at their boundaries, which suggest rearrangement by illegitimate recombination. Among these genomic rearrangements, the previously reported Pina and Pinb genes loss from the Ha locus of polyploid wheat species was caused by a large genomic deletion that probably occurred independently in the A and B genomes. Moreover, the Ha locus in the D genome of hexaploid wheat (T. aestivum) is 29 kb smaller than in the D genome of its diploid progenitor Ae. tauschii, principally because of transposable element insertions and two large deletions caused by illegitimate recombination. Our data suggest that illegitimate DNA recombination, leading to various genomic rearrangements, constitutes one of the major evolutionary mechanisms in wheat species.

374 citations

Journal ArticleDOI
TL;DR: Phylogenetic analyses highlight events in the divergence of the TPS paralogs and suggest orthologous genes and a model for the evolution of theTPS gene family.
Abstract: A family of 40 terpenoid synthase genes (AtTPS) was discovered by genome sequence analysis in Arabidopsis thaliana. This is the largest and most diverse group of TPS genes currently known for any species. AtTPS genes cluster into five phylogenetic subfamilies of the plant TPS superfamily. Surprisingly, thirty AtTPS closely resemble, in all aspects of gene architecture, sequence relatedness and phylogenetic placement, the genes for plant monoterpene synthases, sesquiterpene synthases or diterpene synthases of secondary metabolism. Rapid evolution of these AtTPS resulted from repeated gene duplication and sequence divergence with minor changes in gene architecture. In contrast, only two AtTPS genes have known functions in basic (primary) metabolism, namely gibberellin biosynthesis. This striking difference in rates of gene diversification in primary and secondary metabolism is relevant for an understanding of the evolution of terpenoid natural product diversity. Eight AtTPS genes are interrupted and are likely to be inactive pseudogenes. The localization of AtTPS genes on all five chromosomes reflects the dynamics of the Arabidopsis genome; however, several AtTPS genes are clustered and organized in tandem repeats. Furthermore, some AtTPS genes are localized with prenyltransferase genes (AtGGPPS, geranylgeranyl diphosphate synthase) in contiguous genomic clusters encoding consecutive steps in terpenoid biosynthesis. The clustered organization may have implications for TPS gene evolution and the evolution of pathway segments for the synthesis of terpenoid natural products. Phylogenetic analyses highlight events in the divergence of the TPS paralogs and suggest orthologous genes and a model for the evolution of the TPS gene family.

368 citations


Cited by
More filters
01 Jun 2012
TL;DR: SPAdes as mentioned in this paper is a new assembler for both single-cell and standard (multicell) assembly, and demonstrate that it improves on the recently released E+V-SC assembler and on popular assemblers Velvet and SoapDeNovo (for multicell data).
Abstract: The lion's share of bacteria in various environments cannot be cloned in the laboratory and thus cannot be sequenced using existing technologies. A major goal of single-cell genomics is to complement gene-centric metagenomic data with whole-genome assemblies of uncultivated organisms. Assembly of single-cell data is challenging because of highly non-uniform read coverage as well as elevated levels of sequencing errors and chimeric reads. We describe SPAdes, a new assembler for both single-cell and standard (multicell) assembly, and demonstrate that it improves on the recently released E+V-SC assembler (specialized for single-cell data) and on popular assemblers Velvet and SoapDeNovo (for multicell data). SPAdes generates single-cell assemblies, providing information about genomes of uncultivatable bacteria that vastly exceeds what may be obtained via traditional metagenomics studies. SPAdes is available online ( http://bioinf.spbau.ru/spades ). It is distributed as open source software.

10,124 citations

Journal ArticleDOI
TL;DR: Circos uses a circular ideogram layout to facilitate the display of relationships between pairs of positions by the use of ribbons, which encode the position, size, and orientation of related genomic elements.
Abstract: We created a visualization tool called Circos to facilitate the identification and analysis of similarities and differences arising from comparisons of genomes. Our tool is effective in displaying variation in genome structure and, generally, any other kind of positional relationships between genomic intervals. Such data are routinely produced by sequence alignments, hybridization arrays, genome mapping, and genotyping studies. Circos uses a circular ideogram layout to facilitate the display of relationships between pairs of positions by the use of ribbons, which encode the position, size, and orientation of related genomic elements. Circos is capable of displaying data as scatter, line, and histogram plots, heat maps, tiles, connectors, and text. Bitmap or vector images can be created from GFF-style data inputs and hierarchical configuration files, which can be easily generated by automated tools, making Circos suitable for rapid deployment in data analysis and reporting pipelines.

8,315 citations

01 Jan 2007

4,037 citations

Journal ArticleDOI
Gerald A. Tuskan1, Gerald A. Tuskan2, Stephen P. DiFazio2, Stephen P. DiFazio3, Stefan Jansson4, Joerg Bohlmann5, Igor V. Grigoriev6, Uffe Hellsten6, Nicholas H. Putnam6, Steven G. Ralph5, Stephane Rombauts7, Asaf Salamov6, Jacquie Schein, Lieven Sterck7, Andrea Aerts6, Rishikeshi Bhalerao4, Rishikesh P. Bhalerao8, Damien Blaudez9, Wout Boerjan7, Annick Brun9, Amy M. Brunner10, Victor Busov11, Malcolm M. Campbell12, John E. Carlson13, Michel Chalot9, Jarrod Chapman6, G.-L. Chen2, Dawn Cooper5, Pedro M. Coutinho14, Jérémy Couturier9, Sarah F. Covert15, Quentin C. B. Cronk5, R. Cunningham2, John M. Davis16, Sven Degroeve7, Annabelle Déjardin9, Claude W. dePamphilis13, John C. Detter6, Bill Dirks17, Inna Dubchak6, Inna Dubchak18, Sébastien Duplessis9, Jürgen Ehlting5, Brian E. Ellis5, Karla C Gendler19, David Goodstein6, Michael Gribskov20, Jane Grimwood21, Andrew Groover22, Lee E. Gunter2, Björn Hamberger5, Berthold Heinze, Yrjö Helariutta23, Yrjö Helariutta8, Yrjö Helariutta24, Bernard Henrissat14, D. Holligan15, Robert A. Holt, Wenyu Huang6, N. Islam-Faridi22, Steven J.M. Jones, M. Jones-Rhoades25, Richard A. Jorgensen19, Chandrashekhar P. Joshi11, Jaakko Kangasjärvi23, Jan Karlsson4, Colin T. Kelleher5, Robert Kirkpatrick, Matias Kirst16, Annegret Kohler9, Udaya C. Kalluri2, Frank W. Larimer2, Jim Leebens-Mack15, Jean-Charles Leplé9, Philip F. LoCascio2, Y. Lou6, Susan Lucas6, Francis Martin9, Barbara Montanini9, Carolyn A. Napoli19, David R. Nelson26, C D Nelson22, Kaisa Nieminen23, Ove Nilsson8, V. Pereda9, Gary F. Peter16, Ryan N. Philippe5, Gilles Pilate9, Alexander Poliakov18, J. Razumovskaya2, Paul G. Richardson6, Cécile Rinaldi9, Kermit Ritland5, Pierre Rouzé7, D. Ryaboy18, Jeremy Schmutz21, J. Schrader27, Bo Segerman4, H. Shin, Asim Siddiqui, Fredrik Sterky, Astrid Terry6, Chung-Jui Tsai11, Edward C. Uberbacher2, Per Unneberg, Jorma Vahala23, Kerr Wall13, Susan R. Wessler15, Guojun Yang15, T. Yin2, Carl J. Douglas5, Marco A. Marra, Göran Sandberg8, Y. Van de Peer7, Daniel S. Rokhsar17, Daniel S. Rokhsar6 
15 Sep 2006-Science
TL;DR: The draft genome of the black cottonwood tree, Populus trichocarpa, has been reported in this paper, with more than 45,000 putative protein-coding genes identified.
Abstract: We report the draft genome of the black cottonwood tree, Populus trichocarpa. Integration of shotgun sequence assembly with genetic mapping enabled chromosome-scale reconstruction of the genome. More than 45,000 putative protein-coding genes were identified. Analysis of the assembled genome revealed a whole-genome duplication event; about 8000 pairs of duplicated genes from that event survived in the Populus genome. A second, older duplication event is indistinguishably coincident with the divergence of the Populus and Arabidopsis lineages. Nucleotide substitution, tandem gene duplication, and gross chromosomal rearrangement appear to proceed substantially more slowly in Populus than in Arabidopsis. Populus has more protein-coding genes than Arabidopsis, ranging on average from 1.4 to 1.6 putative Populus homologs for each Arabidopsis gene. However, the relative frequency of protein domains in the two genomes is similar. Overrepresented exceptions in Populus include genes associated with lignocellulosic wall biosynthesis, meristem development, disease resistance, and metabolite transport.

4,025 citations

Journal ArticleDOI
14 Jan 2010-Nature
TL;DR: An accurate soybean genome sequence will facilitate the identification of the genetic basis of many soybean traits, and accelerate the creation of improved soybean varieties.
Abstract: Soybean (Glycine max) is one of the most important crop plants for seed protein and oil content, and for its capacity to fix atmospheric nitrogen through symbioses with soil-borne microorganisms. We sequenced the 1.1-gigabase genome by a whole-genome shotgun approach and integrated it with physical and high-density genetic maps to create a chromosome-scale draft sequence assembly. We predict 46,430 protein-coding genes, 70% more than Arabidopsis and similar to the poplar genome which, like soybean, is an ancient polyploid (palaeopolyploid). About 78% of the predicted genes occur in chromosome ends, which comprise less than one-half of the genome but account for nearly all of the genetic recombination. Genome duplications occurred at approximately 59 and 13 million years ago, resulting in a highly duplicated genome with nearly 75% of the genes present in multiple copies. The two duplication events were followed by gene diversification and loss, and numerous chromosome rearrangements. An accurate soybean genome sequence will facilitate the identification of the genetic basis of many soybean traits, and accelerate the creation of improved soybean varieties.

3,743 citations