Large scale comparison of global gene expression patterns in human and mouse

doi:10.1186/GB-2010-11-12-R124

Citations

PDF

Open Access

More filters

Journal Article•DOI•

PhysioSpace: Relating Gene Expression Experiments from Heterogeneous Sources Using Shared Physiological Processes

[...]

Michael Lenz¹, Bernhard M. Schuldt¹, Franz-Josef Müller, Andreas Schuppert², Andreas Schuppert¹ - Show less +1 more•Institutions (2)

RWTH Aachen University¹, Bayer²

17 Oct 2013-PLOS ONE

TL;DR: Using PhysioSpace with clinical cancer datasets reveals that such data exhibits large heterogeneity in the number of significant signature associations, indicating shared biological functionalities in disease associated processes.

...read moreread less

Abstract: Relating expression signatures from different sources such as cell lines, in vitro cultures from primary cells and biopsy material is an important task in drug development and translational medicine as well as for tracking of cell fate and disease progression. Especially the comparison of large scale gene expression changes to tissue or cell type specific signatures is of high interest for the tracking of cell fate in (trans-) differentiation experiments and for cancer research, which increasingly focuses on shared processes and the involvement of the microenvironment. These signature relation approaches require robust statistical methods to account for the high biological heterogeneity in clinical data and must cope with small sample sizes in lab experiments and common patterns of co-expression in ubiquitous cellular processes. We describe a novel method, called PhysioSpace, to position dynamics of time series data derived from cellular differentiation and disease progression in a genome-wide expression space. The PhysioSpace is defined by a compendium of publicly available gene expression signatures representing a large set of biological phenotypes. The mapping of gene expression changes onto the PhysioSpace leads to a robust ranking of physiologically relevant signatures, as rigorously evaluated via sample-label permutations. A spherical transformation of the data improves the performance, leading to stable results even in case of small sample sizes. Using PhysioSpace with clinical cancer datasets reveals that such data exhibits large heterogeneity in the number of significant signature associations. This behavior was closely associated with the classification endpoint and cancer type under consideration, indicating shared biological functionalities in disease associated processes. Even though the time series data of cell line differentiation exhibited responses in larger clusters covering several biologically related patterns, top scoring patterns were highly consistent with a priory known biological information and separated from the rest of response patterns.

...read moreread less

21 citations

Journal Article•DOI•

Graph Laplacian Regularization With Procrustes Analysis for Sensor Node Localization

[...]

Abhishek Singh¹, Shekhar Verma¹•Institutions (1)

Indian Institute of Information Technology, Allahabad¹

15 Aug 2017-IEEE Sensors Journal

TL;DR: The proposed method works on the observation that noisy data lie on a higher dimension space even though the actual data are embedded in a low dimensional manifold and is able to reduce the noise by around 70%.

...read moreread less

Abstract: In this paper, we investigate the problem of sensor node localization and propose a non-linear semi supervised noise minimization algorithm through iterative manifold learning. The method works on the observation that noisy data lie on a higher dimension space even though the actual data are embedded in a low dimensional manifold. The collective labeled and unlabeled data are represented as a weighted graph. A prediction function is created based on the available labeled data along with manifold learning to exploit the intrinsic geometry. On top of prediction function, iterative feedback mechanism is used, which incrementally flattens the higher dimensional manifold. This reduces the error boundary in each stage for every data point. Result found to converge after a few iterations. This is followed by localized Procrustes analysis to further reduce the error. Experiment using TelosB motes and simulation with labeled and unlabeled data show that the proposed technique is able to reduce the noise, on an average, by around 70%. Results also show that the mechanism is able to localize the sensor nodes with high accuracy and outperforms the baseline method and LapRLS in different conditions.

...read moreread less

20 citations

Cites methods from "Large scale comparison of global ge..."

...Linear manifold learning methods like PCA [22]–[25] and MDS [26]–[29] are not suitable for non-linear data....
[...]

Journal Article•DOI•

ExpressionData - A public resource of high quality curated datasets representing gene expression across anatomy, development and experimental conditions

[...]

Philip Zimmermann, Stefan Bleuler, Oliver Laule, Florian Martin, Nikolai V. Ivanov, Prisca Campanoni, Karen Oishi, Nicolas Lugon-Moulin, Markus Wyss, Tomas Hruz¹, Wilhelm Gruissem¹ - Show less +7 more•Institutions (1)

ETH Zurich¹

31 Aug 2014-Biodata Mining

TL;DR: A new type of standardized datasets representative for the spatial and temporal dimensions of gene expression result from integrating expression data from a large number of globally normalized and quality controlled public experiments.

...read moreread less

Abstract: Reference datasets are often used to compare, interpret or validate experimental data and analytical methods. In the field of gene expression, several reference datasets have been published. Typically, they consist of individual baseline or spike-in experiments carried out in a single laboratory and representing a particular set of conditions. Here, we describe a new type of standardized datasets representative for the spatial and temporal dimensions of gene expression. They result from integrating expression data from a large number of globally normalized and quality controlled public experiments. Expression data is aggregated by anatomical part or stage of development to yield a representative transcriptome for each category. For example, we created a genome-wide expression dataset representing the FDA tissue panel across 35 tissue types. The proposed datasets were created for human and several model organisms and are publicly available at http://www.expressiondata.org .

...read moreread less

20 citations

Cites result from "Large scale comparison of global ge..."

...These results confirm previous findings on comparing human and mouse tissues based on datasets that were normalized differently and in which tissue samples are represented individually [18]....
[...]

Journal Article•DOI•

Global regulatory architecture of human, mouse and rat tissue transcriptomes

[...]

Ajay Prasad¹, Suchitra Suresh Kumar¹, Christophe Dessimoz², Christophe Dessimoz¹, Christophe Dessimoz³, Stefan Bleuler, Oliver Laule, Tomas Hruz¹, Wilhelm Gruissem¹, Philip Zimmermann¹ - Show less +6 more•Institutions (3)

ETH Zurich¹, University College London², Swiss Institute of Bioinformatics³

20 Oct 2013-BMC Genomics

TL;DR: Evidence is shown that tissue expression profiles, if combined with sequence similarity, can improve the correct assignment of functionally related homologs across species and demonstrate that tissue-specific regulation is the main determinant of transcriptome composition and is highly conserved across mammalian species.

...read moreread less

Abstract: Predicting molecular responses in human by extrapolating results from model organisms requires a precise understanding of the architecture and regulation of biological mechanisms across species. Here, we present a large-scale comparative analysis of organ and tissue transcriptomes involving the three mammalian species human, mouse and rat. To this end, we created a unique, highly standardized compendium of tissue expression. Representative tissue specific datasets were aggregated from more than 33,900 Affymetrix expression microarrays. For each organism, we created two expression datasets covering over 55 distinct tissue types with curated data from two independent microarray platforms. Principal component analysis (PCA) revealed that the tissue-specific architecture of transcriptomes is highly conserved between human, mouse and rat. Moreover, tissues with related biological function clustered tightly together, even if the underlying data originated from different labs and experimental settings. Overall, the expression variance caused by tissue type was approximately 10 times higher than the variance caused by perturbations or diseases, except for a subset of cancers and chemicals. Pairs of gene orthologs exhibited higher expression correlation between mouse and rat than with human. Finally, we show evidence that tissue expression profiles, if combined with sequence similarity, can improve the correct assignment of functionally related homologs across species. The results demonstrate that tissue-specific regulation is the main determinant of transcriptome composition and is highly conserved across mammalian species.

...read moreread less

20 citations

Cites background or result from "Large scale comparison of global ge..."

...For example, some studies suggested that orthologous genes have dissimilar expression patterns [3,4,9-11], while others reported congruent expression profiles [5-7,12-17]....
[...]
...[6,20], although here each category in the plot represents an average vector aggregated from a population of samples rather than plotting individual samples in the PCA....
[...]
...[5,7,16]), or to a larger but only partly overlapping set of tissues between human and mouse [6]....
[...]
...Most of these studies were restricted to comparing the human and mouse transcriptomes, thereby limiting the interpretation to a bilateral relationship without evidence from further organisms [3-7]....
[...]
...This result confirms previous findings in a comparison of human and mouse [6,21]....
[...]

Journal Article•DOI•

Rates of evolution in stress-related genes are associated with habitat preference in two Cardamine lineages

[...]

Lino Ometto¹, Mingai Li¹, Luisa Bresadola¹, Claudio Varotto¹•Institutions (1)

Edmund Mach Foundation¹

18 Jan 2012-BMC Evolutionary Biology

TL;DR: In this paper, the authors studied the role of positive and relaxed selection in the evolution of Cardamine genes and concluded that the selective pressures associated with the habitats typical of C. resedifolia and C. impatiens may have caused the rapid evolution of genes involved in cold response.

...read moreread less

Abstract: Elucidating the selective and neutral forces underlying molecular evolution is fundamental to understanding the genetic basis of adaptation. Plants have evolved a suite of adaptive responses to cope with variable environmental conditions, but relatively little is known about which genes are involved in such responses. Here we studied molecular evolution on a genome-wide scale in two species of Cardamine with distinct habitat preferences: C. resedifolia, found at high altitudes, and C. impatiens, found at low altitudes. Our analyses focussed on genes that are involved in stress responses to two factors that differentiate the high- and low-altitude habitats, namely temperature and irradiation. High-throughput sequencing was used to obtain gene sequences from C. resedifolia and C. impatiens. Using the available A. thaliana gene sequences and annotation, we identified nearly 3,000 triplets of putative orthologues, including genes involved in cold response, photosynthesis or in general stress responses. By comparing estimated rates of molecular substitution, codon usage, and gene expression in these species with those of Arabidopsis, we were able to evaluate the role of positive and relaxed selection in driving the evolution of Cardamine genes. Our analyses revealed a statistically significant higher rate of molecular substitution in C. resedifolia than in C. impatiens, compatible with more efficient positive selection in the former. Conversely, the genome-wide level of selective pressure is compatible with more relaxed selection in C. impatiens. Moreover, levels of selective pressure were heterogeneous between functional classes and between species, with cold responsive genes evolving particularly fast in C. resedifolia, but not in C. impatiens. Overall, our comparative genomic analyses revealed that differences in effective population size might contribute to the differences in the rate of protein evolution and in the levels of selective pressure between the C. impatiens and C. resedifolia lineages. The within-species analyses also revealed evolutionary patterns associated with habitat preference of two Cardamine species. We conclude that the selective pressures associated with the habitats typical of C. resedifolia may have caused the rapid evolution of genes involved in cold response.

...read moreread less

18 citations