A multi-source domain annotation pipeline for quantitative metagenomic and metatranscriptomic functional profiling - Sorbonne Université Access content directly
Journal Articles Microbiome Year : 2018

A multi-source domain annotation pipeline for quantitative metagenomic and metatranscriptomic functional profiling

Abstract

Background: Biochemical and regulatory pathways have until recently been thought and modelled within one cell type, one organism and one species. This vision is being dramatically changed by the advent of whole microbiome sequencing studies, revealing the role of symbiotic microbial populations in fundamental biochemical functions. The new landscape we face requires the reconstruction of biochemical and regulatory pathways at the community level in a given environment. In order to understand how environmental factors affect the genetic material and the dynamics of the expression from one environment to another, we want to evaluate the quantity of gene protein sequences or transcripts associated to a given pathway by precisely estimating the abundance of protein domains, their weak presence or absence in environmental samples. Results: MetaCLADE is a novel profile-based domain annotation pipeline based on a multi-source domain annotation strategy. It applies directly to reads and improves identification of the catalog of functions in microbiomes. MetaCLADE is applied to simulated data and to more than ten metagenomic and metatranscriptomic datasets from different environments where it outperforms InterProScan in the number of annotated domains. It is compared to the state-of-the-art non-profile-based and profile-based methods, UProC and HMM-GRASPx, showing complementary predictions to UProC. A combination of MetaCLADE and UProC improves even further the functional annotation of environmental samples. Conclusions: Learning about the functional activity of environmental microbial communities is a crucial step to understand microbial interactions and large-scale environmental impact. MetaCLADE has been explicitly designed for metagenomic and metatranscriptomic data and allows for the discovery of patterns in divergent sequences, thanks to its multi-source strategy. MetaCLADE highly improves current domain annotation methods and reaches a fine degree of accuracy in annotation of very different environments such as soil and marine ecosystems, ancient metagenomes and human tissues.
Fichier principal
Vignette du fichier
s40168-018-0532-2.pdf (2.94 Mo) Télécharger le fichier
Origin : Publication funded by an institution
Loading...

Dates and versions

hal-01875977 , version 1 (18-09-2018)

Licence

Attribution

Identifiers

Cite

Ari Ugarte, Riccardo Vicedomini, Juliana Bernardes, Alessandra Carbone. A multi-source domain annotation pipeline for quantitative metagenomic and metatranscriptomic functional profiling. Microbiome, 2018, 6, pp.149. ⟨10.1186/s40168-018-0532-2⟩. ⟨hal-01875977⟩
162 View
97 Download

Altmetric

Share

Gmail Facebook X LinkedIn More