
MuCoME:
Vielschichtiger Datenkorpus für KI Anwendungen in der Mikrobiellen Ökologie
- Laufzeit:
- 01.01.2027 - 31.12.2028
- Gesamtkoordination:
- none
- Projektleitung (IOW):
- Dr. Christiane Hassenrück
- Finanzierung:
- DFG - Deutsche Forschungsgemeinschaft
- Forschungsbereich:
- Projektpartner:
-
German Centre for Integrative Biodiversity Research (iDiv)Hochschule WismarEuropean Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI)Global Biodiversity Information Facility (GBIF)German Federation for Biological Data (GFBio)Leibniz Institute for Horticultural Sciences (IGZ)
The widespread adoption of metabarcoding for microbial community analyses combined with the existence of centralized sequence archives has resulted in the explosive growth of metabarcoding data in these public archives. However, turning this data into integrated and actionable knowledge has not been straightforward and data reuse remains lower than expected due to two major limitations: metadata essential for data reuse is often unavailable, incomplete, or incorrect, and data comparability is reduced by a lack of consistent, standardized bioinformatics pipelines. MuCoME addresses these limitations by (1) improving methods for the curation and harmonization of sequence data-associated metadata, (2) developing methods for data cleansing, aggregation, and curation of metabarcoding sequence data, and (3) enabling public access to the resulting data corpus through existing data infrastructures, ensuring long-term public availability to entry-level and expert users alike. Specifically, the data corpus will exist of three parts: the manually annotated gold standard that will be used to train a natural language processing model for the extraction and harmonization of metadata from scientific articles and other unstructured natural text associated with a metabarcoding dataset; the enriched metadata resulting from the application of this model to existing metabarcoding datasets made available at the European Nucleotide Archive (ENA); the fully-processed, harmonized microbial community data, integrating hundreds of thousands of metabarcoding datasets, published through the microbial community database (MiCoDa).