Saltar para o conteúdo principal

Access the national information on scientific and technological funding in Portugal

Activity Funded

Automated Annotation and Integration of Biobank Data Using LLMs and Ontologies

Anotação automatica e integração de dados de BioBanco usando LLMs e ontologias

Reference
2024.07694.IACDC
Project Start Date
2025-05-01
Project End Date
2026-01-31
Principal Investigator
Scientific Area
Medical and health sciences
Funding Program
Inteligência Artificial, Ciência dos Dados e Cibersegurança de relevância na Administração Pública

Abstract

In recent years, the proliferation of unstructured data in biobanks has presented significant challenges in data management, integration, and utilization.  This project aims to develop an advanced Natural Language Processing (NLP) framework supported by Large Language Models (LLMs) to automate the extraction and structuring of unstructured text data within a biobank context, aligning it with established biomedical ontologies such as the NCI Thesaurus. The initiative addresses the critical issue of incomplete and non-standardized medical records that hinder interoperability, reusability, and predictive model development. By leveraging the semantic capabilities of state-of-the-art LLMs, we seek to transform unstructured annotations into consistent, structured formats that adhere to specific ontological standards. The information available in biobanks offers a significant and diverse knowledge base to support the creation of ontologies essential for interoperability at a national and international level between health systems and how information is exchanged and understood. The deployment of this NLP framework is anticipated to offer numerous benefits. Primarily, the automation of data management and curation processes is projected to significantly reduce the manual effort required, freeing up resources for more strategic tasks. Furthermore, one key benefit will be on interoperability, allowing datasets to be structured and aligned to the same semantic reference. Finally, it will support predictive analysis by providing researchers with high-quality, structured data that can be used to develop machine learning models for predicting health outcomes and disease progression. The project encompasses several key phases: a data collection phase, a translation phase that will convert Portuguese medical text to English, ensuring compatibility with widely-used English-language ontologies and maximizing the utility of existing NLP resources, and a training phase consisting of  a zero-shot approach to establish a baseline performance benchmark, the creation of high-quality datasets for fine-tuning through expert annotation, and the iterative refinement of the model to increase its accuracy and reliability. In parallel, we will develop an ontology aligner leveraging the semantic abilities of LLMs, by utilizing the embedding distances as similarity measures between semantic hierarchical structures. To demonstrate the practical utility of the developed framework, we will apply it to our biobank cancer dataset and produce a structured version of the dataset in two different ontologies, showcasing the model’s ability to enhance data interoperability and add value to disparate datasets. The outcomes of this project are expected to significantly reduce manual effort in data curation, support predictive analytics, and foster collaboration among researchers by providing a robust, scalable solution for structured data extraction and annotation. This initiative contributes to the broader goal of improving biomedical research through enhanced data consistency and accessibility. One key benefit of this framework will be interoperability between data sources that do not necessarily utilize the same controlled vocabularies. The framework will include an innovative mechanism that leverages the semantic abilities of LLMs to map concepts between ontologies. By utilizing LLM embeddings to capture semantic relationships, the framework can effectively translate and align data from dissimilar sources into a unified structure. This capability ensures seamless integration and comparison of diverse datasets, promoting more robust data sharing and collaboration across various research institutions and clinical settings. Consequently, this will enhance the ability to conduct comprehensive biomedical research and improve the accuracy and efficacy of predictive models, contributing to translational objectives with great impact in public administration. In summary, the proposed NLP framework represents a significant advancement in the management and utilization of medical data available in biobank. By leveraging the capabilities of LLMs, it aims to transform unstructured text into valuable, actionable information, paving the way for improved research outcomes and participant engagement. The project's focus on automation, standardization, and collaboration aligns with the broader goals of modernizing biobank operations and ensuring their alignment with international data management standards. Through these efforts, the project seeks to establish a new paradigm in biobank data management that maximizes the potential of available data while maintaining the highest standards of ethical and legal compliance.

Institutions

Main Institutions

  • Fundação GIMM - Gulbenkian Institute for Molecular Medicine (GIMM)
  • Universidade de Lisboa Instituto de Medicina Molecular João Lobo Antunes (IMM)

Other Institutions

  • Instituto Gulbenkian de Ciência (IGC)

Funding 124.135,25 €

Fundação para a Ciência e a Tecnologia (FCT) - Portugal

0,00 €

União Europeia - Estrutura de Missão Recuperar Portugal (UE - EMRP)

124.135,25 €
  • 2025-05-01 a 2026-01-31
  • 2024.07694.IACDC