Activity Funded
Multilingual Portal for Specialized Languages: mining open data for cross-language information retrieval
Portal multilingue para línguas de especialidade: extração de recursos bilingues a partir de dados em acesso aberto
Details
Reference
PTDC/LLT-LIG/31113/2017
PTDC/LLT-LIG/31113/2017
Project Start Date
2019-11-11
2019-11-11
Project End Date
2023-06-30
2023-06-30
Scientific Area
Humanities
Humanities
Funding Program
Concurso para Financiamento de Projetos de Investigação Científica e Desenvolvimento Tecnológico em Todos os Domínios Científicos - 2017
Concurso para Financiamento de Projetos de Investigação Científica e Desenvolvimento Tecnológico em Todos os Domínios Científicos - 2017
Abstract
The technological revolution of the last decades has contributed to the consolidation of a new social paradigm known as knowledge society or information society. This paradigm is reflected in a globalized and multilingual world, full of economic, commercial, political, social and cultural relations, where professional specialization is a necessity. In this context, specialized languages, with their specific vocabulary, structures and grammatical functions, have an important role to play. As a result of this new paradigm, we have a wealth of text in several areas of specialization, such as the academic field due to the promotion of Open Science, which has received heavy national and European investment for the setting up of scientific repositories. However, the format and structure of these data have limitations in terms of textual processing.
It is our objective to reuse the material available in these academic repositories and in non-academic ones to create a portal for specialized languages that allows for more robust textual analyses. Firstly, we will extract texts in Portuguese and in English from the most representative disciplinary areas in the various repositories. Note that these texts are not translations of each other, but comparable texts, i.e., texts originally written in different languages but on the same topic. We will tag these texts with structural and morphosyntactic annotation and with metadata. Finally, we will build a user-friendly interface where users will be able to insert a term/expression, and then obtain the occurrences of that term in context. For a query in English, for instance, users may view the results in that language and have access to the translation equivalents of the searched term/expression. This second research mode will require the construction of a bilingual search engine capable of identifying translation equivalents from comparable texts. In the future, we will replicate the architecture of the portal for more languages and for the other Portuguese variants, using national networks of repositories from other countries.
This project, which brings together a multidisciplinary team (specialists in terminology, corpus linguistics, natural language processing and artificial intelligence), aims to serve a wide range of users (teachers, students, linguists, researchers, lexicographers, translators, etc.), contributing to improve the apprehension of linguistic conventions, the recognition of monolingual and bilingual discoursive phraseology, the development of specialized writing skills in native or foreign language, teaching and learning languages for specific purposes, specialized discourse analysis in heterogeneous communicative situations, and the creation of lexicographic and terminological resources.
Institutions
Main Institutions
- Universidade do Minho (UMinho)
Other Institutions
- Universidade do Minho (UMinho)
Funding 188.766,72 €
Fundação para a Ciência e a Tecnologia (FCT) - Portugal
32.251,22 €
União Europeia - Portugal2020 (UE - PT2020)
156.515,50 €