Papers with CEF
European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management (L18-1)
Copied to clipboard
Andrea Lösch, Valérie Mapelli, Stelios Piperidis, Andrejs Vasiļjevs, Lilli Smal, Thierry Declerck, Eileen Schnur, Khalid Choukri, Josef van Genabith
| Challenge: | European Language Resource Coordination (ELRC) initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
| Approach: | They propose to initiate actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
| Outcome: | The European Language Resource Coordination (ELRC) consortium initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
Assessing Multilinguality of Publicly Accessible Websites (2022.lrec-1)
Copied to clipboard
| Challenge: | multilingualism on the Web is a problem not only at the world level, but also at the European and regional level. |
| Approach: | They propose a tool that automatically analyses the language diversity of the Web and propose indicators and methodologies to measure multilingualism of European websites. |
| Outcome: | The proposed tool can be independently run at set intervals and concludes that multilingualism on the Web is still a problem not only at the world level, but also at the European and regional level. |
Discovering Parallel Language Resources for Training MT Engines (L18-1)
Copied to clipboard
| Challenge: | Web crawling is an efficient way for compiling the monolingual, parallel and/or domain-specific corpora needed for machine translation and other HLT applications. |
| Approach: | They propose a system for compiling monolingual, parallel and/or domain-specific corpora . ILSP-FC is a web crawling system that generates bilingual lexica and terminology lists . |
| Outcome: | The ILSP Focused Crawler is a system developed by researchers at the IL SP/Athena RIC for the acquisition of such resources. |
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation (2026.acl-long)
Copied to clipboard
Tathagata Raha, Clement Christophe, Nada Saadi, Hamza A Javed, Marco AF Pimentel, Ronnie Rajan, Praveenkumar Kanithi
| Challenge: | Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks. |
| Approach: | They propose a cross-examination framework that generates verifiable questions from each text and performs a Cross-exam to derive three interpretable scores: Coverage, Conformity, and Consistency. |
| Outcome: | The proposed framework detects critical errors across translation, summarization and clinical note-generation and human expert validation shows it is reliable without gold references. |