Challenge: LARA is an open source project that aims to support easy conversion of plain texts into online versions suitable for use by language learners.
Approach: They propose to support easy conversion of plain texts into online versions suitable for use by language learners.
Outcome: The proposed platform is suitable for creating texts in multiple languages via crowdsourcing techniques that can be used for teaching a language via reading and listening.

Similar Papers

Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
LaRA: Large Rank Adaptation for Speech and Text Cross-Modal Learning in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to integrate speech and text capabilities into large language models (LLMs) require significantly larger ranks comparable to the pretrained weights to accommodate the complexities of speech-text cross-modality learning.
Approach: They propose a large-rank adaptive approach for cross-modal integration of speech and text into large language models (LLMs) it uses a Hi-Fi vocoder to synthesize speech waveforms from the generated speech units.
Outcome: The proposed model can be extended to other cross-modal applications.
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Approach: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Outcome: This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)

Copied to clipboard

Challenge: a large corpus of online content has been developed via large-scale crowdsourcing.
Approach: They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform.
Outcome: The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines.
WAT2019: English-Hindi Translation on Hindi Visual Genome Dataset (D19-52)

Copied to clipboard

Challenge: A multimodal translation is a task of translating a source language to a target language . a parallel text corpus and images are used to represent the contextual details of the text .
Approach: They compare a multimodal approach to a parallel text corpus and image caption generation approach to translate text in English to Hindi as a part of WAT2019 shared task.
Outcome: The proposed approach improves the translation of English to Hindi in three tasks . the proposed system is based on a neural network and can be used in healthcare, government, disaster management, etc.
LARA: Linguistic-Adaptive Retrieval-Augmentation for Multi-Turn Intent Classification (2024.emnlp-industry)

Copied to clipboard

Challenge: Multi-turn intent classification is challenging due to the complexity and evolving nature of conversational contexts . lack of data on multi-turn datasets makes it difficult to collect multi-turned datasets a challenge .
Approach: They propose a framework for multi-turn intent classification that integrates a retrieval-augmented mechanism with a fine-tuned smaller model.
Outcome: The proposed framework improves accuracy on multi-turn intent classification tasks across six languages.
CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets in multimodal language are limited and disproportionately affect native speakers of other languages . authors propose a large-scale dataset for Spanish, Portuguese, German and French .
Approach: They propose a large-scale multimodal language dataset for Spanish, Portuguese, German and French.
Outcome: The proposed dataset is the largest of its kind with 40,000 total labelled sentences . it covers a diverse set topics and speakers and carries supervision of 20 labels including sentiment, emotions, and attributes.
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch .
Approach: They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency.
Outcome: The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations