Challenge: Existing methods to learn general language representations from large volumes of unlabeled text have been used to improve multilingual NLP.
Approach: They propose to use a spatial arrangement method to generate large-scale evaluation datasets that balance cross-lingual alignment with language specificity.
Outcome: The proposed method produces semantic verb classes and fine-grained similarity scores for nearly 130 thousand verb pairs.

Similar Papers

Spatial Multi-Arrangement for Clustering and Multi-way Similarity Dataset Construction (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for creating large-scale semantic similarity resources are slow and expensive . a large verb similarity dataset is available for a number of verbs, but not for English.
Approach: They propose a method for fast bottom-up creation of large-scale semantic similarity resources . they leverage semantic intuitions of native speakers and adapt a spatial multi-arrangement approach to lexical stimuli.
Outcome: The proposed approach produces a large-scale verb similarity dataset containing similarity scores for 29,721 unique verb pairs and 825 target verbs.
Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering (L18-1)

Copied to clipboard

Challenge: Existing methods for creating verbal classifications are limited or non-existent in most languages . a range of automatic verb classification approaches have been proposed, but high-quality resources are needed .
Approach: They propose to use top-up semantic clustering to extract syntactic and semantic information from verbs in English, Polish and Croatian.
Outcome: The proposed classifications in English, Polish and Croatian are compared with other languages.
Representing Verbs with Visual Argument Vectors (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes.
Approach: They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities.
Outcome: The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models.
A Closer Look at Clustering Bilingual Comparable Corpora (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for clustering comparable corpora are not suitable for bilingual corpors.
Approach: They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia .
Outcome: The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans .
Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding (D18-1)

Copied to clipboard

Challenge: a new approach to multilingual word embedding is needed to achieve this goal . a multilingual common semantic space is a language-agnostic semantic continuous space .
Approach: They propose a multilingual common semantic space where words from multiple languages are mapped into a shared space so that resources and knowledge can be shared across languages.
Outcome: The proposed approach achieves 14.6% absolute F-score gain over state-of-the-art methods on cross-lingual direct transfer.
Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced significantly in understanding human text, but semantic representations remain crucial for various applications.
Approach: They introduce a multilingual semantic layer which decouples from disambiguation and external inventories and simplifies the task.
Outcome: The proposed model reduces performance gap between languages and annotators by enabling them to understand semantic relations between concepts in any language.
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)

Copied to clipboard

Challenge: MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement.
Approach: They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs.
Outcome: The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs.
Exploring Alignment in Shared Cross-lingual Spaces (2024.acl-long)

Copied to clipboard

Challenge: a new study examines the degree of alignment between languages in multilingual embeddings . cross-lingual embeds are designed to encode linguistic concepts that bridge equivalent semantic meaning . a comprehensive approach is needed to address these questions.
Approach: They employ clustering to uncover latent concepts within multilingual models . they introduce two metrics to quantify alignment and overlap of these concepts .
Outcome: The proposed model can capture linguistic nuances across languages, but is not language-agnostic? the proposed model is able to capture nuances in multiple languages, the authors say.
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)

Copied to clipboard

Challenge: 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process.
Approach: They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments .
Outcome: The proposed tool can be set up for any parallel corpus and explores its quality and properties.
Language Directions in Multilingual LLMs: A Layer-wise Diagnostic Study of Token Alignment and Pretraining Imprint (2026.acl-srw)

Copied to clipboard

Challenge: Using a unified probing framework, we analyze six multilingual LLMs across five languages.
Approach: They analyze multilingual representations across five languages and analyze their behavior . they find that accuracy rises by +73.5 to +80.7 points from L0 to L1 on average .
Outcome: The proposed framework enables a consistent and substantial early jump in accuracy across models . the token–language alignment measures where vocabulary sharing peaks .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations