Challenge: resulting corpus can be useful for different NLP tasks such as natural language understanding or natural language generation.
Approach: They propose to annotate 155k tokens and 1,526 collocations in context in a multilingual corpus in English, Portuguese, and Spanish.
Outcome: The new corpus can be used to evaluate different approaches for collocation identification, which can be useful for different NLP tasks such as natural language understanding or natural language generation.

Similar Papers

The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Evaluating language models for the retrieval and categorization of lexical collocations (2021.eacl-main)

Copied to clipboard

Challenge: Lexical collocations are idiosyncratic combinations of two syntactically bound lexical items.
Approach: They perform an exhaustive analysis of current language models for collocation understanding . they first construct a dataset of apparitions of lexical collocations in context .
Outcome: The proposed models perform well in distinguishing light verb constructions, especially if the collocation’s first argument acts as subject, but often fail to distinguish, first, different syntactic structures within the same semantic category, and second, fine-grained semantic categories which restrict the use of small sets of valid collocates for a given base.
DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for graded contextual word meaning annotation have not been implemented yet.
Approach: They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset.
Outcome: The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments.
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space.
Approach: They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data.
Outcome: The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures.
A New Annotated Portuguese/Spanish Corpus for the Multi-Sentence Compression Task (L18-1)

Copied to clipboard

Challenge: Existing corpus for Multi-sentence Compression (MSC) tasks is limited to English . a dataset is available for MSC tasks in the French language .
Approach: They propose a new corpus for Multi-Sentence Compression task in Portuguese and Spanish.
Outcome: The proposed corpus is compared with two state-of-the-art systems in Portuguese and Spanish.
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers.
Approach: They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch.
Outcome: The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages.
Generation of a Spanish Artificial Collocation Error Corpus (L18-1)

Copied to clipboard

Challenge: collocations are combinations of two elements where one (the base) is freely chosen, despite the limitations of the other (collocate) current tools for collocation error detection and correction focus on collocation validation and identification of miscollocations .
Approach: They propose an algorithm for automatic generation of an artificial collocation error corpus of american English learners of Spanish that includes 17 different types of collocation errors.
Outcome: The proposed algorithm can detect and classify collocation errors in learners' writings . collocation error detection and correction has not received the attention it deserves .
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)

Copied to clipboard

Challenge: Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions.
Approach: They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge.
Outcome: The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian.
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences.
Approach: They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs.
Outcome: The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering .
Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced significantly in understanding human text, but semantic representations remain crucial for various applications.
Approach: They introduce a multilingual semantic layer which decouples from disambiguation and external inventories and simplifies the task.
Outcome: The proposed model reduces performance gap between languages and annotators by enabling them to understand semantic relations between concepts in any language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations