Challenge: a corpus of academic texts provides lexical combinations for the production of academic text.
Approach: They describe the extraction of data from a corpus of academic texts and the use of those data to develop a lexical tool oriented to the production of academic text.
Outcome: The proposed tool will provide indications as to how to use vocabulary typical of the academic genre in order to build complete texts.

Similar Papers

A review of Spanish corpora annotated with negation (C18-1)

Copied to clipboard

Challenge: Existing corpora annotated with negation information are small and not always compatible . negation is a linguistic phenomenon that is not addressed in English .
Approach: They review existing corpora annotated with negation in Spanish and analyze compatibility . they propose to develop a supervised negation processing system for Spanish .
Outcome: The proposed system will not be able to merge the small corpora in Spanish due to lack of compatibility in annotations.
ALEXSIS: A Dataset for Lexical Simplification in Spanish (2022.lrec-1)

Copied to clipboard

Challenge: Lexical Simplification is the process of replacing difficult words with easier synonyms while preserving the original information and meaning.
Approach: They introduce ALEXSIS, a dataset for Lexical Simplification, and use it to benchmark Lexical simplification systems in Spanish.
Outcome: The proposed dataset compares three approaches to Lexical Simplification in Spanish and a previous dataset for English.
PUCP-Metrix: An Open-source and Comprehensive Toolkit for Linguistic Analysis of Spanish Texts (2026.eacl-demo)

Copied to clipboard

Challenge: Existing tools for linguistic analysis of Spanish texts lack linguistic features for interpretability and tasks that involve style, structure, and readability.
Approach: They propose to use PUCP-Metrix to analyze Spanish texts in a language repository.
Outcome: The proposed toolkit performs better on automated readability assessments and machine-generated text detection tasks than existing repositories and strong neural baselines.
Corpus Building and Evaluation of Aspect-based Opinion Summaries from Tweets in Spanish (L18-1)

Copied to clipboard

Challenge: a corpus of Spanish extractive and abstractive summaries of opinions is presented . the goal is to analyze the summary content and to show how different they are written .
Approach: They present a corpus of Spanish extractive and abstractive summaries of opinions . they analyze the summary agreement between them and their aspect coverage and sentiment orientation .
Outcome: The presented corpus of Spanish extractive and abstractive summaries is a reference for academic research.
Pay Attention when you Pay the Bills. A Multilingual Corpus with Dependency-based and Semantic Annotation of Collocations. (P19-1)

Copied to clipboard

Challenge: resulting corpus can be useful for different NLP tasks such as natural language understanding or natural language generation.
Approach: They propose to annotate 155k tokens and 1,526 collocations in context in a multilingual corpus in English, Portuguese, and Spanish.
Outcome: The new corpus can be used to evaluate different approaches for collocation identification, which can be useful for different NLP tasks such as natural language understanding or natural language generation.
A Repository of Corpora for Summarization (L18-1)

Copied to clipboard

Challenge: Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task.
Approach: They propose a repository containing corpora available to train and evaluate automatic summarization systems.
Outcome: The proposed system is based on a repository of corpora available for summarization tasks.
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)

Copied to clipboard

Challenge: Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power .
Approach: They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus.
Outcome: The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences.
The ICoN Corpus of Academic Written Italian (L1 and L2) (L18-1)

Copied to clipboard

Challenge: a corpus of academic written Italian is described in this paper . the corpus includes 2,115,000 tokens written by students having Italian as L2 .
Approach: They describe the ICoN corpus, a corpus of academic written Italian . it includes 2,115,000 tokens written by students having Italian as L2 and 1,769,000 tokens by students with Italian as a L1 .
Outcome: The ICoN corpus includes 2,115,000 tokens written by students having Italian as L2 and 1,769,000 tokens by students with Italian as a L1 . the corpus can be queried online while its complete contents are available on request for research purposes.
LexComSpaL2: A Lexical Complexity Corpus for Spanish as a Foreign Language (2024.lrec-main)

Copied to clipboard

Challenge: 58,240 annotations are available for learners of Spanish as a foreign/second language (L2).
Approach: They propose a corpus which can be employed to train personalised word-level difficulty classifiers for learners of Spanish as a foreign/second language (L2).
Outcome: The proposed model can train personalised word-level difficulty classifiers for learners of Spanish as a foreign/second language (L2) using a customised version of the 5-point lexical complexity prediction scale.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations