Challenge: The Hispanic population in the United States is up to 15-20% of the nation's total population . due to its proximity to the US-Mexico border, Hispanicas have more presence in the Southwest of the country .
Approach: They propose to create a text corpus of the Spanish and English spoken in the US Midwest by different types of bilinguals.
Outcome: The proposed corpus contains short stories narrated in Spanish and in English by 72 speakers representing different types of bilinguals: early simultaneous bilinguals, early sequential bilinguals and late second language learners.

Similar Papers

A Closer Look at Clustering Bilingual Comparable Corpora (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for clustering comparable corpora are not suitable for bilingual corpors.
Approach: They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia .
Outcome: The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans .
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research .
Approach: They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches.
Outcome: The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
EPIC UdS - Creation and Applications of a Simultaneous Interpreting Corpus (2022.lrec-1)

Copied to clipboard

Challenge: EPIC UdS is a multilingual corpus of simultaneous interpreting for English, German and Spanish.
Approach: They describe the creation and annotation of EPIC UdS, a multilingual corpus of simultaneous interpreting for English, German and Spanish.
Outcome: The proposed corpus includes transcripts suitable for research on more than one language pair and on interpreting with regard to German.
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)

Copied to clipboard

Challenge: Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power .
Approach: They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus.
Outcome: The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)

Copied to clipboard

Challenge: Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage.
Approach: They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards.
Outcome: The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French.
Auto-hMDS: Automatic Construction of a Large Heterogeneous Multilingual Multi-Document Summarization Corpus (L18-1)

Copied to clipboard

Challenge: Existing datasets for automatic text summarization are small and focused on newswires.
Approach: They propose to automatically generate a large multilingual multi-document summarization corpus using Wikipedia articles as summaries and to automatically search for appropriate source documents.
Outcome: The proposed corpus contains 7,316 topics in English and German with different summary lengths and number of source documents.
The Natural Stories Corpus (L18-1)

Copied to clipboard

Challenge: Existing corpora of naturalistic text do not contain the low-frequency syntactic constructions needed to distinguish between theories.
Approach: They propose to compare models of language processing by comparing their ability to predict behavioral and neural measures of processing difficulty to corpora of naturalistic text.
Outcome: The proposed corpus contains low-frequency syntactic constructions while sounding fluent to native speakers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations