| Challenge: | The Hispanic population in the United States is up to 15-20% of the nation's total population . due to its proximity to the US-Mexico border, Hispanicas have more presence in the Southwest of the country . |
| Approach: | They propose to create a text corpus of the Spanish and English spoken in the US Midwest by different types of bilinguals. |
| Outcome: | The proposed corpus contains short stories narrated in Spanish and in English by 72 speakers representing different types of bilinguals: early simultaneous bilinguals, early sequential bilinguals and late second language learners. |
Similar Papers
A Closer Look at Clustering Bilingual Comparable Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for clustering comparable corpora are not suitable for bilingual corpors. |
| Approach: | They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia . |
| Outcome: | The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans . |
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)
Copied to clipboard
Nayla Escribano, Jon Ander Gonzalez, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Simón Peña-Fernández, Olatz Perez-de-Viñaspre, Rodrigo Agerri
| Challenge: | a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research . |
| Approach: | They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches. |
| Outcome: | The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
EPIC UdS - Creation and Applications of a Simultaneous Interpreting Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | EPIC UdS is a multilingual corpus of simultaneous interpreting for English, German and Spanish. |
| Approach: | They describe the creation and annotation of EPIC UdS, a multilingual corpus of simultaneous interpreting for English, German and Spanish. |
| Outcome: | The proposed corpus includes transcripts suitable for research on more than one language pair and on interpreting with regard to German. |
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)
Copied to clipboard
| Challenge: | Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power . |
| Approach: | They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus. |
| Outcome: | The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks. |
| Approach: | They present MLSUM, the first large-scale MultiLingual SUMmarization dataset. |
| Outcome: | The proposed dataset contains 1.5M+ article/summary pairs in five different languages. |
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage. |
| Approach: | They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards. |
| Outcome: | The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French. |
Auto-hMDS: Automatic Construction of a Large Heterogeneous Multilingual Multi-Document Summarization Corpus (L18-1)
Copied to clipboard
| Challenge: | Existing datasets for automatic text summarization are small and focused on newswires. |
| Approach: | They propose to automatically generate a large multilingual multi-document summarization corpus using Wikipedia articles as summaries and to automatically search for appropriate source documents. |
| Outcome: | The proposed corpus contains 7,316 topics in English and German with different summary lengths and number of source documents. |
The Natural Stories Corpus (L18-1)
Copied to clipboard
Richard Futrell, Edward Gibson, Harry J. Tily, Idan Blank, Anastasia Vishnevetsky, Steven Piantadosi, Evelina Fedorenko
| Challenge: | Existing corpora of naturalistic text do not contain the low-frequency syntactic constructions needed to distinguish between theories. |
| Approach: | They propose to compare models of language processing by comparing their ability to predict behavioral and neural measures of processing difficulty to corpora of naturalistic text. |
| Outcome: | The proposed corpus contains low-frequency syntactic constructions while sounding fluent to native speakers. |