Challenge: Using corpora for second language acquisition has become more and more common . corporata are used to study morpho-syntactic phenomena in English as a foreign language .
Approach: They propose to use a bilingual corpus of French learners of Korean and Korean learners of French to provide a translated and annotated corpus to the scientific community.
Outcome: The proposed corpus can be used for a wide array of purposes in the field of theoretical but also applied linguistics.

Similar Papers

Constructing Korean Learners’ L2 Speech Corpus of Seven Languages for Automatic Pronunciation Assessment (2024.lrec-main)

Copied to clipboard

Challenge: Multilingual L2 speech corpora for automatic speech assessment are currently available, but lack comprehensive annotations of L2 from non-native speakers of various languages.
Approach: They propose to use Korean learners’ L2 speech corpus of seven languages to develop automatic speech assessment.
Outcome: The proposed corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each in French, German, Spanish, and Russian).
Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)

Copied to clipboard

Challenge: a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes)
Approach: They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data.
Outcome: The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
L1-L2 Parallel Treebank of Learner Chinese: Overused and Underused Syntactic Structures (L18-1)

Copied to clipboard

Challenge: Currently, the treebank consists of 600 L2 sentences and 697 L1 sentences.
Approach: They propose to use "L1-L2 parallel treebanks" to facilitate analyses of learner language.
Outcome: The proposed treebank consists of 600 L2 sentences and 697 L1 sentences.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)

Copied to clipboard

Challenge: Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding.
Approach: They propose to augment an existing (monolingual) corpus: LibriSpeech.
Outcome: The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned.
Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words? (L18-1)

Copied to clipboard

Challenge: a recent study suggests that a classifier trained on unknown words may yield better results for L2 learners.
Approach: They propose to use a supervised learning classifier to predict word complexity in Korean . they propose to train models on annotated corpus of unknown words with 71 % precision .
Outcome: The proposed model recalls 80 % of unknown words with 71 % precision.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)

Copied to clipboard

Challenge: a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources.
Approach: They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset .
Outcome: The proposed subset of the Reuters corpus has balanced class priors for eight languages.
JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages (P19-1)

Copied to clipboard

Challenge: a shortage of parallel data in low-resource languages creates a bottleneck for cross-lingual transfer . a massive collection of parallel texts for over 300 diverse languages is our main contribution .
Approach: They propose a parallel corpus of over 300 languages with 100 thousand parallel sentences per language pair on average.
Outcome: The proposed dataset can be used to build cross-lingual word embeddings and multi-source part-of-speech projections.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations