Rollenwechsel-English: a large-scale semantic role corpus (L18-1)

Copied to clipboard

Challenge: The Rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.
Approach: They present a large corpus of automatically-labelled semantic frames extracted from ukWaC and BNC using Propbank roles.
Outcome: The rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.

Similar Papers

RRGparbank: A Parallel Role and Reference Grammar Treebank (2022.lrec-1)

Copied to clipboard

Challenge: Existing treebanks for Role and Reference Grammar (RRG) are not yet available.
Approach: They propose to use a multilingual parallel treebank for Role and Reference Grammar to apply RRG to large-scale corpus annotations of 1984 and its translations.
Outcome: The proposed treebank contains annotations of Orwell's 1984 and translations thereof.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
A Flexible and Easy-to-use Semantic Role Labeling Framework for Different Languages (C18-2)

Copied to clipboard

Challenge: DAMESRL is an open source framework for deep semantic role labeling . language-specific characteristics and the available amount of training data influence the optimal model structure .
Approach: They propose an open-source framework for deep semantic role labeling that is available under the Apache 2.0 license.
Outcome: The proposed framework is available under the Apache 2.0 license.
Construction of Large-scale English Verbal Multiword Expression Annotated Corpus (L18-1)

Copied to clipboard

Challenge: In this paper, we focus on verbal MWEs, whose accurate recognition is challenging because they could be discontinuous.
Approach: They conduct large-scale annotations of VMWEs on the Wall Street Journal portion of Ontonotes . they first construct a VMwe dictionary based on the english-language Wiktionary .
Outcome: The proposed resource annotates 7,833 VMWE instances belonging to various categories . the authors hope the results will help to develop models for MWE recognition and dependency parsing .
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder (2025.emnlp-main)

Copied to clipboard

Challenge: Prior research on linguistic mechanisms of large language models is limited by coarse granularity, limited analysis scale, and narrow focus.
Approach: They propose a framework for analyzing the linguistic mechanisms of large language models based on Sparse Auto-Encoders.
Outcome: The proposed framework extracts Chinese and English linguistic features across four dimensions . it uncovers intrinsic representations of linguistic knowledge in LLMs and can control outputs .
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
IEPile: Unearthing Large Scale Schema-Conditioned Information Extraction Corpus (2024.acl-short)

Copied to clipboard

Challenge: Large Language Models exhibit a significant performance gap in Information Extraction (IE) high-quality instruction data is the vital key for enhancing LLMs' specific capabilities .
Approach: They propose a bilingual (English and Chinese) IE instruction corpus that contains 0.32B tokens.
Outcome: The proposed model improves the performance of LLMs for IE with zero-shot generalization.
Trained on 100 million words and still in shape: BERT meets British National Corpus (2023.findings-eacl)

Copied to clipboard

Challenge: masked language models are trained on ever larger corpora, but pre-training on a modestly-sized but representative, well-balanced, and publicly available corpus can reach better performance than the original BERT model.
Approach: They propose an optimized LM architecture called LTG-BERT that can be used to train a competitive language model on a small and standardizable corpus.
Outcome: The proposed architecture outperforms the original English BERT model on a representative, well-balanced and publicly available corpus.
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)

Copied to clipboard

Challenge: VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST.
Approach: They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages.
Outcome: The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations