Papers with English-language

28 papers
Thesis Proposal: Detecting Agency Attribution (2024.eacl-srw)

Copied to clipboard

Challenge: 'agency' is the freedom and capacity of an entity to act, and the corresponding Natural Language Processing (NLP) task involves automatically detecting attributions of agency to entities in text.
Approach: They propose a schema to annotate a dataset for agency attribution and formulate additional research questions by applying NLP models.
Outcome: The proposed framework draws on semantic frame analysis, role labelling and related techniques.
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) (2020.emnlp-main)

Copied to clipboard

Challenge: Pretraining on self-supervised linguistic tasks is effective for learning features helpful for language understanding, but it requires more data to learn to prefer linguistic generalizations over surface ones.
Approach: They propose a set of 20 ambiguous binary classification tasks to test whether a pretrained model prefers linguistic or surface generalizations.
Outcome: The proposed model can learn to represent linguistic features with little pretraining data, but requires far more data to learn to prefer linguistic generalizations over surface ones.
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer (2021.naacl-main)

Copied to clipboard

Challenge: Current natural language processing pipelines often use transfer learning, where a model is pre-trained on a data-rich task before being fine-tuned on . this significantly limits their use given that roughly 80% of the world population does not speak English.
Approach: They introduce a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages.
Outcome: The proposed model achieves state-of-the-art on multilingual benchmarks and a simple technique to prevent accidental translation in the zero-shot setting.
CitationIE: Leveraging the Citation Graph for Scientific Information Extraction (2021.acl-long)

Copied to clipboard

Challenge: Existing work on scientific information extraction (SciIE) considers extraction solely based on the content of an individual paper, without considering the paper’s place in the broader literature.
Approach: They propose to automate the extraction of key information from scientific documents by leveraging a complementary source: the citation graph of referential links between citing and cited papers.
Outcome: The proposed model improves on a set of English-language scientific documents.
A Matter of Perspective: Building a Multi-Perspective Annotated Dataset for the Study of Literary Quality (2024.lrec-main)

Copied to clipboard

Challenge: a dataset collecting quality judgments on 9,000 English-language novels is presented . authors include experts opinions and crowd-sourced annotations .
Approach: They propose a dataset collecting quality judgments on 9,000 English-language novels by 3,150 predominantly Anglophone authors.
Outcome: The proposed dataset examines the perceived quality of 9,000 English-language novels by 3,150 predominantly anglophone authors.
PASTA: A Dataset for Modeling PArticipant STAtes in Narratives (2023.tacl-1)

Copied to clipboard

Challenge: Existing models that understand narratives should infer these implicit states and their causal relationships with the narrative's explicit events.
Approach: They propose a dataset that contains inferable participant states, a counterfactual perturbation to each state and the changes to the story that would be necessary if the counterfact was true.
Outcome: The proposed model can reason about the impact of changes to the story that would be necessary if the counterfactual were true.
Decomposed scoring of CCG dependencies (2023.acl-short)

Copied to clipboard

Challenge: a standard evaluation of supertagging errors can result in disproportionate penalization of supertaggers . comparative categorial grammar (ccg) supertaggers can adjust for their own errors to keep sentences parsable .
Approach: They propose a decomposed scoring method based on subcategorial labels to address this problem.
Outcome: The proposed method penalizes supertagging errors and obfuscates erroneous dependencies . the proposed method is based on subcategorial labels .
Identifying Narrative Content in Podcast Transcripts (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods to study narrativity in novels, social media and patient records are limited.
Approach: They propose to process podcast transcripts and extract narrative content from podcasts . they use annotations to enable future research into narrativity within a large corpus of podcast episodes.
Outcome: The proposed methods compare to existing methods and can enable future research into narrativity within a large corpus of approximately 100,000 podcast episodes.
What do tokens know about their characters and how do they know it? (2022.naacl-main)

Copied to clipboard

Challenge: Pre-trained language models that use subword tokenization schemes can succeed at a variety of language tasks that require character-level information.
Approach: They propose to use word tokenization schemes to probe what word pieces encode . they show that larger models can encode character-level information .
Outcome: The proposed models can encode character-level information and perform better on non-Latin alphabets.
Annotating and Analyzing Biased Sentences in News Articles using Crowdsourcing (2020.lrec-1)

Copied to clipboard

Challenge: a lack of publicly available news bias datasets has hindered efforts to detect subtle biases in news articles.
Approach: They propose a news bias dataset which contains sentences with bias labels . they propose to use the dataset to develop and evaluate methods for detecting news bias .
Outcome: The proposed dataset can be used for analyzing news bias and for developing and evaluating methods for news bias detection.
DynaSent: A Dynamic Benchmark for Sentiment Analysis (2021.acl-long)

Copied to clipboard

Challenge: Sentiment analysis is an early success story for NLP, in both a technical and an industrial sense.
Approach: They propose to combine naturally occurring sentences with sentences created using the open-source Dynabench Platform, which facilities human-and-model-in-the-loop dataset creation.
Outcome: The proposed model is more coherent than comparable models and motivates training models from scratch over successive fine-tuning.
A Zero-Shot Monolingual Dual Stage Information Retrieval System for Spanish Biomedical Systematic Literature Reviews (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies have shown that most SRs are skewed towards English databases, excluding databases in Languages other than English (LoE).
Approach: They propose a zero-shot dual information retrieval baseline system that integrates traditional retrieval methods with pre-trained language models and cross-attention re-rankers for enhanced accuracy in Spanish biomedical literature retrieval.
Outcome: The proposed system improves on three real-life case studies in Spanish biomedical literature retrieval using the LILACS database, which is known for its coverage of Latin American and Caribbean biomedically literature.
An annotated dataset of literary entities (N19-1)

Copied to clipboard

Challenge: Existing datasets built on news focus on non-named entities, but not literary texts.
Approach: They propose to annotate 210,532 tokens from 100 different English-language literary texts for ACE entity categories (person, location, geo-political entity, facility, organization, and vehicle).
Outcome: The proposed dataset includes 210,532 tokens drawn from 100 different English-language literary texts.
Multi-task Learning of Negation and Speculation for Targeted Sentiment Classification (2021.naacl-main)

Copied to clipboard

Challenge: Currently, most work on targeted sentiment analysis is focused on improving the overall results.
Approach: They propose a multi-task learning method to incorporate information from syntactic and semantic auxiliary tasks to create English-language models that are more robust to linguistic phenomena.
Outcome: The proposed method improves on negation and speculation datasets but there is room for improvement.
Few-Shot Upsampling for Protest Size Detection (2021.findings-acl)

Copied to clipboard

Challenge: a common task in social science is "upsampling" coarse document labels to finegrained labels . a new task is proposed for "up-samping" coarse labels to more detailed information .
Approach: They propose a task for "upsampling" coarse document labels to finegrained labels or spans . they use a question answering format to provide fine-grained label information .
Outcome: The proposed method outperforms a pre-trained model on a small set of examples but is weaker on fewer examples.
FactAppeal: Identifying Epistemic Factual Appeals in News Media (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods focus on the content of factual statements and ignore the epistemic structures that confer credibility and persuasive force to these claims.
Approach: They propose a task of Epistemic Appeal Identification to identify whether and how factual statements have been anchored by external sources or evidence.
Outcome: The proposed task identifies whether and how factual statements have been anchored by external sources or evidence.
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students’ Hand-Drawn Math Images (2025.naacl-long)

Copied to clipboard

Challenge: DrawEduMath examines the ability of vision language models to handle real-world math problems, such as those encountered in classrooms and tutoring sessions.
Approach: They present DrawEduMath, an English-language dataset of 2,030 images of students’ handwritten responses to math problems.
Outcome: The proposed model can be used to evaluate teachers' QA pairs and 44,362 synthetic QAs derived from teachers' descriptions.
Transfer-Free Data-Efficient Multilingual Slot Labeling (2023.emnlp-main)

Copied to clipboard

Challenge: Slot labeling (SL) is a key component of task-oriented dialogue systems . extending the system to any new language-domain-task configuration requires expensive data annotation .
Approach: They propose a two-stage slot labeling approach which transforms sentence encoders into effective slot labels.
Outcome: The proposed approach is especially effective for the most challenging transfer-free few-shot setups.
Where Do People Tell Stories Online? Story Detection Across Online Communities (2024.acl-long)

Copied to clipboard

Challenge: Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text.
Approach: They propose a toolkit to detect stories in online communities using an annotated reddit dataset and a codebook adapted to social media context.
Outcome: The proposed toolkit includes an annotation-rich dataset of 502 Reddit posts and comments . it also includes a codebook adapted to the social media context and models to predict storytelling at document and span levels.
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
Cross Domain Classification of Education Talk Turns (2025.coling-main)

Copied to clipboard

Challenge: Prior research has focused on the annotation of conversational talk-turns within the classroom, offering a statistical analysis of the various types of discourse prevalent in these environments.
Approach: They examine the generalizability and transferability of text classifiers trained to predict classroom discourse across educational domains by accompanying each talk turn with dialog-level context.
Outcome: The proposed models exhibit high generalizability when training and test datasets originate from the same or similar domains.
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language .
Approach: They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance .
Outcome: The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement.
Using Structured Content Plans for Fine-grained Syntactic Control in Pretrained Language Model Generation (2022.coling-1)

Copied to clipboard

Challenge: Large pretrained language models can generate powerful text but cannot be controlled at a sub-sentential level.
Approach: They propose to make such fine-grained control possible in pretrained LMs by generating text directly from a semantic representation, Abstract Meaning Representation (BART), which is augmented at the node level with syntactic control tags.
Outcome: The proposed method can generate text from a semantic representation, which is augmented at the node level with syntactic control tags.
LCGbank: A Corpus of Syntactic Analyses Based on Proof Nets (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have focused on statistical syntactic parsing with proof nets . however, there has been a paucity of corpora in formalisms for which proof net is applicable .
Approach: They propose a corpus of syntactic analyses based on Lambek categorial grammar . they leverage the relationship between LCG and CCG to address this problem .
Outcome: The proposed method exploits the relationship between LCG and CCG to build an English-language corpus of syntactic analyses based on proof nets . the results suggest that the proposed method is weakly context-free equivalent and NP-complete .
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning (2024.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed a substantial increase in the demand for legal services, especially for individuals with modest means.
Approach: They propose a diagnostic legal large language model which uses adaptive lawyer-like diagnostic questions to collect additional case information and then provides high-quality feedback.
Outcome: The proposed model surpasses classical LLMs by providing outstanding performance and a remarkable user experience in the legal domain.
LLMs Reproduce Stereotypes of Sexual and Gender Minorities (2025.findings-emnlp)

Copied to clipboard

Challenge: a large body of research has found substantial gender bias in NLP systems . authors show that LLMs generate stereotyped representations of sexual and gender minorities in this setting .
Approach: They propose to use a stereotype content model to study gender bias in large language models . they show that LLMs generate stereotyped representations of sexual and gender minorities .
Outcome: The proposed model generates negative stereotypes of sexual and gender minorities in English-language surveys .
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness.
Approach: They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.
Outcome: The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness.
so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Despite popularity of poetry as both an art form and a generation task for large language models, whitespace has not received sufficient attention from the NLP community.
Approach: They examine how 4k poets have used whitespace in their works . they compare it to 51k LLM-generated poems and 12k unpublished poems posted online .
Outcome: The proposed dataset compares 4k poetry poems with 51k LLM-generated poems and 12k unpublished poems posted in an online community.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations