Papers by Michael Roth

15 papers
Combining Discourse Markers and Cross-lingual Embeddings for Synonym–Antonym Classification (N19-1)

Copied to clipboard

Challenge: Recent work shows that distributional semantic approaches have difficulty distinguishing between synonyms and antonyms.
Approach: They propose to use monolingual distributional information available in a target language to transfer supervision to other languages using cross-lingual word embeddings.
Outcome: The proposed method improves the transfer of monolingual distributional information to other languages using co-occurrences with discourse markers indicative of antonymy.
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)

Copied to clipboard

Challenge: Various approaches for script knowledge extraction and processing have been proposed in recent years.
Approach: They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge.
Outcome: The proposed dataset provides test cases for the broader natural language understanding community.
Clarifying Implicit and Underspecified Phrases in Instructional Text (2022.lrec-1)

Copied to clipboard

Challenge: Natural language consists of implicit and underspecified phrases, which can cause misunderstandings.
Approach: They propose to use wikiHow to extract human clarifications that resolve an implicit or underspecified phrase.
Outcome: The proposed model can be used to generate alternate clarifications, which may or may not be compatible with the human clarification.
Toward Implicit Reference in Dialog: A Survey of Methods and Data (2022.aacl-main)

Copied to clipboard

Challenge: In natural language, speakers often leave out information that is understood by the other party through the shared context.
Approach: They propose to use omitted entities as implicit references in dialogs to improve language processing.
Outcome: The proposed method is based on a set of experiments which show that the proposed method has a high level of accuracy and is a success.
“Feels Feminine to Me”: Understanding Perceived Gendered Style through Human Annotations (2025.emnlp-main)

Copied to clipboard

Challenge: Using gender identity-based framing, language–gender associations are often grounded in the author’s gender identity, inferred from their language use.
Approach: They propose to operationalize the language–gender association as a perceived gender expression of language, focusing on how expression is externally interpreted by humans, independent of the author’s gender identity.
Outcome: The first dataset of itskind identifies 5,100 human annotations of perceived gendered style—human-written texts rated on a five-point scale from very feminine to very masculine.
Commonsense Inference in Natural Language Processing (COIN) - Shared Task Report (D19-60)

Copied to clipboard

Challenge: The workshop on Commonsense Inference in NLP (COIN) evaluated text understanding systems' ability to draw inferences about facts that are not mentioned in the text, but that are assumed to be common ground.
Approach: They propose to use commonsense knowledge to evaluate systems' ability to answer questions/queries about a text.
Outcome: The proposed tasks evaluated systems in two contexts: Commonsense Inference and Commonsensible Inference.
wikiHowToImprove: A Resource and Analyses on Edits in Instructional Texts (2020.lrec-1)

Copied to clipboard

Challenge: wikiHow articles are subject to revision edits, but do they provide clarifications? a new study compares changes made across multiple versions of the same set of instructions .
Approach: They use wikiHow to analyze revision histories for 2.7 million sentences from wikihow . they use human annotation to categorize subset of edits and provide models .
Outcome: The proposed model can distinguish between “older” and “newer” revisions of a sentence.
Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences (N18-1)

Copied to clipboard

Challenge: Using a dataset of 6,500+ questions, we found that human solvers achieved an F1-score of 88.1%.
Approach: They propose a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
Outcome: The proposed reading comprehension challenge is based on a reading comprehension dataset with 6,500+ questions and 1000+ paragraphs across 7 domains.
Towards Modeling Revision Requirements in wikiHow Instructions (2020.emnlp-main)

Copied to clipboard

Challenge: wikiHow is a collaboratively edited platform of how-to guides . authors extend existing textual edits with 4 million sentences that remain unedited .
Approach: They extend existing textual edits with a set of 4 million sentences that remain unedited over time.
Outcome: The proposed model can predict the need for edits in wikiHow guides . the authors extend an existing resource of textual edits with a complementary set of 4 million sentences that remain unedited over time .
Clarifying Underspecified Discourse Relations in Instructional Texts (2025.findings-acl)

Copied to clipboard

Challenge: Discourse relations can be optionally realized through explicit connectives such as “but” and “while”.
Approach: They build a corpus of 4,274 text revisions in which a connective was explicitly inserted . they collect plausibility annotations on other connectives to check whether they represent suitable alternatives .
Outcome: The proposed model predicts plausibility of individual connectives with up to 66% accuracy, but is not reliable when multiple relations are plausible.
Generalization in Instruction Following Systems (2021.naacl-main)

Copied to clipboard

Challenge: Understanding and executing natural language instructions in a grounded domain is one of the hallmarks of artificial intelligence.
Approach: They propose a learning strategy that involves data augmentation to improve the model's performance.
Outcome: The proposed learning strategy outperforms state-of-the-art models in the blocks world domain while satisfying our expectations much better.
How-to Guides for Specific Audiences: A Corpus and Initial Findings (2023.acl-srw)

Copied to clipboard

Challenge: wikiHow guides for specific target groups reflect disparate social norms and subtle stereotypes, a new study shows . wikihow guides are subject to subtle biases, and we aim to raise awareness of these inequalities in future work.
Approach: They investigate the extent to which how-to guides from wikiHow differ in practice depending on intended audience.
Outcome: The findings show that how-to guides from wikiHow differ in practice depending on the intended audience.
A Computational Analysis of Vagueness in Revisions of Instructional Texts (2021.eacl-srw)

Copied to clipboard

Challenge: We analyze edits that involve cases of vagueness in instructional texts . we extract and analyze version pairs of an instruction before and after a revision .
Approach: They propose to extract and analyze edits that involve cases of vagueness in instructions . they adopt a pairwise ranking task to show improvements over existing baselines .
Outcome: The proposed model can distinguish between two versions of an instruction in a noisy dataset.
What Can We Learn from Noun Substitutions in Revision Histories? (2020.coling-main)

Copied to clipboard

Challenge: Recent work shows that resulting improvements can be modelled computationally, assuming that each revision contributes to the improvement.
Approach: They propose to model improvements in sentences using wikiHow revision histories by assuming that each revision contributes to the improvement.
Outcome: The proposed model fails in cases where humans can resort to factual knowledge or intuitions about the required level of specificity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations