Papers by Maxwell Forbes
Thinking Like a Skeptic: Defeasible Inference in Natural Language (2020.findings-emnlp)
Copied to clipboard
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, Yejin Choi
| Challenge: | Defeasible inference is a mode of reasoning in which an inference may be weakened or overturned in light of new evidence. |
| Approach: | They propose a dataset for defeasible inference in natural language that includes extensions to existing inference datasets. |
| Outcome: | Defeasible NLI extends existing datasets for defeaasibility inference in natural language . generative models can weaken or strengthen inferences up to 68% of the time, it shows . |
Neural Naturalist: Generating Fine-Grained Image Comparisons (D19-1)
Copied to clipboard
| Challenge: | a dataset of 41k sentences describes fine-grained differences between photographs of birds . human observers are adept at making fine-grain comparisons, but sometimes require aid in distinguishing visually similar classes. |
| Approach: | They propose a model that generates comparative language from a dataset of 41k sentences describing fine-grained differences between photographs of birds. |
| Outcome: | The proposed model can explain differences in visual embedding space using natural language . it evaluates the results with humans who must use the descriptions to distinguish real images . |
Social Chemistry 101: Learning to Reason about Social and Moral Norms (2020.emnlp-main)
Copied to clipboard
| Challenge: | SOCIAL CHEMISTRY is a conceptual formalism to study people’s everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language. |
| Approach: | They propose a new conceptual formalism to study people's everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language. |
| Outcome: | The proposed model can be used to model people's everyday social norms and moral judgments over a rich spectrum of real life situations. |
CLIPScore: A Reference-free Evaluation Metric for Image Captioning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Image captioning relies on reference-based automatic evaluations, but references are expensive to collect and comparing against multiple human-authored captions is insufficient. |
| Approach: | They propose a reference-free metric that can be used for automatic caption evaluation without references. |
| Outcome: | The proposed model outperforms existing metrics on image-text compatibility and a reference-augmented version achieves even higher correlation with human judgements. |
Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences (2021.emnlp-main)
Copied to clipboard
| Challenge: | aaron carroll: in social settings, human behavior is governed by unspoken rules of conduct rooted in societal norms . carroll and colleagues examine whether language generation models can serve as behavioral priors if they are not . they say we examine whether they can generate descriptions of actions that accomplish predefined goals . |
| Approach: | They propose to combine multiple expert models to improve quality of generated actions, consequences, and norms. |
| Outcome: | The proposed models significantly improve the quality of generated actions, consequences, and norms compared to baselines. |
Learning to Write with Cooperative Discriminators (P18-1)
Copied to clipboard
| Challenge: | Despite their local fluency, long-form text generated from RNNs is often generic, repetitive, and even self-contradictory. |
| Approach: | They propose a unified learning framework that can guide a base RNN generator towards more globally coherent generations by combining discriminators with a composite decoding objective. |
| Outcome: | The proposed framework can guide a base RNN generator towards more globally coherent generations by combining discriminators with the base RRN generator through a composite decoding objective. |
Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text (2022.acl-long)
Copied to clipboard
| Challenge: | a recent study has reported that crowdsourcing cannot distinguish between machine-authored and human-authored text. |
| Approach: | They propose a framework called Scarecrow for scrutinizing machine text via crowd annotation . they use crowd annotation to identify redundancy, commonsense errors, and incoherence . |
| Outcome: | The proposed method quantifies gaps between human-authored and machine-generated text . it can detect redundancy, commonsense errors, and incoherence . |
Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation (2021.acl-long)
Copied to clipboard
| Challenge: | Edited media frames are structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation. |
| Approach: | They propose a new formalism to understand visual media manipulation as structured annotations with respect to intents, emotional reactions, attacks on individuals, and the implications of disinformation. |
| Outcome: | The proposed model obtains promising results on a dataset with 56k question-answer pairs written in rich natural language. |