Papers by Jiefu Ou
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Literature review tables are essential for summarizing and comparing collections of scientific papers. |
| Approach: | They propose to generate a database of literature review tables from a pool of papers and to model retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators. |
| Outcome: | The proposed method improves over strong baselines while the absolute scores remain modest, underscoring the task’s difficulty. |
Pragmatic Inference with a CLIP Listener for Contrastive Captioning (2023.findings-acl)
Copied to clipboard
| Challenge: | a new method for contrastive captioning generates discriminative captions that distinguish target images from very similar alternative distractor images. |
| Approach: | They propose a pragmatic inference procedure that formulates captioning as a reference game between a speaker and a listener. |
| Outcome: | The proposed method outperforms previous methods for discriminative captioning by 11% to 15% accuracy in human evaluations. |
InFillmore: Frame-Guided Language Generation with Bidirectional Context (2021.starsem-1)
Copied to clipboard
| Challenge: | Existing methods for automatic story plan generation use coarse-to-fine representations of semantic content. |
| Approach: | They propose a structured extension to bidirectional-context conditional language generation, or "infilling" they propose evocative frame annotations and a method for frame-guided generation that leverages frame semantic lexical units. |
| Outcome: | The proposed method allows for explicit manipulation of intended infill semantics with minimal loss of distinguishability from human-generated text. |
Exploring Discourse Structures for Argument Impact Classification (2021.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that discourse structures influence the persuasiveness of arguments. |
| Approach: | They propose to fuse sentence-level structural discourse information with contextualized features derived from large-scale language models to investigate how discourse relations influence argument impact. |
| Outcome: | The proposed model improves its backbone RoBERTa around 1.67%, compared with other models, but side effects are brought by other models. |
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers? (2025.findings-emnlp)
Copied to clipboard
Jiefu Ou, William Gantt Walden, Kate Sanders, Zhengping Jiang, Kaiser Sun, Jeffrey Cheng, William Jurayj, Miriam Wanner, Shaobo Liang, Candice Morgan, Seunghoon Han, Weiqi Wang, Chandler May, Hannah Recknor, Daniel Khashabi, Benjamin Van Durme
| Challenge: | CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview. |
| Approach: | They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
| Outcome: | The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses. |