Papers by Kate Sanders

7 papers
Grounding Partially-Defined Events in Multimodal Data (2024.findings-emnlp)

Copied to clipboard

Challenge: Evidence suggests prelinguistic infants are capable of recognizing discrete events in real-world stimuli.
Approach: They propose a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task.
Outcome: The proposed approach can extract events from 14.5 hours of annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities.
Core: Robust Factual Precision with Informative Sub-Claim Identification (2025.findings-acl)

Copied to clipboard

Challenge: Using the Decompose-Then-Verify framework, such as FActScore, can be manipulated by adding obvious or repetitive subclaims to artificially inflate scores.
Approach: They propose a decomposition-based tool called Core to filter subclaims based on their uniqueness and informativeness.
Outcome: The proposed evaluation framework supports easy and modular use of Core and various decomposition strategies.
TurkingBench: A Challenge Benchmark for Web Agents (2025.naacl-long)

Copied to clipboard

Challenge: TurkingBench is a benchmark consisting of tasks presented as web pages with textual instructions and multi-modal contexts.
Approach: They propose to use HTML pages to perform various annotation tasks on crowdsourcing platforms.
Outcome: The proposed model outperforms other models on the TurkingBench benchmark.
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic (2024.emnlp-main)

Copied to clipboard

Challenge: Recent language models allow structured reasoning with text, but lack of a clear protocol for discerning entailment causes noisy datasets and limited performance gains.
Approach: They propose a consistent approach to annotating decompositional entailment and evaluate its impact on LLM-based textual inference.
Outcome: The proposed approach has higher internal consistency than prior decompositional entailment datasets and significantly improves proof quality and accuracy.
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers? (2025.findings-emnlp)

Copied to clipboard

Challenge: CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview.
Approach: They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses.
Outcome: The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses.
WikiVideo: Article Generation from Multiple Videos (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for retrieval-augmented generation focus on text rather than video.
Approach: They propose a benchmark to generate Wikipedia-style articles from multiple videos . they propose 'collaborative article generation' that leverages an r1-style reasoning model and a VideoLLM to draw higher-level inferences about the target event than is possible with VideoLLms alone.
Outcome: The proposed method outperforms existing methods in oracle retrieval and RAG settings while suggesting promising avenues for future work.
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: TV-TREES is the first multimodal entailment tree generator for video understanding . it searches for trees of enanglement relationships between text-video evidence and higher-level conclusions that prove question-answer pairs.
Approach: They propose a multimodal entailment tree generator that promotes interpretable joint-modality reasoning by searching for trees of enanglement relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs.
Outcome: The proposed approach performs on the TVQA benchmark and shows that it is state-of-the-art on full clips.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations