Papers by Tingyu Song

9 papers
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models’ Detection of Human risky health behavior Content in Jirai Community (2026.eacl-long)

Copied to clipboard

Challenge: a cross-lingual dataset captures a transnational cultural phenomenon . risky health behaviors (RHB) are often linked to complex mental health conditions .
Approach: They present the first cross-lingual dataset that captures a transnational cultural phenomenon . their dataset of more than 15,000 annotated social media posts forms the core of JiraiBench .
Outcome: The study shows that cultural context can be more influential than linguistic similarity . the study also shows that the Japanese prompts better handle Chinese content .
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) excel in turn-by-turn human-AI collaboration but struggle with simultaneous tasks requiring real-time interaction.
Approach: They propose a language agent framework that integrates *System 1* and *System 2* for efficient real-time simultaneous human-AI collaboration.
Outcome: The proposed framework improves on existing LLM-based agents and human collaborators by integrating Theory of Mind and asynchronous reflection to infer human intentions and perform reasoning-based autonomous decisions.
A Survey of Reasoning-Intensive Retrieval: Progress and Challenges (2026.acl-long)

Copied to clipboard

Challenge: Reasoning-Intensive Retrieval (RIR) targets retrieval settings where relevance is mediated by latent inferential links between a query and supporting evidence, rather than semantic similarity.
Approach: They propose a taxonomy that categorizes methods based on where and how reasoning is integrated into the retrieval pipeline.
Outcome: The proposed method framework provides a detailed analysis of the current landscape and its trade-offs and practical applications.
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing (2026.acl-long)

Copied to clipboard

Challenge: Composed Image Retrieval (CIR) is a complex task in multimodal understanding . current CIR benchmarks lack a robust evaluation pipeline and limited query categories .
Approach: They construct a fine-grained CIR benchmark that allows for precise control over modification types and content.
Outcome: The proposed benchmark covers 5,000 high-quality queries structured across five main categories and fifteen subcategories.
Efficiency-Effectiveness Reranking FLOPs for LLM-based Rerankers (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing studies evaluate the efficiency of LLM-based rerankers using proxy metrics such as latency and the number of forward passes.
Approach: They propose to use a large language model to evaluate the efficiency of LLM-based rerankers . they propose to measure ranking quality and query processing efficiency using an interpretable FLOPs estimator .
Outcome: The proposed metrics evaluate LLM-based rerankers with different architectures without running any experiments.
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval (2025.naacl-long)

Copied to clipboard

Challenge: Current information retrieval systems struggle to handle complex instructions, despite its critical importance . current models struggle to follow complex instructions in real-world applications, resulting in user-specific tasks.
Approach: They propose a benchmark to evaluate instruction-following information retrieval in expert domains.
Outcome: The proposed method improves on existing models and provides valuable insights to guide future advancements in retrieval.
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation benchmarks for retrievers are narrow and evaluate them in isolation . existing evaluation benchmarking frameworks focus on evaluating retrievers in isolation, obscuring their value in real-world applications.
Approach: They propose an evaluation framework that evaluates retrievers in agentic search systems . they provide expert-annotated reasoning aspects, positive documents, a reference response and evaluation rubrics .
Outcome: The proposed framework assesses retrievers in agentic search systems.
LimRank: Less is More for Reasoning-Intensive Information Reranking (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to rerank information require large-scale fine-tuning, which is computationally expensive.
Approach: They propose an open-source pipeline for generating diverse, challenging, and realistic reranking examples.
Outcome: The proposed model performs competitively on two benchmarks, while being trained on less than 5% of the data typically used in prior work.
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are used for video quality assessment, image captioning and video analysis.
Approach: They propose a benchmark to evaluate MLLMs on AIGC videos using coherence validation, error awareness, error type detection and reasoning evaluation tasks.
Outcome: The proposed benchmark evaluates 13 frontier MLLMs on AIGC videos.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations