Papers by Wentao Zhang

55 papers
K-order Ranking Preference Optimization for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing list-wise methods focus on optimizing list ranking consistency for LLMs to improve ranking abilities.
Approach: They propose to extend the Plackett-Luce model to accommodate top-K ranking by extending the DPO’s Plact-Lucer model to dynamically determine appropriate K for different samples.
Outcome: The proposed model can be extended to accommodate top-K ranking and improve training efficiency.
TC–RAG: Turing–Complete RAG’s Case study on Medical LLM Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to RAG neglect system state variables, resulting in poor performance and erroneous knowledge accumulation.
Approach: They propose a framework that incorporates a Turing Complete System to manage state variables and manage retrieval halting.
Outcome: The proposed framework improves on seven real-world healthcare datasets and shows that it is more accurate than existing methods.
Patton: Language Model Pretraining on Text-Rich Networks (2023.acl-long)

Copied to clipboard

Challenge: Existing models for text-rich networks do not take inter-document structure into account.
Approach: They propose a pretraining framework for a text-rich network using a masked language model and a masking node prediction framework.
Outcome: The proposed model outperforms baselines on four tasks in academic and e-commerce domains.
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs’ Responsiveness to Human Feedback (2025.emnlp-main)

Copied to clipboard

Challenge: Existing research focuses on benchmarking LLMs in single-turn dialogues, neglecting the nuanced nature of human feedback within real-world usage scenarios.
Approach: They propose a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese.
Outcome: The proposed benchmarks show that human feedback can significantly impact LLMs’ responsiveness in real-world usage scenarios.
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation (2026.acl-long)

Copied to clipboard

Challenge: Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components with iterative reuse of existing codes.
Approach: They propose a benchmark to evaluate LLMs' ability to perform codeflow by reusing existing functions over multiple turns.
Outcome: The proposed benchmarks show that LLMs perform significantly worse in multi-turn codeflow scenarios and that their performance inversely correlates with dependency complexity.
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch (2026.acl-long)

Copied to clipboard

Challenge: Existing open-source vision language models lack high-quality training data for chart reasoning . current models are simplistic and repetitive, while associated QA pairs are prone to hallucinations .
Approach: They propose a framework to synthesize complex charts and reliable reasoning data from scratch.
Outcome: Experimental results show that ChartVerse-8B surpasses existing models in QA and difficulty . lack of high-quality training data hampers development of open-source models .
Dynamic Parallel Tree Search for Efficient LLM Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Recent methods focus on search accuracy while overlooking computational efficiency.
Approach: They propose a parallelism framework that dynamically optimizes reasoning path in inference.
Outcome: The proposed framework improves efficiency by 2-4 on average while maintaining or even surpassing existing reasoning algorithms in accuracy.
Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration (2026.findings-acl)

Copied to clipboard

Challenge: Structured Query Language (SQL) is the cornerstone for data-driven decision-making.
Approach: They propose a benchmark to rigorously evaluate Large Language Models within a dynamic interaction framework.
Outcome: The proposed benchmark aims to rigorously evaluate LLMs within a dynamic interaction framework.
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing data synthesis methods suffer from limited diversity and lack precise control over problem difficulty, making them insufficient for efficient training paradigms such as curriculum learning.
Approach: They propose a data synthesis paradigm that generates high-quality, difficulty-controllable mathematical reasoning problems through hybrid and decomposed strategies.
Outcome: The proposed paradigm outperforms existing methods and improves mathematical reasoning abilities.
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax (2026.findings-acl)

Copied to clipboard

Challenge: Extending large language models to low-resource languages often incurs an "alignment tax" token-level fine-tuning enforces token-level surface imitation on narrow and biased data distributions.
Approach: They propose a semantic-space alignment paradigm powered by group-level semantic rewards instead of likelihood maximization.
Outcome: The proposed model acquires low-resource capa- bilities while mitigating alignment tax on Tibetan–Chinese machine translation and Ti- betan headline generation.
Inductive Relation Inference of Knowledge Graph Enhanced by Ontology Information (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to inference knowledge graphs lack ontology information, which is often too sparse.
Approach: They propose a knowledge graph inductive inference method that fuses ontology information to learn the semantic information of entities.
Outcome: The proposed method outperforms large language models like ChatGPT on two benchmark datasets and improves the MRR metrics by 15.4% and 44.1%, respectively.
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent work has shown promise by incorporating pixel-level visual information into the reasoning process, enabling VLMs to access high-resolution visual details during their thought process.
Approach: They propose a framework that dynamically determines necessary pixel-level operations based on the input query.
Outcome: The proposed model achieves 73.4% accuracy on HR-Bench 4K while maintaining a tool usage ratio of only 20.1%, improving accuracy and reducing tool usage by 66.5% compared to the previous methods.
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training (2024.acl-long)

Copied to clipboard

Challenge: Experimental results demonstrate that ProtLLM achieves superior performance against protein-specialized baselines on protein-centric tasks and induces zero-shot and in-context learning capabilities on protein language tasks.
Approach: They propose a cross-modal large language model (LLM) that can handle protein-centric and protein-language tasks by using a dynamic protein mounting mechanism.
Outcome: The proposed model can predict proteins from a vast pool of candidates and can also predict natural language and biological papers.
Towards General Agentic Intelligence via Environment Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Diverse real-world APIs require precise, robust function-calling intelligence, which needs agents to develop these capabilities through interaction in varied environments.
Approach: They propose a framework that scales up environments to enable agentic intelligence . they use a two-phase agent fine-tuning strategy to first endow agents with basic agentic capabilities, then specializing them for domain-specific contexts.
Outcome: Experiments on -bench, -Bench, and ACEBench show that the model significantly enhances the models’ function-calling capability.
Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token’s Nature (2026.acl-long)

Copied to clipboard

Challenge: Existing methods that use entropy as a discrete filter or post-hoc regulator are limited in their ability to optimize for reasoning tasks.
Approach: They propose a token-aware algorithm that continuously adapts optimization dynamics based on token-level entropy throughout the entire training process.
Outcome: Extensive experiments on mathematical reasoning, code, and logic tasks across multiple models demonstrate HAPO’s consistent superiority over DAPO.
Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing jamming attacks on RAG systems typically induce explicit refusals or denial-of-service behaviors.
Approach: They propose a black-box attack framework that exploits safety-aligned behaviors of large language models to trigger soft failures.
Outcome: The proposed framework exploits safety-aligned behaviors of large language models to induce soft failures.
Improving Multi-label Malevolence Detection in Dialogues through Multi-faceted Label Correlation Enhancement (2022.acl-long)

Copied to clipboard

Challenge: Current methods for detecting dialogue malevolence neglect label correlation.
Approach: They propose to crowdsource a multi-label dataset for detecting malevolent dialogue responses and a model with label correlation enhanced CRF to measure the correlation between malevolence and negative emotions.
Outcome: The proposed model outperforms the best performing baseline method on precision, recall, F1, and Jaccard score by 16.1%, 11.9%, 12.0%, and 6.1% on malevolence.
Taming LLMs with Gradient Grouping (2025.acl-long)

Copied to clipboard

Challenge: a new study presents scaling with gradient grouping (SGG) the adaptive learning rate scaling approach is based on per-parameter statistics, which incurs memory overhead.
Approach: They propose an optimizer wrapper that improves adaptive learning rate estimation by dynamic grouping and group-specific scaling.
Outcome: The proposed algorithm improves learning rate estimation on diverse models with different model sizes and batch sizes.
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Efficient data selection is crucial to accelerate the pretraining of language models . limited research has addressed the inherent conflicts between data selection methods .
Approach: They propose a multi-actor collaborative data selection mechanism that prioritizes data based on its specific criterion and updates prioritization rules using the current state of the model.
Outcome: The proposed model accelerates convergence in LM pretraining and achieves an average relative performance gain of 10.5% across multiple language model benchmarks.
Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuning (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study highlights unpaired feedback as a key challenge for long-term LLM-based recommenders . unpaired user feedback is crucial for improving LLMs in dynamic user environments, authors say .
Approach: They propose a framework that incorporates unpaired feedback into LLMs to improve long-term recommendation performance.
Outcome: The proposed framework improves long-term recommendation performance by incorporating unpaired feedback without requiring paired supervision.
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work focuses on domain-specific enhancements during fine-tuning, the challenge of which lies in catastrophic forgetting of knowledge across other domains.
Approach: They propose a data composition framework that allows LLMs to enhance their multi-domain capabilities during supervised fine-tuning.
Outcome: The proposed framework improves multi-domain fostering performance by 29.77% compared to uniform weights.
Interactive Training: Feedback-Driven Neural Network Optimization (2025.emnlp-demos)

Copied to clipboard

Challenge: In traditional neural network training, static optimization methods lack flexibility and responsiveness . authors demonstrate that Interactive Training provides superior training stability and reduced sensitivity to initial hyperparameters .
Approach: They propose an open-source framework that enables real-time feedback-driven optimization of neural networks by human experts or automated AI agents.
Outcome: The proposed framework achieves superior training stability, reduced sensitivity to initial hyperparameters, and improved adaptability to evolving user needs.
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for data augmentation neglect fine-grained knowledge, such as entities and quantities, leading to insufficient diversity and high data noise.
Approach: They propose a pipeline-based data augmentation method via LLMs and introduce the Gaussian-decayed gradient-assisted Contrastive Sentence Embedding (GCSE) model to enhance unsupervised sentence embeddings.
Outcome: The proposed method achieves state-of-the-art performance in semantic textual similarity tasks using fewer data samples and smaller LLMs.
Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to sequence labeling are limited due to the scarcity of domain-specific data and semantic distribution biases in domain-based contexts.
Approach: They propose a framework that integrates an LLM-based knowledge enhancement workflow with a span-based Knowledge Fusion for Rich and Efficient Extraction model.
Outcome: The proposed model achieves state-of-the-art performance on multiple domain-specific sequence labeling datasets and is highly efficient.
The Data Frontier for Large Language Models: Selection, Synthesis, and Tools (2026.acl-tutorials)

Copied to clipboard

Challenge: acquiring and curating high-quality training data remains a significant bottleneck . acquiring such high-quality data is a key challenge for researchers and practitioners .
Approach: This tutorial provides a comprehensive and practical guide to the state-of-the-art in data research directions for LLMs.
Outcome: The tutorial covers methods for curating the most valuable information from vast, noisy datasets and the synthetic data revolution.
Can LLMs be Good Graph Judge for Knowledge Graph Construction? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for converting unstructured text into structured Knowledge Graphs (KGs) have limitations such as large amount of noise, inaccurate knowledge, and hallucination .
Approach: They propose a GraphJudge framework to reduce noise in real-world documents . they propose Graphjudge to fine-tune a LLM as a graph judge to enhance quality .
Outcome: The proposed framework eliminates noise in real-world documents and improves the quality of generated KGs.
Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Existing ASR TTA methods struggle with instability under continual and long-term distribution shifts.
Approach: They propose a continuous adaptive model-bank framework that adapts to domain shifts in ASR test-time scenarios.
Outcome: Experiments on diverse, continuously shifting ASR benchmarks show that DMSUTA outperforms existing continual TTA baselines.
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks designed to evaluate the reasoning capabilities of large models are limited in scope and lack flexibility to adapt difficulty according to evolving reasoning capacities of models.
Approach: They propose a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
Outcome: The proposed benchmark incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
Nested Browser-Use Learning for Agentic Information Seeking (2026.acl-long)

Copied to clipboard

Challenge: Existing information-seeking (IS) agents rely on the web for their information acquisition.
Approach: They propose a browser-action framework that decouples interaction control from page exploration through a nested structure.
Outcome: Empirical results show that NestBrowse offers clear benefits in practice.
Data-Centric Perspectives on Agentic Retrieval-Augmented Generation: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel at natural language understanding and generation, yet rely on static pre-training data.
Approach: They propose to augment Large Language Models with external retrieval to ground model outputs . traditional RAG is constrained by a fixed retrieve-then-generate routine . authors aim to guide creation of high-quality datasets for next generation of adaptive LLM agents .
Outcome: The proposed model can decompose tasks, issue exploratory queries, and refine evidence through iterative retrieval.
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts.
Approach: They propose an explicit policy optimization model that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.
Outcome: The proposed model provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences.
Approach: They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge.
Outcome: The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria.
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding. (2026.findings-acl)

Copied to clipboard

Challenge: LongInsightBench is the first benchmark designed to assess models’ ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements.
Approach: They propose a benchmark to assess models’ ability to understand long videos with a focus on human language, viewpoints, actions, and other contextual elements.
Outcome: The proposed model excels in three key areas: a) long-duration, human-centric videos; b) diversifying and challenging task scenarios; c) quality assurance pipeline; and d) reliability.
Knowledge Graph-Driven Memory Editing with Directional Interventions (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are hampered by inaccuracies and outdated information.
Approach: They propose a framework that constructs knowledge graphs using available information to guide the direction of knowledge editing.
Outcome: The proposed framework allows consistent, aligned, and stable information during large-scale editing scenarios.
Training Verifier to Assessing Complex Real-World Tool-Use Trajectories (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for training effective AI agents often resort to synthetic data generation.
Approach: They propose a plug-and-play framework for data quality control in tool-use scenarios . they construct a tool-verify dataset and release a benchmark to assess its performance .
Outcome: The proposed framework surpasses Qwen2.5-72B-Instruct on Tool-V-Bench and the previous APIGen-MT dataset.
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) focus on fragmented constraints or narrow scenarios, but they overlook the comprehensiveness and authenticity of constraints from the user’s perspective.
Approach: They propose a Chinese Comprehensive Constraints Following Benchmark for LLMs that compiles constraints from real-world instructions and constructs a systematic framework for constraint types.
Outcome: The proposed framework integrates multi-dimensional assessment criteria with requirement prioritization, covering various perspectives of constraints, instructions, and requirement fulfillment.
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms.
Approach: They propose a visual-target trajectory collection pipeline that generates trajectories for GUI and embodied tasks using a single formulation.
Outcome: The proposed agent outperforms state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation.
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) integrate visual and textual inputs, yet modality alignment remains one of the most challenging aspects.
Approach: They propose a token-level supervision alignment method that enables more precise visual-text alignment during pretraining.
Outcome: The proposed method improves performance across various model sizes, with smaller models benefiting the most.
Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style Transfer (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for unsupervised text style transfer lack parallel data and difficulties in content preservation.
Approach: They propose a neural approach to unsupervised text style transfer using non-parallel data.
Outcome: The proposed approach can be trained end-to-end on two widely-used public datasets.
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval? (2026.findings-acl)

Copied to clipboard

Challenge: Rapid advances in multimodal large language models have revolutionized cross-modality understanding.
Approach: They propose a method that uses whitening transformations to adjust MLLM representation spaces . they propose ML models that are dominated by textual semantics and visual semantics .
Outcome: The proposed approach improves zero-shot multimodal retrieval performance without fine-tuning efforts.
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to optimize large language models with human preferences suffer from preference conflicts in the data.
Approach: They propose to construct Pareto-optimal responses to resolve preference conflicts by using a self-improving DPO framework that enables LLMs to self-generate and select Paret-optimized responses.
Outcome: The proposed framework achieves superior Pareto Front performance over baselines on two datasets.
UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks do not assess agents’ capabilities across data types . Existing tools only evaluate agents' ability to extract reasonable insights across data formats.
Approach: They propose a multi-source benchmark to evaluate the performance of data analytics agents in handling diverse data sources.
Outcome: The proposed agent performs end-to-end analysis over diverse data sources by automatically discovering cross-source linkages, decomposing goals, and generating robust, self-correcting code to extract actionable insights.
QAEncoder: Towards Aligned Representation Learning in Question Answering Systems (2025.acl-long)

Copied to clipboard

Challenge: Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses, but the inherent gap between user queries and relevant documents hinders precise matching.
Approach: They propose a retrieval-augmented generation (RAG)-based approach to bridge this gap by attaching document fingerprints to the embedding to estimate the expectation of potential queries.
Outcome: Experiments across diverse datasets, languages, and embedding models confirm the proposed solution is simple-yet-effective with zero additional index storage, retrieval latency, training costs, or catastrophic forgetting and hallucination issues.
LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to long reasoning traces are hard to tune and fail to adapt to evolving LLMs.
Approach: They propose a reinforcement learning framework that optimizes the length of reasoning traces by a Lagrangian primal–dual method.
Outcome: The proposed framework reduces the average reasoning length by 60% across diverse tasks while maintaining competitive performance.
From Chat Logs to Collective Insights: Aggregative Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to analyzing large-scale conversation logs treat interactions as independent, missing critical insights.
Approach: They propose a task that requires models to reason explicitly over thousands of user-chatbot interactions to answer aggregational queries.
Outcome: The proposed task requires models to reason over thousands of user-chatbot interactions to answer aggregational queries such as identifying emerging concerns among demographics.
Trust Within? Seek Beyond? Knowledge Boundary Aware Policy Optimization for Agentic Search (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to augment large language models with external knowledge suffer from a lack of calibration regarding the model’s knowledge boundary.
Approach: They propose a reinforcement learning framework that explicitly aligns retrieval decisions with quantified knowledge states.
Outcome: The proposed framework outperforms strong baselines while exhibiting reduced hallucination rates.
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification (2025.acl-long)

Copied to clipboard

Challenge: MM-Verifier and MM Reasoner are a powerful multimodal reasoning model . large language models (LLMs) have demonstrated exceptional performance across tasks spanning myriad domains.
Approach: They propose a method which combines tree search and verification to generate high-quality chain-of-thought data.
Outcome: The proposed method outperforms all larger models on the MathCheck, MathVista, and MathVerse benchmarks.
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks provide only final questions and answers, while lacking intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query.
Approach: They propose to build a multi-hop reasoning model that is primarily constructed automatically by large language models and designed to support step-by-step validation.
Outcome: The proposed benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks.
EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue (2026.acl-long)

Copied to clipboard

Challenge: Existing dialogue models address empathy and ethical safety in isolation . Existing models fail to adapt their behavior as ethical risk and user emotion evolve .
Approach: They propose a risk-aware framework that integrates ethical-emotional alignment in dialogue as an explicit turn-level decision problem.
Outcome: The proposed framework achieves more consistent ethical guidance and emotional engagement than baselines in ethically complex interactions.
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for scientific diagram generation rely on image-centric metrics or evaluation of intermediate symbolic representations rather than final rendered images.
Approach: They propose a structure-first benchmark for evaluating scientific diagram generation from pixel-level outputs.
Outcome: The proposed benchmark evaluates scientific diagram generation directly from pixel-level outputs.
Towards Reverse Engineering of Language Models: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Due to the vast amounts of data and computational resources required for model development, protecting the model’s parameters and training data has become an urgent and crucial concern.
Approach: They define "reverse engineering" techniques as attacks on large language models and provide an in-depth analysis of them.
Outcome: The proposed attacks are described as “reverse engineering” techniques on LMs and provide an introduction to existing protective strategies.
HopRAG: Multi-Hop Reasoning for Logic-Aware Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Traditional retrieval systems focus on lexical or semantic similarity rather than logical relevance.
Approach: They propose a new RAG framework that augments retrieval with logical reasoning . hopRAG uses a retrieve-reason-prune mechanism to explore multi-hop neighbors .
Outcome: The proposed framework outperforms conventional retrieval systems and state-of-the-art benchmarks on multi-hop QA tasks.
Removal of Hallucination on Hallucination: Debate-Augmented RAG (2025.acl-long)

Copied to clipboard

Challenge: erroneous or biased retrieval can mislead generation, compounding hallucinations.
Approach: They propose a framework that integrates multi-agent debates into retrieval and generation stages to improve retrieval reliability.
Outcome: The proposed framework improves retrieval reliability, reduces hallucinations and significantly improves overall factual accuracy.
Bi-Tuning with Collaborative Information for Controllable LLM-based Sequential Recommendation (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to optimize sequential recommendation systems rely on item ID sequences, but they lack collaborative knowledge and limited controllability.
Approach: They propose a simple bi-tuning framework with collaborative information for controllable Large Language Model-based Sequential Recommendation (Laser) they incorporate learnable virtual tokens at prefix and suffix of input text to adapt LLMs with collaborative knowledge .
Outcome: The proposed framework outperforms state-of-the-art recommendations on real-world datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations