Papers by Shiqi Liu

11 papers
On the Universal Truthfulness Hyperplane Inside LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have explored hallucinations through the lens of internal representations, proposing mechanisms to decipher LLMs’ adherence to facts.
Approach: They propose to train a universal truthfulness hyperplane that distinguishes the model’s factually correct and incorrect outputs on a diverse collection of over 40 datasets and examine its cross-task, cross-domain, and in-domain generalization.
Outcome: The proposed model is able to distinguish factual outputs from incorrect outputs on a diverse collection of over 40 datasets.
Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for text classification based on large language models are difficult to apply directly to solve.
Approach: They propose a data quality enhancement method to improve LLMs' performance in classification tasks by using a greedy algorithm to select data and then performing fine-tuning.
Outcome: The proposed method improves the performance of large language models in text classification tasks and significantly improves training efficiency, saving nearly half of the training time.
MalURLBench: A Benchmark Evaluating Agents’ Vulnerabilities When Processing Web URLs (2026.findings-acl)

Copied to clipboard

Challenge: Existing models struggle to detect elaborately disguised malicious URLs, despite their ability to process malicious URL's.
Approach: They propose a benchmark to evaluate LLMs’ vulnerabilities to malicious URLs and a lightweight defense module to mitigate the vulnerability.
Outcome: The proposed framework analyzes 61,845 attack instances spanning 10 real-world scenarios and 7 categories of real malicious websites.
GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement learning fine-tuning methods suffer from inefficient exploration and slow convergence . supervised fine- tuning methods have limited performance ceiling and less solid theoretical foundation .
Approach: They propose a Guess-Think-Answer framework that combines supervised and supervised learning in a unified training paradigm.
Outcome: The proposed framework outperforms both standalone SFT and RL training models on three text classification benchmarks.
Lunar-Bench: Towards Evaluating Task-Oriented Reasoning of LLMs in Lunar Exploration Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on static, context-independent reasoning tasks and fail to capture constraints and dependencies of lunar missions.
Approach: They propose a benchmark to assess the task-oriented reasoning and decision-making performance of large language models through 3,000 tasks derived from mission procedures and documentation.
Outcome: The proposed model achieves 47.8% accuracy compared with 65.1% for human experts on 36 representative missions.
Generating Deep Questions with Commonsense Reasoning Ability from the Text by Disentangled Adversarial Inference (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for commonsense question generation produce shallow questions that can be answered by simple word matching.
Approach: They propose a task of commonsense question generation that aims to yield deep-level questions from the text.
Outcome: The proposed model can yield deep-level and to-the-point questions from the text.
CodeFort: Robust Training for Code Generation Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing research efforts to improve code generation models are inadequate . code generation model performance is degraded under small perturbations .
Approach: They propose a framework to improve the robustness of code generation models by generalizing code perturbations to enrich training data and enabling various robust training strategies.
Outcome: The proposed framework increases pass rates and robustness drop rate against code-syntax perturbations.
TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to enzyme–reaction retrieval suffer from poor generalization across tasks and distributions . TIGER is a text-informed generalized enzyme-reaction retrieval framework that bridges enzymes and biochemical reactions.
Approach: They propose a text-informed generalized enzyme-reaction retrieval framework that leverages protein-to-text generation models to distill textual knowledge from enzyme sequences.
Outcome: The proposed framework outperforms state-of-the-art methods in enzyme–reaction retrieval tasks and distributions.
Domain-Specific Data Generation Framework for RAG Adaptation (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) combines the language understanding and reasoning capabilities of large language models (LLMs) with external retrieval to produce domain-grounded responses.
Approach: They propose a scalable and modular data-centric framework for generating domain-grounded question–answer–context triples tailored to diverse RAG adaptation strategies.
Outcome: The proposed framework generates domain-grounded question–answer–context triples for multiple RAG adaptation strategies.
How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective (2026.acl-long)

Copied to clipboard

Challenge: Low-resource languages are a long-tail problem for multilingual LLMs due to limited high-quality training data.
Approach: They propose a method that translates high-quality, knowledge-rich English data into low-resource languages . they propose SynRank, which leverages synthetic data as positive samples to train a classifier .
Outcome: The proposed method matches handcrafted rule-based filtering by human experts and significantly improves knowledge-intensive tasks with less data.
Improving Context Fidelity via Native Retrieval-Augmented Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to fidelity to contexts rely on expensive supervised fine-tuning to generate evidence post-answer or train models to perform web searches without improving utilization of the given context.
Approach: They propose a native retrieval-augmented reasoning framework that integrates in-context evidence with the model’s own retrieval capabilities.
Outcome: The proposed approach outperforms supervised fine-tuning, retrieval-augmented generation methods, and external retrieval solutions on multiple real-world and counterfactual QA benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations