Papers by Tianyi Ma

15 papers
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models (2025.acl-long)

Copied to clipboard

Challenge: composition of pre-training datasets for large language models remains undisclosed . current methods for evaluating data quality are limited by single-dimensional evaluation or redundancy-focused strategies.
Approach: They propose a multi-dimensional data selection method that integrates dimensions with existing quality metrics through learned optimal weightings.
Outcome: The proposed method doubles convergence speed for 1.3B model models and improves downstream task performance by 3.23%.
Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, but it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns.
Approach: They propose a benchmark that combines a constructionist Out-of-Sample dataset with reverse understanding probes to evaluate large-scale large-format models.
Outcome: The proposed model performs well on classical Chinese poetry benchmarks, but a performance gap persists . the model can complete famous couplets and can be used to understand a variety of texts.
PsyScam: A Benchmark for Psychological Techniques in Real-World Scams (2025.findings-emnlp)

Copied to clipboard

Challenge: PTs are employed by scammers to manipulate victims and cause lasting psychological trauma.
Approach: They propose a benchmark to capture the PTs employed in real-worldscam reports and investigate how LLMs can be utilized to generate variants of scams based on the pts and the contexts provided by thesescams.
Outcome: The proposed model can generate variants of scams based on the PTs employed in real-world scam reports and the contexts provided by these scams.
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing zero-shot LLM-based Vision-and-Language Navigation agents either encode images as textual scene descriptions, potentially oversimplifying visual details, or process raw image inputs, which can fail to capture abstract semantics required for high-level reasoning.
Approach: They propose to integrate large language models into embodied AI models by incorporating textual descriptions that facilitate analogical reasoning across images from multiple perspectives.
Outcome: The proposed approach improves the agent’s contextual understanding on the R2R dataset, showing that it can make better decisions based on the LLMs.
AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to agent routing emphasize cost efficiency while overlooking the fine-grained contextual and relational structure inherent in QA tasks.
Approach: They propose a framework that formulates multi-agent QA as a knowledge-graph-guided routing problem supervised by empirical performance signals.
Outcome: The proposed framework outperforms single-agent and ensemble baselines while generalizing across benchmarks and LLM backbones.
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Diet plays a critical role in human health, but tailoring dietary reasoning to individual health conditions remains a challenge.
Approach: a new benchmark evaluates dietary reasoning using a national health survey data set.
Outcome: The NGQA benchmark evaluates dietary reasoning across three tasks using a set of question complexity settings and baseline models.
Diagnosing Hidden Instabilities in Model Editing via Uncertainty Quantification (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to update large language models (LLMs) without expensive retraining are fragile under single-edit evaluation protocols.
Approach: They propose a framework that characterizes activation-based editing as a constrained intervention on intermediate representations.
Outcome: The proposed method reveals local knowledge conflicts invisible to existing benchmarks.
LLM-Empowered Class Imbalanced Graph Prompt Learning for Online Drug Trafficking Detection (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to combat illicit drug trafficking are impractical due to the scarcity of labeled samples and imbalance of classes.
Approach: They propose a Large Language Model-empowered Heterogeneous Graph Prompt Learning framework for illicit drug trafficking detection that leverages LLM to facilitate heterogeneous graph neural networks to effectively identify minority classes.
Outcome: The proposed framework is able to identify minority classes in class-imbalanced scenarios.
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents (2026.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a subfield of embodied AI that integrates natural language understanding, visual perception, and sequential decision-making to allow autonomous agents to navigate and interact within visual environments.
Approach: They propose a modular framework that introduces structured, skill-based reasoning into Transformer-based VLN agents.
Outcome: The proposed framework decomposes navigation into atomic skills handled by a specialized agent.
Can LLMs Convert Graphs to Text-Attributed Graphs? (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to model graph-structured data are limited by the availability of text-attributed graph data.
Approach: They propose a method to convert existing graphs into text-attributed graphs using large language models.
Outcome: The proposed method outperforms existing approaches that manually design node features on text-free graphs.
TheraAgent: Self-Improving Therapeutic Agent for Precise and Comprehensive Treatment Planning (2026.findings-acl)

Copied to clipboard

Challenge: Existing large language models rely on one-shot output without explicit verification, resulting in rough, incomplete, and potentially unsafe treatment plans.
Approach: They propose an agentic framework that replaces one-shot generation with an iterative generate-judge-refine pipeline.
Outcome: The proposed framework achieves state-of-the-art results on HealthBench, leading in Accuracy and Completeness.
RASPberry: Retrieval-Augmented Monte Carlo Tree Self-Play with Reasoning Consistency for Multi-Hop Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating and analyzing multiple document knowledge are not effective for multi-hop question answering.
Approach: They propose a Monte Carlo tree-based approach to inference-time scaling using RASPberry.
Outcome: Experimental results show that the proposed method achieves better inference-time scaling on smaller LLMs.
GUITester: Enabling GUI Agents for Exploratory Defect Discovery (2026.findings-acl)

Copied to clipboard

Challenge: Exploratory GUI testing is essential for software quality but suffers from high manual costs.
Approach: They propose a framework that decouples navigation from verification via two modules . they propose 143 tasks and a GUITestBench benchmark that features 26 defects .
Outcome: The proposed framework outperforms state-of-the-art benchmarks in 143 tasks and 26 defects.
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning (2026.acl-long)

Copied to clipboard

Challenge: Fill-in-the-Middle (FIM) models suffer from performance degradation and prohibitive latency.
Approach: They propose a search-and-replace infilling framework that integrates agentic verification and editing into a single-pass inference process.
Outcome: The proposed framework harmonizes completion tasks with the instruction-following priors of Chat LLMs, extending the paradigm from static infilling to dynamic context-aware editing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations