Papers by Xuan-Phi Nguyen

13 papers
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly being used for reasoning intensive tasks.
Approach: They propose an algorithm that trains judges to be robust to positional biases . they also propose a benchmark that evaluates judges in diverse reasoning settings .
Outcome: The proposed algorithm outperforms GPT-4o and the next best small judge by 6.7% and 9% on ReasoningJudgeBench and JudgeBench.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations .
Approach: They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization.
Outcome: The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations.
ParaICL: Towards Parallel In-Context Learning (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to improve ICL performance are limited by the length of the input context.
Approach: They propose a method that utilizes all demonstration examples without exceeding the manageable context length.
Outcome: The proposed method can be scaled up to integrate with existing methods.
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are known to perform tasks by simply observing few exemplars, but performance among under-represented languages falls behind due to pre-training data imbalance.
Approach: They propose to assemble synthetic exemplars from high-resource languages to prompt LLMs to translate from any language into English and use them to create intra-lingual exemplar models to perform tasks in target languages.
Outcome: The proposed method outperforms supervised few-shot learning in LLMs of different sizes for translations between English and 13 Indic and 21 African low-resource languages.
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in handling long context inputs, but this comes at the cost of increased computational resources and latency.
Approach: They propose an algorithm that uses early LLM layers as filters to select and compress input tokens, reducing the context length for subsequent processing.
Outcome: The proposed method outperforms existing techniques on the Needle in a Haystack task while demonstrating comparable performance on the LongBench challenge.
Differentiable Window for Dynamic Local Attention (2020.acl-main)

Copied to clipboard

Challenge: Existing general purpose components for learning differentiable windows are hard to optimize.
Approach: They propose a new neural module and general purpose component for dynamic window selection that can enable more focused attentions over the input regions.
Outcome: The proposed approach improves on a myriad of NLP tasks including machine translation, sentiment analysis, subject-verb agreement and language modeling.
RST Parsing from Scratch (2021.naacl-main)

Copied to clipboard

Challenge: Fig. 1 shows a document level discourse parser that performs top-down end-to-end parsing without requiring segmentation .
Approach: They propose a top-down end-to-end formulation of document level discourse parsing in the Rhetorical Structure Theory framework.
Outcome: The proposed model outperforms existing methods in end-to-end parsing and parse with gold segmentation without handcrafted features.
Demystifying Domain-adaptive Post-training for Financial LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Domain-adaptive post-training of large language models (LLMs) has emerged as promising approach for specialized domains such as medicine and finance.
Approach: They propose a system to identify optimal adaptation criteria and training strategies for LLMs for the finance domain.
Outcome: The proposed model achieves state-of-the-art performance across a wide range of financial tasks.
SeaLLMs - Large Language Models for Southeast Asia (2024.acl-demos)

Copied to clipboard

Challenge: Existing large language models favor high-resource languages, such as English, at the expense of low-resourced and regional languages.
Approach: They propose a series of language models that specifically focuses on Southeast Asian languages.
Outcome: SeaLLM models outperform ChatGPT-3.5 in non-Latin languages by large margins . linguistic disparity impedes access to state-of-the-art AI technologies for non-English-speaking populations .
A Conditional Splitting Framework for Efficient Constituency Parsing (2021.acl-long)

Copied to clipboard

Challenge: Developing efficient and effective parsing solutions has always been a key focus in NLP.
Approach: They propose a generic seq2seq parsing framework that casts constituency parsers into a series of conditional splitting decisions.
Outcome: The proposed framework outperforms state-of-the-art (SoTA) methods in discourse parsing . it is based on a syntactic and discourse parsed model and is linear in number of nodes .
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math (2026.acl-long)

Copied to clipboard

Challenge: Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition .
Approach: They propose a human-annotated step-level verification benchmark that measures step- level verifiers at the frontier.
Outcome: The proposed benchmark outperforms closed-source models in step-level verification and the impact of scaling verifier compute.
A Hierarchical Encoding-Decoding Scheme for Abstractive Multi-document Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have been used for abstractive single-document summarization (SDS) but they may not be suitable for multi-document summary (MDS)
Approach: They propose to enforce hierarchy on both encoder and decoder to facilitate multi-document interactions for MDS.
Outcome: Xiao et al. (2019) outperforms or is competitive with the previous best models.
Efficient Constituency Parsing by Pointing (2020.acl-main)

Copied to clipboard

Challenge: Constituency parsing is a core task in natural language processing (NLP) Existing methods for constituency paring are greedy transition-based or globally optimized.
Approach: They propose a constituency parsing model that casts the problem into a series of pointing tasks.
Outcome: The proposed model achieves 92.78 F1 without pre-trained models, which is faster than existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations