Papers by Mingyang Song
Noisy Multi-Label Text Classification via Instance-Label Pair Correction (2024.findings-naacl)
Copied to clipboard
| Challenge: | Noise is a significant challenge for machine learning models, especially deep learning models. |
| Approach: | They propose a holistic selection metric that identifies noisy pairs while considering global loss information and instance-specific ranking information. |
| Outcome: | The proposed approach significantly improves performance in noisy multi-label text classification tasks. |
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study (2025.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation approaches to evaluate Large Language Models are affected by potential biases within LLMs. |
| Approach: | They propose two many-shot In-Context Learning (ICL) prompt templates to help LLM evaluators mitigate potential biases. |
| Outcome: | The proposed templates reduce biases by using in-context examples with model-generated rationales as references. |
Mitigating Over-Generation for Unsupervised Keyphrase Extraction with Heterogeneous Centrality Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing keyphrase extraction models incorrectly determine a keyphrase as a phrase but output other candidates as keyphrases because they contain the same word. |
| Approach: | They propose a new approach that detects both implicit and explicit centrality within a heterogeneous graph as the importance score of each candidate keyphrase. |
| Outcome: | The proposed approach outperforms state-of-the-art keyphrase extraction models on three benchmark datasets. |
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Improving training efficiency remains a challenge in large-scale Reinforcement Learning (RL). |
| Approach: | They propose a curriculum RL framework with stage-wise context scaling to improve RL training efficiency. |
| Outcome: | The proposed framework outperforms state-of-the-art reasoning models on five benchmarks and achieves 49.6% accuracy on AIME 2024. |
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset (2026.acl-long)
Copied to clipboard
YongKang Liu, Jiayang Yu, Mingyang Wang, Yiqun Zhang, Ercong Nie, Shi Feng, Daling Wang, Kaisong Song, Hinrich Schuetze
| Challenge: | Argumentation is a key part of human reasoning and decision-making . existing argumentative corpora focus on single-turn settings, but multi-turn dialogues are often realized as multi-turned dialogues . |
| Approach: | They present a dataset for strategic multi-turn argumentation dialogues . they annotate each utterance with five strategy types, allowing multiple strategies per utterrance . |
| Outcome: | The proposed dataset shows that explicit prompting improves fluency, stylistic coherence and persuasiveness. |
Importance Estimation from Multiple Perspectives for Keyphrase Extraction (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing keyphrase extraction methods focus on the part of phrase that is important . experimental results show that KIEMP outperforms existing keyphrase extracting methods . |
| Approach: | They propose to estimate the importance of keyphrase from multiple perspectives using a chunking module, ranking module and matching module. |
| Outcome: | The proposed method outperforms the state-of-the-art keyphrase extraction methods on six benchmark datasets. |
Match More, Extract Better! Hybrid Matching Model for Open Domain Web Keyphrase Extraction (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing models for keyphrase extraction use noisy information to filter the salient phrases from the document. |
| Approach: | They propose a hybrid matching model that combines representation-focused and interaction-based matching modules into a unified framework for improving keyphrase extraction. |
| Outcome: | The proposed model outperforms state-of-the-art keyphrase extraction models on the OpenKP dataset. |
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving (2025.emnlp-main)
Copied to clipboard
Mihir Parmar, Palash Goyal, Xin Liu, Yiwen Song, Mingyang Ling, Chitta Baral, Hamid Palangi, Tomas Pfister
| Challenge: | Recent studies have shown that decomposing complex problems into simple subtasks has significantly boosted the performance of large language models (LLMs). |
| Approach: | They propose a unified post-training framework that distills synthetic task decompositions and fine-tunes smaller LLMs via supervised and reinforcement-learning objectives to improve complex reasoning. |
| Outcome: | The proposed framework outperforms strong baselines on GSM8k and MATH benchmarks and shows that it can improve generalization capabilities on out-of-domain datasets. |
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks for long-context language models have lagged behind . however, there is still room for improvement as the context window and complexity of the tasks increase. |
| Approach: | They propose a long-context benchmark to evaluate the performance of long-text language models. |
| Outcome: | The proposed benchmarks show that the models perform better in long-context environments as the context window increases and complexity increases. |
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models (2025.acl-long)
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have achieved significant performance in complex reasoning tasks such as mathematics and code generation. |
| Approach: | They propose a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. |
| Outcome: | The proposed model measures the accuracy, soundness, and sensitivity of 25 models across open-source and closed-source large language models. |
Unsupervised Keyphrase Extraction by Learning Neural Keyphrase Set Function (2023.findings-acl)
Copied to clipboard
| Challenge: | Unsupervised keyphrase extraction is a task of extracting a keyphrase set that provides readers with highlevel information about the key ideas or important topics described in the document. |
| Approach: | They propose an unsupervised keyphrase extraction task that is a document-set matching problem instead of modeling the relevance between an individual phrase and the document. |
| Outcome: | The proposed model outperforms the state-of-the-art unsupervised keyphrase extraction baselines by a large margin. |
HyperRank: Hyperbolic Ranking Model for Unsupervised Keyphrase Extraction (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing unsupervised keyphrase extraction models overlook latent hierarchical structures when extracting keyphrases. |
| Approach: | They propose a new ranking model that models global and local contexts to estimate the importance of each candidate keyphrase within the hyperbolic space. |
| Outcome: | The proposed model outperforms state-of-the-art models in keyphrase extraction tasks. |
Hyperbolic Relevance Matching for Neural Keyphrase Extraction (2022.naacl-main)
Copied to clipboard
| Challenge: | Keyphrase extraction is a fundamental task in natural language processing that aims to extract a set of phrases with important information from a source document. |
| Approach: | They propose a hyperbolic matching model to explore keyphrase extraction in hyperbolical space using word embeddings from RoBERTa to capture hierarchical syntactic and semantic structures. |
| Outcome: | The proposed model outperforms the state-of-the-art models on six benchmark datasets and outperformed previous models. |
A Survey on Recent Advances in Keyphrase Extraction from Pre-trained Language Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | Keyphrase extraction is a key component in Natural Language Processing (NLP) systems for selecting a set of phrases from the document that could summarize the important information discussed in the source document. |
| Approach: | They propose to use supervised and unsupervised keyphrase extraction techniques to investigate the state-of-the-art models for keyphrase extracting. |
| Outcome: | The proposed keyphrase extraction system can significantly accelerate the speed of retrieval and help people get first-hand information from a long document quickly and accurately. |
Improving Embedding-based Unsupervised Keyphrase Extraction by Incorporating Structural Information (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing unsupervised keyphrase extraction models ignore the indicative role of the highlights in certain locations, leading to wrong keyphrases extraction. |
| Approach: | They propose a Highlight-Guided Unsupervised Keyphrase Extraction model that models phrase-document relevance via the highlights of documents and calculates cross-phrase relevance between all candidate phrases. |
| Outcome: | The proposed model outperforms the state-of-the-art unsupervised keyphrase extraction models on three benchmarks. |
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks for table reasoning are incomplete due to the complexity of the tables and user questions in real-world applications. |
| Approach: | They propose a Multi-scale spreadsheet benchmark with Meta operations for Table reasoning that incorporates two key features and a new criterion with six categories of meta operations for measuring the difficulty of each question. |
| Outcome: | The proposed model outperforms Claude-3.5-Sonnet with 77.4% accuracy on the existing benchmarks. |
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Podcast script generation is a challenging task for large language models, but evaluation resources are limited. |
| Approach: | They propose a benchmark to evaluate podcast script generation using a multifaceted evaluation framework . PodBench is a prototype that integrates quantitative constraints with LLM-based quality assessment . |
| Outcome: | The proposed framework integrates quantitative constraints with LLM-based quality assessment. |
On Temperature-Constrained Non-Deterministic Machine Translation: Potential and Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on the non-deterministic properties of language models, but these properties remain under-explored in machine translation. |
| Approach: | They propose a method that evaluates MT systems and identifies temperature-constrained non-deterministic MT as a distinct phenomenon. |
| Outcome: | The proposed framework provides higher-quality candidates than Deterministic MT under temperature constraints. |
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable capabilities in machine translation, but most MT-specific LLMs rely heavily on external supervision during training. |
| Approach: | They propose a reinforcement learning framework for machine translation that is reference-free and relies solely on self-judging rewards. |
| Outcome: | The proposed framework outperforms existing LLMs and larger general LLM models on English Chinese translation benchmarks and performs competitively with leading closed-source systems. |