Challenge: Test-time computing approaches that leverage additional computational resources during inference have been proven effective in enhancing large language model performance.
Approach: They propose a linearly scaling approach that leverages local consistency of neighboring unlabeled data to improve test-time predictions.
Outcome: The proposed approach outperforms baseline methods such as prompting and self-consistency across eight datasets and performs robustly across embedding models.

Similar Papers

s1: Simple test-time scaling (2025.emnlp-main)

Copied to clipboard

Challenge: OpenAI’s o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts.
Approach: They curate a small dataset s1K with 1,000 reasoning questions based on three criteria we validate through ablations: difficulty, diversity, and quality.
Outcome: The proposed model exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24).
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for assessing the reliability of Large Language Models (LLMs) by confidence elicitation require expensive computational overhead or suffer from poor calibration, making them unreliable for real-world deployment.
Approach: They propose a Generative Approach to Confidence Elicitation that enables reliable confidence elicitation for Large Language Models.
Outcome: The proposed method achieves the best discriminative capacity and calibration on open-ended tasks without resorting to additional sampling or an auxiliary model.
S*: Test Time Scaling for Code Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: S* is the first hybrid test-time scaling framework that significantly improves the coverage and selection accuracy of generated code.
Approach: They propose a hybrid test-time scaling framework that augments parallel scaling with sequential scaling to further increase the performance.
Outcome: The proposed framework outperforms existing scaling approaches in large-scale modeling and reasoning models.
Exploring Predictive Uncertainty and Calibration in NLP: A Study on the Impact of Method & Data Scarcity (2022.findings-emnlp)

Copied to clipboard

Challenge: Using low-resource languages, we assess the quality of uncertainty estimates from a wide array of approaches, but with more data.
Approach: They train models on sub-sampled datasets in three different languages to assess the confidence of a neural classifier.
Outcome: The proposed models train on sub-sampled datasets in three different languages and show that the quality of uncertainty estimates suffers with more data.
Thought calibration: Efficient and confident test-time scaling (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for teaching language models to be economical with their token budgets have failed to achieve the desired results.
Approach: They propose to calibrate a language model's growing body of thoughts to determine when new reasoning plateaus.
Outcome: The proposed framework preserves model performance with up to 60% reduction in thinking tokens on in-distribution data, and up to 20% in out-of-difference data.
UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for speculative decoding ignore device-specific verification costs and lack of mechanisms to assess draft token quality.
Approach: They propose a training-free, lossless speculative decoding framework that enables robust, plug-and-play LLM acceleration across diverse hardware configurations and languages.
Outcome: The proposed framework outperforms existing training-free methods while maintaining identical output quality across different hardware environments.
Scaling Evaluation-Time Compute with Reasoning Models as Evaluators (2026.findings-acl)

Copied to clipboard

Challenge: Language model (LM) evaluators that generate chain-of-thought reasoning are widely used for the assessment of LM responses.
Approach: They investigate whether increasing LMs' "thinking" time through scaling test-time compute can improve an LM's evaluation capability.
Outcome: The proposed reasoning models improve evaluation performance monotonically with the number of reasoning tokens generated, mirroring trends seen in LM reasoning.
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for scaling test-time computation rely on external models that introduce substantial computational overhead and fail to capture context-aware semantics.
Approach: They propose a method that leverages the generator LLM’s internal hidden states for clustering, eliminating the need for external models.
Outcome: The proposed method improves the computational efficiency of test-time scaling while maintaining or exceeding the performance of existing methods.
Not Far Away, Not So Close: Sample Efficient Nearest Neighbour Data Augmentation via MiniMax (2021.findings-acl)

Copied to clipboard

Challenge: Existing kNN-based augmentation techniques blindly incorporate all samples, but MiniMax-kNN uses a subset of augmented samples to maximize KL-divergence between teacher and student models.
Approach: They propose a semi-supervised approach to augmented data augmentation using kNN.
Outcome: The proposed method outperforms existing kNN-based augmentation techniques on several classification tasks and requires fewer augmented examples and less computation to achieve superior performance.
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of test-time scaling assume that a reasoning system should always give an answer to any question provided.
Approach: They propose to increase compute budget at inference time to increase confidence in correct responses by considering settings with non-zero levels of response risk.
Outcome: The proposed model can answer more questions correctly and have higher confidence in correct responses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations