Challenge: Empirically, PHSIC is learned thousands of times faster than an RNN-based PMI while outperforming PMI in accuracy.
Approach: They propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions with a very short learning time.
Outcome: The proposed measure can be applied to sparse linguistic expressions with a very short learning time, and is called the pointwise HSIC.

Similar Papers

AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for aggregating large-form outputs overlook the nuance of neutral information and suffer from the high computational cost of fine-grained decomposition.
Approach: They propose a UQ framework that uses NLI neutral probabilities as triggers to distinguish irrelevance from uncertainty, reducing computation costs.
Outcome: Experiments on BIO and LongFact show that the proposed framework reduces inference time by 60% compared to full atomic decomposition.
On the Interpretability and Significance of Bias Metrics in Texts: a PMI-based Approach (2023.acl-short)

Copied to clipboard

Challenge: Word embeddings have been used to quantify biases in texts for years, but their statistical properties and advantages have not been studied.
Approach: They propose to use PMI-based metric to quantify bias in corpora by conditional probabilities and odds ratio to approximate it.
Outcome: The proposed measure can be approximated by an odds ratio, which makes statistical inferences cost-effective and meaningful.
Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded Dialogs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing metrics for faithfulness of response are not aligned with human judgments.
Approach: They propose a new metric that utilizes (Conditional) Point-wise Mutual Information (PMI) between the generated response and the source document, conditioned on the dialogue.
Outcome: The proposed metric improves on BEGIN benchmarks and shows that it generates more faithful responses than standard decoding techniques.
Dynamic PMI-Guided Contrastive Decoding Reduces Hallucination in Large Language Models: A Unified Framework of Fine-Grained Input Transformations (2026.findings-acl)

Copied to clipboard

Challenge: Despite the remarkable generation capabilities of large language models, the issue of hallucination remains a critical challenge.
Approach: They propose a contrastive decoding framework based on dynamic pointwise mutual information that disentangles spurious dependencies induced by context priors, lexical co-occurrences, and syntactic structures and prioritizes causal logic.
Outcome: The proposed framework significantly improves the model’s factuality and reasoning robustness while maintaining high computational efficiency.
PMI-Align: Word Alignment With Point-Wise Mutual Information Without Requiring Parallel Training Data (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that using contextualized embeddings from pre-trained multilingual language models could give us high quality word alignments without the need of parallel training data.
Approach: They propose a method which uses contextualized embeddings from pre-trained language models to extract word alignments without parallel training.
Outcome: The proposed method outperforms rival methods on five out of six language pairs.
Unsupervised Extractive Summarization using Pointwise Mutual Information (2021.eacl-main)

Copied to clipboard

Challenge: Unsupervised approaches to extractive summarization rely on notion of sentence importance defined by semantic similarity between a sentence and the document.
Approach: They propose a method to measure relevance and redundancy using PMI between sentences.
Outcome: The proposed method outperforms similarity-based methods on news, medical journal articles, and personal anecdotes.
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to improve language models' performance do not exploit this phenomenon .
Approach: They propose to use contextual information to select and construct prompts that improve model performance.
Outcome: The proposed methods show that the mutual information between a context and a question is an effective gauge for language model performance.
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have been increasing context lengths to enhance their performance, but at long context length, the KV cache becomes the new bottleneck in memory usage during inference.
Approach: They propose an approach which performs the attention operation in a low-rank space and reduces the KV cache memory overhead.
Outcome: The proposed approach reduces the KV cache memory overhead and reduces memory usage with minimal drop in performance over OPT, MPT, and Llama model families.
LEPO: Latent Reasoning Policy Optimization for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing latent reasoning methods that use chain of thought (CoT) are limited to selecting one discrete token at each reasoning step, which potentially induces information loss.
Approach: They propose a framework that injects controllable stochasticity into latent reasoning via Gumbel-Softmax, restoring LLMs' exploratory capacity and enhancing their compatibility with Reinforcement Learning (RL).
Outcome: The proposed framework preserves richer information for more comprehensive reasoning and is compatible with Reinforcement Learning (RL).
Top-n𝜎: Eliminating Noise in Logit Space for Robust Token Sampling of LLM (2025.acl-long)

Copied to clipboard

Challenge: Existing sampling methods that are sensitive to temperature scaling fail to distinguish between diversity and noise.
Approach: They propose a method that identifies informative tokens by eliminating noise directly in logit space and a new sampling method that is temperature-invariant.
Outcome: The proposed method outperforms existing methods with significant improvements in reasoning and creative writing tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations