Papers by Serguei Barannikov

6 papers
Robust AI-Generated Text Detection by Restricted Embeddings (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for artificial text detection are score-based and classifier-based . however, score-driven methods often rely on a score-derived score.
Approach: They investigate the ability of classifier-based detectors to transfer to unseen generators or semantic domains.
Outcome: The proposed methods improve the out-of-distribution classification score by up to 9% and 14%.
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are prone to producing so-called hallucinations, i.e., content that is factually or contextually incorrect.
Approach: They propose a TOpology-based HAllucination detector which quantifies the structural properties of graphs induced by attention matrices.
Outcome: The proposed detector achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources.
Acceptability Judgements via Examining the Topology of Attention Maps (2022.findings-emnlp)

Copied to clipboard

Challenge: Acceptability judgments are a key component of generative linguistics, but their ability to judge grammatical acceptability has not been explored.
Approach: They propose to exploit the geometric properties of the attention graph to evaluate the grammatical acceptability of sentences using topological data analysis.
Outcome: The proposed approach outperforms nine statistical and Transformer LM baselines on the BLiMP benchmark and the human-level performance on the same benchmark.
Artificial Text Detection via Examining the Topology of Attention Maps (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text detection lack interpretability and robustness towards unseen models.
Approach: They propose three new types of interpretable topological features based on topological data analysis which is currently understudied in the field of NLP.
Outcome: The proposed features outperform count- and neural-based baselines up to 10% on three common datasets and tend to be the most robust towards unseen GPT-style generation models.
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders (2025.findings-acl)

Copied to clipboard

Challenge: Existing algorithms for AI text detection lack interpretability, limiting their reliability in highstakes applications.
Approach: They extend existing ATD frameworks by using Sparse Autoencoders to extract features from Gemma-2-2b residual stream.
Outcome: The proposed algorithms can extract human-interpretable features from Gemma-2-2b model.
Quantifying Logical Consistency in Transformers via Query-Key Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Existing solutions for multi-step logical reasoning are unreliable . Existing methods generate intermediate steps but provide no internal check of coherence .
Approach: They propose a method that uses internal Query-Key interactions within transformer attention heads as a proxy for logical consistency.
Outcome: The proposed method reveals latent reasoning structure in large language models and provides a mechanistic alternative to ablation-based analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations