Papers by Serguei Barannikov
Robust AI-Generated Text Detection by Restricted Embeddings (2024.findings-emnlp)
Copied to clipboard
Kristian Kuznetsov, Eduard Tulchinskii, Laida Kushnareva, German Magai, Serguei Barannikov, Sergey Nikolenko, Irina Piontkovskaya
| Challenge: | Existing approaches for artificial text detection are score-based and classifier-based . however, score-driven methods often rely on a score-derived score. |
| Approach: | They investigate the ability of classifier-based detectors to transfer to unseen generators or semantic domains. |
| Outcome: | The proposed methods improve the out-of-distribution classification score by up to 9% and 14%. |
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs (2026.acl-long)
Copied to clipboard
Alexandra Bazarova, Andrei Volodichev, Aleksandr Yugay, Andrey Shulga, Alina Ermilova, Konstantin Polev, Julia Belikova, Rauf Parchiev, Dmitry Simakov, Maxim Savchenko, Andrey Savchenko, Serguei Barannikov, Alexey Zaytsev
| Challenge: | Large language models (LLMs) are prone to producing so-called hallucinations, i.e., content that is factually or contextually incorrect. |
| Approach: | They propose a TOpology-based HAllucination detector which quantifies the structural properties of graphs induced by attention matrices. |
| Outcome: | The proposed detector achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources. |
Acceptability Judgements via Examining the Topology of Attention Maps (2022.findings-emnlp)
Copied to clipboard
Daniil Cherniavskii, Eduard Tulchinskii, Vladislav Mikhailov, Irina Proskurina, Laida Kushnareva, Ekaterina Artemova, Serguei Barannikov, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev
| Challenge: | Acceptability judgments are a key component of generative linguistics, but their ability to judge grammatical acceptability has not been explored. |
| Approach: | They propose to exploit the geometric properties of the attention graph to evaluate the grammatical acceptability of sentences using topological data analysis. |
| Outcome: | The proposed approach outperforms nine statistical and Transformer LM baselines on the BLiMP benchmark and the human-level performance on the same benchmark. |
Artificial Text Detection via Examining the Topology of Attention Maps (2021.emnlp-main)
Copied to clipboard
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev
| Challenge: | Existing methods for text detection lack interpretability and robustness towards unseen models. |
| Approach: | They propose three new types of interpretable topological features based on topological data analysis which is currently understudied in the field of NLP. |
| Outcome: | The proposed features outperform count- and neural-based baselines up to 10% on three common datasets and tend to be the most robust towards unseen GPT-style generation models. |
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders (2025.findings-acl)
Copied to clipboard
Kristian Kuznetsov, Laida Kushnareva, Anton Razzhigaev, Polina Druzhinina, Anastasia Voznyuk, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
| Challenge: | Existing algorithms for AI text detection lack interpretability, limiting their reliability in highstakes applications. |
| Approach: | They extend existing ATD frameworks by using Sparse Autoencoders to extract features from Gemma-2-2b residual stream. |
| Outcome: | The proposed algorithms can extract human-interpretable features from Gemma-2-2b model. |
Quantifying Logical Consistency in Transformers via Query-Key Alignment (2025.emnlp-main)
Copied to clipboard
Eduard Tulchinskii, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
| Challenge: | Existing solutions for multi-step logical reasoning are unreliable . Existing methods generate intermediate steps but provide no internal check of coherence . |
| Approach: | They propose a method that uses internal Query-Key interactions within transformer attention heads as a proxy for logical consistency. |
| Outcome: | The proposed method reveals latent reasoning structure in large language models and provides a mechanistic alternative to ablation-based analysis. |