Papers by Elisei Rykov

5 papers
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images (2025.naacl-srw)

Copied to clipboard

Challenge: Existing methods to measure image common sense inconsistentness are difficult to implement because of their complexity.
Approach: They propose a visual commonsense model that leverages large vision-language models to extract atomic facts from images and a compact attention-pooling classifier to fine-tune it over encoded atomic fact.
Outcome: The proposed method outperforms existing methods on the WHOOPS! and WEIRD datasets while maintaining a compact attention-pooling classifier over encoded atomic facts.
Fine-Grained Semantic Comparison of Legal Documents using LLMs (2026.acl-srw)

Copied to clipboard

Challenge: Existing tools for detecting inconsistencies and contradictions in complex regulatory documents rely on character-level diffs.
Approach: They propose a benchmark to evaluate span-aware semantic comparison of legal documents . legDiff is an annotated pair of legal paragraphs that is automatically generated .
Outcome: The proposed benchmark evaluates span-aware semantic comparisons of legal documents . it generates synthetic training data that aligns with the manual annotations and mirrors the structure and label distribution of the benchmark .
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows (2026.acl-industry)

Copied to clipboard

Challenge: 45% of sessions automated in production without degrading support quality level . traditional automated processes are costly at scale and require manual rule authoring .
Approach: They propose a system that automates end-to-end customer support workflows inside an enterprise BPM platform.
Outcome: The proposed system automates 45% of sessions and reduces average handling time by 39% without degrading support quality level.
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing hallucination detection benchmarks operate at the sequence level and are limited to English . Existing methods lacking fine-grained, multilingual supervision are limited in English based on the sequence .
Approach: They propose a large-scale, multilingual dataset annotated with span-level hallucinations across 14 languages.
Outcome: The proposed dataset annotated with span-level hallucinations across 14 languages is scalable and cost-efficient.
Multimodal Evaluation of Russian-language Architectures (2026.eacl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are at the center of research attention, yet intelligence, limitations, and risks remain insufficiently understood.
Approach: They propose an open multimodal evaluation framework for Russian-spoken architectures . the framework is instruction-based and includes 18 newly constructed evaluation tasks .
Outcome: The proposed framework provides a replicable methodology for constructing multimodal benchmarks in Russian-spoken architectures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations