Papers by Yannis Katsis

8 papers
Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours (2022.emnlp-demos)

Copied to clipboard

Challenge: Label Sleuth is an open source system for labeling and creating text classifiers which does not require coding skills nor machine learning knowledge.
Approach: *Label Sleuth* is an open source system for labeling and creating text classifiers which does not require coding skills nor machine learning knowledge.
Outcome: *Label Sleuth* is an open source system for labeling and creating text classifiers.
AIT-QA: Question Answering Dataset over Complex Tables in the Airline Industry (2022.naacl-industry)

Copied to clipboard

Challenge: Table Question Answering (Table QA) systems have been shown to be highly accurate when trained and tested on open-domain datasets built on top of Wikipedia tables.
Approach: They propose a domain-specific Table QA test dataset to test Table Question Answering systems on open-domain datasets built on top of Wikipedia tables.
Outcome: The proposed methods are highly accurate when tested on open-domain datasets built on top of Wikipedia tables.
Beyond Labels: Empowering Human Annotators with Natural Language Explanations through a Novel Active-Learning Architecture (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing low-resource learning techniques focus on label annotation while neglecting the natural language explanation of a data point.
Approach: They propose a novel architecture that leverages an explanation-generation model to produce explanations guided by human explanations and a prediction model that utilizes generated explanations toward prediction faithfully.
Outcome: The proposed architecture produces explanations guided by human explanations, a prediction model that utilizes generated explanations toward prediction faithfully, and a data diversity-based AL sampling strategy that benefits from the explanation annotations.
Zero-shot Topical Text Classification with LLMs - an Experimental Study (2023.findings-emnlp)

Copied to clipboard

Challenge: Topical text classification is an ancient, yet timely research area in natural language processing.
Approach: They compare the zero-shot performance of a variety of LMs over a large dataset of 23 publicly available TTC datasets.
Outcome: The proposed models outperform their counterparts over a large dataset and show that they perform better in a zero-shot scenario.
A Survey of the State of Explainable AI for Natural Language Processing (2020.aacl-main)

Copied to clipboard

Challenge: Recent years have seen significant advances in the quality of state-of-the-art models, but they have come at the expense of models becoming less interpretable.
Approach: This survey examines the current state of Explainable AI within the domain of NLP . they detail the operations and explainability techniques currently available for generating explanations for NLP models .
Outcome: This survey examines the state of explainable AI (XAI) within the domain of natural language processing . it focuses on the operations and explainability techniques currently available for NLP models .
Development of an Enterprise-Grade Contract Understanding System (2021.naacl-industry)

Copied to clipboard

Challenge: Currently, legal contract review remains an expensive and arduous process.
Approach: They describe a commercial system designed and deployed for contract understanding that enables legal professionals to review contracts.
Outcome: The proposed system is used by a wide range of enterprise users and solves three major challenges.
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations (2026.findings-acl)

Copied to clipboard

Challenge: Several benchmarks have been released to evaluate model performance on multi-turn retrieval augment generation tasks.
Approach: They propose to benchmark 666 conversations with over 2,800 conversation turns across 6 domains and a corpora that focuses on unanswerable questions and later conversation turns.
Outcome: The proposed benchmarks show that retrieval and generation models struggle on conversations with UNanswerable, UNderspecified, and NONstandalone questions and UNclear responses.
Identifying Noise in Human-Created Datasets using Training Dynamics from Generative Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing noise detection techniques for autoencoder models do not generalize to ArLMs due to differences in learning dynamics.
Approach: They propose a method that leverages training dynamics to rank datapoints from easy-to-learn to hard-tolear . TDRanker achieves at least 2x faster denoising than previous techniques .
Outcome: The proposed method demonstrates robustness across multiple model architectures and noise levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations