Papers by Traian Rebedea
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)
Copied to clipboard
Traian Rebedea, Leon Derczynski, Shaona Ghosh, Makesh Narsimhan Sreedhar, Faeze Brahman, Liwei Jiang, Bo Li, Yulia Tsvetkov, Christopher Parisien, Yejin Choi
| Challenge: | Pretrained generative models provide novel ways for users to interact with computers. |
| Approach: | This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol. |
| Outcome: | This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol. |
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails (2023.emnlp-demo)
Copied to clipboard
| Challenge: | NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems. |
| Approach: | They propose to add programmable guardrails to LLMs that are user-defined, independent of the underlying LLM, and interpretable. |
| Outcome: | The proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails. |
Unsupervised Extraction of Dialogue Policies from Conversations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are used to extract dialogue policies from conversational data. |
| Approach: | They propose a method for extracting dialogue policies from conversational data using canonical forms and graph traversal algorithms. |
| Outcome: | The proposed method gives conversation designers greater control and improves the process of developing dialogue policies. |
BART-TL: Weakly-Supervised Topic Label Generation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for topic labeling use weak labelers to train rankers . recent studies show that weakly-supervised methods can produce meaningful labels . |
| Approach: | They propose a weakly-supervised method for assigning topic labels to models by using weak labelers. |
| Outcome: | The proposed model can generate valuable and novel labels in a weakly-supervised manner and can be improved by adding other weak labelers or distant supervision on similar tasks. |
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification (2025.emnlp-main)
Copied to clipboard
| Challenge: | **MultiMatch** is a semi-supervised learning (SSL) algorithm that combines co-training and consistency regularization with pseudo-labeling. |
| Approach: | They propose a semi-supervised learning algorithm that integrates co-training and consistency regularization with pseudo-labeling. |
| Outcome: | The proposed algorithm outperforms the second-best approach on 8 out of 10 setups from 5 natural language processing datasets and outperformed the second best by 3.26%. |
Multimodal Semi-supervised Learning for Disaster Tweet Classification (2022.coling-1)
Copied to clipboard
| Challenge: | During natural disasters, people use social media platforms to post information about casualties and damage . annotating data can be burdensome, subjective and expensive . et al., 2018b; sohn e.t., 2020) proposed semi-supervised multimodal approach to improve performance on multimodal tasks. |
| Approach: | They propose a semi-supervised approach to annotate unlabeled data from Twitter . they extend FixMatch algorithm to a multimodal setting to account for subjective data . |
| Outcome: | The proposed approach improves on multimodal disaster tweet classification tasks. |
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in instruction-tuning datasets focus on specific tasks like mathematical or logical reasoning. |
| Approach: | They propose to use synthetic dialogues to help language models remain focused on the subject at hand during task-oriented interactions. |
| Outcome: | The proposed dataset improves language models' ability to maintain topical coherence compared to general-purpose instruction-tuned LLMs like gpt-4-turbo and Mixtral-Instruct. |
Natural Language Interface for Databases Using a Dual-Encoder Model (C18-1)
Copied to clipboard
| Challenge: | Existing approaches to train data-driven natural language interfaces for databases are limited and lack of large datasets is probably the main reason for the lack of complex machine learning approaches. |
| Approach: | They propose a sketch-based two-step neural model for generating structured queries based on a user’s request in natural language. |
| Outcome: | The proposed model improves on two recent large datasets suitable for data-driven solutions for natural language interfaces for databases. |
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)
Copied to clipboard
Mihai Masala, Denis Ilie-Ablachim, Alexandru Dima, Dragos Georgian Corlatescu, Miruna-Andreea Zavelca, Ovio Olaru, Simina-Maria Terian, Andrei Terian, Marius Leordeanu, Horia Velicu, Marius Popescu, Mihai Dascalu, Traian Rebedea
| Challenge: | Large Language Models (LLMs) have achieved almost human-like performance on various tasks. |
| Approach: | They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian. |
| Outcome: | The proposed model trains, evaluates and releases open-source models tailored for Romanian. |
AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails (2025.naacl-long)
Copied to clipboard
Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, Christopher Parisien
| Challenge: | Existing safety-related content safety models are not well-suited for commercial use. |
| Approach: | They propose a taxonomy that can be used to categorize safety risks . it combines human annotations with a multi-LLM "jury" system to assess safety . they plan to open-source Aegis2.0 data and models to aid in safety guardrailing . |
| Outcome: | The proposed taxonomy can be used to assess the safety of human-LLM interactions . it can be trained on large, non-commercial datasets and is open-source . |
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent research shows that reasoning-based language models offer significant benefits for LLM safety and guardrail applications. |
| Approach: | They conduct an analysis of reasoning-based guardrail models for content moderation . they find reasoning models exhibit strong sample efficiency and inference efficiency . |
| Outcome: | The reasoning-based guardrail models show strong performance across domains . the models achieve competitive performance with significantly fewer training examples . |
Answering questions by learning to rank - Learning to rank by answering questions (D19-1)
Copied to clipboard
| Challenge: | Existing approaches to answer multiple-choice questions with no supporting documents are poor performance. |
| Approach: | They propose a method which can be used to semantically rank documents extracted from Wikipedia . they propose 'semantic ranking' method that latently learns to rank documents by their importance . |
| Outcome: | The proposed model achieves state-of-the-art accuracy on two datasets: ARC Easy and Challenge. |
GunStance: Stance Detection for Gun Control and Gun Regulation (2024.acl-long)
Copied to clipboard
Nikesh Gyawali, Iustin Sirbu, Tiberiu Sosea, Sarthak Khanal, Doina Caragea, Traian Rebedea, Cornelia Caragea
| Challenge: | Social media, especially Twitter, has been a melting pot for such debates. |
| Approach: | They propose to annotate tweets relevant to shooting events into three classes: In-Favor, Against, and Neutral. |
| Outcome: | The proposed approach outperforms supervised, semi-supervised, and LLM-based zero-shot models on the dataset. |
Neural Approaches for Natural Language Interfaces to Databases: A Survey (2020.coling-main)
Copied to clipboard
Radu Cristian Alexandru Iacob, Florin Brad, Elena-Simona Apostol, Ciprian-Octavian Truică, Ionel Alexandru Hosu, Traian Rebedea
| Challenge: | Interest in NLIDBs has resurged in the past years due to the availability of large datasets and improvements to neural sequence-to-sequence models. |
| Approach: | They focus on key design decisions behind current state of the art neural approaches . they highlight linking question tokens to database schema elements . |
| Outcome: | The proposed approaches are grouped into encoder and decoder improvements . they include better architectures for encoding the textual query taking into account the schema and improved generation of structured queries using autoregressive neural models. |
Distilling the Knowledge of Romanian BERTs Using Multiple Teachers (2022.lrec-1)
Copied to clipboard
Andrei-Marius Avram, Darius Catrina, Dumitru-Clementin Cercel, Mihai Dascalu, Traian Rebedea, Vasile Pais, Dan Tufis
| Challenge: | Existing approaches to train pre-trained language models focus on the English language, thus widening the gap when considering low-resource languages. |
| Approach: | They propose three versions of distilled BERT models for the Romanian language . they argue that the models offer performance comparable to their teachers . |
| Outcome: | The proposed models perform comparable to their teachers, while being twice as fast on a GPU and 35% smaller. |