Papers by Traian Rebedea

15 papers
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails (2023.emnlp-demo)

Copied to clipboard

Challenge: NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Approach: They propose to add programmable guardrails to LLMs that are user-defined, independent of the underlying LLM, and interpretable.
Outcome: The proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails.
Unsupervised Extraction of Dialogue Policies from Conversations (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are used to extract dialogue policies from conversational data.
Approach: They propose a method for extracting dialogue policies from conversational data using canonical forms and graph traversal algorithms.
Outcome: The proposed method gives conversation designers greater control and improves the process of developing dialogue policies.
BART-TL: Weakly-Supervised Topic Label Generation (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for topic labeling use weak labelers to train rankers . recent studies show that weakly-supervised methods can produce meaningful labels .
Approach: They propose a weakly-supervised method for assigning topic labels to models by using weak labelers.
Outcome: The proposed model can generate valuable and novel labels in a weakly-supervised manner and can be improved by adding other weak labelers or distant supervision on similar tasks.
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification (2025.emnlp-main)

Copied to clipboard

Challenge: **MultiMatch** is a semi-supervised learning (SSL) algorithm that combines co-training and consistency regularization with pseudo-labeling.
Approach: They propose a semi-supervised learning algorithm that integrates co-training and consistency regularization with pseudo-labeling.
Outcome: The proposed algorithm outperforms the second-best approach on 8 out of 10 setups from 5 natural language processing datasets and outperformed the second best by 3.26%.
Multimodal Semi-supervised Learning for Disaster Tweet Classification (2022.coling-1)

Copied to clipboard

Challenge: During natural disasters, people use social media platforms to post information about casualties and damage . annotating data can be burdensome, subjective and expensive . et al., 2018b; sohn e.t., 2020) proposed semi-supervised multimodal approach to improve performance on multimodal tasks.
Approach: They propose a semi-supervised approach to annotate unlabeled data from Twitter . they extend FixMatch algorithm to a multimodal setting to account for subjective data .
Outcome: The proposed approach improves on multimodal disaster tweet classification tasks.
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in instruction-tuning datasets focus on specific tasks like mathematical or logical reasoning.
Approach: They propose to use synthetic dialogues to help language models remain focused on the subject at hand during task-oriented interactions.
Outcome: The proposed dataset improves language models' ability to maintain topical coherence compared to general-purpose instruction-tuned LLMs like gpt-4-turbo and Mixtral-Instruct.
Natural Language Interface for Databases Using a Dual-Encoder Model (C18-1)

Copied to clipboard

Challenge: Existing approaches to train data-driven natural language interfaces for databases are limited and lack of large datasets is probably the main reason for the lack of complex machine learning approaches.
Approach: They propose a sketch-based two-step neural model for generating structured queries based on a user’s request in natural language.
Outcome: The proposed model improves on two recent large datasets suitable for data-driven solutions for natural language interfaces for databases.
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved almost human-like performance on various tasks.
Approach: They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian.
Outcome: The proposed model trains, evaluates and releases open-source models tailored for Romanian.
AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails (2025.naacl-long)

Copied to clipboard

Challenge: Existing safety-related content safety models are not well-suited for commercial use.
Approach: They propose a taxonomy that can be used to categorize safety risks . it combines human annotations with a multi-LLM "jury" system to assess safety . they plan to open-source Aegis2.0 data and models to aid in safety guardrailing .
Outcome: The proposed taxonomy can be used to assess the safety of human-LLM interactions . it can be trained on large, non-commercial datasets and is open-source .
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent research shows that reasoning-based language models offer significant benefits for LLM safety and guardrail applications.
Approach: They conduct an analysis of reasoning-based guardrail models for content moderation . they find reasoning models exhibit strong sample efficiency and inference efficiency .
Outcome: The reasoning-based guardrail models show strong performance across domains . the models achieve competitive performance with significantly fewer training examples .
Answering questions by learning to rank - Learning to rank by answering questions (D19-1)

Copied to clipboard

Challenge: Existing approaches to answer multiple-choice questions with no supporting documents are poor performance.
Approach: They propose a method which can be used to semantically rank documents extracted from Wikipedia . they propose 'semantic ranking' method that latently learns to rank documents by their importance .
Outcome: The proposed model achieves state-of-the-art accuracy on two datasets: ARC Easy and Challenge.
GunStance: Stance Detection for Gun Control and Gun Regulation (2024.acl-long)

Copied to clipboard

Challenge: Social media, especially Twitter, has been a melting pot for such debates.
Approach: They propose to annotate tweets relevant to shooting events into three classes: In-Favor, Against, and Neutral.
Outcome: The proposed approach outperforms supervised, semi-supervised, and LLM-based zero-shot models on the dataset.
Neural Approaches for Natural Language Interfaces to Databases: A Survey (2020.coling-main)

Copied to clipboard

Challenge: Interest in NLIDBs has resurged in the past years due to the availability of large datasets and improvements to neural sequence-to-sequence models.
Approach: They focus on key design decisions behind current state of the art neural approaches . they highlight linking question tokens to database schema elements .
Outcome: The proposed approaches are grouped into encoder and decoder improvements . they include better architectures for encoding the textual query taking into account the schema and improved generation of structured queries using autoregressive neural models.
Distilling the Knowledge of Romanian BERTs Using Multiple Teachers (2022.lrec-1)

Copied to clipboard

Challenge: Existing approaches to train pre-trained language models focus on the English language, thus widening the gap when considering low-resource languages.
Approach: They propose three versions of distilled BERT models for the Romanian language . they argue that the models offer performance comparable to their teachers .
Outcome: The proposed models perform comparable to their teachers, while being twice as fast on a GPU and 35% smaller.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations