Papers by Leon Derczynski

12 papers
Countering Hateful and Offensive Speech Online - Open Challenges (2024.emnlp-tutorials)

Copied to clipboard

Challenge: a comprehensive understanding of the field is needed to maintain respectful and inclusive online environments.
Approach: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Outcome: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
Anchoring Fine-tuning of Sentence Transformer with Semantic Label Information for Efficient Truly Few-shot Classification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fewshot text classification require substantial computing power and data.
Approach: They propose an efficient method to add task and label information to a sentence transformer model by contrastive learning and a triplet loss to enforce training instances to be closest to their own textual semantic label information.
Outcome: The proposed method achieves strong performance in data-sparse scenarios compared to existing methods across SST-5, Emotion detection, and AG News data even with just two examples per class.
Accelerated High-Quality Mutual-Information Based Word Clustering (2020.lrec-1)

Copied to clipboard

Challenge: Word clustering is a hard hierarchical clustering that uses short-range distributional information to construct clusters.
Approach: They propose to use a hierarchical clustering algorithm with a fixed-width beam to build clusters that outperform other word representations.
Outcome: The proposed method outperforms the original methods in the computation of hierarchical and flat clusters.
RWKV: Reinventing RNNs for the Transformer Era (2023.findings-emnlp)

Copied to clipboard

Challenge: recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability.
Approach: They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs.
Outcome: The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models.
Efficient Methods for Natural Language Processing: A Survey (2023.tacl-1)

Copied to clipboard

Challenge: Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data, but using only scale to improve performance means resource consumption also grows.
Approach: They propose to use data, time, storage, or energy to improve model performance.
Outcome: The proposed methods and findings provide guidance for conducting NLP under limited resources and point towards promising research directions for developing more efficient methods.
Detection and Resolution of Rumors and Misinformation with NLP (2020.coling-tutorials)

Copied to clipboard

Challenge: Detecting false and misleading claims on the web is a sub-field of NLP . this half-day tutorial presents the theory behind each of these steps and the state-of-the-art solutions.
Approach: This half-day tutorial presents the theory behind false and misleading claims detection . it covers the steps involved in identifying check-worthy claims, tracking claims and rumors, rumor collection and annotation, grounding claims against knowledge bases, and using stance to verify claims.
Outcome: This half-day tutorial presents the theory behind each of these steps and the state-of-the-art solutions.
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish .
Approach: They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language .
Outcome: The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 .
Annotating Online Misogyny (2021.acl-long)

Copied to clipboard

Challenge: Online misogyny is a category of online abusive language with serious and harmful social consequences.
Approach: They propose an iterative annotation process and a taxonomy of labels for annotating misogyny in natural written language and cite a high-quality dataset of annotated posts from social media posts.
Outcome: The proposed method aims to identify misogynistic language in natural written language and annotate it in social media posts using a high-quality dataset.
Handling and Presenting Harmful Text in NLP Research (2022.findings-emnlp)

Copied to clipboard

Challenge: Text data can pose a risk of harm, but the risks remain unresolved in the NLP community.
Approach: They propose an analytical framework categorising harms on three axes: harm type, whether harm sought as a feature of research design, whether harmful content is encountered when working on unrelated problems, and who it affects .
Outcome: The proposed framework categorises harms on three axes: harm type, whether harm sought as feature of research design, and whether harmful content is encountered when working on unrelated problems.
Quantifying the morphosyntactic content of Brown Clusters (N19-1)

Copied to clipboard

Challenge: Using corpora representing several language families, we show that word clusters are highly capable at distinguishing Parts of Speech.
Approach: They propose to use Brown and Exchange word clusters to represent morphosyntactic information in NLP systems.
Outcome: The proposed clusters are highly capable at distinguishing Parts of Speech and can be used to perform tasks dependent on morphosyntactic information.
Combating Security and Privacy Issues in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide a summary of risks and vulnerabilities in large language models . a number of studies have focused on security, privacy and copyright aspects of LLMs .
Approach: This tutorial seeks to provide a systematic summary of risks and vulnerabilities in large language models . authors will discuss security, privacy and copyright aspects of LLMs .
Outcome: This tutorial aims to provide a systematic summary of risks and vulnerabilities in large language models . it will also outline emerging challenges in security, privacy and reliability of LLMs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations