Challenge: Existing studies focus on coping with social harms that large language models pose . however, discussions on sensitive issues can become toxic even if the users are well-intentioned.
Approach: They propose to use Korean dataset to test whether LLMs can generate offensive content and propagate prejudices.
Outcome: The proposed dataset shows that acceptable response generation improves for HyperCLOVA and GPT-3.

Similar Papers

A Chinese Dataset for Evaluating the Safeguards in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks.
Approach: They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples .
Outcome: The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset.
Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey (2023.eacl-main)

Copied to clipboard

Challenge: Recent advances in the capacity of large language models to generate human-like text have prompted a heated discourse around the risks of societal harms they introduce.
Approach: They propose a taxonomy of interventions organized around the different phases where they can be adopted to mitigate harms.
Outcome: The proposed methods are based on several prior works’ taxonomies of language model risks and provide an overview of strategies for detecting and ameliorating different kinds of risks/harms.
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation frameworks lack systematic methods to identify weaknesses in LLMs . Existing methods to evaluate LLM responses to sensitive topics are lacking .
Approach: They propose a FINE-grained response evaluation taxonomy for sensitive topics that breaks down helpfulness and harmlessness into errors across three main categories: Content, Logic, and Appropriateness.
Outcome: The proposed model outperforms refinement without guidance on Korean-sensitive questions . FINEST significantly improves the model responses across all three categories .
KoSBI: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Applications (2023.acl-industry)

Copied to clipboard

Challenge: Existing research and resources are not readily applicable in South Korea due to the differences in language and culture, both of which significantly affect the biases and targeted demographic groups.
Approach: They propose a social bias dataset of 34k pairs of contexts and sentences in Korean covering 72 demographic groups in 15 categories.
Outcome: The proposed dataset reduces social biases by 16.47%p on average for HyperClova (30B and 82B), and GPT-3.
Realistic Evaluation of Toxicity in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a large amount of data exposes large language models to toxicity and bias . prompt engineering can be easily bypassed with minimal prompt engineering.
Approach: They propose a dataset that uses manually crafted prompts to nullify protective layers of large language models.
Outcome: The proposed dataset shows that prompts can nullify protective layers of large language models.
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing static benchmarks for harmful content detection face limitations in scalability and diversity.
Approach: They propose a framework for synthesizing harmful content using persona-guided large language model agents.
Outcome: The proposed framework achieves a high success rate in harmful generation tests across multiple detection systems.
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search System (2025.findings-naacl)

Copied to clipboard

Challenge: generative LLMs have been used by industries for various purposes, but limited resources and limited experience hinder their deployment and maintenance.
Approach: They propose a taxonomy for sensitive search queries and outline approaches to generating generative LLMs.
Outcome: The proposed model can be used to analyze sensitive queries from real users.
Exploring Inherent Biases in LLMs within Korean Social Context: A Comparative Analysis of ChatGPT and GPT-4 (2024.naacl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been criticized for perpetuating stereotypes against diverse groups based on race, sexual orientation, and other attributes.
Approach: They devised a set of prompts that reflect major societal issues in Korea and assign varied personas to both ChatGPT and GPT-4 to assess the toxicity of the generated sentences.
Outcome: The proposed model produces twice the level of toxic content as ChatGPT and GPT-4 under certain conditions.
DELPHI: Data for Evaluating LLMs’ Performance in Handling Controversial Issues (2023.emnlp-industry)

Copied to clipboard

Challenge: a recent study of controversy-handling in large language models (LLMs) has shown that people may become increasingly dependent on such systems for information.
Approach: They propose to construct a controversial questions dataset using a subset of a publicly available dataset.
Outcome: The proposed dataset presents challenges concerning knowledge recency, safety, fairness, and bias.
A Large-Scale Dataset for Empathetic Response Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing empathetic datasets are limited in size and cost due to the cost of manual labor.
Approach: They propose to annotate 1M dialogues with 32 fine-grained emotions and eight empathetic response intents and the Neutral category using a silver dataset.
Outcome: The proposed pipeline compares the quality of the proposed dataset with a state-of-the-art gold dataset using offline experiments and visual validation methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations