Challenge: Existing language models that answer recipes better than humans can mitigate risks to users.
Approach: They propose to specialize the analysis to more concrete applications and their plausible users.
Outcome: The proposed model answers recipes as well or better than humans who answered the questions on the web.

Similar Papers

Risk-graded Safety for Handling Medical Queries in Conversational AI (2022.aacl-short)

Copied to clipboard

Challenge: Conversational AI systems can engage in unsafe behaviour when handling medical queries that could lead to death.
Approach: They label medical queries with crowdsourced and expert annotations to identify the seriousness of the prompts and recognise the risk types posed by the responses.
Outcome: The results suggest that these tasks can be automated, but caution should be exercised, as errors can potentially be very serious.
Mitigating Societal Harms in Large Language Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: Recent studies have highlighted societal harms that can be caused by language generation models deployed in the wild.
Approach: They propose to use a typology of technical approaches to mitigating harms of language generation models to provide an overview of potential social issues in language generation including toxicity, social biases, misinformation, factual inconsistency, and privacy violations.
Outcome: The proposed typology addresses toxicity, biases, misinformation, factual inconsistency, and privacy violations in language generation models.
Handling and Presenting Harmful Text in NLP Research (2022.findings-emnlp)

Copied to clipboard

Challenge: Text data can pose a risk of harm, but the risks remain unresolved in the NLP community.
Approach: They propose an analytical framework categorising harms on three axes: harm type, whether harm sought as a feature of research design, whether harmful content is encountered when working on unrelated problems, and who it affects .
Outcome: The proposed framework categorises harms on three axes: harm type, whether harm sought as feature of research design, and whether harmful content is encountered when working on unrelated problems.
On Measures of Biases and Harms in NLP (2022.findings-aacl)

Copied to clipboard

Challenge: Recent studies show that natural language processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality.
Approach: They propose a framework for harms and questions to help practitioners understand biases . they propose measurable measures to detect and mitigate biased groups .
Outcome: The proposed framework provides a framework for harms and questions for practitioners to answer to guide the development of bias measures.
SafetyALFRED: Evaluating Safety-Conscious Planning of Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, but lack a critical gap in evaluating an agent.
Approach: They evaluate multimodal large language models with six categories of kitchen hazards . they propose a safety-based approach that prioritizes multi-step corrective actions .
Outcome: The proposed model can recognize hazards in QA settings, but average mitigation success rates are low . the proposed model is based on the embodied agent benchmark ALFRED .
Can NLI Models Verify QA Systems’ Predictions? (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent question answering systems perform well on benchmark datasets, but are not always well-calibrated to spot spurious answers under distribution shifts.
Approach: They propose to use natural language inference to verify whether answers are correct . they leverage large pre-trained models and recent prior datasets to construct powerful question conversion and decontextualization modules.
Outcome: The proposed approach improves the confidence estimation of a QA model across different domains, evaluated in a selective QA setting.
Do-Not-Answer: Evaluating Safeguards in LLMs (2024.findings-eacl)

Copied to clipboard

Challenge: a dataset evaluating harmful capabilities in large language models is available at https://github.com/Libr-AI/do-not-answer.
Approach: They collect an open-source dataset to evaluate the safeguards in large language models . they find that simple BERT-style classifiers can achieve results comparable to GPT-4 .
Outcome: The proposed dataset compares the safety of six popular LLMs to GPT-4 on automatic safety evaluation.
Designing, Evaluating, and Learning from Humans Interacting with NLP Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Approach: They will provide a systematic overview of key considerations and effective approaches for studying human-NLP model interactions.
Outcome: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
Mitigating Covertly Unsafe Text within Natural Language Systems (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on text safety have focused on overtly unsafe, covertly, or indirectly unsafe statements.
Approach: They propose a method to identify physical harm-causing statements as overtly, covertly or indirectly unsafe and a solution to mitigate the generation of such statements.
Outcome: The proposed methods identify the type of unsafe language that can cause physical harm and identify mitigation strategies to inspire future researchers to tackle this challenging problem.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations