Challenge: Anthropomorphism is commonplace in people's interactions with technology . anthropomorphizing language can suggest undue accountability and agency in technologies .
Approach: They propose an automatic metric of implicit anthropomorphism in language . they use a masked language model to quantify how non-human entities are implicitly framed as human by the surrounding context.
Outcome: The proposed metric measures how non-human entities are implicitly framed as human by the surrounding context.

Similar Papers

Thinking beyond the anthropomorphic paradigm benefits LLM research (2026.acl-long)

Copied to clipboard

Challenge: anthropomorphism is an automatic and unconscious response that occurs even in advanced technical expertise.
Approach: They argue that anthropomorphism is an automatic and unconscious response . they identify and examine five assumptions that shape research across the LLM development lifecycle .
Outcome: The proposed framework challenges assumptions that shape research across the LLM development lifecycle and offers promising directions for LLMs.
Exploring Supervised Approaches to the Detection of Anthropomorphic Language in the Reporting of NLP Venues (2025.findings-acl)

Copied to clipboard

Challenge: anthropomorphisms are used to describe technical contributions to AI . however, they also give potential for incorrect assumptions about LLMs' capacities.
Approach: They undertake a corpus annotation of one year of ACL abstracts and news articles from the same period and train a regression classifier based on BERT to identify anthropomorphic language.
Outcome: The proposed method can automatically label abstracts for their degree of anthropomorphism based on their corpus and reporting on diachronic and inter-venue findings.
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit anthropomorphism characteristics – human-like qualities portrayed across their outlook, language, behavior, and reasoning functions.
Approach: They propose that anthropomorphism should be treated as a design concept that can be intentionally tuned to support user goals.
Outcome: The proposed design should reflect interaction between artifact designers and interpreters, and should be based on cues embedded in the artifactor and the (cognitive) responses of interpreters to the cue.
From Pixels to Personas: Investigating and Modeling Self-Anthropomorphism in Human-Robot Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that robots display human-like characteristics in dialogues . this anthropomorphism raises concerns about the accuracy of AI and its capabilities .
Approach: They propose to use a dataset to analyze self-anthropomorphic and non-self-anthropophilic responses in robots . they propose to combine these two types of responses to create a new category of bot responses .
Outcome: The proposed approach preserves the original dialogues from existing corpora and enhances them with paired responses: self-anthropomorphic and non-self-anthropophilic for each original bot response.
Mirages. On Anthropomorphism in Dialogue Systems (2023.emnlp-main)

Copied to clipboard

Challenge: Automated dialogue systems are anthropomorphised by developers and personified by users.
Approach: They propose to examine linguistic factors that contribute to the anthropomorphism of dialogue systems and the harms that can arise thereof.
Outcome: The proposed systems are anthropomorphised and personified by users . linguistic factors can also be used to reinforce gender stereotypes and conceptions of acceptable language.
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona simulation and Role-Playing Language Agents (RPLAs).
Approach: They propose a framework that treats psychological patterns as interacting causal forces and synthesizes 113 scenarios where 2-5 patterns reinforce, conflict, or modulate each other.
Outcome: The proposed framework outperforms Qwen3-32B on multi-pattern dynamics despite 4 fewer parameters.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy.
Approach: They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method.
Outcome: The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on how text generation systems can lead to harmful outcomes such as over-reliance, emotional dependence, dehumanization, deception, or even physical harm.
Approach: They propose to use an inventory of interventions to help identify possible interventions and provide a conceptual framework to help characterize the landscape of possible interventions.
Outcome: The proposed interventions are based on an inventory of interventions grounded in prior literature and a crowdsourcing study where participants edited system outputs to make them less human-like.
How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future (2025.emnlp-main)

Copied to clipboard

Challenge: Entity alignment (EA) is critical for knowledge graph (KG) integration.
Approach: They propose a taxonomy that categorizes methods in three stages: data preparation, feature embedding, and alignment.
Outcome: The proposed taxonomy categorizes methods in three key stages: data preparation, feature embedding, and alignment.
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for measuring gender stereotypical bias in language models are inconsistencies . lack of explicit standards in data gathering can have detrimental effects on results .
Approach: They propose that currently available benchmarks capture only partial facets of gender stereotypes . they apply a framework from social psychology to balance data across components of gender stereotypes based on stereotypical benchmarks.
Outcome: The proposed framework improves correlation between different benchmarks by using simple balancing techniques.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations