Papers with self-awareness
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety (2024.acl-demos)
Copied to clipboard
Chuang Liu, Linhao Yu, Jiaxuan Li, Renren Jin, Yufei Huang, Ling Shi, Junhui Zhang, Xinmeng Ji, Tingting Cui, Liutao Liutao, Jinwang Song, Hongying Zan, Sun Li, Deyi Xiong
| Challenge: | a rapid development of Chinese large language models poses big challenges for efficient LLM evaluation. |
| Approach: | They propose an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety. |
| Outcome: | The evaluation platform OpenEval benchmarks Chinese LLMs across capability, alignment and safety. |
Self-Aware Feedback-Based Self-Learning in Large-Scale Conversational AI (2022.naacl-industry)
Copied to clipboard
| Challenge: | Large-scale conversational AI systems require constant update to adapt to changing customer behavior and trends . lack of self-awareness in feedback-based systems can cause degradation of performance . et al., e. alderman and scott k. d. argues that such systems are not scalable enough to sustain the rapid update pace of conversational systems. |
| Approach: | They propose a superposition-based model that reactively learns local-adaptive decision boundaries . they propose rewritings with a bi-variate beta setting to improve the model's performance . |
| Outcome: | The proposed model improves the PR-AUC by 27.45% and reduces relative defect reductions by 31.22% . the proposed model can adapt faster to changes in global preferences across a large number of customers . |
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have a profound impact on a wide range of applications. |
| Approach: | They propose a framework to mitigate the tool-abuse behavior of Large Language Models and propose SMARTCAL to mitigate this issue. |
| Outcome: | The proposed framework improves the performance of LLMs on three datasets with two mainstream tool-use frameworks and shows an 8.6% increase in QA performance and 21.6 percent lower expected calibration error (ECE) than existing methods. |
AwarenessBench: Assessing Cognitive Capabilities of Language Models (2026.acl-long)
Copied to clipboard
Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen, Haoyuan Shi, Yu Wang, Wei Xu
| Challenge: | Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities. |
| Approach: | They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness . |
| Outcome: | Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds . |
SMART: Self-Aware Agent for Tool Overuse Mitigation (2025.findings-acl)
Copied to clipboard
Cheng Qian, Emre Can Acikgoz, Hongru Wang, Xiusi Chen, Avirup Sil, Dilek Hakkani-Tür, Gokhan Tur, Heng Ji
| Challenge: | Current Large Language Models (LLMs) lack self-awareness to balance reasoning and tool use, increasing computational overhead. |
| Approach: | They propose a paradigm that enhances an agent’s self-awareness to optimize task handling and reduce tool overuse. |
| Outcome: | The proposed model reduces tool use by 24% while improving performance by over 37%. |
REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to extend knowledge scope of large language models (LLMs) lack internal parametric knowledge, resulting in misusing external knowledge. |
| Approach: | They propose a retrieval-augmented approach that provides LLMs with potentially relevant documents through a module. |
| Outcome: | The proposed approach outperforms existing methods on four open-domain QA tasks. |
InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated the potential to mimic human social intelligence, but most studies focus on static self-report or performance-based tests. |
| Approach: | They propose a framework to assess LLMs' ability to understand and manage intentions by mapping their ability to infer the intentions of others in a game setting. |
| Outcome: | The proposed framework assesses LLMs' ability to understand and manage intentions in a game setting. |
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced natural language processing, demonstrating exceptional reasoning, tool usage, and memory capabilities. |
| Approach: | They propose a competition-based benchmark framework specifically designed to assess LLMs within multi-agent environments. |
| Outcome: | The proposed framework enhances the LLMs’ abilities in navigating complex social and cognitive dimensions by over threefold between the strongest and weakest LLM models. |
TT-SI: Self-Improving LLM Agents with Test-Time Training (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for language model fine-tuning are expensive and inefficient . existing methods rarely assess whether a training sample provides novel information . |
| Approach: | They propose a test-time self-improvement algorithm that generates a sample that model struggles with . they also explore Test-Time Distillation, which leverages 'stronger supervisors' |
| Outcome: | The proposed algorithm improves performance with +5.48% absolute accuracy gain on average across benchmarks. |
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models (MLLMs) have demonstrated exceptional capabilities in visual perception and understanding, but they also suffer from hallucinations, which limit their reliability as AI systems. |
| Approach: | They propose a benchmark to evaluate self-awareness in perception for multimodal large language models (MLLMs) by integrating image information with knowledge quadrants, and propose MM-SAP to evaluate this capability. |
| Outcome: | The proposed benchmark offers detailed analysis of MLLMs with self-awareness in perception. |
Don’t Lose Yourself! Empathetic Response Generation via Explicit Self-Other Awareness (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing attempts to generate empathy with other-awareness ignore to include self-a awareness to consider the own views of the self in their responses. |
| Approach: | They propose to include self-awareness to consider the own views of the self in empathetic response generation by integrating three stages of self-other awareness into the process. |
| Outcome: | The proposed method is superior to existing methods on the benchmark dataset. |
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches fail to fully capture all risks in tool utilization, resulting in financial loss or privacy leaking. |
| Approach: | They propose a framework to assess the safety of LLM tool utilization in a prospective manner, covering malicious user instructions and diverse practical toolsets. |
| Outcome: | The proposed framework significantly enhances LLMs’ self-awareness, enabling a more safer and trustworthy tool utilization. |
Towards Autonomous Tool Utilization in Language Models: A Unified, Efficient and Scalable Framework (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in tool learning for large language models have led to a new trend to allow LLMs to leverage external tools. |
| Approach: | They propose a framework for fine-tuning language models that categorizes queries into three different types . they also introduce an "instruct, execute, and reformat" strategy specifically designed for efficient data annotation . |
| Outcome: | The proposed framework surpasses open-source language models and GPT-3.5/4 on multiple evaluation metrics. |
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs (2026.acl-long)
Copied to clipboard
Siqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang
| Challenge: | Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations. |
| Approach: | They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions . |
| Outcome: | Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning. |