Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances (2026.eacl-long)
Copied to clipboard
| Challenge: | Common proxies such as Mean Length of Utterance (MLU), lexical diversity (vocd-D), and readability indices are dominated by length and ignore conversational context, missing aspects of response quality such as reasoning depth, topic maintenance, and discourse planning. |
| Approach: | They propose a framework that classifies the Previous Adult Utterance Type and scores the child’s response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’ s contribution to advancing the discourse). |
| Outcome: | The proposed framework assesses the child's response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’s contribution to advancing the discourse). |
Similar Papers
ChildEval:WHEN LARGE LANGUAGE MODELS MEET CHILDREN’S PERSONALITIES (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved remarkable success in effectively understanding and generating human language, leading to a revolutionary era in LLMs. |
| Approach: | They propose a benchmark to evaluate LLMs' ability to infer and follow child-centered preferences in long-context conversations. |
| Outcome: | The proposed benchmark spans five top-level and fourteen sub-level categories covering children’s daily lives and development. |
Meaning Beyond Truth Conditions: Evaluating Discourse Level Understanding via Anaphora Accessibility (2025.acl-long)
Copied to clipboard
| Challenge: | Existing assessments of understanding at the lexical and sentence levels are limited to lexica and sentence level, but few of them target whether LLMs accurately represent and update states of natural language discourse. |
| Approach: | They propose anaphora accessibility as a diagnostic for assessing discourse understanding . they use a dataset inspired by theoretical research in dynamic semantics to evaluate human and LLM performance. |
| Outcome: | The proposed dataset shows that humans and LLMs align on some tasks and diverge on others. |
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence (2024.naacl-short)
Copied to clipboard
| Challenge: | Existing lexical or semantic metrics cannot accurately capture the discourse coherence of long-form text generation. |
| Approach: | They propose to use automatic metric to quantify the discourse divergence between two long-form articles . they use a theory of functional discourse structure to model the coherence of long-formed texts . |
| Outcome: | The proposed metric outperforms existing evaluation methods on three datasets from representative domains. |
Measuring Contextual Informativeness in Child-Directed Text (2025.coling-main)
Copied to clipboard
Maria R. Valentini, Téa Y. Wright, Ali Marashian, Jennifer M. Ellis, Eliana Colunga, Katharina von der Wense
| Challenge: | Recent advances in natural language processing (NLP) have made it possible to generate children's stories with a single word. |
| Approach: | They propose a task of measuring contextual informativeness in children's stories and a large language model to automate the task. |
| Outcome: | The proposed method outperforms baselines and can generalize to measuring contextual informativeness in adult-directed text. |
On Measuring Context Utilization in Document-Level MT Systems (2024.findings-eacl)
Copied to clipboard
| Challenge: | Current studies on document-level translation evaluation focus on sentence-level models which are inadequate for capturing improvements in discourse phenomena. |
| Approach: | They propose to complement accuracy-based evaluation with measures of context utilization. |
| Outcome: | The proposed model can be used to handle context-dependent discourse phenomena using an automatic annotation tool. |
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)
Copied to clipboard
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems (2020.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics are not designed to cope with this flexibility. |
| Approach: | They propose to group the qualities into three groups to obtain a single metric called USL-H. |
| Outcome: | The proposed metric achieves good correlations with human judgment and maintains its configurability towards different aspects and metrics. |
Beyond Surprisal: A Dual Metric Framework for Lexical Skill Acquisition in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | Existing learning curves capture when and how a model learns to use words correctly, but they neglect the equally important skill of avoiding incorrect usage. |
| Approach: | They propose a new metric which measures a model's capacity to refrain from using words in unexpected or unexpected contexts. |
| Outcome: | The proposed metric measures the model's ability to refrain from using words in unexpected or unexpected contexts. |
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts (2025.emnlp-main)
Copied to clipboard
Yuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min, Taewon Yun, Minjeong Ban, Kim Yul, Hwanjun Song
| Challenge: | HAMLET is a framework for evaluating the long-context comprehension of large language models. |
| Approach: | They propose a framework for evaluating the long-context comprehension of large language models . HAMLET structures key information of source texts into a three-level hierarchy . |
| Outcome: | HAMLET achieves 90% agreement with expert judgments while reducing evaluation cost by up to 25. |
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement (2024.eacl-short)
Copied to clipboard
| Challenge: | Memorizing and utilizing speakers’ personas is a common practice for response generation in long-term conversations, yet human-authored datasets often provide uninformative persona sentences that hinder response quality. |
| Approach: | They propose a framework that leverages commonsense-based persona expansion to address such issues in long-term conversations. |
| Outcome: | The proposed framework facilitates better response generation via human-like persona refinement. |