Papers by Anthony Sicilia
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Effective interlocutors account for the uncertain goals, beliefs, and emotions of others. |
| Approach: | They propose to calibrate language models to better represent outcome uncertainty . they propose to use two methods to calibrated small open-source models . |
| Outcome: | The proposed fine-tuning strategies can calibrate smaller open-source models to beat pre-trained models 10x their size. |
Accounting for Sycophancy in Language Model Uncertainty Estimation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Effective human-machine collaboration requires machine learning models to externalize uncertainty. |
| Approach: | They propose a generalization of the definition of sycophancy bias and a new algorithm to account for scophancies in uncertainty estimation. |
| Outcome: | The proposed algorithm can account for sycophancy in uncertainty estimation process. |
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts. |
| Approach: | They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users. |
| Outcome: | The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm. |
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | ambiguities in natural language can lead to outputs that seem correct but fail to reflect the speaker’s intent. |
| Approach: | They propose to identify and then resolve ambiguities in natural language and propose metrics to quantify them. |
| Outcome: | The proposed metrics better correlate with human annotations than uncertainty baselines. |
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational Cues (2025.acl-long)
Copied to clipboard
| Challenge: | Typically, beliefs are held or not held, but there are situations where an individual's beliefs are better represented more flexibly. |
| Approach: | They propose a set of tasks that challenge language models to model the uncertainty of participants in a dialogue. |
| Outcome: | The proposed tasks show that language models can model the uncertainty of participants in a conversation. |
Dialogue is the Plan: From Interface to Joint Action in Agentic AI (2026.acl-short)
Copied to clipboard
| Challenge: | Large Language Model agents' language use is often used as an interface for instructing and reporting results. |
| Approach: | They argue that large language models are often used as an interface for instructingactions and reporting results. |
| Outcome: | We show that large-scale language models can be used to plan and act, yet their language is often used as an interface for instructing and reporting results. |
LEATHER: A Framework for Learning to Generate Human-like Text in Dialogue (2022.findings-aacl)
Copied to clipboard
| Challenge: | Generating coherent, human-like text for dialogue remains a challenge . lack of careful design of rewards can lead to mode-collapse in dialogue . |
| Approach: | They propose a theoretical framework for learning to generate text in dialogue . they propose to use data-shift to develop theoretical guarantees for learners . |
| Outcome: | The proposed framework improves both task-success and human-likeness of the generated text. |
Modeling Non-Cooperative Dialogue: Theoretical and Empirical Insights (2022.tacl-1)
Copied to clipboard
| Challenge: | a robust dialogue agent cannot assume a cooperative conversational counterpart when deployed in the wild. |
| Approach: | They propose a theoretical model for identifying non-cooperative interlocutors . they use reinforcement learning to implement multiple communication strategies . |
| Outcome: | The proposed model is validated by using reinforcement learning to implement multiple communication strategies. |
Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the ability of large language models to self-monitor and ask for human intervention. |
| Approach: | They propose a formal analysis of LLM self-reflection for uncertainty estimation using domain adaptation theory. |
| Outcome: | The proposed method improves accuracy and human interpretation on reasoning tasks. |
The Change that Matters in Discourse Parsing: Estimating the Impact of Domain Shift on Parser Error (2022.findings-acl)
Copied to clipboard
| Challenge: | Discourse analysis is very low on texts outside of the training distribution’s coverage, diminishing the practical utility of existing models. |
| Approach: | They propose to use a distribution shift statistic to estimate the error-gap of a discourse model and to use it to estimate it. |
| Outcome: | The proposed model can be estimated via distribution shift but does not correlate with change in the observed error of a classifier (i.e. error-gap). |
Learning to Generate Equitable Text in Dialogue from Biased Training Data (2023.acl-long)
Copied to clipboard
| Challenge: | Absence of equitable and inclusive principles can hinder the formation of common ground, which in turn negatively impacts the overall performance of the system. |
| Approach: | They propose to use theories of computational learning to study equitable text generation in dialogues using augmented data to prove formal definitions of equity in text generation and formal connections between human-likeness and learning equity. |
| Outcome: | The proposed model predicts relative-performance of multiple algorithms in generating equitable text as measured by human and automated evaluation. |
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations (2024.eacl-long)
Copied to clipboard
| Challenge: | Using demographic factors, pre-trained language models can adapt to demographic changes. |
| Approach: | They propose a framework to measure demographic alignment of language models with a target demographic for the first time. |
| Outcome: | The proposed framework outperforms human-machine language models in age-related tasks and outperformed a typical 21-year-old at memorization. |