Papers by Nicholas Deas
Summarization of Opinionated Political Documents with Varied Perspectives (2025.coling-main)
Copied to clipboard
| Challenge: | Political ideologies can lead people to develop misperceptions of groups with opposing opinions, such as the 2024 US presidential election, French legislative election, or the Brexit referendum. |
| Approach: | They propose a dataset and task for independently summarizing political perspectives in a set of opinionated news articles. |
| Outcome: | The proposed dataset and task evaluates models of varying sizes and architectures on a set of opinionated news articles. |
Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work evaluating emotion and affective understanding in large language models rely on predetermined label sets or focus on a singular evaluation task. |
| Approach: | They examine the ability of multilingual language models to predict any term used by an author to label their own feelings or emotions. |
| Outcome: | The proposed models perform poorly on three different tasks in English and Spanish. |
Reranking-based Generation for Unbiased Perspective Summarization (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks rely on traditional metrics for measuring key attributes such as coverage and faithfulness without verifying their applicability. |
| Approach: | They propose to use human annotations to measure perspective summary quality and reranking-based methods yield strong results. |
| Outcome: | The proposed methods show that they perform well with synthetically generated and reranking-labeled data. |
MASIVE: Open-Ended Affective State Identification in English and Spanish (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that fail to understand cultural and language influences the meaning of emotional terms like "love" a new study shows that smaller finetuned models outperform much larger LLMs on region-specific span prediction tasks. |
| Approach: | They propose to use a reddit reddits dataset to identify a set of affective states . they find that smaller finetuned multilingual models outperform larger LLMs . |
| Outcome: | The proposed model outperforms larger models on span prediction task even on region-specific Spanish affective states. |
Rejected Dialects: Biases Against African American Language in Reward Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Preference alignment via reward models can introduce new biases, hindering reward models’ fairness and equity. |
| Approach: | They propose a framework for evaluating dialect biases in reward models and conduct a case study on biase . they compare reward models' preferences and behavior on paired White Mainstream English and machine-translated and human-written AAL corpora. |
| Outcome: | The proposed framework evaluates dialect biases in reward models and compares them with paired White Mainstream English (WME) and machine-translated and human-written AAL corpora. |
Data Caricatures: On the Representation of African American Language in Pretraining Corpora (2025.acl-long)
Copied to clipboard
Nicholas Deas, Blake Vente, Amith Ananthram, Jessica A Grieser, Desmond U. Patton, Shana Kleiner, James R. Shepard Iii, Kathleen McKeown
| Challenge: | Recent work in linguistics and NLP has investigated the quantity and quality of AAL representation in pretraining corpora. |
| Approach: | They examine the quantity and quality of African American Language (AAL) representation in pretraining corpora. |
| Outcome: | The results show that AAL is underrepresented in all evaluated corpora compared to US demographics . they also show that most automated filters are more likely to conserve white Mainstream English (WME) texts over AAL . |
Evaluation of African American Language Bias in Natural Language Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages. |
| Approach: | They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM. |
| Outcome: | The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model. |
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions (2025.emnlp-main)
Copied to clipboard
| Challenge: | We introduce and study artificial impressions–patterns in LLMs’ internal representations of prompts that resemble human impressions and stereotypes based on language. |
| Approach: | They introduce and study artificial impressions–patterns in LLMs’ internal representations of prompts that resemble human impressions and stereotypes based on language. |
| Outcome: | The proposed models predict impressions and model behavior based on the two-dimensional Stereotype Content Model (SCM). |