Papers by Fangru Lin
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)
Copied to clipboard
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei
| Challenge: | a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination . |
| Approach: | They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE . |
| Outcome: | The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE. |
TCP: a Benchmark for Temporal Constraint-Based Planning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks evaluate temporal reasoning and planning in isolation and under limited forms of complexity. |
| Approach: | They propose a temporal constraint-based planning benchmark that assesses temporal reasoning and planning capabilities in large language models. |
| Outcome: | The proposed model fails to perform well under limited constraints and lacks temporal grounding. |
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics (2024.lrec-main)
Copied to clipboard
| Challenge: | Scalar adjectives describe different domain scales and vary in intensity . they can be triggered by scalar adjective and require listeners to reason pragmatically about them. |
| Approach: | They probe different families of Large Language Models for their knowledge of the lexical semantics of scalar adjectives and one specific aspect of their pragmatics. |
| Outcome: | The proposed models encode rich lexical-semantic information about scalar adjectives but lack a good understanding of skalar diversity. |