Papers by Jennifer Hu
I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and Dragons (2023.acl-long)
Copied to clipboard
Pei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu
| Challenge: | Existing dialogue agents, while able to produce human-like responses, often do not model goal-driven and grounded language interactions. |
| Approach: | They propose to decompose and model teacher-student natural language interactions into (1) the DM’s intent to guide players toward a given goal; (2) the dm’s guidance utterance to the players expressing this intent; (3) a theory-of-mind model that anticipates the players’ reaction to the guidance one turn into the future. |
| Outcome: | The proposed task is based on a goal-driven and grounded environment with a teacher-student interaction model and theory-of-mind model. |
On Emergent Social World Models — Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LMs) possess astonishing abilities and prove useful for a plethora of downstream tasks, but controversy persists regarding how to conceptualize their capacities. |
| Approach: | They analyze LMs’ performance across seven subcategories of ToM abilities using a large localizer dataset than used in prior work. |
| Outcome: | The proposed models recruit shared computational mechanisms for general Theory of Mind (ToM) and language-specific pragmatic reasoning on a substantially larger localizer dataset than used in prior work. |
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models (2020.acl-demos)
Copied to clipboard
| Challenge: | SyntaxGym is an online platform and open-source framework for targeted syntactic evaluation of neural network language models. |
| Approach: | They propose to make targeted syntactic evaluations accessible to both experts in NLP and linguistics and reproducible across computing environments. |
| Outcome: | The proposed framework is reproducible across computing environments and standardized following the norms of psycholinguistic experimental design. |
One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies suggest large language models can capture certain behavioral patterns, but there are ongoing debates as to whether they are valid replacements for human subjects. |
| Approach: | They propose to use large language models as replacements for humans in behavioral research by relating the internal variability of simulated individuals to the population-level variability. |
| Outcome: | The proposed model can capture human-like conceptual diversity, but it is unclear whether post-training alignment affects models’ internal diversity. |
Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models (2021.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that structural supervision helps English language models learn generalizations about syntactic phenomena such as subject-verb agreement. |
| Approach: | They train LSTMs, Recurrent Neural Network Grammars, Transformer language models, and Transformer-parameterized generative parsing models on Mandarin Chinese datasets. |
| Outcome: | The proposed models learn aspects of Mandarin Chinese grammar that assess syntactic and semantic relationships. |
What Can String Probability Tell Us About Grammaticality? (2026.tacl-1)
Copied to clipboard
| Challenge: | linguistic theories have argued that language models have largely achieved grammatical competence, but they will assign non-zero probability to all strings. |
| Approach: | They propose a theoretical framework for analyzing string probabilities in linguistics based on simple assumptions about the generative process of corpus data. |
| Outcome: | The proposed framework makes three predictions using 280K sentence pairs in English and Chinese. |
Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling Approaches (2023.findings-emnlp)
Copied to clipboard
| Challenge: | People rely heavily on context to enrich meaning beyond what is literally said. |
| Approach: | They analyze how task goals, environmental contexts, and communicative affordances in each work enrich linguistic meaning. |
| Outcome: | The proposed frameworks are based on linguistic goals, environmental contexts, and communicative affordances to enrich linguistic meaning. |
A Systematic Assessment of Syntactic Generalization in Neural Language Models (2020.acl-main)
Copied to clipboard
| Challenge: | Existing work on syntactic knowledge models has not provided a clear picture of the properties required to produce proper syntaktic generalizations. |
| Approach: | They propose to evaluate syntactic knowledge of language models by varying model architectures . they find substantial differences in syntaktic generalization performance by model architecture . |
| Outcome: | The proposed model architectures outperform other architectures on a set of 34 English-language syntactic test suites. |
Expectations over Unspoken Alternatives Predict Pragmatic Inferences (2023.tacl-1)
Copied to clipboard
| Challenge: | Scalar inferences (SI) are a signature example of how humans interpret language based on unspoken alternatives. |
| Approach: | They propose to use context-driven expectations to explain scale-based inferences . they find that expectedness of a strong scalemate captures SI rates within and across scales - but only under meaning-based view of alternatives. |
| Outcome: | The proposed model captures SI rates by expectedness of a strong scalemate as an alternative, but only under a meaning-based view of alternatives. |
A fine-grained comparison of pragmatic language understanding in humans and language models (2023.acl-long)
Copied to clipboard
| Challenge: | Pragmatics and non-literal language understanding are essential to human communication . a long-standing challenge for artificial language models is to capture pragmatics . |
| Approach: | They compare language models and humans on seven pragmatic phenomena using curated English materials. |
| Outcome: | The proposed model achieves high accuracy and matches human error patterns . the results suggest pragmatic behaviors can emerge in models without explicit representations of mental states . |
Generating Bilingual Pragmatic Color References (N18-1)
Copied to clipboard
| Challenge: | Contextual influences on language often exhibit substantial cross-lingual regularities, but are obscured by semantic and syntactic differences. |
| Approach: | They propose a model that captures language-specific syntax and semantics while also exhibiting responsiveness to contextual difficulty in Chinese and English. |
| Outcome: | The proposed model can identify synonyms between the two languages, even with no exposure to parallel data. |
Prompting is not a substitute for probability measurements in large language models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Prompting is a dominant method for evaluating the linguistic knowledge of large language models (LLMs). |
| Approach: | They compare metalinguistic prompting and direct probability measurements as ways of measuring LLMs’ linguistic knowledge. |
| Outcome: | The results show that the results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization. |