Papers by Ariel Goldstein
Can LLMs Learn Macroeconomic Narratives from Social Media? (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing evaluation strategies for analyzing economic data with narratives are limited due to the complexity of the interplay of numerous factors and the difficulty in isolating causal relationships. |
| Approach: | They propose to use two Twitter datasets to capture economy-related narratives and use them to construct models using large language models. |
| Outcome: | The proposed models are able to predict macroeconomic fluctuations using the extracted or extracted narratives in two Twitter datasets. |
Decoding Stumpers: Large Language Models vs. Human Problem-Solvers (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in large language models have led to the development of systems 2 models that can solve complex tasks and predict human behavior. |
| Approach: | They compare the performance of four state-of-the-art LLMs to human participants and compare their results to stumpers, a unique single-step intuition problem that humans can easily verify. |
| Outcome: | The proposed models excel in solving stumpers and surpass human performance on stumpers, while humans exhibit superior skills in verifying solutions to the same problems. |
Confidence Improves Self-Consistency in LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Modern large language models (LLMs) demonstrate strong reasoning capabilities, driven in part by their capacity to generate a sequence of intermediate reasoning steps that lead them toward a final answer. |
| Approach: | They propose a method that performs a weighted majority vote based on confidence scores obtained directly from the model. |
| Outcome: | The proposed method outperforms self-consistency on nine models and four datasets, reducing the required number of reasoning paths by over 40% on average. |
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies. |
| Approach: | They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
| Outcome: | The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in LLMs have sparked a debate on whether they understand text. |
| Approach: | They propose two working definitions for understanding which explicitly acknowledge the question of consciousness and draw connections with a rich literature in philosophy, psychology and neuroscience. |
| Outcome: | The proposed models achieve impressive results on various benchmarks, seeming to generalize to unseen tasks and domains. |