Papers by Ariel Goldstein

5 papers
Can LLMs Learn Macroeconomic Narratives from Social Media? (2025.findings-naacl)

Copied to clipboard

Challenge: Existing evaluation strategies for analyzing economic data with narratives are limited due to the complexity of the interplay of numerous factors and the difficulty in isolating causal relationships.
Approach: They propose to use two Twitter datasets to capture economy-related narratives and use them to construct models using large language models.
Outcome: The proposed models are able to predict macroeconomic fluctuations using the extracted or extracted narratives in two Twitter datasets.
Decoding Stumpers: Large Language Models vs. Human Problem-Solvers (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have led to the development of systems 2 models that can solve complex tasks and predict human behavior.
Approach: They compare the performance of four state-of-the-art LLMs to human participants and compare their results to stumpers, a unique single-step intuition problem that humans can easily verify.
Outcome: The proposed models excel in solving stumpers and surpass human performance on stumpers, while humans exhibit superior skills in verifying solutions to the same problems.
Confidence Improves Self-Consistency in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Modern large language models (LLMs) demonstrate strong reasoning capabilities, driven in part by their capacity to generate a sequence of intermediate reasoning steps that lead them toward a final answer.
Approach: They propose a method that performs a weighted majority vote based on confidence scores obtained directly from the model.
Outcome: The proposed method outperforms self-consistency on nine models and four datasets, reducing the required number of reasoning paths by over 40% on average.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in LLMs have sparked a debate on whether they understand text.
Approach: They propose two working definitions for understanding which explicitly acknowledge the question of consciousness and draw connections with a rich literature in philosophy, psychology and neuroscience.
Outcome: The proposed models achieve impressive results on various benchmarks, seeming to generalize to unseen tasks and domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations