Papers by Tal August
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for plain language summarization (PLS) lack a dedicated assessment metric and the suitability of text generation evaluation metrics is unclear due to unique transformations. |
| Approach: | They propose a granular meta-evaluation testbed to evaluate PLS metrics . they identify four PLS criteria and define perturbations that sensitive metrics should be able to detect . |
| Outcome: | The proposed testbed assesses performance of 14 existing metrics including scores, features, and prompt-based evaluations. |
Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis (2025.acl-long)
Copied to clipboard
| Challenge: | Existing comparative summarization methods focus on surface-level semantic differences, which may not capture the most relevant distinctions. |
| Approach: | They propose a framework which transforms scientific papers into LLM personas that debate their respective novelties. |
| Outcome: | The proposed framework generates informative arguments and effectively contrasts papers, and supports researchers in their literature review. |
Generating Scientific Definitions with Controllable Complexity (2022.acl-long)
Copied to clipboard
| Challenge: | Unfamiliar terminology and complex language can make understanding science difficult for readers. |
| Approach: | They propose a task and dataset for defining scientific terms and controlling the complexity of generated definitions by a sequence-to-sequence approach. |
| Outcome: | The proposed system is based on a sequence-to-sequence approach and human evaluations show it offers superior fluency while controlling complexity. |
Detecting Urgency in Multilingual Medical SMS in Kenya (2022.aacl-srw)
Copied to clipboard
| Challenge: | Access to mobile phones has increased exponentially over the last 20 years, providing an opportunity to connect patients with healthcare interventions through mobile phones. |
| Approach: | They propose to use natural language processing to improve nurses' management of messages from pregnant and postpartum women in Kenya. |
| Outcome: | The proposed model did not reach the clinical usefulness threshold but could improve nurse workflow and responsiveness to urgent messages. |
Personalized Jargon Identification for Enhanced Interdisciplinary Communication (2024.naacl-long)
Copied to clipboard
| Challenge: | Identifying and translating scientific jargon for individual researchers could speed up research, but current methods of jaron identification rely on corpus-level familiarity indicators rather than modeling researcher-specific needs. |
| Approach: | They collect over 10K term familiarity annotations from 11 computer science researchers and investigate supervised and prompt-based methods to predict individual jargon familiarity. |
| Outcome: | The proposed method improves jargon familiarity prediction by using domain, subdomain, and individual knowledge. |
MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula (2024.findings-emnlp)
Copied to clipboard
| Challenge: | pedagogical experts spend months reviewing published math problems to ensure that they align with critical skills or concepts. |
| Approach: | They propose a novel approach for evaluating language models' mathematical abilities by combining a dataset of 385 fine-grained descriptions of K-12 math skills and concepts with 9.9K math problems labeled with these standards. |
| Outcome: | The proposed model can discern skills and concepts enabled by math content, and it can be used to assess language models' mathematical abilities. |
Writing Strategies for Science Communication: Data and Computational Analysis (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing science communication guides do not provide empirical evidence for how their strategies are used in practice. |
| Approach: | They propose to use prescriptive writing strategies to identify and train human-readable annotations that can be automatically recognized by a corpus of 128k science writing documents in English. |
| Outcome: | The proposed system can be used to detect and suggest writing strategies for scientists by allowing them to automatically recognize them. |
ParseJargon: Personalized Real-time Jargon Support in Online Meetings (2026.acl-demo)
Copied to clipboard
| Challenge: | Recent advances in speech-to-text technologies and large language models (LLMs) have the potential to overcome these limitations with automated, real-time jargon support. |
| Approach: | They built an interactive LLM-powered system that provides real-time personalized jargon support tailored to users’ individual backgrounds in online meetings. |
| Outcome: | The proposed system provides more precise jargon identification and enhanced participants’ comprehension, engagement, and appreciation of colleagues’ work. |
Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance Prediction (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work has modeled functional roles for meeting participants from simple speech features or used behavioral patterns to predict "latent" roles and team outcomes. |
| Approach: | They operationalize a taxonomy of eight communication roles grounded in education literature and annotate a corpus of 6,307 Slack messages from 55 students across 18 teams. |
| Outcome: | The proposed taxonomy outperforms lexical, conversational, and LLM-prompting baselines in predicting team performance after deliberation. |
Research Borderlands: Analysing Writing Across Research Cultures (2025.acl-long)
Copied to clipboard
| Challenge: | a recent study has focused on improving cultural competence of language technologies, but most studies rely on synthetic setups and imperfect proxies of culture. |
| Approach: | They use a human-centered approach to discover and measure language-based cultural norms and cultural competence of large language models (LLMs). |
| Outcome: | The proposed framework identifies cultural norms that vary across research cultures and identifie a lack of cultural competence in LLMs. |
All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text (2021.acl-long)
Copied to clipboard
| Challenge: | evaluators distinguish between human- and machine-authored text in three domains without training . evals' accuracy improved up to 55%, but it did not significantly improve across the three domain. |
| Approach: | They examine the role untrained human evaluations play in NLG evaluation and propose ways to improve their evaluations. |
| Outcome: | The evaluators distinguished between human- and machine-authored text at random chance level without training, but their accuracy did not improve across the three domains. |
Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling (2024.acl-long)
Copied to clipboard
Hang Jiang, Xiajie Zhang, Robert Mahari, Daniel Kessler, Eric Ma, Tal August, Irene Li, Alex Pentland, Yoon Kim, Deb Roy, Jad Kabbara
| Challenge: | a novel application of large language models (LLMs) to legal education helps non-experts learn complex legal concepts . authors find storytelling helps nonexperts understand complex legal terms and concepts compared to definitions . |
| Approach: | They propose a novel application of large language models to legal education . they use LLMs to generate legal stories explaining complex legal concepts . |
| Outcome: | The proposed method improves comprehension and interest among non-native speakers compared to definitions . the novel method also shows that non-experts retain more stories . |