Papers by Tal August

12 papers
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for plain language summarization (PLS) lack a dedicated assessment metric and the suitability of text generation evaluation metrics is unclear due to unique transformations.
Approach: They propose a granular meta-evaluation testbed to evaluate PLS metrics . they identify four PLS criteria and define perturbations that sensitive metrics should be able to detect .
Outcome: The proposed testbed assesses performance of 14 existing metrics including scores, features, and prompt-based evaluations.
Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis (2025.acl-long)

Copied to clipboard

Challenge: Existing comparative summarization methods focus on surface-level semantic differences, which may not capture the most relevant distinctions.
Approach: They propose a framework which transforms scientific papers into LLM personas that debate their respective novelties.
Outcome: The proposed framework generates informative arguments and effectively contrasts papers, and supports researchers in their literature review.
Generating Scientific Definitions with Controllable Complexity (2022.acl-long)

Copied to clipboard

Challenge: Unfamiliar terminology and complex language can make understanding science difficult for readers.
Approach: They propose a task and dataset for defining scientific terms and controlling the complexity of generated definitions by a sequence-to-sequence approach.
Outcome: The proposed system is based on a sequence-to-sequence approach and human evaluations show it offers superior fluency while controlling complexity.
Detecting Urgency in Multilingual Medical SMS in Kenya (2022.aacl-srw)

Copied to clipboard

Challenge: Access to mobile phones has increased exponentially over the last 20 years, providing an opportunity to connect patients with healthcare interventions through mobile phones.
Approach: They propose to use natural language processing to improve nurses' management of messages from pregnant and postpartum women in Kenya.
Outcome: The proposed model did not reach the clinical usefulness threshold but could improve nurse workflow and responsiveness to urgent messages.
Personalized Jargon Identification for Enhanced Interdisciplinary Communication (2024.naacl-long)

Copied to clipboard

Challenge: Identifying and translating scientific jargon for individual researchers could speed up research, but current methods of jaron identification rely on corpus-level familiarity indicators rather than modeling researcher-specific needs.
Approach: They collect over 10K term familiarity annotations from 11 computer science researchers and investigate supervised and prompt-based methods to predict individual jargon familiarity.
Outcome: The proposed method improves jargon familiarity prediction by using domain, subdomain, and individual knowledge.
MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula (2024.findings-emnlp)

Copied to clipboard

Challenge: pedagogical experts spend months reviewing published math problems to ensure that they align with critical skills or concepts.
Approach: They propose a novel approach for evaluating language models' mathematical abilities by combining a dataset of 385 fine-grained descriptions of K-12 math skills and concepts with 9.9K math problems labeled with these standards.
Outcome: The proposed model can discern skills and concepts enabled by math content, and it can be used to assess language models' mathematical abilities.
Writing Strategies for Science Communication: Data and Computational Analysis (2020.emnlp-main)

Copied to clipboard

Challenge: Existing science communication guides do not provide empirical evidence for how their strategies are used in practice.
Approach: They propose to use prescriptive writing strategies to identify and train human-readable annotations that can be automatically recognized by a corpus of 128k science writing documents in English.
Outcome: The proposed system can be used to detect and suggest writing strategies for scientists by allowing them to automatically recognize them.
ParseJargon: Personalized Real-time Jargon Support in Online Meetings (2026.acl-demo)

Copied to clipboard

Challenge: Recent advances in speech-to-text technologies and large language models (LLMs) have the potential to overcome these limitations with automated, real-time jargon support.
Approach: They built an interactive LLM-powered system that provides real-time personalized jargon support tailored to users’ individual backgrounds in online meetings.
Outcome: The proposed system provides more precise jargon identification and enhanced participants’ comprehension, engagement, and appreciation of colleagues’ work.
Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance Prediction (2026.acl-long)

Copied to clipboard

Challenge: Prior work has modeled functional roles for meeting participants from simple speech features or used behavioral patterns to predict "latent" roles and team outcomes.
Approach: They operationalize a taxonomy of eight communication roles grounded in education literature and annotate a corpus of 6,307 Slack messages from 55 students across 18 teams.
Outcome: The proposed taxonomy outperforms lexical, conversational, and LLM-prompting baselines in predicting team performance after deliberation.
Research Borderlands: Analysing Writing Across Research Cultures (2025.acl-long)

Copied to clipboard

Challenge: a recent study has focused on improving cultural competence of language technologies, but most studies rely on synthetic setups and imperfect proxies of culture.
Approach: They use a human-centered approach to discover and measure language-based cultural norms and cultural competence of large language models (LLMs).
Outcome: The proposed framework identifies cultural norms that vary across research cultures and identifie a lack of cultural competence in LLMs.
All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text (2021.acl-long)

Copied to clipboard

Challenge: evaluators distinguish between human- and machine-authored text in three domains without training . evals' accuracy improved up to 55%, but it did not significantly improve across the three domain.
Approach: They examine the role untrained human evaluations play in NLG evaluation and propose ways to improve their evaluations.
Outcome: The evaluators distinguished between human- and machine-authored text at random chance level without training, but their accuracy did not improve across the three domains.
Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling (2024.acl-long)

Copied to clipboard

Challenge: a novel application of large language models (LLMs) to legal education helps non-experts learn complex legal concepts . authors find storytelling helps nonexperts understand complex legal terms and concepts compared to definitions .
Approach: They propose a novel application of large language models to legal education . they use LLMs to generate legal stories explaining complex legal concepts .
Outcome: The proposed method improves comprehension and interest among non-native speakers compared to definitions . the novel method also shows that non-experts retain more stories .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations