Papers with automation
Beyond Abstracts: A New Dataset, Prompt Design Strategy and Method for Biomedical Synthesis Generation (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to automate systematic reviews of papers are slow and incomplete . authors propose a new method to automating the systematic review process . |
| Approach: | They propose a method for automatic synthesis generation using a dataset and prompting-based method. |
| Outcome: | The proposed method improves the existing model and prompts the system to generate high-quality syntheses. |
A Shoulder to Cry on: Towards A Motivational Virtual Assistant for Assuaging Mental Agony (2022.naacl-main)
Copied to clipboard
| Challenge: | Mental health disorders are one of the primary causes of disability worldwide . lack of qualified and competent mental health professionals is a major problem . we propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Approach: | They propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Outcome: | The proposed system outperforms baselines in the evaluation of 7k dyadic conversations from a peer-to-peer support platform. |
UniToolBench: A Benchmark for Tool-Augmented LLMs in Cross-Domain, Universal Task Automation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing benchmarks that focus on manually curated tool graphs lack scalability and diversity across domains. |
| Approach: | They propose a large-scale, cross-domain benchmark to evaluate LLMs' ability to reason over and utilize interconnected tools for automation. |
| Outcome: | The proposed benchmark incorporates automated tool graph construction by formulating link prediction as a probabilistic task, instead of relying on categorical LLM outputs. |
What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods for prompting for large language models have limitations such as being labor-intensive or lacking insights. |
| Approach: | They propose a new approach that automatically distinguishes between random variations and systematic differences in language model outputs by using token patterns. |
| Outcome: | The proposed method combines both automation and human analysis to provide new insights into established prompt data. |