Challenge: a tool for elicitation and management of process metadata is presented . detailed documentation of workflows is an arduous and neglected task .
Approach: They propose a software tool for elicitation and management of process metadata.
Outcome: The proposed tool minimizes the additional effort required for producing a sustainable workflow documentation.

Similar Papers

A Systematic Review of Reproducibility Research in Natural Language Processing (2021.eacl-main)

Copied to clipboard

Challenge: Despite the recent progress in reproducibility, the field is far from reaching a consensus on how reproducibility should be defined, measured and addressed.
Approach: They propose to provide a wide-angle snapshot of current work on reproducibility in NLP.
Outcome: The proposed work will provide a wide-angle snapshot of current work on reproducibility in NLP.
Towards Reproducible Machine Learning Research in Natural Language Processing (2022.acl-tutorials)

Copied to clipboard

Challenge: a tutorial on reproducibility in ML addresses the problem of research results that are not reproducible.
Approach: They propose a tutorial to ensure reproducible research in ML with an emphasis on computational linguistics and NLP.
Outcome: The proposed tutorial focuses on computational linguistics and NLP . it provides a framework for using reproducibility as a teaching tool in university-level computer science programs.
DocAgent: A Multi-Agent System for Automated Code Documentation Generation (2025.acl-demo)

Copied to clipboard

Challenge: Existing methods for generating documentation using Large Language Models (LLMs) produce incomplete, unhelpful, or factually incorrect outputs.
Approach: They propose a novel collaborative system that uses topological code processing for incremental context building to generate documentation by agents.
Outcome: The proposed system outperforms baselines in completeness, helpfulness, and truthfulness evaluations.
CiteLab: Developing and Diagnosing LLM Citation Generation Workflows via the Human-LLM Interaction (2025.acl-demo)

Copied to clipboard

Challenge: Existing frameworks for enabling Large Language Models to generate citations are lacking . however, they can still produce hallucinated responses that are non-factual or irrelevant to the input.
Approach: They propose an open-source and modular framework for enabling LLMs to generate citations in Question-Answering tasks.
Outcome: The proposed framework is extensible and paired with a visual interface, Citefix, facilitating case study and modification of existing citation generation methods.
ReproEvalCard: A Reporting Standard for Reproducible Evaluation of LLM Pipelines (2026.acl-short)

Copied to clipboard

Challenge: Existing evaluation standards for multistage pipelines are inconsistent, leaving the reproducibility and independent validation of published evaluations unclear.
Approach: They propose a lightweight reporting standard that specifies the minimum artifacts required to reproduce and validate LLM evaluations.
Outcome: The proposed standard audits 55 pipeline-based LLM papers published between 2022 and 2025 and quantifies the availability of reproducibility-critical evaluation artifacts.
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to large language models rely on static templates or manual workflows.
Approach: AdaptFlow is a language-based meta-learning framework inspired by model-agnostic meta- learning.
Outcome: AdaptFlow outperforms manual and automated workflows on question answering, code generation and mathematical reasoning benchmarks.
Annotating Research Infrastructure in Scientific Papers: An NLP-driven Approach (2023.acl-industry)

Copied to clipboard

Challenge: a pipeline is used to identify, extract and link research infrastructure used in scientific publications.
Approach: They propose a natural language processing pipeline for the identification, extraction and linking of Research Infrastructure (RI) used in scientific publications.
Outcome: The proposed pipeline can be used to identify, extract and link research infrastructure used in scientific publications.
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
ToolDNA: Autonomous Evolution of Tool Metadata for Robust Dialogue Agents (2026.findings-acl)

Copied to clipboard

Challenge: Task-oriented dialogue systems face labor-intensive manual metadata tuning and sparse reinforcement learning (RL) rewards that fail to diagnose invocation errors.
Approach: They propose a framework that enables auto-evolution of policy networks and tool metadata via RL . a tool metadata loop coordinates metadata through policy-generated candidates during rollouts .
Outcome: The proposed framework achieves +11% problem resolution and +54% accuracy over commercial LLMs with prompt engineering and +25%/+35% over supervised fine-tuning.
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Metadata extraction relies heavily on manual annotation of documents.
Approach: They propose a framework that leverages Large Language Models to automatically extract metadata attributes from scientific papers covering datasets of languages other than Arabic.
Outcome: The proposed framework automates the extraction of metadata attributes from Arabic scientific papers using large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations