| Challenge: | BRAT is a widely used web-based text annotation tool, but lacks robust Python support for effective annotation management and processing. |
| Approach: | They propose an open-source extension of BRAT that introduces a solid Python backend and enables advanced annotation functions such as annotation typings, collection typings with statistical insights, corpus and annotation handling, object modifications, and entity-level evaluation. |
| Outcome: | The proposed extension streamlines annotation workflows, improves usability, and facilitates high-quality NLP research. |
Similar Papers
CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing (2020.lrec-1)
Copied to clipboard
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash
| Challenge: | CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Approach: | They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Outcome: | The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing. |
Fluid Annotation: A Granularity-aware Annotation Tool for Chinese Word Fluidity (L18-1)
Copied to clipboard
| Challenge: | Using word segmentation, we propose a wordhood annotation framework for Chinese language . word segmentations have been used for years in preprocessing NLP tasks for languages without explicit word delimiter. |
| Approach: | They propose a word-granularity-aware annotation framework for Chinese language . they argue that word segmentation is fluid in nature and that it rearranges the boundary of word segmentations and linguistic annotation. |
| Outcome: | The proposed framework rearranges the boundary between word segmentation and linguistic annotation and supports flexible annotation tasks for various linguistic and affective phenomena. |
Thresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Existing tools for fine-grained human evaluation lack adaptability to different domains or languages, or modify annotation settings according to user needs. |
| Approach: | They propose a unified platform for fine-grained evaluation that is customizable and deployable with a single YAML configuration file. |
| Outcome: | The proposed frameworks are based on a single YAML configuration file and can be easily extended to different domains or languages. |
X-AMR Annotation Tool (2024.eacl-demo)
Copied to clipboard
| Challenge: | X-AMR annotation tool is designed for annotating key corpus-level event semantics. |
| Approach: | They propose a new annotation tool for annotation of key corpus-level event semantics using machine assistance. |
| Outcome: | The proposed tool enhances the user experience and improves annotation efficiency. |
ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text Classification (2025.acl-long)
Copied to clipboard
| Challenge: | ProtoLens provides fine-grained, sub-sentence level interpretability for text classification. |
| Approach: | They propose a prototype-based model that provides fine-grained, sub-sentence level interpretability for text classification. |
| Outcome: | Extensive experiments show that ProtoLens outperforms both prototype-based and non-interpretable baselines on multiple text classification benchmarks. |
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI (2024.naacl-demo)
Copied to clipboard
Elron Bandel, Yotam Perlitz, Elad Venezian, Roni Friedman, Ofir Arviv, Matan Orbach, Shachar Don-Yehiya, Dafna Sheinwald, Ariel Gera, Leshem Choshen, Michal Shmueli-Scheuer, Yoav Katz
| Challenge: | Textual data processing pipelines are tailored to specific datasets, task and model combinations. |
| Approach: | They propose a library for customizable textual data preparation and evaluation tailored to generative language models. |
| Outcome: | Unitxt is a library for customizable textual data preparation and evaluation tailored to generative language models. |
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)
Copied to clipboard
| Challenge: | The first workshop on crowdsourcing for NLP is open to all . |
| Approach: | The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks. |
| Outcome: | The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data . |
FITAnnotator: A Flexible and Intelligent Text Annotation System (2021.naacl-demos)
Copied to clipboard
| Challenge: | In this paper, we introduce FITAnnotator, a generic web-based tool for efficient text annotation. |
| Approach: | They propose a generic web-based tool for efficient text annotation. |
| Outcome: | The proposed tool is based on a fully modular architecture and provides three kinds of interfaces to annotate instances, evaluate annotation quality and manage the annotation task for annotators, reviewers and managers. |
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data (2026.acl-long)
Copied to clipboard
Shiping Yang, Jie Wu, Wenbiao Ding, Ning Wu, Shining Liang, Ming Gong, Hongzhi Li, Hengyuan Zhang, Angel X. Chang, Dongmei Zhang
| Challenge: | Existing studies on robustness to explicit noise (e.g., document semantics) but overlook implicit noise (spurious features). |
| Approach: | They propose a framework to quantify the robustness of RAGs against spurious features by integrating a data synthesis pipeline and a taxonomy. |
| Outcome: | The proposed framework quantifies the robustness of RALMs against spurious features. |
PyRater: A Python Toolkit for Annotation Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | PyRater is an open-source Python toolkit for analysing corpora annotations. |
| Approach: | They propose to use PyRater to analyse corpora annotations. |
| Outcome: | The proposed model can be used to identify the best annotations and retrieve the gold standard. |