WASA: A Web Application for Sequence Annotation (L18-1)

Copied to clipboard

Challenge: a major barrier to research on CS has been the lack of large multilingual, multi-genre CS-annotated corpora.
Approach: They propose a web-based annotation system that manages large-scale CS data annotation.
Outcome: The proposed system can manage large-scale multilingual code switching (CS) data annotation.

Similar Papers

Universal Semantic Annotator: the First Unified API for WSD, SRL and Semantic Parsing (2022.lrec-1)

Copied to clipboard

Challenge: Existing approaches to understanding textual information are still far from achieving true natural language understanding (NLU).
Approach: They propose a unified API for high-quality automatic annotations of texts in 100 languages through state-of-the-art systems for Word Sense Disambiguation, Semantic Role Labeling and Semantics Parsing.
Outcome: The proposed system can provide users with rich and diverse semantic information, help second-language learners, and integrate explicit semantic knowledge into downstream tasks and real-world applications.
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotation tools are not efficient for the annotation of corpora and are not error-free.
Approach: They propose to extend existing annotation tools by evaluating their flexibility and efficiency.
Outcome: The proposed system performs platform-independent multimodal annotations and annotates complex textual structures.
Web-based Annotation Tool for Inflectional Language Resources (L18-1)

Copied to clipboard

Challenge: Wasim is a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages.
Approach: They present a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages resources.
Outcome: The tool has high flexibility in segmenting tokens, editing, diacritizing, labelling tokens and segments.
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing (2025.findings-emnlp)

Copied to clipboard

Challenge: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances across five core NLP tasks are annotating by three bilingual annotators .
Approach: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances are annotating by three bilingual annotators .
Outcome: The dataset covers five core NLP tasks, including Token-level Language Identification, Matrix Language Identification and Named Entity Recognition.
mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences (2023.findings-emnlp)

Copied to clipboard

Challenge: a new text-to-text transformer is suitable for multilingual inputs . many of the current models are English-only, making them inapplicable to other languages.
Approach: They propose to extend a multilingual text-to-text transformer to handle long inputs . they use the mC4 dataset to pretrain the model to handle multilingual data .
Outcome: The proposed model performs well on multilingual summarization and question-answering tasks.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
A UIMA Database Interface for Managing NLP-related Text Annotations (L18-1)

Copied to clipboard

Challenge: despite the use of UIMA as a document-based schema, it does not provide native database support.
Approach: They develop a database interface to allow generic use of UIMA documents in database systems.
Outcome: The framework is evaluated in relation to file system-based storage and provides data protection.
FITAnnotator: A Flexible and Intelligent Text Annotation System (2021.naacl-demos)

Copied to clipboard

Challenge: In this paper, we introduce FITAnnotator, a generic web-based tool for efficient text annotation.
Approach: They propose a generic web-based tool for efficient text annotation.
Outcome: The proposed tool is based on a fully modular architecture and provides three kinds of interfaces to annotate instances, evaluate annotation quality and manage the annotation task for annotators, reviewers and managers.
Bringing Emerging Architectures to Sequence Labeling in NLP (2026.eacl-long)

Copied to clipboard

Challenge: Pretrained Transformer encoders are the dominant approach to sequence labeling . however, few have been applied to sequence labels on flat or simplified tasks .
Approach: They propose to use pretrained Transformer encoders to model relations across words . they find that the architectures adapt well across tagging tasks that vary in complexity .
Outcome: The proposed architectures perform well across tagging tasks across languages and datasets.
Translation via Annotation: A Computational Study of Translating Classical Chinese into Japanese (2026.eacl-long)

Copied to clipboard

Challenge: Ancient people translated classical Chinese into Japanese using a system of annotations placed around characters.
Approach: They propose to introduce an LLM-based annotation pipeline and construct a dataset from digitized open-source translation data to improve sequence tagging tasks.
Outcome: The proposed method achieves high scores on direct machine translation, but could serve as a supplement to LLMs to improve the quality of character’s annotation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations