| Challenge: | despite the use of UIMA as a document-based schema, it does not provide native database support. |
| Approach: | They develop a database interface to allow generic use of UIMA documents in database systems. |
| Outcome: | The framework is evaluated in relation to file system-based storage and provides data protection. |
Similar Papers
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing annotation tools are not efficient for the annotation of corpora and are not error-free. |
| Approach: | They propose to extend existing annotation tools by evaluating their flexibility and efficiency. |
| Outcome: | The proposed system performs platform-independent multimodal annotations and annotates complex textual structures. |
Datasets: A Community Library for Natural Language Processing (2021.emnlp-demo)
Copied to clipboard
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, Thomas Wolf
| Challenge: | Contemporary NLP systems use many different datasets at significantly varying scale and level of annotation. |
| Approach: | a community library for contemporary NLP is available at https://github.com/datasets . the library includes more than 650 unique datasets and has more than 250 contributors a year after its initial development . |
| Outcome: | the library includes more than 650 unique datasets and has more than 250 contributors . it supports a variety of cross-dataset research projects and shared tasks . |
Unlocking the Heterogeneous Landscape of Big Data NLP with DUUI (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Automated analysis of large corpora is a complex task, especially in terms of time efficiency. |
| Approach: | They propose a framework for automatic distributed analysis of text corpora that leverages Big Data experience and virtualization with Docker. |
| Outcome: | The proposed framework is scalable, flexible, lightweight, and feature-rich for automatic distributed analysis of text corpora. |
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)
Copied to clipboard
| Challenge: | TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations . |
| Approach: | They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory. |
| Outcome: | The GUI interface is user-friendly and provides two visualization modes. |
INTELMO: Enhancing Models’ Adoption of Interactive Interfaces (2023.emnlp-demo)
Copied to clipboard
| Challenge: | INTELMO is an easy-to-use library to help model developers adopt user-faced interactive interfaces for their language models. |
| Approach: | They propose a library to help model developers adopt user-faced interactive interfaces and articles from real-time RSS sources for their language models. |
| Outcome: | The proposed library categorizes common NLP tasks and provides default style patterns . it provides developers with fine-grained and flexible control over user interfaces . |
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect (2022.coling-1)
Copied to clipboard
| Challenge: | text-to-SQL is a language processing and database-based language processing (NLP) task is to convert natural utterances into SQL queries and its practical application is to build natural language interfaces to database systems. |
| Approach: | They propose to conduct a systematic survey of text-to-SQL to examine the challenges and potential future directions. |
| Outcome: | The proposed system converts natural utterances into SQL queries and is a representative task in semantic parsing. |
Neural Approaches for Natural Language Interfaces to Databases: A Survey (2020.coling-main)
Copied to clipboard
Radu Cristian Alexandru Iacob, Florin Brad, Elena-Simona Apostol, Ciprian-Octavian Truică, Ionel Alexandru Hosu, Traian Rebedea
| Challenge: | Interest in NLIDBs has resurged in the past years due to the availability of large datasets and improvements to neural sequence-to-sequence models. |
| Approach: | They focus on key design decisions behind current state of the art neural approaches . they highlight linking question tokens to database schema elements . |
| Outcome: | The proposed approaches are grouped into encoder and decoder improvements . they include better architectures for encoding the textual query taking into account the schema and improved generation of structured queries using autoregressive neural models. |
LightTag: Text Annotation Platform (2021.emnlp-demo)
Copied to clipboard
| Challenge: | LightTag is a text annotation tool built on the premise of global optimization by addressing annotator as well as project managers and data scientists who manage the work and enforce production quality. |
| Approach: | They propose to use LightTag to optimize the global NLP process by addressing annotators as well as project managers and data scientists who manage the work and enforce production quality. |
| Outcome: | The proposed tool is based on the theory of constraints and is available for free for academic use. |
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources (2026.acl-long)
Copied to clipboard
| Challenge: | Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties. |
| Approach: | They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
| Outcome: | The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
Applications of Natural Language Processing in Clinical Research and Practice (N19-5)
Copied to clipboard
| Challenge: | a tutorial on clinical NLP will introduce students and experts to the field . a focus will be on the use of clinical Nlp in clinical research and practice . |
| Approach: | This tutorial introduces the clinical use of natural language processing (NLP) techniques . it will review techniques and tools developed for the clinical domain . |
| Outcome: | This tutorial will introduce the clinical NLP methodologies and tools at two top universities . the goal of the tutorial is to encourage NLP researchers in the general domain to contribute . |