Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations

39 papers
Using and comparing Rhetorical Structure Theory parsers with rst-workbench (2021.eacl-demos)

Copied to clipboard

Challenge: Rhetorical Structure Theory (RST) parsers are usually only trained on English data .
Approach: rst-workbench is a web-based tool that lets users install and use RST parsers.
Outcome: rst-workbench is a web-based tool that lets users run multiple RST parsers simultaneously.
SF-QA: Simple and Fair Evaluation Library for Open-domain Question Answering (2021.eacl-demos)

Copied to clipboard

Challenge: Open-domain question answering (QA) requires large amounts of resources and is difficult to reproduce results due to complex configurations.
Approach: They propose a simple and fair evaluation framework for open-domain question answering (QA) it modularizes the pipeline open- domain QA system, making it easily accessible .
Outcome: The proposed evaluation framework is publicly available and anyone can contribute to the code and evaluations.
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)

Copied to clipboard

Challenge: a library for low-level processing of brahmic scripts is available for free.
Approach: They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts.
Outcome: The proposed library supports low-level processing of ten major south Asian Brahmic scripts.
CovRelex: A COVID-19 Retrieval System with Relation Extraction (2021.eacl-demos)

Copied to clipboard

Challenge: Existing challenges to making the system more practical include dealing with newly created and unknown data, and solving the performance gap when utilizing present data.
Approach: They propose a scientific paper retrieval system targeting entities and relations via relation extraction on COVID-19 scientific papers.
Outcome: The proposed system can be accessed via https://www.jaist.ac.jp/is/labs/nguyen-lab/systems/covrelex/.
MATILDA - Multi-AnnoTator multi-language InteractiveLight-weight Dialogue Annotator (2021.eacl-demos)

Copied to clipboard

Challenge: MATILDA is the first multi-annotator, multi-language dialogue annotation tool . it allows the creation of corpora, the management of users, the annotation of dialogues, the quick adaptation of the user interface to any language and the resolution of interannotation disagreement.
Approach: They propose to use MATILDA to create corpora, manage users, and annotation dialogues.
Outcome: The proposed tool supports the full pipeline for dialogue annotation, and non-technical people can use it.
AnswerQuest: A System for Generating Question-Answer Items from Multi-Paragraph Documents (2021.eacl-demos)

Copied to clipboard

Challenge: Existing systems that generate and answer questions in a question-and-answer format can facilitate reading comprehension.
Approach: They propose a system that integrates question answering and question generation tasks to produce a list of Q&A items for a text.
Outcome: The proposed system generates a catalog of Q&A items for a text.
T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition (2021.eacl-demos)

Copied to clipboard

Challenge: Language model (LM) pretraining has led to consistent improvements in many downstream tasks, including named entity recognition (NER).
Approach: They propose a Python library for NER LM finetuning that facilitates cross-domain and cross-lingual generalization of LMs finetuned on NER.
Outcome: The proposed library outperforms LMs trained on NERs in cross-domain and cross-lingual generalization tests on nine datasets.
Forum 4.0: An Open-Source User Comment Analysis Framework (2021.eacl-demos)

Copied to clipboard

Challenge: Using Forum 4.0, we analyze, aggregate, and visualize user comments based on labels defined by domain experts.
Approach: They introduce an open-source framework to semi-automatically analyze, aggregate, and visualize user comments based on labels defined by domain experts.
Outcome: The proposed framework can analyze, aggregate, and visualize user comments based on labels defined by domain experts.
SLTEV: Comprehensive Evaluation of Spoken Language Translation (2021.eacl-demos)

Copied to clipboard

Challenge: Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades.
Approach: They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations.
Outcome: The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language.
Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing (2021.eacl-demos)

Copied to clipboard

Challenge: Trankit is a lightweight, pre-trained toolkit for multilingual natural language processing.
Approach: They propose a transformer-based toolkit for multilingual natural language processing that trains pipelines over 100 languages and 90 pretrained pipelines for 56 languages.
Outcome: The proposed tool outperforms existing pipelines over sentence segmentation, part-of-speech tagging, morphological feature tabbing, and dependency parsing while maintaining competitive performance over tokenization, multi-word token expansion, and lemmatization over 90 Universal Dependencies treebanks.
DebIE: A Platform for Implicit and Explicit Debiasing of Word Embedding Spaces (2021.eacl-demos)

Copied to clipboard

Challenge: Recent research has shown that distributional word vector spaces often encode stereotypical human biases, such as racism and sexism.
Approach: They propose a platform that measures and mitigates bias in word embeddings by executing two (mutually composable) debiasing models.
Outcome: The proposed platform can measure and mitiga bias in word embeddings.
A Dashboard for Mitigating the COVID-19 Misinfodemic (2021.eacl-demos)

Copied to clipboard

Challenge: a new public dashboard aims to understand the impact of the COVID-19 misinfodemic on Twitter . the dashboard uses a curated catalog of COVId-19 related facts and debunks of misinformation .
Approach: They propose a public dashboard that matches tweets with COVID-19 misinformation . they also propose experiments to analyze the spread of misinformation on twitter .
Outcome: The proposed dashboard uses a curated catalog of COVID-19 related facts and debunks misinformation . it shows the most prevalent information from the catalog among Twitter users in user-selected geographic regions .
EasyTurk: A User-Friendly Interface for High-Quality Linguistic Annotation with Amazon Mechanical Turk (2021.eacl-demos)

Copied to clipboard

Challenge: Amazon Mechanical Turk (AMT) is one of the most popular crowd-sourcing platforms, allowing researchers from all over the world to create linguistic datasets quickly and at a relatively low cost.
Approach: They propose to improve the potential of Amazon Mechanical Turk by adding some new features to the tool.
Outcome: The proposed tool improves the performance of Amazon Mechanical Turk by adding new features.
ASAD: Arabic Social media Analytics and unDerstanding (2021.eacl-demos)

Copied to clipboard

Challenge: Currently, there are no publicly available tools for analyzing Arabic social media, such as ADIDA and CAMeL, which are not trained with Twitter data.
Approach: They propose to use Arabic social media analysis and unDerstanding to analyze tweets using a web API and a user interface.
Outcome: The proposed system allows users to determine dialects, sentiment, news category, offensiveness, hate speech, adult content, and spam in Arabic tweets.
COCO-EX: A Tool for Linking Concepts from Texts to ConceptNet (2021.eacl-demos)

Copied to clipboard

Challenge: ConceptNet is a semantic network which contains general commonsense facts about the world, e.g., Birds can fly or Computers are used for sending e-mails.
Approach: They propose a tool for Extracting Concepts from texts and linking them to ConceptNet, using the maximum relational information stored in ConceptNet.
Outcome: The proposed method extracts meaningful concepts from natural language texts and links them to conjunct concept nodes in ConceptNet, utilizing the maximum of relational information stored in the KnowledgeGraph.
A description and demonstration of SAFAR framework (2021.eacl-demos)

Copied to clipboard

Challenge: Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language .
Approach: They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework"
Outcome: The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect.
InterpreT: An Interactive Visualization Tool for Interpreting Transformers (2021.eacl-demos)

Copied to clipboard

Challenge: Using Transformer-based models for NLU/NLP tasks is a growing interest . but there are many open questions regarding the behavior of these models .
Approach: They present an interactive visualization tool for interpreting Transformer-based models.
Outcome: The tool can track and visualize token embeddings through each layer of a Transformer, highlight distances between certain token embeds, and identify task-related functions of attention heads using new metrics.
Representing ELMo embeddings as two-dimensional text online (2021.eacl-demos)

Copied to clipboard

Challenge: ELMoViz module adds support for contextualized embedding architectures, in particular for token embeddable word models.
Approach: They propose to add a module to the free and open-source WebVectors toolkit which provides lexical hyperlinks to word representations in static embedding models.
Outcome: The ELMoViz module adds support for contextualized embedding architectures, in particular for ELMa models.
LOME: Large Ontology Multilingual Extraction (2021.eacl-demos)

Copied to clipboard

Challenge: LOME is a system for performing multilingual information extraction with large ontologies.
Approach: They propose a system for multilingual information extraction with a framenet parser . LOME is available as a Docker container on Docker Hub and a lightweight version is available on the web .
Outcome: The proposed system outperforms or is competitive with the (monolingual) state-of-the-art . it can be used to build knowledge graphs with large ontologies and across multiple languages .
MadDog: A Web-based System for Acronym Identification and Disambiguation (2021.eacl-demos)

Copied to clipboard

Challenge: Acronyms and abbreviations are the short-form of longer phrases and are frequently used in writing but they can also present challenges for newcomers.
Approach: They propose to develop a web-based acronym identification and disambiguation system which can process acronyms from various domains including scientific, biomedical, and general domains.
Outcome: The proposed system can process acronyms from scientific, biomedical, and general domains.
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)

Copied to clipboard

Challenge: Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures.
Approach: They propose to use graph rewriting to describe rule-based transformations on linguistic structures.
Outcome: The proposed tools can be used to compute rule-based transformations on linguistic structures.
Massive Choice, Ample Tasks (MaChAmp): A Toolkit for Multi-task Learning in NLP (2021.eacl-demos)

Copied to clipboard

Challenge: Multi-task learning (MTL) has become a standard repertoire in natural language processing (NLP) it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters.
Approach: They propose a toolkit for fine-tuning contextualized embeddings in multi-task settings.
Outcome: The proposed toolkit supports a variety of natural language processing tasks . it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters.
SCoT: Sense Clustering over Time: a tool for the analysis of lexical change (2021.eacl-demos)

Copied to clipboard

Challenge: Sense Clustering over Time (SCoT) is a network-based tool for analysing lexical change . it visualises word formation, change, and demise as clusters of similar words . SCoT has been successfully used in a European study on the changing meaning of ‘crisis’.
Approach: They propose a new network-based tool for analysing lexical change using a dynamic network of word similarities.
Outcome: The proposed tool has been successfully used in a European study on the changing meaning of ‘crisis’.
GCM: A Toolkit for Generating Synthetic Code-mixed Text (2021.eacl-demos)

Copied to clipboard

Challenge: Code-mixing is a spoken language phenomenon and is difficult to train in multilingual communities.
Approach: They propose a tool that can automatically generate code-mixed data given parallel data in two languages.
Outcome: The proposed tool can generate code-mixed data in two languages using two linguistic theories.
T2NER: Transformers based Transfer Learning Framework for Named Entity Recognition (2021.eacl-demos)

Copied to clipboard

Challenge: Named entity recognition (NER) is an important task in information extraction due to large variations in entity names and flexibility in how entities are mentioned.
Approach: They propose a Transformers based Transfer Learning framework for Named Entity Recognition (T2NER) that integrates transformer models with the state-of-the-art in NLP and provides a unified platform for transfer learning.
Outcome: The proposed framework bridges the gap between the state-of-the-art in transformer models and the state of the art in NER with deep transformer models.
A New Surprise Measure for Extracting Interesting Relationships between Persons (2021.eacl-demos)

Copied to clipboard

Challenge: Interesting facts are useful information for a variety of important tasks.
Approach: They propose a method that extracts all personal relationships from dependency trees and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner.
Outcome: The proposed method extracts all personal relationships from dependency trees for the texts and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner.
Paladin: an annotation tool based on active and proactive learning (2021.eacl-demos)

Copied to clipboard

Challenge: Existing tools for active learning focus on the active learning algorithms and provide no user interface thus making it difficult to use for the end-users.
Approach: They present an open-source web-based annotation tool for creating high-quality multi-label document-level datasets that integrates active learning and proactive learning.
Outcome: The proposed tool is designed for multi-label annotation, but it can be adapted to other tasks in single-l Label settings.
Story Centaur: Large Language Model Few Shot Learning as a Creative Writing Tool (2021.eacl-demos)

Copied to clipboard

Challenge: Few shot learning with large language models has the potential to give individuals without formal machine learning training access to a wide range of text to text models.
Approach: They propose a user interface for prototyping few shot models and a set of recombinable web components that deploy them.
Outcome: The proposed interface lets writers build their own co-creation tools that further their own artistic directions.
FrameForm: An Open-source Annotation Interface for FrameNet (2021.eacl-demos)

Copied to clipboard

Challenge: FrameNet is a computational lexicography tool that provides in-depth semantic information regarding the argument structure and thematic relations of a predicate.
Approach: They introduce an open-source annotation tool that can be easily modified to accommodate predicate annotations based on Frame Semantics.
Outcome: The proposed tool can be easily modified to answer the annotation needs of a wide range of languages.
OCTIS: Comparing and Optimizing Topic models is Simple! (2021.eacl-demos)

Copied to clipboard

Challenge: Current topic modeling frameworks focus on preprocessing, evaluation, comparison of models and visualization.
Approach: They propose an evaluation framework for Topic Models with optimal hyper-parameters estimated using Bayesian Optimization approach.
Outcome: The proposed framework integrates several state-of-the-art topic models and evaluation metrics.
ELITR Multilingual Live Subtitling: Demo and Strategy (2021.eacl-demos)

Copied to clipboard

Challenge: Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages.
Approach: They propose an automatic speech translation system aimed at live subtitling of conference presentations.
Outcome: The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages.
Breaking Writer’s Block: Low-cost Fine-tuning of Natural Language Generation Models (2021.eacl-demos)

Copied to clipboard

Challenge: Currently, it is standard procedure to fine-tune large pre-trained language models for information extraction tasks, but this is not the case for generation tasks, which relies on a variety of techniques for controlled language generation.
Approach: They propose a system that fine-tunes a natural language generation model for the problem of solving writer’s block.
Outcome: The proposed system obtains excellent results even with a small number of epochs and a total cost of USD 150.
OPUS-CAT: Desktop NMT with CAT integration and local fine-tuning (2021.eacl-demos)

Copied to clipboard

Challenge: Neural machine translation (NMT) has brought about a dramatic increase in the quality of machine translation in the past five years.
Approach: OPUS-CAT is a collection of software which enables translators to use neural machine translation in computer-assisted translation tools without exposing themselves to security and confidentiality risks.
Outcome: OPUS-CAT is a collection of software which enables translators to use neural machine translation in computer-assisted translation tools without exposing themselves to security and confidentiality risks.
Domain Expert Platform for Goal-Oriented Dialog Collection (2021.eacl-demos)

Copied to clipboard

Challenge: a prerequisite for the creation of a goal-oriented neural network dialogue system is a dataset that represents typical dialogue scenarios and includes various semantic annotations.
Approach: They propose a web-based platform for collecting and writing goal-oriented dialogue samples.
Outcome: The proposed platform is language-independent and is currently being used to collect dialogue samples in Latvian .
Which is Better for Deep Learning: Python or MATLAB? Answering Comparative Questions in Natural Language (2021.eacl-demos)

Copied to clipboard

Challenge: Comparative QA is a challenging task since it requires collecting evidence from many different sources.
Approach: They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices.
Outcome: The proposed system can be used in personal assistants, chatbots, and similar NLP devices.
PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text (2021.eacl-demos)

Copied to clipboard

Challenge: Prior punctuation restoration methods have focused on using lexical features, prosodic features or combination of both.
Approach: They propose a multitask modeling approach to restore punctuation in multiple high resource languages using acoustic models and a computational model.
Outcome: The proposed system can restore punctuation in Germanic, Romanic and low resource languages without extensive knowledge of grammar or syntax.
Conversational Agent for Daily Living Assessment Coaching Demo (2021.eacl-demos)

Copied to clipboard

Challenge: Conversational Agent for Daily Living Assessment Coaching (CADLAC) is a multi-modal conversational agent system designed to impersonate “individuals” with various levels of ability in activities of daily living.
Approach: They propose to use a multi-modal conversational agent system to impersonate individuals with various levels of ability in activities of daily living to train assessors how to conduct interviews .
Outcome: The system is implemented on the MindMeld platform for conversational AI and features a bidirectional long short-term memory topic tracker that allows the agent to navigate conversations spanning 18 different ADL domains.
HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing (2021.eacl-demos)

Copied to clipboard

Challenge: Pretrained models have been taking the lead of many natural language processing benchmarks such as GLUE, but energy efficiency in the process of model training and inference becomes a critical bottleneck.
Approach: They propose a multi-task energy efficiency benchmarking platform for responsible natural language processing that compares pretrained models’ energy efficiency from the perspectives of time and cost.
Outcome: The proposed model improves on the fine-tuning efficiency of pretrained models from the perspectives of time and cost.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations