Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations
Using and comparing Rhetorical Structure Theory parsers with rst-workbench (2021.eacl-demos)
Copied to clipboard
| Challenge: | Rhetorical Structure Theory (RST) parsers are usually only trained on English data . |
| Approach: | rst-workbench is a web-based tool that lets users install and use RST parsers. |
| Outcome: | rst-workbench is a web-based tool that lets users run multiple RST parsers simultaneously. |
SF-QA: Simple and Fair Evaluation Library for Open-domain Question Answering (2021.eacl-demos)
Copied to clipboard
| Challenge: | Open-domain question answering (QA) requires large amounts of resources and is difficult to reproduce results due to complex configurations. |
| Approach: | They propose a simple and fair evaluation framework for open-domain question answering (QA) it modularizes the pipeline open- domain QA system, making it easily accessible . |
| Outcome: | The proposed evaluation framework is publicly available and anyone can contribute to the code and evaluations. |
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)
Copied to clipboard
| Challenge: | a library for low-level processing of brahmic scripts is available for free. |
| Approach: | They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. |
| Outcome: | The proposed library supports low-level processing of ten major south Asian Brahmic scripts. |
CovRelex: A COVID-19 Retrieval System with Relation Extraction (2021.eacl-demos)
Copied to clipboard
| Challenge: | Existing challenges to making the system more practical include dealing with newly created and unknown data, and solving the performance gap when utilizing present data. |
| Approach: | They propose a scientific paper retrieval system targeting entities and relations via relation extraction on COVID-19 scientific papers. |
| Outcome: | The proposed system can be accessed via https://www.jaist.ac.jp/is/labs/nguyen-lab/systems/covrelex/. |
MATILDA - Multi-AnnoTator multi-language InteractiveLight-weight Dialogue Annotator (2021.eacl-demos)
Copied to clipboard
| Challenge: | MATILDA is the first multi-annotator, multi-language dialogue annotation tool . it allows the creation of corpora, the management of users, the annotation of dialogues, the quick adaptation of the user interface to any language and the resolution of interannotation disagreement. |
| Approach: | They propose to use MATILDA to create corpora, manage users, and annotation dialogues. |
| Outcome: | The proposed tool supports the full pipeline for dialogue annotation, and non-technical people can use it. |
AnswerQuest: A System for Generating Question-Answer Items from Multi-Paragraph Documents (2021.eacl-demos)
Copied to clipboard
| Challenge: | Existing systems that generate and answer questions in a question-and-answer format can facilitate reading comprehension. |
| Approach: | They propose a system that integrates question answering and question generation tasks to produce a list of Q&A items for a text. |
| Outcome: | The proposed system generates a catalog of Q&A items for a text. |
T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition (2021.eacl-demos)
Copied to clipboard
| Challenge: | Language model (LM) pretraining has led to consistent improvements in many downstream tasks, including named entity recognition (NER). |
| Approach: | They propose a Python library for NER LM finetuning that facilitates cross-domain and cross-lingual generalization of LMs finetuned on NER. |
| Outcome: | The proposed library outperforms LMs trained on NERs in cross-domain and cross-lingual generalization tests on nine datasets. |
Forum 4.0: An Open-Source User Comment Analysis Framework (2021.eacl-demos)
Copied to clipboard
Marlo Haering, Jakob Smedegaard Andersen, Chris Biemann, Wiebke Loosen, Benjamin Milde, Tim Pietz, Christian Stöcker, Gregor Wiedemann, Olaf Zukunft, Walid Maalej
| Challenge: | Using Forum 4.0, we analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
| Approach: | They introduce an open-source framework to semi-automatically analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
| Outcome: | The proposed framework can analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
SLTEV: Comprehensive Evaluation of Spoken Language Translation (2021.eacl-demos)
Copied to clipboard
| Challenge: | Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades. |
| Approach: | They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations. |
| Outcome: | The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language. |
Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing (2021.eacl-demos)
Copied to clipboard
| Challenge: | Trankit is a lightweight, pre-trained toolkit for multilingual natural language processing. |
| Approach: | They propose a transformer-based toolkit for multilingual natural language processing that trains pipelines over 100 languages and 90 pretrained pipelines for 56 languages. |
| Outcome: | The proposed tool outperforms existing pipelines over sentence segmentation, part-of-speech tagging, morphological feature tabbing, and dependency parsing while maintaining competitive performance over tokenization, multi-word token expansion, and lemmatization over 90 Universal Dependencies treebanks. |
DebIE: A Platform for Implicit and Explicit Debiasing of Word Embedding Spaces (2021.eacl-demos)
Copied to clipboard
| Challenge: | Recent research has shown that distributional word vector spaces often encode stereotypical human biases, such as racism and sexism. |
| Approach: | They propose a platform that measures and mitigates bias in word embeddings by executing two (mutually composable) debiasing models. |
| Outcome: | The proposed platform can measure and mitiga bias in word embeddings. |
A Dashboard for Mitigating the COVID-19 Misinfodemic (2021.eacl-demos)
Copied to clipboard
Zhengyuan Zhu, Kevin Meng, Josue Caraballo, Israa Jaradat, Xiao Shi, Zeyu Zhang, Farahnaz Akrami, Haojin Liao, Fatma Arslan, Damian Jimenez, Mohanmmed Samiul Saeef, Paras Pathak, Chengkai Li
| Challenge: | a new public dashboard aims to understand the impact of the COVID-19 misinfodemic on Twitter . the dashboard uses a curated catalog of COVId-19 related facts and debunks of misinformation . |
| Approach: | They propose a public dashboard that matches tweets with COVID-19 misinformation . they also propose experiments to analyze the spread of misinformation on twitter . |
| Outcome: | The proposed dashboard uses a curated catalog of COVID-19 related facts and debunks misinformation . it shows the most prevalent information from the catalog among Twitter users in user-selected geographic regions . |
EasyTurk: A User-Friendly Interface for High-Quality Linguistic Annotation with Amazon Mechanical Turk (2021.eacl-demos)
Copied to clipboard
| Challenge: | Amazon Mechanical Turk (AMT) is one of the most popular crowd-sourcing platforms, allowing researchers from all over the world to create linguistic datasets quickly and at a relatively low cost. |
| Approach: | They propose to improve the potential of Amazon Mechanical Turk by adding some new features to the tool. |
| Outcome: | The proposed tool improves the performance of Amazon Mechanical Turk by adding new features. |
ASAD: Arabic Social media Analytics and unDerstanding (2021.eacl-demos)
Copied to clipboard
| Challenge: | Currently, there are no publicly available tools for analyzing Arabic social media, such as ADIDA and CAMeL, which are not trained with Twitter data. |
| Approach: | They propose to use Arabic social media analysis and unDerstanding to analyze tweets using a web API and a user interface. |
| Outcome: | The proposed system allows users to determine dialects, sentiment, news category, offensiveness, hate speech, adult content, and spam in Arabic tweets. |
COCO-EX: A Tool for Linking Concepts from Texts to ConceptNet (2021.eacl-demos)
Copied to clipboard
| Challenge: | ConceptNet is a semantic network which contains general commonsense facts about the world, e.g., Birds can fly or Computers are used for sending e-mails. |
| Approach: | They propose a tool for Extracting Concepts from texts and linking them to ConceptNet, using the maximum relational information stored in ConceptNet. |
| Outcome: | The proposed method extracts meaningful concepts from natural language texts and links them to conjunct concept nodes in ConceptNet, utilizing the maximum of relational information stored in the KnowledgeGraph. |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
InterpreT: An Interactive Visualization Tool for Interpreting Transformers (2021.eacl-demos)
Copied to clipboard
Vasudev Lal, Arden Ma, Estelle Aflalo, Phillip Howard, Ana Simoes, Daniel Korat, Oren Pereg, Gadi Singer, Moshe Wasserblat
| Challenge: | Using Transformer-based models for NLU/NLP tasks is a growing interest . but there are many open questions regarding the behavior of these models . |
| Approach: | They present an interactive visualization tool for interpreting Transformer-based models. |
| Outcome: | The tool can track and visualize token embeddings through each layer of a Transformer, highlight distances between certain token embeds, and identify task-related functions of attention heads using new metrics. |
Representing ELMo embeddings as two-dimensional text online (2021.eacl-demos)
Copied to clipboard
| Challenge: | ELMoViz module adds support for contextualized embedding architectures, in particular for token embeddable word models. |
| Approach: | They propose to add a module to the free and open-source WebVectors toolkit which provides lexical hyperlinks to word representations in static embedding models. |
| Outcome: | The ELMoViz module adds support for contextualized embedding architectures, in particular for ELMa models. |
LOME: Large Ontology Multilingual Extraction (2021.eacl-demos)
Copied to clipboard
Patrick Xia, Guanghui Qin, Siddharth Vashishtha, Yunmo Chen, Tongfei Chen, Chandler May, Craig Harman, Kyle Rawlins, Aaron Steven White, Benjamin Van Durme
| Challenge: | LOME is a system for performing multilingual information extraction with large ontologies. |
| Approach: | They propose a system for multilingual information extraction with a framenet parser . LOME is available as a Docker container on Docker Hub and a lightweight version is available on the web . |
| Outcome: | The proposed system outperforms or is competitive with the (monolingual) state-of-the-art . it can be used to build knowledge graphs with large ontologies and across multiple languages . |
MadDog: A Web-based System for Acronym Identification and Disambiguation (2021.eacl-demos)
Copied to clipboard
| Challenge: | Acronyms and abbreviations are the short-form of longer phrases and are frequently used in writing but they can also present challenges for newcomers. |
| Approach: | They propose to develop a web-based acronym identification and disambiguation system which can process acronyms from various domains including scientific, biomedical, and general domains. |
| Outcome: | The proposed system can process acronyms from scientific, biomedical, and general domains. |
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)
Copied to clipboard
| Challenge: | Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures. |
| Approach: | They propose to use graph rewriting to describe rule-based transformations on linguistic structures. |
| Outcome: | The proposed tools can be used to compute rule-based transformations on linguistic structures. |
Massive Choice, Ample Tasks (MaChAmp): A Toolkit for Multi-task Learning in NLP (2021.eacl-demos)
Copied to clipboard
| Challenge: | Multi-task learning (MTL) has become a standard repertoire in natural language processing (NLP) it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters. |
| Approach: | They propose a toolkit for fine-tuning contextualized embeddings in multi-task settings. |
| Outcome: | The proposed toolkit supports a variety of natural language processing tasks . it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters. |
SCoT: Sense Clustering over Time: a tool for the analysis of lexical change (2021.eacl-demos)
Copied to clipboard
| Challenge: | Sense Clustering over Time (SCoT) is a network-based tool for analysing lexical change . it visualises word formation, change, and demise as clusters of similar words . SCoT has been successfully used in a European study on the changing meaning of ‘crisis’. |
| Approach: | They propose a new network-based tool for analysing lexical change using a dynamic network of word similarities. |
| Outcome: | The proposed tool has been successfully used in a European study on the changing meaning of ‘crisis’. |
GCM: A Toolkit for Generating Synthetic Code-mixed Text (2021.eacl-demos)
Copied to clipboard
| Challenge: | Code-mixing is a spoken language phenomenon and is difficult to train in multilingual communities. |
| Approach: | They propose a tool that can automatically generate code-mixed data given parallel data in two languages. |
| Outcome: | The proposed tool can generate code-mixed data in two languages using two linguistic theories. |
T2NER: Transformers based Transfer Learning Framework for Named Entity Recognition (2021.eacl-demos)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is an important task in information extraction due to large variations in entity names and flexibility in how entities are mentioned. |
| Approach: | They propose a Transformers based Transfer Learning framework for Named Entity Recognition (T2NER) that integrates transformer models with the state-of-the-art in NLP and provides a unified platform for transfer learning. |
| Outcome: | The proposed framework bridges the gap between the state-of-the-art in transformer models and the state of the art in NER with deep transformer models. |
European Language Grid: A Joint Platform for the European Language Technology Community (2021.eacl-demos)
Copied to clipboard
Georg Rehm, Stelios Piperidis, Kalina Bontcheva, Jan Hajic, Victoria Arranz, Andrejs Vasiļjevs, Gerhard Backfried, Jose Manuel Gomez-Perez, Ulrich Germann, Rémi Calizzano, Nils Feldhus, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Julian Moreno-Schneider, Dimitris Galanis, Penny Labropoulou, Miltos Deligiannis, Katerina Gkirtzou, Athanasia Kolovou, Dimitris Gkoumas, Leon Voukoutis, Ian Roberts, Jana Hamrlova, Dusan Varis, Lukas Kacena, Khalid Choukri, Valérie Mapelli, Mickaël Rigault, Julija Melnika, Miro Janosik, Katja Prinz, Andres Garcia-Silva, Cristian Berrio, Ondrej Klejch, Steve Renals
| Challenge: | Europe is a multilingual society, in which dozens of languages are spoken. |
| Approach: | They describe the European Language Grid, which is targeted to evolve into the primary platform and marketplace for LT in Europe by providing one umbrella platform for the European LT landscape. |
| Outcome: | The European Language Grid (ELG) will provide access to 1300 services for all European languages as well as thousands of data sets. |
A New Surprise Measure for Extracting Interesting Relationships between Persons (2021.eacl-demos)
Copied to clipboard
| Challenge: | Interesting facts are useful information for a variety of important tasks. |
| Approach: | They propose a method that extracts all personal relationships from dependency trees and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner. |
| Outcome: | The proposed method extracts all personal relationships from dependency trees for the texts and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner. |
Paladin: an annotation tool based on active and proactive learning (2021.eacl-demos)
Copied to clipboard
| Challenge: | Existing tools for active learning focus on the active learning algorithms and provide no user interface thus making it difficult to use for the end-users. |
| Approach: | They present an open-source web-based annotation tool for creating high-quality multi-label document-level datasets that integrates active learning and proactive learning. |
| Outcome: | The proposed tool is designed for multi-label annotation, but it can be adapted to other tasks in single-l Label settings. |
Story Centaur: Large Language Model Few Shot Learning as a Creative Writing Tool (2021.eacl-demos)
Copied to clipboard
| Challenge: | Few shot learning with large language models has the potential to give individuals without formal machine learning training access to a wide range of text to text models. |
| Approach: | They propose a user interface for prototyping few shot models and a set of recombinable web components that deploy them. |
| Outcome: | The proposed interface lets writers build their own co-creation tools that further their own artistic directions. |
FrameForm: An Open-source Annotation Interface for FrameNet (2021.eacl-demos)
Copied to clipboard
| Challenge: | FrameNet is a computational lexicography tool that provides in-depth semantic information regarding the argument structure and thematic relations of a predicate. |
| Approach: | They introduce an open-source annotation tool that can be easily modified to accommodate predicate annotations based on Frame Semantics. |
| Outcome: | The proposed tool can be easily modified to answer the annotation needs of a wide range of languages. |
OCTIS: Comparing and Optimizing Topic models is Simple! (2021.eacl-demos)
Copied to clipboard
| Challenge: | Current topic modeling frameworks focus on preprocessing, evaluation, comparison of models and visualization. |
| Approach: | They propose an evaluation framework for Topic Models with optimal hyper-parameters estimated using Bayesian Optimization approach. |
| Outcome: | The proposed framework integrates several state-of-the-art topic models and evaluation metrics. |
ELITR Multilingual Live Subtitling: Demo and Strategy (2021.eacl-demos)
Copied to clipboard
Ondřej Bojar, Dominik Macháček, Sangeet Sagar, Otakar Smrž, Jonáš Kratochvíl, Peter Polák, Ebrahim Ansari, Mohammad Mahmoudi, Rishu Kumar, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai-Son Nguyen, Felix Schneider, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams
| Challenge: | Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages. |
| Approach: | They propose an automatic speech translation system aimed at live subtitling of conference presentations. |
| Outcome: | The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages. |
Breaking Writer’s Block: Low-cost Fine-tuning of Natural Language Generation Models (2021.eacl-demos)
Copied to clipboard
| Challenge: | Currently, it is standard procedure to fine-tune large pre-trained language models for information extraction tasks, but this is not the case for generation tasks, which relies on a variety of techniques for controlled language generation. |
| Approach: | They propose a system that fine-tunes a natural language generation model for the problem of solving writer’s block. |
| Outcome: | The proposed system obtains excellent results even with a small number of epochs and a total cost of USD 150. |
OPUS-CAT: Desktop NMT with CAT integration and local fine-tuning (2021.eacl-demos)
Copied to clipboard
| Challenge: | Neural machine translation (NMT) has brought about a dramatic increase in the quality of machine translation in the past five years. |
| Approach: | OPUS-CAT is a collection of software which enables translators to use neural machine translation in computer-assisted translation tools without exposing themselves to security and confidentiality risks. |
| Outcome: | OPUS-CAT is a collection of software which enables translators to use neural machine translation in computer-assisted translation tools without exposing themselves to security and confidentiality risks. |
Domain Expert Platform for Goal-Oriented Dialog Collection (2021.eacl-demos)
Copied to clipboard
| Challenge: | a prerequisite for the creation of a goal-oriented neural network dialogue system is a dataset that represents typical dialogue scenarios and includes various semantic annotations. |
| Approach: | They propose a web-based platform for collecting and writing goal-oriented dialogue samples. |
| Outcome: | The proposed platform is language-independent and is currently being used to collect dialogue samples in Latvian . |
Which is Better for Deep Learning: Python or MATLAB? Answering Comparative Questions in Natural Language (2021.eacl-demos)
Copied to clipboard
Viktoriia Chekalina, Alexander Bondarenko, Chris Biemann, Meriem Beloucif, Varvara Logacheva, Alexander Panchenko
| Challenge: | Comparative QA is a challenging task since it requires collecting evidence from many different sources. |
| Approach: | They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices. |
| Outcome: | The proposed system can be used in personal assistants, chatbots, and similar NLP devices. |
PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text (2021.eacl-demos)
Copied to clipboard
| Challenge: | Prior punctuation restoration methods have focused on using lexical features, prosodic features or combination of both. |
| Approach: | They propose a multitask modeling approach to restore punctuation in multiple high resource languages using acoustic models and a computational model. |
| Outcome: | The proposed system can restore punctuation in Germanic, Romanic and low resource languages without extensive knowledge of grammar or syntax. |
Conversational Agent for Daily Living Assessment Coaching Demo (2021.eacl-demos)
Copied to clipboard
| Challenge: | Conversational Agent for Daily Living Assessment Coaching (CADLAC) is a multi-modal conversational agent system designed to impersonate “individuals” with various levels of ability in activities of daily living. |
| Approach: | They propose to use a multi-modal conversational agent system to impersonate individuals with various levels of ability in activities of daily living to train assessors how to conduct interviews . |
| Outcome: | The system is implemented on the MindMeld platform for conversational AI and features a bidirectional long short-term memory topic tracker that allows the agent to navigate conversations spanning 18 different ADL domains. |
HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing (2021.eacl-demos)
Copied to clipboard
| Challenge: | Pretrained models have been taking the lead of many natural language processing benchmarks such as GLUE, but energy efficiency in the process of model training and inference becomes a critical bottleneck. |
| Approach: | They propose a multi-task energy efficiency benchmarking platform for responsible natural language processing that compares pretrained models’ energy efficiency from the perspectives of time and cost. |
| Outcome: | The proposed model improves on the fine-tuning efficiency of pretrained models from the perspectives of time and cost. |