Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

29 papers
SyntaViz: Visualizing Voice Queries through a Syntax-Driven Hierarchical Ontology (D18-2)

Copied to clipboard

Challenge: SyntaViz provides quick access to high-impact failure points of the existing intent understanding system and evidence for data-driven decisions in the development cycle.
Approach: They propose a visualization interface specifically designed for analyzing natural-language queries created by users of a voice-enabled product.
Outcome: The proposed visualization interface helps developers identify multiple action items in a short amount of time without any special training.
TRANX: A Transition-based Neural Abstract Syntax Parser for Semantic Parsing and Code Generation (D18-2)

Copied to clipboard

Challenge: Existing neural semantic parsers only focus on a small subset of tasks, such as SQL queries, robotic commands, and even general-purpose programming languages like Java.
Approach: They propose a transition-based neural semantic parser that maps natural language utterances into formal meaning representations (MRs) they use an abstract syntax description language to constrain the output space and model the information flow.
Outcome: Experiments on four different semantic parsing and code generation tasks show that the proposed system is generalizable, extensible, and effective.
Data2Text Studio: Automated Text Generation from Structured Data (D18-2)

Copied to clipboard

Challenge: Data2Text Studio is a platform for automated text generation from structured data.
Approach: They conduct experiments on RotoWire datasets for template extraction and text generation . they find that the Semi-HMMs model improves interactivity and interpretability .
Outcome: The proposed model improves on template extraction and text generation tasks on RotoWire datasets.
Term Set Expansion based NLP Architect by Intel AI Lab (D18-2)

Copied to clipboard

Challenge: SetExpander is a corpus-based system for expanding a seed set of terms into a more complete set of words belonging to the same semantic class.
Approach: They propose a corpus-based system for expanding a seed set of terms into a more complete set of words that belong to the same semantic class.
Outcome: The proposed system can expand a seed set of terms into a more complete set of words belonging to the same semantic class.
MorAz: an Open-source Morphological Analyzer for Azerbaijani Turkish (D18-2)

Copied to clipboard

Challenge: MorAz is an open-source morphological analyzer for Azerbaijani Turkish . it provides a number of "readings" or analysis for each word, as a part of the overall NLP task.
Approach: They propose an open-source morphological analyzer for Azerbaijani Turkish . they wrap the analyzer with python scripts and implement it in a Django instance .
Outcome: The analyzer is available as a website and as python script in a Django instance.
An Interactive Web-Interface for Visualizing the Inner Workings of the Question Answering LSTM (D18-2)

Copied to clipboard

Challenge: Existing visualisation methods for deep learning models are limited by their low interpretability and lack a tool for interpreting them.
Approach: They propose a visualisation tool which plots heatmaps of neurons’ firings and allows a user to check the dependency between neurons and manual features.
Outcome: The proposed visualisation tool plots heatmaps of neurons’ firings and allows a user to check the dependency between neurons and manual features.
Visual Interrogation of Attention-Based Models for Natural Language Inference and Machine Comprehension (D18-2)

Copied to clipboard

Challenge: Neural networks models have gained popularity due to their state-of-the-art performance but lack of interpretability hinders their deployment and refinement.
Approach: They propose a visual analytic library that provides a user with a customizable visual anallytic environment.
Outcome: The proposed visualization library provides an interactive environment in which the user can investigate and interrogate the relationships between input, model internals and output predictions.
DERE: A Task and Domain-Independent Slot Filling Framework for Declarative Relation Extraction (D18-2)

Copied to clipboard

Challenge: Comparability of models across tasks is lacking in most machine learning systems for natural language processing.
Approach: They propose a framework for declarative specification and compilation of template-based information extraction that uses a generic specification language for the task and for data annotations in terms of spans and frames.
Outcome: The proposed framework enables representation of a large variety of natural language processing tasks.
Demonstrating Par4Sem - A Semantic Writing Aid with Adaptive Paraphrasing (D18-2)

Copied to clipboard

Challenge: a new tool for semantic writing aids collects training examples from usage data.
Approach: They propose a semantic writing aid tool based on adaptive paraphrasing that integrates into a real word application to collect training examples from usage data.
Outcome: The proposed tool is integrated into a real word application to collect training examples from usage data.
Juman++: A Morphological Analysis Toolkit for Scriptio Continua (D18-2)

Copied to clipboard

Challenge: a morphological analyzer is useful for languages without natural word boundaries, but it is difficult to improve it without creating costly annotations.
Approach: They propose a toolkit for developing morphological analyzers for languages without natural word boundaries using lattices and neural nets.
Outcome: The proposed morphological analyzer of Japanese achieves new SOTA on Jumandic-based corpora while being 250 times faster than the previous one.
Visualization of the Topic Space of Argument Search Results in args.me (D18-2)

Copied to clipboard

Challenge: args.me is the first search engine for controversial topics . it ranks pro and con arguments by their relevance to a topic .
Approach: They propose a visualization interface for result exploration that provides an overview of main aspects in a barycentric coordinate system.
Outcome: The proposed search engine is the first dedicated argument search engine on the web.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing (D18-2)

Copied to clipboard

Challenge: Existing subword segmentation tools assume input is pre-tokenized into word sequences, but SentencePiece can train subword models directly from raw sentences.
Approach: They propose a language-independent subword tokenizer and detokenizer for Neural-based text processing.
Outcome: The proposed system achieves comparable accuracy to training from raw sentences.
CogCompTime: A Tool for Understanding Time in Natural Language (D18-2)

Copied to clipboard

Challenge: Existing systems that extract temporal information from text can be useful for natural language understanding.
Approach: They propose a system that extracts temporal information from text and normalizes it to a standard format.
Outcome: The proposed system achieves state-of-the-art performance and incorporates the most recent progress.
A Multilingual Information Extraction Pipeline for Investigative Journalism (D18-2)

Copied to clipboard

Challenge: a new pipeline is being developed to process large collections of unstructured textual data . the pipeline is a key input processor for the upcoming major release of our software .
Approach: a new pipeline is introduced to extract large amounts of unstructured data . the pipeline is used by journalists to process large files containing unknown contents .
Outcome: the pipeline is an input processor for the upcoming major release of our new/s/leak 2.0 software.
Sisyphus, a Workflow Manager Designed for Machine Translation and Automatic Speech Recognition (D18-2)

Copied to clipboard

Challenge: Sisyphus is a workflow manager for Python that can be used for large and complicated workflows.
Approach: Sisyphus is a Python-based workflow manager that can be used to train and test a machine . it maps all jobs to a unique path and can create links bearing descriptive names.
Outcome: Sisyphus is a Python-based workflow manager that can handle large experiments . it can be used without modification to edit, debug, document the workflow .
KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos (D18-2)

Copied to clipboard

Challenge: KT-Speech-Crawler is an automated dataset building tool for speech recognition.
Approach: They propose an approach for automatic dataset construction for speech recognition by crawling YouTube videos.
Outcome: The proposed algorithm can obtain 150 hours of transcribed speech in a day with an estimated 3.5% word error rate.
Visualizing Group Dynamics based on Multiparty Meeting Understanding (D18-2)

Copied to clipboard

Challenge: During a discussion, participants might adjust their own opinions and tune their attitudes towards others’ opinions based on the unfolding interactions.
Approach: They propose a multi-party meeting opinion mining system that visualizes real-time opinion extraction and group dynamics using bipartite graphs.
Outcome: The proposed system extracts opinions from speech and visualizes an influence factor for each participant using current and previous utterances in real time.
An Interface for Annotating Science Questions (D18-2)

Copied to clipboard

Challenge: a new interface for human annotation of science question-answer pairs with their knowledge and reasoning types is proposed . the interface is based on previous work on the ARC dataset, but does not provide clear definitions of these types of knowledge.
Approach: They propose an interface for human annotation of science question-answer pairs with their respective knowledge and reasoning types.
Outcome: The proposed interface improves the classification of science questions in a preliminary study involving 10 participants.
APLenty: annotation tool for creating high-quality datasets using active and proactive learning (D18-2)

Copied to clipboard

Challenge: APLenty is an annotation tool for creating high-quality sequence labeling datasets using active and proactive learning.
Approach: They present APLenty, an annotation tool for creating high-quality sequence labeling datasets using active and proactive learning.
Outcome: The proposed tool is highly flexible and can be adapted to various other tasks.
Interactive Instance-based Evaluation of Knowledge Base Question Answering (D18-2)

Copied to clipboard

Challenge: Existing approaches to Knowledge Base Question Answering are based on semantic parsing.
Approach: They propose a tool that aids in debugging of question answering systems that construct a structured semantic representation for the input question.
Outcome: The proposed system allows debugging of model predictions on individual instances and simplifies manual error analysis.
Magnitude: A Fast, Efficient Universal Vector Embedding Utility Package (D18-2)

Copied to clipboard

Challenge: Magnitude is an open source Python package that performs common operations up to 6,000 times faster than Gensim.
Approach: They present a Python tool for utilizing vector embeddings that performs common operations up to 6,000 times faster than Gensim.
Outcome: The Magnitude package performs common operations up to 6,000 times faster than Gensim and introduces several novel features for improved robustness like out-of-vocabulary lookups.
Integrating Knowledge-Supported Search into the INCEpTION Annotation Platform (D18-2)

Copied to clipboard

Challenge: Annotating entity mentions and linking them to a knowledge resource are essential tasks in many domains.
Approach: a new tool integrates knowledge-supported search and entity linking into INCEpTION . the tool allows users to search the corpus and create cross-document coreferences .
Outcome: a new tool integrates knowledge-supported search and entity linking into INCEpTION . the tool disambiguates mentions, introduces cross-document coreferences, and provides fast queries.
CytonMT: an Efficient Neural Machine Translation Open-source Toolkit Implemented in C++ (D18-2)

Copied to clipboard

Challenge: Neural machine translation (NMT) has made remarkable progress over the past few years.
Approach: They propose to use C++ and NVIDIA’s GPU-accelerated libraries to build an open-source neural machine translation toolkit called CytonMT.
Outcome: The proposed toolkit accelerates the training speed by 64.5% to 110.8% on neural networks of various sizes, and achieves competitive translation quality.
OpenKE: An Open Toolkit for Knowledge Embedding (D18-2)

Copied to clipboard

Challenge: Existing knowledge embedding tools are available for embeddable knowledge graphs.
Approach: They propose a unified framework and various fundamental models to embed knowledge graphs into a continuous low-dimensional space.
Outcome: The toolkit and pre-trained embeddings are available on http://openke.thunlp.org/.
LIA: A Natural Language Programmable Personal Assistant (D18-2)

Copied to clipboard

Challenge: a prototype of an intelligent personal assistant can be programmed using natural language . a user can instruct her assistants using language similar to how humans teach other humans .
Approach: They present LIA, an intelligent personal assistant that can be programmed using natural language. LIA resides on a typical mobile Android device.
Outcome: The proposed system can be programmed using natural language, and it can perceive the external environment through sensors and effectors.
PizzaPal: Conversational Pizza Ordering using a High-Density Conversational AI Platform (D18-2)

Copied to clipboard

Challenge: a pizza ordering bot that can be used to order pizzas is described in this paper.
Approach: They describe PizzaPal, a voice-only agent for ordering pizza, and the Conversational AI architecture built at b4.ai.
Outcome: The pizza ordering bot is based on a dialog framework developed by b4.ai .
Developing Production-Level Conversational Interfaces with Shallow Semantic Parsing (D18-2)

Copied to clipboard

Challenge: In this demo, we demonstrate an end-to-end approach for building conversational interfaces from prototype to production.
Approach: They propose an end-to-end approach for building conversational interfaces from prototype to production that leverages shallow semantic parsing.
Outcome: The proposed approach has proven to work well for a number of applications across diverse verticals.
When science journalism meets artificial intelligence : An interactive demonstration (D18-2)

Copied to clipboard

Challenge: Existing tools for automating science journalism do not provide adequate training for AIs to be trained.
Approach: They propose an online tool that generates titles of blog titles by mimicking a human science journalist.
Outcome: The proposed tool generates blog titles by mimicking a human science journalist . it is evaluated using standard metrics to show its viability .
Universal Sentence Encoder for English (D18-2)

Copied to clipboard

Challenge: TensorFlow Hub sentence embedding models have good task transfer performance . model variants allow for trade-offs between accuracy and compute resources .
Approach: They propose easy-to-use TensorFlow Hub sentence embedding models with good task transfer performance.
Outcome: The proposed models outperform models without transfer learning and those that use only word-level transfer on a number of NLP tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations