Challenge: Structured prediction models often involve complex inference problems for which finding exact solutions is intractable.
Approach: They propose to perform gradient descent with respect to the output structure directly and train a neural network to perform inference.
Outcome: The proposed methods achieve better speed/accuracy/search error trade-off than gradient descent while being faster than exact inference at similar accuracy levels.

Similar Papers

An Exploration of Arbitrary-Order Sequence Labeling via Energy-Based Inference Networks (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work shows that conditional random fields (CRFs) perform well in sequence labeling tasks.
Approach: They propose several high-order energy terms to capture dependencies among labels in sequence labeling . they use convolutional, recurrent, and self-attention networks to construct these energy terms .
Outcome: The proposed approach improves on four sequence labeling tasks while having the same decoding speed as simple classifiers.
Classical Structured Prediction Losses for Sequence to Sequence Learning (N18-1)

Copied to clipboard

Challenge: Recent work on training neural attention models at the sequence level has focused on a series of objective functions commonly used for structured prediction.
Approach: They propose to use objective functions commonly used to train linear models for structured prediction to train neural attention models at the sequence-level using either reinforcement learning-style methods or beam search optimization.
Outcome: The proposed model outperforms beam search optimization on German-English translation and abstractive summarization tasks.
Training Structured Prediction Energy Networks with Indirect Supervision (N18-2)

Copied to clipboard

Challenge: a new rank-based training method for structured prediction energy networks is proposed . structured prediction is important in many domains, including computer vision, computational biology and natural language processing.
Approach: They propose a rank-based training method for structured prediction energy networks . they use a scoring function defined with domain knowledge to train the models .
Outcome: The proposed method minimizes ranking violation of the sampled structures with respect to a scalar scoring function defined with domain knowledge.
Towards a Better Understanding of Label Smoothing in Neural Machine Translation (2020.aacl-main)

Copied to clipboard

Challenge: In recent years, Neural Network (NN) models bring steady and concrete improvements on the task of Machine Translation (MT).
Approach: They propose to penalize over-confident outputs and regularize the model so that its outputs do not diverge too much from some prior distribution.
Outcome: The proposed method is well-motivated and can improve the performance of strong neural machine translation systems.
Randomized Deep Structured Prediction for Discourse-Level Processing (2021.eacl-main)

Copied to clipboard

Challenge: Expressive text encoders have been at the center of recent NLP work . however, some tasks require complex structural dependencies between texts .
Approach: They propose to leverage deep structured prediction and expressive neural encoders for argumentation mining tasks.
Outcome: The proposed framework can be used for argumentation mining tasks without expensive inference tools.
Efficient, Uncertainty-based Moderation of Neural Networks Text Classifiers (2022.findings-acl)

Copied to clipboard

Challenge: A series of benchmarking experiments based on three different datasets and three state-of-the-art classifiers show that our framework can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%)
Approach: They propose a semi-automated approach that passes unconfident, probably incorrect classifications to human moderators to minimize the workload.
Outcome: The proposed approach can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%) while reducing the moderation load up to 73.3% compared to a random moderation.
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast—Choose Three (2020.emnlp-main)

Copied to clipboard

Challenge: Modern neural networks do not always produce wellcalibrated predictions . post-hoc calibration methods require a held-out calibration dataset, which may not be available in all circumstances.
Approach: They validate ensemble distillation framework for producing well-calibrated structured prediction models without the prohibitive inference-time cost of ensembles.
Outcome: The proposed framework produces well-calibrated predictions without the prohibitive inference-time cost of ensembles.
Calibrating Structured Output Predictors for Natural Language Processing (2020.acl-main)

Copied to clipboard

Challenge: Several modern machine-learning based NLP systems can provide a confidence score with their output predictions.
Approach: They propose a general calibration scheme for output entities of interest in NLP applications that can be used to calibrate confidence scores.
Outcome: The proposed calibration scheme outperforms current calibration techniques for Named Entity Recognition, Part-of-speech tagging and Question Answering systems.
An Empirical Investigation of Structured Output Modeling for Graph-based Neural Dependency Parsing (P19-1)

Copied to clipboard

Challenge: In the past few years, graph-based dependency parsers have led to impressive empirical successes on parsing accuracy.
Approach: They propose to use a graph-based dependency parser to model global outputs.
Outcome: The proposed model has been shown to perform better on sentence-level Complete Match metric compared with the previous model.
Rethinking Complex Neural Network Architectures for Document Classification (N19-1)

Copied to clipboard

Challenge: Neural network models for many NLP tasks have grown increasingly complex in recent years . authors of recent papers question the necessity of such architectures and find them quite effective .
Approach: They propose to use regularization techniques borrowed from language modeling to improve model accuracy . they find that a simple biLSTM architecture with appropriate regularization yields competitive results .
Outcome: a simple biLSTM model outperforms the state-of-the-art on four benchmark datasets . authors say that improvements are not real, but are attributed to mundane reasons .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations