Benchmarking Approximate Inference Methods for Neural Structured Prediction (N19-1)
Copied to clipboard
| Challenge: | Structured prediction models often involve complex inference problems for which finding exact solutions is intractable. |
| Approach: | They propose to perform gradient descent with respect to the output structure directly and train a neural network to perform inference. |
| Outcome: | The proposed methods achieve better speed/accuracy/search error trade-off than gradient descent while being faster than exact inference at similar accuracy levels. |
Similar Papers
An Exploration of Arbitrary-Order Sequence Labeling via Energy-Based Inference Networks (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent work shows that conditional random fields (CRFs) perform well in sequence labeling tasks. |
| Approach: | They propose several high-order energy terms to capture dependencies among labels in sequence labeling . they use convolutional, recurrent, and self-attention networks to construct these energy terms . |
| Outcome: | The proposed approach improves on four sequence labeling tasks while having the same decoding speed as simple classifiers. |
Classical Structured Prediction Losses for Sequence to Sequence Learning (N18-1)
Copied to clipboard
| Challenge: | Recent work on training neural attention models at the sequence level has focused on a series of objective functions commonly used for structured prediction. |
| Approach: | They propose to use objective functions commonly used to train linear models for structured prediction to train neural attention models at the sequence-level using either reinforcement learning-style methods or beam search optimization. |
| Outcome: | The proposed model outperforms beam search optimization on German-English translation and abstractive summarization tasks. |
Training Structured Prediction Energy Networks with Indirect Supervision (N18-2)
Copied to clipboard
| Challenge: | a new rank-based training method for structured prediction energy networks is proposed . structured prediction is important in many domains, including computer vision, computational biology and natural language processing. |
| Approach: | They propose a rank-based training method for structured prediction energy networks . they use a scoring function defined with domain knowledge to train the models . |
| Outcome: | The proposed method minimizes ranking violation of the sampled structures with respect to a scalar scoring function defined with domain knowledge. |
Towards a Better Understanding of Label Smoothing in Neural Machine Translation (2020.aacl-main)
Copied to clipboard
| Challenge: | In recent years, Neural Network (NN) models bring steady and concrete improvements on the task of Machine Translation (MT). |
| Approach: | They propose to penalize over-confident outputs and regularize the model so that its outputs do not diverge too much from some prior distribution. |
| Outcome: | The proposed method is well-motivated and can improve the performance of strong neural machine translation systems. |
Randomized Deep Structured Prediction for Discourse-Level Processing (2021.eacl-main)
Copied to clipboard
| Challenge: | Expressive text encoders have been at the center of recent NLP work . however, some tasks require complex structural dependencies between texts . |
| Approach: | They propose to leverage deep structured prediction and expressive neural encoders for argumentation mining tasks. |
| Outcome: | The proposed framework can be used for argumentation mining tasks without expensive inference tools. |
Efficient, Uncertainty-based Moderation of Neural Networks Text Classifiers (2022.findings-acl)
Copied to clipboard
| Challenge: | A series of benchmarking experiments based on three different datasets and three state-of-the-art classifiers show that our framework can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%) |
| Approach: | They propose a semi-automated approach that passes unconfident, probably incorrect classifications to human moderators to minimize the workload. |
| Outcome: | The proposed approach can improve the classification F1-scores by 5.1 to 11.2% (up to approx. 98 to 99%) while reducing the moderation load up to 73.3% compared to a random moderation. |
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast—Choose Three (2020.emnlp-main)
Copied to clipboard
| Challenge: | Modern neural networks do not always produce wellcalibrated predictions . post-hoc calibration methods require a held-out calibration dataset, which may not be available in all circumstances. |
| Approach: | They validate ensemble distillation framework for producing well-calibrated structured prediction models without the prohibitive inference-time cost of ensembles. |
| Outcome: | The proposed framework produces well-calibrated predictions without the prohibitive inference-time cost of ensembles. |
Calibrating Structured Output Predictors for Natural Language Processing (2020.acl-main)
Copied to clipboard
| Challenge: | Several modern machine-learning based NLP systems can provide a confidence score with their output predictions. |
| Approach: | They propose a general calibration scheme for output entities of interest in NLP applications that can be used to calibrate confidence scores. |
| Outcome: | The proposed calibration scheme outperforms current calibration techniques for Named Entity Recognition, Part-of-speech tagging and Question Answering systems. |
An Empirical Investigation of Structured Output Modeling for Graph-based Neural Dependency Parsing (P19-1)
Copied to clipboard
| Challenge: | In the past few years, graph-based dependency parsers have led to impressive empirical successes on parsing accuracy. |
| Approach: | They propose to use a graph-based dependency parser to model global outputs. |
| Outcome: | The proposed model has been shown to perform better on sentence-level Complete Match metric compared with the previous model. |
Rethinking Complex Neural Network Architectures for Document Classification (N19-1)
Copied to clipboard
| Challenge: | Neural network models for many NLP tasks have grown increasingly complex in recent years . authors of recent papers question the necessity of such architectures and find them quite effective . |
| Approach: | They propose to use regularization techniques borrowed from language modeling to improve model accuracy . they find that a simple biLSTM architecture with appropriate regularization yields competitive results . |
| Outcome: | a simple biLSTM model outperforms the state-of-the-art on four benchmark datasets . authors say that improvements are not real, but are attributed to mundane reasons . |