Papers by Kay Rottmann

4 papers
Unsupervised training data re-weighting for natural language understanding with local distribution approximation (2022.emnlp-industry)

Copied to clipboard

Challenge: a distribution mismatch between offline training and live data can cause biases . cyclic seasonality shifts, and changing pool of users can contribute to this problem .
Approach: They propose an unsupervised approach to mitigate offline training data sampling bias . they propose a local distribution approximation in the pre-trained embedding space .
Outcome: The proposed approach mitigates the offline training data sampling bias in multiple NLU tasks without additional annotation.
Semi-supervised Adversarial Text Generation based on Seq2Seq models (2022.emnlp-industry)

Copied to clipboard

Challenge: In contrast, adversarial training has been used in computer vision to improve models’ robustness due to the discrete nature of text.
Approach: They propose a way to generate adversarial samples by using pseudo-labeled in-domain text data to train a seq2seq model for adversarials and combine it with paraphrase detection.
Outcome: The proposed model generates realistic and relevant adversarial samples compared to other state-of-the-art models and recovers up to 70% of errors.
MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages (2023.acl-long)

Copied to clipboard

Challenge: We present the MASSIVE dataset–Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant evaluation.
Approach: They present a 1M-example dataset of Amazon Slu utterances . they localize the dataset into 50 typologically diverse languages .
Outcome: The proposed model includes exact match accuracy, intent classification accuracy, and slot-filling F1 score.
Mitigating the Burden of Redundant Datasets via Batch-Wise Unique Samples and Frequency-Aware Losses (2023.acl-industry)

Copied to clipboard

Challenge: Existing solutions to train deep learning models on redundant datasets are difficult to implement in industrial settings.
Approach: They propose a method to eliminate duplicates at the batch level without altering the data distribution observed by the model.
Outcome: The proposed approach reduces training times on models on redundant datasets by up to 87% and 46% on average, with a drop in model performance of 0.2% relative at worst.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations