Papers by Patrick Huber
CCQA: A New Web-Scale Question Answering Dataset for Model Pre-Training (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing approaches to answer open domain questions rely on unlabeled text or synthetically generated question-answer pairs. |
| Approach: | They propose a large-scale open-domain question-answering dataset based on the Common Crawl project that can be used to in-domain pre-train popular language models. |
| Outcome: | The proposed dataset achieves promising results in zero-shot, low resource and fine-tuned settings across multiple tasks, models and benchmarks. |
MEGA RST Discourse Treebanks with Structure and Nuclearity from Scalable Distant Sentiment Supervision (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing discourse treebanks are limited in the application of data-driven approaches to discourse parsing. |
| Approach: | They propose a method to automatically generate discourse treebanks using distant supervision from sentiment annotated datasets by heuristic beam-search strategy extended with a stochastic component. |
| Outcome: | The proposed method generates discourse trees incorporating structure and nuclearity for documents of arbitrary length using an efficient beam-search strategy, extended with a stochastic component. |
From Sentiment Annotations to Sentiment Prediction through Discourse Augmentation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing sentiment analysis models lack temporal information to capture semantics of long texts. |
| Approach: | They propose a framework to exploit task-related discourse structures for sentiment analysis. |
| Outcome: | The proposed framework improves the performance even beyond existing approaches based on human annotated data. |
Small But Funny: A Feedback-Driven Approach to Humor Distillation (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been used to transfer knowledge from LLMs to smaller, smaller language models (SLMs). |
| Approach: | They propose to assign a dual role to the LLM as a “teacher” generating data, as well as evaluating the student’s performance. |
| Outcome: | The proposed approach narrows the performance gap between LLMs and larger models by incorporating feedback into the data. |
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling (2024.acl-long)
Copied to clipboard
Zekun Li, Zhiyu Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Dong, Adithya Sagar, Xifeng Yan, Paul Crook
| Challenge: | Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. |
| Approach: | They propose a method for solving dialogue state tracking (DST) with large language models through function calling. |
| Outcome: | The proposed approach improves zero-shot DST, allowing adaptation to diverse domains without extensive data collection or model tuning. |
Predicting Discourse Structure using Distant Supervision from Sentiment (D19-1)
Copied to clipboard
| Challenge: | Discourse parsing is a fundamental NLP task known to enhance key downstream tasks, such as sentiment analysis, text classification and summarization. |
| Approach: | They propose a method that uses document supervision to generate abundant data for RST-style discourse structure prediction by using an optimal CKY-style tree generation algorithm. |
| Outcome: | The proposed approach performs well on the more difficult task of inter-domain discourse structure prediction, but it does not match the performance of a parser trained and tested on the same dataset. |
Discourse Structure Extraction from Pre-Trained and Fine-Tuned Language Models in Dialogues (2023.findings-eacl)
Copied to clipboard
| Challenge: | Discourse processing suffers from data sparsity, especially for dialogues . a variety of discourse frameworks have been proposed to extract discourse information from dialogues. |
| Approach: | They propose unsupervised and semi-supervised methods to infer latent discourse structures for dialogues based on attention matrices from Pre-trained Language Models. |
| Outcome: | The proposed methods achieve encouraging results on the STAC corpus, with F1 scores of 57.2 and 59.3 for the unsupervised and semi-supervised methods, respectively. |
Scaling Parameter-Constrained Language Models with Quality Data (2024.emnlp-industry)
Copied to clipboard
Ernie Chang, Matteo Paltenghi, Yang Li, Pin-Jie Lin, Changsheng Zhao, Patrick Huber, Zechun Liu, Rastislav Rabatin, Yangyang Shi, Vikas Chandra
| Challenge: | Scaling laws in language modeling quantify training loss as a function of dataset size and model parameters, but neglect the critical role of data quality in model generalization. |
| Approach: | They propose to use effective training tokens as a combination of text diversity and syntheticity as measured by a teacher model to calculate scaling laws. |
| Outcome: | The proposed term effective training tokens is a combination of two readily-computed indicators of text diversity and syntheticity as measured by a teacher model. |
Unleashing the Power of Neural Discourse Parsers - A Context and Structure Aware Approach Using Large Scale Pretraining (2020.coling-main)
Copied to clipboard
| Challenge: | Discourse parsing is an important upstream task within the area of Natural Language Processing (NLP) . |
| Approach: | They propose a discourse parser that incorporates recent contextual language models to improve the performance of RST-based discourse parses. |
| Outcome: | The proposed parser outperforms existing models on two key RST datasets and on large-scale "silver-standard" discourse treebank MEGA-DT. |
Predicting Discourse Trees from Transformer-based Neural Summarizers (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing extractive summarization tasks use only neural approaches to learn discourse information, but recent work has shown that it is beneficial for summarizing discourse information. |
| Approach: | They propose to generate document-level discourse trees from pre-trained neural summarizers that encode dependency- and constituency-style discourse information. |
| Outcome: | The proposed model learns both, dependency- and constituency-style discourse information, consistent with pre-neural results. |
W-RST: Towards a Weighted RST-style Discourse Framework (2021.acl-long)
Copied to clipboard
| Challenge: | We show that weighted discourse trees from auxiliary tasks can benefit downstream applications . linguistic theories play a less and less critical role in the field of discourse . |
| Approach: | They propose a weighted-RST framework that assigns a binary assessment of importance between text segments by a relation attribute. |
| Outcome: | The proposed framework can be replaced by real-valued scores, the authors show . they show that weighted discourse trees can benefit key NLP downstream applications . |
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers (2025.acl-long)
Copied to clipboard
| Challenge: | In language model training, it is difficult to obtain the right data mixtures for various tasks as the relationship between data and tasks is difficult. |
| Approach: | They propose to identify checkpoint models based on their respective capabilities and leverage them as data mixers by using their aggregated first-order influence approximation over source data. |
| Outcome: | The proposed framework shows significant improvements on eight reasoning benchmarks, with accuracy increases of up to 1.93%. |
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment (2026.acl-industry)
Copied to clipboard
Hanxian Huang, Igor Fedorov, Andrey Gromov, Bernard Beckerman, Naveen Suda, David Eriksson, Maximilian Balandat, Rylan Conway, Patrick Huber, Chinnadhurai Sankar, Ayushi Dalmia, Zechun Liu, Lemeng Wu, Tarek Elgamal, Adithya Sagar, Vikas Chandra, Raghuraman Krishnamoorthi
| Challenge: | MobileLLM-Flash is a family of foundation models for efficient on-device use with strong capabilities. |
| Approach: | They propose a method for designing on-device large language models under mobile latency constraints using hardware-in-the-loop architecture search. |
| Outcome: | The proposed model is amenable to industry-scale deployment and is compatible with mobile runtimes like Executorch. |
Towards Understanding Large-Scale Discourse Structures in Pre-Trained and Fine-Tuned Language Models (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing approaches to pre-training/fine-tuning are focusing on the alignment of pre-trained and fine-tuned PLMs with large-scale discourse structures. |
| Approach: | They propose a novel approach to infer discourse information for arbitrarily long documents using supervised, distantly supervised and simple baselines. |
| Outcome: | The proposed approach shows that the captured discourse information is local and general, even across fine-tuning tasks. |
Automated Evaluation of Out-of-Context Errors (L18-1)
Copied to clipboard
| Challenge: | Existing methods to modify text understanding systems use only one sentence at a time . however, considering a larger context can improve performance for text understanding tasks. |
| Approach: | They propose to modify existing text data to insert out-of-context errors . they use a 2016 TEDTalk corpus to evaluate computational models for text understanding . |
| Outcome: | The proposed method targets real-world problems of transcription and translation systems by inserting authentic out-of-context errors. |