Papers by Junyi Du
Learning to Imagine: Visually-Augmented Natural Language Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for natural language generation are pre-trained on text-only corpora, resulting in visual commonsense. |
| Approach: | They propose a method that makes pre-trained language models learn to imagine for visually-augmented natural language generation. |
| Outcome: | The proposed method is compatible with Transformer-based architecture. |
Eliciting Knowledge from Experts: Automatic Transcript Parsing for Cognitive Task Analysis (P19-1)
Copied to clipboard
| Challenge: | Cognitive task analysis (CTA) is a type of analysis used to elicit and represent the knowledge and thought processes of domain experts. |
| Approach: | They propose a weakly-supervised framework for automated CTA transcript parsing . they partition the parser process into a sequence labeling task and a text span-pair relation extraction task with distant supervision from human-curated protocol files. |
| Outcome: | The proposed framework reduces human labor and scales the task to a small scale. |
Zero-shot Visual Question Answering with Language Model Feedback (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for knowledge-based visual question answering are based on pre-trained language models. |
| Approach: | They propose a language model guided captioning approach that leverages a pre-trained language model to generate captions for an image to help answer a visual question. |
| Outcome: | The proposed method outperforms several competing methods on the knowledge-based VQA task and achieves comparable results to a fine-tuned VLP model. |
Pre3: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation (2025.acl-long)
Copied to clipboard
Junyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu, Chuheng Du, Hailong Yang, Ruihao Gong, Shengzhong Liu, Fan Wu, Guihai Chen
| Challenge: | Existing methods for structured generation of outputs are inefficient under large inference batches. |
| Approach: | They propose a new LLM-based method that parses LR(1) grammars into a pushdown automaton and exploits deterministic pushdown automation to optimize the constrained LLM decoding efficiency. |
| Outcome: | The proposed method improves time per output token (TPOT) by 40% and throughput by 36% . |
Visually Grounded Continual Learning of Compositional Phrases (2020.emnlp-main)
Copied to clipboard
| Challenge: | Modern NLP systems rely on offline training and are inefficient for new tasks. |
| Approach: | They propose a visually grounded ContinuaL learning task which simulates the continual acquisition of compositional phrases from streaming visual scenes. |
| Outcome: | The proposed system improves on existing systems, but it's infeasible to store all possible compositions. |