Papers by Da-Cheng Juan
Neuron-Level Differentiation of Memorization and Generalization in Large Language Models (2025.emnlp-main)
Copied to clipboard
Ko-Wei Huang, Yi-Fu Fu, Ching-Yu Tsai, Yu-Chieh Tu, Tzu-ling Cheng, Cheng-Yu Lin, Yi-Ting Yang, Heng-Yi Liu, Keng-Te Liao, Da-Cheng Juan, Shou-De Lin
| Challenge: | Existing models exhibit memorization and generalization behaviors in ways that are not easily interpretable or controllable. |
| Approach: | They propose to use a GPT-2 and LLaMA-3.2 model to identify distinct neuron subsets responsible for each behavior to steer the model toward memorization or generalization. |
| Outcome: | The proposed models show that inference-time interventions on these neurons can steer the model’s behavior toward memorization or generalization. |
DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback (2025.naacl-long)
Copied to clipboard
Jiao Sun, Deqing Fu, Yushi Hu, Su Wang, Royi Rassin, Da-Cheng Juan, Dana Alon, Charles Herrmann, Sjoerd Van Steenkiste, Ranjay Krishna, Cyrus Rashtchian
| Challenge: | Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text. |
| Approach: | They propose a training algorithm that trains T2I models to be faithful to the input text. |
| Outcome: | The proposed model improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic). |
RARR: Researching and Revising What Language Models Say, Using Language Models (2023.acl-long)
Copied to clipboard
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, Kelvin Guu
| Challenge: | Language models (LMs) excel at many tasks but often produce unsupported or misleading content. |
| Approach: | They propose a system that finds attribution for any text generation model and post-edits it to fix unsupported content. |
| Outcome: | The proposed system improves attribution while preserving the original output. |
Low-Dimensional Hyperbolic Knowledge Graph Embeddings (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for predicting missing facts do not account for hierarchical and logical patterns in KGs. |
| Approach: | They propose a class of hyperbolic KG embedding models that capture hierarchical and logical patterns. |
| Outcome: | Experimental results show that the proposed method improves by 6.1% in mean reciprocal rank in low dimensions over previous methods. |
AirConcierge: Generating Task-Oriented Dialogue via Efficient Large-Scale Knowledge Retrieval (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing neural task-oriented dialogue systems cannot be encoded by memory networks, such as memory networks. |
| Approach: | They propose an end-to-end trainable text-to SQL guided framework to learn a neural agent that interacts with KBs using the generated SQL queries. |
| Outcome: | The proposed method significantly improves on the AirDialogue dataset, which contains the conversations of customers booking flight tickets from the agent. |
On the Robustness of Self-Attentive Models (P19-1)
Copied to clipboard
| Challenge: | Experimental results show that self-attentive neural models are more robust against adversarial perturbations compared to recurrent neural networks. |
| Approach: | They propose an adversarial attack algorithm that generates more natural adversarials . they propose to use the attention mechanism to learn a context-dependent representation . |
| Outcome: | The proposed attack algorithm generates more natural adversarial examples that could mislead models but not humans. |
“Does it Matter When I Think You Are Lying?” Improving Deception Detection by Integrating Interlocutor’s Judgements in Conversations (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for deception detection are based on interrogator's perceptions of truth-bias . despite its frequent occurrences, human is not good at detecting deceptions despite inclination of truth bias . |
| Approach: | They propose a Judgmental-Enhanced Automatic Deception Detection Network that explicitly considers interrogator's perceived truths-deceptions with three types of speechlanguage features extracted during a conversation. |
| Outcome: | The proposed method outperforms the current state-of-the-art approach without conditioning on interrogator's judgements. |
Question Answering with Long Multiple-Span Answers (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing QA systems for question answering are limited by the availability of annotated datasets. |
| Approach: | They propose a dataset for question-answering that extracts information from multiple parts of text . they propose QA-based multi-span neural architecture that captures relevance among multiple answer spans . |
| Outcome: | The proposed model outperforms state-of-the-art QA models in this multi-span QA setting. |
A2N: Attending to Neighbors for Knowledge Graph Inference (P19-1)
Copied to clipboard
| Challenge: | Existing knowledge graph completion methods learn a fixed embedding for every entity, which is suboptimal as it requires memorizing and generalizing to all possible entity relationships. |
| Approach: | They propose a method which learns query-dependent representations of entities by combining relevant neighborhood of an entity. |
| Outcome: | The proposed model performs competitively or better than existing state-of-the-art models for knowledge graph completion. |