Papers by Pengda Qin
Robust Distant Supervision Relation Extraction via Deep Reinforcement Learning (P18-1)
Copied to clipboard
| Challenge: | Distant supervision is an efficient method for relation extraction, but it is noisy. |
| Approach: | They propose a deep reinforcement learning strategy to generate false-positive indicators . they redistribute false positives into negative examples to reduce false positive problem . |
| Outcome: | The proposed method significantly improves the performance of distant supervision compared to state-of-the-art systems. |
DSGAN: Generative Adversarial Training for Distant Supervision Relation Extraction (P18-1)
Copied to clipboard
| Challenge: | Distant supervision can effectively label data for relation extraction, but suffers from the noise labeling problem. |
| Approach: | They propose a sentence-level true-positive generator to learn a true-negative generator from a fuzzy sentence bag. |
| Outcome: | The proposed method significantly improves the performance of distant supervision relation extraction compared to state-of-the-art systems. |
Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention (P19-1)
Copied to clipboard
| Challenge: | Existing models for limited-domain RNNs are difficult to scale due to the complexity of the inputs. |
| Approach: | They propose to use dialog acts to build a multi-layer hierarchical graph with a disentangled self-attention network. |
| Outcome: | The proposed model improves on the baselines on automatic and human evaluation metrics. |
Deep Reinforcement Learning with Distributional Semantic Rewards for Abstractive Summarization (D19-1)
Copied to clipboard
| Challenge: | Abstractive summarization tasks are often based on deep reinforcement learning (RL) but the traditional reward system Rouge-L simply looks for exact n-grams matches between candidates and annotated references, which makes the generated sentences repetitive and incoherent. |
| Approach: | They propose to use distributional semantics to measure matching degrees instead of Rouge-L to generate sentences with n-grams matches. |
| Outcome: | The proposed reward has superiority over the existing reward, despite the incoherence of the generated sentences. |