Papers by Yen-Chen Wu
Clipping Loops for Sample-Efficient Dialogue Policy Optimisation (2021.naacl-main)
Copied to clipboard
| Challenge: | In previous work, a large number of human dialogues are required to train dialogue agents. |
| Approach: | They propose loop-clipping policy optimisation to eliminate useless responses by clipping loops from dialogue history and clipping advantage to distinguish useless actions from others. |
| Outcome: | The proposed method achieves 80% success rate on a Cambridge restaurant dialogue system using 260 training dialogues compared to baseline of 2160 dialogues. |