Papers by Fatemeh Tavakkoli
PerPaDa: A Persian Paraphrase Dataset based on Implicit Crowdsourcing Data Collection (2022.lrec-1)
Copied to clipboard
| Challenge: | In this paper, we present a dataset that is collected from users’ input in a plagiarism detection system. |
| Approach: | They propose to use a Persian paraphrase dataset that is collected from users’ input in a plagiarism detection system to improve the quality of the data. |
| Outcome: | The proposed dataset contains 2446 instances of paraphrasing. |