Papers by Atabak Ashfaq
Re-evaluating Evaluation in Text Summarization (2020.emnlp-main)
Copied to clipboard
| Challenge: | Automated evaluation metrics are an essential part of the development of text-generation tasks such as summarization. |
| Approach: | They propose to use top-scoring system outputs to assess the reliability of automatic evaluation metrics for text summarization. |
| Outcome: | The proposed evaluation method is based on human judgments from 25 top-scoring neural summarization systems. |
Metrics also Disagree in the Low Scoring Range: Revisiting Summarization Evaluation Metrics (2020.coling-main)
Copied to clipboard
| Challenge: | In text summarization evaluation, evaluating the efficacy of automated metrics without human judgments has become popular. |
| Approach: | They revisit their experiments and find that automatic metrics disagree when ranking high-scoring summaries. |
| Outcome: | The proposed method is a human judgment-free method, but it is not a meta-evaluation strategy. |