Papers by Jinglu Hu
Improving Image Captioning Evaluation by Considering Inter References Variance (2020.acl-main)
Copied to clipboard
| Challenge: | Existing one-to-one metrics penalize mismatches without considering the intrinsic variance between ground truth captions. |
| Approach: | They propose a one-to-one metric based on BERTScore that could be extended to include new features for image captioning evaluation. |
| Outcome: | The proposed metric achieves state-of-the-art human judgment correlation while improving performance. |
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods focus on single-language scenarios, overlooking multilingual and cross-lingual contexts. |
| Approach: | They propose a tool to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. |
| Outcome: | MaXIFE evaluates instruction-following capabilities across 23 languages with 1667 verifiable instruction tasks. |