Papers by Jimin Sun
Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated Tasks (2021.eacl-main)
Copied to clipboard
| Challenge: | a large amount of work on cross-lingual transfer learning focused on typological and genealogical similarities between languages. |
| Approach: | They propose three features that capture cross-cultural similarities that manifest in linguistic patterns and quantify distinct aspects of language pragmatics. |
| Outcome: | The proposed features capture cross-cultural similarities manifest in linguistic patterns and quantify aspects of language pragmatics. |
Tools Fail: Detecting Silent Errors in Faulty Tools (2024.emnlp-main)
Copied to clipboard
| Challenge: | a failure in one tool can trigger a cascade of errors, leading to complete task failure. |
| Approach: | They propose a framework for tools more broadly which explores a model’s ability to detect “silent” tool errors and reflect on how to plan. |
| Outcome: | The proposed approach shows that the model can detect "silent" tool errors and plan. |
A Multi-dimensional Evaluation of Tokenizer-free Multilingual Pretrained Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | Recent work on tokenizer-free models shows promising results in cross-lingual transfer . previous work focused on reporting accuracy on a limited set of tasks and data settings . |
| Approach: | They compare tokenizer-free and subword-based models using various dimensions . they find subword models are still the most practical choice in many settings . |
| Outcome: | The proposed model improves cross-lingual transfer and reduces engineering overhead. |
Audio-centric Video Understanding Benchmark without Text Shortcut (2025.emnlp-main)
Copied to clipboard
Yudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Yixuan Li, Peihan Li, Yifan Jiang, Wei Li, Zejun Ma, Chao Zhang
| Challenge: | Recent advances in multimodal large language models (MLLMs) focus on visual abilities, but audio is essential for video understanding. |
| Approach: | They propose an audio-centric video understanding benchmark to evaluate video comprehension capabilities of multimodal LLMs with a particular focus on auditory information. |
| Outcome: | The proposed video understanding benchmarks evaluate video comprehension capabilities of multimodal models with a particular focus on auditory information. |