Papers by Xingyuan Xian
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study (2025.coling-main)
Copied to clipboard
| Challenge: | a new benchmark evaluates video-based optical character recognition (Video OCR) performance of multi-modal models in videos . the benchmark aims to improve video LLMs' ability to extract text from video content . previous benchmarks have focused on video QA, but not video-related QA. |
| Approach: | They propose to evaluate the video OCR performance of multi-modal models in videos . they use a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement . |
| Outcome: | The proposed benchmark includes 1,028 videos and 2,961 question-answer pairs . it integrates the OCR ability of image LLMs with manual refinement . |