Papers by Meizhu Liu
Do Image–Text Metrics Respect Semantic Invariances? (2026.findings-acl)
Copied to clipboard
Amit Agarwal, Hitesh Laxmichand Patel, Meizhu Liu, Jyotika Singh, Karan Dua, Hansa Meghwani, Matthew Rowe, M. Avendi, Yassi Abbasi, Tao Sheng, Sujith Ravi, Dan Roth
| Challenge: | Reference-free image–to–text evaluators are now standard for scoring image–caption alignment, yet it is unclear whether they respect semantic invariances. |
| Approach: | They propose an invariance probe on five popular evaluators under semantics-preserving perturbations along three axes: spatial edits, object changes, and socio-linguistic framing. |
| Outcome: | The proposed invariance probe shows that spatial edits and simple phrasing changes shift scores by ()6% on average and cause ranking flips in up to (),37% of cases. |
No Label? No Problem: Unsupervised Continual Learning for Adaptive Medical ASR (2026.eacl-industry)
Copied to clipboard
| Challenge: | Medical audio often contains specialized terminology, such as medication names, which existing ASR systems struggle to transcribe accurately. |
| Approach: | They propose an unsupervised continual learning ASR framework that adapts to new data while preserving prior knowledge. |
| Outcome: | Experiments on real-world medical audio show that the proposed framework improves over state-of-the-art models. |
Synthetic Doctor-Patient Dialogue Generation for Robust Medical ASR: A Scalable Pipeline for Vocabulary Expansion and Privacy Preservation (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing ASR models struggle with high word error rates (WER) on clinical vocabulary, especially medication names. |
| Approach: | They propose to generate doctor-patient dialogues in both text and audio formats using a curated set of over 124,000 medical terms. |
| Outcome: | The proposed pipeline generated over 1 billion audios with ground truth transcriptions. |