Papers by Haixu Chen
LLMs as Lab Engineers: A Benchmark for Analytical Method Lifecycle Management (2026.findings-acl)
Copied to clipboard
Xiaoyi Chen, Mahsa Monshizadeh, Chaoqi Zhang, Jianjun Lang, Yang Wu, Genevieve Mortensen, Xiaozhong Liu, Haixu Tang
| Challenge: | General-purpose commercial models outperform domain-specialized ones, while RAG and reasoning significantly improve performance. |
| Approach: | They propose a benchmark to evaluate LLMs' capabilities in analytical chemistry scenarios. |
| Outcome: | The proposed framework outperforms existing benchmarks focused on factual knowledge and provides practical guidance for analytical chemistry challenges. |
Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR (2026.acl-long)
Copied to clipboard
| Challenge: | Low-resource multilingual OCR models struggle with complex script structures and data scarcity. |
| Approach: | They propose a framework for multilingual character recognition that integrates visual and linguistic backbones with a novel glyph-aware interface. |
| Outcome: | The proposed framework improves on high-resolution visual and language backbones with glyph-aware interface. |
Hey, That’s My Data! Token-Only Dataset Inference in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing dataset inference methods require logit access, but many modern LLMs restrict such access. |
| Approach: | They propose a token-only dataset inference framework that allows models to overwrite prior knowledge when trained on new data. |
| Outcome: | The proposed framework overwrites prior knowledge when trained on new data. |