Papers by Nick Haber
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images (2025.acl-long)
Copied to clipboard
| Challenge: | Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations. |
| Approach: | They propose a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens. |
| Outcome: | The proposed method achieves up to 22% reduction in hallucinations and significant gains in vision-centric and general tasks while maintaining or improving the model's general abilities. |
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology (2025.emnlp-main)
Copied to clipboard
| Challenge: | State-of-the-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored due to a limited understanding of their performance in high-stakes clinical settings. |
| Approach: | They propose a taxonomy of real-world use cases of multimodal language models in speech-language pathologies to address this gap. |
| Outcome: | The proposed model outperforms 15 state-of-the-art models in speech-language pathologies across five use cases and achieves improvements of over 30% on domain-specific data. |
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing (2026.eacl-long)
Copied to clipboard
| Challenge: | a single prompt can inspire countless valid stories, making objective verification impossible. |
| Approach: | They propose a large-scale benchmark for creative writing evaluation using a reddit corpus and a 2,480-pair test set. |
| Outcome: | The proposed model outperforms existing OTS judges and generative reward models in the evaluation of creative writing. |
Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency (2023.emnlp-main)
Copied to clipboard
| Challenge: | Developing an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses. |
| Approach: | They propose to fine-tune large language models to simulate how previous students would have responded to unseen items to generate high-quality parallel tests. |
| Outcome: | The proposed test forms are designed to be content-equivalent and produce identical individual scores as the original test form. |