Papers by Longfei Yun
The Price of Format: Diversity Collapse in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Instruction-tuned large language models employ structured templates to enforce format consistency during inference. |
| Approach: | They fine-tune instruction-tuning large language models with structured templates and evaluate their results across three axes: downstream task performance, alignment behavior, and output diversity. |
| Outcome: | The proposed model generates semantically similar outputs even under high temperature sampling and structural tokens in templates significantly constrain the model’s output space. |
Deriving Character Logic from Storyline as Codified Decision Trees (2026.acl-long)
Copied to clipboard
| Challenge: | Existing behavioral profiles are unstructured, weakly validated, and unusable . existing models are weakly valid, leading to brittle agent behavior . Using codified decision trees, we show that CDT outperforms previous methods . |
| Approach: | They propose a data-driven framework that induces an executable decision structure from narrative data. |
| Outcome: | The proposed framework outperforms human-written profiles and prior profiles on multiple benchmarks. |
End-to-End Optimization for Multimodal Retrieval-Augmented Generation via Reward Backpropagation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | MM-RAG is a promising approach for enhancing the reliability and factuality of large vision-language models . current methods focus on component-level optimizations and necessitate extensive component-specific training datasets . |
| Approach: | They propose a new paradigm that backpropagates global rewards to each component . this backpropage transforms local losses into specific local losses . |
| Outcome: | The proposed paradigm achieves high training efficiency on knowledge-intensive multimodal benchmarks. |
ULTRABENCH: Benchmarking LLMs under Extreme Fine-grained Text Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks evaluate models on only a few attributes, typically fewer than five . a new benchmark evaluates large language models under dense, multi-attribute constraints . |
| Approach: | They propose a benchmark that evaluates large language models under dense, multi-attribute constraints. |
| Outcome: | The proposed benchmark evaluates large language models under dense, multi-attribute constraints. |