Papers by Ying-Cong Chen
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have focused on test-time scaling to improve reasoning quality but at the cost of efficiency. |
| Approach: | They propose a training-free framework that enhances reasoning accuracy and stability with minimal overhead. |
| Outcome: | The proposed framework yields consistent gains across general, coding, and STEM tasks while remaining highly efficient. |
SceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for synthesising 3D scenes from a single image are text-driven and lack precise metric understanding from images. |
| Approach: | They propose a language-model-based framework that grounds 3D scene synthesis in visual evidence by recovering an executable metric 3D layout directly from a single image. |
| Outcome: | The proposed framework recovers an executable metric 3D layout directly from an RGB image and instantiates, places, and edits objects for iterative refinement. |
Orchestrating Audio: Multi-Agent Framework for Long-Video Audio Synthesis (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for video-to-audio dubbing for long-form content are fragmented and lack dedicated datasets. |
| Approach: | They propose a multi-agent framework that offers a coordinated, multi-component approach to long-video audio generation. |
| Outcome: | The proposed method outperforms state-of-the-art V2A models in audio quality. |
PreGenie: An Agentic Framework for High-quality Visual Presentation Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Visual presentations are vital for effective communication, but they are limited by their complexity and lack of visual understanding. |
| Approach: | a new framework is proposed to generate high-quality visual presentations using multimodal large language models. |
| Outcome: | The proposed framework outperforms existing models in multimodal understanding and content consistency. |