Papers by Jiasen Lu
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent work has adapted vision-and-language models to generative tasks like image captioning. |
| Approach: | They propose an extension to LXMERT with training refinements to generate images from text. |
| Outcome: | The proposed model can generate images from pieces of text while still being comparable to existing models. |
Reducing Token Redundancy in LVLMs: A Systematic Review of Token Pruning Methods (2026.acl-long)
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) excel at visual understanding but face severe computational bottlenecks when processing high-resolution images and long videos due to massive visual token counts. |
| Approach: | They propose a taxonomy categorizing methods into vision-side, LLM-side and hybrid paradigms and analyze token selection mechanisms and pruning strategy. |
| Outcome: | The proposed method selectively removes less informative tokens while maintaining performance. |