Papers by Howard Yen
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent work has identified retrieval heads as a subset of attention heads responsible for retrieving salient information in long-context language models. |
| Approach: | They introduce a retrieval head that uses attention scores to enhance retrieval from long context . they use QRRetriever to select the most relevant parts with the highest retrieval scores . |
| Outcome: | The proposed retrieval heads outperform other retrieval-based retrieval retrievers on BEIR benchmarks. |
How to Train Long-Context Language Models (Effectively) (2025.acl-long)
Copied to clipboard
| Challenge: | a new study shows that language models can process extremely long contexts with minimal training. |
| Approach: | They use supervised fine-tuning and continued training to evaluate a language model's long-context capabilities. |
| Outcome: | The proposed model outperforms Llama-3.1-8B-Instruct on most long-context tasks . the model can process 512K tokens, one of the longest context windows of LMs . |
Enabling Large Language Models to Generate Text with Citations (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work relies on commercial search engines and human evaluation, making it difficult to reproduce and compare different modeling approaches. |
| Approach: | They propose a new generation paradigm that requires large language models to provide citations to one or a few text passages for any statement they generate. |
| Outcome: | The proposed model improves factual correctness and verifiability of large language models by providing citations to a set of questions and retrieval corpora and generating answers with citation. |
Long-Context Language Modeling with Parallel Context Encoding (2024.acl-long)
Copied to clipboard
| Challenge: | Existing long-context models degenerate with retrieved contexts. |
| Approach: | They propose a framework that can be applied to existing decoder-only LLMs for context expansion. |
| Outcome: | The proposed framework can be applied to any existing decoder-only LLMs for context expansion. |