Papers by Minjin Choi

5 papers
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have recently exhibited performance gains owing to a wide variety of prompting techniques, including Retrieval-Augmented Generation (RAG), Chain-of-Thought (CoT), and In-Context Learning (ICL).
Approach: They propose a prompt compression method that captures the global context without compromising semantic consistency while detouring the necessity of pseudo-labels for training the compressor.
Outcome: Empirical results show that the proposed method retains key contexts while reducing the prompt length by 80%.
MelBERT: Metaphor Detection via Contextualized Late Interaction using Metaphorical Identification Theories (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies have developed computational models to recognize metaphorical words in sentences.
Approach: They propose a model that leverages contextualized word representation and linguistic metaphor identification theories to detect whether the target word is metaphorical.
Outcome: The proposed model outperforms baseline models on four benchmark datasets . it leverages contextualized word representation and linguistic metaphor identification theories to detect whether the target word is metaphorical.
GLEN: Generative Retrieval via Lexical Index Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document retrieval bypass auxiliary index structures and can be optimized through end-to-end learning.
Approach: They propose a method to generate a relevant document's identifier using an index learning strategy.
Outcome: The proposed method achieves state-of-the-art or competitive performance on benchmark datasets.
GRAM: Generative Recommendation via Semantic-aware Multi-granular Late Fusion (2025.acl-long)

Copied to clipboard

Challenge: Existing studies rely on item metadata to construct abbreviated item IDs, leading to a loss of valuable details.
Approach: They propose a Generative Recommender via semantic-aware multi-granular late fusion to integrate rich semantics efficiently with minimal information loss.
Outcome: The proposed model outperforms eight state-of-the-art recommendation models on four benchmark datasets and achieves significant improvements of 11.5-16.0% in Recall@5 and 5.3-13.6% in NDCG@5.
TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks fail to evaluate large language models' instruction-following capabilities . current benchmarks lack multilinguality, implicit constraints and multi-turn dialogue .
Approach: a new benchmark is designed to evaluate large language models' instruction-following capabilities . the benchmark features input prompts across 12 languages and includes inter-instance multilingual instructions .
Outcome: a new benchmark for large language models (LLMs) is designed to assess their performance in real-world settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations