Papers by Alex Kim

8 papers
LLMs Behind the Scenes: Enabling Narrative Scene Illustration (2025.emnlp-main)

Copied to clipboard

Challenge: Generative AI has established the ability to readily transform content from one medium to another.
Approach: They propose a pipeline that uses large language models to prompt text-to-image models to generate scenes for story text.
Outcome: The proposed pipeline synthesizes illustrations for scenes in a story corpus using human annotation tasks.
Towards Robust Mathematical Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: IMO-Bench is a suite of advanced reasoning benchmarks that targets the international mathematical Olympiad level.
Approach: They propose IMO-Bench, a suite of advanced reasoning benchmarks that targets the level of the international mathematical Olympiad.
Outcome: IMO-Bench is a suite of advanced reasoning benchmarks that targets the level of the international mathematical Olympiad.
Large Language Models as Realistic Microservice Trace Generators (2025.emnlp-main)

Copied to clipboard

Challenge: Obtaining real-world traces is difficult due to limited public data availability and the difficulty of collecting them at large scale from diverse environments.
Approach: They propose to train a large language model to generate microservice call graphs using a recursive approach to capture hierarchical structures and implicit constraints in such traces.
Outcome: The proposed method outperforms existing methods in accuracy and validity.
Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling (P19-1)

Copied to clipboard

Challenge: State-of-the-art models in natural language processing (NLP) often incorporate sentence encoder functions which generate a sequence of vectors intended to represent the in-context meaning of each word in an input text.
Approach: They conduct the first large-scale systematic study of candidate pretraining tasks, comparing 19 different tasks as alternatives and complements to language modeling.
Outcome: The proposed model can be used to train sentences on language modeling tasks.
DEBATE: Devil’s Advocate-Based Assessment and Text Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for evaluating the quality of machine-generated texts have a relatively low correlation with human performance.
Approach: They propose an NLG evaluation framework based on multi-agent scoring system augmented with a concept of Devil’s Advocate.
Outcome: The proposed evaluation framework outperforms the previous state-of-the-art methods in two meta-evaluation benchmarks in NLG evaluation, SummEval and TopicalChat.
HARE: an entity and relation centric evaluation framework for histopathology reports (2025.findings-emnlp)

Copied to clipboard

Challenge: evaluating the clinical quality of medical domain automated text generation remains a challenge.
Approach: They propose a framework for histopathology automated report evaluation that prioritizes clinically relevant content by aligning critical histo pathology entities and relations between reference and generated reports.
Outcome: The proposed framework outperforms existing metrics in histopathology report evaluations.
Reconstruction Probing (2023.findings-acl)

Copied to clipboard

Challenge: a new analysis method for contextualized representations is proposed . contextualization boosts reconstructability of tokens close to the token being reconstructed .
Approach: They propose a method for contextualized representations based on reconstruction probabilities in masked language models.
Outcome: The proposed method compares reconstruction probabilities of tokens in masked language models . it finds that contextualization boosts reconstructability of token that are close to the token being reconstructed .
Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling (2024.acl-long)

Copied to clipboard

Challenge: a novel application of large language models (LLMs) to legal education helps non-experts learn complex legal concepts . authors find storytelling helps nonexperts understand complex legal terms and concepts compared to definitions .
Approach: They propose a novel application of large language models to legal education . they use LLMs to generate legal stories explaining complex legal concepts .
Outcome: The proposed method improves comprehension and interest among non-native speakers compared to definitions . the novel method also shows that non-experts retain more stories .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations