Papers by Yujin Yang

5 papers
Exploring In-context Example Generation for Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated strong performance across various tasks with just a few examples.
Approach: They propose a method that generates in-context example pairs without external resources.
Outcome: The proposed method builds upon two prior criteria, relevance and diversity, which have been highlighted as key factors for in-context example selection.
Self-Training Elicits Concise Reasoning in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Chain-of-thought reasoning has enabled large language models to use additional computation through intermediate tokens to solve complex tasks, but current models often generate more tokens than necessary to accomplish the task, incurring extraneous inference costs.
Approach: They propose to fine-tune models with self-generated concise reasoning paths obtained by best-of-N sampling and few-shot conditioning in task-specific settings to elicit concise reasoning.
Outcome: The proposed method reduces output tokens by 30% on GSM8K and MATH while maintaining average accuracy.
FinHarmBench: Financial Jailbreak Benchmark and Unsupervised Safety Fine-Tuning via Refusal Steering Distillation (2026.acl-industry)

Copied to clipboard

Challenge: Existing safety benchmarks focus on general harms and lack the granularity needed to capture domain-specific financial threats.
Approach: They propose a benchmark to evaluate financially harmful and confusable benign prompts.
Outcome: The proposed framework improves refusal behavior without annotating refusal responses.
HARE: Explainable Hate Speech Detection with Step-by-Step Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent benchmarks have attempted to identify and explain hate speech but lack the reasoning to supervise detection models.
Approach: They propose a framework that uses large language models to fill in the gaps in hate speech explanations by using existing annotations.
Outcome: The proposed framework outperforms baselines on SBIC and Implicit Hate using model-generated data and improves generalization to unseen datasets.
Towards Formality-Aware Neural Machine Translation by Leveraging Context Information (2023.findings-emnlp)

Copied to clipboard

Challenge: Formality is one of the most important linguistic properties to determine the naturalness of translation.
Approach: They propose a method to explicitly inform neural machine translation models by pinpointing key informative tokens using a formality classifier.
Outcome: The proposed method improves translation quality and conforms to the appropriate syntax.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations