Papers by Noah Lee

7 papers
That was the last straw, we need more: Are Translation Systems Sensitive to Disambiguating Context? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for translation of ambiguous text use context to disambiguate meaning . current models for MTs consistently translate English idioms literally, whereas LMs are context-aware .
Approach: They use a dataset of 512 pairs of English sentences to study semantic ambiguities . they use literal and figurative idioms to disambiguate intended meaning .
Outcome: The results show that current models translate English idioms literally, even when the context suggests a figurative interpretation.
Evaluating the Consistency of LLM Evaluators (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown potential as general evaluators with the benefits of speed and cost.
Approach: They conduct extensive studies on the two aspects of consistency in LLM evaluations, Self-Consistency (SC) and Inter-scale Consistency on different scoring scales and criterion granularity with open-source and proprietary models.
Outcome: The results show that strong proprietary models are not necessarily consistent evaluators, highlighting the importance of considering consistency in assessing the capability of LLM evalueators.
ORPO: Monolithic Preference Optimization without Reference Model (2024.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models with vast training corpora have shown remarkable abilities in diverse natural language processing tasks.
Approach: They propose a model-free monolithic odds ratio preference optimization algorithm, ORPO, to improve preference alignment.
Outcome: The proposed algorithm outperforms state-of-the-art language models with more than 7B and 13B parameters on the ultrafeedback alone.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
Cross-lingual Transfer of Reward Models in Multilingual Alignment (2025.naacl-short)

Copied to clipboard

Challenge: Recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments.
Approach: They investigate cross-lingual transfer of English RMs by representation shifts . they also analyze cross-linguistic transfer of RM through the representation shift .
Outcome: The results show that English RMs can be transferred across languages by 34% .
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Southeast Asia (SEA) is home to over 1,300 indigenous languages and 671 million people . prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA .
Approach: They propose to provide a resource center that provides standardized corpora in nearly 1,000 SEA languages across three modalities.
Outcome: a new benchmark assesses the quality of AI models on 36 SEA languages across 13 tasks . the results highlight the importance of SEA as a culturally diverse region .
Can Large Language Models Capture Dissenting Human Voices? (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks.
Approach: They evaluate the performance and alignment of large language models with humans using Monte Carlo Estimation and Log Probability Estimationic methods to estimate the multinomial distribution.
Outcome: The proposed models fail to capture human disagreement distribution and inference and human alignment performance plunge even further on data samples with high disagreement levels raising concerns about their natural language understanding ability and representativeness to a larger human population.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations