Papers by James Thorne

31 papers
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmark datasets for Korean cultural and linguistic knowledge are derived from the English counterparts through translation, so they overlook cultural contexts.
Approach: They propose to use Korean cultural and linguistic intelligence to assess Korean model performance by providing fine-grained annotations of cultural and cultural knowledge.
Outcome: The proposed dataset includes 1,995 QA pairs and is based on 1,992 Korean exams and textbooks.
Context Filtering with Reward Modeling in Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Question Answering (QA) tasks require a mix of relevant and irrelevant information in these contexts to perform well.
Approach: They propose a context filtering approach that removes non-essential details, summarizing crucial content through Reward Modeling.
Outcome: The proposed approach outperforms baseline models in 6.8-folds.
Capturing the Relationship Between Sentence Triplets for LLM and Human-Generated Texts to Enhance Sentence Embeddings (2024.findings-eacl)

Copied to clipboard

Challenge: Recent advances in building sentence embedding models have centered on replacing traditional human-generated text datasets with those generated by LLMs.
Approach: They propose a loss function that incorporates Positive-Negative sample Augmentation within the contrastive learning objective to enhance sentence embeddings using both human and LLM-generated datasets.
Outcome: The proposed model mitigates the sentence anisotropy problem in Wikipedia corpus and improves Spearman’s correlation in standard Semantic Textual Similarity (STS) tasks (+1.47% compared to CLHAIF).
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions (2025.acl-long)

Copied to clipboard

Challenge: Existing studies show that direct generation of diagram descriptions is costly and biased against blind and low-vision (BLV) users.
Approach: They ask sighted individuals to assess diagram descriptions generated by vision-language models . they use latent supervision to guide the models with latent inference .
Outcome: The results show that visual descriptions generated by vision-language models are effective and useful to educators who are themselves BLV and teach visually impaired learners.
Re3val: Reinforced and Reranked Generative Retrieval (2024.findings-eacl)

Copied to clipboard

Challenge: generative retrieval models encode pointers to information in a corpus as an index within the model’s parameters.
Approach: They propose a generative retrieval model that leverages contextual information to rerank retrieved page titles and utilizes REINFORCE to maximize rewards generated by constrained decoding.
Outcome: The proposed model can't be tuned for the downstream readers as decoding the page title is a non-differentiable operation.
Evaluating the Consistency of LLM Evaluators (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown potential as general evaluators with the benefits of speed and cost.
Approach: They conduct extensive studies on the two aspects of consistency in LLM evaluations, Self-Consistency (SC) and Inter-scale Consistency on different scoring scales and criterion granularity with open-source and proprietary models.
Outcome: The results show that strong proprietary models are not necessarily consistent evaluators, highlighting the importance of considering consistency in assessing the capability of LLM evalueators.
Automated Fact Checking: Task Formulations, Methods and Future Directions (C18-1)

Copied to clipboard

Challenge: Recent research on fact checking has focused on misinformation . however, relevant papers and articles have been published in research communities that are unaware of each other and use inconsistent terminology.
Approach: They propose avenues for future NLP research on automated fact checking . they highlight the use of evidence as an important distinguishing factor .
Outcome: The proposed methods unify the task formulations and methodologies across papers and authors.
Epistemology of Language Models: Do Language Models Have Holistic Knowledge? (2024.findings-acl)

Copied to clipboard

Challenge: et al., 2021) explores whether language models exhibit characteristics consistent with epistemological holism . authors examined the epistle of language models from the perspective of abduction, revision, and argument generation tasks.
Approach: They examine whether language models exhibit characteristics consistent with epistemological holism . they created a scientific reasoning dataset and examined the epistology of language models .
Outcome: The language models showed that they did not distinguish between core and peripheral knowledge, compared with other tasks.
BEnQA: A Question Answering Benchmark for Bengali and English (2024.findings-acl)

Copied to clipboard

Challenge: a dataset of parallel Bengali and English exam questions is used to compare LLMs in low-resource languages.
Approach: They introduce BEnQA, a dataset comprising parallel Bengali and English exam questions . they benchmark several Large Language Models with their parallel dataset and observe performance disparity .
Outcome: The proposed dataset consists of 5K questions covering several subjects in science . the authors find that the models perform poorly in Bengali and English .
Disentangling Structure and Style: Political Bias Detection in News by Inducing Document Hierarchy (2023.findings-emnlp)

Copied to clipboard

Challenge: a new method to detect political bias in news articles overcomes this domain dependency . partisan bias exists in various social issues, including the 2016 presidential election .
Approach: They propose a multi-head hierarchical attention model that encodes the structure of long documents through a diverse ensemble of attention heads.
Outcome: The proposed model outperforms existing methods for detecting political bias in news articles.
Knowledge Corpus Error in Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work in open-domain question answering (QA) has explored generating context passages from large language models (LLMs) however, it is not well understood why generated passages can be more effective than retrieved ones.
Approach: They propose to generate context passages from large language models by paraphrasing human-annotated gold context using LLMs to observe knowledge corpus error.
Outcome: The proposed framework shows that paraphrasing human-annotated gold contexts improves performance over retrieval steps.
FactKG: Fact Verification via Reasoning on Knowledge Graphs (2023.acl-long)

Copied to clipboard

Challenge: knowledge graphs (KGs) have not been fully utilized as a knowledge source for fact verification.
Approach: They propose a dataset to enable the community to better use knowledge graphs . they propose 108k natural language claims with five types of reasoning .
Outcome: The proposed dataset consists of 108k natural language claims with five types of reasoning . authors believe the proposed method can advance reliability and practicality .
FEVER: a Large-scale Dataset for Fact Extraction and VERification (N18-1)

Copied to clipboard

Challenge: 185,445 claims generated by altering sentences from Wikipedia are verified without knowledge of the sentence they were derived from.
Approach: They propose a publicly available dataset for verification against textual sources, FEVER: Fact Extraction and VERification.
Outcome: The proposed dataset achieves 31.87% accuracy on labeling a claim accompanied by the correct evidence, compared to 50.91% if we ignore the evidence.
Generating Token-Level Explanations for Natural Language Inference (N19-1)

Copied to clipboard

Challenge: Existing methods to generate token-level explanations for NLI on single sentences have not been tested.
Approach: They propose to generate token-level explanations for NLI without explicitly annotating training data.
Outcome: The proposed approach is faster and more accurate than the black-box methods.
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that pre-training compute can improve multilingual performance, but is it effective for test-time scaling?
Approach: They propose a multilingual math benchmark with competition-level problems in 55 languages . they propose "test-time scaling" which further lengthens the time it takes to scale .
Outcome: The proposed methods fail to generalize robustly across languages, with no improvements in variance or consistency.
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question.
Approach: They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries.
Outcome: The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images.
ORPO: Monolithic Preference Optimization without Reference Model (2024.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models with vast training corpora have shown remarkable abilities in diverse natural language processing tasks.
Approach: They propose a model-free monolithic odds ratio preference optimization algorithm, ORPO, to improve preference alignment.
Outcome: The proposed algorithm outperforms state-of-the-art language models with more than 7B and 13B parameters on the ultrafeedback alone.
Database reasoning over text (2021.acl-long)

Copied to clipboard

Challenge: Existing models cannot handle database queries such as “List/Count all female athletes who were born in 20th century”.
Approach: They propose a modular architecture to answer database-style queries over multiple spans from text and aggregate them at scale.
Outcome: The proposed architecture scales to databases containing thousands of facts whereas current models are limited by how many facts can be encoded.
HARE: Explainable Hate Speech Detection with Step-by-Step Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent benchmarks have attempted to identify and explain hate speech but lack the reasoning to supervise detection models.
Approach: They propose a framework that uses large language models to fill in the gaps in hate speech explanations by using existing annotations.
Outcome: The proposed framework outperforms baselines on SBIC and Implicit Hate using model-generated data and improves generalization to unseen datasets.
KILT: a Benchmark for Knowledge Intensive Language Tasks (2021.naacl-main)

Copied to clipboard

Challenge: Existing models for knowledge-intensive language tasks require access to large, external knowledge sources.
Approach: They propose a benchmark for knowledge-intensive language tasks (KILT) they test a shared dense vector index coupled with a seq2seq model to generate disambiguated text.
Outcome: The proposed model outperforms tailor-made approaches on fact checking, open-domain question answering and dialog by generating disambiguated text.
Evaluating adversarial attacks against multiple fact verification systems (D19-1)

Copied to clipboard

Challenge: Automated fact verification is progressing due to advances in modeling and availability of large datasets.
Approach: They propose two scoring metrics which take into account the correctness of adversarial instances.
Outcome: The proposed method and paraphrasing method have higher potency and higher resilience than baselines.
Learning to Insert [PAUSE] Tokens for Better Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies have explored incorporating special-purpose tokens into the training process to enhance reasoning capabilities.
Approach: They propose a method for inserting dummy tokens consecutively just before reasoning steps to increase model effectiveness.
Outcome: The proposed method outperforms fine-tuning and previous token insertion methods on multiple datasets and models.
Elastic weight consolidation for better bias inoculation (2021.eacl-main)

Copied to clipboard

Challenge: Recent studies have shown that the lack of suitable inductive biases in sentence-pair classification models can cause misclassifications on training datasets.
Approach: They propose to use elastic weight consolidation (EWC) to fine-tune models to mitigate biases while being less susceptible to catastrophic forgetting.
Outcome: The proposed model improves on fact verification and stress tests while maintaining the original task accuracy.
Evidence-based Factual Error Correction (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to correct factual errors are limited to labeled claims . a recent task of fact verification has attracted significant attention .
Approach: They propose a task of factual error correction that performs edits to a claim so that the generated rewrite is better supported by evidence.
Outcome: The proposed method produces accurate factual error corrections for 5x more instances in human evaluation and a .125 increase in SARI score.
The FEVER2.0 Shared Task (D19-66)

Copied to clipboard

Challenge: Existing deep neural models are becoming more complex and difficult to understand and characterize their behaviour.
Approach: They present the results of the second Fact Extraction and VERification (FEVER2.0) Shared Task.
Outcome: The proposed task was based on the second Fact Extraction and VERification (FEVER2.0) shared task.
Cross-lingual Transfer of Reward Models in Multilingual Alignment (2025.naacl-short)

Copied to clipboard

Challenge: Recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments.
Approach: They investigate cross-lingual transfer of English RMs by representation shifts . they also analyze cross-linguistic transfer of RM through the representation shift .
Outcome: The results show that English RMs can be transferred across languages by 34% .
I0T: Embedding Standardization Method Towards Zero Modality Gap (2025.acl-long)

Copied to clipboard

Challenge: Recent studies on Contrastive Language-Image Pretraining suffer from a *modality gap* . modality gap occurs when image and text embeddings are projected to disparate manifolds .
Approach: They propose a framework that reduces the modality gap by adding two normalization layers to each encoder.
Outcome: The proposed framework reduces the modality gap while preserving the original embedding representations of trained models with their locked parameters.
From Evidence to Belief: A Bayesian Epistemology Approach to Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on language models have focused on factual correctness and justification, but prior research has focused on the factual truth condition and justifier.
Approach: They analyze language models’ responses and confidence using verbalized confidence, token probability, and sampling to examine their knowledge of Bayesian epistemology.
Outcome: The language models that follow the Bayesian confirmation assumption with true evidence show varying performance depending on the degree of irrelevance, indicating they deviate from Bayes' assumptions.
Detrimental Contexts in Open-Domain Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Using the whole passages in QA datasets can improve model accuracy by 10% .
Approach: They analyze how passages can have a detrimental effect on retrieve-then-read architectures used in question answering when evaluated on common question answering datasets.
Outcome: The proposed model accuracy can be improved by 10% on two popular QA datasets by filtering out detrimental passages.
Can Large Language Models Capture Dissenting Human Voices? (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks.
Approach: They evaluate the performance and alignment of large language models with humans using Monte Carlo Estimation and Log Probability Estimationic methods to estimate the multinomial distribution.
Outcome: The proposed models fail to capture human disagreement distribution and inference and human alignment performance plunge even further on data samples with high disagreement levels raising concerns about their natural language understanding ability and representativeness to a larger human population.
Stable Language Model Pre-training by Reducing Embedding Variability (2024.emnlp-main)

Copied to clipboard

Challenge: Stable pre-training is essential for achieving better-performing language models, but tracking pre-train stability is impractical due to high computational costs.
Approach: They propose to use Token Embedding Variability as a proxy to estimate pre-training stability.
Outcome: The proposed method improves stability and lowers perplexities even at deeper layer counts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations