Papers by Tom Hope

24 papers
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Evaluating debate speeches requires a deep understanding of arguments at multiple levels.
Approach: They propose a benchmark task for LLM judges based on annotated debate speeches . they analyze the judgment capabilities and behavior of frontier LLMs .
Outcome: The proposed task requires a comprehensive understanding of argumentation and its arguments.
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback (2026.eacl-long)

Copied to clipboard

Challenge: Novelty assessment is a central yet understudied aspect of peer review . manuscript submissions double roughly every 15 years, and individual reviewers now complete an average of 14 reviews per year.
Approach: They propose a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction, retrieval and synthesis of related work, and structured comparison for evidence-based assessment.
Outcome: The proposed approach outperforms existing LLM-based baselines on 182 ICLR 2025 submissions with human-annotated reviewer novelty assessments.
LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models (2026.findings-acl)

Copied to clipboard

Challenge: Using retrieval augmentation, large vision language models can be used for diagnostic accuracy, but multimodal retrieval-augmented diagnosis is challenging.
Approach: They propose a lightweight mechanism for enhancing diagnostic performance of retrieval-augmented LVLMs by fine-tuning a multimodal retriever and general-purpose backbone models.
Outcome: The proposed mechanism achieves competitive results without medical training compared to pre-trained models with extensive training.
Computational Discovery of Chiasmus in Ancient Religious Text (2025.naacl-short)

Copied to clipboard

Challenge: chiasmus, or chiastic units, is a debated literary device in biblical texts . a computational approach to detect chiastes is shown to be efficient, but not efficient .
Approach: They propose a computational approach to detect chiasmus within Biblical passages . they leverage neural embeddings to capture lexical and semantic patterns associated with chiastics - using annotators to review a subset of the detected patterns.
Outcome: The proposed method achieves high inter-annotator agreement and system accuracy of 0.80 at verse level and 0.60 at half-verse level.
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews (2024.acl-long)

Copied to clipboard

Challenge: Existing systems that can interpret complex writing feedback and edit documents in response are limited on the most demanding writing tasks.
Approach: They propose to use peer feedback to revise scientific papers based on peer feedback . they provide labels linking each reviewer comment to the specific paper edits made by the author .
Outcome: The proposed model fails to identify which edits correspond to a comment and the original paper.
SciMON: Scientific Inspiration Machines Optimized for Novelty (2024.acl-long)

Copied to clipboard

Challenge: Existing literature-based hypothesis generation models focus on binary link prediction, limiting expressivity of hypotheses.
Approach: They propose a framework that uses literature-based hypothesis generation as input . they use literature-derived literature as background and output natural language ideas .
Outcome: The proposed model improves the ability of language models to generate new scientific directions grounded in literature.
Multi-Vector Models with Textual Guidance for Fine-Grained Scientific Document Similarity (2022.naacl-main)

Copied to clipboard

Challenge: Using co-citations, we can train a model that matches aspects of papers to document level similarity.
Approach: They propose a model that matches fine-grained aspects of papers and aggregates them into a document level similarity model using a naturally-occurring source of supervision: co-citations.
Outcome: The proposed model improves performance on document similarity tasks in four datasets and achieves competitive results.
CARE: Extracting Experimental Findings From Clinical Literature (2024.findings-naacl)

Copied to clipboard

Challenge: Existing annotation schemas and datasets fail to capture real-world complexity and nuance of experimental findings.
Approach: They propose a new annotation schema capturing fine-grained findings as n-ary relations between entities and attributes.
Outcome: The proposed schema captures fine-grained findings as n-ary relations between entities and attributes.
A Dataset for N-ary Relation Extraction of Drug Combinations (2022.naacl-main)

Copied to clipboard

Challenge: Combination therapies are becoming standard of care for diseases such as cancer, tuberculosis, malaria and HIV.
Approach: They construct an expert-annotated dataset for extracting drug combinations from the scientific literature.
Outcome: The proposed dataset is the first relation extraction dataset consisting of variable-length relations.
SciSight: Combining faceted navigation and research group detection for COVID-19 exploratory scientific search (2020.emnlp-demos)

Copied to clipboard

Challenge: SciSight is a system for exploratory search of COVID-19 literature . it explores associations between biomedical facets extracted from papers .
Approach: They propose a system for exploratory search of COVID-19 literature that integrates two key capabilities: first, exploring associations between biomedical facets automatically extracted from papers; second, combining textual and network information to search and visualize groups of researchers and their ties.
Outcome: The proposed system has served over 15K users with over 42K page views and 13% returns.
On-the-fly Definition Augmentation of LLMs for Biomedical NER (2024.naacl-long)

Copied to clipboard

Challenge: Despite their general capabilities, LLMs struggle on biomedicalNER tasks due to specialized terminology and lack of training data.
Approach: They propose a new knowledge augmentation approach which incorporates definitions of relevant concepts on-the-fly.
Outcome: The proposed approach improves performance on biomedicalNER tasks by 15% (on average) The proposed method outperforms fine-tuned language models in few-shot settings.
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links (2026.eacl-long)

Copied to clipboard

Challenge: Using retrieval models and LLMs achieves a 73% approval rate for suggested links, more than doubling the acceptance of strong retrievers alone.
Approach: They propose a domain-agnostic framework for bootstrapping sentence-level cross-document links from scratch and apply it to large-scale human-in-the-loop annotation of natural text pairs.
Outcome: The proposed framework generates semi-synthetic datasets and uses them to benchmark and shortlist the best-performing methods and applies them in large-scale human-in-the-loop annotation of natural text pairs.
Literature-Augmented Clinical Outcome Prediction (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to clinical outcome prediction use only clinical notes and general biomedical literature.
Approach: They propose to retrieve patient-specific medical literature and incorporate it into predictive models by combining clinical notes with language models.
Outcome: The proposed approach boosts predictive performance on three important clinical tasks in comparison to strong LM baselines, increasing F1 by up to 5 points and precision@Top-K by a large margin of over 25%.
In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis (2026.acl-long)

Copied to clipboard

Challenge: citation counts are a shallow view that fails to capture how a paper has influenced subsequent work.
Approach: They propose a task to generate nuanced, expressive, and time-aware impact summaries . they analyze fine-grained confirmatory and correction citation intents to generate summary .
Outcome: The proposed task shows moderate to strong human correlation on subjective metrics such as insightfulness.
ACCoRD: A Multi-Document Approach to Generating Diverse Descriptions of Scientific Concepts (2022.emnlp-demos)

Copied to clipboard

Challenge: Current systems that automatically define unfamiliar terms only surface a single "best" description for all users, which may not be accessible for all readers, given varying background knowledge.
Approach: They propose an end-to-end system that generates sets of descriptions of scientific concepts . ACCoRD corpus includes 1,275 labeled contexts and 1,787 expert-authored concept descriptions .
Outcome: The proposed system produces diverse descriptions of concepts in terms of reference concepts.
Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces (2025.naacl-short)

Copied to clipboard

Challenge: a new study characterizes the preservation of intertextuality across human and machine translations . intertextual references can range from direct quotation to semantic resemblance, both within and between texts .
Approach: They use multilingual embedding spaces to characterize preservation of intertextuality . they use biblical texts, which are both full of inter textual references .
Outcome: The proposed method characterizes preservation of intertextuality across human and machine translations.
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation (2026.acl-long)

Copied to clipboard

Challenge: a hallmark of human innovation is recombination.
Approach: They propose a task to extract recombination instances from scientific literature . they analyze patterns of recombined concepts and apply it to a broad corpus of AI papers .
Outcome: The proposed model can predict cross-disciplinary research directions . it can predict recombinations across areas and link methods and concepts .
Beyond Good Intentions: Reporting the Research Landscape of NLP for Social Good (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have created a vast number of applications that are aimed at social good applications.
Approach: They propose a dataset with three tasks that can help identify NLP4SG papers and characterize the NLP landscape by: (1) identifying the papers that address a social problem, (2) mapping them to the corresponding UN Sustainable Development Goals, and (3) identifying their methods.
Outcome: The proposed dataset can help identify NLP4SG papers and characterize the NLP landscape by: (1) identifying the papers that address a social problem, (2) mapping them to the corresponding UN Sustainable Development Goals (SDGs), and (3) identifying their methods.
Extracting a Knowledge Base of Mechanisms from COVID-19 Papers (2021.naacl-main)

Copied to clipboard

Challenge: COVID-19 has spawned a diverse body of scientific literature that is challenging to navigate . researchers are using automated tools to help find useful knowledge .
Approach: They develop a schema to extract mechanism relations from scientific papers . their search engine, dataset and code are publicly available .
Outcome: The proposed schema outperforms PubMed search in clinical trials.
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation (2025.findings-acl)

Copied to clipboard

Challenge: Automated scientific discovery (ASD) systems are limited in their evaluation of software artifacts and large volumes of research artifs are typically evaluated using conference-style paper review with limited evaluation of code.
Approach: They propose a novel ASD system that frames ideation and experiment construction as a form of genetic search jointly over combinations of research articles and codeblocks defining common actions in a domain.
Outcome: The proposed system returns 19 discoveries on machine-generated ideas in the domain of agents and virtual environments.
Language (Re)modelling: Towards Embodied Language Understanding (2020.acl-main)

Copied to clipboard

Challenge: Despite the rapid progress in NLU, current systems lack the rich mental representations that people use for language understanding.
Approach: They propose an approach to representation and learning based on the tenets of embodied cognitive linguistics (ECL) they propose a system architecture along with a roadmap towards realizing this vision.
Outcome: The proposed approach will improve the performance of existing systems and provide a roadmap towards realizing this vision.
CHAMP: Efficient Annotation and Consolidation of Cluster Hierarchies (2023.emnlp-demo)

Copied to clipboard

Challenge: Various annotation tasks require a complex hierarchical structure over nodes, where each node is a cluster of items.
Approach: They propose an open source tool that incrementally constructs clusters and hierarchy simultaneously over any type of text.
Outcome: The proposed approach significantly reduces annotation time and guarantees transitivity at the cluster and hierarchy levels.
Towards a Human-Computer Collaborative Scientific Paper Lifecycle: A Pilot Study and Hands-On Tutorial (2024.lrec-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an overview of the scientific paper lifecycle . large language models (LLMs) have increasingly played an important role in academic writing .
Approach: They propose to provide an overview of the scientific paper lifecycle using large language models.
Outcome: The tutorial will provide an overview of the scientific paper lifecycle, including scientific literature understanding, experiment development, manuscript draft writing, and finally draft evaluation.
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature (2025.emnlp-main)

Copied to clipboard

Challenge: ScIRIFF is the only entirely expert-written instruction-following dataset for scientific literature understanding . it features complex instructions with long input contexts, detailed task descriptions, and structured outputs.
Approach: They present a dataset of 137K instruction-following instances for training and evaluation . they finetuned large language models using a mix of general domain and ScIRIFF instructions .
Outcome: The proposed dataset shows that on nine out-of-distribution held-out tasks, the model performs better than baselines trained on general domain instructions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations