Xuan Wang, Yingjun Guan, Weili Liu, Aabhas Chauhan, Enyi Jiang, Qi Li, David Liem, Dibakar Sigdel, John Caufield, Peipei Ping, Jiawei Han
| Challenge: | EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences. |
| Approach: | They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
| Outcome: | EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
Similar Papers
BiomedCurator: Data Curation for Biomedical Literature (2022.aacl-demo)
Copied to clipboard
Mohammad Golam Sohrab, Khoa N.A. Duong, Ikeda Masami, Goran Topić, Yayoi Natsume-Kitatani, Masakata Kuroda, Mari Nogami Itoh, Hiroya Takamura
| Challenge: | BiomedCurator uses state-of-the-art natural language processing techniques to extract structured data from scientific articles. |
| Approach: | They propose a web application that extracts structured data from PubMed and ClinicalTrials.gov . the application uses a combination of natural language processing techniques and a pattern-based extraction approach . |
| Outcome: | The proposed system extracts the structured data from PubMed and ClinicalTrials.gov datasets. |
NLATool: an Application for Enhanced Deep Text Understanding (C18-2)
Copied to clipboard
Markus Gärtner, Sven Mayer, Valentin Schwind, Eric Hämmerle, Emine Turcan, Florin Rheinwald, Gustav Murawski, Lars Lischke, Jonas Kuhn
| Challenge: | a wide range of subfields in natural language processing see systems solving their tasks with sufficiently high-quality levels. |
| Approach: | They propose a web application that supports text annotation and enriches the text with additional information from a number of sources directly within the application. |
| Outcome: | The proposed web application is based on a human-centered design process . it offers a rich visualization of texts and the entities mentioned in them through an easy to use interface. |
DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions (2023.acl-long)
Copied to clipboard
| Challenge: | Modern machine learning relies on datasets to develop and validate research ideas. |
| Approach: | They propose a dataset recommendation system that uses a training set and an evaluation set to help people find relevant datasets. |
| Outcome: | The proposed model finds more relevant search results than existing third-party search engines. |
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have raised concerns about reliability and trustworthiness of the models. |
| Approach: | They analyze 134 papers and introduce a taxonomy of evidence-based text generation with LLMs. |
| Outcome: | The proposed methods highlight open challenges and outline promising directions for future work. |
BioReader: a Retrieval-Enhanced Text-to-Text Transformer for Biomedical Literature (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has equipped language models with the ability to attend over relevant and factual information from non-parametric external sources, drawing a complementary path to architectural scaling. |
| Approach: | They propose a retrieval-enhanced text-to-text model that augments the input prompt by fetching and assembling relevant scientific literature chunks from a neural database centered on PubMed. |
| Outcome: | The proposed model outperforms state-of-the-art models on a broad array of downstream tasks while using up to 3x fewer parameters. |
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)
Copied to clipboard
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John Mariani, Paul Rayson, Keith Suderman
| Challenge: | a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery. |
| Approach: | They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus. |
| Outcome: | The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive. |
TexSmart: A System for Enhanced Natural Language Understanding (2021.acl-demo)
Copied to clipboard
Lemao Liu, Haisong Zhang, Haiyun Jiang, Yangming Li, Enbo Zhao, Kun Xu, Linfeng Song, Suncong Zheng, Botong Zhou, Dick Zhu, Xiao Feng, Tao Chen, Tao Yang, Dong Yu, Feng Zhang, ZhanHui Kang, Shuming Shi
| Challenge: | TexSmart supports fine-grained named entity recognition (NER) Large-scale fine-granular entity types are expected to provide richer semantic information for downstream NLP applications. |
| Approach: | They introduce TexSmart, a text understanding system that supports fine-grained named entity recognition (NER) and enhanced semantic analysis functionalities. |
| Outcome: | The proposed system supports fine-grained named entity recognition (NER) and enhanced semantic analysis functions. |
Named Entity Recognition for Chinese biomedical patents (2020.coling-main)
Copied to clipboard
| Challenge: | Existing attempts to address NER for Chinese biomedical texts have been limited due to the amount of Chinese biomedicine discoveries being patented. |
| Approach: | They train and evaluate Chinese biomedical patents NER models based on BERT . their model is optimized for Chinese bio-patent data and scored an F1 . |
| Outcome: | The proposed model achieves an F1 score of 0.540.15 for Chinese biomedical patent data. |
Rethinking Text-based Protein Understanding: Retrieval or LLM? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have focused on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment. |
| Approach: | They propose a retrieval-enhanced method which significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. |
| Outcome: | The proposed method significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. |
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis (2025.emnlp-demos)
Copied to clipboard
| Challenge: | SciClaims is an interactive web-based system for scientific claim analysis in the biomedical domain. |
| Approach: | They present SciClaims, an interactive web-based system for scientific claim analysis in the biomedical domain. |
| Outcome: | The system extracts factual claims from scientific texts and retrieves evidence from PubMed . it also verifies the validity of each claim using large language models . the system is optimized to run efficiently on a single GPU and is publicly available . |