| Challenge: | Scholarly document processing (SDP) is a powerful tool for researchers to process knowledge stored in research papers. |
| Approach: | They propose a system that allows researchers to search papers published in ACL conferences with metadata and text embeddings. |
| Outcome: | The proposed system is simple and efficient to reduce maintenance and financial costs and is extensible to support open development and transparency. |
Similar Papers
GenGO Ultra: an LLM-powered ACL Paper Explorer (2025.acl-demo)
Copied to clipboard
| Challenge: | The main repository of natural language processing (NLP) has grown its number of stored papers by 70% from 2019 to 2023. |
| Approach: | They propose an extension to GenGO Ultra which exploits large language models to dynamically generate responses grounded by published papers. |
| Outcome: | The proposed system exploits large language models to generate responses grounded by published papers and performs multi-granularity experiments. |
NLP Scholar: An Interactive Visual Explorer for Natural Language Processing Literature (2020.acl-demos)
Copied to clipboard
| Challenge: | aCL Anthology and Google Scholar provide a single dataset of NLP papers and their meta-information . authors describe interactive visualizations that present various aspects of the data . |
| Approach: | They propose to use citation data from the ACL Anthology and Google Scholar to create a unified dataset of NLP papers and their meta-information. |
| Outcome: | The proposed dataset includes papers published in the area of their interest and by specified authors. |
The ACL OCL Corpus: Advancing Open Science in Computational Linguistics (2023.emnlp-main)
Copied to clipboard
| Challenge: | ACL OCL is a scholarly corpus derived from the ACL Anthology . it provides metadata, PDF files, citation graphs and additional structured full texts . |
| Approach: | They present ACL OCL, a scholarly corpus derived from the ACL Anthology . it integrates metadata, PDF files, citation graphs and additional structured full texts . they highlight how it applies to observe trends in computational linguistics . |
| Outcome: | The ACL OCL spans seven decades and contains 73,285 papers . the scholarly corpus is based on the ACL Anthology and is available from HuggingFace . |
ACL-rlg: A Dataset for Reading List Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing tools for searching the literature return an overwhelming number of results, making familiarization process daunting and inefficient. |
| Approach: | They propose to use ACL-rlg as the largest open expert-annotated reading list dataset to help researchers navigate key literature. |
| Outcome: | The proposed dataset outperforms existing search engines and indexing methods and shows signs of data contamination. |
Automatic Annotation of Semantic Term Types in the Complete ACL Anthology Reference Corpus (L18-1)
Copied to clipboard
| Challenge: | a recent increase in quantitative studies of scientific text collections has led to a significant increase in the use of semantic labeling techniques. |
| Approach: | They propose to use semantic class labels to enhance a well-known resource . they use semantic labels to assign semantic class labeling to technical terms . |
| Outcome: | The proposed approach enhances the ACL Anthology Reference Corpus with semantic class labels for 20,000 technical terms . the goal is to use this information as one feature in the profiling of scientific papers, communities, and disciplines. |
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)
Copied to clipboard
| Challenge: | Summarization is the task of shortening a text while preserving the most important information it contains. |
| Approach: | They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution. |
| Outcome: | The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work. |
A Survey of AMR Applications (2024.emnlp-main)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a semantic representation that takes the form of a rooted, directed graph. |
| Approach: | They analyze more than 100 papers which use Abstract Meaning Representation (AMR) they highlight the range of applications for which AMR has been harnessed and techniques for incorporating it . they also highlight broader AMR engineering patterns and outline areas of future work that seem ripe for AMR incorporation. |
| Outcome: | The results highlight the range of applications for which AMR has been harnessed and the techniques for incorporating it into those applications. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)
Copied to clipboard
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John Mariani, Paul Rayson, Keith Suderman
| Challenge: | a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery. |
| Approach: | They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus. |
| Outcome: | The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive. |
FoRC4CL: A Fine-grained Field of Research Classification and Annotated Dataset of NLP Articles (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing systems for categorising scientific knowledge are lacking in many digital repositories. |
| Approach: | They propose to classify papers in the ACL Anthology using a hierarchical taxonomy of core CL/NLP topics and sub-topics. |
| Outcome: | The proposed corpus of 1,500 ACL Anthology publications is annotated with their main contributions using a hierarchical taxonomy of core CL/NLP topics and sub-topics. |