ACT2: A multi-disciplinary semi-structured dataset for importance and purpose classification of citations (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for classifying citations rely on bibliometric measures to consider the semantics of citation. |
| Approach: | They propose to use a Citation Context Classification (3C) shared task dataset to classify citations according to their purpose and importance. |
| Outcome: | The proposed model can be used to link research works to graphs and enable efficient knowledge discovery. |
Similar Papers
Dynamic Context Extraction for Citation Classification (2022.aacl-main)
Copied to clipboard
| Challenge: | Prior studies have focused on the application of fixed-size contiguous citation contexts or manually curated citation contextual contexts. |
| Approach: | They propose an automated unsupervised approach for the selection of a dynamic-size and potentially non-contiguous citation context based on transformer-based document representations and embedding similarities. |
| Outcome: | The proposed model improves on the domain-specific and multi-disciplinary datasets, irrespective of the dataset's domain. |
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)
Copied to clipboard
| Challenge: | Academic papers are an increasingly important textual domain for natural language processing (NLP) research. |
| Approach: | They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text. |
| Outcome: | The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers. |
Towards a Unified Framework for Reference Retrieval and Related Work Generation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for related work generation use human-annotated references as information sources. |
| Approach: | They propose a model which combines reference retrieval and related work generation processes in a unified framework based on the large language model. |
| Outcome: | The proposed model outperforms the state-of-the-art models on two wide-applied datasets. |
GLiNER2: Schema-Driven Multi-Task Learning for Structured Information Extraction (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing solutions for information extraction (IE) require specialized models for different tasks or require expensive large language models. |
| Approach: | They propose a framework that enhances the original GLiNER architecture to support named entity recognition, text classification, and hierarchical structured data extraction within a single efficient model. |
| Outcome: | The proposed framework improves performance across diverse IE tasks and accessibility compared to LLM-based alternatives. |
A Context-based Framework for Modeling the Role and Function of On-line Resource Citations in Scientific Literature (D19-1)
Copied to clipboard
| Challenge: | Existing academic search engines cannot detect relevant papers where a resource is mentioned. |
| Approach: | They propose a framework to model the role and function of on-line resource citations . they construct a dataset SciRes, which includes 3,088 manually annotated resource contexts based on a multi-task framework . |
| Outcome: | The proposed model achieves the best results on both the classification task and recommendation task. |
Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification (2024.findings-emnlp)
Copied to clipboard
Zeren Shui, Petros Karypis, Daniel Karls, Mingjian Wen, Saurav Manchanda, Ellad Tadmor, George Karypis
| Challenge: | Prior research has shown that pretrained language models (PLMs) can achieve state-of-the-art performance on CIC benchmarks. |
| Approach: | They propose a multi-task learning framework that fine-tunes pretrained language models on a dataset of primary interest together with multiple auxiliary CIC datasets to take advantage of additional supervision signals. |
| Outcome: | The proposed framework outperforms current state-of-the-art models on small datasets while aligning with the best-performing model on a large dataset. |
A High-Quality Gold Standard for Citation-based Tasks (L18-1)
Copied to clipboard
| Challenge: | Citation recommendation tasks involve recommending citations within their specific contexts. |
| Approach: | They propose to use arXiv.org's citation-dependent evaluation data set to evaluate citations . their data set is characterized by the fact that it exhibits almost zero noise in its extracted content . |
| Outcome: | The proposed data set exhibits almost zero noise in extracted content and all citations are linked to their correct publications. |
FoRC4CL: A Fine-grained Field of Research Classification and Annotated Dataset of NLP Articles (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing systems for categorising scientific knowledge are lacking in many digital repositories. |
| Approach: | They propose to classify papers in the ACL Anthology using a hierarchical taxonomy of core CL/NLP topics and sub-topics. |
| Outcome: | The proposed corpus of 1,500 ACL Anthology publications is annotated with their main contributions using a hierarchical taxonomy of core CL/NLP topics and sub-topics. |
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)
Copied to clipboard
| Challenge: | The first workshop on crowdsourcing for NLP is open to all . |
| Approach: | The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks. |
| Outcome: | The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data . |
Word Attribute Prediction Enhanced by Lexical Entailment Tasks (2020.lrec-1)
Copied to clipboard
| Challenge: | a semantic attribute is associated with a designated dimension in attribute-based vector representations . semantic attributes are created by psychological experimental settings involving human annotators . a conceptual attribute of a concept dictates a specific semantic aspect of the concept . |
| Approach: | They propose a two-stage neural network architecture that fine-tunes attribute representations by employing supervised entailment tasks. |
| Outcome: | The proposed method improves performance of semantic/visual similarity/relatedness evaluation tasks. |