| Challenge: | citation-based metrics are not as transparent as once thought, says a new study . citation metrics are a key component of measuring scientific impact in society . |
| Approach: | They propose a tool to build corpora of news articles linked to scientific papers . citation-based metrics have catalysed research funding councils' interest in impact . |
| Outcome: | The proposed tool can build corpora of news articles linked to scientific papers . it integrates with 3 large external citation networks to surface relevant examples of scientific literature . |
Similar Papers
CitationIE: Leveraging the Citation Graph for Scientific Information Extraction (2021.acl-long)
Copied to clipboard
| Challenge: | Existing work on scientific information extraction (SciIE) considers extraction solely based on the content of an individual paper, without considering the paper’s place in the broader literature. |
| Approach: | They propose to automate the extraction of key information from scientific documents by leveraging a complementary source: the citation graph of referential links between citing and cited papers. |
| Outcome: | The proposed model improves on a set of English-language scientific documents. |
SciLit: A Platform for Joint Scientific Literature Discovery, Summarization and Citation Generation (2023.acl-demo)
Copied to clipboard
| Challenge: | Scientific writing involves retrieving, summarizing, and citing relevant papers. |
| Approach: | They propose a pipeline that automatically recommends relevant papers, extracts highlights, and suggests a reference sentence as a citation of a paper. |
| Outcome: | The proposed pipeline recommends relevant papers from large databases of hundreds of millions of papers . it provides extractive summaries and abstractively-generated citation sentences . authors question whether it is possible to partly automate this process to reduce cognitive load . |
When science journalism meets artificial intelligence : An interactive demonstration (D18-2)
Copied to clipboard
| Challenge: | Existing tools for automating science journalism do not provide adequate training for AIs to be trained. |
| Approach: | They propose an online tool that generates titles of blog titles by mimicking a human science journalist. |
| Outcome: | The proposed tool generates blog titles by mimicking a human science journalist . it is evaluated using standard metrics to show its viability . |
Annotating Research Infrastructure in Scientific Papers: An NLP-driven Approach (2023.acl-industry)
Copied to clipboard
Seyed Amin Tabatabaei, Georgios Cheirmpos, Marius Doornenbal, Alberto Zigoni, Veronique Moore, Georgios Tsatsaronis
| Challenge: | a pipeline is used to identify, extract and link research infrastructure used in scientific publications. |
| Approach: | They propose a natural language processing pipeline for the identification, extraction and linking of Research Infrastructure (RI) used in scientific publications. |
| Outcome: | The proposed pipeline can be used to identify, extract and link research infrastructure used in scientific publications. |
‘Don’t Get Too Technical with Me’: A Discourse Structure-Based Framework for Automatic Science Journalism (2023.emnlp-main)
Copied to clipboard
| Challenge: | Science journalism is the production of journalistic content covering scientific topics that are not covered in the scientific literature. |
| Approach: | They propose to use a dataset to generate a scientific paper's tuples, a summary snippet and a novel technical framework to integrate a paper' s discourse structure with its metadata to guide generation. |
| Outcome: | The proposed system outperforms baseline methods in elaborating a content plan meaningful for the target audience, simplifying the information selected, and producing a coherent final report in a layman’s style. |
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization. |
| Approach: | They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references. |
| Outcome: | The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)
Copied to clipboard
| Challenge: | To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English. |
| Approach: | They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English. |
| Outcome: | The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one. |
Event Coreference Data (Almost) for Free: Mining Hyperlinks from Online News (2021.emnlp-main)
Copied to clipboard
| Challenge: | Annotating CDCR data is laborious and expensive, explaining why existing corpora are small and lack domain coverage. |
| Approach: | They use hyperlinks to extract event coreference data from online news articles . they find that models trained on small subsets of HyperCoref are highly competitive . |
| Outcome: | The proposed system frees up CDCR research from costly human-annotated training data and opens up possibilities beyond English. |
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches. |
| Approach: | They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries. |
| Outcome: | The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications. |