HarriGT: A Tool for Linking News to Science (P18-4)

Copied to clipboard

Challenge: citation-based metrics are not as transparent as once thought, says a new study . citation metrics are a key component of measuring scientific impact in society .
Approach: They propose a tool to build corpora of news articles linked to scientific papers . citation-based metrics have catalysed research funding councils' interest in impact .
Outcome: The proposed tool can build corpora of news articles linked to scientific papers . it integrates with 3 large external citation networks to surface relevant examples of scientific literature .

Similar Papers

CitationIE: Leveraging the Citation Graph for Scientific Information Extraction (2021.acl-long)

Copied to clipboard

Challenge: Existing work on scientific information extraction (SciIE) considers extraction solely based on the content of an individual paper, without considering the paper’s place in the broader literature.
Approach: They propose to automate the extraction of key information from scientific documents by leveraging a complementary source: the citation graph of referential links between citing and cited papers.
Outcome: The proposed model improves on a set of English-language scientific documents.
SciLit: A Platform for Joint Scientific Literature Discovery, Summarization and Citation Generation (2023.acl-demo)

Copied to clipboard

Challenge: Scientific writing involves retrieving, summarizing, and citing relevant papers.
Approach: They propose a pipeline that automatically recommends relevant papers, extracts highlights, and suggests a reference sentence as a citation of a paper.
Outcome: The proposed pipeline recommends relevant papers from large databases of hundreds of millions of papers . it provides extractive summaries and abstractively-generated citation sentences . authors question whether it is possible to partly automate this process to reduce cognitive load .
When science journalism meets artificial intelligence : An interactive demonstration (D18-2)

Copied to clipboard

Challenge: Existing tools for automating science journalism do not provide adequate training for AIs to be trained.
Approach: They propose an online tool that generates titles of blog titles by mimicking a human science journalist.
Outcome: The proposed tool generates blog titles by mimicking a human science journalist . it is evaluated using standard metrics to show its viability .
Annotating Research Infrastructure in Scientific Papers: An NLP-driven Approach (2023.acl-industry)

Copied to clipboard

Challenge: a pipeline is used to identify, extract and link research infrastructure used in scientific publications.
Approach: They propose a natural language processing pipeline for the identification, extraction and linking of Research Infrastructure (RI) used in scientific publications.
Outcome: The proposed pipeline can be used to identify, extract and link research infrastructure used in scientific publications.
‘Don’t Get Too Technical with Me’: A Discourse Structure-Based Framework for Automatic Science Journalism (2023.emnlp-main)

Copied to clipboard

Challenge: Science journalism is the production of journalistic content covering scientific topics that are not covered in the scientific literature.
Approach: They propose to use a dataset to generate a scientific paper's tuples, a summary snippet and a novel technical framework to integrate a paper' s discourse structure with its metadata to guide generation.
Outcome: The proposed system outperforms baseline methods in elaborating a content plan meaningful for the target audience, simplifying the information selected, and producing a coherent final report in a layman’s style.
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)

Copied to clipboard

Challenge: Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization.
Approach: They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references.
Outcome: The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)

Copied to clipboard

Challenge: To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English.
Approach: They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English.
Outcome: The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one.
Event Coreference Data (Almost) for Free: Mining Hyperlinks from Online News (2021.emnlp-main)

Copied to clipboard

Challenge: Annotating CDCR data is laborious and expensive, explaining why existing corpora are small and lack domain coverage.
Approach: They use hyperlinks to extract event coreference data from online news articles . they find that models trained on small subsets of HyperCoref are highly competitive .
Outcome: The proposed system frees up CDCR research from costly human-annotated training data and opens up possibilities beyond English.
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches.
Approach: They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries.
Outcome: The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations