Automatic Identification of Research Fields in Scientific Papers (L18-1)

Copied to clipboard

Challenge: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Approach: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Outcome: The proposed method will help scientists identify geographical territories areas from scientific papers available in digital versions within and outside the ISTEX library.

Similar Papers

Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions (2024.lrec-main)

Copied to clipboard

Challenge: SciDMT is an enhanced and expanded corpus for scientific mention detection . existing corpora are limited by their small volume and entity linking capabilities .
Approach: They propose to enhance SciDMT, an annotated scientific corpus for scientific mention detection.
Outcome: The proposed corpus is the largest for scientific entity mention detection . it is based on deep learning architectures like SciBERT and GPT-3.5 .
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)

Copied to clipboard

Challenge: Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization.
Approach: They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references.
Outcome: The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section.
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)

Copied to clipboard

Challenge: To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English.
Approach: They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English.
Outcome: The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one.
GeospaCy: A tool for extraction and geographical referencing of spatial expressions in textual data (2024.eacl-demo)

Copied to clipboard

Challenge: Spatial information in text enables to understand the geographical context and relationships within text for location-sensitive applications.
Approach: They propose to use spatial information extracted from textual data to perform geoparsing and geocoding tasks.
Outcome: The GeospaCy software tool is designed for the extraction and georeferencing of spatial information present in textual data.
Annotating Research Infrastructure in Scientific Papers: An NLP-driven Approach (2023.acl-industry)

Copied to clipboard

Challenge: a pipeline is used to identify, extract and link research infrastructure used in scientific publications.
Approach: They propose a natural language processing pipeline for the identification, extraction and linking of Research Infrastructure (RI) used in scientific publications.
Outcome: The proposed pipeline can be used to identify, extract and link research infrastructure used in scientific publications.
A New Public Corpus for Clinical Section Identification: MedSecId (2022.coling-1)

Copied to clipboard

Challenge: a study aims to segment sections of clinical medical domain documentation . section identification is a process by which sections are demarcated and labeled .
Approach: They use a set of 2,002 fully annotated medical notes from the MIMIC-III to segment sections in clinical medical domain documentation.
Outcome: The proposed model shows that medical concepts are related across sections using principal component analysis.
A Survey on Open Information Extraction (C18-1)

Copied to clipboard

Challenge: Existing approaches to open information extraction (Open IE) focus on narrow, well-defined requests over a predefined set of target relations on small, homogeneous corpora.
Approach: They propose to use unsupervised methods to extract all types of relations found in text . they propose to implement a system that can be automated to detect possible relations .
Outcome: The proposed approaches have been compared with existing methods and are based on the results of a literature review.
A Scientific Information Extraction Dataset for Nature Inspired Engineering (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to extract relevant biological information from scientific literature are difficult and require domain-specific knowledge.
Approach: They describe a dataset of 1,500 manually-annotated sentences that express domain-independent relations between central concepts in a scientific biology text.
Outcome: The proposed dataset allows for training and evaluation of Relation Extraction algorithms that aim for coarse-grained typing of scientific biological documents, enabling a high-level filter for engineers.
FigEx: Aligned Extraction of Scientific Figures and Captions (2025.findings-emnlp)

Copied to clipboard

Challenge: FigEx is a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers.
Approach: They propose a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers.
Outcome: The proposed model improves subfigure detection APb over Grounding DINO by 0.023 and boosts caption separation BLEU over Llama-2-13B by 0.465.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations