Eric Kergosien, Amin Farvardin, Maguelonne Teisseire, Marie-Noëlle Bessagnet, Joachim Schöpfel, Stéphane Chaudiron, Bernard Jacquemin, Annig Lacayrelle, Mathieu Roche, Christian Sallaberry, Jean Philippe Tonneau
| Challenge: | TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. |
| Approach: | TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. |
| Outcome: | The proposed method will help scientists identify geographical territories areas from scientific papers available in digital versions within and outside the ISTEX library. |
Similar Papers
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions (2024.lrec-main)
Copied to clipboard
| Challenge: | SciDMT is an enhanced and expanded corpus for scientific mention detection . existing corpora are limited by their small volume and entity linking capabilities . |
| Approach: | They propose to enhance SciDMT, an annotated scientific corpus for scientific mention detection. |
| Outcome: | The proposed corpus is the largest for scientific entity mention detection . it is based on deep learning architectures like SciBERT and GPT-3.5 . |
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization. |
| Approach: | They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references. |
| Outcome: | The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section. |
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)
Copied to clipboard
| Challenge: | To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English. |
| Approach: | They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English. |
| Outcome: | The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one. |
GeospaCy: A tool for extraction and geographical referencing of spatial expressions in textual data (2024.eacl-demo)
Copied to clipboard
| Challenge: | Spatial information in text enables to understand the geographical context and relationships within text for location-sensitive applications. |
| Approach: | They propose to use spatial information extracted from textual data to perform geoparsing and geocoding tasks. |
| Outcome: | The GeospaCy software tool is designed for the extraction and georeferencing of spatial information present in textual data. |
Annotating Research Infrastructure in Scientific Papers: An NLP-driven Approach (2023.acl-industry)
Copied to clipboard
Seyed Amin Tabatabaei, Georgios Cheirmpos, Marius Doornenbal, Alberto Zigoni, Veronique Moore, Georgios Tsatsaronis
| Challenge: | a pipeline is used to identify, extract and link research infrastructure used in scientific publications. |
| Approach: | They propose a natural language processing pipeline for the identification, extraction and linking of Research Infrastructure (RI) used in scientific publications. |
| Outcome: | The proposed pipeline can be used to identify, extract and link research infrastructure used in scientific publications. |
A New Public Corpus for Clinical Section Identification: MedSecId (2022.coling-1)
Copied to clipboard
| Challenge: | a study aims to segment sections of clinical medical domain documentation . section identification is a process by which sections are demarcated and labeled . |
| Approach: | They use a set of 2,002 fully annotated medical notes from the MIMIC-III to segment sections in clinical medical domain documentation. |
| Outcome: | The proposed model shows that medical concepts are related across sections using principal component analysis. |
A Survey on Open Information Extraction (C18-1)
Copied to clipboard
| Challenge: | Existing approaches to open information extraction (Open IE) focus on narrow, well-defined requests over a predefined set of target relations on small, homogeneous corpora. |
| Approach: | They propose to use unsupervised methods to extract all types of relations found in text . they propose to implement a system that can be automated to detect possible relations . |
| Outcome: | The proposed approaches have been compared with existing methods and are based on the results of a literature review. |
A Scientific Information Extraction Dataset for Nature Inspired Engineering (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to extract relevant biological information from scientific literature are difficult and require domain-specific knowledge. |
| Approach: | They describe a dataset of 1,500 manually-annotated sentences that express domain-independent relations between central concepts in a scientific biology text. |
| Outcome: | The proposed dataset allows for training and evaluation of Relation Extraction algorithms that aim for coarse-grained typing of scientific biological documents, enabling a high-level filter for engineers. |
FigEx: Aligned Extraction of Scientific Figures and Captions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | FigEx is a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Approach: | They propose a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Outcome: | The proposed model improves subfigure detection APb over Grounding DINO by 0.023 and boosts caption separation BLEU over Llama-2-13B by 0.465. |