PDF-to-Text Reanalysis for Linguistic Data Mining (L18-1)

Copied to clipboard

Challenge: In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem.
Approach: They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining.
Outcome: The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes.

Similar Papers

PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles (L18-1)

Copied to clipboard

Challenge: Existing tools to extract structured textual content from PDFs are essential to enable scientific text mining.
Approach: They propose a PDF-to-XML textual content extraction tool that extracts structured textual contents from scientific articles in PDF format.
Outcome: The proposed tool extracts structured textual content from scientific articles in PDF format while preserving both the textual contents and layout details of the input PDF document.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the Web (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial explores the commonalities in the challenges and solutions developed to address information extraction from the World Wide Web.
Approach: This tutorial examines methods for extracting information from the World Wide Web . it explores the commonalities in the challenges and solutions developed to address these different forms of text .
Outcome: This paper examines the commonalities in the challenges and solutions developed to address the World Wide Web.
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)

Copied to clipboard

Challenge: To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English.
Approach: They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English.
Outcome: The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one.
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)

Copied to clipboard

Challenge: Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures.
Approach: They propose to use graph rewriting to describe rule-based transformations on linguistic structures.
Outcome: The proposed tools can be used to compute rule-based transformations on linguistic structures.
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect (2022.coling-1)

Copied to clipboard

Challenge: text-to-SQL is a language processing and database-based language processing (NLP) task is to convert natural utterances into SQL queries and its practical application is to build natural language interfaces to database systems.
Approach: They propose to conduct a systematic survey of text-to-SQL to examine the challenges and potential future directions.
Outcome: The proposed system converts natural utterances into SQL queries and is a representative task in semantic parsing.
An efficient method for Natural Language Querying on Structured Data (2023.acl-industry)

Copied to clipboard

Challenge: a new approach to NLQ on structured data is based on text-to-SQL type semantic parsing . domain classification, domain classification and domain classification are the main tasks . semantic parsed queries are less common when information is in structured form .
Approach: They propose an efficient and reliable approach to natural language Querying on databases . they use domain classification, domain classification and slot/entity extraction to query a DB .
Outcome: The proposed approach simplifies the NLQ on structured data problem to the following "bread and butter" tasks.
Profiling-UD: a Tool for Linguistic Profiling of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Profiling–UD is a text analysis tool that can be used to characterize language variation from different perspectives.
Approach: They introduce Profiling–UD, a text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives.
Outcome: The proposed tool is specifically designed to be multilingual since it is based on the Universal Dependencies framework.
Measurement Extraction with Natural Language Processing: A Review (2022.findings-emnlp)

Copied to clipboard

Challenge: Information extraction (IE) is a task in natural language processing that extracts information from documents.
Approach: They describe different approaches to measurement extraction and outline challenges posed by this task.
Outcome: The proposed methods are compared with the literature on the extraction of quantitative data from documents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations