| Challenge: | In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem. |
| Approach: | They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining. |
| Outcome: | The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes. |
Similar Papers
PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles (L18-1)
Copied to clipboard
| Challenge: | Existing tools to extract structured textual content from PDFs are essential to enable scientific text mining. |
| Approach: | They propose a PDF-to-XML textual content extraction tool that extracts structured textual contents from scientific articles in PDF format. |
| Outcome: | The proposed tool extracts structured textual content from scientific articles in PDF format while preserving both the textual contents and layout details of the input PDF document. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022.tacl-1)
Copied to clipboard
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofetoluwa Adeyemi
| Challenge: | Lower-resource corpora have systematic issues, including mislabeled or nonstandard/ambiguous language codes. |
| Approach: | They manually audit the quality of 205 language-specific corpora released with five major public datasets. |
| Outcome: | The results show that lower-resource corpora have systematic issues even for non-proficient speakers. |
Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the Web (2020.acl-tutorials)
Copied to clipboard
| Challenge: | a tutorial explores the commonalities in the challenges and solutions developed to address information extraction from the World Wide Web. |
| Approach: | This tutorial examines methods for extracting information from the World Wide Web . it explores the commonalities in the challenges and solutions developed to address these different forms of text . |
| Outcome: | This paper examines the commonalities in the challenges and solutions developed to address the World Wide Web. |
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)
Copied to clipboard
| Challenge: | To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English. |
| Approach: | They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English. |
| Outcome: | The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one. |
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)
Copied to clipboard
| Challenge: | Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures. |
| Approach: | They propose to use graph rewriting to describe rule-based transformations on linguistic structures. |
| Outcome: | The proposed tools can be used to compute rule-based transformations on linguistic structures. |
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect (2022.coling-1)
Copied to clipboard
| Challenge: | text-to-SQL is a language processing and database-based language processing (NLP) task is to convert natural utterances into SQL queries and its practical application is to build natural language interfaces to database systems. |
| Approach: | They propose to conduct a systematic survey of text-to-SQL to examine the challenges and potential future directions. |
| Outcome: | The proposed system converts natural utterances into SQL queries and is a representative task in semantic parsing. |
An efficient method for Natural Language Querying on Structured Data (2023.acl-industry)
Copied to clipboard
| Challenge: | a new approach to NLQ on structured data is based on text-to-SQL type semantic parsing . domain classification, domain classification and domain classification are the main tasks . semantic parsed queries are less common when information is in structured form . |
| Approach: | They propose an efficient and reliable approach to natural language Querying on databases . they use domain classification, domain classification and slot/entity extraction to query a DB . |
| Outcome: | The proposed approach simplifies the NLQ on structured data problem to the following "bread and butter" tasks. |
Profiling-UD: a Tool for Linguistic Profiling of Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | Profiling–UD is a text analysis tool that can be used to characterize language variation from different perspectives. |
| Approach: | They introduce Profiling–UD, a text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives. |
| Outcome: | The proposed tool is specifically designed to be multilingual since it is based on the Universal Dependencies framework. |
Measurement Extraction with Natural Language Processing: A Review (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Information extraction (IE) is a task in natural language processing that extracts information from documents. |
| Approach: | They describe different approaches to measurement extraction and outline challenges posed by this task. |
| Outcome: | The proposed methods are compared with the literature on the extraction of quantitative data from documents. |