Papers by Ryan Georgi

1 papers
PDF-to-Text Reanalysis for Linguistic Data Mining (L18-1)

Copied to clipboard

Challenge: In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem.
Approach: They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining.
Outcome: The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations