Papers by Adrien Barbaresi
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction (2021.acl-demo)
Copied to clipboard
| Challenge: | Existing tools for text extraction and web corpus construction are not enough to extract and pre-process web data to meet scientific expectations with respect to text quality. |
| Approach: | They propose a text discovery and extraction tool published under open-source license that allows for main text, comments and metadata extraction while also providing building blocks for web crawling tasks. |
| Outcome: | The proposed tool performs significantly better than other open-source solutions on real-world data and in external benchmarks. |
A database of German definitory contexts from selected web sources (L18-1)
Copied to clipboard
| Challenge: | a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts. |
| Approach: | They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction . |
| Outcome: | The proposed method is based on a web corpus and a robust pattern-based extraction method. |
A corpus of German political speeches from the 21st century (L18-1)
Copied to clipboard
| Challenge: | a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level . |
| Approach: | a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level . |
| Outcome: | The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts . |