| Challenge: | a set of general linguistic features are used to identify conceptually-oral historical texts . linguists recognize that there is also a lot of variation within discourse modes . |
| Approach: | They propose to use general linguistic features to identify conceptually-oral historical texts . they find they are useful for determining conceptuality of historical data as for modern data . |
| Outcome: | The proposed features are used to identify conceptually-oral historical German texts . the features are useful in determining conceptuality of historical data as they are for modern data . |
Similar Papers
Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)
Copied to clipboard
| Challenge: | Existing methods to improve transcription and indexing quality of Oral History interviews are not available. |
| Approach: | They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks . |
| Outcome: | The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions. |
Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)
Copied to clipboard
Zbyněk Zajíc, Lucie Skorkovská, Petr Neduchal, Pavel Ircing, Josef V. Psutka, Marek Hrúz, Aleš Pražák, Daniel Soutner, Jan Švec, Lukáš Bureš, Luděk Müller
| Challenge: | a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities . |
| Approach: | They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs. |
| Outcome: | The proposed archive will be accessible online and provide multifaceted search capabilities. |
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible. |
| Approach: | They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries. |
| Outcome: | The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country. |
Automatic Focus Annotation: Bringing Formal Pragmatics Alive in Analyzing the Information Structure of Authentic Data (N18-1)
Copied to clipboard
| Challenge: | Using focus-background dichotomy, discourse and information structure of sentences are being studied in context. |
| Approach: | They propose to automate the analysis of focus in authentic written data by using a range of lexical, syntactic, and semantic features to achieve an accuracy of 78.1%. |
| Outcome: | The proposed approach achieves 78.1% accuracy for identifying focus in authentic written data. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)
Copied to clipboard
Piroska Lendvai, Maarten van Gompel, Anna Jouravel, Elena Renje, Uwe Reichel, Achim Rabus, Eckhart Arnold
| Challenge: | a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well. |
| Approach: | They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data . |
| Outcome: | The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP . |
Detecting Syntactic Change with Pre-trained Transformer Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a fine-tuned BERT model can distinguish between text from the early 1800s and late 1900s . we use it to identify specific instances of syntactic change and specific words for which a new part of speech was introduced. |
| Approach: | They propose to use a BERT-based model to find syntactic differences between English of the early 1800s and that of the late 1900s. |
| Outcome: | The proposed model can distinguish between English of the early 1800s and that of the late 1900s using only syntactic information. |
Annotation and Automatic Classification of Aspectual Categories (P19-1)
Copied to clipboard
| Challenge: | Annotated resource for aspectual classification of German verb tokens in context. |
| Approach: | They present a resource for aspectual classification of German verb tokens in their clausal context. |
| Outcome: | The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
A Diachronic Corpus for Literary Style Analysis (L18-1)
Copied to clipboard
| Challenge: | Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data . |
| Approach: | They propose a resource for diachronic style analysis in particular the analysis of literary authors over time. |
| Outcome: | The proposed resource can be used to analyze literary authors over time. |