Automatic Orality Identification in Historical Texts (2020.lrec-1)

Copied to clipboard

Challenge: a set of general linguistic features are used to identify conceptually-oral historical texts . linguists recognize that there is also a lot of variation within discourse modes .
Approach: They propose to use general linguistic features to identify conceptually-oral historical texts . they find they are useful for determining conceptuality of historical data as for modern data .
Outcome: The proposed features are used to identify conceptually-oral historical German texts . the features are useful in determining conceptuality of historical data as they are for modern data .

Similar Papers

Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)

Copied to clipboard

Challenge: Existing methods to improve transcription and indexing quality of Oral History interviews are not available.
Approach: They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks .
Outcome: The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions.
Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)

Copied to clipboard

Challenge: a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities .
Approach: They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs.
Outcome: The proposed archive will be accessible online and provide multifaceted search capabilities.
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.
Automatic Focus Annotation: Bringing Formal Pragmatics Alive in Analyzing the Information Structure of Authentic Data (N18-1)

Copied to clipboard

Challenge: Using focus-background dichotomy, discourse and information structure of sentences are being studied in context.
Approach: They propose to automate the analysis of focus in authentic written data by using a range of lexical, syntactic, and semantic features to achieve an accuracy of 78.1%.
Outcome: The proposed approach achieves 78.1% accuracy for identifying focus in authentic written data.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)

Copied to clipboard

Challenge: a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well.
Approach: They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data .
Outcome: The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP .
Detecting Syntactic Change with Pre-trained Transformer Models (2023.findings-emnlp)

Copied to clipboard

Challenge: a fine-tuned BERT model can distinguish between text from the early 1800s and late 1900s . we use it to identify specific instances of syntactic change and specific words for which a new part of speech was introduced.
Approach: They propose to use a BERT-based model to find syntactic differences between English of the early 1800s and that of the late 1900s.
Outcome: The proposed model can distinguish between English of the early 1800s and that of the late 1900s using only syntactic information.
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
A Diachronic Corpus for Literary Style Analysis (L18-1)

Copied to clipboard

Challenge: Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data .
Approach: They propose a resource for diachronic style analysis in particular the analysis of literary authors over time.
Outcome: The proposed resource can be used to analyze literary authors over time.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations