Challenge: a recent study has focused on the syntactic development of scientific discourse in English and German.
Approach: They present two comparable diachronic corpora of scientific English and German from the Late Modern Period (17th c.–19th d.) annotated with Universal Dependencies.
Outcome: The presented corpora are comparable to existing studies on grammatical change in English and German . the results show that the pre-processing steps significantly improve parsing accuracy .

Similar Papers

A Diachronic Corpus for Literary Style Analysis (L18-1)

Copied to clipboard

Challenge: Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data .
Approach: They propose a resource for diachronic style analysis in particular the analysis of literary authors over time.
Outcome: The proposed resource can be used to analyze literary authors over time.
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)

Copied to clipboard

Challenge: SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives.
Approach: They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications.
Outcome: The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general.
Exploring the Effect of Nominal Compound Structure in Scientific Texts on Reading Times of Experts and Novices (2025.acl-srw)

Copied to clipboard

Challenge: Using a corpus of eye-tracking data of German native speakers, we find that some compound types are associated with longer reading times.
Approach: They use a corpus containing eye-tracking data of german native speakers reading scientific texts.
Outcome: The authors show that some compound types are associated with longer reading times and that experts may have an advantage while reading in-domain texts, but also while reading out-of-domain.
The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study (2020.lrec-1)

Copied to clipboard

Challenge: a new version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts .
Approach: They present a new version of the Royal Society Corpus, a diachronic corpus of scientific English covering 300+ years of scientific writing.
Outcome: The extended version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts .
Detecting Syntactic Change with Pre-trained Transformer Models (2023.findings-emnlp)

Copied to clipboard

Challenge: a fine-tuned BERT model can distinguish between text from the early 1800s and late 1900s . we use it to identify specific instances of syntactic change and specific words for which a new part of speech was introduced.
Approach: They propose to use a BERT-based model to find syntactic differences between English of the early 1800s and that of the late 1900s.
Outcome: The proposed model can distinguish between English of the early 1800s and that of the late 1900s using only syntactic information.
Universal Dependencies: Extensions for Modern and Historical German (2024.lrec-main)

Copied to clipboard

Challenge: a new UD treebank is being developed for Middle High German annotations . the annotation scheme is inconsistent with other treebanks for this period .
Approach: They propose to extend the UD scheme for modern and historical German by a range of tokens . they propose to use a treebank that is the first UD treebank for Middle High German .
Outcome: The proposed extensions relate in part to differences between arguments and modifiers . the proposed treebank is the first UD treebank for Middle High German .
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)

Copied to clipboard

Challenge: outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels .
Approach: They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words .
Outcome: The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
The Low Saxon LSDC Dataset at Universal Dependencies (2024.lrec-main)

Copied to clipboard

Challenge: Low Saxon is a low-resource language that lacks a common standard . dialectal variation in morphological categories can cause problems .
Approach: They extend the Low Saxon Universal Dependencies dataset to include 8 of the 9 major dialects.
Outcome: The proposed dataset covers the last 200 years and 8 of the 9 major dialects.
Diachronic word embeddings and semantic shifts: a survey (C18-1)

Copied to clipboard

Challenge: Existing methods for tracing time-related semantic shifts with word embedding models lack the cohesion, common terminology and shared practices of more established areas of natural language processing.
Approach: They propose several axes along which these methods can be compared and propose a framework for comparison.
Outcome: The proposed methods are compared with existing methods and outline their main challenges and potential applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations