Tracing Syntactic Change in the Scientific Genre: Two Universal Dependency-parsed Diachronic Corpora of Scientific English and German (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on the syntactic development of scientific discourse in English and German. |
| Approach: | They present two comparable diachronic corpora of scientific English and German from the Late Modern Period (17th c.–19th d.) annotated with Universal Dependencies. |
| Outcome: | The presented corpora are comparable to existing studies on grammatical change in English and German . the results show that the pre-processing steps significantly improve parsing accuracy . |
Similar Papers
A Diachronic Corpus for Literary Style Analysis (L18-1)
Copied to clipboard
| Challenge: | Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data . |
| Approach: | They propose a resource for diachronic style analysis in particular the analysis of literary authors over time. |
| Outcome: | The proposed resource can be used to analyze literary authors over time. |
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)
Copied to clipboard
Dimitrios Roussis, Vassilis Papavassiliou, Prokopis Prokopidis, Stelios Piperidis, Vassilis Katsouros
| Challenge: | SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives. |
| Approach: | They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications. |
| Outcome: | The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general. |
Exploring the Effect of Nominal Compound Structure in Scientific Texts on Reading Times of Experts and Novices (2025.acl-srw)
Copied to clipboard
| Challenge: | Using a corpus of eye-tracking data of German native speakers, we find that some compound types are associated with longer reading times. |
| Approach: | They use a corpus containing eye-tracking data of german native speakers reading scientific texts. |
| Outcome: | The authors show that some compound types are associated with longer reading times and that experts may have an advantage while reading in-domain texts, but also while reading out-of-domain. |
The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study (2020.lrec-1)
Copied to clipboard
| Challenge: | a new version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts . |
| Approach: | They present a new version of the Royal Society Corpus, a diachronic corpus of scientific English covering 300+ years of scientific writing. |
| Outcome: | The extended version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts . |
Detecting Syntactic Change with Pre-trained Transformer Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a fine-tuned BERT model can distinguish between text from the early 1800s and late 1900s . we use it to identify specific instances of syntactic change and specific words for which a new part of speech was introduced. |
| Approach: | They propose to use a BERT-based model to find syntactic differences between English of the early 1800s and that of the late 1900s. |
| Outcome: | The proposed model can distinguish between English of the early 1800s and that of the late 1900s using only syntactic information. |
Universal Dependencies: Extensions for Modern and Historical German (2024.lrec-main)
Copied to clipboard
| Challenge: | a new UD treebank is being developed for Middle High German annotations . the annotation scheme is inconsistent with other treebanks for this period . |
| Approach: | They propose to extend the UD scheme for modern and historical German by a range of tokens . they propose to use a treebank that is the first UD treebank for Middle High German . |
| Outcome: | The proposed extensions relate in part to differences between arguments and modifiers . the proposed treebank is the first UD treebank for Middle High German . |
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
The Low Saxon LSDC Dataset at Universal Dependencies (2024.lrec-main)
Copied to clipboard
| Challenge: | Low Saxon is a low-resource language that lacks a common standard . dialectal variation in morphological categories can cause problems . |
| Approach: | They extend the Low Saxon Universal Dependencies dataset to include 8 of the 9 major dialects. |
| Outcome: | The proposed dataset covers the last 200 years and 8 of the 9 major dialects. |
Diachronic word embeddings and semantic shifts: a survey (C18-1)
Copied to clipboard
| Challenge: | Existing methods for tracing time-related semantic shifts with word embedding models lack the cohesion, common terminology and shared practices of more established areas of natural language processing. |
| Approach: | They propose several axes along which these methods can be compared and propose a framework for comparison. |
| Outcome: | The proposed methods are compared with existing methods and outline their main challenges and potential applications. |