The Subject Annotations of the Danish Parliament Corpus (2009-2017) - Evaluated with Automatic Multi-label Classification (2022.lrec-1)
Copied to clipboard
| Challenge: | The interest in analysing and automatically processing large amounts of political data has increased in the past decades. |
| Approach: | They address the semi-automatic annotation of subjects in the Danish Parliament Corpus (2009-2017) v.2 and describe multi-label classification experiments to verify the consistency of the subject annotation. |
| Outcome: | The proposed method improves on the baseline classifier, which is a majority classifier. |
Similar Papers
Annotation and Automatic Classification of Aspectual Categories (P19-1)
Copied to clipboard
| Challenge: | Annotated resource for aspectual classification of German verb tokens in context. |
| Approach: | They present a resource for aspectual classification of German verb tokens in their clausal context. |
| Outcome: | The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications. |
A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier (L18-1)
Copied to clipboard
| Challenge: | a Danish FrameNet is a lexicon based on the Danish Thesaurus . it is significantly faster than building a new one from scratch . |
| Approach: | They propose a way to efficiently compile a Danish FrameNet based on the Danish Thesaurus . they present the corresponding corpus annotations of frames and roles and show how this can be used for a semantic frame classifier . |
| Outcome: | The proposed approach is faster than building a lexicon from scratch. |
How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates (2024.lrec-main)
Copied to clipboard
| Challenge: | a new perspective on framing through the lens of speech acts investigates how politicians make use of different pragmatic speech act functions in political debates. |
| Approach: | They propose a new framework for framing through the lens of speech acts and an annotation scheme for political debates. |
| Outcome: | The proposed framework can predict speech acts with an avg. F1 of around 82.0% . the proposed framework is based on a dataset of German parliamentary debates . |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
Developing A Multilabel Corpus for the Quality Assessment of Online Political Talk (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index . |
| Approach: | They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries . |
| Outcome: | The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english . |
Out of the Mouths of MPs: Speaker Attribution in Parliamentary Debates (2024.lrec-main)
Copied to clipboard
| Challenge: | Identifying who says what to whom is an essential prerequisite for analysing human communication. |
| Approach: | They propose a new corpus for speaker attribution in german parliamentary debates . the data includes more than 7,700 manually annotated events of speech, thought and writing . they then apply their model to predict speech events in 20 years of debates and investigate the use of factives in the rhetoric of MPs. |
| Outcome: | The proposed model predicts speech events in 20 years of debates and investigates the use of factives in the rhetoric of MPs. |
DDisCo: A Discourse Coherence Dataset for Danish (2022.lrec-1)
Copied to clipboard
| Challenge: | Discourse coherence models have been developed using randomly shuffled texts instead of highly edited and coherent data. |
| Approach: | They propose to annotate Danish Wikipedia and Reddit for discourse coherence using real-world text instead of artificially incoherent text for training and testing models. |
| Outcome: | The proposed model performs well on annotated texts from the Danish Wikipedia and Reddit dataset. |
Corpus for Automatic Structuring of Legal Documents (2022.lrec-1)
Copied to clipboard
Prathamesh Kalamkar, Aman Tiwari, Astha Agarwal, Saurabh Karn, Smita Gupta, Vivek Raghavan, Ashutosh Modi
| Challenge: | In populous countries, pending legal cases are growing exponentially. |
| Approach: | They propose a corpus of legal judgment documents in English that is annotated with a label coming from a list of pre-defined rhetorical roles. |
| Outcome: | The proposed corpus of legal judgment documents is annotated with a label coming from a list of pre-defined rhetorical roles. |
An Environment for Relational Annotation of Political Debates (P19-3)
Copied to clipboard
| Challenge: | Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities. |
| Approach: | They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors. |
| Outcome: | The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors. |
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions . |
| Approach: | They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models . |
| Outcome: | The proposed model outperforms more complex models on a given dataset. |