Challenge: The interest in analysing and automatically processing large amounts of political data has increased in the past decades.
Approach: They address the semi-automatic annotation of subjects in the Danish Parliament Corpus (2009-2017) v.2 and describe multi-label classification experiments to verify the consistency of the subject annotation.
Outcome: The proposed method improves on the baseline classifier, which is a majority classifier.

Similar Papers

Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier (L18-1)

Copied to clipboard

Challenge: a Danish FrameNet is a lexicon based on the Danish Thesaurus . it is significantly faster than building a new one from scratch .
Approach: They propose a way to efficiently compile a Danish FrameNet based on the Danish Thesaurus . they present the corresponding corpus annotations of frames and roles and show how this can be used for a semantic frame classifier .
Outcome: The proposed approach is faster than building a lexicon from scratch.
How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates (2024.lrec-main)

Copied to clipboard

Challenge: a new perspective on framing through the lens of speech acts investigates how politicians make use of different pragmatic speech act functions in political debates.
Approach: They propose a new framework for framing through the lens of speech acts and an annotation scheme for political debates.
Outcome: The proposed framework can predict speech acts with an avg. F1 of around 82.0% . the proposed framework is based on a dataset of German parliamentary debates .
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
Developing A Multilabel Corpus for the Quality Assessment of Online Political Talk (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index .
Approach: They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries .
Outcome: The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english .
Out of the Mouths of MPs: Speaker Attribution in Parliamentary Debates (2024.lrec-main)

Copied to clipboard

Challenge: Identifying who says what to whom is an essential prerequisite for analysing human communication.
Approach: They propose a new corpus for speaker attribution in german parliamentary debates . the data includes more than 7,700 manually annotated events of speech, thought and writing . they then apply their model to predict speech events in 20 years of debates and investigate the use of factives in the rhetoric of MPs.
Outcome: The proposed model predicts speech events in 20 years of debates and investigates the use of factives in the rhetoric of MPs.
DDisCo: A Discourse Coherence Dataset for Danish (2022.lrec-1)

Copied to clipboard

Challenge: Discourse coherence models have been developed using randomly shuffled texts instead of highly edited and coherent data.
Approach: They propose to annotate Danish Wikipedia and Reddit for discourse coherence using real-world text instead of artificially incoherent text for training and testing models.
Outcome: The proposed model performs well on annotated texts from the Danish Wikipedia and Reddit dataset.
Corpus for Automatic Structuring of Legal Documents (2022.lrec-1)

Copied to clipboard

Challenge: In populous countries, pending legal cases are growing exponentially.
Approach: They propose a corpus of legal judgment documents in English that is annotated with a label coming from a list of pre-defined rhetorical roles.
Outcome: The proposed corpus of legal judgment documents is annotated with a label coming from a list of pre-defined rhetorical roles.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)

Copied to clipboard

Challenge: Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions .
Approach: They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models .
Outcome: The proposed model outperforms more complex models on a given dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations