Challenge: Identifying who says what to whom is an essential prerequisite for analysing human communication.
Approach: They propose a new corpus for speaker attribution in german parliamentary debates . the data includes more than 7,700 manually annotated events of speech, thought and writing . they then apply their model to predict speech events in 20 years of debates and investigate the use of factives in the rhetoric of MPs.
Outcome: The proposed model predicts speech events in 20 years of debates and investigates the use of factives in the rhetoric of MPs.

Similar Papers

How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates (2024.lrec-main)

Copied to clipboard

Challenge: a new perspective on framing through the lens of speech acts investigates how politicians make use of different pragmatic speech act functions in political debates.
Approach: They propose a new framework for framing through the lens of speech acts and an annotation scheme for political debates.
Outcome: The proposed framework can predict speech acts with an avg. F1 of around 82.0% . the proposed framework is based on a dataset of German parliamentary debates .
An Attribution Relations Corpus for Political News (L18-1)

Copied to clipboard

Challenge: Existing resources for recognizing attributions in context are limited in size and completeness.
Approach: They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction .
Outcome: The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date.
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
A corpus of German political speeches from the 21st century (L18-1)

Copied to clipboard

Challenge: a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level .
Approach: a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level .
Outcome: The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts .
AttributionBench: How Hard is Automatic Attribution Evaluation? (2024.findings-acl)

Copied to clipboard

Challenge: generative search engines enhance the reliability of large language model responses by providing cited evidence.
Approach: They propose to use a benchmark to evaluate whether a large language model supports the generated responses or not .
Outcome: The proposed benchmark shows that even a fine-tuned GPT-3.5 only achieves around 80% macro-F1 under a binary classification formulation.
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
What does the sea say to the shore? A BERT based DST style approach for speaker to dialogue attribution in novels (2022.acl-long)

Copied to clipboard

Challenge: Existing systems can perform the first two tasks accurately, but attributing characters to direct speech is a challenging problem due to the narrator’s lack of explicit character mentions and the frequent use of nominal and pronominal coreference when such explicit mentions are made.
Approach: They propose a pipeline to extract characters and link them to their direct-speech utterances by using a novel's list of characters and a list of attributing them to the speaking characters.
Outcome: The proposed pipeline improves state-of-the-art models by 50% in F1-score compared with existing models .
Who Sides with Whom? Towards Computational Construction of Discourse Networks for Political Debates (P19-1)

Copied to clipboard

Challenge: a vision of computational construction of discourse networks from newspaper reports is essential for understanding democratic political decision making.
Approach: They propose to use a requirements analysis and an annotated pilot corpus of migration claims to build a computationally-based model of political debates from newspaper reports.
Outcome: The proposed framework could be scaled up to a large scale and be useful for political scientists.
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)

Copied to clipboard

Challenge: specialized corpora on a variety of topics are available online, but online corporates are scarce.
Approach: They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists .
Outcome: The proposed corpus contains more than six million speeches in English and Chinese and is available for free online.
Who’s in, who’s out? Predicting the Inclusiveness or Exclusiveness of Personal Pronouns in Parliamentary Debates (2022.lrec-1)

Copied to clipboard

Challenge: clusivity properties of personal pronouns are captured in context, including/excluding audience and/or non-speech act participants.
Approach: They propose a compositional annotation scheme to capture the clusivity properties of personal pronouns in context, which is their ability to construct and manage in-groups and out-group.
Outcome: The proposed schema achieves high inter-annotator agreement with a Cohen’s in the range of 89.7-93.2 and a percentage agreement of > 96%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations