Papers with kappa
Gold Standard Annotations for Preposition and Verb Sense with Semantic Role Labels in Adult-Child Interactions (C18-1)
Copied to clipboard
| Challenge: | Existing corpus of child-directed speech augments existing corpus for semantic role labels . sense and number of arguments were open to multiple interpretations due to rapidly changing discourse . |
| Approach: | They propose to augment an existing corpus of child-directed speech to provide supervised learning of semantic role labels. |
| Outcome: | The resulting corpus is a gold standard for supervised learning of semantic role labels in child-directed speech. |
Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability (2021.acl-long)
Copied to clipboard
| Challenge: | Respectable journals typically require reporting quantitative evidence for inter-rater reliability (IRR) of the data. |
| Approach: | They propose to benchmark IRR against baseline measures in a replication dataset and use Cohen's (1960) kappa to measure inter-rater reliability. |
| Outcome: | The proposed framework can be used to measure the quality of crowdsourced datasets. |
TwiUSD: A Benchmark Dataset and Structure-Aware LLM Framework for User Stance Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for political user-level stance detection rely on noisy heuristics or distant supervision. |
| Approach: | They propose a large-scale, expert-annotated benchmark for political user-level stance detection with explicit social network structure that integrates user content and followee signals. |
| Outcome: | The proposed framework outperforms baselines in terms of quality and reliability. |