A Study in Contradiction: Data and Annotation for AIDA Focusing on Informational Conflict in Russia-Ukraine Relations (2022.lrec-1)
Copied to clipboard
| Challenge: | This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives (AIDA) program . AIDA systems must extract entities, events, and relations from multimedia documents, aggregate that information across documents and languages, and produce multiple “hypotheses” about what has happened. |
| Approach: | This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives program . the program aims to develop language technology that can help humans manage large volumes of conflicting information . |
| Outcome: | The proposed corpus focuses on the domain of Russia-Ukraine relations and contains source data in English, Russian and Ukrainian . it is designed to support the development and evaluation of systems that extract entities, events, and relations from individual multimedia documents, aggregate the information across documents and languages, and produce multiple “hypotheses” about what has happened. |
Similar Papers
Schema Learning Corpus: Data and Annotation Focused on Complex Events (2024.lrec-main)
Copied to clipboard
| Challenge: | The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data. |
| Approach: | The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian. |
| Outcome: | The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set. |
Conflicts in Texts: Data, Implications and Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Conflicts in data could reflect complexity of situations, changes that need to be explained and dealt with, difficulties in data annotation, and mistakes in generated outputs. |
| Approach: | This survey categorizes conflicting information into three key areas . they identify the areas where conflicting data can be ignored and undermine models' reliability and trustworthiness. |
| Outcome: | The findings highlight key challenges and future directions for developing conflict-aware NLP systems that can reason over and reconcile conflicting information more effectively. |
Ukrainian Resilience: A Dataset for Detection of Help-Seeking Signals Amidst the Chaos of War (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new dataset is being developed to categorize posts that show distress or urgency . the dataset could improve humanitarian efforts, allowing for quicker and more targeted help . |
| Approach: | They propose a dataset that brings together social media posts in the Ukrainian language for the detection of help-seeking posts in times of war. |
| Outcome: | The proposed dataset can be used to improve humanitarian efforts . it can be compared with existing datasets and achieve an accuracy of 81.15% . |
How Diplomats Dispute: The UN Security Council Conflict Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Until now, there has been little work on how to formalize conflicts in a diplomatic setting. |
| Approach: | They present a corpus of 87 UNSC speeches that are annotated for conflicts and demonstrate the difficulty when dealing with diplomatic language. |
| Outcome: | The proposed method demonstrates that diplomatic language is complex and often implicit along various dimensions. |
An Environment for Relational Annotation of Political Debates (P19-3)
Copied to clipboard
| Challenge: | Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities. |
| Approach: | They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors. |
| Outcome: | The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors. |
Semi-automatically Annotated Learner Corpus for Russian (2022.lrec-1)
Copied to clipboard
| Challenge: | Revita Learner Corpus is a semi-automatically annotated learner corpus for Russian . it is used for research in second language acquisition and foreign language teaching . |
| Approach: | They propose a semi-automatically annotated learner corpus for Russian that detects errors automatically and annotates errors by type. |
| Outcome: | The proposed corpus detects errors automatically and is annotated by type . the data is made public and the process is much cheaper and faster . |
Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory information (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive capabilities in retrieving and analyzing complex information, but their reliability in conflicting contexts remains poorly understood. |
| Approach: | They propose an adversarial extension of the Needle-in-a-Haystack framework in which three mutually exclusive “needles” are embedded within long documents. |
| Outcome: | The proposed framework highlights critical limitations in the robustness of current LLMs—including commercial systems—to contradiction. |
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing datasets and models fail to address the complexities of multilingual data, authors say . detection of radical content on online platforms has become an increasingly pressing concern . |
| Approach: | They propose a publicly available multilingual dataset annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic. |
| Outcome: | The proposed dataset is annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic. |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
AnnoCTR: A Dataset for Detecting and Linking Entities, Tactics, and Techniques in Cyber Threat Reports (2024.lrec-main)
Copied to clipboard
Lukas Lange, Marc Müller, Ghazaleh Haratinezhad Torbati, Dragan Milchevski, Patrick Grau, Subhash Chandra Pujari, Annemarie Friedrich
| Challenge: | Abstract: Natural language processing can help with managing large amounts of unstructured information. |
| Approach: | They propose to annotate a CC-BY-SA-licensed dataset of cyber threat reports . they use named entities, temporal expressions, and cybersecurity-specific concepts . |
| Outcome: | The proposed dataset annotates reports with named entities, temporal expressions, and cybersecurity-specific concepts including implicitly mentioned techniques and tactics. |