Papers by Ann Bies
Schema Learning Corpus: Data and Annotation Focused on Complex Events (2024.lrec-main)
Copied to clipboard
| Challenge: | The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data. |
| Approach: | The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian. |
| Outcome: | The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set. |
Simple Semantic Annotation and Situation Frames: Two Approaches to Basic Text Understanding in LORELEI (L18-1)
Copied to clipboard
| Challenge: | Existing annotations for low resource languages are under-resourced for human language technology, but lack of resources does not correlate with lack of need for such technologies. |
| Approach: | They propose two types of semantic annotation for the DARPA Low Resource Languages for Emerging Incidents program: Simple Semantic Annotation (SSA) and Situation Frames (SF). |
| Outcome: | The proposed approaches are aimed at labeling basic semantic information relevant to humanitarian aid and disaster relief scenarios. |
Cross-Document, Cross-Language Event Coreference Annotation Using Event Hoppers (L18-1)
Copied to clipboard
| Challenge: | Defined event hoppers for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task. |
| Approach: | They propose an approach for cross-document, cross-lingual event coreference for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task. |
| Outcome: | The proposed approach is based on the definition of event hoppers for the DEFT rich entities, relations, events and their attributes . it yields 389 cross-document event hoppings in 505 documents in three languages . |
Morphological Segmentation for Low Resource Languages (2020.lrec-1)
Copied to clipboard
Justin Mott, Ann Bies, Stephanie Strassel, Jordan Kodner, Caitlin Richter, Hongzhi Xu, Mitchell Marcus
| Challenge: | a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages . |
| Approach: | This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program. |
| Outcome: | The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications. |
Spanless Event Annotation for Corpus-Wide Complex Event Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for annotating multilingual, multimedia data are limited by the availability of multilingual corpora for schema-based event representation. |
| Approach: | They propose a new approach to event annotation to promote whole-corpus understanding of complex events in multilingual, multimedia data. |
| Outcome: | The proposed method is part of the DARPA Knowledge-directed Artificial Intelligence Reasoning Over Schemas (KAIROS) Program. |
A Study in Contradiction: Data and Annotation for AIDA Focusing on Informational Conflict in Russia-Ukraine Relations (2022.lrec-1)
Copied to clipboard
| Challenge: | This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives (AIDA) program . AIDA systems must extract entities, events, and relations from multimedia documents, aggregate that information across documents and languages, and produce multiple “hypotheses” about what has happened. |
| Approach: | This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives program . the program aims to develop language technology that can help humans manage large volumes of conflicting information . |
| Outcome: | The proposed corpus focuses on the domain of Russia-Ukraine relations and contains source data in English, Russian and Ukrainian . it is designed to support the development and evaluation of systems that extract entities, events, and relations from individual multimedia documents, aggregate the information across documents and languages, and produce multiple “hypotheses” about what has happened. |