Papers by Jennifer Tracey

7 papers
Schema Learning Corpus: Data and Annotation Focused on Complex Events (2024.lrec-main)

Copied to clipboard

Challenge: The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data.
Approach: The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian.
Outcome: The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set.
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)

Copied to clipboard

Challenge: Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events.
Approach: They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements .
Outcome: The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events.
Simple Semantic Annotation and Situation Frames: Two Approaches to Basic Text Understanding in LORELEI (L18-1)

Copied to clipboard

Challenge: Existing annotations for low resource languages are under-resourced for human language technology, but lack of resources does not correlate with lack of need for such technologies.
Approach: They propose two types of semantic annotation for the DARPA Low Resource Languages for Emerging Incidents program: Simple Semantic Annotation (SSA) and Situation Frames (SF).
Outcome: The proposed approaches are aimed at labeling basic semantic information relevant to humanitarian aid and disaster relief scenarios.
VAST: A Corpus of Video Annotation for Speech Technologies (L18-1)

Copied to clipboard

Challenge: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Approach: The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Outcome: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition .
BeSt: The Belief and Sentiment Corpus (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of propositional content is a set of cognitive attitudes of different agents towards a text . propositional attitudes are a cognitive attitude, including belief and sentiment, towards .
Approach: They propose a corpus which records cognitive state: who believes what, who has what sentiment . they use newswire and discussion forums in Chinese, English, and Spanish .
Outcome: The proposed corpus records who believes what (i.e., factuality) and who has what sentiment towards what.
Spanless Event Annotation for Corpus-Wide Complex Event Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for annotating multilingual, multimedia data are limited by the availability of multilingual corpora for schema-based event representation.
Approach: They propose a new approach to event annotation to promote whole-corpus understanding of complex events in multilingual, multimedia data.
Outcome: The proposed method is part of the DARPA Knowledge-directed Artificial Intelligence Reasoning Over Schemas (KAIROS) Program.
A Study in Contradiction: Data and Annotation for AIDA Focusing on Informational Conflict in Russia-Ukraine Relations (2022.lrec-1)

Copied to clipboard

Challenge: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives (AIDA) program . AIDA systems must extract entities, events, and relations from multimedia documents, aggregate that information across documents and languages, and produce multiple “hypotheses” about what has happened.
Approach: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives program . the program aims to develop language technology that can help humans manage large volumes of conflicting information .
Outcome: The proposed corpus focuses on the domain of Russia-Ukraine relations and contains source data in English, Russian and Ukrainian . it is designed to support the development and evaluation of systems that extract entities, events, and relations from individual multimedia documents, aggregate the information across documents and languages, and produce multiple “hypotheses” about what has happened.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations