Papers by Stephanie Strassel

15 papers
Schema Learning Corpus: Data and Annotation Focused on Complex Events (2024.lrec-main)

Copied to clipboard

Challenge: The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data.
Approach: The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian.
Outcome: The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set.
Call My Net 2: A New Resource for Speaker Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Call My Net 2 (CMN2) corpus features Tunisian Arabic conversations between friends and family . call recordings include speech in various realistic and natural acoustic settings, both noisy and non-noisy.
Approach: They introduce the Call My Net 2 (CMN2) corpus, a new resource for speaker recognition featuring Tunisian Arabic conversations between friends and family.
Outcome: The Call My Net 2 (CMN2) corpus contains data from over 400 Tunisian Arabic speakers . each speaker made 10 or more calls each lasting up to 10 minutes .
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)

Copied to clipboard

Challenge: Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events.
Approach: They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements .
Outcome: The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events.
The SAFE-T Corpus: A New Resource for Simulated Public Safety Communications (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium developed the SAFE-T Corpus to support the NIST OpenSAT evaluation series.
Approach: They introduce a new resource, the SAFE-T Corpus, designed to simulate first-responder communications by inducing high vocal effort and urgent speech with situational background noise.
Outcome: The SAFE-T Corpus was developed to support the NIST OpenSAT (Speech Analytic Technologies) evaluation series.
Reflections on 30 Years of Language Resource Development and Sharing (2022.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development.
Approach: They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs.
Outcome: The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans.
Simple Semantic Annotation and Situation Frames: Two Approaches to Basic Text Understanding in LORELEI (L18-1)

Copied to clipboard

Challenge: Existing annotations for low resource languages are under-resourced for human language technology, but lack of resources does not correlate with lack of need for such technologies.
Approach: They propose two types of semantic annotation for the DARPA Low Resource Languages for Emerging Incidents program: Simple Semantic Annotation (SSA) and Situation Frames (SF).
Outcome: The proposed approaches are aimed at labeling basic semantic information relevant to humanitarian aid and disaster relief scenarios.
CAMIO: A Corpus for OCR in Multiple Languages (2022.lrec-1)

Copied to clipboard

Challenge: CAMIO is a corpus of 70,000 images of machine printed text for optical character recognition (OCR) it covers 35 languages across 24 unique scripts.
Approach: CAMIO is a corpus of annotated multilingual images for optical character recognition . the corpus includes nearly 70,000 images of machine printed text .
Outcome: The corpus includes nearly 70,000 images of machine printed text . most images have been exhaustively annotated for text localization .
Cross-Document, Cross-Language Event Coreference Annotation Using Event Hoppers (L18-1)

Copied to clipboard

Challenge: Defined event hoppers for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task.
Approach: They propose an approach for cross-document, cross-lingual event coreference for the DEFT Rich Entities, Relations and Events (Rich ERE) annotation task.
Outcome: The proposed approach is based on the definition of event hoppers for the DEFT rich entities, relations, events and their attributes . it yields 389 cross-document event hoppings in 505 documents in three languages .
VAST: A Corpus of Video Annotation for Speech Technologies (L18-1)

Copied to clipboard

Challenge: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Approach: The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Outcome: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition .
From ‘Solved Problems’ to New Challenges: A Report on LDC Activities (L18-1)

Copied to clipboard

Challenge: This paper reports on the activities of the Linguistic Data Consortium .
Approach: This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report .
Outcome: The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world .
Morphological Segmentation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages .
Approach: This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program.
Outcome: The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications.
Spanless Event Annotation for Corpus-Wide Complex Event Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for annotating multilingual, multimedia data are limited by the availability of multilingual corpora for schema-based event representation.
Approach: They propose a new approach to event annotation to promote whole-corpus understanding of complex events in multilingual, multimedia data.
Outcome: The proposed method is part of the DARPA Knowledge-directed Artificial Intelligence Reasoning Over Schemas (KAIROS) Program.
A Study in Contradiction: Data and Annotation for AIDA Focusing on Informational Conflict in Russia-Ukraine Relations (2022.lrec-1)

Copied to clipboard

Challenge: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives (AIDA) program . AIDA systems must extract entities, events, and relations from multimedia documents, aggregate that information across documents and languages, and produce multiple “hypotheses” about what has happened.
Approach: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives program . the program aims to develop language technology that can help humans manage large volumes of conflicting information .
Outcome: The proposed corpus focuses on the domain of Russia-Ukraine relations and contains source data in English, Russian and Ukrainian . it is designed to support the development and evaluation of systems that extract entities, events, and relations from individual multimedia documents, aggregate the information across documents and languages, and produce multiple “hypotheses” about what has happened.
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources.
Approach: a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation.
Outcome: 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation.
WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition (2022.lrec-1)

Copied to clipboard

Challenge: The WeCanTalk corpus is a multi-modal, multi-language resource for speaker recognition.
Approach: The WeCanTalk corpus is a multi-modal resource for speaker recognition.
Outcome: The corpus contains data from 202 native speakers in Hong Kong who were fluent in at least one other language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations