Papers by Esther Seyffarth

4 papers
Corpus-based Identification of Verbs Participating in Verb Alternations Using Classification and Manual Annotation (2020.coling-main)

Copied to clipboard

Challenge: Verb alternations allow verbs to appear in a set of syntactically different constructions whose associated semantic frames are systematically related.
Approach: They use ENCOW and VerbNet data to train classifiers to predict the instrument subject alternation and the causative-inchoative alternation . they use count-based and vector-based features as well as perplexity-based language model features to reflect each alternation’s felicity by simulating it.
Outcome: The proposed approach reduces the required annotation effort by only presenting annotators with the highest-scoring candidates from the previous classification.
Verb Alternations and Their Impact on Frame Induction (N18-4)

Copied to clipboard

Challenge: Frame induction is the automatic creation of frame-semantic resources similar to FrameNet or PropBank, which map lexical units of a language to frame representations of each lexical unit’s semantics.
Approach: They propose to use frames to map lexical units to frame representations of each lexical unit's semantics.
Outcome: The proposed framework compares the semantics of alternating verbs and their similarities and differences.
AET: Web-based Adjective Exploration Tool for German (L18-1)

Copied to clipboard

Challenge: AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties.
Approach: They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database .
Outcome: The proposed tool can be transferred to other languages and modification phenomena.
The Maaloula Aramaic Speech Corpus (MASC): From Printed Material to a Lemmatized and Time-Aligned Corpus (2022.lrec-1)

Copied to clipboard

Challenge: The corpus contains 64,845 words, including lemmas, tokens, types, lemmas, sentences, narratives, and speakers.
Approach: They present the first electronic speech corpus of Maaloula Aramaic . it is a Western Neo-Aramaic variety spoken in three Syrian villages .
Outcome: The corpus contains transcriptions, lemmatized transcriptions and audio files . it is available in four formats: transcriptions with audio and phonetic transcriptions .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations