Papers by Laura Kallmeyer
Do you Feel Certain about your Annotation? A Web-based Semantic Frame Annotation Tool Considering Annotators’ Concerns and Behaviors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools for manual annotations are resourceintensive and complex, and experienced annotators and tools specialized for the purpose of the annotation task are required. |
| Approach: | They propose to use a web-based application with a responsive design for modular semantic frame annotation (SFA) the proposed application keeps track of the time and changes during the annotation process and stores the users’ confidence with the current annotation. |
| Outcome: | The proposed system can be used to build a manually annotated corpus and its arguments for task 2 of SemEval 2019 regarding unsupervised lexical frame induction. |
On the Relation Between Fine-Tuning, Topological Properties, and Task Performance in Sense-Enhanced Embeddings (2025.acl-long)
Copied to clipboard
| Challenge: | Enhanced word embeddings do not align well with word senses, resulting in poor performance on word sense identification tasks. |
| Approach: | They propose to use two methods to fine-tune embeddings to identify the topological properties that contribute to sense-enhanced embeddables. |
| Outcome: | The proposed methods improve the embeddings’ ability to capture nuanced semantic distinctions while reducing their expressiveness. |
Multilingual Nonce Dependency Treebanks: Understanding how Language Models Represent and Process Syntactic Structure (2024.naacl-long)
Copied to clipboard
| Challenge: | a number of studies have focused on making explicit the linguistic information encoded in language models (LMs) however, this method has been criticized for various reasons. |
| Approach: | They introduce a framework for creating nonce treebanks for multilingual UD corpora . they investigate word co-occurrence statistics and show how nonce data affects the performance of syntactic dependency probes. |
| Outcome: | The proposed framework satisfies syntactic argument structure and ensures grammaticality via language-specific rules. |
Dissecting Paraphrases: The Impact of Prompt Syntax and supplementary Information on Knowledge Retrieval from Pretrained Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Pre-trained language models contain various kinds of knowledge. |
| Approach: | They designed a probe that allows comparison of 34 million distinct paraphrases that follow a unified meta-template enabling the controlled variation of syntax and semantics across arbitrary relations. |
| Outcome: | Extensive knowledge retrieval experiments show that prompts following clausal syntax have several desirable properties in comparison to appositive syntax. |
DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification (2023.acl-long)
Copied to clipboard
| Challenge: | Current text simplification research mostly focuses on English and on sentencelevel simplification. |
| Approach: | They propose to use a dataset of parallel, professionally written and manually aligned simplifications in plain German "plain DE" and "Einfache Sprache" they build a web harvester and experiment with automatic alignment methods to facilitate integration of non-aligned and to be-published parallel documents. |
| Outcome: | The proposed dataset of parallel, professionally written and manually aligned simplifications in plain German is extended to 750 document pairs and 3.5k sentence pairs. |
Improving Word Sense Induction through Adversarial Forgetting of Morphosyntactic Information (2024.starsem-1)
Copied to clipboard
| Challenge: | Contextualized word representations from pre-trained language models encode more information than is necessary for the identification of word senses and some of this information affect performance negatively in unsupervised settings. |
| Approach: | They propose to use a framework to erase specific information from pre-trained word models and create feature-invariant representations that are invariant to these ‘nuisance features’. |
| Outcome: | The proposed framework erases information from the representations of pre-trained language models, thereby creating feature-invariant representations. |
German and French Neural Supertagging Experiments for LTAG Parsing (P18-3)
Copied to clipboard
| Challenge: | Lexicalized Tree Adjoining Grammars are a linguistically motivated grammar formalism that allows parsers to express linguistic generalizations that are not captured by statistical parsing. |
| Approach: | They propose a supertagging approach combined with deep learning to extract LTAG supertags from the French Treebank and propose n-best supertailing for German and French. |
| Outcome: | The proposed supertagging approach is able to extract LTAG supertags from the French Treebank and n-best supertracking for German and German. |
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
TS-ANNO: An Annotation Tool to Build, Annotate and Evaluate Text Simplification Corpora (2022.acl-demo)
Copied to clipboard
| Challenge: | Currently, high-quality corpora of this type are rare and often of comparably small size. |
| Approach: | They propose an open-source web application for automatic text simplification. |
| Outcome: | TS-ANNO can be used for i) sentence–wise alignment, ii) rating alignment pairs, w.r.t. simplification transformations, and iv) manual simplification of complex documents. |
Statistical Parsing of Tree Wrapping Grammars (2020.coling-main)
Copied to clipboard
| Challenge: | Using tree-wrapping grammars, we propose a statistical parsing algorithm for the grammar. |
| Approach: | They propose a statistical parsing algorithm based on neural supertagging and A* parse for Tree-Wrapping Grammars (TWG) they extract a grammar for English from constituency treebanks and discuss first parser results with this grammar. |
| Outcome: | The proposed algorithm is based on neural supertagging and A* parsing. |
Improving Low-resource RRG Parsing with Cross-lingual Self-training (2022.coling-1)
Copied to clipboard
| Challenge: | a theoretical framework for low-resource parsing is understudied in computational linguistics but widely used in typological research . a novel approach uses Role and Reference Grammar to parse low-source languages . |
| Approach: | They propose to extend an existing RRG parser into a cross-lingual parsing model . they also adopt self-training to adapt the model to a related language with no trees . |
| Outcome: | The proposed model extends into a cross-lingual parser, and iteratively expands the training data. |
Probing for Constituency Structure in Neural Language Models (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Using standard probing techniques, we examine whether contextual neural language models implicitly learn syntactic structure. |
| Approach: | They investigate to which extent contextual neural language models implicitly learn syntactic structure. |
| Outcome: | The proposed model is able to represent constituents of different categories within the neuron activations of a LM such as RoBERTa with high performance even on manipulated data. |
Corpus-based Identification of Verbs Participating in Verb Alternations Using Classification and Manual Annotation (2020.coling-main)
Copied to clipboard
| Challenge: | Verb alternations allow verbs to appear in a set of syntactically different constructions whose associated semantic frames are systematically related. |
| Approach: | They use ENCOW and VerbNet data to train classifiers to predict the instrument subject alternation and the causative-inchoative alternation . they use count-based and vector-based features as well as perplexity-based language model features to reflect each alternation’s felicity by simulating it. |
| Outcome: | The proposed approach reduces the required annotation effort by only presenting annotators with the highest-scoring candidates from the previous classification. |
RRGparbank: A Parallel Role and Reference Grammar Treebank (2022.lrec-1)
Copied to clipboard
Tatiana Bladier, Kilian Evang, Valeria Generalova, Zahra Ghane, Laura Kallmeyer, Robin Möllemann, Natalia Moors, Rainer Osswald, Simon Petitjean
| Challenge: | Existing treebanks for Role and Reference Grammar (RRG) are not yet available. |
| Approach: | They propose to use a multilingual parallel treebank for Role and Reference Grammar to apply RRG to large-scale corpus annotations of 1984 and its translations. |
| Outcome: | The proposed treebank contains annotations of Orwell's 1984 and translations thereof. |
Multilingual Multi-class Sentiment Classification Using Convolutional Neural Networks (L18-1)
Copied to clipboard
| Challenge: | a new language-independent model for sentiment analysis is proposed for social media . a sentiment dictionary cannot list all the possible ways people can express their opinions . |
| Approach: | They propose a language-independent model for multi-class sentiment analysis using a neural network architecture. |
| Outcome: | The proposed model does not rely on language-specific features such as ontologies, dictionaries, or morphological or syntactic pre-processing. |