Papers by Chris Biemann
Copied to clipboard
| Challenge: | a new tool for semantic writing aids collects training examples from usage data. |
| Approach: | They propose a semantic writing aid tool based on adaptive paraphrasing that integrates into a real word application to collect training examples from usage data. |
| Outcome: | The proposed tool is integrated into a real word application to collect training examples from usage data. |
Copied to clipboard
| Challenge: | Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication. |
| Approach: | They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings . |
| Outcome: | The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture . |
Copied to clipboard
| Challenge: | polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks . |
| Approach: | They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. |
| Outcome: | The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context. |
Copied to clipboard
| Challenge: | Using a knowledge graph question answering task, we replace the entire SPARQL vocabulary with alternate vocabularies. |
| Approach: | They replace the entire SPARQL vocabulary with alternate vocabularies . they find absolute gains in the range of 17% on the GrailQA dataset . |
| Outcome: | The proposed substitutions show that the model performance improves on the GrailQA dataset. |
Copied to clipboard
| Challenge: | Comparative QA is a challenging task since it requires collecting evidence from many different sources. |
| Approach: | They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices. |
| Outcome: | The proposed system can be used in personal assistants, chatbots, and similar NLP devices. |
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only. |
| Approach: | They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. |
| Outcome: | The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions. |
Copied to clipboard
| Challenge: | Existing QA datasets containing text-and-table data typically contain context-dependent questions, which may yield multiple correct answers depending on the provided context. |
| Approach: | They propose a benchmark to evaluate RAG methods on text-and-table data. |
| Outcome: | The proposed method evaluates RAG methods on real-world text-and-table data. |
Copied to clipboard
| Challenge: | WebAnno focuses on document-level annotation, which is complicated. |
| Approach: | They propose to create hierarchical codebooks that allow to move and sort categories in the hierarchy. |
| Outcome: | The proposed system is based on the existing WebAnno annotation tools and is compatible with existing spreadsheet applications. |
Copied to clipboard
| Challenge: | a thesis aims to explore the use of event extraction in literary texts . event extraction is a challenging domain based on its variety of genres . |
| Approach: | They propose to use event extraction to extract semantic information from literary texts . they propose to build on sequences of event embeddings to form schema embeddables . |
| Outcome: | The proposed approach will allow comparisons between sections of documents and entire literary works. |
Copied to clipboard
| Challenge: | Graph measures, such as node distances, are inefficient to compute. |
| Approach: | They propose a way to learn graph embeddings by using vector operations instead of a graph structure. |
| Outcome: | The proposed method outperforms other graph embeddings on word similarity and word sense disambiguation tasks. |
Copied to clipboard
| Challenge: | Sense Clustering over Time (SCoT) is a network-based tool for analysing lexical change . it visualises word formation, change, and demise as clusters of similar words . SCoT has been successfully used in a European study on the changing meaning of ‘crisis’. |
| Approach: | They propose a new network-based tool for analysing lexical change using a dynamic network of word similarities. |
| Outcome: | The proposed tool has been successfully used in a European study on the changing meaning of ‘crisis’. |
Copied to clipboard
| Challenge: | The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded. |
| Approach: | They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH) |
| Outcome: | The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research. |
Copied to clipboard
| Challenge: | Recent research in large vision-language models has shown promising results, but the issue of hallucination remains. |
| Approach: | They propose an instruction-based method to reduce hallucinations in large vision-language models . they use disturbance instructions to exacerbate hallucinosity in multimodal fusion modules . |
| Outcome: | The proposed method reduces hallucinations in multimodal fusion modules by reducing alignment uncertainty and subtracting hallucines from the original distribution. |
Copied to clipboard
| Challenge: | Comparative Question Answering (cQA) is the task of providing accurate answers to questions . most question answering systems focus on answering factoid questions, but they fail at answering comparative questions in an efficient argumentative manner. |
| Approach: | They propose two new open-domain datasets for identifying and labeling comparative questions . they use a binary classification task and an unsupervised sequence labeling task . |
| Outcome: | The proposed datasets reach close-to-human results on a binary classification task with a neural model using ALBERT embeddings. |
Copied to clipboard
| Challenge: | a new pipeline is being developed to process large collections of unstructured textual data . the pipeline is a key input processor for the upcoming major release of our software . |
| Approach: | a new pipeline is introduced to extract large amounts of unstructured data . the pipeline is used by journalists to process large files containing unknown contents . |
| Outcome: | the pipeline is an input processor for the upcoming major release of our new/s/leak 2.0 software. |
Copied to clipboard
| Challenge: | In primary school, children's books, as well as in modern language learning apps, multi-modal learning strategies like illustrations of terms and phrases are used to support reading comprehension. |
| Approach: | They propose to use multi-modal transformers to train multi-dimensional models on text-image retrieval to support a user's reading comprehension of arbitrary text. |
| Outcome: | The proposed model performs poorly because of the short and relatively simple textual data that the current models are trained with. |
Copied to clipboard
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
Copied to clipboard
| Challenge: | Standard word embeddings lack the ability to distinguish senses of a word by projecting them to exactly one vector. |
| Approach: | They propose to retrofit standard word embeddings to produce sense-aware embeddable vectors using external resources as sense inventories. |
| Outcome: | The proposed method improves word similarity and relatedness scores on multiple word embeddings and established word similarities, sometimes up to an impressive margin of +0.15 Spearman correlation score. |
Copied to clipboard
| Challenge: | Multitask learning and transfer learning are techniques to overcome data scarcity . finding suitable auxiliary datasets for multitask learning is a trial-and-error approach . |
| Approach: | They propose to automatically assess the similarity of sequence tagging datasets to identify beneficial auxiliary data for MTL or TL setups. |
| Outcome: | The proposed methods can compute similarity between two sequence tagging datasets . they show that the same measures correlate with the change in test score of the auxiliary dataset . |
Copied to clipboard
| Challenge: | Using Forum 4.0, we analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
| Approach: | They introduce an open-source framework to semi-automatically analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
| Outcome: | The proposed framework can analyze, aggregate, and visualize user comments based on labels defined by domain experts. |
Copied to clipboard
| Challenge: | Existing benchmarks address single tables or non-visual data, leaving a critical gap . MTabVQA comprises 3,745 complex question-answer pairs . |
| Approach: | They propose a benchmark specifically designed for multi-tabular visual question answering that measures the ability to parse diverse table images and correlate information across them. |
| Outcome: | The proposed benchmarks show that fine-tuning VLMs with MTabVQA-Instruct significantly improves their reasoning abilities. |
Copied to clipboard
| Challenge: | Existing approaches to manage hate speech rely on reactive measures such as blocking or suspending offensive messages . despite regulations imposed by nations and social media platforms, hateful content remains a challenge . |
| Approach: | They propose a framework for automated hate speech moderation based on different strategies . they examine hate speech regulations and strategies from three perspectives . |
| Outcome: | The proposed framework could be based on a combination of country regulations, social platform policies, and NLP research datasets. |
Copied to clipboard
| Challenge: | idiomatic phrases have a non-compositional meaning, meanings of which can be derived from constituents and their grammatical relations. |
| Approach: | They propose to combine hierarchical and distributional information to blend hierarchic and distribution-based hierarchies to detect compositionality for noun phrases. |
| Outcome: | The proposed technique achieves significant improvements over state-of-the-art models based on distributional information and a weighted average of the distributional similarity and p-like function. |
Copied to clipboard
| Challenge: | Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from high-resource languages to low-resourced ones. |
| Approach: | They propose a framework that performs translation and label projection via XML tags. |
| Outcome: | The proposed framework outperforms baselines and improves translation quality across languages and annotation complexity. |
Copied to clipboard
| Challenge: | Existing multilingual vision-language (VL) benchmarks typically only cover a handful of languages, underscoring the need for evaluation data for low-resource languages. |
| Approach: | They propose a multilingual vision-language benchmark that evaluates cross-modal and text-only topical matching across 205 languages. |
| Outcome: | The proposed model performs better in cross-modal and text-only topical matching in lower-resource languages than the most multilingual benchmarks. |
Copied to clipboard
| Challenge: | Pretrained language models (PTLMs) are used for many tasks including syntax, semantics and commonsense. |
| Approach: | They propose to integrate semantic attributes and their values into pretrained language models to improve their performance on many natural language processing tasks. |
| Outcome: | The proposed model performs better on masked tokens than humans on this task. |
Copied to clipboard
| Challenge: | Existing web-based platform for qualitative discourse analysis is limited to text, image, audio, video, and other multimodal data. |
| Approach: | They propose to extend existing web-based platform for digital qualitative discourse analysis with a new extension, Whiteboards, which offers a customizable view of the material and a wide range of actions that enable new ways of interacting with it. |
| Outcome: | The proposed extension facilitates reflection of the research process through sampling maps, creation of actor networks, and refining code taxonomies. |
Copied to clipboard
| Challenge: | Existing methods for extracting hypernyms focus on the acquisition of binary hypernies . |
| Approach: | They propose a distributionally-induced semantic class for extracting hypernyms . they also use distributional semantics to induce sense-aware semantic classes . |
| Outcome: | The proposed method improves the quality of the hypernymy extraction in terms of precision and recall. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can be effective at rewriting toxic content, but they often default to overly polite rewrites, distorting the emotional tone and communicative intent. |
| Approach: | They evaluate 17 large language models with variant architectures to evaluate their ability to rewrite toxic content while preserving the speaker's original intent. |
| Outcome: | The first Chinese detoxification dataset explicitly designed to preserve sentiment polarity is evaluated across five real-world scenarios. |
Copied to clipboard
| Challenge: | Existing approaches to model fictional narratives have focused on the aspect of "what" rather than "how" they are being told. |
| Approach: | They propose a model that embeds stories such that similar stories will result in similar embeddings. |
| Outcome: | The proposed model shows state-of-the-art performance on multiple retrieval tasks and a narrative understanding task. |
Copied to clipboard
| Challenge: | In hierarchical multi-label classification, samples are classified into one or multiple class labels organized in a structured label hierarchy. |
| Approach: | They apply and compare shallow capsule networks for hierarchical multi-label text classification and introduce a new real-world scenario dataset. |
| Outcome: | The proposed model outperforms neural networks and non-neural network architectures on a real-world scenario dataset. |
Copied to clipboard
| Challenge: | Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach. |
| Approach: | They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity. |
| Outcome: | The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds. |
Copied to clipboard
| Challenge: | Existing studies evaluate RAG methods in isolation and focus on single-turn settings. |
| Approach: | They compare retrieval-augmented generation methods for multi-turn conversational QA with those that use dialogue history and coreference to ground large language models. |
| Outcome: | The proposed methods outperform vanilla RAG and advanced methods fail to yield gains and can even degrade performance below the No-RAG baseline. |
Copied to clipboard
| Challenge: | Argumentation is a multi-disciplinary field that extends from philosophy and psychology to linguistics as well as to artificial intelligence. |
| Approach: | They propose to use TARGER to tagging arguments in free text and keyword-based retrieval of arguments from a web-scale corpus. |
| Outcome: | The proposed framework can be used without any reproducibility effort on the user's side and is easily portable to other domains and use cases. |
Copied to clipboard
| Challenge: | Concept Over Time Analysis is a machine-learning-based feature that allows users to define, refine, and visualize concepts of interest within an interactive interface. |
| Approach: | They propose to extend the Discourse Analysis Tool Suite with Concept Over Time Analysis extension that allows users to define, refine, and visualize their concepts of interest within an interactive interface. |
| Outcome: | The proposed system allows users to define, refine, and visualize their concepts of interest within an interactive interface. |
Copied to clipboard
| Challenge: | lexical resource that enriches Framester knowledge graph with semantic features from text corpora . paves way for development of novel, deeper semantic-aware applications . |
| Approach: | They propose a lexical resource that enriches the Framester knowledge graph with semantic features from text corpora. |
| Outcome: | The proposed resource enables the development of deeper semantic-aware applications . it combines knowledge from text and symbolic representations of events and participants . |
Copied to clipboard
| Challenge: | Comparative Question Answering (CQA) is a task that involves processing information and diverse viewpoints. |
| Approach: | They construct a dataset of arguments annotated with their relevance and use it to answer comparative questions. |
| Outcome: | The proposed dataset contains arguments annotated with their relevance and enables precise traceability and faithfulness. |
Copied to clipboard
| Challenge: | Commercial AI-assisted programming Chatbots may generate incorrect information when requests go beyond the model training data or require additional knowledge. |
| Approach: | They propose to implement different self-alignment processes and retrieval-augmented generation pipelines to improve the copilot performance. |
| Outcome: | The proposed model improves the copilot performance on repository-level semantics, dependency between files, and meta-information about the repository. |
Copied to clipboard
| Challenge: | 45% of job posting traffic is driven by recommender systems for job postings . a large-scale job recommendation system is needed to detect similarity between job posting and item-to-item based recommendations. |
| Approach: | They propose to use dense vector representations to enhance a large-scale job recommendation system and rank job advertisements regarding similarity. |
| Outcome: | The proposed method increases the click-through rate on job recommendations by 8.0%. |
Copied to clipboard
| Challenge: | a systematic review of academic writing aids aims to build a writing aid system that automatically edits a text to adhere to the academic style of writing. |
| Approach: | They propose to build a writing aid system that automatically edits a text to adhere to the academic style of writing. |
| Outcome: | The proposed system outperforms existing academic resources in terms of word identification and ranking . the informal word identification component achieves an F-1 score of 82% . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used in numerous NLP tasks, including counterspeech generation. |
| Approach: | They propose three different prompting strategies for generating different types of counterspeech and propose a set of prompting techniques for counterspeak generation. |
| Outcome: | The proposed prompting strategies improve the performance of the models for counterspeech generation in two datasets, but with high toxicity with increase in model size. |
Copied to clipboard
| Challenge: | LT Expertfinder is a web-based tool for expert finding and expert profiling. |
| Approach: | They propose a web-application that enables qualitative comparison between different ranking methods . LT Expertfinder provides detailed expert profiles linked to Wikidata and Google Scholar . |
| Outcome: | The LT Expertfinder is a web-based tool for expert finding and evaluation. |
Copied to clipboard
| Challenge: | Entity Disambiguation (ED) is the task of linking an ambiguous entity mention to a corresponding entry in a knowledge base. |
| Approach: | They propose a method that integrates structured information from the knowledge base with unstructured information from text-based representations. |
| Outcome: | The proposed method improves on a graph of hyperlinks between Wikipedia articles and a state-of-the-art neural ED model. |
Copied to clipboard
| Challenge: | Existing approaches to low-resource languages are limited to 500 languages . a lot of tasks for low-rsource languages remain unsolved . |
| Approach: | They propose a new approach called MeritOpt that can be applied to Natural Language Tasks with heterogeneous data. |
| Outcome: | The proposed approach can be applied to a low-resource machine translation task using the datasets of South East Asian and Finno-Ugric languages. |
Copied to clipboard
| Challenge: | a new challenge is learning from a real-world data stream and continuously updating the model without explicit supervision. |
| Approach: | They develop an adaptive learning system for text simplification which improves the underlying ranking model from usage data. |
| Outcome: | The proposed system improves the learning-to-rank model from usage data over time. |
Copied to clipboard
| Challenge: | Existing approaches to domain-specific taxonomy induction from text are relying on distributional semantics for hyponym-hypernym relationships, but many of them learn prototypical hypernymes, not taking into account the relation between both terms in classification. |
| Approach: | They propose to use Poincaré embeddings to improve existing approaches to domain-specific taxonomy induction from text as a signal for relocating wrong hyponym terms and attaching disconnected terms in a taxonomies. |
| Outcome: | The proposed method significantly improves state-of-the-art methods on the SemEval-2016 Task 13 on taxonomy extraction. |
Copied to clipboard
| Challenge: | Recent work on frame-semantics has enabled the development of wide-coverage frame parsers using supervised learning. |
| Approach: | They propose to use dependency triples to perform unsupervised frame induction on a Web-scale corpus. |
| Outcome: | The proposed approach performs state-of-the-art on a FrameNet-derived dataset and performs on par with competitive methods on . verb class clustering task. |
Copied to clipboard
| Challenge: | Lack of annotated data for quotation attribution in news articles severely limits the quality and usability of possible systems. |
| Approach: | They propose a dataset for quotation attribution in German news articles using WIKINEWS and manually annotated quotes from 1000 articles. |
| Outcome: | The proposed dataset provides curated, high-quality annotations across 1000 documents (250,000 tokens) in a fine-grained annotation schema enabling various downstream uses for the dataset. |
Copied to clipboard
| Challenge: | DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing . |
| Approach: | They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus . |
| Outcome: | The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset. |
Copied to clipboard
| Challenge: | Existing systems for word sense disambiguation are limited to the Russian language and lack of resources to address the problem. |
| Approach: | They propose an unsupervised system for word sense disambiguation that uses a traditional vector space model to estimate the most similar word sense corresponding to its context. |
| Outcome: | The proposed system outperforms the sparse mode on all datasets according to the adjusted Rand index. |
Copied to clipboard
| Challenge: | Comparative Question Answering is a Natural Language Processing task that combines Question Answers and Argument Mining. |
| Approach: | They propose a system for answering comparative questions called CAM 2.0 and a public leaderboard called CompUGE that unifies existing datasets under a single easy-to-use evaluation suite. |
| Outcome: | The proposed system is compared with previous web-form-based systems . it features question identification, object and aspect labeling, stance classification, summarization . the proposed system has a user-friendly interface and is available for free on the web . |
Copied to clipboard
| Challenge: | Existing approaches to represent narratives on short-form texts are limited as narrative semantics are an open class. |
| Approach: | They propose to use Wikipedia summaries as a proxy for entire stories or for analysis of the summary itself. |
| Outcome: | The proposed dataset contains 96,831 individual summaries across 29,505 stories. |
Copied to clipboard
| Challenge: | Existing crowdsourcing platforms do not support sentiment analysis for Amharic, and there are no expert researchers in the area. |
| Approach: | They propose to build a social-network-friendly Amharic sentiment analysis tool using the Telegram bot and collect 9.4k tweets where each tweet is annotated by three Telegram users. |
| Outcome: | The proposed system outperforms existing classifiers in Amharic and other low-resource languages due to the widespread use of sarcasm and figurative speech. |
Copied to clipboard
| Challenge: | Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. |
| Approach: | They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data . |
| Outcome: | The proposed model outperforms existing models in 14 tasks and 56 languages. |
Copied to clipboard
| Challenge: | a dataset containing source code solutions to algorithmic programming exercises solved by students at the University of Hamburg is available under the permissive CC BY-NC 4.0 license. |
| Approach: | They present a dataset containing source code solutions to algorithmic programming exercises solved by students at the University of Hamburg. |
| Outcome: | The proposed dataset contains solutions to 21 programming tasks written in Java and Python and over 1500 individual solutions. |
Copied to clipboard
| Challenge: | De-identification is the task of detecting protected health information (PHI) in medical text. |
| Approach: | They propose to create shareable representations of medical text that contain no PHI and can be shared between organizations to create unified datasets for training de-identification models. |
| Outcome: | The proposed representation allows training a simple LSTM-CRF model to an F1 score of 97.4%. |
Copied to clipboard
| Challenge: | Existing tools for document-level annotation lack document-based quality and flexibility. |
| Approach: | a new annotation tool is being developed for industry and research use cases . a configurable user interface and a RESTful API are included . authors propose to use ACTIVEANNO as default for document-level annotation . |
| Outcome: | ACTIVEANNO is an annotation tool for industry and research use cases. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) achieve excellent performance through pretraining on extensive data. |
| Approach: | They propose an efficient selective layer intervention based on parameter-efficient fine-tuning methods to select the optimal steering layer to modulate LLM semantics. |
| Outcome: | The proposed approach is based on a model-agnostic framework and is safe to deploy. |
Copied to clipboard
| Challenge: | Comparative Question Answering systems help users make informed decisions by generating comparative information. |
| Approach: | They propose a comprehensive benchmark designed to evaluate Comparative Question Answering systems. |
| Outcome: | The proposed benchmark is available on HuggingFace Spaces . it unifies multiple datasets and provides a robust evaluation platform . |
Copied to clipboard
| Challenge: | Existing studies have shown that multimodal information is crucial for concept formation, accordingly for language acquisition. |
| Approach: | They collect a multimodal dataset enriched with complex word annotations and validated image match. |
| Outcome: | The proposed dataset contains 1125 comprehension texts retrieved from Wikipedia Simple Corpus . |
Copied to clipboard
| Challenge: | Existing large language models are not designed for semantic retrieval and PDF-based legislative sources introduce substantial noise due to imperfect text extraction. |
| Approach: | They propose a large-scale multilingual corpus of EU environmental legislation constructed from 24,953 official EUR-Lex PDF documents covering 25 languages. |
| Outcome: | The proposed model improves Top-k retrieval accuracy in monolingual and bilingual settings . it also improves accuracy in low- and high-resource languages . |
Copied to clipboard
| Challenge: | Existing methods of disambiguation of word senses are based on knowledge bases, taxonomies, and other externally built resources. |
| Approach: | They propose a method that takes a pre-trained word embedding model and induces a fully-fledged word sense inventory for 158 languages. |
| Outcome: | The proposed model is based on a pre-trained word embedding model and induces a fully-fledged word sense inventory in 158 languages. |