Papers by Chris Biemann

62 papers
Demonstrating Par4Sem - A Semantic Writing Aid with Adaptive Paraphrasing (D18-2)

Copied to clipboard

Challenge: a new tool for semantic writing aids collects training examples from usage data.
Approach: They propose a semantic writing aid tool based on adaptive paraphrasing that integrates into a real word application to collect training examples from usage data.
Outcome: The proposed tool is integrated into a real word application to collect training examples from usage data.
Modeling Referential Gaze in Task-oriented Settings of Varying Referential Complexity (2022.findings-aacl)

Copied to clipboard

Challenge: Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication.
Approach: They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings .
Outcome: The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture .
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)

Copied to clipboard

Challenge: polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks .
Approach: They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events.
Outcome: The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context.
The Role of Output Vocabulary in T2T LMs for SPARQL Semantic Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Using a knowledge graph question answering task, we replace the entire SPARQL vocabulary with alternate vocabularies.
Approach: They replace the entire SPARQL vocabulary with alternate vocabularies . they find absolute gains in the range of 17% on the GrailQA dataset .
Outcome: The proposed substitutions show that the model performance improves on the GrailQA dataset.
Which is Better for Deep Learning: Python or MATLAB? Answering Comparative Questions in Natural Language (2021.eacl-demos)

Copied to clipboard

Challenge: Comparative QA is a challenging task since it requires collecting evidence from many different sources.
Approach: They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices.
Outcome: The proposed system can be used in personal assistants, chatbots, and similar NLP devices.
GIMMICK: Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only.
Approach: They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions.
Outcome: The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions.
T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation (2026.eacl-long)

Copied to clipboard

Challenge: Existing QA datasets containing text-and-table data typically contain context-dependent questions, which may yield multiple correct answers depending on the provided context.
Approach: They propose a benchmark to evaluate RAG methods on text-and-table data.
Outcome: The proposed method evaluates RAG methods on real-world text-and-table data.
CodeAnno: Extending WebAnno with Hierarchical Document Level Annotation and Automation (2023.eacl-demo)

Copied to clipboard

Challenge: WebAnno focuses on document-level annotation, which is complicated.
Approach: They propose to create hierarchical codebooks that allow to move and sort categories in the hierarchy.
Outcome: The proposed system is based on the existing WebAnno annotation tools and is compatible with existing spreadsheet applications.
Towards Layered Events and Schema Representations in Long Documents (2021.naacl-srw)

Copied to clipboard

Challenge: a thesis aims to explore the use of event extraction in literary texts . event extraction is a challenging domain based on its variety of genres .
Approach: They propose to use event extraction to extract semantic information from literary texts . they propose to build on sequences of event embeddings to form schema embeddables .
Outcome: The proposed approach will allow comparisons between sections of documents and entire literary works.
Making Fast Graph-based Algorithms with Graph Metric Embeddings (P19-1)

Copied to clipboard

Challenge: Graph measures, such as node distances, are inefficient to compute.
Approach: They propose a way to learn graph embeddings by using vector operations instead of a graph structure.
Outcome: The proposed method outperforms other graph embeddings on word similarity and word sense disambiguation tasks.
SCoT: Sense Clustering over Time: a tool for the analysis of lexical change (2021.eacl-demos)

Copied to clipboard

Challenge: Sense Clustering over Time (SCoT) is a network-based tool for analysing lexical change . it visualises word formation, change, and demise as clusters of similar words . SCoT has been successfully used in a European study on the changing meaning of ‘crisis’.
Approach: They propose a new network-based tool for analysing lexical change using a dynamic network of word similarities.
Outcome: The proposed tool has been successfully used in a European study on the changing meaning of ‘crisis’.
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)

Copied to clipboard

Challenge: The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded.
Approach: They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH)
Outcome: The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research.
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding (2024.findings-acl)

Copied to clipboard

Challenge: Recent research in large vision-language models has shown promising results, but the issue of hallucination remains.
Approach: They propose an instruction-based method to reduce hallucinations in large vision-language models . they use disturbance instructions to exacerbate hallucinosity in multimodal fusion modules .
Outcome: The proposed method reduces hallucinations in multimodal fusion modules by reducing alignment uncertainty and subtracting hallucines from the original distribution.
Elvis vs. M. Jackson: Who has More Albums? Classification and Identification of Elements in Comparative Questions (2022.lrec-1)

Copied to clipboard

Challenge: Comparative Question Answering (cQA) is the task of providing accurate answers to questions . most question answering systems focus on answering factoid questions, but they fail at answering comparative questions in an efficient argumentative manner.
Approach: They propose two new open-domain datasets for identifying and labeling comparative questions . they use a binary classification task and an unsupervised sequence labeling task .
Outcome: The proposed datasets reach close-to-human results on a binary classification task with a neural model using ALBERT embeddings.
A Multilingual Information Extraction Pipeline for Investigative Journalism (D18-2)

Copied to clipboard

Challenge: a new pipeline is being developed to process large collections of unstructured textual data . the pipeline is a key input processor for the upcoming major release of our software .
Approach: a new pipeline is introduced to extract large amounts of unstructured data . the pipeline is used by journalists to process large files containing unknown contents .
Outcome: the pipeline is an input processor for the upcoming major release of our new/s/leak 2.0 software.
Towards Multi-Modal Text-Image Retrieval to improve Human Reading (2021.naacl-srw)

Copied to clipboard

Challenge: In primary school, children's books, as well as in modern language learning apps, multi-modal learning strategies like illustrations of terms and phrases are used to support reading comprehension.
Approach: They propose to use multi-modal transformers to train multi-dimensional models on text-image retrieval to support a user's reading comprehension of arbitrary text.
Outcome: The proposed model performs poorly because of the short and relatively simple textual data that the current models are trained with.
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
Retrofitting Word Representations for Unsupervised Sense Aware Word Similarities (L18-1)

Copied to clipboard

Challenge: Standard word embeddings lack the ability to distinguish senses of a word by projecting them to exactly one vector.
Approach: They propose to retrofit standard word embeddings to produce sense-aware embeddable vectors using external resources as sense inventories.
Outcome: The proposed method improves word similarity and relatedness scores on multiple word embeddings and established word similarities, sometimes up to an impressive margin of +0.15 Spearman correlation score.
Estimating the influence of auxiliary tasks for multi-task learning of sequence tagging tasks (2020.acl-main)

Copied to clipboard

Challenge: Multitask learning and transfer learning are techniques to overcome data scarcity . finding suitable auxiliary datasets for multitask learning is a trial-and-error approach .
Approach: They propose to automatically assess the similarity of sequence tagging datasets to identify beneficial auxiliary data for MTL or TL setups.
Outcome: The proposed methods can compute similarity between two sequence tagging datasets . they show that the same measures correlate with the change in test score of the auxiliary dataset .
Forum 4.0: An Open-Source User Comment Analysis Framework (2021.eacl-demos)

Copied to clipboard

Challenge: Using Forum 4.0, we analyze, aggregate, and visualize user comments based on labels defined by domain experts.
Approach: They introduce an open-source framework to semi-automatically analyze, aggregate, and visualize user comments based on labels defined by domain experts.
Outcome: The proposed framework can analyze, aggregate, and visualize user comments based on labels defined by domain experts.
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks address single tables or non-visual data, leaving a critical gap . MTabVQA comprises 3,745 complex question-answer pairs .
Approach: They propose a benchmark specifically designed for multi-tabular visual question answering that measures the ability to parse diverse table images and correlate information across them.
Outcome: The proposed benchmarks show that fine-tuning VLMs with MTabVQA-Instruct significantly improves their reasoning abilities.
HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to manage hate speech rely on reactive measures such as blocking or suspending offensive messages . despite regulations imposed by nations and social media platforms, hateful content remains a challenge .
Approach: They propose a framework for automated hate speech moderation based on different strategies . they examine hate speech regulations and strategies from three perspectives .
Outcome: The proposed framework could be based on a combination of country regulations, social platform policies, and NLP research datasets.
On the Compositionality Prediction of Noun Phrases using Poincaré Embeddings (P19-1)

Copied to clipboard

Challenge: idiomatic phrases have a non-compositional meaning, meanings of which can be derived from constituents and their grammatical relations.
Approach: They propose to combine hierarchical and distributional information to blend hierarchic and distribution-based hierarchies to detect compositionality for noun phrases.
Outcome: The proposed technique achieves significant improvements over state-of-the-art models based on distributional information and a weighted average of the distributional similarity and p-like function.
Just Use XML: Revisiting Joint Translation and Label Projection (2026.findings-acl)

Copied to clipboard

Challenge: Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from high-resource languages to low-resourced ones.
Approach: They propose a framework that performs translation and label projection via XML tags.
Outcome: The proposed framework outperforms baselines and improves translation quality across languages and annotation complexity.
MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching (2025.findings-acl)

Copied to clipboard

Challenge: Existing multilingual vision-language (VL) benchmarks typically only cover a handful of languages, underscoring the need for evaluation data for low-resource languages.
Approach: They propose a multilingual vision-language benchmark that evaluates cross-modal and text-only topical matching across 205 languages.
Outcome: The proposed model performs better in cross-modal and text-only topical matching in lower-resource languages than the most multilingual benchmarks.
Probing Pre-trained Language Models for Semantic Attributes and their Values (2021.findings-emnlp)

Copied to clipboard

Challenge: Pretrained language models (PTLMs) are used for many tasks including syntax, semantics and commonsense.
Approach: They propose to integrate semantic attributes and their values into pretrained language models to improve their performance on many natural language processing tasks.
Outcome: The proposed model performs better on masked tokens than humans on this task.
Extending the Discourse Analysis Tool Suite with Whiteboards for Visual Qualitative Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Existing web-based platform for qualitative discourse analysis is limited to text, image, audio, video, and other multimodal data.
Approach: They propose to extend existing web-based platform for digital qualitative discourse analysis with a new extension, Whiteboards, which offers a customizable view of the material and a wide range of actions that enable new ways of interacting with it.
Outcome: The proposed extension facilitates reflection of the research process through sampling maps, creation of actor networks, and refining code taxonomies.
Improving Hypernymy Extraction with Distributional Semantic Classes (L18-1)

Copied to clipboard

Challenge: Existing methods for extracting hypernyms focus on the acquisition of binary hypernies .
Approach: They propose a distributionally-induced semantic class for extracting hypernyms . they also use distributional semantics to induce sense-aware semantic classes .
Outcome: The proposed method improves the quality of the hypernymy extraction in terms of precision and recall.
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can be effective at rewriting toxic content, but they often default to overly polite rewrites, distorting the emotional tone and communicative intent.
Approach: They evaluate 17 large language models with variant architectures to evaluate their ability to rewrite toxic content while preserving the speaker's original intent.
Outcome: The first Chinese detoxification dataset explicitly designed to preserve sentiment polarity is evaluated across five real-world scenarios.
Story Embeddings — Narrative-Focused Representations of Fictional Stories (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model fictional narratives have focused on the aspect of "what" rather than "how" they are being told.
Approach: They propose a model that embeds stories such that similar stories will result in similar embeddings.
Outcome: The proposed model shows state-of-the-art performance on multiple retrieval tasks and a narrative understanding task.
Hierarchical Multi-label Classification of Text with Capsule Networks (P19-2)

Copied to clipboard

Challenge: In hierarchical multi-label classification, samples are classified into one or multiple class labels organized in a structured label hierarchy.
Approach: They apply and compare shallow capsule networks for hierarchical multi-label text classification and introduce a new real-world scenario dataset.
Outcome: The proposed model outperforms neural networks and non-neural network architectures on a real-world scenario dataset.
Word Complexity is in the Eye of the Beholder (2021.naacl-main)

Copied to clipboard

Challenge: Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach.
Approach: They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity.
Outcome: The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds.
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA (2026.eacl-srw)

Copied to clipboard

Challenge: Existing studies evaluate RAG methods in isolation and focus on single-turn settings.
Approach: They compare retrieval-augmented generation methods for multi-turn conversational QA with those that use dialogue history and coreference to ground large language models.
Outcome: The proposed methods outperform vanilla RAG and advanced methods fail to yield gains and can even degrade performance below the No-RAG baseline.
TARGER: Neural Argument Mining at Your Fingertips (P19-3)

Copied to clipboard

Challenge: Argumentation is a multi-disciplinary field that extends from philosophy and psychology to linguistics as well as to artificial intelligence.
Approach: They propose to use TARGER to tagging arguments in free text and keyword-based retrieval of arguments from a web-scale corpus.
Outcome: The proposed framework can be used without any reproducibility effort on the user's side and is easily portable to other domains and use cases.
Concept Over Time Analysis: Unveiling Temporal Patterns for Qualitative Data Analysis (2024.naacl-demo)

Copied to clipboard

Challenge: Concept Over Time Analysis is a machine-learning-based feature that allows users to define, refine, and visualize concepts of interest within an interactive interface.
Approach: They propose to extend the Discourse Analysis Tool Suite with Concept Over Time Analysis extension that allows users to define, refine, and visualize their concepts of interest within an interactive interface.
Outcome: The proposed system allows users to define, refine, and visualize their concepts of interest within an interactive interface.
Enriching Frame Representations with Distributionally Induced Senses (L18-1)

Copied to clipboard

Challenge: lexical resource that enriches Framester knowledge graph with semantic features from text corpora . paves way for development of novel, deeper semantic-aware applications .
Approach: They propose a lexical resource that enriches the Framester knowledge graph with semantic features from text corpora.
Outcome: The proposed resource enables the development of deeper semantic-aware applications . it combines knowledge from text and symbolic representations of events and participants .
How to Compare Things Properly? A Study of Argument Relevance in Comparative Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Comparative Question Answering (CQA) is a task that involves processing information and diverse viewpoints.
Approach: They construct a dataset of arguments annotated with their relevance and use it to answer comparative questions.
Outcome: The proposed dataset contains arguments annotated with their relevance and enables precise traceability and faithfulness.
On Improving Repository-Level Code QA for Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Commercial AI-assisted programming Chatbots may generate incorrect information when requests go beyond the model training data or require additional knowledge.
Approach: They propose to implement different self-alignment processes and retrieval-augmented generation pipelines to improve the copilot performance.
Outcome: The proposed model improves the copilot performance on repository-level semantics, dependency between files, and meta-information about the repository.
Document-based Recommender System for Job Postings using Dense Representations (N18-3)

Copied to clipboard

Challenge: 45% of job posting traffic is driven by recommender systems for job postings . a large-scale job recommendation system is needed to detect similarity between job posting and item-to-item based recommendations.
Approach: They propose to use dense vector representations to enhance a large-scale job recommendation system and rank job advertisements regarding similarity.
Outcome: The proposed method increases the click-through rate on job recommendations by 8.0%.
Automatic Compilation of Resources for Academic Writing and Evaluating with Informal Word Identification and Paraphrasing System (2020.lrec-1)

Copied to clipboard

Challenge: a systematic review of academic writing aids aims to build a writing aid system that automatically edits a text to adhere to the academic style of writing.
Approach: They propose to build a writing aid system that automatically edits a text to adhere to the academic style of writing.
Outcome: The proposed system outperforms existing academic resources in terms of word identification and ranking . the informal word identification component achieves an F-1 score of 82% .
On Zero-Shot Counterspeech Generation by LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used in numerous NLP tasks, including counterspeech generation.
Approach: They propose three different prompting strategies for generating different types of counterspeech and propose a set of prompting techniques for counterspeak generation.
Outcome: The proposed prompting strategies improve the performance of the models for counterspeech generation in two datasets, but with high toxicity with increase in model size.
LT Expertfinder: An Evaluation Framework for Expert Finding Methods (N19-4)

Copied to clipboard

Challenge: LT Expertfinder is a web-based tool for expert finding and expert profiling.
Approach: They propose a web-application that enables qualitative comparison between different ranking methods . LT Expertfinder provides detailed expert profiles linked to Wikidata and Google Scholar .
Outcome: The LT Expertfinder is a web-based tool for expert finding and evaluation.
Improving Neural Entity Disambiguation with Graph Embeddings (P19-2)

Copied to clipboard

Challenge: Entity Disambiguation (ED) is the task of linking an ambiguous entity mention to a corresponding entry in a knowledge base.
Approach: They propose a method that integrates structured information from the knowledge base with unstructured information from text-based representations.
Outcome: The proposed method improves on a graph of hyperlinks between Wikipedia articles and a state-of-the-art neural ED model.
Low-Resource Machine Translation through the Lens of Personalized Federated Learning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to low-resource languages are limited to 500 languages . a lot of tasks for low-rsource languages remain unsolved .
Approach: They propose a new approach called MeritOpt that can be applied to Natural Language Tasks with heterogeneous data.
Outcome: The proposed approach can be applied to a low-resource machine translation task using the datasets of South East Asian and Finno-Ugric languages.
Par4Sim – Adaptive Paraphrasing for Text Simplification (C18-1)

Copied to clipboard

Challenge: a new challenge is learning from a real-world data stream and continuously updating the model without explicit supervision.
Approach: They develop an adaptive learning system for text simplification which improves the underlying ranking model from usage data.
Outcome: The proposed system improves the learning-to-rank model from usage data over time.
Every Child Should Have Parents: A Taxonomy Refinement Algorithm Based on Hyperbolic Term Embeddings (P19-1)

Copied to clipboard

Challenge: Existing approaches to domain-specific taxonomy induction from text are relying on distributional semantics for hyponym-hypernym relationships, but many of them learn prototypical hypernymes, not taking into account the relation between both terms in classification.
Approach: They propose to use Poincaré embeddings to improve existing approaches to domain-specific taxonomy induction from text as a signal for relocating wrong hyponym terms and attaching disconnected terms in a taxonomies.
Outcome: The proposed method significantly improves state-of-the-art methods on the SemEval-2016 Task 13 on taxonomy extraction.
Unsupervised Semantic Frame Induction using Triclustering (P18-2)

Copied to clipboard

Challenge: Recent work on frame-semantics has enabled the development of wide-coverage frame parsers using supervised learning.
Approach: They propose to use dependency triples to perform unsupervised frame induction on a Web-scale corpus.
Outcome: The proposed approach performs state-of-the-art on a FrameNet-derived dataset and performs on par with competitive methods on . verb class clustering task.
Dataset of Quotation Attribution in German News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Lack of annotated data for quotation attribution in news articles severely limits the quality and usability of possible systems.
Approach: They propose a dataset for quotation attribution in German news articles using WIKINEWS and manually annotated quotes from 1000 articles.
Outcome: The proposed dataset provides curated, high-quality annotations across 1000 documents (250,000 tokens) in a fine-grained annotation schema enabling various downstream uses for the dataset.
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)

Copied to clipboard

Challenge: DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing .
Approach: They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus .
Outcome: The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset.
An Unsupervised Word Sense Disambiguation System for Under-Resourced Languages (L18-1)

Copied to clipboard

Challenge: Existing systems for word sense disambiguation are limited to the Russian language and lack of resources to address the problem.
Approach: They propose an unsupervised system for word sense disambiguation that uses a traditional vector space model to estimate the most similar word sense corresponding to its context.
Outcome: The proposed system outperforms the sparse mode on all datasets according to the adjusted Rand index.
CAM 2.0: End-to-End Open Domain Comparative Question Answering System (2024.lrec-main)

Copied to clipboard

Challenge: Comparative Question Answering is a Natural Language Processing task that combines Question Answers and Argument Mining.
Approach: They propose a system for answering comparative questions called CAM 2.0 and a public leaderboard called CompUGE that unifies existing datasets under a single easy-to-use evaluation suite.
Outcome: The proposed system is compared with previous web-form-based systems . it features question identification, object and aspect labeling, stance classification, summarization . the proposed system has a user-friendly interface and is available for free on the web .
Tell Me Again! a Large-Scale Dataset of Multiple Summaries for the Same Story (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to represent narratives on short-form texts are limited as narrative semantics are an open class.
Approach: They propose to use Wikipedia summaries as a proxy for entire stories or for analysis of the summary itself.
Outcome: The proposed dataset contains 96,831 individual summaries across 29,505 stories.
Exploring Amharic Sentiment Analysis from Social Media Texts: Building Annotation Tools and Classification Models (2020.coling-main)

Copied to clipboard

Challenge: Existing crowdsourcing platforms do not support sentiment analysis for Amharic, and there are no expert researchers in the area.
Approach: They propose to build a social-network-friendly Amharic sentiment analysis tool using the Telegram bot and collect 9.4k tweets where each tweet is annotated by three Telegram users.
Outcome: The proposed system outperforms existing classifiers in Amharic and other low-resource languages due to the widespread use of sarcasm and figurative speech.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
Dataset of Student Solutions to Algorithm and Data Structure Programming Assignments (2022.lrec-1)

Copied to clipboard

Challenge: a dataset containing source code solutions to algorithmic programming exercises solved by students at the University of Hamburg is available under the permissive CC BY-NC 4.0 license.
Approach: They present a dataset containing source code solutions to algorithmic programming exercises solved by students at the University of Hamburg.
Outcome: The proposed dataset contains solutions to 21 programming tasks written in Java and Python and over 1500 individual solutions.
Adversarial Learning of Privacy-Preserving Text Representations for De-Identification of Medical Records (P19-1)

Copied to clipboard

Challenge: De-identification is the task of detecting protected health information (PHI) in medical text.
Approach: They propose to create shareable representations of medical text that contain no PHI and can be shared between organizations to create unified datasets for training de-identification models.
Outcome: The proposed representation allows training a simple LSTM-CRF model to an F1 score of 97.4%.
ActiveAnno: General-Purpose Document-Level Annotation Tool with Active Learning Integration (2021.naacl-demos)

Copied to clipboard

Challenge: Existing tools for document-level annotation lack document-based quality and flexibility.
Approach: a new annotation tool is being developed for industry and research use cases . a configurable user interface and a RESTful API are included . authors propose to use ACTIVEANNO as default for document-level annotation .
Outcome: ACTIVEANNO is an annotation tool for industry and research use cases.
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) achieve excellent performance through pretraining on extensive data.
Approach: They propose an efficient selective layer intervention based on parameter-efficient fine-tuning methods to select the optimal steering layer to modulate LLM semantics.
Outcome: The proposed approach is based on a model-agnostic framework and is safe to deploy.
CompUGE-Bench: Comparative Understanding and Generation Evaluation Benchmark for Comparative Question Answering (2025.coling-demos)

Copied to clipboard

Challenge: Comparative Question Answering systems help users make informed decisions by generating comparative information.
Approach: They propose a comprehensive benchmark designed to evaluate Comparative Question Answering systems.
Outcome: The proposed benchmark is available on HuggingFace Spaces . it unifies multiple datasets and provides a robust evaluation platform .
MOTIF: Contextualized Images for Complex Words to Improve Human Reading (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have shown that multimodal information is crucial for concept formation, accordingly for language acquisition.
Approach: They collect a multimodal dataset enriched with complex word annotations and validated image match.
Outcome: The proposed dataset contains 1125 comprehension texts retrieved from Wikipedia Simple Corpus .
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval (2026.eacl-srw)

Copied to clipboard

Challenge: Existing large language models are not designed for semantic retrieval and PDF-based legislative sources introduce substantial noise due to imperfect text extraction.
Approach: They propose a large-scale multilingual corpus of EU environmental legislation constructed from 24,953 official EUR-Lex PDF documents covering 25 languages.
Outcome: The proposed model improves Top-k retrieval accuracy in monolingual and bilingual settings . it also improves accuracy in low- and high-resource languages .
Word Sense Disambiguation for 158 Languages using Word Embeddings Only (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods of disambiguation of word senses are based on knowledge bases, taxonomies, and other externally built resources.
Approach: They propose a method that takes a pre-trained word embedding model and induces a fully-fledged word sense inventory for 158 languages.
Outcome: The proposed model is based on a pre-trained word embedding model and induces a fully-fledged word sense inventory in 158 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations