Papers by Kalina Bontcheva
On the Impact of Temporal Concept Drift on Model Explanations (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Explanation faithfulness of model predictions is typically evaluated on held-out data from the same temporal distribution as the training data. |
| Approach: | They examine the impact of temporal variation on model explanations extracted by eight feature attribution methods and three select-then-predict models across six text classification tasks. |
| Outcome: | The proposed method shows the most robust faithfulness scores across datasets and in asynchronous settings. |
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe (2020.lrec-1)
Copied to clipboard
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajič, Khalid Choukri, Andrejs Vasiļjevs, Gerhard Backfried, Christoph Prinz, José Manuel Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriūtė, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavriilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette Pedersen, Inguna Skadiņa, Marko Tadić, Dan Tufiș, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
| Challenge: | Language Technologies (LTs) are a powerful means to break down language barriers impacting business, cross-lingual and cross-cultural communication in Europe. |
| Approach: | They present an overview of the European LT landscape and the current state of play in industry and the LT market. |
| Outcome: | The present study outlines funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. |
Can Rumour Stance Alone Predict Veracity? (C18-1)
Copied to clipboard
| Challenge: | Existing studies of automatic veracity classification of social media rumours have not explored the effectiveness of crowd stance to determine veracity. |
| Approach: | They propose to use stance as an additional feature to those commonly used in earlier studies to model the veracity of a rumour using Hidden Markov Models and collective stance information to model a social media rumor. |
| Outcome: | The proposed models outperform those using crowd stance and tweets’ times as the only features for modelling true and false rumours. |
Using Deep Neural Networks with Intra- and Inter-Sentence Context to Classify Suicidal Behaviour (2020.lrec-1)
Copied to clipboard
Xingyi Song, Johnny Downs, Sumithra Velupillai, Rachel Holden, Maxim Kikoler, Kalina Bontcheva, Rina Dutta, Angus Roberts
| Challenge: | Mental health problems are a major risk factor for suicide attempts. |
| Approach: | They propose to integrate information from sentences to left and right of the target sentence into the model to improve classification accuracy. |
| Outcome: | The proposed model was able to classify suicidal behaviour in autism spectrum disorder patient records significantly better than previous approaches. |
European Language Grid: A Joint Platform for the European Language Technology Community (2021.eacl-demos)
Copied to clipboard
Georg Rehm, Stelios Piperidis, Kalina Bontcheva, Jan Hajic, Victoria Arranz, Andrejs Vasiļjevs, Gerhard Backfried, Jose Manuel Gomez-Perez, Ulrich Germann, Rémi Calizzano, Nils Feldhus, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Julian Moreno-Schneider, Dimitris Galanis, Penny Labropoulou, Miltos Deligiannis, Katerina Gkirtzou, Athanasia Kolovou, Dimitris Gkoumas, Leon Voukoutis, Ian Roberts, Jana Hamrlova, Dusan Varis, Lukas Kacena, Khalid Choukri, Valérie Mapelli, Mickaël Rigault, Julija Melnika, Miro Janosik, Katja Prinz, Andres Garcia-Silva, Cristian Berrio, Ondrej Klejch, Steve Renals
| Challenge: | Europe is a multilingual society, in which dozens of languages are spoken. |
| Approach: | They describe the European Language Grid, which is targeted to evolve into the primary platform and marketplace for LT in Europe by providing one umbrella platform for the European LT landscape. |
| Outcome: | The European Language Grid (ELG) will provide access to 1300 services for all European languages as well as thousands of data sets. |
A Browser-based Open Source Assistant for Multimodal Content Verification (2026.eacl-demo)
Copied to clipboard
Rosanna Milner, Michael Foster, Twin Karmakharm, Olesya Razuvayevskaya, Valentin Porcellini, Denis Teyssou, Ian Roberts, Kalina Bontcheva
| Challenge: | Disinformation and advanced generative AI content pose a significant challenge for journalists and fact-checkers who must rapidly verify digital media. |
| Approach: | They propose to integrate a browser-based tool that automatically extracts content from a suite of backend NLP classifiers and presents actionable credibility signals and AI-generation likelihood in an easy-to-digest format. |
| Outcome: | The Verification Assistant is a browser-based tool that extracts content and routes it to a suite of backend NLP classifiers, presenting actionable credibility signals, AI-generation likelihood, and other verification advice in an easy-to-digest format. |
Toxic Language Detection in Social Media for Brazilian Portuguese: New Dataset and Multilingual Analysis (2020.aacl-main)
Copied to clipboard
| Challenge: | Hate speech and toxic comments are a common concern of social media platform users . identifying toxic comments is important for studying and preventing the proliferation of toxicity in social media. |
| Approach: | They propose to use Brazilian Portuguese to analyze toxic or non-toxic tweets . they propose to analyze tweets as toxic or in different types of toxicity . |
| Outcome: | The proposed model achieves 76% macro-F1 score using monolingual data in the binary case. |
Analysing State-Backed Propaganda Websites: a New Dataset and Linguistic Study (2023.emnlp-main)
Copied to clipboard
| Challenge: | a network of doppelganger websites (impersonating genuine news sites) was discovered in 2022 . a novel dataset enables studies of disinformation networks and the training of NLP tools for disinformation detection. |
| Approach: | They analyze two hitherto unstudied sites sharing state-backed disinformation . they perform cross-site topic clustering and perform linguistic and temporal analysis . |
| Outcome: | The proposed dataset includes 14,053 articles, annotated with each language version, and additional metadata such as links and images. |
Journalist-in-the-Loop: Continuous Learning as a Service for Rumour Analysis (D19-3)
Copied to clipboard
| Challenge: | Existing rumour analysis tools do not scale due to the large volume and velocity of user generated content. |
| Approach: | They propose to use a web-based rumour analysis tool that can continuously learn from journalists and integrate it into existing tools and platforms. |
| Outcome: | The proposed system can be easily integrated as a service into existing tools and platforms used by journalists using a REST API. |
European Language Grid: An Overview (2020.lrec-1)
Copied to clipboard
Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitris Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, David Jones, Ian Roberts, Jan Hajič, Jana Hamrlová, Lukáš Kačena, Khalid Choukri, Victoria Arranz, Andrejs Vasiļjevs, Orians Anvari, Andis Lagzdiņš, Jūlija Meļņika, Gerhard Backfried, Erinç Dikici, Miroslav Janosik, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuel Gómez-Pérez, Andres Garcia Silva, Christian Berrío, Ulrich Germann, Steve Renals, Ondrej Klejch
| Challenge: | European LT business is dominated by hundreds of SMEs and a few large players, with technologies that outperform the global players. |
| Approach: | European Language Grid (ELG) project addresses this by establishing the ELG as the primary platform for LT in Europe. |
| Outcome: | European Language Grid (ELG) will be primary platform for LT in Europe . it will provide access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. |
Measuring the Impact of Readability Features in Fake News Detection (2020.lrec-1)
Copied to clipboard
Roney Santos, Gabriela Pedro, Sidney Leal, Oto Vale, Thiago Pardo, Kalina Bontcheva, Carolina Scarton
| Challenge: | Recent efforts to detect fake news use language-based approaches to detect news articles . authors show that readability features can improve classification accuracy . |
| Approach: | They propose to use readability features to detect fake news in the Brazilian Portuguese language . they show that such features can achieve up to 92% classification accuracy . |
| Outcome: | The proposed features achieve up to 92% accuracy and may improve previous classification results. |
Measuring What Counts: The Case of Rumour Stance Classification (2020.aacl-main)
Copied to clipboard
| Challenge: | Numerous methods have been proposed to predict the stance of replies towards a given rumour, but their performance is not optimal for the four-class imbalanced task of rumor stance classification. |
| Approach: | They propose to use a four-class problem to predict the stance of replies towards a given rumour to help identify the most informative minority classes. |
| Outcome: | The proposed methods are robust to imbalanced data and score higher systems capable of recognising the two most informative minority classes (support and deny). |
It’s about Time: Rethinking Evaluation on Rumor Detection Benchmarks using Chronological Splits (2023.findings-eacl)
Copied to clipboard
| Challenge: | Current rumor detection benchmarks use random splits as training, development and test sets which results in topical overlaps. |
| Approach: | They propose to use chronological rather than random splits for rumor classification . they propose to always use chronological splits to minimize topical overlaps . |
| Outcome: | The proposed model overestimates performance on four popular rumor detection benchmarks considering chronological instead of random splits. |
Large Language Models Offer an Alternative to the Traditional Approach of Topic Modelling (2024.lrec-main)
Copied to clipboard
| Challenge: | Topic modelling has found extensive use in automatically detecting significant topics within a corpus of documents, but there are certain drawbacks. |
| Approach: | They propose a framework that prompts large language models to generate topics from a given set of documents and establish evaluation protocols to assess the clustering efficacy of LLMs. |
| Outcome: | The proposed model generates relevant topic titles and adheres to human guidelines to refine and merge topics. |
Don’t waste a single annotation: improving single-label classifiers through soft labels (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for annotating data are limited by ambiguity and lack of context in data samples. |
| Approach: | They challenge the traditional approach of annotating data by only providing a single label for each sample and annotator disagreement is discarded . instead, they use additional annotation information such as confidence, secondary label and disagreement to generate soft labels. |
| Outcome: | The proposed method improves model performance and calibration on the hard label test set. |
Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science (2024.lrec-main)
Copied to clipboard
Yida Mu, Ben P. Wu, William Thorne, Ambrose Robinson, Nikolaos Aletras, Carolina Scarton, Kalina Bontcheva, Xingyi Song
| Challenge: | Existing instruction-tuned Large Language Models (LLMs) have impressive language understanding and the capacity to generate responses that follow specific prompts. |
| Approach: | They evaluate the zero-shot performance of two publicly accessible LLMs, ChatGPT and OpenAssistant, in the context of six Computational Social Science classification tasks. |
| Outcome: | The proposed LLMs perform better than state-of-the-art models on social science tasks. |
Examining Temporalities on Stance Detection towards COVID-19 Vaccination (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have highlighted the importance of vaccination as an effective strategy to control the transmission of the COVID-19 virus. |
| Approach: | They evaluate a range of transformer-based models using chronological and random splits of social media data to examine the impact of temporal concept drift on stance detection towards COVID-19 vaccination. |
| Outcome: | The proposed models show that the models performed better with chronological and random splits than with random split models. |
GATE Teamware 2: An open-source tool for collaborative document classification annotation (2023.eacl-demo)
Copied to clipboard
| Challenge: | GATE Teamware 2 is an open-source web-based platform for managing teams of annotators working on document classification tasks. |
| Approach: | They present GATE Teamware 2: an open-source web-based platform for managing teams of annotators working on document classification tasks. |
| Outcome: | GATE Teamware 2 is an open-source web-based platform for managing teams of annotators working on document classification tasks. |
Examining the Limitations of Computational Rumor Detection Models Trained on Static Datasets (2024.lrec-main)
Copied to clipboard
| Challenge: | Past research has indicated that content-based rumor detection models perform less effectively on unseen rumors. |
| Approach: | They propose to use data split strategies to minimize the effects of temporal concept drift in static datasets during the training of rumor detection methods. |
| Outcome: | The proposed model over-relys on the information derived from the rumors’ source post and overlooks the significant role that contextual information can play. |