HebID: Detecting Social Identities in Hebrew-language Political Text (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing NLP datasets focus on coarse-grained identity categories . existing datasets are mostly English-centric and focus on fine-grain categories based on cultural contexts. |
| Approach: | They introduce the first multilabel Hebrew corpus for social identity detection . they use Hebrew-tuned encoders alongside 2B-9B-parameter decoders . |
| Outcome: | The proposed classifier is based on a national public survey and uses Hebrew-tuned encoders to analyze political discourse and political speeches. |
Similar Papers
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)
Copied to clipboard
| Challenge: | Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us . |
| Approach: | They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages. |
| Outcome: | The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning. |
Framing Political Bias in Multilingual LLMs Across Pakistani Languages (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) shape public discourse, yet most evaluations of economic and political bias focus on high-resource Western languages and contexts. |
| Approach: | They propose to use a culturally adapted Political Compass Test to evaluate political bias in 13 state-of-the-art LLMs across five Pakistani languages. |
| Outcome: | The proposed framework captures ideological stance (economic/social axes) and stylistic framing (content, tone, emphasis) in 13 state-of-the-art LLMs across five Pakistani languages. |
The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech (2025.emnlp-main)
Copied to clipboard
| Challenge: | a new computational model for political delegitimization discourse is proposed for analysis of democratic discourse . we identify the importance of PDD as a powerful tool in political competition . |
| Approach: | They propose a computational classification pipeline for political delegitimization discourse . they annotate a Hebrew-language corpus of 10,410 sentences from parliamentary speeches, facebook posts and leading news outlets . |
| Outcome: | The proposed model achieves an F1 of 0.74 for binary detection and a macro-F1 of 0.6 for classification of delegitimization characteristics. |
Developing A Multilabel Corpus for the Quality Assessment of Online Political Talk (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index . |
| Approach: | They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries . |
| Outcome: | The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english . |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Social biases manifest in language agency, but there is no comprehensive benchmark for evaluating such biase in language models. |
| Approach: | They propose a benchmark to evaluate language agency biases in large language models . they propose 'Mitigation via Selective Rewrite' to selectively revise parts of generated texts . |
| Outcome: | The proposed language agency bias evaluation benchmark identifies gender, racial, and intersectional biases in 3 recent LLMs. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
IsraParlTweet: The Israeli Parliamentary and Twitter Resource (2024.lrec-main)
Copied to clipboard
| Challenge: | IsraParlTweet is a linked corpus of parliamentary discussions from the Knesset between 1992-2023 and Twitter posts made by Members of the Kneset between 2008-2023. |
| Approach: | They propose a linked corpus of parliamentary discussions from the Knesset between 1992-2023 and Twitter posts made by Members of the Kneset between 2008-2023. |
| Outcome: | IsraParlTweet can be used to conduct quantitative and qualitative analyses and provide valuable insights into political discourse in Israel. |
The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses (2022.lrec-1)
Copied to clipboard
| Challenge: | Gender bias in natural language processing (NLP) applications has been receiving increasing attention, largely due to the lack of datasets and resources. |
| Approach: | They propose a corpus for gender identification and rewriting in contexts involving one or two target users with independent grammatical gender preferences. |
| Outcome: | The proposed corpus expands on Habash et al.'s Arabic Parallel Gender Corpus (APGC) by adding second person targets and increasing the total number of sentences over 6.5 times, reaching over 590K words. |
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |