Papers with CSS
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks (2024.eacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are not perfect generalists as they often underperform traditional fine-tuning methods. |
| Approach: | They compare human-labeled and synthetically generated data in CSS classification tasks . they leverage large language models such as OpenAI's GPT-4 for zero-shot classification . |
| Outcome: | The proposed models perform better on human-labeled data than synthetically augmented models on rare classes within multi-class tasks. |
Community lexical access for an endangered polysynthetic language: An electronic dictionary for St. Lawrence Island Yupik (N19-4)
Copied to clipboard
| Challenge: | a new electronic dictionary for St. Lawrence Island Yupik is developed to facilitate language-learning on the island . the endangered language is spoken primarily on St. lisa's St.liss island, Alaska . |
| Approach: | They propose a morphologically-aware electronic dictionary for St. Lawrence Island Yupik . the dictionary is set in an uncluttered interface and uses HTML, Javascript, and CSS . |
| Outcome: | The proposed dictionary is set in an uncluttered interface and is available in English and in Yupik . it is based on the morphologically-aware version of the Badten et al. paper dictionary . |
CSS: Combining Self-training and Self-supervised Learning for Few-shot Dialogue State Tracking (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing few-shot dialogue state tracking (DST) methods transfer knowledge from labeled data into DST, but collecting large amount of labeles is laborious. |
| Approach: | They propose a few-shot dialogue state tracking framework that integrates self-training and self-supervised learning methods into the framework. |
| Outcome: | The proposed framework achieves competitive performance in several few-shot scenarios. |
Identifying Power Relations in Conversations using Multi-Agent Social Reasoning (2025.naacl-short)
Copied to clipboard
| Challenge: | Existing approaches to understanding power relationships in conversations are based on task-specific supervised learning. |
| Approach: | They propose a multi-agent social reasoning framework that leverages social science tools to generate and evaluate reasons from multiple perspectives and construct a factor graph for inference. |
| Outcome: | The proposed framework outperforms standard prompting baselines on power dynamics in conversations. |
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification (2026.acl-long)
Copied to clipboard
| Challenge: | generative large language models (LLMs) are used extensively for text classification in computational social science . conceptualization of categories to classify and using LLM predictions can tempt analysts to skip conceptualization altogether. |
| Approach: | They argue that LLMs can tempt analysts to skip conceptualization altogether . they argue that conceptualization failures induce downstream inferential bias . |
| Outcome: | The proposed model can tempt analysts to skip conceptualization altogether . the proposed model is a first-order concern in the LLM-era . |
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis (2025.findings-acl)
Copied to clipboard
| Challenge: | Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. |
| Approach: | They propose a framework that aligns synthesized speech with the emotional context of user-agent interactions to achieve empathy. |
| Outcome: | The proposed framework produces more expressive speech than existing methods on three datasets. |
Transc&Anno: A Graphical Tool for the Transcription and On-the-Fly Annotation of Handwritten Documents (L18-1)
Copied to clipboard
| Challenge: | Transc&Anno is a web-based collaboration tool for linguists to facilitate the transcription of text images and their shallow on-the-fly annotation. |
| Approach: | They propose a web-based collaboration tool that allows the transcription of text images and their shallow on-the-fly annotation. |
| Outcome: | The Transc&Anno tool can be used for any type of corpora requiring transcription and shallow on-the-fly annotation resulting in inline XML. |
Can Unconfident LLM Annotations Be Used for Confident Conclusions? (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown high agreement with human raters across a variety of tasks, demonstrating potential to ease the challenges of human data collection. |
| Approach: | They propose a method that combines LLM annotations and LLM confidence indicators to strategically select which human annotations to use. |
| Outcome: | The proposed method produces accurate estimates and valid confidence intervals while reducing the number of human annotations by over 25%. |
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)
Copied to clipboard
Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Prathiba Dhanesh, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi
| Challenge: | Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue. |
| Approach: | They propose two benchmarks that provide a framework for evaluating large language models for mental health support. |
| Outcome: | The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments. |
Improving Neural Political Statement Classification with Class Hierarchical Information (2022.findings-acl)
Copied to clipboard
Erenay Dayanik, Andre Blessing, Nico Blokker, Sebastian Haunss, Jonas Kuhn, Gabriella Lapesa, Sebastian Pado
| Challenge: | skewed classification of fine-grained categories in text-based computational social science is challenging on the NLP side. |
| Approach: | They propose to use hierarchical relations among categories in codebooks to create constraints on the learned model. |
| Outcome: | The proposed model improves on two datasets and multiple languages. |
Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing factor mining models are inefficient and inefficient, resulting in a significant challenge to extract interpretable factors. |
| Approach: | They propose a model that integrates the strengths of both neural and symbolic models for factor mining. |
| Outcome: | The proposed model surpasses the SOTA RankIC and RankICIR in predicting S&P 500 returns on real-world stock market data. |
Fake Alignment: Are LLMs Really Aligned Well? (2024.naacl-long)
Copied to clipboard
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, Yingchun Wang
| Challenge: | Existing studies on large language models have shown that they are poorly aligned in practice. |
| Approach: | They propose a framework to evaluate safety in large language models . they propose two new metrics to quantify fake alignment and obtain corrected performance estimation. |
| Outcome: | The proposed framework and two metrics show that some models with purported safety are poorly aligned in practice. |
NLP Systems That Can’t Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing language models fail to distinguish use from mention, leading to misinformation and hate speech detection, resulting in censorship of counterspeech. |
| Approach: | They propose prompting mitigations that teach the use-mention distinction and show they reduce these errors. |
| Outcome: | The proposed model reduces misinformation and hate speech detection errors by reducing misinformation, and reducing hate speech. |
Masking Actor Information Leads to Fairer Political Claims Detection (2020.acl-main)
Copied to clipboard
| Challenge: | In recent years, NLP methods have found increasing adoption in the social sciences . however, CSS must be crucially interested in the algorithmic fairness of the underlying methods . |
| Approach: | They propose two methods which mask proper names and pronouns during training of the model, thus removing personal information bias. |
| Outcome: | The proposed methods decrease frequency bias while keeping the overall performance stable. |
CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset (2023.findings-acl)
Copied to clipboard
| Challenge: | a cross-domain text-to-SQL task aims to parse user questions into SQL on complete unseen databases . a single-domain task evaluates the performance on identical databases based on the same domain . |
| Approach: | They propose a cross-domain text-to-SQL task that parses user questions into SQL on unseen databases. |
| Outcome: | The proposed system can parse user questions into SQL on complete unseen databases. |
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods overlook the fine-grained semantic and prosodic interaction modeling at the word level. |
| Approach: | They propose a novel approach to generate conversational prosody by understanding multimodal dialogue history (MDH) using fine-grained semantic and prosodic interaction modeling, they construct specialized multimodal fine-grain dialogue interaction graphs that encode interaction between word-level semantics and prosody. |
| Outcome: | The proposed system outperforms baseline models in terms of prosodic expressiveness. |
A New Dataset and Empirical Study for Sentence Simplification in Chinese (2023.acl-long)
Copied to clipboard
| Challenge: | Sentence simplification is a valuable technique that can benefit language learners and children. |
| Approach: | They propose a dataset for assessing sentence simplification in Chinese using manual simplifications from human annotators. |
| Outcome: | The proposed dataset shows that Chinese sentences are more accessible to children and nonnative readers than English sentences. |
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)
Copied to clipboard
| Challenge: | a large corpus of German newspaper articles is available for free in other languages, such as English. |
| Approach: | They propose to use taz2024full to analyse gender representation across four decades of reporting. |
| Outcome: | The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies. |
Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science Research (2024.emnlp-main)
Copied to clipboard
| Challenge: | Social media data exhibits distinctive characteristics such as rapid and continual topic evolution. |
| Approach: | They propose new protocols and best practices for improving dataset development from social media data and its usage. |
| Outcome: | The proposed protocols and best practices improve the performance of social media datasets and their usage. |