Papers with CSS

19 papers
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks (2024.eacl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are not perfect generalists as they often underperform traditional fine-tuning methods.
Approach: They compare human-labeled and synthetically generated data in CSS classification tasks . they leverage large language models such as OpenAI's GPT-4 for zero-shot classification .
Outcome: The proposed models perform better on human-labeled data than synthetically augmented models on rare classes within multi-class tasks.
Community lexical access for an endangered polysynthetic language: An electronic dictionary for St. Lawrence Island Yupik (N19-4)

Copied to clipboard

Challenge: a new electronic dictionary for St. Lawrence Island Yupik is developed to facilitate language-learning on the island . the endangered language is spoken primarily on St. lisa's St.liss island, Alaska .
Approach: They propose a morphologically-aware electronic dictionary for St. Lawrence Island Yupik . the dictionary is set in an uncluttered interface and uses HTML, Javascript, and CSS .
Outcome: The proposed dictionary is set in an uncluttered interface and is available in English and in Yupik . it is based on the morphologically-aware version of the Badten et al. paper dictionary .
CSS: Combining Self-training and Self-supervised Learning for Few-shot Dialogue State Tracking (2022.aacl-short)

Copied to clipboard

Challenge: Existing few-shot dialogue state tracking (DST) methods transfer knowledge from labeled data into DST, but collecting large amount of labeles is laborious.
Approach: They propose a few-shot dialogue state tracking framework that integrates self-training and self-supervised learning methods into the framework.
Outcome: The proposed framework achieves competitive performance in several few-shot scenarios.
Identifying Power Relations in Conversations using Multi-Agent Social Reasoning (2025.naacl-short)

Copied to clipboard

Challenge: Existing approaches to understanding power relationships in conversations are based on task-specific supervised learning.
Approach: They propose a multi-agent social reasoning framework that leverages social science tools to generate and evaluate reasons from multiple perspectives and construct a factor graph for inference.
Outcome: The proposed framework outperforms standard prompting baselines on power dynamics in conversations.
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification (2026.acl-long)

Copied to clipboard

Challenge: generative large language models (LLMs) are used extensively for text classification in computational social science . conceptualization of categories to classify and using LLM predictions can tempt analysts to skip conceptualization altogether.
Approach: They argue that LLMs can tempt analysts to skip conceptualization altogether . they argue that conceptualization failures induce downstream inferential bias .
Outcome: The proposed model can tempt analysts to skip conceptualization altogether . the proposed model is a first-order concern in the LLM-era .
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis (2025.findings-acl)

Copied to clipboard

Challenge: Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding.
Approach: They propose a framework that aligns synthesized speech with the emotional context of user-agent interactions to achieve empathy.
Outcome: The proposed framework produces more expressive speech than existing methods on three datasets.
Transc&Anno: A Graphical Tool for the Transcription and On-the-Fly Annotation of Handwritten Documents (L18-1)

Copied to clipboard

Challenge: Transc&Anno is a web-based collaboration tool for linguists to facilitate the transcription of text images and their shallow on-the-fly annotation.
Approach: They propose a web-based collaboration tool that allows the transcription of text images and their shallow on-the-fly annotation.
Outcome: The Transc&Anno tool can be used for any type of corpora requiring transcription and shallow on-the-fly annotation resulting in inline XML.
Can Unconfident LLM Annotations Be Used for Confident Conclusions? (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown high agreement with human raters across a variety of tasks, demonstrating potential to ease the challenges of human data collection.
Approach: They propose a method that combines LLM annotations and LLM confidence indicators to strategically select which human annotations to use.
Outcome: The proposed method produces accurate estimates and valid confidence intervals while reducing the number of human annotations by over 25%.
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue.
Approach: They propose two benchmarks that provide a framework for evaluating large language models for mental health support.
Outcome: The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments.
Improving Neural Political Statement Classification with Class Hierarchical Information (2022.findings-acl)

Copied to clipboard

Challenge: skewed classification of fine-grained categories in text-based computational social science is challenging on the NLP side.
Approach: They propose to use hierarchical relations among categories in codebooks to create constraints on the learned model.
Outcome: The proposed model improves on two datasets and multiple languages.
Can Large Language Models Mine Interpretable Financial Factors More Effectively? A Neural-Symbolic Factor Mining Agent Model (2024.findings-acl)

Copied to clipboard

Challenge: Existing factor mining models are inefficient and inefficient, resulting in a significant challenge to extract interpretable factors.
Approach: They propose a model that integrates the strengths of both neural and symbolic models for factor mining.
Outcome: The proposed model surpasses the SOTA RankIC and RankICIR in predicting S&P 500 returns on real-world stock market data.
Fake Alignment: Are LLMs Really Aligned Well? (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies on large language models have shown that they are poorly aligned in practice.
Approach: They propose a framework to evaluate safety in large language models . they propose two new metrics to quantify fake alignment and obtain corrected performance estimation.
Outcome: The proposed framework and two metrics show that some models with purported safety are poorly aligned in practice.
NLP Systems That Can’t Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps (2024.naacl-long)

Copied to clipboard

Challenge: Existing language models fail to distinguish use from mention, leading to misinformation and hate speech detection, resulting in censorship of counterspeech.
Approach: They propose prompting mitigations that teach the use-mention distinction and show they reduce these errors.
Outcome: The proposed model reduces misinformation and hate speech detection errors by reducing misinformation, and reducing hate speech.
Masking Actor Information Leads to Fairer Political Claims Detection (2020.acl-main)

Copied to clipboard

Challenge: In recent years, NLP methods have found increasing adoption in the social sciences . however, CSS must be crucially interested in the algorithmic fairness of the underlying methods .
Approach: They propose two methods which mask proper names and pronouns during training of the model, thus removing personal information bias.
Outcome: The proposed methods decrease frequency bias while keeping the overall performance stable.
CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset (2023.findings-acl)

Copied to clipboard

Challenge: a cross-domain text-to-SQL task aims to parse user questions into SQL on complete unseen databases . a single-domain task evaluates the performance on identical databases based on the same domain .
Approach: They propose a cross-domain text-to-SQL task that parses user questions into SQL on unseen databases.
Outcome: The proposed system can parse user questions into SQL on complete unseen databases.
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods overlook the fine-grained semantic and prosodic interaction modeling at the word level.
Approach: They propose a novel approach to generate conversational prosody by understanding multimodal dialogue history (MDH) using fine-grained semantic and prosodic interaction modeling, they construct specialized multimodal fine-grain dialogue interaction graphs that encode interaction between word-level semantics and prosody.
Outcome: The proposed system outperforms baseline models in terms of prosodic expressiveness.
A New Dataset and Empirical Study for Sentence Simplification in Chinese (2023.acl-long)

Copied to clipboard

Challenge: Sentence simplification is a valuable technique that can benefit language learners and children.
Approach: They propose a dataset for assessing sentence simplification in Chinese using manual simplifications from human annotators.
Outcome: The proposed dataset shows that Chinese sentences are more accessible to children and nonnative readers than English sentences.
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)

Copied to clipboard

Challenge: a large corpus of German newspaper articles is available for free in other languages, such as English.
Approach: They propose to use taz2024full to analyse gender representation across four decades of reporting.
Outcome: The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies.
Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science Research (2024.emnlp-main)

Copied to clipboard

Challenge: Social media data exhibits distinctive characteristics such as rapid and continual topic evolution.
Approach: They propose new protocols and best practices for improving dataset development from social media data and its usage.
Outcome: The proposed protocols and best practices improve the performance of social media datasets and their usage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations