Papers by João Sedoc
An “Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives” (2023.emnlp-main)
Copied to clipboard
| Challenge: | Mental health conversational agents (a.k.a. chatbots) are widely studied for their potential to offer accessible support to those experiencing mental health challenges. |
| Approach: | They review 534 papers on building mental health-related conversational agents . they recommend a few recommendations to bridge the disciplinary divide . |
| Outcome: | The systematic review reveals 136 key papers on building mental health-related conversational agents with diverse characteristics of modeling and experimental design techniques. |
SMRT Chatbots: Improving Non-Task-Oriented Dialog with Simulated Multiple Reference Training (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Simulated Multiple Reference Training (SMRT) improves non-task-oriented dialog models by reducing the need for related-domain dialog data. |
| Approach: | They apply Simulated Multiple Reference Training (SMRT) to chatbots to overcome sparse dialog data. |
| Outcome: | The proposed model outperforms pretraining on human evaluation quality and lexical diversity without requiring related-domain dialog data. |
Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas (2024.findings-emnlp)
Copied to clipboard
Salvatore Giorgi, Tingting Liu, Ankit Aich, Kelsey Isman, Garrick Sherman, Zachary Fried, João Sedoc, Lyle Ungar, Brenda Curtis
| Challenge: | Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. |
| Approach: | They propose to prompt LLMs with human-like personas and ask them to answer as if they were a specific human, either explicitly, with exact demographics, political beliefs, and lived experiences, or implicitly via names prevalent in specific populations. |
| Outcome: | The proposed model is based on explicit, explicit, and implicit personas, and fails to show implicit biases. |
Modeling Empathy and Distress in Reaction to News Stories (D18-1)
Copied to clipboard
| Challenge: | a recent work on empathy prediction has underestimated the complexity of the phenomenon and lacks a shared corpus. authors present a novel annotation methodology which reliably captures empathy assessments by the writer of a statement using multi-item scales. |
| Approach: | They propose a method which captures empathy assessments by the writer of a statement using multi-item scales. |
| Outcome: | The proposed method distinguishes between multiple forms of empathy, empathic concern, and personal distress, as recognized throughout psychology. |
A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization (2023.acl-long)
Copied to clipboard
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc
| Challenge: | Using crowdsourcing, it is difficult to obtain high-quality annotations for difficult tasks. |
| Approach: | They propose a recruitment pipeline to recruit high-quality Amazon Mechanical Turk workers . they filter out subpar workers before they carry out the evaluations . |
| Outcome: | The proposed method can filter out subpar workers before they carry out evaluations and obtain high-agreement annotations with similar constraints on resources. |
Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also be due . poor experimental practice can be mitigated by bringing multiple comparable studies together in systematic reviews that draw conclusions beyond the level of the individual studies. |
| Approach: | They propose to assess NLP/ML practitioners' views and experience of reproducibility over the past two years. |
| Outcome: | The results of two identical surveys show that views and experience of reproducibility have changed over the past two years. |
Automatic Document Selection for Efficient Encoder Pretraining (2022.emnlp-main)
Copied to clipboard
| Challenge: | Pretraining language models is expensive and data-intensive, but can it be improved? Several studies have found that directly pretraining on task data is more effective . |
| Approach: | They propose to automatically identify smaller yet domain-representative subsets by pretraining a model on a target domain. |
| Outcome: | The proposed method outperforms random selection on perplexity and downstream tasks with 20x less data and 3x fewer training iterations and 2x less estimated cloud compute cost. |
Measuring the Language of Self-Disclosure across Corpora (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing models that estimate self-disclosure from language are poorly generalized due to variations in corpora and labeling instructions. |
| Approach: | They build single-task models on five self-disclosure corpora and use them to predict self-declaration across corpors. |
| Outcome: | The proposed model predicts self-disclosure across corpora, but the results are poor for out-of-corpora models. |
Conceptor-Aided Debiasing of Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained large language models reflect inherent social biases of their training corpus. |
| Approach: | They propose two methods to identify and remove the bias subspace in pre-trained large language models such as BERT and GPT by applying conceptors to a conceptor NOT operation. |
| Outcome: | The proposed method achieves state-of-the-art (SoTA) debiasing while maintaining LLMs’ performance on the GLUE benchmark. |
Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification (N19-1)
Copied to clipboard
Reno Kriz, João Sedoc, Marianna Apidianaki, Carolina Zheng, Gaurav Kumar, Eleni Miltsakaki, Chris Callison-Burch
| Challenge: | Recent research has applied sequence-to-sequence (Seq2Sequen) models to text simplification . generic models tend to copy directly from the original sentence, resulting in outputs that are long and complex. |
| Approach: | They propose to incorporate word complexities into the loss function during training and generate a large set of diverse candidate simplifications at test time. |
| Outcome: | The proposed model can perform competitively with state-of-the-art systems while generating simpler sentences. |
ChatEval: A Tool for Chatbot Evaluation (N19-4)
Copied to clipboard
| Challenge: | open-domain dialog systems are difficult to evaluate due to lack of standardization and standardization in evaluation procedures. |
| Approach: | They propose a framework for human evaluation of chatbots that augments existing tools . researchers can submit their trained models to the ChatEval web interface . reproducibility and model assessment for opendomain dialog systems is challenging . |
| Outcome: | The proposed framework provides a web-based hub for researchers to compare their models with baselines and prior work. |
Large Human Language Models: A Need and the Challenges (2024.naacl-long)
Copied to clipboard
| Challenge: | a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences . |
| Approach: | et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group . |
| Outcome: | a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say . |
Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG (2026.acl-long)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains where safety depends on how a system answers . out-of-domain (OOD) queries can impair performance and safety . |
| Approach: | They propose to use lightweight, KB-aligned OOD detection as an always-on gate for RAG systems. |
| Outcome: | The proposed method scores queries in a compact subspace selected either by explained-variance retention (EVR) or by a separability-driven -test ranking. |
On the Role of Summary Content Units in Text Summarization Evaluation (2024.naacl-short)
Copied to clipboard
Marcel Nawrath, Agnieszka Nowak, Tristan Ratz, Danilo Walenta, Juri Opitz, Leonardo Ribeiro, João Sedoc, Daniel Deutsch, Simon Mille, Yixin Liu, Sebastian Gehrmann, Lining Zhang, Saad Mahamood, Miruna Clinciu, Khyathi Chandu, Yufang Hou
| Challenge: | a human written summary content unit (SCU) is used to judge the quality of a summary . a pyramid evaluation method is based on SCUs that decompose a reference summary into concise sentences . |
| Approach: | They propose to use automated SCUs to evaluate the quality of a candidate summary . they propose to generate SCU approximations from AMR meaning representations and large language models . |
| Outcome: | The proposed method can be fully automated, but lacks the human effort to validate it. |
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)
Copied to clipboard
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina Mcmillan-major, Anna Shvets, Ashish Upadhyay, Bernd Bohnet, Bingsheng Yao, Bryan Wilie, Chandra Bhagavatula, Chaobin You, Craig Thomson, Cristina Garbacea, Dakuo Wang, Daniel Deutsch, Deyi Xiong, Di Jin, Dimitra Gkatzia, Dragomir Radev, Elizabeth Clark, Esin Durmus, Faisal Ladhak, Filip Ginter, Genta Indra Winata, Hendrik Strobelt, Hiroaki Hayashi, Jekaterina Novikova, Jenna Kanerva, Jenny Chim, Jiawei Zhou, Jordan Clive, Joshua Maynez, João Sedoc, Juraj Juraska, Kaustubh Dhole, Khyathi Raghavi Chandu, Laura Perez Beltrachini, Leonardo F . R. Ribeiro, Lewis Tunstall, Li Zhang, Mahim Pushkarna, Mathias Creutz, Michael White, Mihir Sanjay Kale, Moussa Kamal Eddine, Nico Daheim, Nishant Subramani, Ondrej Dusek, Paul Pu Liang, Pawan Sasanka Ammanamanchi, Qi Zhu, Ratish Puduppully, Reno Kriz, Rifat Shahriyar, Ronald Cardenas, Saad Mahamood, Salomey Osei, Samuel Cahyawijaya, Sanja Štajner, Sebastien Montella, Shailza Jolly, Simon Mille, Tahmid Hasan, Tianhao Shen, Tosin Adewumi, Vikas Raunak, Vipul Raheja, Vitaly Nikolaev, Vivian Tsai, Yacine Jernite, Ying Xu, Yisi Sang, Yixin Liu, Yufang Hou
| Challenge: | Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work. |
| Approach: | They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations. |
| Outcome: | The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work. |
COD3S: Diverse Generation with Discrete Semantic Signatures (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating semantically diverse sentences are based on locality-sensitive hash (LSH)-based semantic sentence codes that explicitly capture meaningful semantic differences. |
| Approach: | They propose a method for generating semantically diverse sentences using neural sequence-to-sequence models by conditioned on locality-sensitive hash-based semantic sentence codes whose Hamming distances correlate with human judgments of semantic textual similarity. |
| Outcome: | The proposed method improves output diversity without degrading performance on causal generation tasks. |
Benchmark Data and Evaluation Framework for Intent Discovery Around COVID-19 Vaccine Hesitancy (2023.findings-eacl)
Copied to clipboard
Shai Gretz, Assaf Toledo, Roni Friedman, Dan Lahav, Rose Weeks, Naor Bar-Zeev, João Sedoc, Pooja Sangha, Yoav Katz, Noam Slonim
| Challenge: | As COVID-19 vaccines were rolled out, they were met with widespread hesitancy. |
| Approach: | They propose a new framework for intent discovery that leverages existing intent classifiers to provide a real-world conversational dataset of conversations conducted by actual users with VIRA. |
| Outcome: | The proposed framework enables users to find out what they are doing and why they are hesitant. |
Comparison of Diverse Decoding Methods from Conditional Language Models (P19-1)
Copied to clipboard
| Challenge: | Conditional language models can generate a diverse set of outputs, but for open-ended tasks, beam search is ill-suited to generating a set of diverse sequences. |
| Approach: | They propose a method where we over-sample candidates and use clustering to remove similar sequences to achieve high diversity without sacrificing quality. |
| Outcome: | The proposed method over-samples candidates and removes similar sequences to achieve high diversity without sacrificing quality. |
Incremental Neural Coreference Resolution in Constant Memory (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on coreference resolution has focused on improving pairwise span scoring functions and methods for decoding into globally consistent clusters. |
| Approach: | They extend an incremental clustering algorithm to utilize contextualized encoders and neural components to generate a high-performing model. |
| Outcome: | The proposed model reduces memory usage to constant space with only a 0.3% relative loss in F1 on OntoNotes 5.0. |
Measuring the ‘I don’t know’ Problem through the Lens of Gricean Quantity (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing work on analyzing chatbots for generic, safe responses has not addressed the ‘I don’t know’ problem, but lack of analysis leaves it unclear if a method improves chatbot models by mitigating this problem, or another. |
| Approach: | They propose to use Relative Utterance Quantity to diagnose the ‘I don’t know’ problem, in which a dialog system produces generic responses. |
| Outcome: | The proposed method allows for the direct analysis of the ‘I don’t know’ problem, which has been addressed but not analyzed by prior work. |
Inducing Generalizable and Interpretable Lexica (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Lexica are widely used as generalizable language features to predict sentiment, emotions, mental health, and personality. |
| Approach: | They propose to induce lexica using context-oblivious and context-aware approaches and compare their performance using crowd-worker assessment. |
| Outcome: | The proposed models can be induced using context-oblivious and context-aware approaches and evaluate their quality using crowd-worker assessment. |
From Text to Context: Contextualizing Language with Humans, Groups, and Communities for Socially Aware NLP (2024.naacl-tutorials)
Copied to clipboard
Adithya V Ganesan, Siddharth Mangalik, Vasudha Varadarajan, Nikita Soni, Swanie Juhng, João Sedoc, H. Andrew Schwartz, Salvatore Giorgi, Ryan L Boyd
| Challenge: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Approach: | This tutorial will cover the latest techniques and libraries for doing so at each level of analysis. |
| Outcome: | The tutorial covers human-centered techniques that provide benefit to traditional document- or word-level NLP tasks. |
Item Response Theory for Natural Language Processing (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial introduces the wider NLP community to Item Response Theory (IRT) existing software for fitting IRT models is limited by human-data sized constraints. |
| Approach: | They will introduce IRT and the mathematical foundations which make IRT models. |
| Outcome: | This tutorial aims to introduce the wider NLP community to Item Response Theory and show its benefits for a number of NLP tasks. |
Common Law Annotations: Investigating the Stability of Dialog System Output Annotations (2023.findings-acl)
Copied to clipboard
Seunggun Lee, Alexandra DeLucia, Nikita Nangia, Praneeth Ganedi, Ryan Guan, Rubing Li, Britney Ngaw, Aditya Singhal, Shalaka Vaidya, Zijun Yuan, Lining Zhang, João Sedoc
| Challenge: | High agreement is often used to show reliability of annotation procedures, but it is insufficient to ensure or reproducibility. |
| Approach: | They propose a protocol that increases Inter-Annotator Agreement among annotators and a standardized and codified protocol that strictly enforces transparency in the annotation process. |
| Outcome: | The proposed protocol ensures transparency in the annotation process, which ensures reproducibility of annotation guidelines. |
Conditioning on Dialog Acts improves Empathy Style Transfer (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing research has focused on empathetic response generation, but it is not applicable to more sensitive cases such as medicine and therapy where the content of the responses requires the supervision of medical experts. |
| Approach: | They propose two new style transfer strategies that use only examples of the target style and dialog-act-conditioned prompting to make a sentence more empathetic. |
| Outcome: | The proposed methods improve empathy more effectively while maintaining semantic similarity and preserving both semantics and the dialog-act type. |
Continual Learning for Sentence Representations Using Conceptors (N19-1)
Copied to clipboard
| Challenge: | Existing sentence encoders for distributed representations of sentences are limited in their performance on fixed corpora. |
| Approach: | They propose a continual learning scenario for distributed representations of sentences . they initialize sentence encoders with corpus-independent features and update them sequentially . |
| Outcome: | The proposed sentence encoder can learn features from new corpora while maintaining its competence on previously encountered corporales. |
Learning Word Ratings for Empathy and Distress from Document-Level User Responses (2020.lrec-1)
Copied to clipboard
| Challenge: | Emotion analysis of text is increasing in popularity in NLP, however, manually creating lexica for psychological constructs such as empathy has proven difficult. |
| Approach: | They compare different approaches to learning word ratings from higher-level supervision and use a Mixed-Level Feed Forward Network to create the first-ever empathy lexicon. |
| Outcome: | The proposed model automatically creates empathy word ratings from document-level ratings. |