Papers by Lorena Calvo-Bartolomé
Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic Models (2025.acl-long)
Copied to clipboard
Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan Lee Boyd-Graber
| Challenge: | a common use of NLP is to facilitate the understanding of large document collections. |
| Approach: | They propose to use large language models to replace probabilistic topic models in real-world applications. |
| Outcome: | The proposed model generates more human-readable topics and shows higher average win probabilities than traditional models for data exploration. |
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification (2025.emnlp-demos)
Copied to clipboard
Chenfei Xiong, Jingwei Ni, Yu Fan, Vilém Zouhar, Donya Rooein, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Zhijing Jin, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Mennatallah El-Assady, Elliott Ash
| Challenge: | Social scientists often need to develop codebooks that can be reliable but require significant human effort. |
| Approach: | They propose a mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language models. |
| Outcome: | The proposed framework integrates human expertise with automatic annotation guided by large language models. |
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering (2025.emnlp-main)
Copied to clipboard
Lorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa, Jerónimo Arenas-García, Jordan Lee Boyd-Graber
| Challenge: | Multilingual question answering systems must ensure factual consistency across languages while also accounting for cultural variation in subjective responses. |
| Approach: | They propose a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. |
| Outcome: | The proposed tool detects factual and cultural discrepancies in bilingual question answering systems. |
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering (2025.acl-long)
Copied to clipboard
| Challenge: | Topic models and document clustering evaluations often use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. |
| Approach: | They propose a protocol for evaluating topic models and document clustering evaluations that uses crowdworker annotations to validate automated proxies. |
| Outcome: | The proposed protocol is scalable and easy to adapt to an LLM prompt. |
pAtChWoRK: Patching the Pieces of Public Procurement Documents (2026.acl-demo)
Copied to clipboard
| Challenge: | pAtChWoRK corrects manual classification errors and extracts complex unstructured fields such as award and solvency criteria and tenders’ objectives. |
| Approach: | pAtChWoRK corrects manual classification errors and extracts complex unstructured fields such as award and solvency criteria and tenders’ objectives. |
| Outcome: | pAtChWoRK corrects manual classification errors and extracts complex unstructured fields such as award and solvency criteria and tenders’ objectives. |