Papers by David Dobolyi
Textagon: Boosting Language Models with Theory-guided Parallel Representations (2025.acl-demo)
Copied to clipboard
| Challenge: | Pretrained language models do not account for the wide variety of available expert-generated language resources and lexicons that explicitly encode linguistic/domain knowledge. |
| Approach: | They propose a Python package for generating parallel representations for text based on predefined lexicons and selecting representations that provide the most information. |
| Outcome: | The proposed model can generate parallel representations of text based on predefined lexicons and select representations that provide the most information. |
Constructing a Psychometric Testbed for Fair Natural Language Processing (2021.emnlp-main)
Copied to clipboard
| Challenge: | Psychometric dimensions are important for understanding user behavior in various contexts including health, security, e-commerce, and finance. |
| Approach: | They propose to construct a corpus for psychometric natural language processing related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain. |
| Outcome: | The proposed corpus includes 8,502 user-generated responses from 8,502-person survey datasets and includes self-reported demographic information, including race, sex, age, income, and education. |