Papers by Agnieszka Karlińska

2 papers
PLLuM-Align: Polish Preference Dataset for Large Language Model Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models generate preferred responses while avoiding harmful or inappropriate outputs, despite their ability to generate cross-language transferability.
Approach: They introduce the first Polish preference dataset PLLuM-Align, created entirely through human annotation to reflect Polish language and cultural nuances.
Outcome: The proposed dataset lays the groundwork for more aligned Polish LLMs and contributes to the broader goal of multilingual alignment in underrepresented languages.
Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse (2025.acl-long)

Copied to clipboard

Challenge: specialized Polish language models are more effective at detecting harmful content than traditional methods.
Approach: They propose a Polish-language dataset for erotic content detection that captures ambiguity, violence, and socially unacceptable behaviors.
Outcome: The proposed dataset shows that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations