Papers by Petr Sojka

5 papers
Concept-aware Data Construction Improves In-context Learning of Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent work curating in-context learners assumes that ICL emerges from vast over-parametrization or the scale of multitask training.
Approach: They propose a framework for constructing training scenarios that make it beneficial for the LM to learn to utilize the analogical reasoning concepts from demonstrations.
Outcome: The proposed framework makes it beneficial for the LM to learn to utilize the analogical reasoning concepts from demonstrations and fares comparably to previous in-context learners trained in large-scale multitask learning requiring magnitudes of more training data.
Soft Alignment Objectives for Robust Adaptation of Language Generation (2023.acl-long)

Copied to clipboard

Challenge: Domain adaptation is a common approach for generative language models, but it is notorious for over-specialization to the target domain, resulting in catastrophic forgetting.
Approach: They propose to build training objectives on a semantic similarity of predicted tokens to the reference and avoid catastrophic forgetting of adaptation by preserving adaptation in-domain quality.
Outcome: The proposed objectives mitigate catastrophic forgetting while preserving the adaptation in-domain quality while reducing computational costs.
Adaptor: Objective-Centric Adaptation Framework for Language Models (2022.acl-demo)

Copied to clipboard

Challenge: Adaptor library aims to simplify complex training processes requiring customizations.
Approach: They introduce Adaptor library which transposes traditional model-centric approach to objective-centric training pipeline with Objective as central abstraction.
Outcome: The proposed framework simplifies training processes and improves reproducibility.
Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models (2024.eacl-long)

Copied to clipboard

Challenge: Existing work shows that Large Language Models (LLMs) are not robust to complex language understanding tasks due to reliance on spurious correlations of training datasets.
Approach: They propose a method for measuring model reliance on spurious features by exploiting chosen biases on out-of-distribution (OOD) datasets.
Outcome: The proposed method shows that the reported OOD gains of debiasing methods can't be explained by mitigated reliance on biased features, suggesting that biases are shared among different QA datasets.
Towards the Roots of the Negation Problem: A Multilingual NLI Dataset and Model Scaling Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: Negations are key to determining sentence meaning, making them essential for logical reasoning.
Approach: They construct and publish two new textual entailment datasets in four languages with paired examples differing in negation.
Outcome: The results show that increasing the model size may improve the models’ ability to handle negations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations