Papers by Tamara Czinczoll
Scientific and Creative Analogies in Pretrained Language Models (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing analogy datasets focus on a limited set of analogical relations with a high similarity of the two domains between which the analogy holds. |
| Approach: | They propose a dataset that encodes analogy in pretrained language models . they use a system that maps attributes and relational structures across dissimilar domains . |
| Outcome: | The proposed dataset shows that state-of-the-art models achieve low performance on analogy tasks . |
NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents (2024.acl-long)
Copied to clipboard
| Challenge: | (large) language models struggle to process long sequences due to the quadratic scaling of the underlying attention mechanism. |
| Approach: | They propose a Masked Language Model operating on higher-level semantic representations in the form of text embeddings to solve this problem. |
| Outcome: | The proposed model outperforms larger embedding models on three types of tasks. |