Papers by Samuel Weinbach
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings (2024.emnlp-main)
Copied to clipboard
| Challenge: | Tokenizers are crucial for encoding information in Large Language Models, but their development has stagnated. |
| Approach: | They propose a tokenizer that embeds words through sparse activation patterns over character triplets . they show competitive downstream performance with a parameter reduction of more than 85% . |
| Outcome: | The proposed approach achieves competitive downstream performance with a parameter reduction of more than 85% on embedding layers. |
Tokenizer Choice For LLM Training: Negligible or Crucial? (2024.findings-naacl)
Copied to clipboard
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, Nicolas Flores-Herr
| Challenge: | Recent success of large language models has been driven by curating the training dataset composition, scaling of model architectures and advancements in pretraining objectives, leaving tokenizer influence as a blind spot. |
| Approach: | They conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale. |
| Outcome: | The proposed model can significantly impact the model's downstream performance and training costs. |
MAGMA – Multimodal Augmentation of Generative Models through Adapter-based Finetuning (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale pretraining is becoming the norm in Vision-Language (VL) modeling. |
| Approach: | They propose a method for augmenting generative language models with additional modalities using adapter-based finetuning. |
| Outcome: | The proposed method outperforms Frozen on open-ended generative tasks while maintaining the language model weights. |