Papers by Samuel Weinbach

3 papers
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings (2024.emnlp-main)

Copied to clipboard

Challenge: Tokenizers are crucial for encoding information in Large Language Models, but their development has stagnated.
Approach: They propose a tokenizer that embeds words through sparse activation patterns over character triplets . they show competitive downstream performance with a parameter reduction of more than 85% .
Outcome: The proposed approach achieves competitive downstream performance with a parameter reduction of more than 85% on embedding layers.
Tokenizer Choice For LLM Training: Negligible or Crucial? (2024.findings-naacl)

Copied to clipboard

Challenge: Recent success of large language models has been driven by curating the training dataset composition, scaling of model architectures and advancements in pretraining objectives, leaving tokenizer influence as a blind spot.
Approach: They conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale.
Outcome: The proposed model can significantly impact the model's downstream performance and training costs.
MAGMA – Multimodal Augmentation of Generative Models through Adapter-based Finetuning (2022.findings-emnlp)

Copied to clipboard

Challenge: Large-scale pretraining is becoming the norm in Vision-Language (VL) modeling.
Approach: They propose a method for augmenting generative language models with additional modalities using adapter-based finetuning.
Outcome: The proposed method outperforms Frozen on open-ended generative tasks while maintaining the language model weights.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations