Papers by Phillip Rust
Text Rendering Strategies for Pixel Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent approaches to rendering text use a large set of almost-equivalent input patches, which may prove sub-optimal for downstream tasks due to redundancy in the input representations. |
| Approach: | They propose four approaches to rendering text in a PIXEL model using character bigrams and patch frequency biases. |
| Outcome: | The proposed models perform better on sentence-level tasks without compromising performance on token-level or multilingual tasks. |
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models (2021.acl-long)
Copied to clipboard
| Challenge: | Using pretraining data, we find that a designated monolingual tokenizer plays an equally important role in the downstream performance of the model. |
| Approach: | They propose to compare pretrained multilingual models with their monolingual counterparts on a set of five diverse monolingual downstream tasks. |
| Outcome: | The proposed models offer previously unmatched performance in all NLP tasks. |
Multilingual Pretraining for Pixel Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | PIXEL-M4 model pretrains on four visually and linguistically diverse languages . previous work on pixel-based language models focused on monolingual pretraining on English data . |
| Approach: | They propose a pixel-based language model that is pretrained on four visually diverse languages. |
| Outcome: | The proposed model outperforms an English-only counterpart on non-Latin scripts on semantic and syntactic tasks. |
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation (2026.acl-long)
Copied to clipboard
Mahi Luthra, Jiayi Shen, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, null Surya, Vanessa Stark, Rashel Moritz, Juan Pino, Yann LeCun, Emmanuel Dupoux
| Challenge: | Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability and downstream spoken language modeling scores . current self-supervised learning models require thousands of hours of training data to learn meaningful linguistic representations. |
| Approach: | They propose a bi-level optimization framework for rapid adaptation of speech units to new languages using minimal unlabeled data. |
| Outcome: | The proposed model achieves rapid gains in phonemic discriminability and spoken language modeling scores . it surpasses in-domain toplines after training on less than 1h of target-language audio . |
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation (2025.findings-acl)
Copied to clipboard
| Challenge: | Direct prompting fails to detect ambiguity while linear probes can decode ambiguities with high accuracy, sometimes exceeding 90%. |
| Approach: | They introduce an adversarial ambiguity dataset that includes syntactic, lexical, and phonological ambiguities along with adversarials. |
| Outcome: | The proposed dataset includes syntactic, lexical, and phonological ambiguities along with adversarial variations. |
PHD: Pixel-Based Language Modeling of Historical Documents (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent years have seen a boom in efforts to digitise historical documents in numerous languages and sources, leading to a transformation in the way historians work. |
| Approach: | They propose a method for generating synthetic scans to resemble real historical documents by pre-training a model to reconstruct masked patches instead of predicting token distributions. |
| Outcome: | The proposed model can reconstruct masked patches and understand language well. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users (2025.acl-long)
Copied to clipboard
Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard
| Challenge: | Despite high adoption rate of Large Language Models, there are limitations related to contextual understanding, cultural sensitivity, and complex scene understanding. |
| Approach: | They conduct a user survey to identify adoption patterns and key challenges users face with such technologies. |
| Outcome: | The proposed models have high adoption rates but still face limitations in visual aids. |
Towards Privacy-Aware Sign Language Translation at Scale (2024.acl-long)
Copied to clipboard
| Challenge: | Existing sign language training systems require detailed and time aligned annotations to be effective. |
| Approach: | They propose a two-stage framework for privacy-aware SLT at scale that leverages self-supervised video pretraining on anonymized and unannotated videos followed by supervised SLT finetuning on a curated parallel dataset. |
| Outcome: | The proposed framework outperforms baselines on the How2Sign dataset and achieves state-of-the-art finetuned and zero-shot gloss-free SLT performance. |
PuzzLing Machines: A Challenge on Learning From Small Data (2020.acl-main)
Copied to clipboard
| Challenge: | a benchmark dataset of 81 languages is released to test deep neural models' human-like reasoning and generalization skills. |
| Approach: | They propose a challenge on learning from small data using Rosetta Stone puzzles from Linguistic Olympiads for high school students. |
| Outcome: | The proposed benchmark consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students. |