Papers by Phillip Rust

10 papers
Text Rendering Strategies for Pixel Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to rendering text use a large set of almost-equivalent input patches, which may prove sub-optimal for downstream tasks due to redundancy in the input representations.
Approach: They propose four approaches to rendering text in a PIXEL model using character bigrams and patch frequency biases.
Outcome: The proposed models perform better on sentence-level tasks without compromising performance on token-level or multilingual tasks.
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models (2021.acl-long)

Copied to clipboard

Challenge: Using pretraining data, we find that a designated monolingual tokenizer plays an equally important role in the downstream performance of the model.
Approach: They propose to compare pretrained multilingual models with their monolingual counterparts on a set of five diverse monolingual downstream tasks.
Outcome: The proposed models offer previously unmatched performance in all NLP tasks.
Multilingual Pretraining for Pixel Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: PIXEL-M4 model pretrains on four visually and linguistically diverse languages . previous work on pixel-based language models focused on monolingual pretraining on English data .
Approach: They propose a pixel-based language model that is pretrained on four visually diverse languages.
Outcome: The proposed model outperforms an English-only counterpart on non-Latin scripts on semantic and syntactic tasks.
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability and downstream spoken language modeling scores . current self-supervised learning models require thousands of hours of training data to learn meaningful linguistic representations.
Approach: They propose a bi-level optimization framework for rapid adaptation of speech units to new languages using minimal unlabeled data.
Outcome: The proposed model achieves rapid gains in phonemic discriminability and spoken language modeling scores . it surpasses in-domain toplines after training on less than 1h of target-language audio .
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Direct prompting fails to detect ambiguity while linear probes can decode ambiguities with high accuracy, sometimes exceeding 90%.
Approach: They introduce an adversarial ambiguity dataset that includes syntactic, lexical, and phonological ambiguities along with adversarials.
Outcome: The proposed dataset includes syntactic, lexical, and phonological ambiguities along with adversarial variations.
PHD: Pixel-Based Language Modeling of Historical Documents (2023.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen a boom in efforts to digitise historical documents in numerous languages and sources, leading to a transformation in the way historians work.
Approach: They propose a method for generating synthetic scans to resemble real historical documents by pre-training a model to reconstruct masked patches instead of predicting token distributions.
Outcome: The proposed model can reconstruct masked patches and understand language well.
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)

Copied to clipboard

Challenge: Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages.
Approach: They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices .
Outcome: The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices.
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users (2025.acl-long)

Copied to clipboard

Challenge: Despite high adoption rate of Large Language Models, there are limitations related to contextual understanding, cultural sensitivity, and complex scene understanding.
Approach: They conduct a user survey to identify adoption patterns and key challenges users face with such technologies.
Outcome: The proposed models have high adoption rates but still face limitations in visual aids.
Towards Privacy-Aware Sign Language Translation at Scale (2024.acl-long)

Copied to clipboard

Challenge: Existing sign language training systems require detailed and time aligned annotations to be effective.
Approach: They propose a two-stage framework for privacy-aware SLT at scale that leverages self-supervised video pretraining on anonymized and unannotated videos followed by supervised SLT finetuning on a curated parallel dataset.
Outcome: The proposed framework outperforms baselines on the How2Sign dataset and achieves state-of-the-art finetuned and zero-shot gloss-free SLT performance.
PuzzLing Machines: A Challenge on Learning From Small Data (2020.acl-main)

Copied to clipboard

Challenge: a benchmark dataset of 81 languages is released to test deep neural models' human-like reasoning and generalization skills.
Approach: They propose a challenge on learning from small data using Rosetta Stone puzzles from Linguistic Olympiads for high school students.
Outcome: The proposed benchmark consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations