Papers by Shaked Yehezkel
Incorporating Context into Subword Vocabularies (2023.eacl-main)
Copied to clipboard
| Challenge: | Current tokenizers are trained on word frequency statistics over a corpus without considering information about co-occurrence or context. |
| Approach: | They propose a tokenizer that bakes in contextualized signal at the vocabulary creation phase to tailor subwords for their downstream use. |
| Outcome: | The proposed tokenizer is able to keep token contexts cohesive while not incurring a large price in terms of encoding efficiency or domain robustness. |