Papers by Shohei Taniguchi
Infinity-MoE: Generalizing Mixture of Experts to Infinite Experts (2026.eacl-short)
Copied to clipboard
| Challenge: | Existing methods to increase the number of experts are -MoE and . |
| Approach: | They propose a mixture of experts that selects a few feed-forward networks per token to increase the number of experts. |
| Outcome: | The proposed model improves on a GPT-2 Small model with 129M active and 186M total parameters by 2.5% over the current model. |
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers (2026.findings-eacl)
Copied to clipboard
| Challenge: | Length generalization is the ability of language models to maintain performance on inputs longer than those seen during pretraining. |
| Approach: | They propose a position encoding strategy that uses random float sampling to generalize to unseen lengths. |
| Outcome: | The proposed strategy can generalize to lengths unseen during training and in benchmarks. |