Papers by Florian Schottmann
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models that target a single language are not seen during finetuning, but are able to respond in multiple languages once deployed in downstream applications. |
| Approach: | They investigate the minimal amount of multilinguality required during finetuning to elicit effective cross-lingual generalisation in English-centric LLMs. |
| Outcome: | The proposed model can respond in as few as two to three languages to a user's query in English, but the degree to which a target language is seen during pretraining is limiting. |
Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model (2023.acl-long)
Copied to clipboard
| Challenge: | Existing work has explored using sequence-to-sequence rewriting models to transform biased outputs into more gender-fair language by creating pseudo training data through linguistic rules. |
| Approach: | They propose to use machine translation models to create gender-biased text from real gender-fair text via round-trip translation to eliminate rule-based data creation. |
| Outcome: | The proposed approach matches the performance of state-of-the-art rewriting models for English. |