Papers by Barry O'Sullivan
LaCoMSA: Language-Consistency Multilingual Self-Alignment with Latent Representation Rewarding (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing multilingual alignment methods mitigate these issues but rely on external supervision, such as translation systems or English-biased signal. |
| Approach: | They propose a preference optimization framework that leverages an LLM’s own latent representations as intrinsic supervision signals and rewards lower-resource language outputs based on their alignment with high-resourced (English) counterparts in the "semantic hub". |
| Outcome: | The proposed framework improves a Llama 3 8B model multilingual win rates by up to 6.8% absolute (55.0% relative) on X-AlpacaEval and achieves consistent gains across benchmarks and models. |