Papers by Mathieu Rita

    2 papers
    Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning (2024.findings-acl)

    Copied to clipboard

    Challenge: Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning.
    Approach: They propose a reinforcement learning approach that leverages human demonstrations and a reward model to recalibrate the reward objective.
    Outcome: The proposed approach achieves comparable performance to carefully tuned baselines while mitigating ROO in three RL language tasks.
    On the Correspondence between Compositionality and Imitation in Emergent Neural Communication (2023.findings-acl)

    Copied to clipboard

    Challenge: a study examining compositionality and imitation learning in a Lewis game demonstrates that it is difficult to imitate compositional languages.
    Approach: They explore the link between compositionality and imitation in a Lewis game . they show that the learning algorithm used to imitate is crucial .
    Outcome: The proposed model improves compositionality and imitation in a Lewis game . the study shows that compositional languages are easier to imitate .

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations