Papers by Shoutao Guo

9 papers
LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis (2025.acl-long)

Copied to clipboard

Challenge: LLaMA-Omni 2 is a series of speech language models (SpeechLMs) based on large language models.
Approach: They introduce a series of speech language models capable of real-time speech interaction . LLaMA-Omni 2 trains on 200K multi-turn speech dialogue samples .
Outcome: The proposed speech language models surpass state-of-the-art models on spoken question answering and speech instruction.
Turning Fixed to Adaptive: Integrating Post-Evaluation into Simultaneous Machine Translation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to perform adaptive and fixed translations lack evaluation before taking actions.
Approach: They propose a method to perform adaptive translation policy via post-evaluation into fixed policy . their method evaluates rationality of next action by measuring change in source content .
Outcome: The proposed method exceeds strong baselines under all latency.
Learning Optimal Policy for Simultaneous Machine Translation via Binary Search (2023.acl-long)

Copied to clipboard

Challenge: Simultaneous machine translation model needs a precise translation policy to achieve good latency-quality trade-offs.
Approach: They propose a method for building the optimal translation policy online via binary search by employing explicit supervision.
Outcome: Experiments on four translation tasks show that the proposed method exceeds strong baselines across all latency scenarios.
Wait-info Policy: Balancing Source and Target at Information Level for Simultaneous Machine Translation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to balance source and target information at the token level are limited by the number of received source tokens.
Approach: They propose a Wait-info Policy to balance source and target at the information level . they quantify the amount of info contained in each token and compare it with previous outputs .
Outcome: The proposed method outperforms baselines under and achieves better balance . it is based on comparisons between the total info of previous target outputs and received source inputs .
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning (2024.acl-long)

Copied to clipboard

Challenge: Existing simultaneous translation methods focus on text-to-text and speech-totext translation.
Approach: They propose a Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.
Outcome: The proposed model can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation (2024.acl-long)

Copied to clipboard

Challenge: Existing translation pipelines require additional cascade components to achieve speech-to-speech translation.
Approach: They propose a non-autoregressive generation framework for simultaneous speech translation . it integrates both text-to-text and speech-tospeech tasks into a unified framework .
Outcome: The proposed framework outperforms state-of-the-art models in speech-to-text and speech- to-speech tasks.
Non-autoregressive Streaming Transformer for Simultaneous Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Simultaneous machine translation models are trained to strike a balance between latency and translation quality.
Approach: They propose a non-autoregressive streaming Transformer which generates blank tokens and decodes repetitive tokens to adjust its READ/WRITE strategy flexibly.
Outcome: The proposed model outperforms previous strong autoregressive models on various benchmarks on siMT.
Decoder-only Streaming Transformer for Simultaneous Translation (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for siMT focus on the Encoder-Decoder architecture, but there are limitations in training and inference.
Approach: They propose a model that generates translation while reading source tokens . they propose Streaming Self-Attention mechanism tailored for the Decoder-only architecture .
Outcome: The proposed model achieves state-of-the-art performance on three translation tasks.
Simultaneous Machine Translation with Tailored Reference (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing SiMT models are trained using the same reference disregarding the varying amounts of available source information at different latency.
Approach: They propose a method that provides tailored reference for the SiMT models trained at different latency by rephrasing ground-truth to the tailored reference.
Outcome: The proposed method achieves state-of-the-art translation performance on three translation tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations