Papers by Hongyu Gong
WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia (2021.eacl-main)
Copied to clipboard
| Challenge: | a new approach to extract parallel sentences from Wikipedia articles is proposed . the approach is based on multilingual sentence embeddings, but does not limit it to English . |
| Approach: | They propose to automatically extract parallel sentences from Wikipedia articles in 96 languages . they train neural MT baseline systems on the mined data and evaluate them on the TED corpus . |
| Outcome: | The proposed approach extracts parallel sentences from Wikipedia articles in 96 languages . the extracted sentences achieve strong BLEU scores for many language pairs . |
PaRe: A Paper-Reviewer Matching Approach Using a Common Topic Space (D19-1)
Copied to clipboard
| Challenge: | Existing approaches to reviewer-paper matching are less effective to deal with the vocabulary mismatch and partial topic overlap between the submission and reviewer. |
| Approach: | They propose to combine the common topic model and abstract topic vectors to model the topics common to the submission and the reviewer's profile while relying on abstract topic vectors. |
| Outcome: | The proposed model improves on the existing model on two datasets. |
Preposition Sense Disambiguation and Representation (D18-1)
Copied to clipboard
| Challenge: | Prepositions are highly polysemous and their variegated senses encode significant semantic information. |
| Approach: | They match each preposition’s context and their interplay to the geometry of the word vectors to the left and right of the preposition. |
| Outcome: | The proposed algorithm is comparable to and better than state-of-the-art on two benchmark datasets. |
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input (2026.acl-long)
Copied to clipboard
| Challenge: | AV-Dialog uses audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses. |
| Approach: | They propose a multimodal dialog framework that uses both audio and visual cues to track the target speaker. |
| Outcome: | AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction and human-rated dialogue quality. |
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations (2023.acl-long)
Copied to clipboard
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk
| Challenge: | SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Approach: | They present a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Outcome: | The proposed model can train bilingual models on 136 language pairs with 418 thousand hours of speech. |
Reinforcement Learning Based Text Style Transfer without Parallel Training Corpus (N19-1)
Copied to clipboard
| Challenge: | Existing methods for text style transfer have demonstrated considerable success, but a parallel corpus may not always be available for a transfer task. |
| Approach: | They propose a text style transfer model that uses an attention-based encoder-decoder to transfer a sentence from the source style to the target style. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two different style transfer tasks. |
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)
Copied to clipboard
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, Ann Lee
| Challenge: | a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data. |
| Approach: | They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language . |
| Outcome: | The proposed system can be used to train models in languages without standard writing systems. |
Non-compositional Expression Generation Based on Curriculum Learning and Continual Learning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Non-compositional expressions are a classic ‘pain in the neck’ for NLP systems because of their non-composibility and limited data resources. |
| Approach: | They propose a dynamic curriculum learning framework which learns training examples from easy ones to harder ones but suffers from the forgetting problem. |
| Outcome: | The proposed framework improves on idiomatic expression generation and metaphor generation. |
Textless Speech-to-Speech Translation on Real Data (2022.naacl-main)
Copied to clipboard
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu
| Challenge: | Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output . |
| Approach: | They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data. |
| Outcome: | The proposed system can translate speech from one language into another without text data. |
Document Similarity for Texts of Varying Lengths via Hidden Topics (P18-1)
Copied to clipboard
| Challenge: | Existing approaches to measure document similarity are inadequate for document pairs with non-comparable lengths, such as a long document and its summary. |
| Approach: | They propose a document matching approach to bridge the gap between long documents and their abstract information in a common space of hidden topics. |
| Outcome: | The proposed approach outperforms strong baselines on two matching tasks and incorporates domain knowledge to gain further performance improvement. |
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available. |
| Approach: | They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream. |
| Outcome: | The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism. |
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora (2026.acl-long)
Copied to clipboard
Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, JI Yi, Gong Zhang, Renhai Chen, Yangqiu Song
| Challenge: | Existing knowledge graph construction frameworks require predefined schemas, limiting their scalability and domain coverage. |
| Approach: | They propose a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. |
| Outcome: | The proposed framework outperforms state-of-the-art models on multi-hop QA tasks and enhances LLM factuality. |
Embedding Syntax and Semantics of Prepositions via Tensor Decomposition (N18-1)
Copied to clipboard
| Challenge: | Existing methods on preposition representation treat prepositions no different from content words (e.g., word2vec and GloVe). |
| Approach: | They propose to use word-triple counts to capture a preposition’s interaction with its attachment and complement and derive preposition embeddings via tensor decomposition on a large unlabeled corpus. |
| Outcome: | The proposed model is comparable to or better than the state-of-the-art on multiple standardized datasets. |
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent expressive speech-to-speech translation systems have achieved impressive expressivity preservation performances by cascading unit-to speech (U2S) generator to the speech- to-unit translation model. |
| Approach: | They propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST) They aim to address this limitation by incorporating a distillation with no label (DINO) self-controlled training strategy into the model’s pretraining process. |
| Outcome: | The proposed model significantly improved the expressive speech-to-speech translation system in noisy environments while maintaining competitive performance in clean environments. |
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing spoken dialogue models are half-duplex in nature and require explicit prompting by the user or implicit tracking of interruption or silence events. |
| Approach: | They propose to integrate time information into Llama3-8b so that they run synchronously with the real-world clock. |
| Outcome: | The proposed model outperforms state-of-the-art in dialogue meaningfulness while maintaining naturalness. |
Recurrent Chunking Mechanisms for Long-Text Machine Reading Comprehension (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to machine reading comprehension (MRC) on long texts typically chunk text into equally-spaced segments without considering information from other segments. |
| Approach: | They propose to let a model learn to chunk in a more flexible way via reinforcement learning. |
| Outcome: | The proposed model extracts a text span from document and query as answer . previous models can only take a fixed-length (e.g., 512) text as input . |
T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to perform zero-shot cross-modal transfer between speech and text are limited to a very small number of language pairs. |
| Approach: | They propose a method to perform zero-shot cross-modal transfer between speech and text for translation tasks by using a speech decoder. |
| Outcome: | The proposed model significantly improves state-of-the-art for zero-shot speech translation on Must-C. |
Unified Speech-Text Pre-training for Speech Translation and Recognition (2022.acl-long)
Copied to clipboard
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Pino
| Challenge: | Existing methods to pre-train speech and text use unlabeled data to learn universal feature representations. |
| Approach: | They propose a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. |
| Outcome: | The proposed method achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task. |