Papers by Yingqiang Gao
SwissADT: An Audio Description Translation System for Swiss Languages (2025.naacl-industry)
Copied to clipboard
| Challenge: | despite advances in multilingual machine translation, lack of well-crafted AD data impedes development of audio description translation systems. |
| Approach: | They propose an audio description translation system for three main Swiss languages and English . they combine human expertise with the power of Large Language Models to improve quality . |
| Outcome: | The proposed system is designed to enhance accessibility for multilingual populations in Switzerland. |
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive (2025.emnlp-demos)
Copied to clipboard
Yingqiang Gao, Fabian Winiger, Patrick Montjourides, Anastassia Shaitarova, Nianlong Gu, Simon Peng-Keller, Gerold Schneider
| Challenge: | Religion and spirituality (R/S) are complex and domain-dependent concepts that have long confounded researchers and policymakers. |
| Approach: | They propose an interactive question-answering system based on Retrieval-Augmented Generation (RAG) SpiritRAG allows researchers and policymakers to conduct complex, context-sensitive database searches of large datasets . |
| Outcome: | SpiritRAG is an interactive Q&A system based on Retrieval-Augmented Generation (RAG) built using 7,500 UN resolution documents related to religion and spirituality in the domains of health and education. |
DETECT: Determining Ease and Textual Clarity of German Text Simplifications (2026.eacl-long)
Copied to clipboard
| Challenge: | Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore. |
| Approach: | They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency. |
| Outcome: | The proposed metric achieves higher correlations with human judgments than widely used ATS metrics. |
SwiLTra-Bench: The Swiss Legal Translation Benchmark (2025.acl-long)
Copied to clipboard
Joel Niklaus, Jakob Merane, Luka Nenadic, Sina Ahmadi, Yingqiang Gao, Cyrill A. H. Chevalley, Claude Humbel, Christophe Gösken, Lorenzo Tanzi, Thomas Lüthi, Stefan Palombo, Spencer Poff, Boling Yang, Nan Wu, Matthew Guillod, Robin Mamié, Daniel Brunner, Julio Pereyra, Niko Grupen
| Challenge: | In Switzerland legal translation relies on legal experts who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice. |
| Approach: | They propose a multilingual benchmarking system that evaluates Swiss legal translation systems based on 180K aligned Swiss legal translator pairs . they show frontier models achieve superior translation performance across all document types while specialized translation systems excel specifically in laws but under-perform in headnotes. |
| Outcome: | The proposed model outperforms specialized models in laws but underperform in headnotes. |
GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information (2023.emnlp-main)
Copied to clipboard
| Challenge: | Abstracts of scientific papers typically contain premises and conclusions, but in non-structured abstracts the concluding information is not marked. |
| Approach: | They propose to use Normalized Mutual Information (NMI) to optimize the NMI score between two segments by assuming that conclusions are strongly semantically linked with preceding premises. |
| Outcome: | The proposed approach outperforms baseline methods on structured abstracts and on non-structured abstracts. |
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords (2025.acl-long)
Copied to clipboard
Sina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich
| Challenge: | Lexical borrowing is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors. |
| Approach: | They propose a novel contrastive dataset comprising sentences with and without loanwords across 10 languages to examine how machine translation and language models process loanword . |
| Outcome: | The proposed dataset shows that state-of-the-art models prefer loanwords over native terms and exhibit varying performance across languages. |
Evaluating Unsupervised Argument Aligners via Generation of Conclusions of Structured Scientific Abstracts (2024.eacl-short)
Copied to clipboard
| Challenge: | Scientific abstracts provide a concise summary of research findings. |
| Approach: | They evaluate unsupervised approaches for extracting scientific arguments as aligned premise-conclusion pairs . they find mutual information outperforms other measures on this task . |
| Outcome: | The proposed methods outperform language models on the task of extracting scientific arguments from abstracts. |
Character-Level Translation with Self-attention (2020.acl-main)
Copied to clipboard
| Challenge: | Existing models for character-level neural machine translation operate on word-level, which makes them memory inefficient because of large vocabulary sizes. |
| Approach: | They propose a transformer-based model and a novel variant that uses convolutions to combine information from nearby characters to facilitate character interactions. |
| Outcome: | The proposed model outperforms the standard transformer model and learns more robust character alignments on bilingual and multilingual translation datasets. |
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)
Copied to clipboard
| Challenge: | Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content. |
| Approach: | They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future . |
| Outcome: | The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future . |