Papers by Yingqiang Gao

9 papers
SwissADT: An Audio Description Translation System for Swiss Languages (2025.naacl-industry)

Copied to clipboard

Challenge: despite advances in multilingual machine translation, lack of well-crafted AD data impedes development of audio description translation systems.
Approach: They propose an audio description translation system for three main Swiss languages and English . they combine human expertise with the power of Large Language Models to improve quality .
Outcome: The proposed system is designed to enhance accessibility for multilingual populations in Switzerland.
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive (2025.emnlp-demos)

Copied to clipboard

Challenge: Religion and spirituality (R/S) are complex and domain-dependent concepts that have long confounded researchers and policymakers.
Approach: They propose an interactive question-answering system based on Retrieval-Augmented Generation (RAG) SpiritRAG allows researchers and policymakers to conduct complex, context-sensitive database searches of large datasets .
Outcome: SpiritRAG is an interactive Q&A system based on Retrieval-Augmented Generation (RAG) built using 7,500 UN resolution documents related to religion and spirituality in the domains of health and education.
DETECT: Determining Ease and Textual Clarity of German Text Simplifications (2026.eacl-long)

Copied to clipboard

Challenge: Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore.
Approach: They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency.
Outcome: The proposed metric achieves higher correlations with human judgments than widely used ATS metrics.
SwiLTra-Bench: The Swiss Legal Translation Benchmark (2025.acl-long)

Copied to clipboard

Challenge: In Switzerland legal translation relies on legal experts who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice.
Approach: They propose a multilingual benchmarking system that evaluates Swiss legal translation systems based on 180K aligned Swiss legal translator pairs . they show frontier models achieve superior translation performance across all document types while specialized translation systems excel specifically in laws but under-perform in headnotes.
Outcome: The proposed model outperforms specialized models in laws but underperform in headnotes.
GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual Information (2023.emnlp-main)

Copied to clipboard

Challenge: Abstracts of scientific papers typically contain premises and conclusions, but in non-structured abstracts the concluding information is not marked.
Approach: They propose to use Normalized Mutual Information (NMI) to optimize the NMI score between two segments by assuming that conclusions are strongly semantically linked with preceding premises.
Outcome: The proposed approach outperforms baseline methods on structured abstracts and on non-structured abstracts.
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords (2025.acl-long)

Copied to clipboard

Challenge: Lexical borrowing is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors.
Approach: They propose a novel contrastive dataset comprising sentences with and without loanwords across 10 languages to examine how machine translation and language models process loanword .
Outcome: The proposed dataset shows that state-of-the-art models prefer loanwords over native terms and exhibit varying performance across languages.
Evaluating Unsupervised Argument Aligners via Generation of Conclusions of Structured Scientific Abstracts (2024.eacl-short)

Copied to clipboard

Challenge: Scientific abstracts provide a concise summary of research findings.
Approach: They evaluate unsupervised approaches for extracting scientific arguments as aligned premise-conclusion pairs . they find mutual information outperforms other measures on this task .
Outcome: The proposed methods outperform language models on the task of extracting scientific arguments from abstracts.
Character-Level Translation with Self-attention (2020.acl-main)

Copied to clipboard

Challenge: Existing models for character-level neural machine translation operate on word-level, which makes them memory inefficient because of large vocabulary sizes.
Approach: They propose a transformer-based model and a novel variant that uses convolutions to combine information from nearby characters to facilitate character interactions.
Outcome: The proposed model outperforms the standard transformer model and learns more robust character alignments on bilingual and multilingual translation datasets.
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)

Copied to clipboard

Challenge: Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content.
Approach: They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future .
Outcome: The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations