Papers by Kayo Yin

11 papers
When Does Translation Require Context? A Data-driven, Multilingual Exploration (2023.acl-long)

Copied to clipboard

Challenge: Recent studies in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way.
Approach: They develop a multilingual discourse-aware benchmark to evaluate model performance on discourse phenomena in a given dataset.
Outcome: The proposed model improves on previously studied phenomena while uncovering others which were not addressed.
When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical Selection (2021.emnlp-main)

Copied to clipboard

Challenge: Using manual content to learn languages is expensive and time consuming.
Approach: They propose a method for automatically identifying fine-grained lexical distinctions and extracting rules explaining them in a human- and machine-readable format.
Outcome: The proposed method is able to identify fine-grained distinctions and explain them in a human- and machine-readable format.
Interpreting Language Models with Contrastive Explanations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding .
Approach: They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another.
Outcome: The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers .
Do Context-Aware Translation Models Pay the Right Attention? (2021.acl-long)

Copied to clipboard

Challenge: Context-aware machine translation models fail to leverage contextual information to resolve ambiguous words and pronouns.
Approach: They propose a new dataset that includes supporting context words for 14K translations that professional translators found useful for pronoun disambiguation.
Outcome: The proposed model can automatically disambiguate pronouns and polysemous words when they are not in the same context.
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles (2024.emnlp-main)

Copied to clipboard

Challenge: Deaf and hard-of-hearing students face significant barriers in accessing STEM education due to the scarcity of STEM resources in signed languages.
Approach: They develop models to identify fingerspelled words in American Sign Language (ASL) given an English sentence and a video, the model detects which English phrase is fingerspelled in the clip.
Outcome: ASL STEM Wiki is the first continuous signing dataset focused on STEM . it detects fingerspelled words and queries them for appropriate signs to suggest to interpreters.
Using Language Models to Disambiguate Lexical Choices in Translation (2024.emnlp-main)

Copied to clipboard

Challenge: In translation, a concept represented by a single word can have multiple variations in a target language.
Approach: They evaluate language models that can be used to generate English rules for lexical selection . they find weaker models with high-quality lexicals improve accuracy .
Outcome: The proposed model outperforms existing models on the lexical selection task in English and with native speakers.
Better Sign Language Translation with STMC-Transformer (2020.coling-main)

Copied to clipboard

Challenge: Current SLT approaches use a sign language recognition system to extract sign language glosses from videos.
Approach: They propose to use a Sign Language Recognition system to extract sign language glosses from videos and a translation system to generate spoken language translations from the glossed sign language.
Outcome: The proposed system outperforms existing methods on gloss-to-text and video-to text translations on the ASLG-PC12 corpus.
Including Signed Languages in Natural Language Processing (2021.acl-long)

Copied to clipboard

Challenge: Existing research in Sign Language Processing (SLP) rarely explores signed languages . authors urge adoption of an efficient tokenization method and the collection of real-world signed language data .
Approach: They propose to include signed languages as a research area with high social and scientific impact . they review the limitations of current SLP models and identify the open challenges .
Outcome: The proposed model should include signed languages as a research area with high social and scientific impact.
Measuring and Increasing Context Usage in Context-Aware Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Recent work in neural machine translation has demonstrated the necessity and feasibility of using inter-sentential context, but it is often not clear how much they actually utilize it at translation time.
Approach: They propose a conditional cross-mutual information metric to quantify usage of context by model architectures that can use it at translation time.
Outcome: The proposed method increases context usage and improves translation quality according to BLEU and COMET metrics.
Signed Coreference Resolution (2021.emnlp-main)

Copied to clipboard

Challenge: Sign Language Processing is based on linguistic theories of spoken languages and expect either speech or written text as input.
Approach: They propose a new challenge for coreference modeling and Sign Language Processing to solve this problem.
Outcome: The proposed models will be linguistically informed and can address the complexities of the challenge effectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations