Papers by Yuki Saito

4 papers
Static Word Embeddings for Sentence Semantic Representation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn fixed-length embeddings for sentence semantics require large computational cost, making it difficult to process billions of sentences cost-efficiently or deploy models on resource-constrained devices such as smartphones.
Approach: They propose to extract word embeddings from a pre-trained Sentence Transformer and improve them with sentence-level principal component analysis followed by knowledge distillation or contrastive learning.
Outcome: The proposed model outperforms existing models on sentence semantic tasks and surpasses a basic Sentence Transformer model (SimCSE) on a text embedding benchmark.
Dialogue Corpus Construction Considering Modality and Social Relationships in Building Common Ground (2022.lrec-1)

Copied to clipboard

Challenge: Several studies have examined the process of building common ground in text chat, but none have investigated the process in depth.
Approach: They constructed a dialogue corpus to investigate the process of building common ground with a particular focus on the modality of dialogue and the social relationship between workers.
Outcome: The results suggest that adding the modality or developing the relationship between workers speeds up the building of common ground.
SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay (2020.lrec-1)

Copied to clipboard

Challenge: Developing a spontaneous speech corpus is important for spoken language research . a corpus of spontaneous speech is needed to develop these techniques .
Approach: They propose to use Japanese male commentators' spontaneous speech to construct a SMASH corpus . they use transcriptions and topic tags to annotate the commentaries and report some results .
Outcome: The proposed corpus includes spontaneous speech of two Japanese male commentators . the authors report that the annotations yielded a better corpus than the previous methods .
DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Experimental evaluation results show that rich annotations enhance the reproducibility of paralinguistic features of synthetic speech.
Approach: They investigate the effectiveness of using rich annotations in deep neural network-based statistical speech synthesis.
Outcome: The proposed method improves reproducibility of paralinguistic features of synthetic speech . the corpus of spontaneous Japanese (CSJ) has large annotations on paralinguistic and nonlinguistic features .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations