Papers by Shigehiko Schamoni
Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)
Copied to clipboard
| Challenge: | Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query. |
| Approach: | They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages. |
| Outcome: | The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese. |
Sample, Translate, Recombine: Leveraging Audio Alignments for Data Augmentation in End-to-end Speech Translation (2022.acl-short)
Copied to clipboard
| Challenge: | End-to-end speech translation relies on data that pair source-language speech inputs with corresponding translations. |
| Approach: | They propose a method that augments transcriptions by sampling from suffix memory and translating them into target languages. |
| Outcome: | The proposed method delivers up to 0.9 and 1.1 BLEU points on top of augmentation with knowledge distillation on languages on CoVoST 2 and Europarl-ST. |
Embedding Meta-Textual Information for Improved Learning to Rank (2020.coling-main)
Copied to clipboard
| Challenge: | a neural representation learning approach has not been extended to meta-textual information that is readily available for many IR tasks. |
| Approach: | They propose a framework that learns embeddings for meta-textual categories and optimizes a pairwise ranking objective for improved matching based on combined embedds of textual and meta-tactile information. |
| Outcome: | The proposed framework improves cross-lingual retrieval in the Wikipedia domain and Patent domain. |