Papers by Dávid Javorský
Assessing Word Importance Using Models Trained for Semantic Tasks (2023.findings-acl)
Copied to clipboard
| Challenge: | Many NLP tasks require to automatically identify the most significant words in a text. |
| Approach: | They propose to use attribution methods to explain the predictions of two NLP tasks to derive word significance from models trained to solve semantic tasks. |
| Outcome: | The proposed method is robust to the initial task and is able to identify important words in sentences without explicit word importance labeling in training. |
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)
Copied to clipboard
Matthias Sperber, Ondřej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polák, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
| Challenge: | a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics . |
| Approach: | They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context. |
| Outcome: | The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements. |
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines (2025.acl-long)
Copied to clipboard
| Challenge: | Existing parallel corpora of translated texts fail to model long-range interactions between speech segments or specific types of divergences. |
| Approach: | They propose to use MockConf to analyze simultaneous interpreting and to develop a student interpretation dataset that was collected from Mock Conferences. |
| Outcome: | The proposed dataset contains 7 hours of recordings in 5 European languages, transcribed and aligned at the level of spans and words. |