Papers by Dávid Javorský

3 papers
Assessing Word Importance Using Models Trained for Semantic Tasks (2023.findings-acl)

Copied to clipboard

Challenge: Many NLP tasks require to automatically identify the most significant words in a text.
Approach: They propose to use attribution methods to explain the predictions of two NLP tasks to derive word significance from models trained to solve semantic tasks.
Outcome: The proposed method is robust to the initial task and is able to identify important words in sentences without explicit word importance labeling in training.
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)

Copied to clipboard

Challenge: a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics .
Approach: They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context.
Outcome: The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements.
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines (2025.acl-long)

Copied to clipboard

Challenge: Existing parallel corpora of translated texts fail to model long-range interactions between speech segments or specific types of divergences.
Approach: They propose to use MockConf to analyze simultaneous interpreting and to develop a student interpretation dataset that was collected from Mock Conferences.
Outcome: The proposed dataset contains 7 hours of recordings in 5 European languages, transcribed and aligned at the level of spans and words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations