Challenge: Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words . a recent study shows that frequency from YouTube subtitles is comparable to and often better than the best available resources.
Approach: They use YouTube subtitles to construct frequency norms for five languages . they find they are comparable to and often better than the best currently available resources .
Outcome: The proposed method improves on the best currently available resources for Chinese, English, Indonesian, Japanese, and Spanish.

Similar Papers

MuST-Cinema: a Speech-to-Subtitles corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for subtitling are laborious and costly, says aaron sanchez . he says the current methods are laboriously complex and require manual work .
Approach: They propose to use TED subtitles to build a multilingual speech translation corpus . they propose to annotate existing subtitling corpora with subtitle breaks .
Outcome: The proposed model can be used to segment sentences into subtitles and reduces human work . the proposed model reduces the time and cost of human subtitling tasks .
OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora (L18-1)

Copied to clipboard

Challenge: Movie and TV subtitles are a valuable resource for the compilation of parallel corpora . however, the quality of the resulting sentence alignments is often lower than for other parallel corpoora.
Approach: They propose to use movie and TV subtitles to extract parallel corpora from 3.7 million subtitles spread over 60 languages to obtain explicit quality scores for each sentence alignment.
Outcome: The proposed model predicts translation probabilities with a root mean square error of 0.07 . the results show that the model can prune out low-quality alignments .
Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing (2023.tacl-1)

Copied to clipboard

Challenge: a large-scale study of human dubbing in practice is lacking in qualitative literature on human dubs . authors argue for vocal naturalness and translation quality over isometric constraints . a data-driven examination of the way humans perform this task is needed .
Approach: They analyze 319.57 hours of video from 54 professionally produced titles . they argue for vocal naturalness and translation quality over isometric constraints . authors say they need to preserve speech characteristics and transfer of semantic properties .
Outcome: The study challenges assumptions in qualitative and machine-learning literature on dubbing . it also finds that source-side audio influences human dubbing through other channels .
Subjective Evaluation of Comprehensibility in Movie Interactions (2020.lrec-1)

Copied to clipboard

Challenge: Various studies have dealt with the comprehensibility of textual, audio, or audiovisual documents.
Approach: They aim to build a corpus of human annotations that could help to study human perceptions of comprehensibility of audiovisual documents.
Outcome: The proposed corpus of human annotations will help to study human perceptions of comprehensibility of audiovisual documents.
Word Frequency Does Not Predict Grammatical Knowledge in Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Neural language models learn the grammatical properties of natural languages to varying degrees of accuracy.
Approach: They focus on subject-verb agreement and reflexive anaphora to investigate whether there are systematic sources of variation in the language models’ accuracy.
Outcome: The proposed model can learn grammatical properties from training data.
Toward Informal Language Processing: Knowledge of Slang in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have offered a strong potential for natural language systems to process informal language.
Approach: They propose to use movie subtitles to evaluate slang in large language models . they find that smaller LLMs finetuned on the dataset achieve comparable performance .
Outcome: The proposed dataset can be used to evaluate LLMs on slang detection and identification of regional and historical sources for interpretive insights.
More than just Frequency? Demasking Unsupervised Hypernymy Prediction Methods (2021.findings-acl)

Copied to clipboard

Challenge: Using unsupervised methods of hypernymy prediction, we show that the predictions of three methods overlap and are highly correlated with frequency-based predictions.
Approach: They compare unsupervised methods of hypernymy prediction to supervised methods . they show that the methods overlap and are highly correlated with frequency-based predictions .
Outcome: The proposed methods overlap and are highly correlated with frequency-based predictions across English and German datasets.
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)

Copied to clipboard

Challenge: Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired.
Approach: They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks.
Outcome: The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)

Copied to clipboard

Challenge: Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory.
Approach: They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction.
Outcome: Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations