Papers by Farhan Samir

8 papers
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language (2024.naacl-long)

Copied to clipboard

Challenge: a recent study shows that multilingual speech processing systems can generalize to unseen languages without adaptation.
Approach: They propose a phoneme-based phoneme embedding model that can be generalized to unseen languages by using a neural forced aligner.
Outcome: The proposed model can generalize to unseen languages without adaptation.
ZIPA: A family of efficient models for multilingual phone recognition (2025.acl-long)

Copied to clipboard

Challenge: IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions.
Approach: They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition.
Outcome: The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation.
Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection (2023.emnlp-main)

Copied to clipboard

Challenge: a data augmentation technique that generates synthetic examples by randomly substituting stem characters in existing training examples is still poorly understood.
Approach: They propose a data augmentation strategy that generates synthetic examples by randomly substituting stem characters in existing training examples.
Outcome: The proposed method generates synthetic examples by randomly substituting stem characters in existing training examples.
Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia (2024.emnlp-main)

Copied to clipboard

Challenge: a recent study focuses on comparative text analyses to explain social phenomena and identify systematic biases.
Approach: They evaluate InfoGap method to locate information gaps and inconsistencies in articles at the fact level, across languages.
Outcome: The method identifies discrepancies in factual coverage across languages and biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia.
Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages (2022.findings-acl)

Copied to clipboard

Challenge: Recent progress in NLP is driven by pretrained models leveraging massive datasets.
Approach: They argue that IGT data can be leveraged provided target language expertise is available and that it can be used to create effective models.
Outcome: The proposed model can be leveraged provided that target language expertise is available.
An Inflectional Database for Gitksan (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we build a database of partial inflection tables for Gitksan, a low-resource Indigenous language of Canada.
Approach: They use Gitksan data in interlinear glossed format to build a database of partial inflection tables and enrich it with neural transformer reinflection models.
Outcome: The proposed model improves the performance of the experimental data hallucination and back-translation techniques.
A Formidable Ability: Detecting Adjectival Extremeness with DSMs (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies on distributional semantic models capture abstract semantic properties across domains . abstract properties can form the basis for abstract semantic classes .
Approach: They propose to use distributional semantic models to capture cross-domain properties . they use extremeness to model emergence of intensifier meaning in adverbs .
Outcome: The proposed model can capture extremeness and intensifier meaning in adverbs.
Quantifying Cognitive Factors in Lexical Decline (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on lexical decline suggest that cognitive and linguistic factors play a role in the survival of words and their success in the linguistic ecosystem.
Approach: They propose a variety of psycholinguistic factors that are predictive of lexical decline, in which words greatly decrease in frequency over time.
Outcome: The proposed factors show significant differences in the expected direction between each curated set of declining words and their matched stable words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations