Papers by Mert Inan

11 papers
How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations? (2025.naacl-long)

Copied to clipboard

Challenge: despite the growing need for advanced signing technologies, signed language resources remain scarce.
Approach: They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model .
Outcome: The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation.
Accounting for Sycophancy in Language Model Uncertainty Estimation (2025.findings-naacl)

Copied to clipboard

Challenge: Effective human-machine collaboration requires machine learning models to externalize uncertainty.
Approach: They propose a generalization of the definition of sycophancy bias and a new algorithm to account for scophancies in uncertainty estimation.
Outcome: The proposed algorithm can account for sycophancy in uncertainty estimation process.
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts.
Approach: They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users.
Outcome: The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm.
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation (2025.emnlp-main)

Copied to clipboard

Challenge: ambiguities in natural language can lead to outputs that seem correct but fail to reflect the speaker’s intent.
Approach: They propose to identify and then resolve ambiguities in natural language and propose metrics to quantify them.
Outcome: The proposed metrics better correlate with human annotations than uncertainty baselines.
COSMic: A Coherence-Aware Generation Metric for Image Descriptions (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to evaluate captions have limited learning of their output . previous methods focused on n-gram measures of similarity to reference output based on a ngram of similarities to the output metric.
Approach: They propose a first discourse-aware learned generation metric for evaluating image descriptions.
Outcome: The proposed metric predicts human ratings of captions on out-of-domain images.
Dialogue is the Plan: From Interface to Joint Action in Agentic AI (2026.acl-short)

Copied to clipboard

Challenge: Large Language Model agents' language use is often used as an interface for instructing and reporting results.
Approach: They argue that large language models are often used as an interface for instructingactions and reporting results.
Outcome: We show that large-scale language models can be used to plan and act, yet their language is often used as an interface for instructing and reporting results.
Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied Dialogue (2023.emnlp-main)

Copied to clipboard

Challenge: Embodied task completion requires an agent to predict environment actions to complete tasks based on natural language instructions and egocentric visual observations.
Approach: They propose a method to generate human-human dialogues and use them as training data for plan prediction.
Outcome: The proposed model outperforms language-only models but falls short of oracle plans.
Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) are capable of generating well-formed responses, but they struggle in goal-oriented settings.
Approach: They propose a discourse-aware multimodal task-oriented dialogue system that combines discourse theories with offline LLM generation.
Outcome: The proposed system reduces misunderstandings in the dialect of African-American Vernacular English from 93% to 57%.
Modeling Intensification for Sign Language Generation: A Computational Approach (2022.findings-acl)

Copied to clipboard

Challenge: End-to-end sign language generation models do not accurately represent prosody in sign language.
Approach: They propose to model intensification in a data-driven manner to improve prosody in generated sign languages by modeling temporal and spatial variations.
Outcome: The proposed models improve the prosody of generated sign languages by using data-driven models.
Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production (2024.lrec-main)

Copied to clipboard

Challenge: Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link.
Approach: They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts.
Outcome: The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts.
Including Facial Expressions in Contextual Embeddings for Sign Language Generation (2023.starsem-1)

Copied to clipboard

Challenge: State-of-the-art sign language generation frameworks lack expressivity and naturalness . current systems focus on manual signs, neglecting affective, grammatical and semantic functions of facial expressions . communication between the Deaf and Hard of Hearing (DHH) individuals may be facilitated by emerging language technologies .
Approach: They propose a Dual Encoder Transformer capable of generating manual signs and facial expressions by capturing similarities and differences found in text and sign gloss annotations.
Outcome: The proposed model improves the quality of automatically generated sign language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations