Papers by Mert Inan
How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations? (2025.naacl-long)
Copied to clipboard
| Challenge: | despite the growing need for advanced signing technologies, signed language resources remain scarce. |
| Approach: | They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model . |
| Outcome: | The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation. |
Accounting for Sycophancy in Language Model Uncertainty Estimation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Effective human-machine collaboration requires machine learning models to externalize uncertainty. |
| Approach: | They propose a generalization of the definition of sycophancy bias and a new algorithm to account for scophancies in uncertainty estimation. |
| Outcome: | The proposed algorithm can account for sycophancy in uncertainty estimation process. |
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts. |
| Approach: | They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users. |
| Outcome: | The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm. |
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | ambiguities in natural language can lead to outputs that seem correct but fail to reflect the speaker’s intent. |
| Approach: | They propose to identify and then resolve ambiguities in natural language and propose metrics to quantify them. |
| Outcome: | The proposed metrics better correlate with human annotations than uncertainty baselines. |
COSMic: A Coherence-Aware Generation Metric for Image Descriptions (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to evaluate captions have limited learning of their output . previous methods focused on n-gram measures of similarity to reference output based on a ngram of similarities to the output metric. |
| Approach: | They propose a first discourse-aware learned generation metric for evaluating image descriptions. |
| Outcome: | The proposed metric predicts human ratings of captions on out-of-domain images. |
Dialogue is the Plan: From Interface to Joint Action in Agentic AI (2026.acl-short)
Copied to clipboard
| Challenge: | Large Language Model agents' language use is often used as an interface for instructing and reporting results. |
| Approach: | They argue that large language models are often used as an interface for instructingactions and reporting results. |
| Outcome: | We show that large-scale language models can be used to plan and act, yet their language is often used as an interface for instructing and reporting results. |
Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied Dialogue (2023.emnlp-main)
Copied to clipboard
| Challenge: | Embodied task completion requires an agent to predict environment actions to complete tasks based on natural language instructions and egocentric visual observations. |
| Approach: | They propose a method to generate human-human dialogues and use them as training data for plan prediction. |
| Outcome: | The proposed model outperforms language-only models but falls short of oracle plans. |
Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are capable of generating well-formed responses, but they struggle in goal-oriented settings. |
| Approach: | They propose a discourse-aware multimodal task-oriented dialogue system that combines discourse theories with offline LLM generation. |
| Outcome: | The proposed system reduces misunderstandings in the dialect of African-American Vernacular English from 93% to 57%. |
Modeling Intensification for Sign Language Generation: A Computational Approach (2022.findings-acl)
Copied to clipboard
| Challenge: | End-to-end sign language generation models do not accurately represent prosody in sign language. |
| Approach: | They propose to model intensification in a data-driven manner to improve prosody in generated sign languages by modeling temporal and spatial variations. |
| Outcome: | The proposed models improve the prosody of generated sign languages by using data-driven models. |
Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production (2024.lrec-main)
Copied to clipboard
| Challenge: | Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link. |
| Approach: | They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts. |
| Outcome: | The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts. |
Including Facial Expressions in Contextual Embeddings for Sign Language Generation (2023.starsem-1)
Copied to clipboard
| Challenge: | State-of-the-art sign language generation frameworks lack expressivity and naturalness . current systems focus on manual signs, neglecting affective, grammatical and semantic functions of facial expressions . communication between the Deaf and Hard of Hearing (DHH) individuals may be facilitated by emerging language technologies . |
| Approach: | They propose a Dual Encoder Transformer capable of generating manual signs and facial expressions by capturing similarities and differences found in text and sign gloss annotations. |
| Outcome: | The proposed model improves the quality of automatically generated sign language. |