Papers by Yujia Tian

4 papers
Enhancing Hindi Feature Representation through Fusion of Dual-Script Word Embeddings (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models often neglect the integration of different scripts within a language, constraining their ability to capture richer semantic information.
Approach: They propose a dual-script enhanced feature representation method for Hindi . they combine features from Devanagari and Romanized Hindi Roberta .
Outcome: The proposed method improves model performance across multiple natural language processing tasks.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
DebugBench: Evaluating Debugging Capability of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated exceptional coding capabilities, but their debugging capabilities remain relatively unexplored.
Approach: They propose a debugging benchmark consisting of 4,253 LLMs with four major bug categories and 18 minor types in C++, Java, and Python.
Outcome: The proposed benchmark covers four major bug categories and 18 minor types in C++, Java, and Python.
An Effective Deployment of Diffusion LM for Data Augmentation in Low-Resource Sentiment Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models for textual data augmentation (DA) are highly data-hungry and struggle to perform satisfactorily under noisy conditions.
Approach: They propose to leverage a diffusion language model to capture in-domain knowledge and generate pseudo samples by reconstructing strong label-related tokens.
Outcome: The proposed method captures in-domain knowledge and generates pseudo samples by reconstructing strong label-related tokens.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations