Papers by Jason He

11 papers
Robustification of Multilingual Language Models to Real-world Noise in Crosslingual Zero-shot Settings with Robust Contrastive Pretraining (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies on robustness of pretrained multilingual models are limited to the English language.
Approach: They propose to use data augmentation and contrastive loss term to boost robustness of multilingual models in cross-lingual settings.
Outcome: The proposed model outperforms existing models on clean and noisy data in the cross-lingual setting.
Active Generalized Category Discovery with Diverse LLM Feedback (2026.eacl-long)

Copied to clipboard

Challenge: Generalized Category Discovery (GCD) is a practical and challenging open-world task that aims to recognize both known and novel categories in unlabeled data using limited labeled data from known categories.
Approach: They propose a framework for generalized category discovery that actively learns from diverse and collaborative feedback.
Outcome: The proposed framework improves instance-level contrastive features, generates category descriptions, and aligns uncertain instances with LLM-selected category descriptions.
Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt (2024.emnlp-industry)

Copied to clipboard

Challenge: Financial documents are filled with specialized terminology, arcane jargon, and curious acronyms that pose challenges for general-purpose text embeddings.
Approach: They propose to fine tune financial text embeddings finetuned on a carefully constructed dataset of 14.3M query-passage pairs including both public and proprietary financial documents.
Outcome: The proposed embeddings achieve Recall@1 of 62.8% on a held-out test set, vs. only 39.2% for the best general-purpose text embeddING from OpenAI.
REST: Retrieval-Based Speculative Decoding (2024.naacl-long)

Copied to clipboard

Challenge: Retrieval-based speculative decoding (REST) is a new language model generation algorithm . it uses existing knowledge to generate draft tokens, allowing for seamless integration and acceleration of any language model.
Approach: They propose a new algorithm that uses a draft language model to generate tokens from existing knowledge.
Outcome: The proposed method achieves a speedup of 1.62 to 2.36 on code or text generation.
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing (2025.naacl-long)

Copied to clipboard

Challenge: Existing models that memorize past tokens have “flat” memory architectures that restrict the context window.
Approach: They propose a framework that imitates human memorization behavior by preserving tokens from early input segments, passing memory embeddings along the sequence, and recalling relevant information from history.
Outcome: The proposed framework outperforms existing models in language modeling and question-answering tasks and achieves comparable or superior generation quality to long-context models with 2 57 fewer parameters and 2.5 116 less inference memory.
PAWS: Paraphrase Adversaries from Word Scrambling (N19-1)

Copied to clipboard

Challenge: Existing paraphrase identification datasets lack sentence pairs with high word overlap without being paraphrases.
Approach: They propose a workflow for generating pairs of sentences with high word overlap . they use controlled word swapping and back translation followed by fluency and paraphrase judgments .
Outcome: The proposed dataset has 108,463 well-formed paraphrase and non-paraphrase pairs with high lexical overlap.
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection (2024.naacl-long)

Copied to clipboard

Challenge: Semi-supervised dialogue summarization (SSDS) leverages model-generated summaries to reduce reliance on human-labeled data.
Approach: They propose a scoring approach that encapsulates three primary dimensions of summarization model quality.
Outcome: The proposed method reduces reliance on human-labeled data and improves the performance of summarization models.
QuALITY: Question Answering with Long Input Texts, Yes! (2022.naacl-main)

Copied to clipboard

Challenge: Existing models for natural language understanding are limited to processing only a few hundred words at a time.
Approach: They propose a dataset with context passages in English that have an average length of 5,000 tokens.
Outcome: a new dataset with long-text comprehension questions is used to test models on long-document comprehension . the questions are validated by contributors who have read the entire passage, not just excerpts . only half of the questions can be answered by annotators working under tight time constraints .
Leveraging Implicit Feedback from Deployment Data in Dialogue (2024.eacl-short)

Copied to clipboard

Challenge: Xu et al., 2023) and Bai ed., 2019) use crowdworkers to collect signals from natural dialogue episodes.
Approach: They use the publicly released BlenderBot deployment data to extract signals from conversations to implicitly measure the quality of a machine-generated utterance.
Outcome: The proposed model improves over baseline models, but some proxy signals can lead to undesirable generations.
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training (2024.eacl-long)

Copied to clipboard

Challenge: End-to-end (E2E) spoken language understanding models are constrained by the cost of collecting speech-semantics pairs.
Approach: They propose a model that learns E2E SLU without speech-semantics pairs . they propose cross-modal selective self-training (CMSST) to address imbalance and noise issues .
Outcome: The proposed model learns E2E SLU without speech-semantics pairs . the proposed model requires the domains of speech-text and text-sensitization to match .
Mapping Natural Language Instructions to Mobile UI Action Sequences (2020.acl-main)

Copied to clipboard

Challenge: a new problem of grounding natural language instructions to mobile UI actions is emerging . we use a Transformer to extract action phrase tuples from long-range natural language instruction .
Approach: They propose a dataset that pairs English instructions with actions performed by people on a mobile UI emulator.
Outcome: The proposed model achieves 70.59% accuracy on predicting complete ground-truth action sequences in PixelHelp.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations