Papers by Shi Zong

9 papers
L2Dir: Integrating L_2-Norm and Directional Alignment for Unsupervised Contrastive Representation Learning in Multimodal Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to multimodal representation learning focus on directional alignment and embedding magnitudes (L2-norm) however, these methods often fail to account for the intrinsic role of L2-norm in the contrastive process.
Approach: They propose a plug-and-play framework that optimizes L2-norm alignment and Directional consistency jointly.
Outcome: The proposed framework achieves consistent and significant performance gains over established baselines across 95 tasks using UniIR and VLM2Vec-V2 frameworks.
Extracting a Knowledge Base of COVID-19 Events from Social Media (2022.coling-1)

Copied to clipboard

Challenge: a flood of COVID-19 related information has appeared on social media since December 2019 . this includes reports on public figures who have tested positive/negative for the virus .
Approach: They construct a corpus of 10,000 tweets with annotated public reports of five COVID-19 events, using slot-filling questions to fill in slots.
Outcome: The proposed method can be quickly applied to develop knowledge bases for new domains in response to emerging crises, including natural disasters or future disease outbreaks.
Analyzing the Perceived Severity of Cybersecurity Threats Reported on Social Media (N19-1)

Copied to clipboard

Challenge: 6,000 tweets describe software vulnerabilities, which are shared across a range of websites and social media platforms.
Approach: They propose a method to link software vulnerabilities reported in tweets to CVEs in the National Vulnerability Database (NVD) a Precision@50 of 0.86 is achieved when forecasting high severity vulnerabilities, they show .
Outcome: The proposed method outperforms baseline methods based on tweet volume and the language used to describe them online.
Measuring Forecasting Skill from Text (2020.acl-main)

Copied to clipboard

Challenge: Prior studies have shown that some individuals can make accurate predictions with consistently better accuracy.
Approach: They examine linguistic factors associated with people's predictions including uncertainty, readability, and emotion.
Outcome: The proposed model can accurately predict forecasting skill using only language.
Analyzing the Intensity of Complaints on Social Media (2022.findings-naacl)

Copied to clipboard

Challenge: Prior studies on identifying the existence or the type of complaints focus on building automatic classification models for identifying complaints.
Approach: They propose to measure the intensity of complaints from text using Best-Worst Scaling method to estimate the popularity of posts on social media.
Outcome: The proposed model can estimate the popularity of complaints on social media with best-worst scaling (BWS) method.
Doctor Recommendation in Online Health Forums via Expertise Learning (2022.acl-long)

Copied to clipboard

Challenge: Currently, manual doctor allocations are used to handle large volumes of queries, limiting the efficiency to help patients in sheer quantities.
Approach: They propose to use patient queries to model doctor recommendation using their profiles and past dialogues to estimate their capabilities.
Outcome: The proposed model outperforms baseline models on a Chinese online health forum, outperforming baseline models.
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective (2022.findings-emnlp)

Copied to clipboard

Challenge: In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks.
Approach: They propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models.
Outcome: The proposed method analyzes captions generated by five popular VLP models to reveal how well they align with visual words and how well these align with images.
FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have enhanced or replaced traditional non-player characters in video games.
Approach: They propose a benchmark to evaluate social biases across three interaction patterns: transaction, cooperation, and competition.
Outcome: The proposed benchmark assesses four bias types across transaction, cooperation, and competition using a novel metric, FairMCV.
How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts? (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used for question answering over long contexts . high computational costs and latency hinder the process .
Approach: They explore the capabilities of Large Language Models to answer multiple questions based on the same conversational context.
Outcome: The proposed models outperform proprietary and public models in question answering . their results show that they can be cost-effective and transparent .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations