Papers by Jianping Shen
Mention Extraction and Linking for SQL Query Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing text-to-SQL systems take a slot-filling approach, but they are limited in capturing inter-dependencies among SQL clauses. |
| Approach: | They propose an extraction-linking approach where a unified extractor recognizes all types of slot mentions appearing in the question sentence before a linker maps the recognized columns to the table schema to generate executable SQL queries. |
| Outcome: | The proposed method achieves the first place on the WikiSQL benchmark. |
Make Templates Smarter: A Template Based Data2Text System Powered by Text Stitch Model (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Neural network based data2text models drop or modify information in inputs and it is hard to control the generated contents. |
| Approach: | They propose a template-based data2text system powered by a text stitch model that automatically stitches adjacent template units. |
| Outcome: | The proposed system outperforms template-based systems in fidelity and human involvement on a benchmark dataset. |
MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing personalized product search methods assume that users’ query fully captures their real motivation, but in practice, user's queries do not always articulate the requirements. |
| Approach: | They propose a Motivation-Aware Personalized Search method that embeds queries and consultations into a unified semantic space via LLMs and utilizes a Mixture of Attention Experts (MoAE) to prioritize critical semantics. |
| Outcome: | Extensive experiments on real and synthetic data show that the proposed method outperforms existing methods in retrieval and ranking tasks. |
SQL Generation via Machine Reading Comprehension (2020.coling-main)
Copied to clipboard
| Challenge: | Text-to-SQL systems can generate SQL queries given natural language questions. |
| Approach: | They propose a method that formulates a question answering problem as a query answering problem where different slots are predicted by a unified machine reading comprehension (MRC) model. |
| Outcome: | The proposed method can achieve competitive results on WikiSQL, suggesting it being a promising direction for text-to-SQl. |
Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling Network (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing models of reinforcement learning use background planning and may suffer from low-quality simulated experiences. |
| Approach: | They propose a Monte Carlo Tree Search with Double-q Dueling network framework for task-completion dialogue policy learning. |
| Outcome: | The proposed method outperforms the previous model-based reinforcement learning methods and is robust to simulation errors. |
FASTMATCH: Accelerating the Inference of BERT-based Text Matching (2020.coling-main)
Copied to clipboard
| Challenge: | Recent pre-trained language models have shown state-of-the-art accuracies in text matching. |
| Approach: | They propose a BERT-based text matching model where representations and interactions are decoupled . they propose generating final matching scores using a lightweight attention network . |
| Outcome: | Experiments show that the proposed model can achieve up to 100X speed-up to BERT and RoBERTa while keeping more up to 98.7% of the performance. |
Similarity = Value? Consultation Value-Assessment and Alignment for Personalized Search (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods rely on semantic similarity to align historical consultations with current queries due to the absence of ‘value’ labels, but this lacks exploration of needs in user consultations. |
| Approach: | They propose a consultation value assessment framework that evaluates historical consultations from three novel perspectives: (1) Scenario Scope Value, (2) Posterior Action Value, and (3) Time Decay Value. |
| Outcome: | The proposed model outperforms baselines on public and commercial datasets on both retrieval and ranking tasks. |