Challenge: Recent large-scale datasets such as Spider and WikiSQL facilitated novel modeling techniques for text-to-SQl parsing.
Approach: They propose a new cross-domain evaluation dataset of real Web databases . they examine the choice of evaluation tasks for text-to-SQL parsers .
Outcome: The proposed model improves accuracy by 13.2% over state-of-the-art parsers in real-life environments.

Similar Papers

Exploring Unexplored Generalization Challenges for Cross-Database Semantic Parsing (2020.acl-main)

Copied to clipboard

Challenge: Existing evaluation datasets such as Spider are used to support cross-database semantic parsing . XSP systems that map natural language utterances to SQL queries are evaluated on databases unseen during training.
Approach: They propose a setup that uses eight well-studied datasets to evaluate cross-database semantic parsing systems.
Outcome: The proposed system performs well on spider, but struggles to generalize to the repurposed set.
Exploring Underexplored Limitations of Cross-Domain Text-to-SQL Generalization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing text-to-SQL models do not generalize when faced with domain knowledge that does not frequently appear in training data.
Approach: They propose a human-curated dataset based on the Spider benchmark for text-to-SQL translation.
Outcome: The proposed model performs better on unseen domains than existing models on public benchmarks.
A Review of Cross-Domain Text-to-SQL Models (2020.aacl-srw)

Copied to clipboard

Challenge: WikiSQL and Spider are cross-domain text-to-SQl datasets that have attracted much attention from the research community.
Approach: They propose to divide top models into two paradigms and evaluate their models for schema linking, pretrained word embeddings, reasoning assistance modules.
Outcome: The proposed models have over 90% execution accuracy, the authors show . the proposed models are more complex and more complex than the proposed ones .
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect (2022.coling-1)

Copied to clipboard

Challenge: text-to-SQL is a language processing and database-based language processing (NLP) task is to convert natural utterances into SQL queries and its practical application is to build natural language interfaces to database systems.
Approach: They propose to conduct a systematic survey of text-to-SQL to examine the challenges and potential future directions.
Outcome: The proposed system converts natural utterances into SQL queries and is a representative task in semantic parsing.
Evaluating Cross-Domain Text-to-SQL Models and Benchmarks (2023.emnlp-main)

Copied to clipboard

Challenge: Text-to-SQL benchmarks are used to evaluate progress made in the field . however, matching a model-generated SQL query to a reference SQL query fails due to various reasons.
Approach: They conduct an extensive evaluation of text-to-SQL benchmarks and re-evaluate some of the top-performing models.
Outcome: The results show that a recent model surpasses the gold standard reference queries in the Spider benchmark in human evaluation.
Towards Generalizable and Robust Text-to-SQL Parsing (2022.findings-emnlp)

Copied to clipboard

Challenge: Text-to-SQL parsers must be generalizable and robust against input perturbations.
Approach: They propose a novel framework to learn text-to-SQL parsing in stages to improve parser's ability to acquire general SQL knowledge instead of capturing spurious patterns.
Outcome: The proposed framework achieves state-of-the-art performance on the Spider, SParC, and CoSQL datasets.
“What Do You Mean by That?” A Parser-Independent Interactive Approach for Enhancing Text-to-SQL (2020.emnlp-main)

Copied to clipboard

Challenge: In Natural Language Interfaces to Databases systems, text-to-SQL parsers allow users to query databases by using natural language questions.
Approach: They propose a parser-independent interactive approach that interacts with users using multi-choice questions and can easily work with arbitrary parsers.
Outcome: The proposed approach improves performance with limited interaction turns by using simulation and human evaluation on two cross-domain datasets with five state-of-the-art parsers.
Bridging the Generalization Gap in Text-to-SQL Parsing with Schema Expansion (2022.acl-long)

Copied to clipboard

Challenge: Existing text-to-SQL parsers struggle with out-of-domain generalization problems, arguing that they lack the ability to match domain specific phrases to composite operations over columns.
Approach: They propose to use a synthetic dataset and a re-purposed train/test split to quantify out-of-domain generalization over column operations to address this problem.
Outcome: The proposed method outperforms baseline parsers on the domain generalization problem, while boosting the underlying parser’ overall performance by 13.8% relative accuracy gain (5.1% absolute).
Global Reasoning over Database Structures for Text-to-SQL Parsing (D19-1)

Copied to clipboard

Challenge: Existing semantic parsers only select a set of database constants at training time . current models only consider local information, not global ones .
Approach: They propose a semantic parser that globally reasons about the structure of the query to make a more contextually-informed selection of database constants.
Outcome: The proposed model increases accuracy from 39.4% to 47.4% on a zero-shot semantic parsing dataset with complex databases.
Improving Text-to-SQL Evaluation Methodology (P18-1)

Copied to clipboard

Challenge: Current evaluations of text-to-SQL systems are limited by the way they divide data into training and test sets.
Approach: They propose to standardize and improve existing and new text-to-SQL datasets . they propose a template-based slot-filling baseline that cannot generalize to new queries .
Outcome: The proposed system is competitive with prior work on multiple datasets and can be used on training and test sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations