Papers with DB

9 papers
An efficient method for Natural Language Querying on Structured Data (2023.acl-industry)

Copied to clipboard

Challenge: a new approach to NLQ on structured data is based on text-to-SQL type semantic parsing . domain classification, domain classification and domain classification are the main tasks . semantic parsed queries are less common when information is in structured form .
Approach: They propose an efficient and reliable approach to natural language Querying on databases . they use domain classification, domain classification and slot/entity extraction to query a DB .
Outcome: The proposed approach simplifies the NLQ on structured data problem to the following "bread and butter" tasks.
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing language agents powered by large language models face resource-constrained environments . proprietary models raise concerns in cost and service dependency, while large-scale open-source models require substantial computational resources.
Approach: They propose a Korean chemical toxicity information agent that reduces token consumption . they propose 'scenario-based dialogue generation' methodology that distills tool-using capabilities from larger models.
Outcome: The proposed language agent outperforms untuned models and baseline approaches in DB faithfulness and preference.
CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases (D19-1)

Copied to clipboard

Challenge: CoSQL is a corpus for building cross-domain, general-purpose database querying dialogue systems.
Approach: They present a corpus for building cross-domain, general-purpose database querying dialogue systems . they use a Wizard-of-Oz collection of 3k turns plus 10k+ annotated SQL queries .
Outcome: The proposed corpus is based on a Wizard-of-Oz dataset of 3k dialogues querying 200 complex DBs spanning 138 domains.
Unsupervised Paraphrasing with Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Paraphrase generation has benefited from recent advances in the design of training objectives and model architectures, but previous studies focused on supervised methods that require a large amount of labeled data that is costly to collect.
Approach: They propose a transfer learning approach that enables pre-trained language models to generate high-quality paraphrases in an unsupervised setting.
Outcome: The proposed model performs state-of-the-art on the Quora Question Pair and ParaNMT datasets and is robust to domain shift between the two datasets.
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing (2020.findings-emnlp)

Copied to clipboard

Challenge: BRIDGE is a powerful sequential architecture for cross-modal semantic parsing . BRidege captures cross-modal dependencies between natural language questions and relational databases .
Approach: They propose a sequential architecture that captures cross-modal dependencies between questions and relational databases in cross-DB semantic parsing.
Outcome: The proposed architecture performs well on the well-studied Spider benchmark (65.5% dev, 59.2% test).
Representing Schema Structure with Graph Neural Networks for Text-to-SQL Parsing (P19-1)

Copied to clipboard

Challenge: Semantic parsing to SQL has largely ignored the structure of the database schema . a recent study used a simple DB that was observed at both training and test time.
Approach: They propose a semantic parser where the schema structure is encoded with a graph neural network and used at both encoding and decoding time.
Outcome: The proposed parser improves from 33.8% to 39.4%, dramatically above the current state of the art, which is at 19.7%.
Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and Database (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches on semantic parsing suffer from exponential growth of logical form candidates and can hardly generalize to unseen data.
Approach: They propose a unified semantic parser for question answering on KB and DB . they define the primitive as the essential element in their framework .
Outcome: The proposed framework can predict logical forms by altering and composing top-ranked primitives with different operations.
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction (2025.acl-long)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are critical for identifying misconceptions and gaps in knowledge and accurately assessing students' understanding.
Approach: They propose to train a model to generate distractors that are more likely to be selected by students by a pairwise ranker and a distractor generator via Direct Preference Optimization.
Outcome: The proposed model outperforms baseline models and performs comparable to humans in various metrics including pairwise rank accuracy and distractor plausibility.
GRAD: Generalizing RAG Adaptation with Decoding (2026.acl-long)

Copied to clipboard

Challenge: Using GRAD, we can steer Retrieval-augmented generation objectives without retraining large language models.
Approach: They propose an adaptive decoding-time framework that keeps the base generator fixed and composes small, objective-specific guidance at inference.
Outcome: The proposed framework improves accuracy with favorable latency across public benchmarks and private settings with no in-domain labels while reliably activating helpful objectives and suppressing harmful ones, adaptively to tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations