Challenge: Structured data is prevalent in tables, databases, and knowledge graphs, but there is a gap in our understanding of how these linearization-based methods handle structured data, which is inherently non-linear.
Approach: They investigate the linear handling of structured data in encoder-decoder language models, specifically T5.
Outcome: The proposed model can mimic human-designed processes such as schema linking and syntax prediction, and it can be compressed due to modality fusion redundancy.

Similar Papers

Unleashing the True Potential of Sequence-to-Sequence Models for Sequence Tagging and Structure Parsing (2023.tacl-1)

Copied to clipboard

Challenge: Sequence-to-Sequence (S2S) models have been successful on text generation tasks . however, learning complex structures with S2S models remains challenging .
Approach: They propose to use constrained decoding to model part-of-speech tagging, named entity recognition, constituency, and dependency parsing tasks with 3 lexically diverse linearization schemas and corresponding constrained coding methods.
Outcome: The proposed methods outperform the state-of-the-art on four core tasks.
CodecLM: Aligning Language Models with Tailored Synthetic Data (2024.findings-naacl)

Copied to clipboard

Challenge: Recent work on generating diverse instructions and applying LLM to increase instruction complexity neglects downstream use cases.
Approach: They propose a framework for generating high-quality synthetic data for LLM alignment with different downstream instruction distributions and LLMs.
Outcome: Experiments on four open-domain instruction using the proposed framework validate the effectiveness of CodecLM over the current state-of-the-art.
Hierarchical Bracketing Encodings Work for Dependency Graphs (2025.emnlp-main)

Copied to clipboard

Challenge: Sequence labeling (SL) is a simple yet effective paradigm for a wide range of natural language problems.
Approach: They propose a new bracketing approach for dependency graph parsing that encodes graphs as sequences and n tagging actions.
Outcome: The proposed approach significantly reduces label space while preserving structural information.
SR-LLM: Rethinking the Structured Representation in Large Language Model (2025.acl-long)

Copied to clipboard

Challenge: Structured representations have long been pivotal in computational linguistics, but their role remains ambiguous in the Large Language Models (LLMs) era.
Approach: They propose a framework that integrates structured representations into LLMs from training-free and training-dependent perspectives.
Outcome: The proposed framework integrates structured representations through natural language descriptions in LLM prompts while augmenting the model’s inference capability through fine-tuning on linguistically described structured representation.
A Thorough Examination of Decoding Methods in the Era of LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Decoding methods are essential for converting language models from next-token predictors into practical task solvers.
Approach: They propose to evaluate decoding methods in general-purpose large language models . they find that decoding method performance is notably task-dependent .
Outcome: The proposed methods perform task-dependently and are influenced by alignment, model size, and quantization.
SKILL: Structured Knowledge Infusion for Large Language Models (2022.naacl-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated human-level performance on a vast spectrum of natural language tasks.
Approach: They propose a method to infuse structured knowledge into large language models by directly training T5 models on factual triples of knowledge graphs (KGs).
Outcome: The proposed method outperforms baseline models on FreebaseQA and WikiHop, as well as the Wikidata-answerable subset of TriviaQA and NaturalQuestions.
Autoregressive Structured Prediction with Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent years have seen a paradigm shift in NLP towards using pretrained language models for a wide range of tasks.
Approach: They propose to model structures as sequences of actions in autoregressive manner with PLMs . their approach allows in-structure dependencies to be learned without any loss .
Outcome: The proposed approach achieves state-of-the-art on all structured prediction tasks.
How Much Pretraining Does Structured Data Need? (2026.eacl-long)

Copied to clipboard

Challenge: Large language models are increasingly adopted for handling structured data, despite pretraining on unstructured text.
Approach: They propose to re-initialize subsets of layers with random weights before fine-tuning on structured datasets.
Outcome: The proposed models are compared to unstructured datasets and show that they perform well over structured data.
Large Language Models are Good Relational Learners (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to serialize large language models disregard critical relational structures and creates redundancies.
Approach: They propose a graph neural network encoder to create structured relational prompts for large language models within a retrieval-augmented generation framework.
Outcome: The proposed architecture preserves relational structure of databases while enabling LLMs to process and reason over complex entity relationships.
Leveraging AMR Graph Structure for Better Sequence-to-Sequence AMR Parsing (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies on AMR parsing often regard this task as a seq2seq translation problem.
Approach: They propose to translate AMR graphs into AMR token sequences in pre-processing and recover AMR from sequences after decoding.
Outcome: The proposed approach outperforms baseline and achieves 85.5 0.1 and 84.2 0.2 Smatch scores on AMR 2.0 and AMR 3.0.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations