DeepStruct: Pretraining of Language Models for Structure Prediction (2022.findings-acl)
Copied to clipboard
| Challenge: | Pretrained language models perform structural understanding tasks that focus on understanding one aspect of the text. |
| Approach: | They propose a method for improving the structural understanding abilities of language models by pretraining them to generate structures from the text on task-agnostic corpora. |
| Outcome: | The proposed model performs state-of-the-art on 21 of 28 datasets. |
Similar Papers
Pretraining with Artificial Language: Studying Transferable Knowledge in Language Models (2022.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that pretraining with an artificial language with nesting dependency structure provides some knowledge transferable to natural language. |
| Approach: | They propose to pretrain artificial languages with structural properties that mimic natural language and then test their performance on downstream tasks. |
| Outcome: | The proposed language models show strong performance across languages and languages. |
Autoregressive Structured Prediction with Language Models (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Recent years have seen a paradigm shift in NLP towards using pretrained language models for a wide range of tasks. |
| Approach: | They propose to model structures as sequences of actions in autoregressive manner with PLMs . their approach allows in-structure dependencies to be learned without any loss . |
| Outcome: | The proposed approach achieves state-of-the-art on all structured prediction tasks. |
Leveraging pre-trained language models for linguistic analysis: A case of argument structure constructions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Argument structure constructions (ASCs) are lexicogrammatical patterns at the clausal level. |
| Approach: | They evaluate the effectiveness of pre-trained language models in identifying argument structure constructions . they use supervised training with RoBERTa and prompt-guided annotation with GPT-4 . |
| Outcome: | The proposed model outperforms the gold-standard model on three methods . the results show that the model performs better on gold-standardized data . |
Structure-Aware Language Model Pretraining Improves Dense Retrieval on Structured Data (2023.findings-acl)
Copied to clipboard
| Challenge: | Structure Aware Dense Retrieval (SANTA) model encodes user queries and structured data in one universal embedding space for retrieving structured data. |
| Approach: | They propose to use structured data and unstructured data to encode queries and structured data in one universal embedding space for retrieving structured data. |
| Outcome: | The proposed model achieves state-of-the-art on code search and product search and conducts convincing results in the zero-shot setting. |
How Much Pretraining Does Structured Data Need? (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models are increasingly adopted for handling structured data, despite pretraining on unstructured text. |
| Approach: | They propose to re-initialize subsets of layers with random weights before fine-tuning on structured datasets. |
| Outcome: | The proposed models are compared to unstructured datasets and show that they perform well over structured data. |
Patton: Language Model Pretraining on Text-Rich Networks (2023.acl-long)
Copied to clipboard
| Challenge: | Existing models for text-rich networks do not take inter-document structure into account. |
| Approach: | They propose a pretraining framework for a text-rich network using a masked language model and a masking node prediction framework. |
| Outcome: | The proposed model outperforms baselines on four tasks in academic and e-commerce domains. |
Synthetic Pre-Training Tasks for Neural Machine Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | toxicity and bias can be addressed by pre-training with synthetic resources . BLEU scores are used to compare methods with real-world data . |
| Approach: | They propose several ways to generate obfuscated data from large parallel corpus and concatenating phrase pairs from small word-aligned corpus with synthetic parallel data without real human language corpora. |
| Outcome: | The proposed methods can be used to generate obfuscated data or synthetic parallel data without real human language corpora even with high levels of oblication. |
Pretrained Language Models for Sequential Sentence Classification (D19-1)
Copied to clipboard
| Challenge: | Recent successful models for document-level understanding have used hierarchical encoding and CRFs to capture dependencies between subsequent labels. |
| Approach: | They propose a pretrained language model that captures contextual dependencies without hierarchical encoding nor a CRF. |
| Outcome: | The proposed model captures contextual dependencies without hierarchical encoding nor a CRF on four datasets, including a new dataset of structured scientific abstracts. |
Latent Structure Models for Natural Language Processing (P19-4)
Copied to clipboard
| Challenge: | Latent structure models are a powerful tool for compositional data modeling and pipelines. |
| Approach: | This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations . |
| Outcome: | This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations . |
Prompting Language Models for Linguistic Structure (2023.acl-long)
Copied to clipboard
| Challenge: | Existing prompting methods can test this hypothesis on autoregressive PLMs. |
| Approach: | They propose a structured prompting approach for linguistic structured prediction tasks that performs zero- and few-shot sequence tagging with autoregressive PLMs. |
| Outcome: | The proposed approach shows that the model can perform few-shot sequence tagging on part-of-speech taging, named entity recognition, and sentence chunking tasks. |