PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Structured information extraction from unstructured text is critical for Software 3.0 systems . current approaches to extract structured information from unstructed text are static contracts . |
| Approach: | They propose a system that automates JSON schemas for LLM consumption and optimizes them for LRM consumption. |
| Outcome: | The proposed system improves extraction accuracy and reduces errors by 92% within the first retry and maintaining practical latency. |
Similar Papers
Learning to Extract Structured Entities Using Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Language Models (LMs) play a pivotal role in extracting structured information from unstructured text. |
| Approach: | They propose to reformulate the task to be entity-centric, enabling the use of diverse metrics that can provide more insights from various perspectives. |
| Outcome: | The proposed model outperforms baselines and human evaluations on the extracted entities. |
Learning to Generate Structured Output with Schema Reinforcement Learning (2025.acl-long)
Copied to clipboard
Yaxi Lu, Haolun Li, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Zhiyuan Liu, Fangming Liu, Maosong Sun
| Challenge: | Recent advances in large language models have facilitated the development of intelligent applications like automatic web search (Qin et al., 2023) Several methods exist for generating JSON strings from LLMs, including Prompting but often miss certain schemas. |
| Approach: | They propose to use 40K different JSON schemas to assess models' ability to generate valid JSON outputs. |
| Outcome: | The proposed model improves both in generating JSON outputs and downstream tasks. |
Zero-Shot Open-Schema Entity Structure Discovery (2026.eacl-long)
Copied to clipboard
Xueqiang Xu, Jinfeng Xiao, James Barry, Mohab Elkaref, Jiaru Zou, Pengcheng Jiang, Yunyi Zhang, Maxwell J Giammona, Geeth De Mel, Jiawei Han
| Challenge: | Existing methods based on large language models (LLMs) rely heavily on predefined entity attribute schemas or annotated datasets, often leading to incomplete extraction results. |
| Approach: | They propose a novel approach to entity structure extraction that does not require any schema or annotated datasets. |
| Outcome: | Experiments show that ZOES improves LLMs’ ability to extract more complete entity structures across three different domains, showcasing both the effectiveness and generalizability of the method. |
Entity Exchange in the Wild: A Diagnostic Study of LLM Based Real-World Conversational Entity Extraction (2026.acl-industry)
Copied to clipboard
| Challenge: | Prior work has examined the impact of transcription noise and cross-turn reasoning, but it has not systematically analyzed how entity-exchange phenomena themselves shape extraction performance. |
| Approach: | They evaluate 16 large language models on 6,387 real-world customer–agent conversations spanning 12 entity types across numeric, alphanumeric, temporal, and free-text categories. |
| Outcome: | The proposed model improves on the extracted entities across all three axes yielding average gains of up to 6.4% across models. |
Schema-Driven Information Extraction from Heterogeneous Tables (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on information extraction from tables has focused on developing custom pipelines for each table collection. |
| Approach: | They propose a task that transforms tabular data into structured records following a human-authored schema. |
| Outcome: | The proposed task achieves F1 scores ranging from 74.2 to 96.1 while maintaining cost efficiency. |
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction (2026.acl-industry)
Copied to clipboard
| Challenge: | Structured information extraction (IE) pairs values from unstructured text with schema-defined keys. |
| Approach: | They propose a retrieval-augmented generation framework that prunes the output schema space for schema-conditioned information extraction tasks by leveraging schema metadata and few-shot examples. |
| Outcome: | The proposed framework can achieve up to 8.8% increase in micro-F1, 47% reduction in latency, and 48% reduction in token costs on real-world healthcare and e-commerce datasets. |
LLM4RE: A Data-centric Feasibility Study for Relation Extraction (2025.coling-main)
Copied to clipboard
| Challenge: | Relation Extraction (RE) is a critical step in information extraction due to its wide-scale applicability for downstream applications such as Knowledge Base creation and Question Answering (QA). |
| Approach: | They propose to conduct the first feasibility analysis to explore the viability of Large Language Models for RE by investigating their robustness to various RE scenarios stemming from data-specific characteristics. |
| Outcome: | The proposed models are robust to various RE scenarios stemming from data-specific characteristics, but their performance is not yet fully understood. |
ITAKE: Interactive Unstructured Text Annotation and Knowledge Extraction System with LLMs and ModelOps (2024.acl-demos)
Copied to clipboard
| Challenge: | Unstructured text data contains a large amount of valuable knowledge, but there are many tools that do not meet the needs of actual business. |
| Approach: | They propose an unstructured text annotation and knowledge extraction system that integrates Large Language Models and ModelOps to improve model supervision and performance. |
| Outcome: | The proposed system integrates large language models and ModelOps to improve performance in low-resource contexts. |
LMDX: Language Model-based Document Information Extraction and Localization (2024.findings-acl)
Copied to clipboard
Vincent Perot, Kai Kang, Florian Luisier, Guolong Su, Xiaoyu Sun, Ramya Sree Boppana, Zilong Wang, Zifeng Wang, Jiaqi Mu, Hao Zhang, Chen-Yu Lee, Nan Hua
| Challenge: | Large Language Models have revolutionized Natural Language Processing but their application in extracting information from visually rich documents has not been successful. |
| Approach: | They propose a language model-based document information extraction and localization methodology to reframe the document information extract task for a LLM. |
| Outcome: | The proposed method enables extraction of singular, repeated, and hierarchical entities with and without training data. |
A Simple but Effective Approach to Improve Structured Language Model Output for Information Extraction (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have impressive abilities in generating unstructured natural language . performance inconsistent when tasked with producing text that adheres to structured formats . |
| Approach: | They propose a method to generate unstructured natural language using intermediate responses . they use the intermediate responses to organize the output into the desired structure . |
| Outcome: | The proposed method improves performance on NER and RE tasks with minimal effort. |