Papers with JSON
From Paper to Structured JSON: An Agentic AI Workflow for Compliant BMR Digital Transformation (2026.eacl-industry)
Copied to clipboard
| Challenge: | Agentic AI workflow converts noisy pharmaceutical batch records into validated JSON preserving GMP-critical structure. |
| Approach: | They propose a workflow that transforms unstructured batch records into validated JSON using hybrid OCR, vision–language and schema-guided LLMs. |
| Outcome: | The proposed workflow cuts QA review time from hours to minutes while preserving key GMP-critical structure. |
Evaluating Structured Output Robustness of Small Language Models for Open Attribute-Value Extraction from Clinical Notes (2025.acl-srw)
Copied to clipboard
| Challenge: | Comparative analysis of structured outputs generated by small language models for open attribute-value extraction from clinical notes . structure of outputs improves with targeted prompting and larger models, but declines for longer documents and certain note types. |
| Approach: | They compare the parsability of structured outputs generated by small language models for open attribute-value extraction from clinical notes. |
| Outcome: | The proposed model performs well in open attribute-value extraction tasks, but fails to parse for longer documents and note types. |
IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages (2026.eacl-industry)
Copied to clipboard
| Challenge: | Indic Jailbreak Robustness (IJR) is a judge-free benchmark for adversarial safety across 12 languages. |
| Approach: | They propose a judge-free benchmark for adversarial safety across 12 languages . they find contracts inflate refusals but do not stop jailbreaks . |
| Outcome: | The proposed benchmarks cover 45,216 prompts in JSON and Free tracks. |
SLENDER: Structured Outputs for SLM-based NER in Low-Resource Englishes (2025.acl-industry)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) for low-resource variants of English remains challenging, as most models are trained on datasets predominantly focused on American or British English. |
| Approach: | They propose a new output format for Named Entity Recognition (NER) that achieves a three-fold reduction in inference time compared to JSON format. |
| Outcome: | The proposed output format achieves a three-fold reduction in inference time on average compared to JSON format, which is widely used for structured outputs. |
Open Political Corpora: Structuring, Searching, and Analyzing Political Text Collections with PoliCorp (2025.emnlp-demos)
Copied to clipboard
| Challenge: | PoliCorp provides researchers with access to rich textual data, enabling in-depth analysis of parliamentary discourse over time. |
| Approach: | They present a web portal that allows researchers to search political text corpora . the platform currently contains a collection of transcripts from the german parliament . |
| Outcome: | The proposed platform provides researchers with access to rich textual data, enabling in-depth analysis of parliamentary discourse over time. |
Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance. (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Structured generation is used to extract key output information from large language models (LLMs). |
| Approach: | They examine whether constraints on generation space impact LLMs’ abilities, including reasoning and domain knowledge comprehension. |
| Outcome: | The proposed model is based on a few-shot in-context learning and instruction-following capabilities. |
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages (2025.acl-industry)
Copied to clipboard
Mehant Kammakomati, Sameer Pimparkhede, Srikanth G. Tamilselvam, Prince Kumar, Pushpak Bhattacharyya
| Challenge: | Large Language Models (LLMs) have demonstrated potential in code generation and natural language understanding, but they struggle with code constraints. |
| Approach: | They propose to use Large Language Models to handle constraints represented in code . they use JSON, YAML, XML, Python, and natural language to test their effectiveness . |
| Outcome: | The proposed benchmark shows that LLMs can handle code constraints better than natural language . the results suggest that conscious choice of representations can lead to optimal use of LLM in enterprise use cases involving code constraints. |
How Good Are LLMs at Processing Tool Outputs? (2026.eacl-long)
Copied to clipboard
Kiran Kate, Yara Rizk, Poulami Ghosh, Ashu Gulati, Tathagata Chakraborti, Zidane Wright, Mayank Agarwal
| Challenge: | Real-world task automation tasks require large language models to call tools, which often return complex JSON responses. |
| Approach: | They evaluated 15 open and closed weight models using multiple prompting approaches to evaluate their tool response processing task and their ability to process structured (JSON) responses. |
| Outcome: | The proposed model can process structured (JSON) responses with 3% to 50% performance differences. |
Wiktextract: Wiktionary as Machine-Readable Structured Data (2022.lrec-1)
Copied to clipboard
| Challenge: | Unlike previous Wiktionary extractions, the new extractor, Wiktextract, fully interprets and expands templates and Lua modules in Wiktionaries. |
| Approach: | They propose a machine-readable structured version of Wiktionary that interprets and expands templates and Lua modules. |
| Outcome: | The extracted data is multilingual and includes lemmas, inflected forms, translations, etymology, usage examples, pronunciations, and various morphological, syntactic, semantic, topical, and dialectal annotations. |
PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Structured information extraction from unstructured text is critical for Software 3.0 systems . current approaches to extract structured information from unstructed text are static contracts . |
| Approach: | They propose a system that automates JSON schemas for LLM consumption and optimizes them for LRM consumption. |
| Outcome: | The proposed system improves extraction accuracy and reduces errors by 92% within the first retry and maintaining practical latency. |
Verifiable Format Control for Large Language Model Generations (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods focus on benchmarking general instruction following while overlooking how to improve specific format following ability for small LLMs. |
| Approach: | They propose to synthesize massive datasets to improve LLMs' format following abilities by using a verifiable format following feature. |
| Outcome: | The proposed method improves the format following ability of small LLMs with about 7B parameters. |
Learning to Generate Structured Output with Schema Reinforcement Learning (2025.acl-long)
Copied to clipboard
Yaxi Lu, Haolun Li, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Zhiyuan Liu, Fangming Liu, Maosong Sun
| Challenge: | Recent advances in large language models have facilitated the development of intelligent applications like automatic web search (Qin et al., 2023) Several methods exist for generating JSON strings from LLMs, including Prompting but often miss certain schemas. |
| Approach: | They propose to use 40K different JSON schemas to assess models' ability to generate valid JSON outputs. |
| Outcome: | The proposed model improves both in generating JSON outputs and downstream tasks. |
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)
Copied to clipboard
| Challenge: | a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research . |
| Approach: | a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary . |
| Outcome: | a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web . |
PhotoshopQuiA: A Corpus of Non-Factoid Questions and Answers for Why-Question Answering (L18-1)
Copied to clipboard
| Challenge: | Community Question Answering web sites are used for non-factoid question answering . however, there is a scarcity of available datasets for this task . cnn.com's john m. sutter is releasing a dataset for why-QA . |
| Approach: | They propose a dataset of 2,854 why-question and answer(s) pairs related to Adobe Photoshop usage from five CQA web sites. |
| Outcome: | The new dataset is the first English dataset for Why-QA that focuses on a product . it can be used to build Why-Q systems, evaluate approaches and develop new models . |
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models excel at few-shot learning but their direct application in real-world scenarios is often hindered by their high computational cost. |
| Approach: | They propose a framework that uses Large Language Models for data generation and scoring to improve encoder model performance. |
| Outcome: | The proposed approach improves accuracy from 28.9% to 39.3% on a few-shot MCQA task . |
Eesthetic: A Paralex Lexicon of Estonian Paradigms (2024.lrec-main)
Copied to clipboard
| Challenge: | Eesthetic is a comprehensive Estonian noun and verb lexicon . it documents 5475 nouns inflecting for 28 paradigm cells and 5076 verbs inflection for 51 cells. |
| Approach: | They propose to use Ekilex to generate an Estonian noun and verb lexicon with a set of rules for automatic transcription. |
| Outcome: | The Estonian lexicon is based on the Ekilex database and is openly accessible . it contains a total of 452885 inflected forms and is structured and formatted as a set of CSV tables linked by formal relationships. |
Pre3: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation (2025.acl-long)
Copied to clipboard
Junyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu, Chuheng Du, Hailong Yang, Ruihao Gong, Shengzhong Liu, Fan Wu, Guihai Chen
| Challenge: | Existing methods for structured generation of outputs are inefficient under large inference batches. |
| Approach: | They propose a new LLM-based method that parses LR(1) grammars into a pushdown automaton and exploits deterministic pushdown automation to optimize the constrained LLM decoding efficiency. |
| Outcome: | The proposed method improves time per output token (TPOT) by 40% and throughput by 36% . |
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment (2025.acl-long)
Copied to clipboard
Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, Francois Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila
| Challenge: | Large language models such as GPT-4 have limited their deployment in clinical settings . a novel framework for adapting SLMs into high-performing clinical models is needed . |
| Approach: | They propose a framework for adapting large language models into high-performing clinical models . they pre-instruct experts on relevant medical and clinical corpora and model merging . |
| Outcome: | The proposed framework outperforms the existing model on the CLUE+ benchmark on medical entities and radiology reports. |
DocCGen: Document-based Controlled Code Generation (2024.emnlp-main)
Copied to clipboard
Sameer Pimparkhede, Mehant Kammakomati, Srikanth Tamilselvam, Prince Kumar, Ashok Kumar, Pushpak Bhattacharyya
| Challenge: | Large language models (LLMs) produce state-of-the-art performance on natural language to code generation for resource-rich general-purpose languages like C++, Java, and Python. |
| Approach: | They propose a framework that breaks the NL-to-Code generation task into two steps . they use library documentation to detect the correct libraries and schema rules extracted from the documentation to constrain the decoding . |
| Outcome: | The proposed framework improves different sized language models across all six evaluation metrics, reducing syntactic and semantic errors in structured code. |
BNLP: A Text Annotation Platform for Quality Control of LLM-Generated Annotations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing annotation tools lack support for Large Language Models (LLMs) or use LLMs as one-off preannotation engines, compromising data reliability. |
| Approach: | They propose a text annotation platform that embeds LLM-assisted labeling into a quality-aware collaborative workflow. |
| Outcome: | Experiments show that BNLP reduces annotation time by 74.3% and improves annotation quality by 11.6% over purely manual annotation in LLM-assisted settings. |
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated strong potential for understanding user intent . paper describes system architecture, agent roles, retrieval and scoring methods, knowledge graph schema, and evaluation interfaces . |
| Approach: | They propose a multi-agent research discovery and analysis system that integrates multiple agents to reduce the effort required to find, assess, organize, and understand academic literature. |
| Outcome: | The proposed system reduces the effort required to find, assess, organize, and understand academic literature. |