Papers with XML

9 papers
Evaluating Structured Output Robustness of Small Language Models for Open Attribute-Value Extraction from Clinical Notes (2025.acl-srw)

Copied to clipboard

Challenge: Comparative analysis of structured outputs generated by small language models for open attribute-value extraction from clinical notes . structure of outputs improves with targeted prompting and larger models, but declines for longer documents and certain note types.
Approach: They compare the parsability of structured outputs generated by small language models for open attribute-value extraction from clinical notes.
Outcome: The proposed model performs well in open attribute-value extraction tasks, but fails to parse for longer documents and note types.
Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance. (2024.emnlp-industry)

Copied to clipboard

Challenge: Structured generation is used to extract key output information from large language models (LLMs).
Approach: They examine whether constraints on generation space impact LLMs’ abilities, including reasoning and domain knowledge comprehension.
Outcome: The proposed model is based on a few-shot in-context learning and instruction-following capabilities.
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated potential in code generation and natural language understanding, but they struggle with code constraints.
Approach: They propose to use Large Language Models to handle constraints represented in code . they use JSON, YAML, XML, Python, and natural language to test their effectiveness .
Outcome: The proposed benchmark shows that LLMs can handle code constraints better than natural language . the results suggest that conscious choice of representations can lead to optimal use of LLM in enterprise use cases involving code constraints.
Transc&Anno: A Graphical Tool for the Transcription and On-the-Fly Annotation of Handwritten Documents (L18-1)

Copied to clipboard

Challenge: Transc&Anno is a web-based collaboration tool for linguists to facilitate the transcription of text images and their shallow on-the-fly annotation.
Approach: They propose a web-based collaboration tool that allows the transcription of text images and their shallow on-the-fly annotation.
Outcome: The Transc&Anno tool can be used for any type of corpora requiring transcription and shallow on-the-fly annotation resulting in inline XML.
A Parallel Corpus of Arabic-Japanese News Articles (L18-1)

Copied to clipboard

Challenge: a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs.
Approach: They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus.
Outcome: The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it .
The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used to label data across many domains and for myriad tasks.
Approach: They ask large language models to label data using a series of decisions by practitioners . they find that even the smallest perturbations can change the LLM's answer .
Outcome: The proposed model can be used to quickly get a response for arbitrary tasks.
Ranking-Based Autoencoder for Extreme Multi-label Classification (N19-1)

Copied to clipboard

Challenge: Existing methods to solve label dependency and noisy labeling problems are limited . experimental results show the proposed method is competitive to state-of-the-art methods .
Approach: They propose a deep learning XML method with word-vector-based self-attention followed by ranking-based AutoEncoder architecture to solve these problems.
Outcome: The proposed method is competitive to state-of-the-art methods on benchmark datasets.
An Architecture of resolving a multiple link path in a standoff-style data format to enhance the mobility of language resources (2022.lrec-1)

Copied to clipboard

Challenge: XML data formats that represent a semantic data model are difficult to convert into other formats . this difficulty causes a problem in the reuse of data especially in a personal data management environment.
Approach: They propose a new approach to transform a link structure into an instance structure on a marked-up scheme.
Outcome: The proposed format is based on a so-called standoff-style data format in XML . the proposed format injures the mobility of data because it is hard to convert it into other formats .
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)

Copied to clipboard

Challenge: a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research .
Approach: a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary .
Outcome: a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations