Challenge: Substantial resources are typically spent on unitizing, the task of identifying precise span boundaries for entity mentions.
Approach: They propose a method that focuses manual efforts on typed position annotations instead of full concept annotation.
Outcome: The proposed procedure reduces the cost of concept annotations by focusing on typed positions instead of full concept annotation.

Similar Papers

Semantic Span Annotation: An Exploratory Study of LLM Annotation (2026.acl-srw)

Copied to clipboard

Challenge: Structured span extraction research is siloed by context length, annotation task, and domain . Identifying a span within a natural language text and affixing it with a semantic label has been considered a core task in NLP .
Approach: They propose a framework for structured span annotation that integrates five datasets under a common JSONL format with character-level offsets.
Outcome: The proposed framework can generalize across four domains under three prompting configurations.
A Span Selection Model for Semantic Role Labeling (D18-1)

Copied to clipboard

Challenge: Existing models for semantic role labeling use BIO tags to predict argument spans . but performance of these approaches is weak .
Approach: They propose a span-based model that takes into account all possible argument spans and scores them for each label.
Outcome: The proposed model achieves state-of-the-art results on the CoNLL-2005 and 2012 datasets.
An Empirical Study on Finding Spans (2022.emnlp-main)

Copied to clipboard

Challenge: Various information extraction tasks require a span finding component, which either directly yields the output or serves as an essential component of downstream linking.
Approach: They propose methods for span finding, the selection of consecutive tokens in text for some downstream tasks.
Outcome: The proposed methods perform better on masked language models and pre-trained encoders than on encoder-decoder models.
Improving Span Representation by Efficient Span-Level Attention (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for generating high-quality span representations are limited by subset of tokens . span-span interactions should play an important role in span encoding, authors argue .
Approach: They propose to introduce span-span interactions and more comprehensive span-token interactions to improve span representations.
Outcome: The proposed model outperforms baseline models on span-related tasks and shows superior performance.
Sent2Span: Span Detection for PICO Extraction in the Biomedical Text without Span Annotations (2021.findings-emnlp)

Copied to clipboard

Challenge: Experiments show that PICO span detection results achieve much higher results for recall when compared to fully supervised methods.
Approach: They propose to extract and then normalise PICO information from clinical trial articles and use crowdsourced sentence-level annotations to detect spans.
Outcome: The proposed method achieves much higher results for recall when compared to fully supervised methods with PICO sentence detection at least as good as human annotations.
Improving Generalization in Language Model-based Text-to-SQL Semantic Parsing: Two Simple Semantic Boundary-based Techniques (2023.acl-short)

Copied to clipboard

Challenge: Pre-trained language models (LMs)2 have been adopted for semantic parsing due to their promising performance and straightforward architectures.
Approach: They propose to use token preprocessing to preserve semantic boundaries of tokens produced by LM tokenizers and special tokens to mark the boundaries of aligned components.
Outcome: The proposed techniques improve the performance of pre-trained language models on two text-to-SQL semantic parsing datasets.
Segmenting Natural Language Sentences via Lexical Unit Analysis (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent work on sequence segmentation models suffer from invalid predictions and a lack of consistency.
Approach: They propose a unified span-based model that embeds every span and computes a score for each segmentation candidate.
Outcome: The proposed model achieves state-of-the-art on 6 of the 3 tasks tested.
SpanBERT: Improving Pre-training by Representing and Predicting Spans (2020.tacl-1)

Copied to clipboard

Challenge: Pre-training methods like BERT mask individual words or subword units, but many tasks involve reasoning about relationships between two or more spans of text.
Approach: They propose a pre-training method that masks contiguous random spans instead of random tokens to train the span boundary representations to predict the entire content of the masked span.
Outcome: The proposed method outperforms BERT and its better-tuned baselines on span selection tasks and on coreference resolution tasks.
Decoding Text Spans for Efficient and Accurate Named-Entity Recognition (2026.acl-industry)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a key component in industrial information extraction pipelines, where systems must satisfy strict latency and throughput constraints in addition to strong accuracy.
Approach: They propose a span-based NER framework that can be used to compute span representations at the final transformer stage, avoiding redundant computation in earlier layers.
Outcome: The proposed framework matches competitive baselines while improving throughput and reducing computational cost.
Enhanced Language Representation with Label Knowledge for Span Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to extract text spans from plain text do not fully exploit label knowledge.
Approach: They propose a model to integrate label knowledge into text representations by encoding texts and annotations independently and then integrating label knowledge with an elaborate-designed semantics fusion module.
Outcome: The proposed model achieves state-of-the-art performance on four benchmarks and reduces training time and inference time by 76% and 77% on average compared with the existing paradigm.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations