Decomposing Unitization and Typing for Efficient and Consistent Span-Bound Concept Annotation (2026.findings-acl)
Copied to clipboard
| Challenge: | Substantial resources are typically spent on unitizing, the task of identifying precise span boundaries for entity mentions. |
| Approach: | They propose a method that focuses manual efforts on typed position annotations instead of full concept annotation. |
| Outcome: | The proposed procedure reduces the cost of concept annotations by focusing on typed positions instead of full concept annotation. |
Similar Papers
Semantic Span Annotation: An Exploratory Study of LLM Annotation (2026.acl-srw)
Copied to clipboard
| Challenge: | Structured span extraction research is siloed by context length, annotation task, and domain . Identifying a span within a natural language text and affixing it with a semantic label has been considered a core task in NLP . |
| Approach: | They propose a framework for structured span annotation that integrates five datasets under a common JSONL format with character-level offsets. |
| Outcome: | The proposed framework can generalize across four domains under three prompting configurations. |
A Span Selection Model for Semantic Role Labeling (D18-1)
Copied to clipboard
| Challenge: | Existing models for semantic role labeling use BIO tags to predict argument spans . but performance of these approaches is weak . |
| Approach: | They propose a span-based model that takes into account all possible argument spans and scores them for each label. |
| Outcome: | The proposed model achieves state-of-the-art results on the CoNLL-2005 and 2012 datasets. |
An Empirical Study on Finding Spans (2022.emnlp-main)
Copied to clipboard
| Challenge: | Various information extraction tasks require a span finding component, which either directly yields the output or serves as an essential component of downstream linking. |
| Approach: | They propose methods for span finding, the selection of consecutive tokens in text for some downstream tasks. |
| Outcome: | The proposed methods perform better on masked language models and pre-trained encoders than on encoder-decoder models. |
Improving Span Representation by Efficient Span-Level Attention (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for generating high-quality span representations are limited by subset of tokens . span-span interactions should play an important role in span encoding, authors argue . |
| Approach: | They propose to introduce span-span interactions and more comprehensive span-token interactions to improve span representations. |
| Outcome: | The proposed model outperforms baseline models on span-related tasks and shows superior performance. |
Sent2Span: Span Detection for PICO Extraction in the Biomedical Text without Span Annotations (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Experiments show that PICO span detection results achieve much higher results for recall when compared to fully supervised methods. |
| Approach: | They propose to extract and then normalise PICO information from clinical trial articles and use crowdsourced sentence-level annotations to detect spans. |
| Outcome: | The proposed method achieves much higher results for recall when compared to fully supervised methods with PICO sentence detection at least as good as human annotations. |
Improving Generalization in Language Model-based Text-to-SQL Semantic Parsing: Two Simple Semantic Boundary-based Techniques (2023.acl-short)
Copied to clipboard
| Challenge: | Pre-trained language models (LMs)2 have been adopted for semantic parsing due to their promising performance and straightforward architectures. |
| Approach: | They propose to use token preprocessing to preserve semantic boundaries of tokens produced by LM tokenizers and special tokens to mark the boundaries of aligned components. |
| Outcome: | The proposed techniques improve the performance of pre-trained language models on two text-to-SQL semantic parsing datasets. |
Segmenting Natural Language Sentences via Lexical Unit Analysis (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on sequence segmentation models suffer from invalid predictions and a lack of consistency. |
| Approach: | They propose a unified span-based model that embeds every span and computes a score for each segmentation candidate. |
| Outcome: | The proposed model achieves state-of-the-art on 6 of the 3 tasks tested. |
SpanBERT: Improving Pre-training by Representing and Predicting Spans (2020.tacl-1)
Copied to clipboard
| Challenge: | Pre-training methods like BERT mask individual words or subword units, but many tasks involve reasoning about relationships between two or more spans of text. |
| Approach: | They propose a pre-training method that masks contiguous random spans instead of random tokens to train the span boundary representations to predict the entire content of the masked span. |
| Outcome: | The proposed method outperforms BERT and its better-tuned baselines on span selection tasks and on coreference resolution tasks. |
Decoding Text Spans for Efficient and Accurate Named-Entity Recognition (2026.acl-industry)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a key component in industrial information extraction pipelines, where systems must satisfy strict latency and throughput constraints in addition to strong accuracy. |
| Approach: | They propose a span-based NER framework that can be used to compute span representations at the final transformer stage, avoiding redundant computation in earlier layers. |
| Outcome: | The proposed framework matches competitive baselines while improving throughput and reducing computational cost. |
Enhanced Language Representation with Label Knowledge for Span Extraction (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to extract text spans from plain text do not fully exploit label knowledge. |
| Approach: | They propose a model to integrate label knowledge into text representations by encoding texts and annotations independently and then integrating label knowledge with an elaborate-designed semantics fusion module. |
| Outcome: | The proposed model achieves state-of-the-art performance on four benchmarks and reduces training time and inference time by 76% and 77% on average compared with the existing paradigm. |