Papers with annotation
Copied to clipboard
| Challenge: | DoTAT is a domain-oriented text annotation tool that can reduce the time for event annotation by 19.7% . the tool supports multi-person collaborative process with automatically merging and review . |
| Approach: | They propose a domain-oriented text annotation tool called DoTAT . it provides multi-person collaborative process with automatic merging and review . |
| Outcome: | The proposed tool can reduce the time for event annotation by 19.7% compared with existing tools. |
Copied to clipboard
| Challenge: | Mention detection is an important preprocessing step for downstream applications such as NER and coreference resolution. |
| Approach: | They propose and compare three approaches to mention detection using ELMO embeddings and a biaffine classifier. |
| Outcome: | The proposed model outperforms state-of-the-art models on the GENIA corpora and improves on mention recall. |
Copied to clipboard
| Challenge: | Mapping and navigation services struggle to handle natural language geospatial queries. |
| Approach: | They introduce an extensible open-source framework that streamlines the creation of reproducible, traceable map-based QA datasets. |
| Outcome: | a new open-source framework streamlines the creation of reproducible, traceable map-based QA datasets. |
Copied to clipboard
| Challenge: | Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names . |
| Approach: | They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor. |
| Outcome: | The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty. |
Copied to clipboard
| Challenge: | Recent studies have focused on the problem of generalizing from a few examples per category. |
| Approach: | They propose to use feature space data augmentation methods to improve intent classification performance in few-shot setting. |
| Outcome: | The proposed methods improve intent classification performance in few-shot setting beyond transfer learning approaches. |
Copied to clipboard
| Challenge: | a new annotation tool is designed to fill the niche of a lightweight interface for terminal users . current tools are built with direct manipulation via a Graphical User Interface (GUI) this approach is time-consuming and difficult to modify . |
| Approach: | They propose a terminal-based annotation tool that supports multiple annotations . they use a text-based interface that uses almost the entire screen to display documents . |
| Outcome: | The proposed tool is designed to fill the niche of a lightweight interface for users with a terminal-based workflow. |
Copied to clipboard
| Challenge: | Annotation conflict resolution is crucial for machine learning, says a new study . past work on annotation conflict resolution assumed data is collected at once . a a supervised neural model can resolve conflicts in data annotation but requires access to high-quality data . |
| Approach: | They propose an approach to resolve annotation conflicts in a real-world context using a German dialog system. |
| Outcome: | The proposed approach improves on a real-world dataset with 3.5M utterances in German. |
Copied to clipboard
| Challenge: | NLP is based on high-quality annotated datasets, but many other tasks have unique characteristics not considered by standard models of annotation. |
| Approach: | tutorial aims to connect NLP researchers with state-of-the-art aggregation models for canonical language annotation tasks. |
| Outcome: | This tutorial aims to connect NLP researchers with state-of-the-art aggregation models for a diverse set of canonical language annotation tasks. |
Copied to clipboard
| Challenge: | Active learning (AL) is a widely-used training strategy for maximizing predictive performance subject to a fixed annotation budget. |
| Approach: | They propose to use active learning to optimize predictive performance . they find that current approaches do not generalize reliably across models and tasks . |
| Outcome: | The proposed approach outperforms training on i.i.d. datasets on supervised learning tasks. |
Copied to clipboard
| Challenge: | Several approaches to active learning are available, including confidence-based, diversity-based and committee-based. |
| Approach: | They propose to use a baseline and a skyline to measure the accuracy of the unannotated sample pool. |
| Outcome: | The proposed model outperforms a random selection baseline and a skyline approach. |
Copied to clipboard
| Challenge: | a recent study shows that annotating sentiments is difficult and difficult. |
| Approach: | They propose to integrate holder and expression information into sentiment analysis to improve target extraction . they perform experiments on eight English datasets to determine whether annotating expressions improves target extraction. |
| Outcome: | The proposed approach improves target extraction and classification on English datasets. |
Copied to clipboard
| Challenge: | Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored. |
| Approach: | They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
| Outcome: | The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
Copied to clipboard
| Challenge: | Using a neural network, we annotate whether tension is increasing, decreasing, or staying unchanged. |
| Approach: | They propose a machine-assisted method for the identification of tension development using a neural network based prediction model. |
| Outcome: | The proposed method is compared with other methods in in-house and crowdsourced environments. |
Copied to clipboard
| Challenge: | Edu-ConvoKit is an open-source library for analyzing education conversation data. |
| Approach: | They introduce Edu-ConvoKit, an open-source library for conversation data analysis. |
| Outcome: | The open-source library handles pre-processing, annotation and analysis of education conversation data. |
Copied to clipboard
| Challenge: | Annotating scientific literature directly on PDF documents can greatly improve the labeling efficiency of scientists whose annotation costs are very high. |
| Approach: | They propose an integrated onsite scientific literature annotation tool for natural scientists and Natural Language Processing (NLP) researchers. |
| Outcome: | The proposed tool supports the whole lifecycle of corpus generation including i)project management, ii)resource management, and iv)ontology management, as well as manual annotation, onsite auto annotation, and vi)task statistic. |
Copied to clipboard
| Challenge: | Existing RE surveys focus on modeling techniques, but there are few that are based on real-world scenarios. |
| Approach: | They propose to survey RE datasets and revisit the task definition and its adoption by the community. |
| Outcome: | The proposed approach improves the reliability of RE evaluations across multiple datasets and reveals significant discrepancies in annotations. |
Copied to clipboard
| Challenge: | Annotators are asked to annotate coreferent spans of text, which is unnatural . we present an alternative in which annotators can preprocess documents and assign pronouns to entities. |
| Approach: | They propose an alternative in which annotators are asked to assign pronouns to entities and preprocess documents to create a knowledge base. |
| Outcome: | The proposed model-based approach leads to faster annotation and higher inter-annotator agreement and opens up an alternative approach to coreference resolution. |
Copied to clipboard
| Challenge: | a new language-independent treebank annotation tool supports rich annotations with discontinuous constituents and function tags. |
| Approach: | They propose a language-independent treebank annotation tool supporting rich annotations with discontinuous constituents and function tags. |
| Outcome: | The proposed tool supports rich annotations with discontinuous constituents and function tags. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) tasks in the Indonesian language are still lacking data for the majority of languages, including Indonesian. |
| Approach: | They re-annotated an open dataset with 2,000 sentences and compared the results with a bidirectional long short-term memory and conditional random field approach. |
| Outcome: | The proposed approach improved the prediction score and consistent organization tag for the Indonesian language. |
Copied to clipboard
| Challenge: | Formulating effective search queries can be a daunting task for users when they lack expertise in a specific domain or are not proficient in the language of the content. |
| Approach: | QueryExplorer is an interactive query generation, reformulation, and retrieval interface with support for Hug-gingFace generation models and PyTerrier’sretrieval pipelines and datasets. |
| Outcome: | QueryExplorer is an interactive query generation, reformulation, and retrieval interface with support for Hug-gingFace generation models and PyTerrier’sretrieval pipelines and datasets, and extensivelogging of human feedback. |
Copied to clipboard
| Challenge: | a prototype model trained on a small amount of data is not available, leading to limited prediction performance. |
| Approach: | They propose a human-in-the-loop process that generates a set of valid candidates and allows users to quickly traverse the set and filter incorrect parses. |
| Outcome: | The proposed process can be used to efficiently traverse the candidate set and select the correct parse, with minimal modification when necessary. |
Copied to clipboard
| Challenge: | Existing tools for document-level annotation lack document-based quality and flexibility. |
| Approach: | a new annotation tool is being developed for industry and research use cases . a configurable user interface and a RESTful API are included . authors propose to use ACTIVEANNO as default for document-level annotation . |
| Outcome: | ACTIVEANNO is an annotation tool for industry and research use cases. |
Copied to clipboard
| Challenge: | Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. |
| Approach: | They introduce a Question Decomposition Meaning Representation (QDMR) for questions . they demonstrate that QDMRs can be annotated at scale using a hotpotQA dataset . |
| Outcome: | The proposed model outperforms several natural baselines in the open-domain question answering hotpotQA dataset and can be deterministically converted to a pseudo-SQL formal language. |
Copied to clipboard
| Challenge: | Existing systems for technologyenhanced learning address skills on recalling, explaining, and applying knowledge, e.g., in automatically generated language learning exercises and math word problems. |
| Approach: | They propose to leverage a NLP model to support experts in their further data annotation with automatic suggestions and provide automatic feedback for students. |
| Outcome: | The proposed system improves on two user studies on diagnostic reasoning in medicine and teacher education and can be extended to further use cases. |
Copied to clipboard
| Challenge: | Behavioral therapy notes are important for legal compliance and patient care, but quality standards for them remain underdeveloped. |
| Approach: | They propose a rubric for evaluating therapy notes across key dimensions: completeness, conciseness, faithfulness. |
| Outcome: | The proposed evaluation framework improves on therapist-written notes and LLM-generated notes. |
Copied to clipboard
| Challenge: | Recent work on few-shot classification has addressed the issue of data prioritisation of unlabelled data. |
| Approach: | They propose a weighted approach that uses a set of pattern-exploiting training models to actively select unlabelled data as candidates for annotation. |
| Outcome: | The proposed approach shows consistent improvement over baseline methods on two technical fact-checking datasets and using six different pretrained language models. |
Copied to clipboard
| Challenge: | Despite its importance, this direction of research has not been explored as much. |
| Approach: | They propose to use counterfactual simulations to evaluate paper novelty detection models . they ask models to differentiate papers at time t and counterf actual paper from future time . |
| Outcome: | The proposed models can be compared against a set of papers with a given date and with different annotations. |
Copied to clipboard
| Challenge: | Using corpus annotation, we show huge differences in metaphor usage between different registers and specific properties of registers. |
| Approach: | They present their work on corpus annotation for metaphor in germany . they focus on metaphors that can serve as register markers and be reliably indentified . |
| Outcome: | The proposed corpus annotations show huge differences in metaphor usage between different registers and specific properties of registers. |
Copied to clipboard
| Challenge: | Efficient data collection is important for advancing research and building time-sensitive applications. |
| Approach: | They propose an open-source platform that standardizes the data collection pipeline . it includes customizable user interface components, automated annotator qualification, and saved pipelines . |
| Outcome: | The proposed platform simplifies data annotation significantly on diverse datasets . it can be used by researchers and engineers to improve reproducibility and minimize overhead . |
Copied to clipboard
| Challenge: | Existing work on dialogue meaning representations is limited in scalability for complex expressions. |
| Approach: | They propose a pliable and easily extendable representation for task-oriented dialogue . they propose an inheritance hierarchy mechanism focusing on domain extensibility . |
| Outcome: | The proposed representation can be easily extended to a task-oriented dialogue dataset. |
Copied to clipboard
| Challenge: | Existing methods to obtain high-quality annotations under limited budgets focus on selecting informative data for expert annotations while the rest of the data is assigned to model annotation. |
| Approach: | They propose a semi-automatic annotation framework that uses error-aware triage and bi-weighting mechanisms to obtain high-quality annotations under limited budget. |
| Outcome: | The proposed framework outperforms baselines in the data annotation problem under limited budgets. |
Copied to clipboard
| Challenge: | a new dataset of eye-tracking and electroencephalography captures language understanding . eye movement data provides millisecond-accurate records of where humans look when reading . |
| Approach: | They recorded and preprocessed eye-tracking and electroencephalography data during natural reading and during annotation. |
| Outcome: | The study combines eye-tracking and electroencephalography to capture the reading process . the data can be used to evaluate state-of-the-art machine learning systems . |
Copied to clipboard
| Challenge: | Sanskrit Voyager enables users to search for words and phrases as they actually appear in texts . evaluation shows over 92% parsing accuracy on complex compounds compared to BuddhaNexus . |
| Approach: | Sanskrit Voyager is a web application for searching, reading and analyzing the Sanskrt literary corpus. |
| Outcome: | Sanskrit Voyager is a web application for searching, reading, and analyzing the Sanskrt literary corpus. |
Copied to clipboard
| Challenge: | Appraise is an open-source framework for crowd-based annotation tasks . it is used for shared tasks at the conference on machine translation and at IWSLT 2017 . |
| Approach: | They present an open-source framework for crowd-based annotation tasks . they describe the entire lifecycle of an Appraise evaluation campaign . |
| Outcome: | The proposed framework is used to run evaluation campaigns at the WMT Conference on Machine Translation and at IWSLT 2017 . it has been adopted by the translator team at Microsoft Translator for internal quality monitoring . |
Copied to clipboard
| Challenge: | LOME is a system for performing multilingual information extraction with large ontologies. |
| Approach: | They propose a system for multilingual information extraction with a framenet parser . LOME is available as a Docker container on Docker Hub and a lightweight version is available on the web . |
| Outcome: | The proposed system outperforms or is competitive with the (monolingual) state-of-the-art . it can be used to build knowledge graphs with large ontologies and across multiple languages . |
Copied to clipboard
| Challenge: | Existing approaches to detect novel intents have been tested in the last decade. |
| Approach: | They propose a framework to detect multiple novel intents with budgeted human annotation cost. |
| Outcome: | The proposed framework outperforms baseline methods in terms of accuracy and F1-score on a set of benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to literature analysis lack transparency and information retrieval module. |
| Approach: | GraphMind is an easy-to-use interactive web tool designed to assist users in evaluating novelty of scientific papers or drafted ideas. |
| Outcome: | GraphMind enables users to capture the main structure of a scientific paper, explore related ideas through various perspectives, and assess novelty via providing verifiable contextual insights. |
Copied to clipboard
| Challenge: | Document workflow copilot system that can understand user intent and execute tasks accordingly to help users streamline their workflows. |
| Approach: | They propose an AI-assisted document workflow copilot system capable of understanding user intent and executing tasks accordingly. |
| Outcome: | The proposed system can understand user intent and execute tasks accordingly to help users streamline their workflows. |
Copied to clipboard
| Challenge: | Lexical normalisation is the task of identifying and normalising non-canonical tokens (e.g. erroneous spelling, acronyms, etc.) in noisy, non-standard, corpora. |
| Approach: | They propose to use LexiClean to annotate multiple tasks in noisy corpora using in situ token modification and annotation that can be rapidly applied corpus wide. |
| Outcome: | The proposed tool can be rapidly applied corpus wide and can identify and normalise noisy, non-standard, and domain specific corpora. |
Copied to clipboard
| Challenge: | a new approach to knowledge extraction (KE) is needed for the health domain. |
| Approach: | They propose an approach to extracting knowledge about antidepressant drug nonadherence from health forums. |
| Outcome: | The proposed approach can be used to extract knowledge about antidepressant drug nonadherence from health forums. |
Copied to clipboard
| Challenge: | Active learning (AL) is a special family of machine learning algorithms designed to reduce labeling costs and improve accuracy. |
| Approach: | They developed an open-source annotation system for NLP tasks equipped with features to make AL effective in real-world annotation projects. |
| Outcome: | ALANNO is an open-source annotation system for NLP tasks equipped with features to make AL effective in real-world annotation projects. |
Copied to clipboard
| Challenge: | Empirical natural language processing (NLP) systems involve interoperation among multiple components . a wealth of NLP toolkits exist ( 4), such as spaCy, DKPro, CoreNLP. |
| Approach: | They propose a unified open-source framework that supports fast development of NLP workflows . framework includes processors for NLP tasks, visualization, and annotation . |
| Outcome: | The framework offers processors for NLP tasks, visualization, and annotation, and is extensible . it is delivered through two modularized yet integratable open-source projects, Forte and Stave . |
Copied to clipboard
| Challenge: | Using a web-based coreference annotation suite, we demonstrate that non-expert annotators can be trained to perform and review coreference resolution tasks. |
| Approach: | They propose a web-based coreference annotation suite oriented for crowdsourcing that provides guided onboarding and a novel algorithm for a reviewing phase. |
| Outcome: | The proposed tool provides guided onboarding and a novel algorithm for a review phase. |
Copied to clipboard
| Challenge: | Recent advances in natural language processing (NLP) are fuelled by high quality annotated datasets. |
| Approach: | They introduce Redcoat, a web-based annotation tool that supports collaborative hierarchical entity typing. |
| Outcome: | The proposed annotation tool reduces the time it takes for project creators to set up and distribute projects to annotators and scales the workload depending on the number of active annotator. |
Copied to clipboard
| Challenge: | Fallacy detection is an open challenge in NLP and has shown to be intrinsically difficult for both humans and machines. |
| Approach: | They propose a framework that minimizes annotation errors whilst keeping signals of human label variation. |
| Outcome: | The proposed framework minimizes annotation errors while keeping signals of human label variation. |
Copied to clipboard
| Challenge: | In this paper, we automatically create sentiment dictionaries for predicting financial outcomes excess return and volatility. |
| Approach: | They propose to automatically adapt a domain-general dictionary to a financial domain and then manually adapt it to dictionaries for the finance domain. |
| Outcome: | The proposed dictionary outperforms the previous state of the art in predicting financial variables excess return and volatility. |
Copied to clipboard
| Challenge: | Existing methods for obtaining pathway information from biomedical literature rely on simplifying assumptions that limit their ability to capture true complexity of biological reactions. |
| Approach: | They propose a web-based platform to facilitate collaborative pathway graph annotation. |
| Outcome: | The platform supports multi-user collaboration with real-time monitoring, curation, and interactive pathway graph visualization. |
Copied to clipboard
| Challenge: | a prerequisite for the creation of a goal-oriented neural network dialogue system is a dataset that represents typical dialogue scenarios and includes various semantic annotations. |
| Approach: | They propose a web-based platform for collecting and writing goal-oriented dialogue samples. |
| Outcome: | The proposed platform is language-independent and is currently being used to collect dialogue samples in Latvian . |
Copied to clipboard
| Challenge: | Identifying and understanding the argumentative discourse structure in text has been a critical task in argument mining. |
| Approach: | They propose a context-aware Transformer-based argument structure prediction model that outperforms models that rely on features or only encode limited contexts. |
| Outcome: | The proposed model outperforms models that rely on features or encode limited contexts on five domains and on peer reviews on five different domains. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive zero-shot performance on a wide range of NLP tasks. |
| Approach: | They propose to use large language models to augment extractive reading comprehension datasets by fine-tuning their annotations and comparing their performance to human annotators. |
| Outcome: | The proposed model can be used to augment extractive reading comprehension datasets. |
Copied to clipboard
| Challenge: | Clinical letters are written by doctors and typically contain complex medical language that is beyond the scope of the lay reader. |
| Approach: | They propose to augment existing neural text simplification software with a phrase table that links medical terminology to simpler vocabulary by mining SNOMED-CT data. |
| Outcome: | The proposed system is easier to understand than the existing system and the phrase table without it. |
Copied to clipboard
| Challenge: | Annotated data is still central to NLP and Generative AI, yet the demands on annotation have grown in both scale and complexity. |
| Approach: | They introduce Potato 2.0, an open source annotation platform for easy deployment and customization. |
| Outcome: | The new version of potato supports 39 different types of annotation tasks and multiple AI-assistance features. |
Copied to clipboard
| Challenge: | Existing methods for event extraction require expensive annotation and are not extensible to new event ontologies. |
| Approach: | They propose to use textual entailment and/or question answering queries to extract a zero-shot event from a set of TE and/ or QA queries. |
| Outcome: | The proposed method achieves acceptable results on ACE-2005 and ERE, but there is still a large gap from supervised approaches. |
Copied to clipboard
| Challenge: | Existing methods degrade under disfluencies, interruptions, and speaker overlap, yet large real-call corpora are rarely shareable. |
| Approach: | They propose a benchmark and semantic synthetic data generation pipeline that generates linguistically varied training data via (1) LLM-sampled entity values, (2) curated linguistic verbalization patterns covering diverse disfluencies and entity-specific readout styles, and (3) a value–transcript consistency filter. |
| Outcome: | The proposed pipeline outperforms zero-shot baselines and matches or closely approaches human-tuned prompts on real customer transcripts. |
Copied to clipboard
| Challenge: | Existing NLP tasks and benchmarks do not cover all NP-mediated relations . we aim to enrich each NP in a text with all the preposition-mediated relationships that hold between it and other NPs in the text. |
| Approach: | They propose a task to enrich NPs with preposition-mediated relations that hold between them . they build a large-scale dataset and analyze the data to test the task . |
| Outcome: | The proposed task is based on a large-scale dataset and fine-tuned language models. |
Copied to clipboard
| Challenge: | Clinical notes contain important information about medical decisions embedded within unstructured text. |
| Approach: | They propose an open-source interactive system that automatically extracts medical decisions from clinical text. |
| Outcome: | The open-source system extracts and visualizes medical decisions from clinical text. |
Copied to clipboard
| Challenge: | Existing efforts to bolster LLMs’ proficiency in Cypher generation are hindered by the lack of annotated datasets of Query-Cypher pairs. |
| Approach: | They propose a method for constructing a synthetic Query-Cypher pair dataset using LLM prompting and template-filling. |
| Outcome: | The proposed method enhances the performance of LLMs on Text2Cypher task via SFT. |
Copied to clipboard
| Challenge: | Lack of aspect-level labeled data is a major obstacle in sentiment classification due to high cost . document-level labels like reviews are easily accessible from online websites . |
| Approach: | They propose a transfer capsule network model for transferring document-level knowledge to aspect-level sentiment classification by encapsulating sentence-level semantic representations into semantic capsules. |
| Outcome: | The proposed model can transfer document-level knowledge to aspect-level sentiment classification. |
Copied to clipboard
| Challenge: | Named-entity recognition is a task that uses named entities to classify texts . annotated data are time and money consuming, since they need to be created by experts of the domain of the annotation that is going to be done . |
| Approach: | They present an Italian dataset for Named-entity recognition with manual annotations and a semi-automatically annotated part. |
| Outcome: | The proposed dataset covers different styles and language uses, and is the largest in Italy. |
Copied to clipboard
| Challenge: | Existing document embedding approaches focus on capturing sequences of words in documents . however, some document classification and regression tasks need to consider discourse structure of text . |
| Approach: | They propose an unsupervised approach to capture discourse structure in terms of coherence and cohesion for document embedding that does not require expensive parsers or annotation. |
| Outcome: | The proposed method improves essay Organization scoring and Argument Strength scoring. |
Copied to clipboard
| Challenge: | Learning a second language requires exposition to texts, especially for the acquisition of vocabulary. |
| Approach: | They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia . |
| Outcome: | The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency. |
Copied to clipboard
| Challenge: | Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development. |
| Approach: | They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs. |
| Outcome: | The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans. |
Copied to clipboard
| Challenge: | SzegedKoref is a treebank of Hungarian that contains manual annotation at several linguistic layers. |
| Approach: | They introduce a Hungarian corpus in which coreference relations are manually annotated. |
| Outcome: | The proposed corpus can be used in training and testing machine learning based coreference resolution systems. |
Copied to clipboard
| Challenge: | Recent efforts on text-to-audio generation are exploring fine-grained controllability . however, their performance at scale is limited due to data scarcity . |
| Approach: | They propose a multi-task learning problem for high-controllability text-to-audio generation . they propose scalable diffusion transformers that augment condition information in sequence . |
| Outcome: | The proposed method outperforms existing methods on objective and subjective evaluations. |
Copied to clipboard
| Challenge: | Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature. |
| Approach: | They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models. |
| Outcome: | The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models. |
Copied to clipboard
| Challenge: | Existing deep learning models for morphological processing require a large amount of annotated data. |
| Approach: | They propose a deep active learning method that uses only informative samples to reduce the need for annotated data. |
| Outcome: | The proposed method achieves the same results as the state-of-the-art model on Egyptian Arabic with only about 30% of annotated data. |
Copied to clipboard
| Challenge: | Existing methods fail to adequately capture the temporal volatility inherent in policy-related sentiments, arguing that continuous time-series clustering and model merging achieve superior performance. |
| Approach: | They propose to use continuous time-series clustering to select data points for annotation based on temporal trends and then apply model merging techniques. |
| Outcome: | The proposed methods outperform existing methods by an average F1-score of 2.71% on temporally representative data. |
Copied to clipboard
| Challenge: | Aspect-level sentiment classification aims to detect the sentiment polarity of a given opinion target in a sentence. |
| Approach: | They propose a novel attention transfer network which can exploit attention from document-level sentiment datasets to improve the attention capability of the aspect-level classification task. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two ASC benchmark datasets. |
Copied to clipboard
| Challenge: | Fully data driven Chatbots suffer from inconsistent behaviour across their turns due to a general difficulty in controlling parameters like their assumed background personality and knowledge of facts. |
| Approach: | They propose a model that is based on pre-specified facts and opinions and validates the dialogues for adherence to their given fact and opinion profile. |
| Outcome: | The proposed model is able to generate opinionated responses that are judged to be natural and knowledgeable and show attentiveness. |
Copied to clipboard
| Challenge: | a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented . |
| Approach: | They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations . |
| Outcome: | The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages . |
Copied to clipboard
| Challenge: | Existing methods to annotate large language models rely on a fixed set of human-annotated exemplars, which are not always the most effective for different tasks. |
| Approach: | They propose a method to adapt large language models to different tasks with task-specific example prompts (annotated with human-designed CoT reasoning) they introduce several metrics to characterize uncertainty so as to select the most uncertain questions for annotation. |
| Outcome: | The proposed method significantly improves performance on eight complex reasoning tasks. |
Copied to clipboard
| Challenge: | Existing studies on large-scale labeled support sets are not feasible in practical scenarios. |
| Approach: | They introduce a language model-based determinant point process that considers uncertainty and diversity of unlabeled instances for optimal selection. |
| Outcome: | The proposed method can effectively select canonical examples on 9 NLU and 2 Generation datasets. |
Copied to clipboard
| Challenge: | Existing approaches for few-shot transfer show significant gain over zero-shot transfers . language resource distribution is skewed across the world's languages . proposed methods use multiple measures such as data entropy and gradient embedding . |
| Approach: | They propose a loss embedding method for sequence labeling tasks that induces diversity and uncertainty sampling similar to gradient embeddment. |
| Outcome: | The proposed methods outperform baseline methods for POS tagging, NER, and NLI tasks for up to 20 languages. |
Copied to clipboard
| Challenge: | Large-scale pretrained language models have brought substantial advances to the natural language processing field. |
| Approach: | They present an internationalized annotation and human evaluation bundle, called Textinator, along with documentation and video tutorials. |
| Outcome: | The proposed tool is compared to other tools along 9 different axes and is available in multiple languages. |
Copied to clipboard
| Challenge: | Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated. |
| Approach: | They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance. |
| Outcome: | The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings. |
Copied to clipboard
| Challenge: | Existing video captioning algorithms are heavily dependent on supervised training data. |
| Approach: | They propose to train the video captioning model on labeled and unlabeled data jointly in a semi-supervised learning manner. |
| Outcome: | The proposed model outperforms state-of-the-art semi-supervised learning approaches on VATEX, MSR-VTT and MSVD datasets. |
Copied to clipboard
| Challenge: | a generic transformer-based model can achieve competitive performance with minimal code-generation-specific inductive bias design. |
| Approach: | They investigate whether a generic transformer-based seq2seq model can achieve competitive performance with minimal code-generation-specific inductive bias design. |
| Outcome: | The proposed model achieves 81.03% exact match accuracy on Django and 32.57 BLEU score on CoNaLa. |
Copied to clipboard
| Challenge: | Existing diagnostic approaches rely on expensive expert annotations and ”LLM-as-a-judge” paradigms. |
| Approach: | They propose a framework for semantic failure attribution that identifies responsible agents and the originating error step. |
| Outcome: | The proposed framework outperforms baselines in step-level localization and validation. |
Copied to clipboard
| Challenge: | Existing studies show that labeling in crowdsourcing annotations is not an annotation artifact but rather a core linguistic phenomenon. |
| Approach: | They propose to retrieve unlabeled data with a local sensitivity and hardness-aware acquisition function. |
| Outcome: | The proposed method achieves consistent gains over the commonly used active learning strategies in various classification tasks. |
Copied to clipboard
| Challenge: | Text generated by Large Language Models (LLMs) may contain plausible but incorrect information known as hallucinations. |
| Approach: | They extend the label set for verdict prediction to capture claim-evidence relationships humans would commonly interpret as supported or refuted. |
| Outcome: | The proposed system improves F1 by 4 percentage points compared to baseline. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models ignore the potential of unlabeled data. |
| Approach: | They propose a framework that allows users to unleash the power of unlabeled data via self-training. |
| Outcome: | The proposed framework outperforms active learning and self-training baselines and improves the label efficiency of PLM fine-tuning by 56.2% on average. |
Copied to clipboard
| Challenge: | Using Ontonotes, documents in certain genres were split into smaller parts for ease of annotation. |
| Approach: | They propose to merge annotations from documents split into smaller parts in Ontonotes for ease of annotation. |
| Outcome: | The proposed corpus restores documents to their original form, revealing dramatic increases in length in certain genres. |
Copied to clipboard
| Challenge: | concessive discourse relations are crucial for understanding language, and annotation is performed on polysemous conjunctions. |
| Approach: | They present an annotation method for the Japanese conjunctions nagara and tsutsuke, and an annotation for the ambiguous Japanese conjunction tokorode. |
| Outcome: | The annotations for nagara and tsutsuke were performed on Japanese conjunctions and revealed the characteristics of concession that became apparent during annotation. |
Copied to clipboard
| Challenge: | Few speech resources describe interruption phenomena, especially for TV and media content. |
| Approach: | They propose to annotation Transition-Relevance Places (TRPs) and Floor-Taking event types on an existing French TV and Radio broadcast corpus to facilitate studies of interruptions and turn-taking. |
| Outcome: | The proposed annotations on an existing French TV and Radio broadcast corpus show they are reliable and reliable . |
Copied to clipboard
| Challenge: | Current foundation models have shown impressive performance across various tasks, but they are not effective for everyone due to the imbalanced geographical and economic representation of the data used in the training process. |
| Approach: | They propose to identify the data to be annotated to balance model performance and annotation costs by finding countries with visual similarity for the topics. |
| Outcome: | The proposed methods improve model performance and reduce annotation costs by using data from countries with higher visual similarity for these topics. |
Copied to clipboard
| Challenge: | Existing annotation tools are not efficient for the annotation of corpora and are not error-free. |
| Approach: | They propose to extend existing annotation tools by evaluating their flexibility and efficiency. |
| Outcome: | The proposed system performs platform-independent multimodal annotations and annotates complex textual structures. |
Copied to clipboard
| Challenge: | Existing work on event extraction relies on labor-intensive annotation, ignoring semantic meaning of event types' labels. |
| Approach: | They propose a zero-shot event extraction approach that first identifies events with existing tools and then maps them to a given taxonomy of event types in a no-shot manner. |
| Outcome: | The proposed approach doubles the performance of previous approaches on a ACE-2005 dataset . it leverages label representations induced by pre-trained language models and maps events to the target types . |
Copied to clipboard
| Challenge: | Detecting fine-grained differences in content conveyed in different languages is expensive and hard to scale. |
| Approach: | They propose a training strategy for multilingual BERT models by learning to rank divergent examples of varying granularity. |
| Outcome: | The proposed model improves the prediction and annotation of fine-grained semantic divergences. |
Copied to clipboard
| Challenge: | Existing temporal relation (TempRel) annotation schemes have low inter-annotator agreements even between experts, suggesting that the current annotation task needs a better definition. |
| Approach: | They propose to annotate temporal relation (TempRel) annotation schemes based on event start-points instead of a conventional 60’s-80’s model. |
| Outcome: | The proposed model improves IAA from the conventional 60’s to 80’s and can be used by crowdsourcing to alleviate labor intensity. |
Copied to clipboard
| Challenge: | a corpus of 2016 debates and commentary contains 4,648 argumentative propositions annotated with fine-grained proposition types. |
| Approach: | They propose a machine learning-human workflow for annotating for four complex proposition types . they demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates . |
| Outcome: | The proposed method can be used by technical researchers seeking more nuanced representations of argument . it can also be used to analyze rhetorical strategies and structure in presidential debates . |
Copied to clipboard
| Challenge: | Existing active learning approaches for natural language processing ignore the characteristics of natural language. |
| Approach: | They propose a pre-trained language model based active learning approach for sentence matching that provides linguistic criteria to measure instances and help select more effective instances for annotation. |
| Outcome: | The proposed approach can achieve greater accuracy with fewer labeled training instances. |
Copied to clipboard
| Challenge: | Despite extensive research showing the positive impact of uptake on student learning and achievement, there is little evidence that it is effective in teaching. |
| Approach: | They propose a framework for computationally measuring uptake by releasing a dataset of student-teacher exchanges extracted from US math classroom transcripts annotated for uptake . they formalize uptake as pointwise Jensen-Shannon Divergence (pJSD) and conduct a linguistically-motivated comparison of different unsupervised measures. |
| Outcome: | The proposed framework outperforms baseline measures in identifying uptake phenomena like question answering and reformulation. |
Copied to clipboard
| Challenge: | Recent research attention in task-oriented dialogue systems focuses on end-to-end neural models. |
| Approach: | They present a dataset that combines annotated corpora from four domains to provide a unified ontology and annotation schema for task-oriented dialogues. |
| Outcome: | The proposed dataset improves language, information content and performance in dialogues with two recent models. |
Copied to clipboard
| Challenge: | Existing studies on pre-trained language models assume they encode metaphorical knowledge useful for NLP systems. |
| Approach: | They propose to probing metaphoricity information in PLMs and measure their generalization . they find that contextual representations in PMLs encode metaphorical knowledge . |
| Outcome: | The proposed model can encode metaphorical knowledge across languages and datasets . the model can be used to train and test NLP systems . |
Copied to clipboard
| Challenge: | Dementia is one of the most pressing healthcare concerns as median age rises . a conversational agent capable of conducting cognitive health screening interviews could be an inexpensive, flexible, low-stress alternative . |
| Approach: | They propose an annotation schema for assigning dialogue act labels to utterances in patient-interviewer conversations collected as part of a clinically-validated cognitive health screening task. |
| Outcome: | The proposed system is characterized by high inter-annotator agreement and is able to perform clinically-validated cognitive health screening tasks. |
Copied to clipboard
| Challenge: | Discussions on an appropriate annotation scheme for large and complex information are ongoing . multi-layer system allows a comprehensive description of relations between morphological properties, syntactic function and expressed meaning. |
| Approach: | They propose a multi-layer annotation scheme for the Prague Dependency Treebank . they propose morphological properties, syntactic function and expressed meaning as multi-layered systems . |
| Outcome: | The proposed scheme is sound and serves well for complex annotations. |
Copied to clipboard
| Challenge: | Existing methods for stance detection are not diversified or inconsistent with the given target and label information. |
| Approach: | They propose to augment a text with a conditional masked word prediction task . they propose to replace a target mention with 'target-aware' sentences by replacing a reference word with . |
| Outcome: | The proposed method outperforms existing methods on 11 targets. |
Copied to clipboard
| Challenge: | RED is a machine learning-based resource developed for the automatic detection of emotions in Romanian texts. |
| Approach: | They propose an open-source extension of RED by adding trust and surprise . they propose two variants of ground truth suitable for multi-label classification and text regression . |
| Outcome: | The proposed model is based on two models with two transformer models, the Romanian BERT and the multilingual XLM-Roberta model, in categorical and regression settings. |
Copied to clipboard
| Challenge: | Existing methods to learn correspondence between visual segments and texts require temporal coordinates for training, which leads to high costs of annotation. |
| Approach: | They propose weakly supervised language localization networks to detect events in untrimmed videos . they train with only video-sentence pairs without accessing to temporal locations of events . |
| Outcome: | Experiments on ActivityNet Captions and DiDeMo show that WSLLN performs state-of-the-art. |
Copied to clipboard
| Challenge: | Using retrieval models and LLMs achieves a 73% approval rate for suggested links, more than doubling the acceptance of strong retrievers alone. |
| Approach: | They propose a domain-agnostic framework for bootstrapping sentence-level cross-document links from scratch and apply it to large-scale human-in-the-loop annotation of natural text pairs. |
| Outcome: | The proposed framework generates semi-synthetic datasets and uses them to benchmark and shortlist the best-performing methods and applies them in large-scale human-in-the-loop annotation of natural text pairs. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area. |
| Approach: | They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach. |
| Outcome: | The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues. |
Copied to clipboard
| Challenge: | Knowledge graphs are a graph of information organized as entities, relations, and entities. |
| Approach: | They propose a method to calibrate a scoring model over (entity, relation, entity)-tuples . they use an annotated set of tuple truncated by Logistic Regression or Gaussian Process classifiers . |
| Outcome: | The proposed method finds good per-relation thresholds efficiently based on a limited set of annotated tuples. |
Copied to clipboard
| Challenge: | Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. |
| Approach: | They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets. |
| Outcome: | The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others. |
Copied to clipboard
| Challenge: | Existing methods for ED in IE and NLP focus on feature-based models to feature-driven models. |
| Approach: | They propose to use a multilingual dataset to annotate events for 8 different languages . they demonstrate the challenges and transferability of ED across languages in MINION . |
| Outcome: | a new dataset that consistently annotates events for 8 different languages is released . the new dataset will promote future research on multilingual ED . |
Copied to clipboard
| Challenge: | Existing methods for scientific fact checking require domain expertise and time consuming. |
| Approach: | They propose a new supervised method for generating claims from scientific sentences and a novel method for negating claims. |
| Outcome: | The proposed method improves on existing methods on biomedical claims and negations. |
Copied to clipboard
| Challenge: | Existing tools for hate speech detection and sentiment analysis cannot detect veiled offensiveness of microaggressions . linguistic subtlety of micro-aggressives has made it difficult to analyze their exact nature . |
| Approach: | They propose a typology of microaggressions based on a subset of data . they propose an objective criterion for annotation and an active-learning procedure . |
| Outcome: | The proposed typology of microaggressions is based on a subset of social media data. |
Copied to clipboard
| Challenge: | Semantic similarity is a measure of the level of semantic overlap between texts of different lengths. |
| Approach: | They present a cross-level semantic similarity (CLSS) dataset in Serbian and compare it to its English counterpart, SemEval CLSS. They also use pre-trained language models to fine-tune the dataset. |
| Outcome: | The proposed dataset is compared to its preexisting counterpart in English, SemEval CLSS. The results are presented and state-of-the-art pre-trained language models are evaluated on the CLSS task in Serbian. |
Copied to clipboard
| Challenge: | Entrainment is the phenomenon of conversational partners adapting to one another to become more similar. |
| Approach: | They propose to use dyadic conversations to identify speaker traits and conversation contexts that cause variations in entrainment behavior. |
| Outcome: | The proposed corpus identifies speaker traits and conversation contexts that cause variations in entrainment behavior. |
Copied to clipboard
| Challenge: | Existing efforts to automate wet lab workflows are focusing on graph-prediction models that capture both concrete, exact quantities ("30 minutes") and vague instructions ("swirl") |
| Approach: | They manually annotate PEGs in a corpus of complex lab protocols with a novel interactive textual simulator that keeps track of entity traits and semantic constraints during annotation. |
| Outcome: | The proposed graph-prediction models are good at entity identification and local relation extraction while addressing challenges such as cross-sentence relations and long-range coreference. |
Copied to clipboard
| Challenge: | Prior work focused on using sentiment lexicons or leveraging large language models for annotation . lexiconics are often unavailable for historical texts due to limited linguistic resources . |
| Approach: | They propose a role-guided annotation strategy that prompts LLMs to simulate historical perspectives when labeling sentiment. |
| Outcome: | The proposed method outperforms state-of-the-art baselines across historical literature datasets. |
Copied to clipboard
| Challenge: | Active learning (AL) techniques reduce labeling costs for training neural machine translation models by selecting smaller representative subsets from unlabeled data for annotation. |
| Approach: | They propose an AL strategy that combines uncertainty and diversity for sentence selection. |
| Outcome: | The proposed method prioritizes diverse instances having high model uncertainty for annotation in early iterations. |
Copied to clipboard
| Challenge: | standardized tests are used to assess and screen developmental language impairments but require manual laborious transcription, annotation and calculation. |
| Approach: | They propose to use the correct sentence and the sentence produced by patients to evaluate the level of verbal production and return a score. |
| Outcome: | The proposed system evaluates the level of the verbal production and returns a score. |
Copied to clipboard
| Challenge: | End-to-end systems rely on dialogue state tracking and annotations to fulfill user requests . modularized systems require multiple steps, including a direct interaction with the KB . |
| Approach: | They propose a method to embed the KB directly into the model parameters . they evaluate five task-oriented dialogue datasets with small, medium, and large KBs . |
| Outcome: | The proposed model can embed the KB directly into the model parameters without any DST or template responses, nor the kb as input. |
Copied to clipboard
| Challenge: | Existing methods to generate annotated corpora for coreference are expensive and limited. |
| Approach: | They propose a model of annotation for aggregating crowdsourced anaphoric annotations. |
| Outcome: | The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation. |
Copied to clipboard
| Challenge: | Semantic textual similarity (STS) is a task of assigning a numerical score to short texts based on the level of semantic equivalence between them. |
| Approach: | They propose to annotate Serbian STS dataset with fine-grained similarity scores . they propose a supervised bag-of-words model that combines part-of speech weighting with term frequency weighting . |
| Outcome: | The proposed model outperforms existing models on the Serbian STS News Corpus . the proposed model is based on a new morphologically rich language . |
Copied to clipboard
| Challenge: | Existing dependency parsing models for Arabic use complementary annotations, CATiB and UD treebanks, and partially created trees for one annotation are also available to the other as features for the score function. |
| Approach: | They propose to use Arabic dependency annotations to parse projective dependency trees using CATiB and UD treebanks. |
| Outcome: | The proposed model gives 9.9% error reduction on CATiB and 6.1% on UD compared to a strong baseline and ablation tests show that the main contribution is given by sharing tree representation between tasks, and not simply sharing biLSTM layers as is often performed in NLP multitask systems. |
Copied to clipboard
| Challenge: | Prior research has shown the need to consider community language norms when studying taboo text classification and annotations. |
| Approach: | They propose to use special classifiers tuned for each community's language to study bias in taboo classification and annotation where a community perspective is front and center. |
| Outcome: | The proposed method shows that biases are strongest against African Americans and South Asians . a community perspective is front and center in the proposed method . |
Copied to clipboard
| Challenge: | Past researches have shown the superiority of statistical/ML approaches over the rule based approaches. |
| Approach: | They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning. |
| Outcome: | The proposed corpus contains 5,160 clinical documents from forty different clinical specialties. |
Copied to clipboard
| Challenge: | Existing data pruning methods for active learning are expensive and time-consuming. |
| Approach: | They propose a plug-and-play data pruning strategy that leverages language models to prune the unlabeled pool. |
| Outcome: | The proposed pruning strategy outperforms existing pruning methods on translation, sentiment analysis, topic classification, and summarization tasks on diverse datasets. |
Copied to clipboard
| Challenge: | Scholars in interdisciplinary fields like the Digital Humanities are increasingly interested in semantic annotation of specialized corpora. |
| Approach: | They propose an active learning solution for named entity recognition that maximizes a custom model’s improvement per additional unit of manual annotation. |
| Outcome: | The proposed model reduces required annotation by 20-60% and outperforms a competitive active learning baseline. |
Copied to clipboard
| Challenge: | The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody . |
| Approach: | They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language . |
| Outcome: | The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies. |
Copied to clipboard
| Challenge: | Recent code large language models have demonstrated impressive performance on code-related tasks. |
| Approach: | They propose a paradigm that learns from expert battles to address these limitations . they create an arena where leading LLMs challenge each other with evaluations . |
| Outcome: | The proposed model improves on existing models by leveraging expert battles . it achieves state-of-the-art performance even without relying on proprietary models . |
Copied to clipboard
| Challenge: | Spectral sampling strategies that minimize the number of annotations required to train a model are proposed. |
| Approach: | They propose a method that maximizes the amount of information useful for the learning algorithm by minimizing redundancy of samples in the selection. |
| Outcome: | The proposed method maximizes the amount of information useful for the learning algorithm or minimizes redundancy of samples in the selection. |
Copied to clipboard
| Challenge: | Ellipsis is an important challenge for natural language processing systems, says a new paper . previous work on ellipsis focused on news data, but sluicing presents a challenge for dialogue systems . |
| Approach: | They describe a corpus of 4100 sluice occurrences from the NYTimes Gigaword corpus . they build a classifier model to automatically classify slujce . |
| Outcome: | The proposed corpus contains 4100 sluice occurrences, with an accuracy of 67% . the work will support empirical research into slujcing in dialogue systems . |
Copied to clipboard
| Challenge: | Existing studies on the induced emotional states of the driver in a car have not been published. |
| Approach: | They used three sensor systems to collect emotional multimodal data while driving . they defined neutral, positive, frustrated and anxious states of the driver . |
| Outcome: | The collected data were analyzed using a Wizard-of-Oz technique . the participants were asked to fill out questionnaires and annotate the data . |
Copied to clipboard
| Challenge: | Existing methods for training effective PRMs focus on the first incorrect step and all preceding steps, assuming that all subsequent steps are incorrect. |
| Approach: | They propose a data annotation method specifically designed to score the long CoT reasoning process by using an LLM-based judger for annotation. |
| Outcome: | The proposed method improves PRMs' ability to identify effective self-correction behaviors and reasoning based on erroneous steps. |
Copied to clipboard
| Challenge: | MMLU is widely adopted but its ground truth errors obscure the true capabilities of LLMs. |
| Approach: | They propose a framework for identifying dataset errors using a novel error annotation protocol and a subset of 5,700 manually re-annotated questions. |
| Outcome: | The proposed framework is based on 5,700 re-annotated questions from the MMLU benchmark. |
Copied to clipboard
| Challenge: | Relation Extraction (RE) evaluation is limited to in-domain setups . despite the drought of research on cross-domain RE, its practical importance remains . |
| Approach: | They propose a cross-domain benchmark for relation extraction which includes multi-label annotations and meta-data to include explanations and flags of difficult instances. |
| Outcome: | The proposed model includes explanations and flags of difficult instances. |
Copied to clipboard
| Challenge: | Existing active learning models for text spam detection tasks are based on pool-based active learning, but the annotating process is laborious and time consuming for humans. |
| Approach: | They propose a semi-supervised active learning model to address spam imbalances . they propose masked attention learning approach and character variation graph-enhanced augmentation procedure . |
| Outcome: | The proposed model can improve the performance of existing models for Chinese spam detection task. |
Copied to clipboard
| Challenge: | Existing approaches to cognate detection use orthographic, phonetic and semantic similarity based features sets. |
| Approach: | They propose a method for enriching feature sets with cognitive features extracted from gaze behaviour data from human readers’ gaze behaviour. |
| Outcome: | The proposed method improves cognate detection performance by 10% and 12% over existing methods. |
Copied to clipboard
| Challenge: | *Warmth* (W) and *Competence (C) are central dimensions along which people evaluate individuals and social groups. |
| Approach: | They propose a first sentence-level dataset annotated for warmth and competence . they analyze sentences that express attitudes and opinions about individuals or social groups . |
| Outcome: | The first sentence-level dataset annotated for warmth and competence is presented in this paper. |
Copied to clipboard
| Challenge: | Existing methods for temporal relations annotation and management are not widely used because they are too complex from the computational perspective. |
| Approach: | They propose a system for the annotation and management of temporal relations that combines the richness and expressiveness of Freksa’s approach with the simplicity of Allen’s notation. |
| Outcome: | The proposed system achieves more agreeable representations of temporal relations without increasing the complexity of the labeling process. |
Copied to clipboard
| Challenge: | a corpus of 16th century letters from and to the Zurich reformer Heinrich Bullinger has been preserved . a recent study investigated code-switching in these 8600 letters . |
| Approach: | They investigate the automatic detection of code-switching in a 16th century letter exchange . they use a popular language identifier to bootstrap a word-based language classifier . |
| Outcome: | The proposed language classifier bootstraps with a popular identifier on a small training corpus of 150 sentences per language. |
Copied to clipboard
| Challenge: | Annotators using pre-annotation are less efficient at producing high quality annotations. |
| Approach: | They propose to use an automatic pre-annotation for a task to judge annotation quality . they also evaluate the effect of automatic linguistically-based checks on the same data . |
| Outcome: | The proposed method improves the quality of annotated sentences without reducing quality. |
Copied to clipboard
| Challenge: | End-to-end deep learning methods that focus on user satisfaction are challenging due to the required annotation costs and turnaround times. |
| Approach: | They propose a self-supervised contrastive learning approach that leverages the pool of unlabeled data to learn user-agent interactions. |
| Outcome: | The proposed approach reduces the required number of annotations while improving generalization on unseen skills. |
Copied to clipboard
| Challenge: | Xu et al., 2019; Lewis e t al, 2019) show that Bayesian summarization methods can generate high quality summaries but suffer from a couple of issues when inputs lie far from the training data distribution. |
| Approach: | They propose to extend state-of-the-art summarization models with Monte Carlo dropout and perform multiple stochastic forward passes to approximate Bayesian inference. |
| Outcome: | The proposed method outperforms deterministic summarization models on multiple benchmark datasets. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) methods require a substantial quantity of high-quality annotation for training models. |
| Approach: | They propose a method to reduce the number of incorrect pseudo labels in self-training . they propose 'uncertainty-aware teacher learning' and 'student-student collaboration' |
| Outcome: | The proposed method is superior to state-of-the-art DS-NER denoising methods. |
Copied to clipboard
| Challenge: | Prompt-based use of Large Language Models is becoming popular . specialized domains such as entity extraction are expensive to annotate . |
| Approach: | They propose to use a prompt set-up to provide training examples along with the inference request. |
| Outcome: | The proposed methods improve on a fully supervised transformer-based baseline. |
Copied to clipboard
| Challenge: | Social media are integrated with our daily life and are used to circulate information. |
| Approach: | They develop and publicly release the first largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign. |
| Outcome: | The proposed dataset is the largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign, covering many countries in the Arab region. |
Copied to clipboard
| Challenge: | Existing deep learning approaches require huge amounts of data to be trained properly. |
| Approach: | They propose to use Persian as a model to choose the samples for annotation instead of labeling the whole dataset. |
| Outcome: | The proposed models achieve the baseline performance with a significantly lower amount of labeled data. |
Copied to clipboard
| Challenge: | Using word-level linguistic annotations in under-resourced neural machine translation is challenging for many languages. |
| Approach: | They propose to use word-level linguistic annotations to label source-language (SL) or target-language words to improve translation performance. |
| Outcome: | The proposed language annotations outperform part of speech and morphological description tags in the target language, while the morpho-syntactic description tags improve the grammaticality of the output. |
Copied to clipboard
| Challenge: | Successful conversations often rest on common understanding, says a researcher . despite recent advances in dialog systems, there is a noticeable deficit in their grounding capabilities . |
| Approach: | They propose to use a framework to build conversational grounding in dialogs . they propose to analyze two dialog corpora using grounding acts and grounding units . |
| Outcome: | The proposed model shows that language models are not enough to ground dialogs with machines . the proposed model can be used to test the performance of existing Language Models . |
Copied to clipboard
| Challenge: | Using the AnnCor CHILDES Treebank, we assign adult grammar syntactic structures to children's utterances. |
| Approach: | They propose a partially manually verified treebank for Dutch CHILDES corpora . they argue that human annotation and automatic checks on this annotation must go hand in hand . |
| Outcome: | The AnnCor CHILDES Treebank is the first partially manually verified treebank for Dutch CHILdes corpora. |
Copied to clipboard
| Challenge: | a dataset of 16 TV and movie series is filled with challenging multi-party dialogues. |
| Approach: | They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues. |
| Outcome: | The proposed dataset is a step towards better multi-party dialogue structuring and understanding. |
Copied to clipboard
| Challenge: | 7.6% of the words in the original OCR text contain an error; fully manual correction would take thousands of hours due to the size of the corpus. |
| Approach: | They propose a post-processing system to efficiently correct OCR errors in a 2.7 million word Faroese corpus. |
| Outcome: | The proposed method reduces the word error rate to 1.3% with around 65 hours of human annotator work. |
Copied to clipboard
| Challenge: | Existing methods for active learning rely on model uncertainty or disagreement to pick unlabeled data, leading to over-confidence in superficial patterns and lack of exploration. |
| Approach: | They propose to use a bi-directional encoder and a uni-directional decoder to generate and score an explanation for low-resource text classification. |
| Outcome: | The proposed model improves on 9 strong baselines on six datasets and can generate explanations for its predictions. |
Copied to clipboard
| Challenge: | Historical dictionaries of the pre-digital period are important resources for the study of older languages. |
| Approach: | They propose to use printed dictionaries to create a more easily accessible and more sustainable lexical database by automating the conversion process. |
| Outcome: | The ‘Altfranzösisches Wörterbuch’, an Old French dictionary published from 1925 onwards, shows how the printed dictionaries can be turned into a more easily accessible and more sustainable lexical database. |
Copied to clipboard
| Challenge: | judicial opinions use language to comment on or draw attention to other language . a recent case involving a federal anti-discrimination law requires that justices determine the meaning of just one word or phrase in a specific context. |
| Approach: | They identify 9 categories prominent in metalinguistic discussions, including key terms, definitions, and different kinds of sources. |
| Outcome: | The results show that the annotated concepts are well-defined and frequent, and that they differ between majority, concurring, and dissenting opinions. |
Copied to clipboard
| Challenge: | Existing approaches to active learning maximize the system performance by sampling unlabeled instances for annotation that yield the most efficient training. |
| Approach: | They propose an active learning approach that integrates active learning with an end-user application to optimize the user's training and receiving useful instances. |
| Outcome: | The proposed approach satisfies both objectives when alternative methods lead to many unsuitable exercises for end users. |
Copied to clipboard
| Challenge: | Medical professionals often query over clinical notes to find information that can support their decision making. |
| Approach: | They propose to use expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering based on clinical notes. |
| Outcome: | The proposed system can answer clinical questions without using domain knowledge. |
Copied to clipboard
| Challenge: | Existing work on semantic role labeling treats symbolic labels as symbolic . labeled data is costly and often lacking in many tasks, domains, and languages. |
| Approach: | They propose to retrieve and leverage semantic role labels from annotation guidelines . argument classification is at the core of Semantic Role Labeling . |
| Outcome: | The proposed model achieves state-of-the-art on a CoNLL09 dataset injected with label definitions given the predicate senses. |
Copied to clipboard
| Challenge: | Social event detection relies on labeled data, but annotation is costly and labor-intensive. |
| Approach: | They propose a plug-and-play dual augmentation framework that combines explicit text-based and implicit feature-space augmentation to enhance data diversity and model robustness. |
| Outcome: | The proposed framework outperforms the best baseline model by 17.67% on the Twitter2012 dataset and 15.57% on the twitter2018 dataset in terms of the average F1 score. |
Copied to clipboard
| Challenge: | Recent work has sought to reduce the annotation costs through the use of active learning and data sampling. |
| Approach: | They propose to estimate the training sample size needed to achieve a targeted model performance based on small amount of training samples. |
| Outcome: | The proposed approach predicts model performance within a small margin of mean absolute error (0.9%) with only 10% data. |
Copied to clipboard
| Challenge: | An important factor that can affect the Inter Annotator Agreement (IAA) is the presence of annotator bias. |
| Approach: | They propose a new interpretation and application of the Item Response Theory to detect annotators' bias and characterise annotation disagreement. |
| Outcome: | The proposed method can be used to spot outliers, improve annotation guidelines and provide a better picture of the annotation reliability. |
Copied to clipboard
| Challenge: | Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources. |
| Approach: | a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation. |
| Outcome: | 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental step in scientific literature analysis to build AI-driven systems for molecular discovery, synthetic strategy designing, and manufacturing. |
| Approach: | They propose an ontology-guided method for fine-grained named entity recognition (NER) it leverages the chemistry type ontologies to generate distant labels with flexible KB-matching . |
| Outcome: | The proposed method significantly outperforms the state-of-the-art methods with a .25 absolute F1 improvement. |
Copied to clipboard
| Challenge: | Behavioral coding is a procedure that requires human intervention to be performed manually. |
| Approach: | They propose to use a publicly available conversation-based dataset to transfer knowledge to a low-resource behavioral coding task by meta-learning. |
| Outcome: | The proposed framework predicts target behaviors more accurately than baseline models. |
Copied to clipboard
| Challenge: | Question Answering datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. |
| Approach: | They propose a method for generating and validating QA datasets for low-resource languages . they use English data as context to generate synthetic multiple-choice (MC) question-answer pairs . |
| Outcome: | The proposed method maintains quality, reduces likelihood of factual errors, and circumvents costly annotation. |
Copied to clipboard
| Challenge: | Existing uncertainty sampling methods are time-consuming and can't be executed frequently. |
| Approach: | They propose adversarial uncertainty sampling in discrete space to find informative unlabeled text samples for annotation using adversarials. |
| Outcome: | The proposed approach outperforms baselines on effectiveness on five datasets. |
Copied to clipboard
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
Copied to clipboard
| Challenge: | a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion . |
| Approach: | They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity . |
| Outcome: | The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity. |
Copied to clipboard
| Challenge: | a lack of high quality conversational data is limiting progress in dialog systems . we present a dataset of 13,215 task-based dialogs . |
| Approach: | They propose a task-based dialog dataset which includes 13,215 task-related dialogs . they use a two-person, spoken "Wizard of Oz" approach and a "self-dialog" approach . |
| Outcome: | The taskmaster-1 dataset contains 13,215 task-based dialogs comprising six domains. |
Copied to clipboard
| Challenge: | Existing work on pre-trained generative models often fails to detect non-existent or incorrect content . Existing studies have attempted to detect hallucinations based on oracle references . |
| Approach: | They propose a token-level, reference-free hallucination detection task based on Wikipedia annotations to detect non-existent or incorrect content. |
| Outcome: | The proposed task is token-level, reference-free hallucination detection task and dataset . authors argue that the proposed task can be used in real-time to detect hallucines . |
Copied to clipboard
| Challenge: | Multiword Expressions (MWEs) are a pervasive phenomenon in all natural languages and challenge NLP applications because of their unpredictable morpho-syntactic and lexico--semantic behaviour. |
| Approach: | They propose to use linguistic resources to improve MWE translation and MWE generation by up to 5.09 BLEU points on MWE test sets. |
| Outcome: | The proposed annotation and data augmentation improve translation quality and increase performance by up to 5.09 BLEU points on MWE test sets. |
Copied to clipboard
| Challenge: | Existing approaches to training entity disambiguation models require a small labeling budget . a defense research analyst might need to map military equipment to a knowledge base describing emergent defense technologies. |
| Approach: | They propose a method that combines feature diversity with low rank correction . they use bilinear tensor models to train a model that uses a rich representation of context . |
| Outcome: | The proposed approach reduces the amount of labeled data necessary to achieve a given performance. |
Copied to clipboard
| Challenge: | The training of new tagger models for Serbian is motivated by the enhancement of the existing tagset with the grammatical category of a gender. |
| Approach: | They propose to use TreeTagger and spaCy taggers to train new Serbian tagger models and to align Serbian morphological dictionaries with the grammatical category of a gender. |
| Outcome: | The proposed models achieve 98% PoS-tagging precision per token, and the annotated dataset will be published. |
Copied to clipboard
| Challenge: | Question Answering models typically use retrieval and reasoning components to identify relevant information for reasoning. |
| Approach: | They propose a retrieval parameterization method that marginalizes unanswerable queries . they show that marginalization allows a model to mitigate false negatives in annotations . |
| Outcome: | The proposed model improves on two multi-document question answering datasets and shows that marginalization improves performance. |
Copied to clipboard
| Challenge: | SpiCE is a corpus of conversational Cantonese-English bilingual speech recorded in Vancouver, Canada . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantoneses . |
| Approach: | They describe the design, collection, orthographic transcription, and phonetic annotation of SpiCE . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . |
| Outcome: | The SpiCE corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . the corpus will promote bilingualism research for a typologically distinct pair of languages . |
Copied to clipboard
| Challenge: | Questions under Discussion (QUD) are emerging as a useful approach to spelling out the connection between information structure of sentences and nature of discourse. |
| Approach: | They propose a framework for QUD annotation based on explicit pragmatic principles . they propose generating all potentially relevant questions for a given sentence . |
| Outcome: | The proposed framework supports more reliable discourse structure annotation based on explicit questions . but the proposed approach is not robust enough for authentic data . |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) was designed to represent sentence meaning in English text, but recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture. |
| Approach: | They propose to annotate a multimodal (speech and gesture) AMR corpus in a task-based setting and capture coreference relationships across modalities. |
| Outcome: | The proposed corpus captures coreference relationships across modalities, enabling fine-grained analysis of how gesture and natural language interact. |
Copied to clipboard
| Challenge: | This paper reports on the activities of the Linguistic Data Consortium . |
| Approach: | This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report . |
| Outcome: | The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world . |
Copied to clipboard
| Challenge: | The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information. |
| Approach: | They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian. |
| Outcome: | The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway. |
Copied to clipboard
| Challenge: | Annotation quality is often framed as post-hoc cleanup of annotator-caused issues . authors argue that this narrative limits the scope of improving annotation . |
| Approach: | They propose to consider annotation as a procedural collaboration . they propose to capture the nuance and describe the full procedure to resolve issues . |
| Outcome: | The proposed study examines whether and why annotation quality is often framed as post-hoc cleanup of annotator-caused issues. |
Copied to clipboard
| Challenge: | Abstractive text summarization (ATS) requires laborious data annotation and time-consuming model training. |
| Approach: | They propose a novel active learning framework that asks large language models to rate difficulty of instances and then uses certainty gain maximization to select instances with a distribution that aligns well with the overall distribution. |
| Outcome: | The proposed framework improves stability, effectiveness, and efficiency of abstractive text summarization backbones. |
Copied to clipboard
| Challenge: | Semantic Storytelling is the goal of the future to generate stories based on extracted, processed, classified and annotated information from large content resources. |
| Approach: | They propose to create an automatic classifier for semantic relations between extracted text segments from different news articles. |
| Outcome: | The proposed method has high accuracy scores and is validated by a trained model. |
Copied to clipboard
| Challenge: | Conventional supervised methods cannot generalize to event types out of the pre-defined ontology. |
| Approach: | They propose to use two separate transformer models to model the definition semantics of an event type name into the same embedding space and then minimize their embeddable distance via contrastive learning. |
| Outcome: | The proposed model outperforms all previous zero-shot EE methods with fast inference speed due to the disjoint design. |
Copied to clipboard
| Challenge: | Abstractive text summarization has primarily focused on modeling news articles . lack of standardized datasets for summarizing online conversations is a major problem . |
| Approach: | They propose to crowdsource four new datasets for summarizing online conversations . they incorporate argument mining through graph construction to directly model issues, viewpoints, and assertions present in a conversation. |
| Outcome: | The proposed models are compared against widely-used conversation summarization datasets and show comparable or improved results. |
Copied to clipboard
| Challenge: | Existing approaches to model coherence are limited to small newswire corpora . evaluators need to be trained on lexical and document levels to perform evaluations . |
| Approach: | They propose four generic evaluation tasks that capture coherence-specific properties . they aim at capturing correct use of discourse connectives and lexical cohesion . |
| Outcome: | The proposed tasks capture coherence-specific properties, including correct use of discourse connectives, lexical cohesion, temporal consistency among events and participants in a story. |
Copied to clipboard
| Challenge: | Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties. |
| Approach: | They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
| Outcome: | The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental task in natural language processing (NLP). |
| Approach: | They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus. |
| Outcome: | The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed approach outperforms both masked language models and large language models. |
| Approach: | They propose a model-based scoring approach to quantify sentence quality . they propose 'loss function' that optimizes alignment between model predictions and sentence scores . |
| Outcome: | The proposed approach outperforms masked language models and large language models in the quantitative analysis of word substitutions. |
Copied to clipboard
| Challenge: | Existing methods to select unlabeled examples for annotation require a long time due to their complexity, hindering their practical viability. |
| Approach: | They propose a graph-based selection method to efficiently identify high-quality instances while minimizing computational overhead. |
| Outcome: | The proposed method significantly reduces selection time and improves performance on different tasks. |
Copied to clipboard
| Challenge: | Text simplification is a method for improving the accessibility of text by converting complex sentences into simple sentences. |
| Approach: | They propose to use Hindi knowledge annotators to capture the annotator’s language knowledge to build an automatic complex word classifier using a soft voting approach. |
| Outcome: | The proposed dataset shows that native and non-native annotators perceive complex words differently depending on their language knowledge. |
Copied to clipboard
| Challenge: | Existing semantic parsing frameworks rely on nontrivial human labor to generate canonical utterances. |
| Approach: | They propose a framework that uses an unsupervised paraphrase model to parse canonical utterances. |
| Outcome: | The proposed framework is effective and compatible with supervised training. |
Copied to clipboard
| Challenge: | Recent work casts GEC as a translation problem using encoder-decoder models to map bad (ungrammatical) sentences into good (grammatically) sentences. |
| Approach: | They propose to use a pretrained language model to define an LM-Critic that judges a sentence to be grammatical if the LM assigns it a higher probability than its local perturbations. |
| Outcome: | The proposed method outperforms existing methods in both the unsupervised and supervised setting. |
Copied to clipboard
| Challenge: | UD is a community project that maintains a standard scheme for the annotation of grammar in a cross-lingually consistent manner. |
| Approach: | They propose a Universal Dependencies treebank for Punjabi written in the Gurmukhi script and discuss corpus design and linguistic phenomena encountered in annotation. |
| Outcome: | The proposed treebank covers a variety of genres and has been annotated for POS tags, dependency relations, and graph-based Enhanced Dependencies. |
Copied to clipboard
| Challenge: | Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines. |
| Approach: | They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement . |
| Outcome: | The proposed dataset is the first of its kind for Norwegian and is available online. |
Copied to clipboard
| Challenge: | a non-standardized language such as Middle Low German has special requirements for annotating part of speech and morphology. |
| Approach: | They describe a tagset for annotating parts-of-speech and morphology in Middle Low German texts . they describe two special features of the tagse, and prove their usefulness . |
| Outcome: | The proposed tagset can be used to annotate parts-of-speech and morphology in Middle Low German texts. |
Copied to clipboard
| Challenge: | Existing studies on active learning methods focus on the out-of-distribution generalization of out- of-distortion samples. |
| Approach: | They propose a counterfactual active learning approach that empowers active learning with counterfact thinking to bridge the seen samples with unseen cases. |
| Outcome: | The proposed approach outperforms existing active learning methods on public datasets with comparable IID performance. |
Copied to clipboard
| Challenge: | Existing methods for Event Extraction are limited for non-English languages . lack of high-quality multilingual datasets has been the main hindrance . |
| Approach: | They propose a multilingual event extraction dataset that provides annotation for more than 50K event mentions in 8 typologically different languages. |
| Outcome: | The proposed dataset provides annotation for more than 50K event mentions in 8 languages . the proposed dataset will be publicly available to foster future research . |
Copied to clipboard
| Challenge: | Evaluating conversational information retrieval systems requires a significant amount of human labor for annotation. |
| Approach: | They propose to use human annotation to calibrate evaluation results to eliminate evaluation biases. |
| Outcome: | The proposed method consumes less than 1% of human labor and achieves a consistency rate of 95%-99% with human evaluation results. |
Copied to clipboard
| Challenge: | Question answering (QA) is an intuitive means to query text data. |
| Approach: | They propose a radiology question-answer-evidence-pair dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians. |
| Outcome: | The proposed dataset has 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians. |
Copied to clipboard
| Challenge: | Humor plays important role in human communication, which makes it important problem for natural language processing. |
| Approach: | They propose a novel annotation scheme to give scenarios of how humor arises in text . they report reasonable agreement between annotators and analyze the dataset . |
| Outcome: | The proposed scheme gives scenarios of how humor arises in text . it contains key words that trigger humor, character relationship, scene, and humor categories . |
Copied to clipboard
| Challenge: | HateCheck test cases are generic and have simplistic sentence structures that do not match the real-world data. |
| Approach: | They propose a framework to generate more diverse and realistic functional tests from scratch by instructing large language models. |
| Outcome: | The proposed framework generates more diverse and realistic functional tests from scratch by instructing large language models (LLMs). |
Copied to clipboard
| Challenge: | a framework/annotation schema is being developed to assess the offline harm potential of social media texts. |
| Approach: | They propose to annotate social media texts with their potential for triggering offline harm . they propose to use mood and modality as relevant categories to mark the speaker's intention, intended goal and their own evaluation of whether what they are saying is 'necessary' and 'possible' |
| Outcome: | The proposed framework can be used to annotate social media texts with their potential for triggering offline harm. |
Copied to clipboard
| Challenge: | FigAN data is a collection of isolated phrases with only literal and metaphorical meanings . FigSen corpus contains 1833 short fragments of texts containing at least one phrase from Figan data . |
| Approach: | They describe two resources of Polish data focused on literal and metaphorical meanings of adjective-noun phrases. |
| Outcome: | The proposed methods are compared with FigAN and FigSen corpus in Polish . the authors show that the methods are more accurate and more accurate than previous methods . |
Copied to clipboard
| Challenge: | Existing annotation resources for Discourse Dependency Parsing tasks are limited due to their complexity and annotation schema differences. |
| Approach: | They propose a code-based unified dependency parsing method that uses code to model dependency parses under different annotation schemas. |
| Outcome: | The proposed method improves on two Chinese DDP tasks. |
Copied to clipboard
| Challenge: | a paper argues that human label variation impacts all stages of the ML pipeline . human label variations are often considered noise due to disagreement, subjectivity in annotation or multiple plausible answers. |
| Approach: | They propose to reconcile different notions of human label variation and propose a repository of publicly-available datasets with un-aggregated labels. |
| Outcome: | The proposed approaches are compared with publicly available datasets with un-aggregated labels and identify gaps. |
Copied to clipboard
| Challenge: | Structured belief states are crucial for goal tracking and database query in task-oriented dialog systems. |
| Approach: | They propose a probabilistic dialog model where belief states are represented as discrete latent variables and jointly modeled with system responses given user inputs. |
| Outcome: | The proposed model outperforms supervised-only and semi-supervised baselines on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing tools for lexical normalization of social media data are designed with canonical texts in mind, and this makes it difficult to process data in multiple languages. |
| Approach: | They propose to create a lexical normalization dataset for Italian and analyze the inter-annotator agreement for this task. |
| Outcome: | The proposed model improves the parsing of social media data in Italian and shows that it can be used to translate non-standard social media content to canonical language. |
Copied to clipboard
| Challenge: | Social media data is a valuable data resource for natural language processing tasks. |
| Approach: | They propose to adapt input text to a more standard form, a task also referred to as normalization. |
| Outcome: | The proposed system scores 94.29 accuracy on the test data compared to 95.22 when trained on human-annotated data. |
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have shown promise in automating discourse annotation for conversations. |
| Approach: | They propose a pipeline that uses large language models to construct and perform annotations using speech functions and the Switchboard-DAMSL taxonomies. |
| Outcome: | The proposed pipeline outperforms existing tree annotation schemes and can match or surpass human annotations while significantly reducing time required for annotation. |
Copied to clipboard
| Challenge: | Personalized active learning techniques can be used to learn subjective NLP problems . to acquire training data, texts are often randomly assigned to users for annotation . |
| Approach: | They propose to apply an active learning paradigm to a personalized context to learn preferences . they validated their techniques on a Wiki discussion text labeled with aggression and toxicity . |
| Outcome: | The proposed methods outperform random selection and random selection by 30% on three datasets. |
Copied to clipboard
| Challenge: | a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible. |
| Approach: | They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets . |
| Outcome: | The proposed model performs better on similar datasets and worse on more non-offensive samples. |
Copied to clipboard
| Challenge: | Existing datasets for automatic speech recognition (ASR) in the endangered Kichwa language have been limited. |
| Approach: | They present Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. |
| Outcome: | The proposed dataset shows that it can be used to build an automatic speech recognition system for the endangered language with reliable quality despite its small size. |
Copied to clipboard
| Challenge: | Structured prediction is a fundamental problem in NLP, wherein the label space consists of complex structured outputs with groups of interdependent variables. |
| Approach: | They propose a partial annotation approach that selects only the most informative sub-structures for annotation and a method that incorporates the current model's automatic predictions as pseudo-labels for un-annotated sub-structurals. |
| Outcome: | The proposed approach reduces annotation cost over strong full annotation baselines under a fair comparison scheme that takes reading time into consideration. |
Copied to clipboard
| Challenge: | Existing supervised learning methods in natural language processing require large amounts of data. |
| Approach: | They propose an active learning loop that takes LLMs as annotators and puts them into an active loop to determine what to annotate efficiently. |
| Outcome: | The proposed model outperforms existing models with few-shot performance in two NLP tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)’s influential work on the New Yorker Cartoon Caption Contest. |
| Approach: | They propose to decompose humor understanding into three components and improve each by enhancing visual understanding through improved annotation and utilizing LLM-generated humor reasoning and explanations. |
| Outcome: | The proposed approach achieves 82.4% accuracy in caption ranking, significantly better than the previous 67% benchmark and matches the performance of world-renowned human experts in this domain. |
Copied to clipboard
| Challenge: | Contemplata is dedicated to the annotation of constituency trees. |
| Approach: | They propose to use Contemplata to build treebanks and treebank enrichment with relations between syntactic nodes. |
| Outcome: | The proposed solution is dedicated to the annotation of constituency trees and provides a balanced strategy between automatic parsing and manual revision. |
Copied to clipboard
| Challenge: | nudges are a choice architecture that alters people's behavior without forbidding any options or significantly changing their economic incentives. |
| Approach: | They describe a data collection methodology and emotion annotation of dyadic interactions between a human, a Pepper robot, . a Google Home smart-speaker, and other humans. |
| Outcome: | The collected 16-hour audio recordings show that humans change their opinions on more questions with a human than with nudges, even against mainstream ideas. |
Copied to clipboard
| Challenge: | a recent study examines the use of climate-related natural language processing (NLP) for climate-relevant tasks. |
| Approach: | They perform a reproducibility study on 8 tasks and 29 datasets, testing 6 models. |
| Outcome: | The proposed models are based on 8 tasks and 29 datasets. |
Copied to clipboard
| Challenge: | a new study examines the interaction between natural language use and gambling disorders. |
| Approach: | They build a new corpus of sentences that are searched and compared using top-k pooling to form the assessment pools of sentences. |
| Outcome: | The proposed model is based on a new corpus of sentences in spanish . |
Copied to clipboard
| Challenge: | Domain adaptation is underexplored in multimodal learning environments due to expensive data collection and annotation. |
| Approach: | They propose a bi-alignment scheme to perform drift-drift and anchor-driving matching with partially shifting anchors. |
| Outcome: | The proposed approach achieves superior performance compared with state-of-the-art approaches. |
Copied to clipboard
| Challenge: | Motivational interviewing (MI) is a counseling approach that aims to increase intrinsic motivation and commitment to change. |
| Approach: | They propose to annotate MI therapy sessions written in English from public sources . they explore the potential use of the dataset for training MI language models . |
| Outcome: | The proposed dataset includes 242 MI demonstration transcripts annotated with therapist behavioral codes and global scores and client language EAsy Rating (CLEAR) tags for client speech. |
Copied to clipboard
| Challenge: | Prior work has treated the style of a text as separable from the content. |
| Approach: | They use prompting to perform stylometry on a large number of texts to generate a synthetic stylometric dataset. |
| Outcome: | The proposed model trains human-interpretable representations on a large stylometric dataset and a linguistic model for style representation learning. |
Copied to clipboard
| Challenge: | e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text. |
| Approach: | They propose a timeline-based framework that achieves full coverage of all possible TLINKs. |
| Outcome: | The proposed framework achieves full coverage of all possible TLINKs in a text. |
Copied to clipboard
| Challenge: | Language models have boosted the performance of Question Answering, but data annotation is costly. |
| Approach: | They propose to use large language models to improve Question Answering performance . they argue that domain-agnostic knowledge from LMs is sufficient to create a well-curated dataset. |
| Outcome: | The proposed model outperforms state-of-the-art approaches on few-shot Question Answering. |
Copied to clipboard
| Challenge: | Temporal relation annotation in the clinical domain is crucial but challenging due to its workload and the medical expertise required. |
| Approach: | They propose an annotation method that integrates event start-points ordering and question-answering as the annotation format. |
| Outcome: | The proposed method achieves a 0.72 F1 score and enables collaboration among medical experts and non-experts. |
Copied to clipboard
| Challenge: | Human Label Variation (HLV) refers to legitimate disagreement in annotation . current preference-learning datasets routinely collapse multiple annotations into a single label . |
| Approach: | They propose to preserve human label variation as an embodiment of pluralism . they argue that disagreement in annotations should be treated as a selfzweck . |
| Outcome: | The proposed approach preserves pluralism and human pluralismos, the authors argue . they argue that disagreements in annotations should be treated as a selfzweck . |
Copied to clipboard
| Challenge: | Data annotation is a resourceintensive endeavor, necessitating human involvement and expertise. |
| Approach: | They propose to annotate instances to rebalance label distribution by judiciously selecting and limiting the data to be annotated. |
| Outcome: | The proposed method mitigates biases, improves model performance and reduces strategy-dependent disparities. |
Copied to clipboard
| Challenge: | Experimental results show that Active Learning methods ignore example groups whose prevalence may vary . supervised fine-tuning remains a critical component of model development, authors say . |
| Approach: | They propose an approach that uses interpolations to create anchors between examples . they propose to use the model to identify informative examples that counteract shortcuts . |
| Outcome: | The proposed model outperforms state-of-the-art active learning methods on six datasets . it prioritizes high-certainty instances that integrate representations from different example groups . |
Copied to clipboard
| Challenge: | Despite the growing demand for digital therapeutics for children with autism spectrum disorder, there is currently no speech corpus for Korean children with ASD. |
| Approach: | They propose to use Korean children with ASD to improve pronunciation and severity evaluation by transcribed speech and language evaluation sessions to assess their articulatory and linguistic characteristics. |
| Outcome: | The proposed corpus will be 300 children with ASD and 50 typically developing (TD) children. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive performance in many annotation tasks, including subjective tasks common in content moderation and text analysis in the social sciences. |
| Approach: | They propose to give crowdworkers LLM-generated annotation suggestions to "review" LLMs for subjective tasks can impact model performance and analysis downstream . |
| Outcome: | The proposed approach improves self-reported confidence in annotators and models . it also significantly improves model performance by analyzing human-approved datasets. |
Copied to clipboard
| Challenge: | Existing SLU resources are limited in high-resource languages such as English, Mandarin and French. |
| Approach: | They propose to use a Tunisian dialect dataset to build a semantic model of the system that is continuously annotated with dialogue acts and slots. |
| Outcome: | The proposed dataset is based on train-based and ASR-based models of train-driven conversations in Tunisian dialect. |
Copied to clipboard
| Challenge: | Multidialectal Arabic POS tagging is challenging due to the morphological richness and high variability among dialects. |
| Approach: | They propose an active learning approach for multidialectal Arabic POS tagging . they annotate approximately 15,000 tokens, reducing the annotation requirement by about 2,000 tokens . |
| Outcome: | The proposed approach achieves 97.6% accuracy on the Emirati corpus. |
Copied to clipboard
| Challenge: | Existing resources for AE extraction are limited due to complexity, variability, and ambiguity of clinical narratives. |
| Approach: | They present a manually annotated corpus for Adverse Event (AE) extraction from discharge summaries of elderly patients. |
| Outcome: | The proposed model performs well on coarse-grained extraction, but drops notably for rare events and complex attributes. |
Copied to clipboard
| Challenge: | polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks . |
| Approach: | They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. |
| Outcome: | The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context. |
Copied to clipboard
| Challenge: | well annotated corpora have been shown to have great value in linguistic and non-linguistic research . minority languages suffer from fewer available language resources than majority languages . a new method for evaluation of semantic annotation is being developed for Irish . |
| Approach: | They propose to build a tool-set for semantic annotation of Irish using semantic tags . they propose to use a lexicon built from a variety of sources to evaluate the tool . |
| Outcome: | a new method for evaluation of semantic annotation has been developed for Irish . the proposed method has 90% lexical coverage and almost 80% accuracy . |
Copied to clipboard
| Challenge: | Existing methods rely on output-level signals for sample identification, such as predictive entropy or semantic similarities with test-time data, which overlook models’ internal dynamics which could pinpoint specific knowledge gaps. |
| Approach: | They propose a Neuron-Aware Active Few-Shot Learning framework that shifts the selection paradigm from output-level proxies to models’ internal dynamics. |
| Outcome: | Experiments on three datasets show that NeuFS outperforms existing AFSL baselines. |
Copied to clipboard
| Challenge: | Almost 50% of depression patients face the risk of going into relapse. |
| Approach: | They propose to validate a social media dataset on depression relapse using cognitive theories of depression. |
| Outcome: | The first clinically validated social media dataset focused on depression relapse comprises 204 Reddit users annotated by mental health professionals. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) demonstrates remarkable zero-shot annotation tasks . but, they struggle with the specialized conventions of gold-standard benchmarks . |
| Approach: | They propose to reuse and refine annotation guidelines as an alignment mechanism . they propose to use iterative moderation framework to simulate early phases of annotation projects . |
| Outcome: | The proposed framework shows a good potential in effectively refining guidelines, but there is room for improvement. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) have achieved remarkable progress in recent years, yet their ability to perform left–right reasoning in mirror contexts remains underexplored. |
| Approach: | They propose a benchmark to evaluate MLLMs' ability to distinguish left from right from a subject-centered perspective. |
| Outcome: | The proposed benchmarks show that even the best performing models achieve only 65.40% accuracy, far below the 99.28% accuracy of humans. |
Copied to clipboard
| Challenge: | Large language models are transforming biomedical discovery by linking molecular patterns with knowledge encoded in text. |
| Approach: | They propose to map 58 foundation and agentic models developed for single-cell research into eight key analytical tasks. |
| Outcome: | The proposed models are applied to eight key analytical tasks including annotation, trajectory inference, perturbation modeling, and drug-response prediction. |
Copied to clipboard
| Challenge: | Evaluating software engineering capabilities is a core component of large language models (LLMs). |
| Approach: | They propose a benchmark to evaluate LLM-generated test suites that introduces mutated solutions that attempt to "fool" them. |
| Outcome: | The proposed test suites are based on 2,636 mutated variants derived from 800 original instances and include a multilingual subset spanning nine programming languages. |
Copied to clipboard
| Challenge: | Automating systematic reviews is expensive and time consuming, a study finds . automatic approaches are being explored but their performance has been poor . |
| Approach: | They propose to use reasoning-enhanced fine-tuning and DAPO reinforcement learning to automate systematic reviews. |
| Outcome: | The proposed methods significantly improve the performance of LLMs, the authors find . they find that reasoning-enhanced fine-tuning reduces time required for annotation by 80% . |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) extends the capabilities of large language models (LLMs) by providing access to external knowledge. |
| Approach: | They propose a framework that emulates human interactive reading through annotation and re-reading by integrating a thought bubble module that offloads internal cognition into external bookmark tokens, which are then annotated back into the context. |
| Outcome: | The proposed framework offloads internal cognition into external bookmark tokens, which are then annotated back into the context. |