Papers with crowdsourcing
Copied to clipboard
| Challenge: | Developing a theory of crowdsourcing for practical language problems remains an open challenge . |
| Approach: | This tutorial exposes NLP researchers to data collection crowdsourcing methods and principles through case studies. |
| Outcome: | This tutorial exposes NLP researchers to various data collection crowdsourcing methods and practices through case studies. |
Copied to clipboard
| Challenge: | Existing methods of data annotation are time-consuming and expensive . complexity of crowdsourcing increases when dealing with low-resource languages . |
| Approach: | They propose an autonomous method to gather unlabeled data and label them using large language models. |
| Outcome: | The proposed method is cost-efficient and applicable for low-resource language annotation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have undergone considerable development and can solve various natural language processing tasks. |
| Approach: | They aimed to achieve both attractiveness and factuality in a dialogue response by crowdsourcing a dataset and performing classification tasks on several models. |
| Outcome: | The proposed model with the highest classification accuracy could yield about 88% accurate classification results. |
Copied to clipboard
| Challenge: | a lack of gold datasets and knowledge about PAS analysis makes it difficult to create accurate PAS analyses. |
| Approach: | They construct a Japanese blog-QA dataset and a reading comprehension QA dataset using crowdsourcing. |
| Outcome: | The proposed method is most effective, pre-training model to acquire domain knowledge and fine-tuning model based on PAS-QA dataset. |
Copied to clipboard
| Challenge: | a substantial sector of the gig economy is the use of crowdworkers to annotate data for machine learning and analysis. |
| Approach: | They propose a narrative-sorting annotation task that sorts tweets chronologically by topic, emotional content, and length. |
| Outcome: | The proposed task enables readers to sort sequential, target-topical, and emotionally emotional tweets. |
Copied to clipboard
| Challenge: | Using a web-based coreference annotation suite, we demonstrate that non-expert annotators can be trained to perform and review coreference resolution tasks. |
| Approach: | They propose a web-based coreference annotation suite oriented for crowdsourcing that provides guided onboarding and a novel algorithm for a reviewing phase. |
| Outcome: | The proposed tool provides guided onboarding and a novel algorithm for a review phase. |
Copied to clipboard
| Challenge: | Using crowdsourcing, we can detect implicitly abusive comparisons . Abusive language is defined as hurtful, derogatory or obscene utterances made by one person to another . |
| Approach: | They propose to use crowdsourcing to generate a dataset for detecting implicitly abusive comparisons . they also use a range of linguistic features to better understand abusive comparison mechanisms . |
| Outcome: | The proposed dataset includes measures to obtain representative and unbiased comparisons. |
Copied to clipboard
| Challenge: | ComQA dataset captures question phenomena and the diverse ways in which they are formulated. |
| Approach: | They propose a large dataset of real user questions that captures question phenomena and the diverse ways in which they are formulated. |
| Outcome: | The proposed dataset can be a driver of future research on factoid question answering (QA). |
Copied to clipboard
| Challenge: | Indirect speech acts (ISAs) involve utterances whose literal meanings are not identical to their intended meanings. |
| Approach: | They propose a formal representation of ISA Schemas required for such testing, including a measure of the difficulty of a particular schema. |
| Outcome: | The proposed model minimizes the amount of expert authoring needed and maximizes realism. |
Copied to clipboard
| Challenge: | a common approach to quality estimation is to ask multiple reviewers to evaluate the same artifacts. |
| Approach: | They propose a probabilistic model for subjective classification tasks that incorporates the qualities of artifacts as well as the abilities and biases of creators and reviewers as latent variables to be jointly inferred. |
| Outcome: | The proposed model estimates the quality of speech more effectively than a vote aggregation, measured by correlation with a fine-grained classification by experts. |
Copied to clipboard
| Challenge: | Using current methods, the construction of multilingual FrameNets is expensive and complex. |
| Approach: | They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts . |
| Outcome: | The results are now available in Korean FrameNet 1.1. |
Copied to clipboard
| Challenge: | a growing number of documents are needed for multi-document summarization. |
| Approach: | They propose crowdsourcing to evaluate intrinsic and extrinsic quality of extractive text summaries . they conduct intensive comparative crowdsourcing and laboratory experiments . |
| Outcome: | The proposed crowdsourcing task evaluates intrinsic and extrinsic quality of extractive text summaries. |
Copied to clipboard
| Challenge: | Developing specialized dialogue systems for mental health support requires multi-turn conversation data . data privacy protection, time and cost involved in crowdsourcing are challenges . a new method for rewriting public single-turn dialogues into multi-turned ones is needed . |
| Approach: | They propose a single-turn to multi-turn inclusive language expansion technique that prompts ChatGPT to rewrite public single-turned dialogues into multi-turned ones. |
| Outcome: | The proposed method generates a large-scale, lifelike, and diverse dialogue dataset . it also develops SMILECHAT, a mental health chatbot . |
Copied to clipboard
| Challenge: | Clinical letters are written by doctors and typically contain complex medical language that is beyond the scope of the lay reader. |
| Approach: | They propose to augment existing neural text simplification software with a phrase table that links medical terminology to simpler vocabulary by mining SNOMED-CT data. |
| Outcome: | The proposed system is easier to understand than the existing system and the phrase table without it. |
Copied to clipboard
| Challenge: | aggregation has been a common strategy for dealing with unreliable data since the inception of crowdsourcing . however, many applications that rely on aggregate ratings only report the reliability of individual ratings, which is the incorrect unit of analysis. |
| Approach: | They propose to use k-rater reliability (kRR) to determine the correct data reliability for crowdsourced datasets. |
| Outcome: | The proposed method produces similar results on WordSim-353 datasets. |
Copied to clipboard
| Challenge: | a number of people engage in unsound argumentation techniques to prove a claim on online platforms . fallacies are weak arguments that seem convincing, but their evidence does not prove or disprove the conclusion . |
| Approach: | They propose to use user comments containing fallacy mentions as noisy labels to classify fallacies . they use the pragma-dialectical theory of argumentation to study the most common fallacias on Reddit . |
| Outcome: | The proposed dataset of fallacies on reddit shows that neural models perform better in conversational context. |
Copied to clipboard
| Challenge: | Disagreement in natural language annotation has been studied from a perspective of biases introduced by the annotators and the annotation frameworks. |
| Approach: | They propose to analyze task design bias in crowdsourced annotations where lay annotators are used to elicit interpretations. |
| Outcome: | The proposed methods can push annotators towards certain relations and some discourse relation senses can be better elicited with one or the other approach. |
Copied to clipboard
| Challenge: | Manual evaluation methods are perceived as insufficient due to the high cost of the Pyramid method and the required expertise. |
| Approach: | They propose a crowdsourced method that compares system summaries to references and uses crowdsourced scripts to analyze the results. |
| Outcome: | The proposed method shows higher correlation relative to the original Pyramid method. |
Copied to clipboard
| Challenge: | a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented . |
| Approach: | They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations . |
| Outcome: | The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages . |
Copied to clipboard
| Challenge: | Lexical-semantic resources like WordNet are a fundamental resource for many NLP and semantic applications. |
| Approach: | They propose a crowdsourcing workflow that consists of synset localization and validation . they use inter-rater agreement metrics to estimate the precision of the results . |
| Outcome: | The proposed method is cost-effective and provides a good trade-off between quality and speed of progress. |
Copied to clipboard
| Challenge: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing . |
| Approach: | They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners. |
| Outcome: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy . |
Copied to clipboard
| Challenge: | Existing models for named entity recognition (NER) are based on large-scale labeled datasets, which always obtain using crowdsourcing. |
| Approach: | They propose a CONfidence-based partial Label Learning method to integrate prior and posterior confidences for crowd-annotated named entity recognition models. |
| Outcome: | The proposed model improves on real-world and synthetic datasets compared with baselines. |
Copied to clipboard
| Challenge: | Split and Rephrase is a text simplification task that requires a strong evaluation benchmark and metric . despite its relatively new nature, the benchmark dataset contains easily exploitable syntactic cues . |
| Approach: | They propose to use crowdsourced datasets to evaluate split and rephrase models . they find that the widely used benchmark dataset universally contains exploitable syntactic cues . |
| Outcome: | The proposed model performs better than the state-of-the-art model, the authors say . they show that the datasets contain significantly more diverse syntax . |
Copied to clipboard
| Challenge: | generative large language models (LLMs) are replacing human workers for some tasks . crowdsourcing has several downsides: 1) the workforce is costly, 2) output quality is difficult to achieve, and 3) there are overheads related to the design and organization of the process. |
| Approach: | They investigate whether ChatGPT-created paraphrases are more diverse and robust . they use a crowdsourcing tool to collect training or validation examples . |
| Outcome: | The proposed models are more diverse and robust than the existing models. |
Copied to clipboard
| Challenge: | Abusive language is often defined as hurtful, derogatory or obscene utterances made by one person to another. |
| Approach: | They propose to use a dataset to detect abusive sentences in identity groups . they also report on classification experiments. |
| Outcome: | The proposed dataset includes 7 identity groups and includes classification experiments. |
Copied to clipboard
| Challenge: | DialCrowd 2.0 helps requesters obtain higher quality data from human intelligence tasks. |
| Approach: | They propose to use DialCrowd 2.0 to help requesters obtain higher quality data . they aim to improve the way requesters present tasks and facilitate effective communication with workers. |
| Outcome: | The proposed toolkit enables requesters to obtain higher quality data by presenting tasks more clearly and facilitating effective communication with workers. |
Copied to clipboard
| Challenge: | Annotated corpora are often assigned to internet workers whose judgments are reconciled by crowdsourcing models. |
| Approach: | They propose a framework for learning from rich prior knowledge to combine annotations with different structures. |
| Outcome: | The proposed model compares favorably with previous work and enables active sample selection to reduce annotation effort. |
Copied to clipboard
| Challenge: | In this paper, we evaluate use of different attribution methods for aiding identification of training data artifacts. |
| Approach: | They propose hybrid methods that combine saliency maps and instance attribution methods to aid in identifying training data artifacts. |
| Outcome: | The proposed methods can be used to efficiently uncover artifacts in training data when a challenging validation set is available. |
Copied to clipboard
| Challenge: | a vision-language benchmark for human activity planning is designed for humans . the task is easy for humans, but challenging for SOTA deep learning models . |
| Approach: | They propose a vision-language benchmark for human activity planning that extends Charades with intents and builds on a multi-choice question test set. |
| Outcome: | The proposed benchmark evaluates the ability of systems to anticipate and plan human actions in a multimodal visionlanguage setting. |
Copied to clipboard
| Challenge: | et al., 2021) show that instruction models can be trained on crowdsourced datasets with task instructions to achieve superior performance. |
| Approach: | They examine security concerns of emergent instruction tuning paradigm that models are trained on crowdsourced datasets with task instructions to achieve superior performance. |
| Outcome: | The proposed model can achieve 90% success rate across four commonly used datasets. |
Copied to clipboard
| Challenge: | Existing methods for data collection and annotation are costly and prevent launching new dialogue systems. |
| Approach: | They asked crowd workers to create persuasive dialogue systems using emotional expressions . they annotated emotional states and users' acceptance for system persuasion . |
| Outcome: | The proposed system has sufficient agreement even without training, the researchers found . the experiment showed that the collected data are comparable to real-world dialogue recording methods . |
Copied to clipboard
| Challenge: | a dataset of 4,000 pieces of art has annotations for emotions evoked in the observer . the dataset can help answer questions about what makes art evocative, how does art convey different emotions, what attributes of a painting make it well liked, and how much does the title impact the affectual response to art. |
| Approach: | They create a dataset of 4,000 western art pieces that has annotations for emotions . they use crowdsourcing to annotate the art for one or more of twenty emotion categories . fear, happiness, love, sadness were the dominant emotions that obtained consistent annotations . |
| Outcome: | The dataset shows that the most popular emotions are fear, happiness, love and sadness . the dataset can be used to develop systems that detect emotions evoked by art . |
Copied to clipboard
| Challenge: | a new study shows that literature enables engagement in a broader range of complex and subtle emotions. |
| Approach: | They propose to use multiple emotion labels to capture mixed emotions in poetry . they evaluate an annotation experiment with experts and crowdsourcing . |
| Outcome: | The proposed method shows that identifying aesthetic emotions is challenging in the German subset. |
Copied to clipboard
| Challenge: | Recent advances on abstractive summarization have led to fluent summaries, but factual errors in generated summary still severely limit their use in practice. |
| Approach: | They evaluate summaries produced by state-of-the-art models via crowdsourcing and show that factual errors occur frequently. |
| Outcome: | The proposed models can detect errors and reduce them by reranking alternative summaries. |
Copied to clipboard
| Challenge: | Recent development of spoken dialog systems aims at allowing a natural input style. |
| Approach: | They investigate how crowdsourced data can be assessed with respect to its naturalness and usefulness by using a word based language model to identify valid data. |
| Outcome: | The proposed methods show that valid data can be identified with the help of a word based language model. |
Copied to clipboard
| Challenge: | Existing methods to generate annotated dialogues require crowdsourcing, which is expensive and time-consuming. |
| Approach: | They propose a dialogue simulation method based on large language model in-context learning that generates new dialogues and annotations in a controllable way. |
| Outcome: | The proposed method can expand a small set of dialogue data with minimum or zero human involvement and parameter update. |
Copied to clipboard
| Challenge: | Existing datasets for conversation summarization are small due to the lack of large-scale datasets. |
| Approach: | They propose three approaches to generate summary grounded conversations, and evaluate the generated conversations using automatic measures and human judgements. |
| Outcome: | The proposed models can generate entire conversations with only a summary of a conversation as the input. |
Copied to clipboard
| Challenge: | Existing treebanks are limited in size, genre, and topic coverage, making manual annotation time-consuming and expensive. |
| Approach: | They propose a web-based interactive tool for editing dependency trees that uses machine learning to accelerate annotation. |
| Outcome: | CROWDTREE is a web-based interactive tool for editing dependency trees . it can train a parsing model during the annotation process and can even be compatible with Mechanical Turk. |
Copied to clipboard
| Challenge: | Existing benchmarks for end-to-end neural dialog systems lack a key component: natural variation. |
| Approach: | They propose new and more effective testbeds by introducing naturalistic variation by the user. |
| Outcome: | The proposed testbeds incorporate natural variation by the user. |
Copied to clipboard
| Challenge: | a new computational approach to exaggeration detection is needed for non-literal phenomena . a corpus of overstatements (or hyperboles) is used to detect exaggrements . |
| Approach: | They propose a computational approach to detect exaggerated sentences using crowdsourcing data . they build a corpus containing overstatements and then evaluate models trained on HYPO . |
| Outcome: | The proposed approach can detect exaggerated sentences using a crowdsourced dataset. |
Copied to clipboard
| Challenge: | Existing evaluation methods for large language models are labor-intensive and lack efficiency. |
| Approach: | They propose a framework dedicated to assessing long-text generation that includes in-depth human-curated meta-questions spanning various domains . they use a set of proxy-quests with pre-annotated answers to assess the content's quality by incorporating the generated texts as contextual background. |
| Outcome: | The proposed framework assesses the quality of long-text content by matching it with references through human evaluation or automated metrics. |
Copied to clipboard
| Challenge: | Compositional solutions for phrase sentiment are not able to handle idioms because their sentiment is not derived from the sentiment of the individual words. |
| Approach: | They propose a crowdsourcing approach for collecting sentiment annotations of idiomatic expressions using crowdsourcing. |
| Outcome: | The proposed approach is able to capture sentiment strength and ambiguity in idiomatic expressions using crowdsourcing. |
Copied to clipboard
| Challenge: | a new dataset of human-written summaries of bar charts is available in english . a chart summary is a textual description of a data point, which is often analytical . |
| Approach: | They propose a dataset of human-written summaries describing bar charts in english . a total of 47 charts are presented in the dataset, which includes 47 charts . |
| Outcome: | a new dataset of human-written summaries describing bar charts is presented in english . the dataset shows that human speakers often include such statements into chart summary . |
Copied to clipboard
| Challenge: | Experimental results show that noise correction in fine-grained entity typing improves quality of training samples. |
| Approach: | They propose a method that leverages multiple prediction results to correct noisy labels . they integrate prediction results and utilize a differentiated margin to identify inaccurate labels a . |
| Outcome: | The proposed model improves quality of training samples annotated using distant supervision, ChatGPT, and crowdsourcing. |
Copied to clipboard
| Challenge: | In this paper, we focus on verbal MWEs, whose accurate recognition is challenging because they could be discontinuous. |
| Approach: | They conduct large-scale annotations of VMWEs on the Wall Street Journal portion of Ontonotes . they first construct a VMwe dictionary based on the english-language Wiktionary . |
| Outcome: | The proposed resource annotates 7,833 VMWE instances belonging to various categories . the authors hope the results will help to develop models for MWE recognition and dependency parsing . |
Copied to clipboard
| Challenge: | Experimental results show that crowdsourced annotations are highly effective under supervised conditions. |
| Approach: | They propose an annotator-aware representation learning model that is inspired by domain adaptation methods which attempt to capture effective domain-alike features. |
| Outcome: | The proposed model is highly effective on a benchmark dataset and achieves state-of-the-art performance with only a very small scale of expert annotations. |
Copied to clipboard
| Challenge: | Existing tools for examining and fixing missing captions are lacking in mobile UIs. |
| Approach: | They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces. |
| Outcome: | The proposed task can generate captions from image and structural representations of UI elements. |
Copied to clipboard
| Challenge: | Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain. |
| Approach: | They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation . |
| Outcome: | The proposed model reduces the reliance on human validation without degrading the quality of the final data. |
Copied to clipboard
| Challenge: | a large number of large language models are being used to protect user privacy . sanitizing sensitive text using two common strategies is the answer . |
| Approach: | They propose sanitizing sensitive text using deleting expressions and abstracting them . they propose a tool for text rewriting that uses crowdsourcing and large language models . |
| Outcome: | The proposed approach protects privacy before sending sensitive data to large language models . it combines crowdsourcing and large language modeling to create a text rewrite tool . |
Copied to clipboard
| Challenge: | a recent study has reported that crowdsourcing cannot distinguish between machine-authored and human-authored text. |
| Approach: | They propose a framework called Scarecrow for scrutinizing machine text via crowd annotation . they use crowd annotation to identify redundancy, commonsense errors, and incoherence . |
| Outcome: | The proposed method quantifies gaps between human-authored and machine-generated text . it can detect redundancy, commonsense errors, and incoherence . |
Copied to clipboard
| Challenge: | Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader. |
| Approach: | They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically. |
| Outcome: | The proposed method can be used to develop and evaluate hierarchical summarization systems. |
Copied to clipboard
| Challenge: | Existing tasks require only a small set of attributes to track state changes in procedural text. |
| Approach: | They propose a task where given a procedural text as input, the task is to generate a set of state change tuples for each step. |
| Outcome: | The proposed task generates state change tuples from a set of pre-defined attributes for each step and predicts them from an open vocabulary. |
Copied to clipboard
| Challenge: | Using crowdsourcing to train neural machine translation models is expensive and expensive . professional outsourcing of bilingual data is expensive if the translations are of a lower quality . |
| Approach: | They analyze the impact of crowdsourcing on the quality of in-domain training data . they use translations of MOOCs from English to eleven languages to fine-tune machine translation models . |
| Outcome: | The proposed method improves on general-domain training data and with pre-existing in-domain corpora. |
Copied to clipboard
| Challenge: | Despite being large and generic, some languages such as Sinhala are left to underutilize the technology due to the lack of adequate resources. |
| Approach: | They propose to derive a corpus from a publicly available corpus for Sinhala speech recognition using crowdsourcing and web scraping techniques. |
| Outcome: | The proposed corpus reduces the Word-Error-Rate by 15.9%. |
Copied to clipboard
| Challenge: | The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic. |
| Approach: | They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available. |
| Outcome: | The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi. |
Copied to clipboard
| Challenge: | a dataset of 2,437 dialogues and 10,917 QA pairs is used to access domain-specific FAQ information. |
| Approach: | They present a dataset with 2,437 dialogues and 10,917 QA pairs for FAQs . they use the Wizard of Oz method with crowdsourcing to create dialogues using the original post and the original reply. |
| Outcome: | The proposed system can access domain-specific FAQ information without training data. |
Copied to clipboard
| Challenge: | a recent study shows that crowdsourcing is becoming mainstream to create bilingual dictionaries . the number of people who can speak multiple low-resource languages is limited and the average ability of workers is low. |
| Approach: | They propose a method to aggregate the answers of evaluation tasks by majority voting . they use hyper questions to evaluate the reliability of workers and task-allocation method to select high-quality workers . |
| Outcome: | The proposed method improves quality of bilingual dictionaries by integrating answers by majority voting. |
Copied to clipboard
| Challenge: | generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are paraphrased and then used to fine-tune downstream models. |
| Approach: | They propose to use taboo words, hints by previous outlier solutions, and chaining on previous outliest solutions to augment text datasets as part of instructions to LLMs augmenting text dataset. |
| Outcome: | The proposed methods increase diversity of generated texts, but performance is highest with hints. |
Copied to clipboard
| Challenge: | Manual annotation methods, such as crowdsourcing, are costly and require intricate task design skills. |
| Approach: | They propose to use open source LLMs to annotate parallel data for text detoxification . they generate a pseudo-parallel detoxification dataset using activation patching . |
| Outcome: | The proposed model performs comparable to the original dataset in automatic detoxification evaluation metrics and superior quality in manual evaluation and side-by-side comparisons. |
Copied to clipboard
| Challenge: | Konkani is a low-resource language spoken by 2.5 million speakers . idiomatic sense processing is challenging due to the nature of idioms . |
| Approach: | They propose to use crowdsourced idiomatic sentence identification to build a corpus for idioms in the Konkani language. |
| Outcome: | The proposed corpus consists of 6520 sentences written in the Konkani language. |
Copied to clipboard
| Challenge: | Southeast Asia is underrepresented in vision-language research . SEA-VL is an open-source initiative dedicated to developing culturally relevant datasets for SEA languages. |
| Approach: | They propose to use crowdsourced, automated image crawling and synthetic image generation to develop culturally relevant datasets for SEA languages. |
| Outcome: | The proposed datasets capture SEA cultural nuances and contexts better than existing datasets. |
Copied to clipboard
| Challenge: | Existing studies on the prevalence of mental disorders on the Web are limited to the English language. |
| Approach: | They propose to use user messages posted on Telegram groups to annotate the corpus for natural language processing and to conduct experiments on text classification and regression. |
| Outcome: | The proposed corpus contains over 1,300 subjects with more than 45,000 messages posted in different public Telegram groups. |
Copied to clipboard
| Challenge: | Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web. |
| Approach: | They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances. |
| Outcome: | The proposed classifier augments training data with automatically-generated GPT-3 completions. |
Copied to clipboard
| Challenge: | LLM-as-a-Judge uses LLMs to evaluate open-ended questions . however, the discrepancy between LLM generated evaluations and human evaluations remains a critical problem in this field . |
| Approach: | They propose a framework that orchestrates evaluations across multiple criteria using multiple LLMs. |
| Outcome: | The proposed framework achieves superior alignment with human evaluations compared to baselines. |
Copied to clipboard
| Challenge: | Recent studies have shown that multi-task instruction tuning after pre-training greatly improves the model’s robustness and transfer ability, which is crucial for building a high-quality dialog system. |
| Approach: | They propose to use Task-aware Automatic Prompt generation (TAP) to automatically generate high-quality prompts from 15 dialog-related tasks. |
| Outcome: | The proposed model is robust to input prompts and capable of various dialog-related tasks. |