Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track
Copied to clipboard
| Challenge: | Recent text-to-image models require multiple passes of prompt engineering by humans to produce satisfactory results for real-world applications. |
| Approach: | They propose a deep generative model to generate high-quality prompts from raw descriptions using visual feedback. |
| Outcome: | The proposed model produces high-quality prompts from simple raw descriptions . it can be integrated to a cloud-native AI platform to provide better image generation service in the cloud. |
Copied to clipboard
| Challenge: | Recent studies have shown that generative language models lack functional correctness, which is a critical aspect of regular expressions. |
| Approach: | They propose a method that takes functional correctness into account and transforms it into a differentiable gradient feedback using policy gradient techniques. |
| Outcome: | The proposed method has been used in a regulatory scenario to ensure that all online content is free from non-compliant elements, thereby significantly reducing the workload of relevant personnel. |
Copied to clipboard
| Challenge: | Large language models are often inefficient for real-world deployment due to expensive inference costs. |
| Approach: | They propose to use knowledge distillation to transfer the knowledge of the original model to a smaller, more efficient student model. |
| Outcome: | The proposed method is the best for multi-lingual and multilingual student architectures. |
Copied to clipboard
| Challenge: | Existing debt collection agents fail to tailor strategies to debtor personas, leading to ineffective collection. |
| Approach: | They present a commercial practice on debt collection agents that organizes debtor personas into a taxonomy and constructs a persona-aware conversation dataset. |
| Outcome: | The proposed agent increases recovery rate by 3.31% and collects additional 100K RMB after two months of testing. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) require massive amounts of computation and storage, such an approach incurs network and high execution cost. |
| Approach: | They propose a model gatekeeper to stop LLM calls that result in incorrect predictions . they show it can save 46.6% of COGS and improve user experience by not showing incorrect predictions. |
| Outcome: | The proposed model gatekeeper saves 46.6% of COGS and improves user experience . it also improves the suggestion rate of the proposed model by 73% . |
Copied to clipboard
| Challenge: | Pretrained transformer language models have been gaining popularity in the field of natural language processing . however, there is no study into the intersection of these two fields . |
| Approach: | They propose a method to extract knowledge from transformers to produce high-performing efficient attention models with low costs. |
| Outcome: | The proposed model compression method preserves up to 98.6% of original model performance across short-context tasks and up to 95.8% on long-concept Named Entity Recognition tasks while decreasing inference times by up to 57%. |
Copied to clipboard
| Challenge: | Recent research has focused on predicting crimes, predicting outcomes of judicial debates, and extracting information from legal documents. |
| Approach: | They propose to use a large-size Court Debate Dataset to analyze court debates . they invite experienced judges to design appropriate labels for data records . |
| Outcome: | The proposed dataset includes 30,481 court cases, totaling 1,144,425 utterances. |
Copied to clipboard
| Challenge: | Modern speech technologies have moved towards end-to-end models that constitutes black box systems that do not allow for explainability of the prediction or decisions. |
| Approach: | They propose a method for linguistic feature extraction that uses phonetic transcriptions and a forced alignment tool to extract phonetic features. |
| Outcome: | The proposed method is compatible with a forced-alignment tool, the Montreal Forced Aligner. |
Copied to clipboard
| Challenge: | Constrained retrieval is limited to entities in recent user history, which offers low coverage of future requests. |
| Approach: | They propose a personalized entity retrieval system that is robust to phonetic noise and ambiguity but is not limited to a customized index. |
| Outcome: | The proposed system corrects multiple error modes and shows 91% improvement over baseline on the entity retrieval task. |
Copied to clipboard
| Challenge: | In the digital age, large-scale online travel platforms face the challenge of extracting valuable insights from massive volumes of textual data. |
| Approach: | They propose a model that uses a bi-encoder transformer architecture to extract structured information from textual data. |
| Outcome: | The proposed model outperforms state-of-the-art models with 92.9% micro mAP and 75.8% macro mA score compared to baseline models . the proposed model can be used to find hotel facilities and hotel rooms based on positive reviews . |
Copied to clipboard
| Challenge: | e-commerce search engines use customer behavior signals to augment lexical matching and improve search relevance. |
| Approach: | They propose a method to identify duplicate and near-duplicate products across stores . they use Hierarchical Ranked Multi Similarity Loss to learn hierarchical metric space . |
| Outcome: | The proposed model outperforms baselines in terms of catalog coverage and precision of the mappings. |
Copied to clipboard
| Challenge: | Earlier studies have shown that domain-specific LMs are crucial for domain-specific applications. |
| Approach: | They propose a new BERT model for the cybersecurity domain, CTI-BERT . they find it significantly outperforms general-domain and security-domain models . |
| Outcome: | The proposed model outperforms general-domain and security-domain models for cybersecurity tasks. |
Copied to clipboard
| Challenge: | Existing methods for quantization of models are too complicated and can cause performance damage. |
| Approach: | They propose a self-adaptive mixed-precision (SAMP) toolkit to automatically control quantization rate by a mixed-presence architecture to balance model accuracy and efficiency. |
| Outcome: | The proposed toolkit has a higher speedup than PyTorch and FasterTransformer while ensuring the required accuracy. |
Copied to clipboard
| Challenge: | Existing SOTA techniques for semantic matching are mostly based on Siamese networks. |
| Approach: | They propose a novel knowledge distillation algorithm designed for real-time semantic matching . they train low latency accurate student models by leveraging soft labels from a teacher model . |
| Outcome: | The proposed algorithm outperforms teacher and SOTA knowledge distillation benchmarks on e-commerce datasets. |
Copied to clipboard
| Challenge: | a multilingual spelling correction model is needed to meet the tight latency requirements of multilingual NLP . a monolingual teacher model is trained for each language/locale, and individual models are distilled into a single student model . |
| Approach: | They propose a multilingual approach to spelling correction using multi-teacher distillation . they train a monolingual teacher model for each language and distill them into a single model . |
| Outcome: | The proposed model can meet the tight latency requirements of deployed services. |
Copied to clipboard
| Challenge: | scalability of attribute-value extraction (AVE) task is key for a large number of products . a question-answering (QA)-based approach is better for AVE, but requires a larger number of classes to be scalable. |
| Approach: | They propose a question-answering-based approach that additionally inputs the target attribute as a query to extract its values. |
| Outcome: | The proposed approach outperforms a classical approach on real-word e-commerce datasets in accuracy and speed. |
Copied to clipboard
| Challenge: | Existing table-to-text generation techniques that transform complex tabular data into comprehensible narratives are lacking in real-world applications. |
| Approach: | They investigate the table-to-text capabilities of different LLMs using four datasets within two real-world information seeking scenarios. |
| Outcome: | The proposed models can generate table-to-text data in two real-world information seeking scenarios and perform better than existing models. |
Copied to clipboard
| Challenge: | Annually, e-commerce platforms incur substantial financial losses due to trademark infringements. |
| Approach: | They propose a dataset to detect trademark infringement in merchant registrations . they use legal rules and contextual information from Alipay to gather contextual information with annotations from legal experts. |
| Outcome: | The proposed dataset is sourced from Alipay, one of the world’s largest e-commerce and digital payment platforms. |
Copied to clipboard
| Challenge: | Utilizing natural language processing in clinical conversations is effective to improve the efficiency of workflows for medical staff and patients. |
| Approach: | They propose a model for dialogue segmentation and topic categorization that integrates natural language processing techniques into a joint model. |
| Outcome: | The proposed model improves on follow-up calls for diabetes management and reduces computational complexity and cost. |
Copied to clipboard
| Challenge: | Recent work on learning from multiple tasks has shown that adding an extra fusion layer to implement knowledge composition is non-scalable for some applications. |
| Approach: | They propose a two-stage knowledge distillation algorithm to extract task specific knowledge by using local data to train a student adapter. |
| Outcome: | Experiments on frequently asked question retrieval in task-oriented dialog systems validate the efficiency of AdapterDistillation. |
Copied to clipboard
| Challenge: | a new study examines email marketing performance by considering email content and metadata. |
| Approach: | They propose a model that incorporates semantic and structural information from email data to generate latent exemplars for email response prediction. |
| Outcome: | The proposed model outperforms baseline models on two real-world email datasets . it provides interpretability through prototypes at different granularity levels while maintaining comparable performance to non-interpretable models. |
Copied to clipboard
| Challenge: | Recent work has proposed a dual encoder for product matching due to its high performance and computation efficiency. |
| Approach: | They propose retrieval-enhanced dual encoder training to improve product matching . they use public and real-world product matching datasets to train the dual encoded model . |
| Outcome: | The proposed approach improves on a public and real-world product matching datasets. |
Copied to clipboard
| Challenge: | Existing typography solutions lack adaptability, creativity, and computational efficiency. |
| Approach: | They propose a user-driven framework for artistic typography synthesis based on the Large Language Model (LLM) the LLM Engine interprets user inputs and generates actionable prompts for the other modules, transforming abstract concepts into tangible designs. |
| Outcome: | The proposed framework incorporates four key modules: the LLM Engine, SemTypo, StyTyPo, and TexTyPO. |
Copied to clipboard
| Challenge: | Existing methods to extract misspelling-correction pairs from Japanese query logs are not effective due to the unique input methods. |
| Approach: | They propose a romanization-aware edit distance that utilizes romanization lattices to efficiently consider all possible romanized forms of input strings. |
| Outcome: | Empirical results show lattice path edit distance outperforms standard edit distance in Japanese . latticae path editing distance outpersforms existing methods even with romanization . |
Copied to clipboard
| Challenge: | Experimental results on multilingual similarity search and bitext mining tasks show the effectiveness of our approach. |
| Approach: | They propose a multilingual sentence representation model that aligns different languages in a shared representation space. |
| Outcome: | The proposed model performs better than LASER3 on similarity searches and bitext mining tasks. |
Copied to clipboard
| Challenge: | Existing datasets focused on gender or racial biases are not designed for the gaming industry, a concern for models built for toxicity detection in videogames’ written chat. |
| Approach: | They propose to use reactivity analysis to highlight oversensitive terms using a language model developed by Ubisoft for toxicity detection on videogame’s written chat and Perspective API to generate a list of terms that trigger the models to varying degrees. |
| Outcome: | The proposed model can detect and amplify identity biases in annotated language models and is compared with a language model developed by Ubisoft for toxicity detection on videogames’ written chat and Perspective API. |
Copied to clipboard
| Challenge: | Existing methods for text-video retrieval focus on informative representations and delicate matching mechanisms, but real-world scenarios often involve brief, ambiguous queries and low-quality videos. |
| Approach: | They propose a novel method to learn informative embeddings for queries and videos . they use a watch-time-aware contrastive learning paradigm to capture dependencies . |
| Outcome: | The proposed method is effective in a real-world video-search service. |
Copied to clipboard
| Challenge: | churn occurs when retraining models yields different predictions despite using the same data and hyper-parameters. |
| Approach: | They propose a method that pairs semantic parses based on their “function call signature” and encourages similarity through an additional loss based upon Jensen-Shannon Divergence. |
| Outcome: | The proposed method improves in academic, noisy, and industry settings. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have gained popularity but lack specific domain knowledge in domain-specific tasks. |
| Approach: | They propose a model interaction paradigm that empowers LLM to achieve better performance on domain-specific tasks where it is not proficient. |
| Outcome: | The proposed approach outperforms the commonly used LLM with retrieval methods in domain-specific tasks. |
Copied to clipboard
| Challenge: | Existing approaches to extreme multi-label text classification face inherent challenges in terms of model, data, and evaluation. |
| Approach: | They propose a label ranking model as an alternative to the conventional SciBERT-based classification model and an active learning-based pipeline that addresses the data scarcity of new labels during the update of a classification system. |
| Outcome: | The proposed model enables efficient handling of large-scale labels and accommodates new labels. |
Copied to clipboard
| Challenge: | Existing relevance ranking methods focus on text modality, incapable of fully exploiting cross-modal cues present in video. |
| Approach: | They propose a QUery-Aware pre-training model with multi-modality that integrates video tag information as alignment targets and enhances ranking optimization method based on ordinal regression. |
| Outcome: | The proposed model significantly improves video search performance. |
Copied to clipboard
| Challenge: | Continual Federated Learning (CFL) combines decentralized learning with continuous learning . ubiquity of personal devices with a network connection offers rich source of data for learning problems . |
| Approach: | They propose to combine decentralized learning with a continuous learning approach . they propose to coordinate gradient-based replay sample selection across clients . |
| Outcome: | The proposed method shows gains early in the low replay size regime, when the budget for storing past data is small. |
Copied to clipboard
| Challenge: | a study examines how to build meeting summarization systems using large language models . closed-source models are generally better in terms of performance, but open-source ones are more advantageous for industrial use . |
| Approach: | They compare closed-source and open-source meeting summarization models for real-world use . they find that closed-sourced models are generally better in terms of performance . however, smaller open-sourced LLMs could still achieve comparable performance if they are open . |
| Outcome: | The proposed model is more efficient for industrial use than closed-source models due to privacy concerns and high cost. |
Copied to clipboard
| Challenge: | In tweets, people refer to the content it delivers, but also to the person behind it. |
| Approach: | They examine how creator context can be used to advance tweet understanding by recommending relevant tweets to news articles. |
| Outcome: | The proposed model can improve a news article's relevance by recommending relevant tweets to news articles. |
Copied to clipboard
| Challenge: | End-to-end (E2E) automatic speech recognition models struggle to recognize out-of-domain words such as proper nouns and domain-specific terms. |
| Approach: | They propose a domain adaptation technique that relies solely on textual data to adapt to out-of-domain words. |
| Outcome: | The proposed method outperforms the base model by up to 14% relative word error rate improvement on several out-of-domain, publicly available datasets. |
Copied to clipboard
| Challenge: | Large amount of companies' data is stored in relational databases . quick hypotheses validation is rarely, if ever, possible for majority of nontechnical business stakeholders. |
| Approach: | They propose a hybrid NLQ system for conversational DB querying that allows non-technical users to formulate data requests as natural language questions. |
| Outcome: | The proposed system is based on a hybrid NLQ (Natural Language Querying) system for conversational DB querying. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are rapidly becoming more and more popular, but dealing with the potential harms associated with their deployment in real-world scenarios is still an open research question. |
| Approach: | They propose an automated approach for automated generation of adversarial evaluation datasets to test the safety of LLM generations on new downstream applications. |
| Outcome: | AART generates evaluation datasets with high diversity of content characteristics critical for effective adversarial testing. |
Copied to clipboard
| Challenge: | Speakerly TM is a voice-based writing assistance system that works across the different stages of writing. |
| Approach: | They propose a voice-based writing assistance system that helps users with text composition across various use cases such as emails, instant messages, and notes. |
| Outcome: | The proposed system can be used for email, instant messages, and notes. |
Copied to clipboard
| Challenge: | Recent large language models such as ChatGPT and GPT-4 have shown exceptional capabilities of generalist models . however, their applicability and effectiveness in specific domains like finance needs a better understanding . |
| Approach: | They conduct empirical studies to compare the performance of ChatGPT and GPT-4 on financial text analytical problems using eight benchmark datasets from five categories of tasks. |
| Outcome: | The proposed models outperform the state-of-the-art models on a wide range of financial text analytical tasks. |
Copied to clipboard
| Challenge: | Existing QR systems that reformulate defective user queries are limited in English due to the scarcity of non-English QR labels. |
| Approach: | They propose a query reformulation method which reformulates defective user queries to improve non-English QR performance. |
| Outcome: | The proposed framework improves non-English QR performance by leveraging abundant reformulation resources in English. |
Copied to clipboard
| Challenge: | Contextual query rewriting (CQR) is a crucial component in Conversational AI agents, leveraging contextual information from previous user-agent conversations to improve comprehension of current user intent. |
| Approach: | They propose a framework to enhance the CQR model's capability in generating user preference-aligned rewrites. |
| Outcome: | The proposed framework improves the CQR model's ability to generate user preference-aligned rewrites. |
Copied to clipboard
| Challenge: | Inverse Text Normalisation (ITN) is a textrewriting task that converts verbalized text to written form. |
| Approach: | They propose to use a seq2seq model, a non-autoregressive text editor and a sequence tagger + rules combination to fine-tune three pre-trained neural models. |
| Outcome: | The proposed model improves with bootstrapping and data augmentation, and bootstrap alone shows a percentage improvement of 14.12 %. |
Copied to clipboard
| Challenge: | Empirical evaluation shows that the input embedding layer occupies a large portion of the model size. |
| Approach: | They propose an approach for compression of transformer-based models with minimal impact on downstream tasks by replacing the input embedding layer with dynamic embeddable computations. |
| Outcome: | Empirical evaluation shows that the proposed model is 15x smaller (1.2 MB) compared to the traditional model. |
Copied to clipboard
| Challenge: | Existing datasets designed specifically for the Bengali language have been limited. |
| Approach: | They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR. |
| Outcome: | The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) has been used to adapt Large Language Models to a variety of tasks, but it requires substantial computational resources to perform. |
| Approach: | They propose a low-rank adaptive learning approach that leverages LoRA's in-context learning capability through prompt matching via reinforcement learning in resource-constrained environments. |
| Outcome: | The proposed model improves LoRA performance on evaluation metrics and utilises consumer-grade GPU resources. |
Copied to clipboard
| Challenge: | Key Point Analysis (KPA) extracts the main points from opinions and quantifies their prevalence. |
| Approach: | They propose a key point analysis framework that extracts the main points from opinions and quantifies their prevalence. |
| Outcome: | The proposed system is able to match sentences to key points over five datasets and demonstrate its performance. |
Copied to clipboard
| Challenge: | a number of legal documents are archived in the UK, including the Supreme Court's decisions and video recordings of court hearings. |
| Approach: | They propose to link segments in the text judgement to semantically relevant timespans in the videos of the hearings. |
| Outcome: | The proposed tool links segments in the text judgement to semantically relevant timespans in the videos of the hearings. |
Copied to clipboard
| Challenge: | Existing recommendation system invites experts to write marketing themes and select relevant commodities, which suffer from difficulty in mass production, poor timeliness and low online indicators. |
| Approach: | They propose to use pretrained language model to generate marketing themes and commodity consistency module to select relevant commodities for the generative theme. |
| Outcome: | The proposed system can generate popular marketing themes and select relevant commodities automatically and improve theme online effectiveness compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing datasets for incident management tasks are labor-intensive and time-consuming. |
| Approach: | They propose a new IncidentAI dataset for safety prevention that includes three tasks . they argue that NLP techniques are beneficial for analyzing incident reports . |
| Outcome: | The proposed dataset shows that NLP techniques are beneficial for analyzing incident reports to prevent future failures. |
Copied to clipboard
| Challenge: | Existing approaches to service account retrieval have limited human annotation, resulting in labor-intensive and time-consuming tasks. |
| Approach: | They propose an Auxiliary task Boosted Multi-Task Learning method which introduces multiple auxiliary tasks and enhances the performance of the main task, service account retrieval. |
| Outcome: | The proposed method improves the performance of the main task, service account retrieval. |
Copied to clipboard
| Challenge: | Existing methods for extracting structured information from videos are coarse-grained at segment level and unable to capture finegrained information at the entity level. |
| Approach: | They propose a task for extracting hierarchical key information from visual texts on videos . they decouple the task into four subtasks and propose two implementation solutions . |
| Outcome: | The proposed solutions achieve remarkable performance and efficient inference speed on a well-defined dataset. |
Copied to clipboard
| Challenge: | Existing studies have focused on disfluency detection and removal, with limited studies into its impact on downstream tasks. |
| Approach: | They propose to incorporate disfluency in summarization models to reduce the impact of replacement disfluencies on natural language processing tasks. |
| Outcome: | The proposed model improves on both public and real-life datasets and shows that it can handle disfluent data with up to 6.99-point degradation in Rouge-L score and replacement disfluencies have the highest negative impact. |
Copied to clipboard
| Challenge: | Existing methods for extracting structured insights from reviews suffer from drawbacks . lack of structure, non-standard aspect names, lack of abundant training data limit their effectiveness and applicability. |
| Approach: | They propose a semi-supervised multi-level taxonomy from raw customer reviews and a semantic similarity heuristic approach to generate labelled data. |
| Outcome: | The proposed approach outperforms existing methods in structure, hierarchy and completeness. |
Copied to clipboard
| Challenge: | Extensive research has been done to recognize entities in spoken input. |
| Approach: | They propose to fine-tune pre-trained speech encoders to extract spoken entities directly from speech without the need for text transcription. |
| Outcome: | The proposed approach outperforms the 2-step approach for extracting spoken entities from human-computer conversations. |
Copied to clipboard
| Challenge: | generative models are used for product attribute extraction, a new field in information extraction and e-commerce. |
| Approach: | They analyze generative models for product attribute extraction and demonstrate their utility . they perform experiments on Amazon and MAVE product attribute datasets . |
| Outcome: | The proposed model can generate implicit attribute values, which state-of-the-art models are unable to extract. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable performance by following natural language instructions without fine-tuning them on domain-specific tasks and data. |
| Approach: | They propose an in-car retrieval-augmented conversational question-answering system that uses large language models to generate natural, safe and domain-specific answers. |
| Outcome: | The proposed system outperforms state-of-the-art LLMs in generating safe and domain-specific answers. |
Copied to clipboard
| Challenge: | Natural Language Processing has seen major breakthroughs in the last few years, but transferring these advances into industry applications can be difficult. |
| Approach: | They propose to use a BUSiness Transaction Entity Recognition dataset to support industry-oriented research by exploiting both general-purpose and domain-specific language models. |
| Outcome: | The proposed model is the best performing model and an additional silver corpus to BUSTER. |
Copied to clipboard
| Challenge: | Large Language Models have proven successful at modelling tasks, but they are expensive and slow to scale. |
| Approach: | They propose a Multi-Word Tokenizer that represents frequent multi-word expressions as single tokens. |
| Outcome: | The proposed tokenizer is more robust across shorter sequence lengths, allowing for major speedups via early sequence truncation. |
Copied to clipboard
| Challenge: | Tabular data analysis is an important application task of large language models, but advanced models are not yet on par with expert level performance. |
| Approach: | They propose to employ Large Language Models to facilitate an automated guide and execute high-precision data analyzes on tabular datasets. |
| Outcome: | The proposed framework is based on large language models and an automated machine learning pipeline for predictive modeling. |
Copied to clipboard
| Challenge: | End-to-end ASR models struggle to recognize uncommon domain-specific words due to limited audio context. |
| Approach: | They propose a "Retrieve and Copy" mechanism to improve latency while retaining the accuracy even when scaled to a large catalog. |
| Outcome: | The proposed method achieves 6% more word error rate reduction and 3.6% improvement in F1 when scaled to a large catalog size while retaining the accuracy. |
Copied to clipboard
| Challenge: | Existing training datasets for steering use cases are limited due to the cold-start problem. |
| Approach: | They propose a steering detection model that predicts whether a follow-up turn is a user’s attempt to steer the previous command. |
| Outcome: | The proposed model outperforms existing models on human-graded evaluation sets and shows that it can identify steering intent with over 95% accuracy. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models are useful, honest, harmless (HHH) however, RLHF requires high hardware resources and human efforts. |
| Approach: | They propose a framework that allows LLMs to align themselves with HHH . they use IF and reinforcement learning from human feedback to fine-tune their models . |
| Outcome: | The proposed framework achieves similar performance to RLHF and human-generated models with a minimal alignment tax. |
Copied to clipboard
| Challenge: | E-commerce product catalogs contain billions of items with lengthy titles . this leads to a gap between how customers refer to these unnatural titles - and how they are used . |
| Approach: | They propose a novel approach to product title summarization that uses a fine-tuned instruction strategy to train a highly accurate model. |
| Outcome: | The proposed approach can generate more accurate product title summaries with an improvement of over 14 and 8 BLEU and ROUGE points. |
Copied to clipboard
| Challenge: | Existing methods to perform visualization recommendation require a large corpus of dataset-visualization pairs for training and lack natural explanations for their results. |
| Approach: | They propose a new method that uses a ChatGPT-based prompting approach to perform visualization recommendation and return human-like explanations using very few demonstration examples. |
| Outcome: | The proposed method outperforms or performs similarly to supervised learning models like Random Forest, Decision Tree, and MLP, in both few-shot and zero-shot settings. |
Copied to clipboard
| Challenge: | DUBLIN is a pixel-based visual document understanding model that does not rely on OCR. |
| Approach: | They propose a pixel-based visual document understanding model that does not rely on OCR. |
| Outcome: | The proposed model performs on extractive tasks such as DocVQA, InfoVQA and AI2D, and strong performance on abstraction datasets such as VisualMRC and text captioning. |
Copied to clipboard
| Challenge: | Document understanding tasks are a tedious task that requires extensive training and privacy constraints. |
| Approach: | They propose a method to collect weakly labeled data from the web to benefit VDER training . the collected dataset does not depend on specific document types or entity sets . |
| Outcome: | The proposed method does not depend on specific document types or entity sets, making it universally applicable to all VDER tasks. |
Copied to clipboard
| Challenge: | Despite strong in-domain performance, dense retrievers have shown poor generalization to out-of-domain zero-shot tasks where no training queries are available. |
| Approach: | They propose to generate domain-specific pseudo queries for fine-tuning with domain-relevant relevance between PQ and documents. |
| Outcome: | The proposed approach is more robust to domain shifts, validated on BEIR zero-shot tasks. |
Copied to clipboard
| Challenge: | Existing product question answering models do not provide labelled data for the task and description information for products is very lengthy. |
| Approach: | They propose a distant supervision-based NLI model to prepare training data without manual efforts. |
| Outcome: | The proposed model outperforms standard multi-task fine-tuning and improves 6% in human evaluation over baselines. |
Copied to clipboard
| Challenge: | Recent advances in machine learning and artificial intelligence have opened up numerous opportunities and challenges in financial time series forecasting. |
| Approach: | They propose to use Large Language Models for explainable financial time series forecasting to leverage cross-sequence information and extract insights from text and price time series. |
| Outcome: | The proposed model outperforms ARMA-GARCH and gradient-boosting tree models while underperforming on other models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) and their applications in low-resource languages are limited due to lack of training data and benchmarking datasets. |
| Approach: | They propose a question-response system for Vietnamese that uses LLMs . they propose to open-source the model and train it on benchmark datasets based on Vietnamese data . |
| Outcome: | The proposed question answering system for Vietnamese is open-source and performant . it can learn and capture human-like text, but there is a gap in evaluations for Vietnamese . |
Copied to clipboard
| Challenge: | a weather search system is used to retrieve weather data from a massive weather database . a lack of navigation and time-consuming navigation hinders accurate weather forecasting . |
| Approach: | They propose a weather search system that allows users to retrieve weather data from a massive weather database with simple queries. |
| Outcome: | The proposed system achieves an average MRR and Recall of 0.82 on 4 million data points . it is based on a weather database at the Korea Meteorological Administration . |
Copied to clipboard
| Challenge: | Existing methods for deep semantic retrieval are highly sensitive to hyper-parameters . a novel adaptive metric learning method is proposed to overcome this limitation . |
| Approach: | They propose a method that adaptively obtains hyper-parameters without fixed or extra-trainable hyper-parmeters . they adopt a symmetric metric learning method to mitigate model collapse issues . |
| Outcome: | The proposed method outperforms existing methods on a real-world dataset and brings economic benefits. |
Copied to clipboard
| Challenge: | Existing approaches to code generation rely on rejection sampling to generate multiple code snippets then select the best. |
| Approach: | They propose a framework that prioritizes sampling on test problems that models can solve. |
| Outcome: | The proposed framework reduces sampling costs while maintaining comparable code generation performance. |
Copied to clipboard
| Challenge: | Performing inference on large volumes of samples can be computationally and financially costly. |
| Approach: | They propose a prompting approach that enables large language models to run inference in batches instead of one sample at a time. |
| Outcome: | The proposed prompting reduces both token and time costs while retaining downstream performance. |
Copied to clipboard
| Challenge: | Defective queries impact the robustness of conversational AI systems such as Alexa, Siri or Google Assistant. |
| Approach: | They propose a Personalized Query Rewriting system that takes into account individual preferences or unique error patterns identified from a user's historical interactions with the conversational AI. |
| Outcome: | The proposed approach has been proven on a large-scale real-world dataset and online A/B experiments. |
Copied to clipboard
| Challenge: | a recent study of controversy-handling in large language models (LLMs) has shown that people may become increasingly dependent on such systems for information. |
| Approach: | They propose to construct a controversial questions dataset using a subset of a publicly available dataset. |
| Outcome: | The proposed dataset presents challenges concerning knowledge recency, safety, fairness, and bias. |
Copied to clipboard
| Challenge: | Non-profit industry needs a system for accurately matching fund-seekers with fund-givers aligned in cause and target beneficiary group. |
| Approach: | They propose a search system that takes a fund-giver’s mission description as input and returns a ranked list of fund-seekers as output. |
| Outcome: | The proposed system improves on the non-profit evaluation dataset and the state-of-the-art model. |