Papers with Amazon

40 papers
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) (N18-2)

Copied to clipboard

Challenge: NAACL HLT 2018 is the biggest NAAPL conference to date . this year's conference highlights the vibrancy and vitality of the field .
Approach: a new review form and an opportunity for authors to review the reviewers were introduced at this year's conference . the test-of-time awards are named in memory of Aravind Joshi, who died this year .
Outcome: the biggest NAACL conference to date features a new review form and the Test-of-Time awards . the industrial track features papers that focus on scalable, interpretable, reliable and customer facing methods for industrial applications .
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (N18-1)

Copied to clipboard

Challenge: NAACL HLT 2018 is the biggest NAAPL conference to date . this year's conference highlights the vibrancy and vitality of the field .
Approach: a new review form and an opportunity for authors to review the reviewers were introduced at this year's conference . the test-of-time awards are named in memory of Aravind Joshi, who died this year .
Outcome: the biggest NAACL conference to date features a new review form and the Test-of-Time awards . the industrial track features papers that focus on scalable, interpretable, reliable and customer facing methods for industrial applications .
Proceedings of the Second Workshop on Fact Extraction and VERification (FEVER) (D19-66)

Copied to clipboard

Challenge: a workshop on fact extraction and verification is being held at EMNLP 2019 .
Approach: the authors propose a workshop to promote research in joint Fact Extraction and VERification . they received 25 submissions, five of which were system descriptions .
Outcome: the proposed workshop promotes research in joint Fact Extraction and VERification (FEVER) the updated leaderboard will be presented at the second workshop.
Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (N19-1)

Copied to clipboard

Challenge: 2019 Program Co-Chairs have managed the largest number of submissions at any NAACL to date . submissions in 2019 almost doubled from the previous year .
Approach: NAACL's 2019 Program Co-Chairs have managed the largest number of submissions to date . 2019 will see a diversity committee, a careers in NLP panel and a panel on diversity .
Outcome: The 2019 Program Co-Chairs have managed the largest number of submissions at any NAACL to date.
Towards Opinion Summarization of Customer Reviews (P18-3)

Copied to clipboard

Challenge: Existing methods to summarize text are limited to small, homogeneous datasets . authors outline future directions to solve these problems .
Approach: They propose to use neural networks to generate summaries of user-generated travel reviews . they aim to take into account shifting opinions over time and address these issues .
Outcome: The proposed method will make it easier for users of review sites to make more informed decisions.
Pay-Per-Request Deployment of Neural Network Models Using Serverless Architectures (N18-5)

Copied to clipboard

Challenge: Using Amazon’s Lambda service for feedforward evaluation and DynamoDB for word embeddings, we demonstrate a serverless deployment of neural networks for NLP applications.
Approach: They propose a pay-per-request pricing model for neural network deployment in NLP applications using Amazon’s Lambda service for feedforward evaluation and DynamoDB for storing word embeddings.
Outcome: The proposed architecture is scalable and inexpensive.
Multilingual Continual Learning using Attention Distillation (2025.coling-industry)

Copied to clipboard

Challenge: Existing models for Query-product relevance classification are not accurate across multiple languages.
Approach: They propose a multilingual continual learning framework that adds adapters for each new language and incorporates a fusion layer above language-specific adapters.
Outcome: The proposed approach reduces trainable parameters by 80% while outperforming SOTA CL methods on proprietary and external datasets.
Ask-and-Verify: Span Candidate Generation and Verification for Attribute Value Extraction (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing reading comprehension models can over-generate attribute values which hinders precision.
Approach: They propose a product attribute value extraction task that captures key factual information from product descriptions and a new end-to-end pipeline framework called Ask-and-Verify.
Outcome: The proposed framework outperforms existing models by up to 3.1% F1 absolute improvement points while scaling to thousands of attributes.
Multi-Modal Generative Adversarial Network for Short Product Title Generation in Mobile E-Commerce (N19-2)

Copied to clipboard

Challenge: Existing methods for short product title generation only consider textual information from long titles . MM-GAN incorporates image information and attribute tags from product, as well as textual info from original long titles.
Approach: They propose a multi-modal generative adversarial network for short product title generation in E-commerce . they incorporate image information and attribute tags from product, as well as textual information from original long titles .
Outcome: The proposed model outperforms state-of-the-art methods on a large-scale E-commerce dataset.
Deep Metric Learning to Hierarchically Rank - An Application in Product Retrieval (2023.emnlp-industry)

Copied to clipboard

Challenge: e-commerce search engines use customer behavior signals to augment lexical matching and improve search relevance.
Approach: They propose a method to identify duplicate and near-duplicate products across stores . they use Hierarchical Ranked Multi Similarity Loss to learn hierarchical metric space .
Outcome: The proposed model outperforms baselines in terms of catalog coverage and precision of the mappings.
Rationale-Guided Distillation for E-Commerce Relevance Classification: Bridging Large Language Models and Lightweight Cross-Encoders (2025.coling-industry)

Copied to clipboard

Challenge: Large-scale e-commerce search systems typically follow a multi-step process to retrieve relevant products for a given query.
Approach: They propose a distillation approach that uses "rationales" generated by Large Language Models to guide smaller cross-encoder models.
Outcome: The proposed model achieves ROC-AUC improvements of 1.4% on 9 multilingual e-commerce datasets, 2.4% on 3 ESCI datasets and 6% on GLUE datasets while being 50 times faster per sample.
EasyTurk: A User-Friendly Interface for High-Quality Linguistic Annotation with Amazon Mechanical Turk (2021.eacl-demos)

Copied to clipboard

Challenge: Amazon Mechanical Turk (AMT) is one of the most popular crowd-sourcing platforms, allowing researchers from all over the world to create linguistic datasets quickly and at a relatively low cost.
Approach: They propose to improve the potential of Amazon Mechanical Turk by adding some new features to the tool.
Outcome: The proposed tool improves the performance of Amazon Mechanical Turk by adding new features.
Are the Tools up to the Task? an Evaluation of Commercial Dialog Tools in Developing Conversational Enterprise-grade Dialog Systems (N19-2)

Copied to clipboard

Challenge: Existing toolsets are incomplete in meeting the goal of building effective dialog systems, authors say .
Approach: They compare dialog tools available from a number of companies to determine their strengths and weaknesses . they provide quantitative and qualitative results in three main areas: natural language understanding, dialog, and text generation .
Outcome: The toolsets are incomplete, but they are compared to other tools to determine their strengths and weaknesses.
Large Scale Generative Multimodal Attribute Extraction for E-commerce Attributes (2023.acl-industry)

Copied to clipboard

Challenge: E-commerce websites often don’t label or mislabel attributes of products .
Approach: They propose a multi-modal product attribute generation system that extracts product attributes from the product pages of eCommerce stores by using both text and images.
Outcome: The proposed model improves the recall@90P accuracy by 10.16% and 6.9 from the state-of-the-art models.
IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing (2024.emnlp-industry)

Copied to clipboard

Challenge: Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers.
Approach: They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc.
Outcome: The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination.
Generative Models for Product Attribute Extraction (2023.emnlp-industry)

Copied to clipboard

Challenge: generative models are used for product attribute extraction, a new field in information extraction and e-commerce.
Approach: They analyze generative models for product attribute extraction and demonstrate their utility . they perform experiments on Amazon and MAVE product attribute datasets .
Outcome: The proposed model can generate implicit attribute values, which state-of-the-art models are unable to extract.
From Disjoint Sets to Parallel Data to Train Seq2Seq Models for Sentiment Transfer (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for sentiment transfer have relied on unsupervised methods due to lack of parallel corpora.
Approach: They propose a method for creating parallel data to train Seq2Seq neural networks for sentiment transfer.
Outcome: The proposed method outperforms existing unsupervised methods in sentiment transfer tasks.
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is a widely studied task in natural language processing.
Approach: They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors.
Outcome: The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes.
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on product search tasks, but ignore potential risks.
Approach: They propose a data generation pipeline that leverages webpage content and interactive elements to create diverse, functionality-grounded user queries.
Outcome: The proposed framework assesses the performance and safety of web agents under dynamic, real-world e-commerce environments.
“Alexa in the wild” – Collecting Unconstrained Conversations with a Modern Voice Assistant in a Public Environment (2020.lrec-1)

Copied to clipboard

Challenge: Currently, many studies on human-machine interactions focus on private usage, short pre-defined tasks or specific domains.
Approach: They propose to collect 40 hours of device directed utterances during a science exhibition in germany and extract transcripts of both visitors requests and Alexa answers.
Outcome: The proposed dataset provides an unconstrained, unscripted public interaction with a voice assistant during a science exhibition in germany.
Unpaired Sentiment-to-Sentiment Translation: A Cycled Reinforcement Learning Approach (P18-1)

Copied to clipboard

Challenge: Existing studies for sentiment-to-sentiment "translation" only change the underlying sentiment and fail to keep the semantic content.
Approach: They propose a cycled reinforcement learning method that combines neutralization module and emotionalization module.
Outcome: The proposed method outperforms state-of-the-art systems on Yelp and Amazon review datasets.
Efficient Few-Shot Fine-Tuning for Opinion Summarization (2022.findings-naacl)

Copied to clipboard

Challenge: Abstractive summarization models are typically pre-trained on large amounts of generic texts . large annotated datasets of reviews paired with reference summaries are not available .
Approach: They propose a few-shot method which uses adapters to store in-domain knowledge . they pre-train adapters on unannotated customer reviews and fine-tune them on annotated datasets .
Outcome: The proposed method can store in-domain knowledge and improves on large annotated reviews . it improves coherence and redundancies on the Amazon and Yelp datasets .
Aggregating Crowd of LLMs for Cost-Effective Data Annotation (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have shown promise for automated data annotation, yet reliance on expensive commercial models like GPT-4 limits accessibility.
Approach: They propose to build a crowd of LLMs which aggregates annotations from multiple sLLMs using label aggregation algorithms.
Outcome: The proposed approach outperforms individual sLLMs and human crowd labels yields superior results compared to either method alone.
Combining Deep Learning and Topic Modeling for Review Understanding in Context-Aware Recommendation (N18-1)

Copied to clipboard

Challenge: Existing models for user reviews are limited by data sparsity and lack of data.
Approach: They propose to integrate LSTM and Topic Modeling to extract review information for recommender systems by utilizing user reviews.
Outcome: The proposed model outperforms existing models on Amazon review dataset and shows better ability on making topic clustering than traditional topic model based method.
Product Description and QA Assisted Self-Supervised Opinion Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to generate opinion summarization without supervised training data are limited due to the lack of additional sources.
Approach: They propose a synthetic dataset creation strategy that leverages reviews and additional sources to generate a pseudo-summary.
Outcome: The proposed approach achieves 14.5% improvement in ROUGE-1 F1 over existing models.
Delete, Retrieve, Generate: a Simple Approach to Sentiment and Style Transfer (N18-1)

Copied to clipboard

Challenge: Previous work using adversarial methods has struggled to produce high-quality outputs.
Approach: They propose a method that transforms a sentence to alter a specific attribute while preserving its attribute-independent content.
Outcome: The proposed method generates grammatical and appropriate responses on 22% more inputs than the best previous system, averaged over three attribute transfer datasets.
Multi-Domain Targeted Sentiment Analysis (2022.naacl-main)

Copied to clipboard

Challenge: Targeted Sentiment Analysis (TSA) is a task for generating insights from consumer reviews.
Approach: They propose a multi-domain TSA system that augments a given training set with diverse weak labels from assorted domains and augments it with Yelp reviews.
Outcome: The proposed model outperforms manual methods on three evaluation datasets across different domains and shows that it performs well.
Opinion Summarization by Weak-Supervision from Mix-structured Data (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for opinion summarization of multiple reviews lack reference summaries . OAs and ISs are often mismatched between review input and summary .
Approach: They propose a method to generate mixed-structured synthetic training data for opinion summarization.
Outcome: The proposed method outperforms existing methods on Yelp, Amazon and RottenTomatos datasets.
Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models’ Memories (2023.acl-long)

Copied to clipboard

Challenge: Pre-trained language models demonstrate excellent abilities to understand texts in the generic domain while struggling in a specific domain.
Approach: They propose to decouple the feed-forward networks of the Transformer architecture into two parts to maintain old-domain knowledge and a mixture-of-adapters gate to inject domain-specific knowledge in parallel.
Outcome: The proposed method achieves superior performance on in-domain, out-of-domain and knowledge-intensive tasks.
Few-Shot Learning for Opinion Summarization (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that abstractive summarization models fail to capture their essential properties due to the high cost of summary production.
Approach: They propose a few-shot framework for abstractive opinion summarization that bootstraps the output of an unsupervised model.
Outcome: The proposed framework outperforms extractive and abstractive methods on Amazon and Yelp datasets.
Rethinking Style Transformer with Energy-based Interpretation: Adversarial Unsupervised Style Transfer using a Pretrained Model (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to train text style transfer models with adversarial loss degrade fluency compared to other metrics.
Approach: They propose a method which leverages a pretrained language model to improve fluency by restructuring the discriminator and the model itself.
Outcome: The proposed model achieves state-of-the-art on three public benchmarks and achieved state-outperformance on the overall metrics.
Unsupervised Opinion Summarization as Copycat-Review Generation (2020.acl-main)

Copied to clipboard

Challenge: Recent work on opinion summarization has focused on extracting fragments from reviews, but we use novel sentences to generate abstractive summaries.
Approach: They propose an abstractive summarizer which does not use summaries in training and is trained end-to-end on a large collection of reviews.
Outcome: The proposed model produces fluent and coherent summaries reflecting consensus opinions on Amazon and Yelp reviews.
Decode with Template: Content Preserving Sentiment Transfer (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to transfer sentiments for text use only explicit sentiments and templates to remove them from input sentences.
Approach: They propose a method to transfer sentiments from input sentences to output sentences using templates.
Outcome: The proposed model significantly outperforms state-of-the-art models in content preservation.
Prompted Opinion Summarization with GPT-3.5 (2023.findings-acl)

Copied to clipboard

Challenge: Recent years have seen several shifts in summarization research, including extractive models.
Approach: They propose a pipeline method for applying GPT-3.5 to summarize user reviews . they propose three new metrics targeting faithfulness, factuality, and genericity .
Outcome: The proposed methods perform well in opinion summarization, the authors show . they also show that standard evaluation metrics do not reflect this performance .
YASO: A Targeted Sentiment Analysis Evaluation Dataset for Open-Domain Reviews (2021.emnlp-main)

Copied to clipboard

Challenge: YASO contains 2,215 English sentences from dozens of review domains, annotated with target terms and their sentiment.
Approach: They propose a new TSA evaluation dataset of open-domain user reviews in English . YASO contains 2,215 English sentences annotated with target terms and their sentiment .
Outcome: The proposed dataset verifies the reliability of the annotations and explores the characteristics of the collected data.
Robustifying Sentiment Classification by Maximally Exploiting Few Counterfactuals (2022.emnlp-main)

Copied to clipboard

Challenge: a recent study found that finetuned language models rely on spurious patterns in training data . this limitation limits their performance on out-of-distribution (OOD) test data.
Approach: They propose a method that only requires annotation of a small fraction of training data . they add 1% manual counterfactuals to training data and generate extra counterfacts in vector space .
Outcome: The proposed approach improves sentiment classification using IMDb data and other sets for OOD tests.
Synthesize, if you do not have: Effective Synthetic Dataset Creation Strategies for Self-Supervised Opinion Summarization in E-commerce (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to generate general and aspect-specific opinion summarization are limited due to their reliance on human-specified aspects and seed words.
Approach: They propose synthetic dataset creation approaches for general and aspect-specific opinion summarization . general opinion summaries struggle to generate faithful to the input reviews, they say . aspect- specific opinion summarisation models are limited due to reliance on human-specified aspects .
Outcome: The proposed approach outperforms existing models on three e-commerce test sets on general and aspect-specific opinion summarization.
Federated Meta-Learning for Emotion and Sentiment Aware Multi-modal Complaint Identification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on complaint identification are limited to text.
Approach: They propose a meta-learning-based multi-modal multi-task framework for identifying complaints using emotion recognition and sentiment analysis as auxiliary tasks.
Outcome: The proposed framework outperforms baselines and state-of-the-art approaches in centralized and federated meta-learning settings.
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation.
Approach: They propose to use a dataset to investigate a new type of bias in Large Language Models for code generation, provider bias, to determine whether the model favors specific providers.
Outcome: The proposed model favors services from Google and Amazon, but without explicit directives, and can modify input code to incorporate their preferred providers without user requests.
R-Fairness: Assessing Fairness of Ranking in Subjective Data (2025.acl-long)

Copied to clipboard

Challenge: Subjective data, reflecting individual opinions, permeates platforms like Yelp and Amazon . despite the prevalence of such platforms, little attention has been given to fairness in their context .
Approach: They propose a fairness assessment pipeline that starts with data collection phase and then iterates through rated items.
Outcome: The proposed approach favors groups writing best-ranked reviews over others on collaborative rating platforms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations