Generalized Embedding Models for Industry 4.0 Applications (2025.emnlp-industry)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) to automate tasks has emerged as the next frontier of innovation.
Approach: They propose a model that generalizes to queries involving similar assets and retrieves relevant items from natural language tasks.
Outcome: The proposed model can be used to generalize to queries involving similar assets, such as identifying sensors relevant to an asset’s failure mode.

Similar Papers

Towards Unified Task Embeddings Across Multiple Models: Bridging the Gap for Prompt-Based Large Language Models and Beyond (2024.findings-acl)

Copied to clipboard

Challenge: Existing task embedding methods rely on fine-tuned, task-specific language models, which hinders their adaptability to prompt-guided Large Language Models (LLMs).
Approach: They propose a framework for unified task embedding that harmonizes task embeds from various models within a single vector space.
Outcome: The proposed framework harmonizes task embeddings from various models within a single vector space.
Understanding the Influence of Synthetic Data for Text Embedders (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in general purpose text embedders have been driven by training on synthetic training data.
Approach: They propose to use GPT-4 to produce high quality synthetic data that expands existing training datasets for embeddings to new tasks.
Outcome: The proposed dataset is high quality and leads to consistent improvements in performance.
Embedding Strategies for Specialized Domains: Application to Clinical Entity Recognition (P19-2)

Copied to clipboard

Challenge: Off-the-shelf word embeddings tend to perform poorly on texts from specialized domains such as clinical reports.
Approach: They combine off-the-shelf contextual embeddings with static word2vec embedders trained on a small in-domain corpus built from task data to reach and sometimes outperform representations learned from a large corpus in the medical domain.
Outcome: The proposed embedding strategies outperform representations learned from a large corpus in the medical domain.
Group, Embed and Reason: A Hybrid LLM and Embedding Framework for Semantic Attribute Alignment (2025.emnlp-industry)

Copied to clipboard

Challenge: a framework to align attributes that refer to the same concept but differ across schemas is challenging in schema only settings where no instance data is available due to ambiguous names, inconsistent descriptions, and domain-specific terminologies.
Approach: They propose a framework that combines contextual reasoning and embedding-based similarity to address token limitations and hallucinations.
Outcome: The proposed framework scales to large schemas and shows strong performance on healthcare schemas.
Semantic search with domain-specific word-embedding and production monitoring in Fintech (2020.coling-demos)

Copied to clipboard

Challenge: a novel system with domain-specific custom language models for accurate search terms expansion addresses several challenges faced in an industry-setting . many industry-grade search engines exist, but their accuracy can be sub-optimal due to domain specificity and terms provided by the users.
Approach: They propose an end-to-end information retrieval system with domain-specific custom language models for accurate search terms expansion.
Outcome: The proposed system is used in risk management and has wide applicability to other domains.
Improving Text Embeddings with Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for obtaining text embeddings require complex training pipelines . authors leverage proprietary LLMs to generate diverse synthetic data for text embeds based on 93 languages .
Approach: They propose a method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.
Outcome: The proposed method achieves strong performance on competitive text embedding benchmarks without using any labeled data.
Probing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case (2020.coling-main)

Copied to clipboard

Challenge: Semantic embeddings have advanced the state of the art for natural language processing tasks . but their inner workings are poorly understood and there is a shortage of analysis tools .
Approach: They propose to extend visual-semantic embeddings to multimodal domains by defining probing tasks for embeddable image-caption pairs and testing them with classifiers.
Outcome: The proposed probing tasks show up to 16% more accurate on visual-semantic embeddings compared to unimodal embedders . the proposed extensions to multimodal domains have been lauded as promising in natural language processing .
Is Language Modeling Enough? Evaluating Effective Embedding Combinations (2020.lrec-1)

Copied to clipboard

Challenge: specialized embeddings are not available for tasks like entity linking or paragraph classification.
Approach: They evaluate whether universal embeddings can be complemented by specialized embeddables.
Outcome: The proposed embeddings outperform state-of-the-art embeddables without any fine-tuning.
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal embedding models encode multimedia inputs into latent vector representations.
Approach: They propose to synthesize multimodal multilingual data using a multimodal large language model . they identify three criteria for high-quality synthetic multimodal data .
Outcome: The proposed model outperforms existing models on the MMEB Benchmark and the XTD benchmark.
Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to augment LLMs with Knowledge Graphs (KGs) Knowledge-intensive tasks are prone to errors and require a large amount of knowledge to be understood.
Approach: They propose a framework for augmenting LLMs through Knowledge Graphs (KGs) they propose KGs can be used to enhance performance in knowledge-intensive tasks .
Outcome: Experimental results show that a small domain-specific KG can benefit from a performance boost in downstream tasks when linked to a substantial general-purpose KG.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations