Challenge: a new method for mapping job tasks to labor market ontologies is proposed . a top configuration of the method achieved a notable performance improvement .
Approach: They use ontological data with Multiple Negatives Ranking loss to extract job tasks from job postings . they integrate labeled job advertisement data into training to improve their mapping .
Outcome: The proposed method improves on the German job ads and their ontology . it can be used to map job tasks to established labor market ontologies or taxonomies .

Similar Papers

Improving Online Job Advertisement Analysis via Compositional Entity Extraction (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on IE in OJAs has focused on skills extraction, but other information is extracted using job tasks, job titles, and work tools.
Approach: They propose a compositional entity modeling framework for requirement extraction from online job advertisements (OJAs) they annotate a manually annotated dataset of 500 German job ads that captures roles, tools, experience levels, attitudes, and their functional context.
Outcome: The proposed framework can extract requirements from a manually annotated dataset of 500 German job ads.
Evaluation of Transfer Learning and Domain Adaptation for Analyzing German-Speaking Job Advertisements (2022.lrec-1)

Copied to clipboard

Challenge: a paper presents text mining approaches on German-speaking job advertisements . transfer learning and domain adaptation are used to build text mining applications .
Approach: They propose text mining approaches on German-speaking job advertisements . they use transfer learning and domain adaptation to build language models adapted to job ads .
Outcome: The proposed approaches outperform general-domain language models pre-trained on ten times more data.
SkiLLens: Recognising and Mapping Novel Skills from Millions of Job Ads Across Europe Using Language Models (2026.eacl-industry)

Copied to clipboard

Challenge: Online job ads (OJAs) provide a real-time view of changing demands but require first retrieving skill mentions from unstructured text and then solving the entity linking problem of connecting them to standardized skill taxonomies.
Approach: They propose a multilingual human-in-the-loop pipeline that extracts candidate skills from national OJA corpora using country-specific word embeddings.
Outcome: The proposed pipeline enables timely, multilingual monitoring of emerging skills, supporting agile policy-making and targeted training initiatives.
Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for creating robust occupation taxonomies are slow and expensive . a robust taxonomy is critical for job recommendation and labor market intelligence applications .
Approach: They propose a framework that automates creation of occupation taxonomies from job postings . they use global semantic clustering to distill core occupations, then a reflection-based multi-agent system to iteratively build a coherent hierarchy.
Outcome: The proposed framework produces taxonomies that capture unique regional characteristics.
Generalized Embedding Models for Industry 4.0 Applications (2025.emnlp-industry)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) to automate tasks has emerged as the next frontier of innovation.
Approach: They propose a model that generalizes to queries involving similar assets and retrieves relevant items from natural language tasks.
Outcome: The proposed model can be used to generalize to queries involving similar assets, such as identifying sensors relevant to an asset’s failure mode.
Document-based Recommender System for Job Postings using Dense Representations (N18-3)

Copied to clipboard

Challenge: 45% of job posting traffic is driven by recommender systems for job postings . a large-scale job recommendation system is needed to detect similarity between job posting and item-to-item based recommendations.
Approach: They propose to use dense vector representations to enhance a large-scale job recommendation system and rank job advertisements regarding similarity.
Outcome: The proposed method increases the click-through rate on job recommendations by 8.0%.
Retrieving Skills from Job Descriptions: A Language Model Based Extreme Multi-label Classification Framework (2020.coling-main)

Copied to clipboard

Challenge: 65% of job descriptions miss describing a significant number of relevant skills, a problem we address with a deep learning model based on a multi-label classification problem .
Approach: They propose a deep learning model to learn the set of enumerated job skills associated with a job description.
Outcome: The proposed model improves on an existing baseline solution by over 9% and 7% absolute improvements in terms of recall and normalized discounted cumulative gain.
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)

Copied to clipboard

Challenge: Existing studies on term annotation show that even experts differ in their understanding of termhood .
Approach: They propose a new dataset of term annotation that examines the common understanding of what constitutes a term.
Outcome: The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term .
German SRL: Corpus Construction and Model Training (2024.lrec-main)

Copied to clipboard

Challenge: Existing semantic role annotation resources are lacking for German.
Approach: They propose a translation-based approach to train German semantic role models using semantic annotations and alignment models.
Outcome: The proposed method achieves competitive evaluation scores, but avoids limitations of previous approaches.
USB: A Unified Summarization Benchmark Across Tasks and Domains (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing summarization benchmarks lack the rich annotations needed to address important problems related to control and reliability.
Approach: They propose a Wikipedia-derived summarization benchmark with crowd-sourced annotations . they find that fine-tuned models outperform larger few-shot prompted language models .
Outcome: The proposed model outperforms many-shot prompted language models on multiple tasks . the proposed model is based on Wikipedia annotations and can be used in other domains .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations