Challenge: Anonymity on the Darknet allows vendors to stay undetected by using multiple vendor aliases or frequently migrating between markets.
Approach: They propose an NLP-based approach that examines writing patterns to verify, identify, and link unique vendor accounts across text advertisements on seven public Darknet markets.
Outcome: The proposed approach can help law enforcement agencies make more informed decisions by verifying and identifying migrating vendors and their potential aliases on existing and Low-Resource (LR) emerging Darknet markets.

Similar Papers

NLP-ADBench: NLP Anomaly Detection Benchmark (2025.findings-emnlp)

Copied to clipboard

Challenge: Anomaly detection (AD) is an important machine learning task, but its effectiveness in detecting harmful content, phishing attempts, and spam reviews is limited.
Approach: They introduce NLP-ADBench, the most comprehensive NLP anomaly detection benchmark to date . it includes eight curated datasets and 19 state-of-the-art algorithms .
Outcome: The NLP-ADBench benchmark includes 19 state-of-the-art methods and 8 curated datasets . no single model dominates across all datasets, indicating need for automated model selection .
IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort Advertisements (2023.emnlp-main)

Copied to clipboard

Challenge: a significant number of human trafficking cases are associated with online advertisements . identification of HT vendors is challenging for law enforcement agencies .
Approach: IDTraffickers uses 87,595 text ads and 5,244 vendor labels to link HT vendors . a macro-F1 score is achieved in a closed-set classification environment .
Outcome: IDTraffickers is a dataset that enables verification and identification of HT vendors . the model achieves a macro-F1 score in a closed-set classification environment .
Investigating Links between Illicit Massage Businesses through Natural Language Processing and Graph Machine Learning (2026.findings-acl)

Copied to clipboard

Challenge: Illicit massage businesses exploit vulnerable individuals through forced sex or labor . identifying key indicators from vast volume of data associated with these businesses poses significant challenge .
Approach: They propose a multi-stream data integration approach focusing on Yelp reviews . they propose bespoke subgraph extraction strategies to detect links between massage businesses .
Outcome: The proposed approach outperforms baseline methods in a multi-stream data integration framework based on consumer reviews on Yelp.com and contextual data from the U.S. Census and business license records.
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)

Copied to clipboard

Challenge: The first workshop on crowdsourcing for NLP is open to all .
Approach: The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks.
Outcome: The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data .
SYSML: StYlometry with Structure and Multitask Learning: Implications for Darknet Forum Migrant Analysis (2021.emnlp-main)

Copied to clipboard

Challenge: Crypto markets are forums where goods and services are exchanged between parties who use encryption to conceal their identities.
Approach: They propose a stylometry-based multitask learning approach for natural language and model interactions using graph embeddings.
Outcome: The proposed approach outperforms existing methods in four darknet forums with a lift of up to 2.5X on the mean retrieval rank and 2X on recall@10.
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) (D19-61)

Copied to clipboard

Challenge: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China .
Approach: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . call for papers for this second workshop met with a strong response .
Outcome: the EMNLP-IJCNLP 2019 workshop on deep learning approaches for low-resource natural language processing takes place in Hong Kong, China.
Context-specific Language Modeling for Human Trafficking Detection from Online Advertisements (P19-1)

Copied to clipboard

Challenge: Human trafficking is a worldwide crisis.
Approach: They propose a method to detect trafficking ads on online sites using natural language processing using a pre-trained textual language model.
Outcome: The proposed classifier significantly outperforms any single feature set alone.
AD-NLP: A Benchmark for Anomaly Detection in Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: Methods for Anomaly Detection in text have shown strong empirical results on ad-hoc anomaly setups that are usually made by downsampling some classes of a labeled dataset.
Approach: They propose a unified benchmark for detecting various types of anomalies . they evaluate two strong shallow baselines and two current state-of-the-art neural approaches .
Outcome: The proposed benchmarks provide insights into the knowledge the neural models are learning when performing the task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations