Papers by Steven Derby

4 papers
SPICED: News Similarity Detection Dataset with Multiple Topics and Complexity Levels (2024.lrec-main)

Copied to clipboard

Challenge: Existing semantic textual similarity (STS) datasets are not suitable for news similarity detection due to their specificity to a single topic.
Approach: They propose to segment news similarity datasets into topics to improve model training . they propose four different levels of complexity specifically designed for news similarities detection task .
Outcome: The proposed dataset includes seven topics: Crime & Law, Culture & Entertainment, Disasters & Accidents, Economy & Business, Politics / Conflicts, Science & Technology, Sports.
Feature2Vec: Distributional semantic modelling of human property knowledge (D19-1)

Copied to clipboard

Challenge: Existing distributional semantic models of word meaning are limited in size due to constraints associated with exhaustively listing properties for large numbers of words.
Approach: They propose a method for mapping human property knowledge onto a distributional semantic space and adapt it to the task of modelling concept features.
Outcome: The proposed model performs better on evaluation tasks and improves on other evaluation tasks.
Encoding Lexico-Semantic Knowledge using Ensembles of Feature Maps from Deep Convolutional Neural Networks (2020.coling-main)

Copied to clipboard

Challenge: Semantic models derived from visual information have overcome some of the limitations of text-based distributional semantic models.
Approach: They build image-based meta-embeddings from computer vision models which incorporate information from all layers of the network and encode a richer set of semantic attributes.
Outcome: The proposed representations encode a richer set of semantic attributes and yield a more complete representation of human conceptual knowledge.
Topics as Entity Clusters: Entity-based Topics from Large Language Models and Graph Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Topic models aim to reveal latent structures within corpus of text through term-frequency statistics over bag-of-words representations.
Approach: They propose to use bimodal vector representations of entities to extract latent representations from large language models and graph neural networks trained on symbolic relations to derive the most salient aspects of these conceptual units.
Outcome: The proposed approach is better suited to working with entities than state-of-the-art models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations