Papers by Shijing Si

4 papers
Methods for Numeracy-Preserving Word Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Word embedding models capture semantic relationships between words but fail to capture numerical properties associated with numbers.
Approach: They propose a method to assign and learn embeddings for numbers using word embedders.
Outcome: The proposed model outperforms pre-trained word embedding models across multiple examples of two tasks.
Integrating Task Specific Information into Pretrained Language Models for Low Resource Fine Tuning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing pretrained language models are agnostic to downstream information and can overfit when fine-tuned with low resource datasets.
Approach: They integrate label information as a task-specific prior into the self-attention component of pretrained BERT models.
Outcome: Experiments on benchmarks and real-word datasets show that the proposed approach can improve the performance of pretrained models when fine-tuned with small datasets.
Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization (2022.emnlp-main)

Copied to clipboard

Challenge: Existing semantic hashing methods only learn a binary code for each document and use Hamming distance to evaluate document distances.
Approach: They propose to leverage BERT embeddings to perform efficient retrieval based on product quantization technique . they transform original BERT embedded codewords and feed it into a probabilistic product quantizer module .
Outcome: The proposed method outperforms current state-of-the-art methods on three benchmarks.
Leveraging BERT and TFIDF Features for Short Text Clustering via Alignment-Promoting Co-Training (2024.emnlp-main)

Copied to clipboard

Challenge: Existing clustering methods rely on keyword information, but they lack this information.
Approach: They propose a CO**-**T**raining **C**lustering framework to make use of BERT and TFIDF features.
Outcome: The proposed framework outperforms existing SOTA methods on eight datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations