Papers by Alexander Min

4 papers
A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: a study shows that DWT can be effective in the vision domain and natural language processing pre-training stages.
Approach: They examine three key factors to optimize Distillation from Weak Teacher (DWT) DWT is a method of transferring knowledge from a weaker teacher model to a larger student model to improve its performance.
Outcome: a new study examines three key factors to optimize DWT in NLP pre-training scenarios . the impact of teacher model quality and guidelines for adjusting the weighting value for DW T loss are examined .
ExcavatorCovid: Extracting Events and Relations from Text Corpora for Temporal and Causal Analysis for COVID-19 (2021.emnlp-demo)

Copied to clipboard

Challenge: a new machine reading system ingests open-source text documents to analyze COVID-19 events . the system extracts COVId-19 related events and relations between them .
Approach: They propose a machine reading system that ingests open-source text documents and extracts COVID-19 related events and relations between them.
Outcome: The proposed system extracts COVID-19 related events and relations from open-source text . it will help government agencies alleviate the information overload and respond to COVId-19 .
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets are often informed by established research directions in the NLP community.
Approach: They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Outcome: The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Co-training and Co-distillation for Quality Improvement and Compression of Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Knowledge Distillation (KD) compresses expensive pre-trained language models . however, most smaller models fail to surpass performance of larger model .
Approach: They propose a framework that co-trains two models while mutually distilling knowledge to improve performance and inference speed together.
Outcome: The proposed framework outperforms the original larger model by 1.66 on the GLUE benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations