Papers by Hsuan-Tien Lin
Preserving Zero-shot Capability in Supervised Fine-tuning for Multi-label Text Classification (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods that assume label descriptions ensure zero-shot capability lose their zero-shot capability during training. |
| Approach: | They propose a method that preserves the zero-shot capabilities of powerful dual encoders and label-wise attention networks by freezing the label encoder. |
| Outcome: | The proposed methods preserve the zero-shot capabilities of powerful dual encoder and label-wise attention network architectures by freezing the label encoder. |
Understanding and Mitigating Spurious Correlations in Text Classification with Neighborhood Analysis (2024.findings-eacl)
Copied to clipboard
| Challenge: | Recent research has revealed that machine learning models have a tendency to leverage spurious correlations that exist in the training set but may not hold true in general circumstances. |
| Approach: | They propose a metric to detect spurious tokens and a family of regularization methods to mitigate spurious correlations in text classification. |
| Outcome: | The proposed method prevents spurious clusters and significantly improves the robustness of classifiers without auxiliary data. |
Even the Simplest Baseline Needs Careful Re-investigation: A Case Study on XML-CNN (2022.naacl-main)
Copied to clipboard
| Challenge: | XML-CNN has been a popular research topic in NLP due to its superior performance . however, the increasing complexity brings difficulties to ensure the true architectural progress . |
| Approach: | They propose to re-examine an influential multi-label text classification method . they propose suitable baselines for multi-level text classification tasks . |
| Outcome: | The proposed method performs better than the original model, the authors show . they show that the re-implementation reveals contradictory results to the original work . |
Cold-start Active Learning through Self-supervised Language Modeling (2020.emnlp-main)
Copied to clipboard
| Challenge: | Labeling data is a fundamental bottleneck in machine learning due to annotation cost and time. |
| Approach: | They propose a strategy that uses the pre-training loss to find examples that surprise the model and minimize labeling costs. |
| Outcome: | The proposed approach reduces labeling costs and costs by using pre-trained language models. |