Papers by Chieh-Yang Huang
Assessing the Helpfulness of Learning Materials with Inference-Based Learner-Like Agent (2020.emnlp-main)
Copied to clipboard
| Challenge: | Prior work uses hand-crafted scores to recommend sentences but has difficulty adopting such scores to all the near-synonyms as near-near-sonyms differ in various ways. |
| Approach: | They propose an inference-based learner-like agent to mimic learner behavior and identify good learning materials by examining the agent's performance. |
| Outcome: | The proposed agent achieves the best performance in fill-in-the-blank and good example sentence selection tasks. |
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing algorithms that generate captions for scientific figures are costly and dependent on author-written captions. |
| Approach: | They constructed a human evaluation dataset that contains human judgments for 3,600 scientific figure captions for 600 arXiv figures. |
| Outcome: | The proposed model outperforms all other models and outperformed undergraduates in achieving a Kendall correlation score of 0.401 with Ph.D. students’ rankings. |
Semantic Frame Forecast (2021.naacl-main)
Copied to clipboard
| Challenge: | Prior work focused on predicting the immediate future of a story, such as one to a few sentences ahead. |
| Approach: | They propose a task that predicts the semantic frames that will occur in the next 10, 100, or even 1,000 sentences in a running story. |
| Outcome: | The proposed model outperforms random, prior, and replay baselines when the block size is over 150 sentences. |
Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language Varieties (2025.naacl-short)
Copied to clipboard
| Challenge: | Of the world's 7,000 languages, sixty (60) million people speak British English, 23 million speak Taiwan Mandarin, and 10 million speak European Portuguese. |
| Approach: | They propose a contextually aligned dataset that captures comments in different languages from real-world scenarios. |
| Outcome: | The proposed approach shows that large language models underperform in Taiwan Mandarin in a sentiment analysis task. |
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023 (2026.tacl-1)
Copied to clipboard
Ting-Yao Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu, Lun-Wei Ku, Clyde Lee Giles, Ting-Hao Huang
| Challenge: | SciCap dataset launched in 2021 aims to generate high-quality captions for scientific figures. |
| Approach: | They propose to use the SciCap dataset to develop models for captioning diverse figure types across various academic fields. |
| Outcome: | The proposed models showed impressive performance on the SciCap dataset and in various vision-and-language tasks. |
Visual Story Post-Editing (P19-1)
Copied to clipboard
| Challenge: | a dataset for human edits of machine-generated visual stories is released . it includes 14,905 human-edited versions of 2,981 machine- generated visual stories . |
| Approach: | They introduce the first dataset for human edits of machine-generated visual stories . they explore how edits may be used for the visual story post-editing task . |
| Outcome: | The proposed dataset includes 14,905 human-edited versions of 2,981 machine-generated visual stories. |