Papers by Ruixiang Cui
Compositional Generalization in Multilingual Semantic Parsing over Wikidata (2022.tacl-1)
Copied to clipboard
| Challenge: | Semantic parsers are mostly designed for and evaluated on English resources, such as CFQ. |
| Approach: | They propose a method for creating a multilingual, parallel question-query dataset . they analyze compositional generalization of parsers in Hebrew, Kannada, Chinese, and English . |
| Outcome: | The proposed method analyzes compositional generalization of parsers in Hebrew, Kannada, Chinese, and English. |
What does the Failure to Reason with “Respectively” in Zero/Few-Shot Settings Tell Us about Language Models? (2023.acl-long)
Copied to clipboard
| Challenge: | In the context of natural language inference, we examine how language models reason with respective readings from two perspectives: syntactic-semantic and commonsense-world knowledge. |
| Approach: | They propose a controlled synthetic dataset WikiResNLI and a naturally occurring dataset NatResLI to encompass various explicit and implicit realizations of "respectively". |
| Outcome: | The proposed datasets include explicit and implicit readings of "respectively" the proposed dataset shows that fine-tuned models struggle with understanding readings without explicit supervision. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models (2024.findings-naacl)
Copied to clipboard
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, Nan Duan
| Challenge: | Traditional benchmarks for evaluating foundation models often fail to accurately represent their general abilities for human-centric tasks. |
| Approach: | They propose a bilingual benchmark to assess foundation models in the context of human-centric standardized exams such as college entrance exams, law school admission tests, and math competitions. |
| Outcome: | The proposed benchmark exceeds the average human performance on SAT, LSAT, and math competitions with 95% accuracy and 92.5% on the Chinese college entrance English exam. |
How Conservative are Language Models? Adapting to the Introduction of Gender-Neutral Pronouns (2022.naacl-main)
Copied to clipboard
| Challenge: | a recent study shows that gender-neutral pronouns are not associated with processing difficulties . linguistic scholars have observed how technology has altered the course of language evolution . |
| Approach: | They show that gender-neutral pronouns in Danish, English and Swedish are not associated with processing difficulties. |
| Outcome: | a new study shows that gender-neutral pronouns are not associated with human processing difficulties . the findings suggest that such conservativity in language models may limit widespread adoption . |
Can AMR Assist Legal and Logical Reasoning? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) has been shown to be useful for many downstream tasks. |
| Approach: | They propose neural architectures that utilize linearised AMR graphs in combination with pre-trained language models to capture logical relationships on multiple choice question answering tasks. |
| Outcome: | The proposed models outperform text-only baselines but outperformed text models, suggesting complementary abilities. |
Generalized Quantifiers as a Source of Error in Multilingual NLU Benchmarks (2022.naacl-main)
Copied to clipboard
| Challenge: | Quantifiers are pervasive in NLU benchmarks and their occurrence at test time is associated with performance drops. |
| Approach: | They propose a generalized quantifier NLI task to quantify their contribution to the errors of NLU models. |
| Outcome: | The proposed model is based on a generalized quantifier theory and is compared with pre-trained models. |