Papers by Shabnam Behzad
To Ask LLMs about English Grammaticality, Prompt Them in a Different Language (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a study focuses on questions about grammar and fluency in multilingual LLMs . english is the dominant training language for all three models, but prompting in a different language often yields better results. |
| Approach: | They ask three multilingual language models in multiple languages to test their model's grammatical accuracy. |
| Outcome: | The language of the prompt can significantly affect model performance, the study finds . english is the dominant training language for all three models, the researchers show . |
LEAF: Language Learners’ English Essays and Feedback Corpus (2024.naacl-short)
Copied to clipboard
| Challenge: | Current automated essay scoring models lack the granularity desired by learners and instructors seeking more detailed insights. |
| Approach: | They present a corpus of English essays and their corresponding feedback from the “essayforum” website. |
| Outcome: | The LEAF corpus provides valuable feedback for students and teachers . it provides insights on argumentative aspects and organizational coherence . |
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)
Copied to clipboard
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Levine, Jessica Lin, Devika Tiwari, Amir Zeldes
| Challenge: | Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old. |
| Approach: | They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing. |
| Outcome: | The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain. |
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)
Copied to clipboard
| Challenge: | ELQA corpus is metalinguistic—it consists of language about language. |
| Approach: | They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity . |
| Outcome: | The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models . |
MultiMUC: Multilingual Template Filling on MUC-4 (2024.eacl-long)
Copied to clipboard
William Gantt, Shabnam Behzad, Hannah An, Yunmo Chen, Aaron White, Benjamin Van Durme, Mahsa Yarmohammadi
| Challenge: | We present multilingual parallel template filling datasets for MUCs . systems were required to extract one template per incident, containing details about perpetrators, victims, weapons used . |
| Approach: | They introduce MultiMUC, the first multilingual parallel corpus for template filling . they obtain automatic translations from a strong multilingual machine translation system . |
| Outcome: | The proposed dataset includes translations of the classic MUC-4 template filling benchmark into Arabic, Chinese, Farsi, Korean, and Russian. |
Assessing Online Writing Feedback Resources: Generative AI vs. Good Samaritans (2024.lrec-main)
Copied to clipboard
| Challenge: | Providing constructive feedback on student essays presents significant challenges . large language models (LLMs) such as ChatGPT can facilitate this process . |
| Approach: | They compare essayforum.com and large language models such as ChatGPT for students . they argue that both can mutually reinforce each other and provide constructive feedback . |
| Outcome: | The findings highlight the potential of AI in advancing the field of automated essay evaluation. |
AMALGUM – A Free, Balanced, Multilayer English Web Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 4M tokens is available online with a large number of high-quality annotation layers. |
| Approach: | They propose to use a genre-balanced English web corpus with multiple annotation layers . they harness knowledge from multiple annotation layer to achieve a "better than NLP" benchmark . |
| Outcome: | The proposed corpus is genre-balanced and features high-quality automatic annotation layers. |