Challenge: generative AI is expanding in education, yet empirical analyses of large-scale and real-world interactions between students and AI systems remain limited.
Approach: They present a dataset based on a semester-long experiment with 212 college students in English as Foreign Language (EFL) writing courses.
Outcome: The proposed dataset includes conversation logs, students’ intent, students' self-rated satisfaction, and students’ essay edit histories.

Similar Papers

Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are being used by end-users for various tasks, including sensitive ones such as health counseling, disregarding potential safety concerns.
Approach: They use ChatGPT to crowd-source dietary struggles and work with nutrition experts to generate supportive text using ChatGPS.
Outcome: The proposed model outperforms other models on dietary struggles and mental health tasks.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study validates the effectiveness of chat language models by fine-tuning instruction data.
Approach: They propose to use a large-scale dataset of instructional conversations to fine-tune a conversational model on instruction data.
Outcome: The proposed model outperforms open-source models in key metrics including scale, average length, diversity, coherence, etc.
Improved Instruction Ordering in Recipe-Grounded Conversation (2023.acl-long)

Copied to clipboard

Challenge: In this paper, we explore the task of instructional dialogue and focus on the cooking domain.
Approach: They propose to explore two auxiliary subtasks to support response generation with improved instruction grounding by incorporating user intent and instruction state information into the model.
Outcome: The proposed model lacks understanding of user intent and inability to track instruction state (i.e., which step was last instructed) incorporating user intent information helps the response generation model mitigate the incorrect order issue.
Cooking Up a Neural-based Model for Recipe Classification (2020.lrec-1)

Copied to clipboard

Challenge: a dataset of cooking recipes in French is highly imbalanced due to collaborative nature of the dataset . authors propose a neural-based model to address the first task of the DEFT 2013 shared task .
Approach: They propose a neural-based model to address the first task of the DEFT 2013 shared task . they use state-of-the-art embedding approaches and deep architectures to address imbalanced dataset .
Outcome: The proposed model outperforms models that use only pretrained embeddings in micro and macro F1 scores.
RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes (D18-1)

Copied to clipboard

Challenge: Existing comprehension tests for QA are limited by the text sources and questionanswer formats.
Approach: They propose a dataset for multimodal comprehension of cooking recipes . preliminary results indicate RecipeQA will serve as a challenging test bed .
Outcome: The proposed dataset will serve as a test bed and ideal benchmark for evaluating machine comprehension systems.
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)

Copied to clipboard

Challenge: OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown.
Approach: They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning.
Outcome: The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods.
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets (2023.findings-acl)

Copied to clipboard

Challenge: Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth.
Approach: They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets.
Outcome: The proposed model performs well on 140 tasks and generates 255K responses in these datasets.
Is ChatGPT a Good Multi-Party Conversation Solver? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are powerful tools for multi-party conversations, but their capacity to handle multi-parties remains unexplored.
Approach: They propose to evaluate ChatGPT and GPT-4's zero-shot learning capabilities within the context of multi-party conversations (MPCs) they also propose to incorporate MPC structures, encompassing both speaker and addressee architecture.
Outcome: The proposed models perform poorly on a number of MPC tasks while GPT-4 performs well on speaker and addressee architecture.
IntrEx: A Dataset for Modeling Engagement in Educational Conversations (2025.findings-emnlp)

Copied to clipboard

Challenge: IntrEx is the first large dataset annotated for interestingness and expected interestingness in teacher-student interactions.
Approach: They propose a large dataset annotated for interestingness and expected interestingness in teacher-student interactions.
Outcome: The proposed dataset is the first large dataset annotated for interestingness and expected interestingness in teacher-student interactions.
Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows (2022.coling-1)

Copied to clipboard

Challenge: a new dataset enables us to learn a cooking action result for each object in a recipe text.
Approach: They propose a multimodal dataset that enables us to learn a cooking action result for each object in a recipe text.
Outcome: The proposed dataset reduces human annotation costs by allowing multimodal information retrieval.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations