Papers by Dokyong Lee
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies focus on image-sharing behavior in singular sessions, leading to limited long-term social interaction. |
| Approach: | They propose a large-scale long-term multi-modal dialogue dataset that generates long-time multi-modity dialogue distilled from ChatGPT and proposed image aligner. |
| Outcome: | The proposed framework generates long-term multi-modal dialogue from ChatGPT and image aligner. |
Large Language Models can Share Images, Too! (2024.findings-acl)
Copied to clipboard
| Challenge: | Using a zero-shot prompting, large language models can be used to share images in a multi-tasking environment. |
| Approach: | They introduce a dataset that includes enriched annotations and a framework to evaluate LLMs. |
| Outcome: | The proposed framework unlocks image-sharing capability of LLMs in zero-shot prompting, with ChatGPT achieving the best performance. |