Papers by Ivan Koychev

6 papers
Fact-Checking Meets Fauxtography: Verifying Claims About Images (D19-1)

Copied to clipboard

Challenge: Recent explosion of false claims in social media has led to manual fact-checking initiatives . however, existing methods are inadequate to deal with the growing number of false content claims.
Approach: They propose to model claims about images using a new dataset to examine the relationship between the image and the claim.
Outcome: The proposed method improves on the baseline and will enable future research on fact-checking claims about images.
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks emphasize final numerical answers while neglecting intermediate reasoning steps.
Approach: They propose a symbolic benchmark for verifiable Chain-of-Thought evaluation in finance . FINCHAIN spans 58 topics across 12 financial domains and three difficulty levels .
Outcome: The proposed benchmark aims to bridge symbolic reasoning and factual verification.
bgGLUE: A Bulgarian General Language Understanding Evaluation Benchmark (2023.acl-long)

Copied to clipboard

Challenge: bgGLUE is a benchmark for evaluating language models on natural language understanding (NLU) tasks in Bulgarian.
Approach: They propose to use a benchmark to evaluate language models on NLU tasks in Bulgarian.
Outcome: The proposed model performs well on sequence labeling tasks, but there is room for improvement for tasks that require more complex reasoning.
EXAMS: A Multi-subject High School Examinations Dataset for Cross-lingual and Multilingual Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: EXAMS is a benchmark dataset for cross-lingual and multilingual question answering for high school examinations.
Approach: They propose to use EXAMS to evaluate cross-lingual and multilingual question answering for high school examinations.
Outcome: The proposed model can be used to explore multilingual reasoning and knowledge transfer methods and pre-trained models in schools in different languages, which was not possible by now.
CrowdChecked: Detecting Previously Fact-Checked Claims in Social Media (2022.aacl-main)

Copied to clipboard

Challenge: Existing systems to automate fact-checking lack credibility in the eyes of the users.
Approach: They propose to perform automatic fact-checking by verifying whether an input claim has been fact- checked by professional fact- checkers and to return back an article that explains their decision.
Outcome: The proposed method improves on the CLEF’21 CheckThat! test set by two points absolute.
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for vision language models are outdated and unable to accurately assess their performance.
Approach: They propose a multi-discipline multimodal multilingual exam benchmark for vision language models . they collect multiple-choice questions across 20 disciplines across 11 languages from 7 language families .
Outcome: The EXAMS-V exam includes 20,932 multiple-choice questions across 20 disciplines . the questions come in 11 languages from 7 language families and require advanced reasoning skills .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations