Papers by Mathieu Sibue

6 papers
“What is the value of templates?” Rethinking Document Information Extraction Datasets for LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing work on prompt-response datasets for visually rich document understanding (VRDU) is labor-intensive.
Approach: They propose a set of questions that are transformed from a key information extraction template to a prompt-response format using a plethora of bespoke templates.
Outcome: The proposed datasets are compared with baseline models on K2Q with zero-shot prompting.
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding (2024.acl-long)

Copied to clipboard

Challenge: Documents with rich layouts are a significant portion of enterprise corpora and document AI is still a challenge.
Approach: They propose a lightweight extension to traditional large language models for reasoning over visual documents that takes into account both textual semantics and spatial layout.
Outcome: The proposed model outperforms existing large language models on 14 out of 16 datasets and generalizes well to 4 out of 5 previously unseen datasets.
AfroCS-xs: Creating a Compact, High-Quality, Human-Validated Code-Switched Dataset for African Languages (2025.acl-long)

Copied to clipboard

Challenge: AfroCS-xs is a low-quality dataset for code-switching in multilingual communities . code-witching is prevalent in multicultural societies but lacks high-quality data for model development .
Approach: They propose to use human-validated synthetic code-switched datasets to generate code-witched sentences for four African languages and English within a specific domain—agriculture.
Outcome: The proposed model improves translation accuracy on the high-quality dataset for four African languages and English within a specific domain—agriculture.
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images (2026.eacl-long)

Copied to clipboard

Challenge: Existing models for structured information extraction are limited by narrow entity ontologies, simple queries, or homogeneous document types.
Approach: They propose a benchmark dataset for structured Information Extraction (IE) from document images . they analyze open and closed VLMs on this benchmark .
Outcome: The proposed model can perform fine-grained structured extraction across document types and schemas.
The State of the Art of Large Language Models on Chartered Financial Analyst Exams (2024.emnlp-industry)

Copied to clipboard

Challenge: Chartered Financial Analyst (CFA) program is widely recognized globally . study compares state-of-the-art large language models with open-source models . proprietary models pass levels I and II, but fail at level III due to essay questions .
Approach: They benchmark five leading proprietary models and eight open-source models on mock CFA exams to provide an overview of their financial analysis capabilities.
Outcome: The models on the mock CFA exams pass the highest scores, but fail at the lowest levels due to essay questions.
Advanced Messaging Platform (AMP): Pipeline for Automated Enterprise Email Processing (2025.acl-industry)

Copied to clipboard

Challenge: a lack of publicly available datasets for training and benchmarking limits current AI techniques' effectiveness in industry-specific applications.
Approach: They propose an email automation pipeline that automates email response generation at scale in real-world enterprise settings.
Outcome: The proposed pipeline automates email response generation at scale in real-world environments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations