Arctic-TILT. Business Document Understanding at Sub-Billion Scale (2025.acl-industry)
Copied to clipboard
Łukasz Borchmann, Michał Pietruszka, Wojciech Jaśkowski, Dawid Jurkiewicz, Piotr Halama, Paweł Józiak, Łukasz Garncarek, Paweł Liskowski, Karolina Szyndler, Andrzej Gretkowski, Julita Ołtusek, Gabriela Nowakowska, Artur Zawłocki, Łukasz Duhr, Paweł Dyda, Michał Turski
| Challenge: | General-purpose LLMs and their multimodal counterparts provide a crucial advantage in process automation. |
| Approach: | They propose a model that can be finetuned and deployed on a single 24GB GPU . it provides reliable confidence scores and quick inferences for processing files in large-scale or time-sensitive environments. |
| Outcome: | The proposed model achieves state-of-the-art results on seven diverse benchmarks and provides reliable confidence scores and quick inferences. |
Similar Papers
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? (2025.emnlp-main)
Copied to clipboard
An-Lan Wang, Jingqun Tang, Lei Liao, Hao Feng, Qi Liu, Xiang Fei, Jinghui Lu, Han Wang, Hao Liu, Yuliang Liu, Xiang Bai, Can Huang
| Challenge: | Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world . |
| Approach: | They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions. |
| Outcome: | The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents. |
SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks have highlighted character-level tasks as lacking practical relevance . many real-world applications rely heavily on precise sub-token understanding . |
| Approach: | They propose a benchmark that assesses sub-token understanding through practical tasks . they examine the impact of test-time scaling on sub-word reasoning . |
| Outcome: | The proposed benchmark assesses sub-token understanding through practical tasks . it includes ten tasks across four domains and isolates tokenization-related failures . |
GroundCocoa: A Benchmark for Evaluating Compositional & Conditional Reasoning in Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing LLMs excel and often surpass human performance on benchmarks, but they are known to falter in simple tasks and under seemingly straightforward circumstances. |
| Approach: | They propose a benchmark to assess compositional and conditional reasoning within a flight booking task. |
| Outcome: | The proposed model outperforms existing models on the flight booking task with a 67% accuracy rate. |
SkyLLM: Cross-LLM-APIs Federation for Cost-effective Query Processing (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated exceptional capabilities across a wide range of tasks, from text generation to complex problem-solving. |
| Approach: | They propose a system which federates multiple LLM APIs and dynamically assigns a non-empty subset of these APIs to each query prior to inference. |
| Outcome: | The proposed system can match the most accurate LLM with the lowest cost while cutting costs by 67.8%. |
SAJA: A Simple Approach to Judge Alignment for LLM-as-a-Judge (2026.acl-industry)
Copied to clipboard
| Challenge: | Current approaches to evaluate text at scale require multiple calls and per-dataset prompt tuning. |
| Approach: | They propose a model-agnostic approach to evaluate judge alignment that uses a lightweight calibration head. |
| Outcome: | a new model with SAJA matches more complex systems across four evaluation paradigms . it outperforms uncalibrated models on MT-Bench pairwise preference and competitive performance on five classification benchmarks compared to uncalibred models . |
TAIL: A Toolkit for Automatic and Realistic Long-Context Large Language Model Evaluation (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Existing evaluation methods for long-context large language models are overly simplistic and require extensive human annotations. |
| Approach: | They propose an automatic toolkit to create realistic evaluation benchmarks . they use a document-grounded benchmark to generate question-answer pairs . |
| Outcome: | The proposed toolkit provides a way to create realistic evaluation benchmarks and visualize performance metrics of evaluated models. |
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction (2024.acl-long)
Copied to clipboard
Zixuan Li, Yutao Zeng, Yuxin Zuo, Weicheng Ren, Wenxuan Liu, Miao Su, Yucan Guo, Yantao Liu, Lixiang Lixiang, Zhilei Hu, Long Bai, Wei Li, Yidan Liu, Pan Yang, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
| Challenge: | None. None.. None! |
| Approach: | None. None.. None! |
| Outcome: | None. None. No. : |
Attribute or Abstain: Large Language Models as Long Document Assistants (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to attribution have only been evaluated in RAG settings, where initial retrieval confounds performance. |
| Approach: | They propose to use a benchmark to evaluate attribution on long document tasks . they find that citations and additional retrieval perform best for large models . |
| Outcome: | The proposed approach performs best on large and fine-tuned models, while additional retrieval can help for small, prompted models. |
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for extracting structured procedural knowledge from unstructured business documents are limited by simplistic schemas and shallow logical dependencies. |
| Approach: | They propose a framework for extracting structured procedural knowledge from unstructured business documents . they propose BREX, a carefully curated benchmark comprising 409 real-world business documents and 2,855 expert-annotated rules . |
| Outcome: | The proposed framework outperforms standard prompts in rule extraction and execution. |
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations for Structured Knowledge (SK) understanding are non-rigorous and focus on a single type of SK. |
| Approach: | They propose a structured knowledge understanding benchmark that includes four widely used structured knowledge forms. |
| Outcome: | The proposed benchmark is based on four widely used structured knowledge forms . it includes a question, an answer, positive knowledge units, and noisy knowledge units . |