Papers with o3-mini

3 papers
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing medical benchmarks for diagnostic reasoning are limited in their ability to perform complex tasks.
Approach: They propose to benchmark diagnostic capabilities of large language models to assess their accuracy and generalization bottlenecks.
Outcome: The proposed model achieves 45.82%, 31.09%, and 17.79% accuracy, compared to current models, o3-mini, e1 and DeepSeek-R1 .
Stronger Universal and Transferable Attacks by Suppressing Refusals (2025.naacl-long)

Copied to clipboard

Challenge: Efforts have focused on aligning models to human preferences (RLHF) . yet, it is believed that such optimization-based attacks are sample-specific.
Approach: They propose an algorithm to embed a "safety feature" into models to make them safe for mass deployment.
Outcome: The proposed attack achieves 25% success rate against the state-of-the-art Circuit Breaker defense, compared to 2.5% by white-box GCG.
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts (2025.findings-emnlp)

Copied to clipboard

Challenge: ScholarBench evaluates domain-specific knowledge of large language models (LLMs) prior benchmarks lack the scalability to handle complex academic tasks.
Approach: ScholarBench evaluates the academic reasoning ability of large language models . the benchmark is constructed through a three-step process .
Outcome: ScholarBench evaluates the academic reasoning ability of large language models . the benchmark comprises 5,031 examples in Korean and 5,309 examples in English .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations