Papers by Kosuke Arima
Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement (2026.acl-industry)
Copied to clipboard
Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma, Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Takuto Asakura, Chung-Chi Chen, Tatsuya Ishigaki
| Challenge: | Large language models (LLMs) make it easy to generate large numbers of product ideas. |
| Approach: | They propose to use a dataset of 3,000 individual scores across 300 patent-grounded product ideas to assess whether an automatic judge approximates an aggregate consensus. |
| Outcome: | The proposed model evaluators disagree on fine-grained ordinal scores, suggesting structured heterogeneity rather than random noise. |