Decision Bench
Can a small model make this decision for you?
An open benchmark of bounded decisions software asks language models to make, built from real records in open-licence datasets. Compare accuracy, latency, and cost across engineering, agents, safety, support, commerce, finance, legal, product, data, documents, and design.
Every row includes its source, licence, and answer key. Model results are published for inspection and reproduction.
Download benchmark rows · Explore dataset sources · Read the benchmark protocol