Show HN: JevBench, a reproducible benchmark for typed decision models
JevBench is a benchmark for evaluating typed decision models, which return bounded choices and probabilities instead of text, and are faster and cheaper than LLMs. The benchmark assesses accuracy, latency, and cost in a weighted way, with a leaderboard ranking various models. The current leader is Jev with a score of 74.4, followed by SemIf, djev, Winnow-12B Q8, and reflex 4B. The benchmark is restricted to English-only and uses a German server, with a latency adjustment for local and demo results.
Read the full article at benchmarkheaven.com →