Hesperan

(Benchmarks)

Benchmarks

On decision tasks Hesperan 1 matches or beats Jev 1.13 on 8 of 12 public benchmarks. On knowledge-heavy tasks it is weaker — they are listed too.

Decision tasks — accuracy in %
BenchmarknHesperan 1Jev 1.13Δ vs JevBest other published
JevBench public23188.786.6
+2.1
87.0 Reflex-27B
JevBench public — hard11177.573.0
+4.5
75.7 Reflex-27B
jabr classifier v17898.797.4
+1.3
92.3 Von 1.0.1
jabr classifier v286696.296.7
0.5
68.9 GLiNER2
jabr support tickets50097.497.8
0.4
86.0 Von 1.0.1
Nimble public (13 sets, macro)3,88076.176.0
+0.1
77.4 Decider-35B
typed-decisions2,00073.072.7
+0.3
36.0 Laya
Kev transfer-v4 dev65685.585.7
0.2
81.2 Kev-9B
Kev transfer-v9 dev1,04679.985.4
5.5
OOD support tickets (EN/KO)90077.875.1
+2.7
Phishing (trifleen)10083.082.0
+1.0
BoolQ validation3,27091.891.6
+0.2
Knowledge-heavy tasks — accuracy in %
BenchmarknHesperan 1Jev 1.13Δ vs JevBest other published
OpenBookQA50091.894.2
2.4
CommonsenseQA1,22186.388.1
1.8
HellaSwag2,00084.286.1
1.9
MMLU-Pro (1k sample)1,00061.282.9
21.7

Hesperan 1 measured by us on 22 Sep 2026 on the public items of each benchmark, with the benchmark's own metric. All other numbers are the values published by the benchmark authors.

† The Hesperan 1 training mix contains the train split of this source (in-distribution).

No Jev outputs were used for training, distillation, labelling or model selection. Independent evaluation; not affiliated with or endorsed by TypeSafe AI. Jev and TypeSafe are trademarks of TypeSafe AI, Inc.

Try it yourself →How judgments work