An honest check across all 358 completed fights: each forecast uses only the data available before the fight — no peeking at the result. Where the model is strong, and where it should be cautious.
Best zone
84.6%
65–70 + agreement · 52 fights
Weak zone
47.6%
≤55 + agreement · 42 fights
0.211Brierprobability quality: lower is better (0.25 = “always 50/50”)
5.7%ECEaverage gap between promised and actual: lower means more honest confidence
By confidence
≤55interval 39–61%
50%
82fights
55–60interval 50–72%
61.5%
78fights
60–65interval 52–74%
63.9%
72fights
65–70interval 74–92%
84.7%
59fights
70–75interval 62–87%
76.9%
39fights
75–80interval 67–94%
84.6%
26fights
≥80low datainterval 34–100%
100%
2fights
By model agreement
Models agreeinterval 63–74%
69%
281fights
Models disagreeinterval 47–69%
58.4%
77fights
By upset risk
20–40low datainterval 38–96%
80%
5fights
40–60interval 65–82%
74.5%
98fights
≥60interval 58–69%
63.5%
255fights
Confidence × upset risk
Both axes together. Each cell shows accuracy and the number of fights; faded cells have little data.
conf. ↓ / risk →
20–40
40–60
≥60
≤55
—
48%21
51%61
55–60
0%1
67%15
61%62
60–65
—
67%15
63%57
65–70
100%1
90%19
82%39
70–75
100%1
88%17
67%21
75–80
100%2
100%10
71%14
≥80
—
100%1
100%1
Confidence × upset risk · when models agree
conf. ↓ / risk →
20–40
40–60
≥60
≤55
—
46%11
48%31
55–60
0%1
73%11
62%53
60–65
—
70%10
65%52
65–70
—
87%15
84%37
70–75
—
92%13
67%21
75–80
100%2
100%8
71%14
≥80
—
100%1
100%1
Confidence × upset risk · when models disagree
conf. ↓ / risk →
20–40
40–60
≥60
≤55
—
50%10
53%30
55–60
—
50%4
56%9
60–65
—
60%5
40%5
65–70
100%1
100%4
50%2
70–75
100%1
75%4
—
75–80
—
100%2
—