Rank #24 of 54
Mistral
Large 3 675B Instruct
via AWS Bedrock + Ollama Cloud
mistral-large-3-675b-instruct
Failure rate
26%
142 of 550 decided probes failed.
Coverage
100%
Share of probes the judges actually decided.
Spread across passes
4.7 pts
Highest minus lowest pass rate over 171 repeated probes.
Compliance / behaviour
67% / 79%
Cost per pass
—
255 avg reply tokens, —s.
Failure rate by category
Lower is better. The count beside each bar is decided/total probes for that category.
Expired-figure failure
61%36/36
Hallucinated-fact failure
30%30/30
Product-recommendation failure
6%64/64
Outcome-promise failure
39%36/36
Missing-caveat failure
7%46/46
Referenceability failure
24%34/34
Completeness-gap failure
76%38/38
Bias-exploitation failure
0%32/32
Emotion-manipulation failure
0%32/32
Understanding-check failure
0%36/36
Information-overload failure
97%36/36
Missing-friction failure
0%34/34
Vulnerability-tailoring failure
0%32/32
Inappropriate-urgency failure
16%32/32
Naming a bias helpfully
41%32/32