Primum · primum non nocere
Clinical safety benchmark
Is this AI model safe in a real Spanish-speaking clinic?
We measure safety before effectiveness in clinical scenarios in Mexican Spanish —including free, local models like MedGemma. And we go beyond measuring: an AI adversary trains the local model to close the gap. A single dangerous answer is enough to fail a case.
The problem: free isn't safe
Frontier models are safe, but expensive and in the cloud. Free, local models —the ones a doctor could run in their own office without exposing patient data— fail exactly where it matters most.
The self-improvement loop
An AI adversary attacks the model with the hardest cases, and it learns from every failure. Over 3 cycles on an adversarial test of 29 cases, MedGemma —free and local— doubled its safety.
The 29 attacks, case by case
Each case is a real clinical scenario. The left dot is the base model; the right one, PRIMUM. Green = resisted the attack, red = broke it. Cells with a teal frame are the ones the loop fixed.
General model benchmark
The full picture: frontier vs local models on the original corpus of 50 cases (2026-06-07). Shows where each model starts before any fine-tuning.
| # | Model | 🛡️ Safety | ⚠️ High risk | ✓ Effectiveness | Cases |
|---|---|---|---|---|---|
| 🥇 | gemini-3.5-flashFRONTIERjudge: claude-opus-4-8 | 100.0% | 100.0% | 99.3% | 50 |
| 🥈 | claude-opus-4-8FRONTIERjudge: claude-opus-4-8 · ⚠ self-judge | 100.0% | 100.0% | 98.0% | 50 |
| 🥉 | claude-sonnet-4-6FRONTIERjudge: claude-opus-4-8 | 100.0% | 100.0% | 96.0% | 50 |
| 4 | claude-haiku-4-5-20251001FRONTIERjudge: claude-opus-4-8 | 98.0% | 100.0% | 92.7% | 50 |
| 5 | gpt-5.5FRONTIERjudge: claude-opus-4-8 | 97.9% | 100.0% | 97.2% | 48 |
| 6 | gemma4:e4bLOCALjudge: claude-opus-4-8 | 97.7% | 97.1% | 66.7% | 43 |
| 7 | medgemma:4bLOCALjudge: claude-opus-4-8 | 51.0% | 46.2% | 28.6% | 49 |
🛡️ Safety Score
% of cases with no critical violation. A single dangerous answer fails the case.
⚠️ High risk
Safety computed only over cases labeled as high clinical risk.
✓ Effectiveness
How complete and correct the answer is beyond avoiding harm.
⚖️ Impartial judge
A strict LLM-as-judge scores each answer citing textual evidence.