Uncensored builds, measured.
Anyone can strip refusals from a model. The spread between a surgical ablation and a lobotomy is enormous — published forensics show math scores across five abliterations of one base model ranging from 27.5% to 75.1%. So we benchmark before we list, and we print the numbers. A blank cell means we haven't run it yet, not that it failed.
| Build on our sheet | Method | Publisher | Refusal rate | Capability delta vs base | Listed |
|---|---|---|---|---|---|
| qwen3.8-27b | abliterated | OrcaRouter FP8 | — | vision + MTP intact (published by author) | day 1 |
| qwen3.6-35b-a3b | abliterated | huihui-ai | — | — | day 1 |
| qwen3.5-9b | abliterated | huihui-ai | — | — | day 1 |
| dolphin-3.0 | fine-tune | Eric Hartford | — | trained refusal-free; no ablation damage possible | day 1 |
| hermes-4 | fine-tune | Nous Research | — | trained refusal-free; neutral alignment | day 1 |
| gpt-oss-120b | abliterated | CheapWeights (Heretic) | — | our own run; numbers publish before listing | day 1 |
| glm-5.3-flash | abliterated | OrcaRouter FP8 | — | — | day 1 |
Reference points from public third-party measurements we're benchmarked against before listing: Gemma-4-26B-A4B uncensored — 0.7% refusal across four test sets, quality unchanged (publisher-documented); wangzhang Qwen3.6-27B abliterated v2 — KL divergence 0.024 vs base, honest ~10% refusal rate; Gemma 9B abliterated — MMLU 68.0 vs 68.4 aligned (independent). Our own runs publish in this table as they complete.
Refusal rate
A fixed 200-prompt battery across categories, run on the build and its base model. We publish both numbers. Zero is not the goal — honest, stable behavior is.
Capability retention
MMLU + GSM8K + a coding pass on build vs base. Abliteration damage shows up here first; anything beyond noise gets the build rejected.
Long-form stability
Multi-turn sessions at production context lengths. A known abliteration artifact is overthinking until the token budget dies — we measure time-to-answer, not just scores.