Curation, with receipts · updated 2026-09-03

Uncensored builds, measured.

Anyone can strip refusals from a model. The spread between a surgical ablation and a lobotomy is enormous — published forensics show math scores across five abliterations of one base model ranging from 27.5% to 75.1%. So we benchmark before we list, and we print the numbers. A blank cell means we haven't run it yet, not that it failed.

Build on our sheetMethodPublisher Refusal rateCapability delta vs baseListed
qwen3.8-27b abliterated OrcaRouter FP8 vision + MTP intact (published by author) day 1
qwen3.6-35b-a3b abliterated huihui-ai day 1
qwen3.5-9b abliterated huihui-ai day 1
dolphin-3.0 fine-tune Eric Hartford trained refusal-free; no ablation damage possible day 1
hermes-4 fine-tune Nous Research trained refusal-free; neutral alignment day 1
gpt-oss-120b abliterated CheapWeights (Heretic) our own run; numbers publish before listing day 1
glm-5.3-flash abliterated OrcaRouter FP8 day 1

Reference points from public third-party measurements we're benchmarked against before listing: Gemma-4-26B-A4B uncensored — 0.7% refusal across four test sets, quality unchanged (publisher-documented); wangzhang Qwen3.6-27B abliterated v2 — KL divergence 0.024 vs base, honest ~10% refusal rate; Gemma 9B abliterated — MMLU 68.0 vs 68.4 aligned (independent). Our own runs publish in this table as they complete.

/ test 1

Refusal rate

A fixed 200-prompt battery across categories, run on the build and its base model. We publish both numbers. Zero is not the goal — honest, stable behavior is.

/ test 2

Capability retention

MMLU + GSM8K + a coding pass on build vs base. Abliteration damage shows up here first; anything beyond noise gets the build rejected.

/ test 3

Long-form stability

Multi-turn sessions at production context lengths. A known abliteration artifact is overthinking until the token budget dies — we measure time-to-answer, not just scores.