Uncensored weights.
Honest prices.
Curated uncensored builds of the best open models, behind one API key. Prepaid credits and per-token billing: no subscriptions, no expiry, no logs.
Why it’s this cheap
MoE math
2026 models are mixture-of-experts: huge total parameters, tiny active parameters. A 320B model that only activates 18B per token serves at small-model speed on commodity GPUs. We pass that arithmetic straight to you.
Cost-plus, in writing
Every price obeys one rule and it’s printed right here. If a build can’t meet it on the GPU it needs, it doesn’t go on the sheet.
price ≥ 2.5 × our cost @ 30% utilizationNo idle-GPU tax
We launch models on scale-to-zero serverless GPUs and only rent dedicated hardware when a model stays >40% busy for 30 days. You never pay for a card sitting idle, because we don’t either.
Curated, benchmarked, no refusals.
Two build methods, both labelled: abliterations (huihui-ai, OrcaRouter, our own Heretic runs) and refusal-free fine-tunes (Dolphin, Hermes). We measure refusal rate and capability retention against the base model before listing, and publish the numbers. Nothing sits between your prompt and the model: no policy gateway, no classifier.
See the benchmarks →Privacy-first
Requests are processed in memory and discarded. We keep four things: your account email, your credit ledger, the token counts of each request, and hashed keys. Content: zero. Not a policy promise — the schema has nowhere to put it.
Read the policy →Plugs into your agent
Qwen Code, Claude Code, OpenCode, Aider, Cline: if it can change a base URL, it works: streaming, tool calling and JSON mode included, on models that don’t refuse.
OPENAI_BASE_URL="https://api.cheapweights.ai/v1"