Sovereign
Uncensored. Accelerated.
Abliterated weights. Splash acceleration. Compact Repair.
Brought together for Apple Silicon.
Open weights for code, tools, and answers your software can use. Sovereign combines an uncensored Qwen foundation with fast generation, checked editing workflows, and ready-to-use profiles.
| Fix the code | Generate less to change less | Keep code moving |
|---|---|---|
| 36/36 repairs passed | 79.2% fewer output tokens | Up to 130 tokens/s observed |
| Compact Repair on 36 fresh synthetic Python tasks, with executable checks. | 2,303 versus 11,076 output tokens on those same tasks. No retry was needed; all responses are counted. | A 16,020-token structured-output run reached 130.3 tokens/s at 128K context. A separate Python module reached 114.8 source tokens/s and passed 20 checks. |
These are author-run measurements with published outputs and protocols, not a model ranking or guaranteed throughput. The repair gain comes from the optional editing workflow, not newly trained weights. Repair comparison · Python speed series.
Get started · API examples · Inspect the evaluations
Long work, checked at every step
On an M4 Max with 48 GB unified memory, the same Splash Q4 weights ran with a 262,144-token runtime context. Starting from 16,123 input tokens, twelve stateful continuations produced 201,677 output tokens in one checkpointed synthetic task. An independent pass verified all 9,216/9,216 consecutively numbered NDJSON records and three planted codes.
| Measured long-work result | Observed value |
|---|---|
| Output across 12 checked continuations | 201,677 tokens |
| Last turn: accumulated input → output | 202,277 → 16,896 tokens |
| Decode speed from first to last turn | 68.1 → 50.9 tokens/s |
| Whole-run output rate, including prefill | 38.0 tokens/s over 88.4 minutes |
This is a 12-response workflow, not a one-shot 201K-token completion. It used deterministic structured output, so it does not establish general code, prose, or reasoning quality at that length. For extended jobs, request bounded segments of about 16K–17K output tokens with a 20K per-segment cap, save each response, and validate before continuing via LM Studio's stateful chat API. The 256K measurements used a temporary runtime override; the later named profile uses the same weights and context setting. 64K context / 16K output remains the practical default. 128K/256K measurements, protocol and limitations.
Download
Native Splash Q4 package (about 17.38 GB). Download with the Hugging Face CLI:
hf download Proofofvitalik/Sovereign-Uncensored-Splash-Q4 --local-dir Sovereign-Uncensored-Splash-Q4
Then follow the LM Studio setup or standalone instructions. These weights require Splash; they are not an MLX or GGUF checkpoint.
What Sovereign adds
A direct path from ARA to native Q4. The target is quantized once from Heretic's BF16 checkpoint to affine Q4/group64, then packed for Splash without another quantization pass. Source revisions, conversion code, and artifact hashes make the build traceable.
Profiles that turn supported features into working tasks. The included extraction recipe combines a precise selection policy with JSON Schema. The tool workflow recipe covers both actions and decisions to stop or ask for missing information. Fast and Low let you choose how much reasoning to use. Workflow recipes and recorded examples.
Optional Python validation. Sovereign Code Guard checks syntax and isolated module imports, then allows one corrective generation when needed. It saves the original response, correction, token usage and latency. A four-case development trial finished with 4/4 modules passing 46 external functional checks after one repair; that repair added 11.85 seconds. CLI, scope and evidence.
Smaller edits, checked before use. Compact Repair asks for precise changes instead of rewriting the file. It preserves the original, rejects ambiguous or stale edits, and checks the candidate before saving it. On 36 fresh synthetic repairs it passed 36/36 supplied test suites, using 79.2% fewer output tokens and 41.8% less measured wall time than full-file generation. Including prompt tokens, the reduction was 20.5%. On the 12 longer sources, output savings reached 86.2%. Both modes used one request per task, and all responses are counted. These results establish workflow efficiency, not a coding-ability advantage. CLI and usage · All 36 cases, raw outputs and method.
One package for chat, images, and longer sessions. Original BF16 vision weights and the stock Inco DFlash2 draft accompany the target. LM Studio load entries and generation presets cover everyday use, a tested 64K configuration, and experimental 128K and 256K configurations on 48 GB hardware.
The 17.38 GB native package uses Qwen3.8-27B, adapted by Heretic ARA. ARA means Arbitrary-Rank Ablation, the upstream method used to reduce refusals. “Uncensored” identifies that lineage; it is not a guarantee about every response. Inco supplies the acceleration and structured-output engine; Sovereign's contribution is the conversion, integration, editing and validation tools, workflow recipes, and their measured behavior.
Choose your profile
| Load profile + generation preset | Context | Output cap | Completed example on M4 Max / 48 GB |
|---|---|---|---|
| Everyday Fast / Low | 8,192 | 2,048 | Short chat and tool requests |
| Long 64K Fast / Low | 65,536 | 8,192 / 16,384 | Tested long-input and 16K-output workflows |
| Long 128K — 32K Fast | 131,072 | 32,768 | 92,113 input → 32,717 output; 1,536/1,536 exact records |
| Long 256K — 16K Fast | 262,144 | 16,384 | 135,103 input → 16,052 output; 768/768 exact records |
| Long 256K — 32K Fast | 262,144 | 32,768 | 80,127 input → 32,717 output; 1,536/1,536 exact records |
Output caps include reasoning. The 128K and 256K rows are experimental structured-output examples, not general long-context quality or a promise that every request can reach its cap. Each measured pair is one completed request; a separate 180K-input/32K-output attempt timed out. Context must cover input, history, system prompt, tools and generated tokens. A generation preset alone does not enlarge the loaded context, and all Long entries use the same weights. For practical long work, use checked continuations of about 16K–17K output tokens; 64K context / 16K output remains the default recommendation. Install and select the 64K, 128K and 256K LM Studio profiles.
Tested setup: LM Studio 0.4.25+1 with official Splash runtime 0.0.5, or standalone Splash 1.0.1; M4 Max / 48 GB / macOS 26.6.2. The integration provides Sovereign Uncensored Native and same-weights Long 64k, Long 128k and Long 256k load definitions. The 256K measurements used a temporary override of the 128K entry; the new named 256K definition reproduces that setting but has not been separately run. Sovereign is the display brand; existing runtime identifiers remain compatible. Installation and presets.
The native package loads through Splash. The separate MLX companion uses a different file format; these packed files are not GGUF or Transformers checkpoints.
Speed you can inspect
The table makes each result's workload explicit. Rates divide output by full request time with the model already loaded; unlike task types are not directly comparable.
| Observed run | Output rate | Result and scope |
|---|---|---|
| Structured NDJSON, 16K output, 128K context | 130.3 tokens/s | 16,020 tokens; 768/768 records exact; repetitive synthetic output |
| Python, single Fast example | 114.8 source tokens/s | 2,131 tokens in 18.56 s; 20/20 functional checks |
| Python, single Greedy example | 109.8 source tokens/s | 2,039 tokens in 18.57 s; 20/20 functional checks |
| Structured NDJSON, 32K output cap, 128K context | 109.7 tokens/s | 32,685 tokens in 298.05 s; 1,536/1,536 records exact |
| Initial 18-request Python series | 72.47 source tokens/s average | 46,867 tokens in 646.72 s; range 55.44–114.79 |
| Checkpointed structured task, 256K context | 38.0 tokens/s overall | 201,677 tokens across 12 responses in 88.4 min; includes repeated prefill; 9,216/9,216 records exact |
The Python examples are executable source without Markdown wrappers. The structured-output examples are formatting stress tests, not code or prose-quality results. The initial series had three repeats per task/profile; later repeats were slower. Its six-request 8K reload follow-up averaged 68.66 source tokens/s and does not isolate context size from time or machine state. These are observed examples, not sustained throughput.
For the Python rows: Thinking off, 6,144-token output cap; Fast temperature 0.6 / top-p 0.95 / top-k 20, or greedy temperature 0. M4 Max, 40 GPU cores, 48 GB; LM Studio + Splash 0.0.5. Exact package tokenization counts code, docstrings and generated unit tests, excluding special tokens and any enclosing Markdown fence; the full original request time remains in the denominator. Functional results and raw Python formatting are recorded separately. All 24 Python requests, checks, and method · Long-run protocol.
Earlier short-request speed samples
| Request | Output tokens | Full request time | Output / full time |
|---|---|---|---|
| Russian explanation | 469 | 11.64 s | 40.30 tok/s |
| English explanation | 386 | 8.00 s | 48.27 tok/s |
| Python LRU cache | 241 | 2.49 s | 96.73 tok/s |
| All three | 1,096 | 22.13 s | 49.53 tok/s |
Fast mode, temperature 0.6, top-p 0.95, top-k 20, 2,048-token output cap. One request per workload, already-loaded model, active desktop, no fixed seed or cache reset. First-token latency was 0.65–0.92 seconds. End-to-end time includes request processing and generation, but not model loading. These samples measure completion and speed, not coding correctness or sustained throughput. Raw measurements.
What we checked
Tools: complete the workflow. The 20 episodes include lookups, dependent calls, missing-input clarification, cancellation, bounded retries, and conflict recovery. In the version-conflict example, Sovereign reads the configuration, attempts an update, reads the new version after a conflict, and retries with the new token. These are simulated tools, including episodes where the correct behavior is to make no call.
Structured records: get the answer right, too. In 32 synthetic English/Russian cases, Sovereign selected IDs where active=true and quantity met a threshold. With the fixed policy and JSON Schema, 32/32 answers and 32/32 JSON outputs passed. This is record filtering, not a general document-extraction benchmark.
Selected completed workflow checks
| Check | Result | What was measured |
|---|---|---|
| Compact Python repairs | 36/36 | Fresh synthetic repair tasks with executable checks; optional editing workflow |
| Native tool workflows | 20/20 | Low-mode development episodes using simulated tools |
| Structured record filtering | 32/32 | Synthetic cases using a fixed policy and JSON Schema |
| Code Guard validation | 4/4 modules; 46 functional checks | Four-case development trial, including one corrective generation |
These are selected author-run development checks, not an overall accuracy score. The tool episodes used mock tools; no external actions were performed. The complete evaluation ledger retains every outcome, protocol and raw response.
Long-context capacity on M4 Max
At the saved 65,536-token profile, the model recovered three planted facts from 54,854 input tokens; an explicit-schema follow-up on the same document passed strict JSON at 54,917 input tokens. These cold requests took about five to six minutes.
Experimental overrides also loaded the same weights at 131,072 and 262,144 context tokens. Synthetic requests combined long retrieval input with checked structured output:
| Runtime context | Input → output tokens | Exact result |
|---|---|---|
| 128K | 100,119 → 16,052 | Three planted codes; 768/768 records |
| 128K | 92,113 → 32,717 | Three planted codes; 1,536/1,536 records |
| 256K | 80,127 → 32,717 | Three planted codes; 1,536/1,536 records |
| 256K | 135,103 → 16,052 | Three planted codes; 768/768 records |
| 256K | 180,099 → 51 | Three planted codes; one requested record |
| 256K, stateful last turn | 202,277 → 16,896 | Continuation of the 12-response task; cumulative history, not a fresh 202K-token document |
The table lists completed checks only. In the twelve-stage run, the lowest sampled free memory was 20%, memory pressure reached level 2, and swap did not grow. These fixtures test capacity and exact formatting, not general full-window understanding. The included 128K and 256K definitions point to the same weights; the 256K tests used a temporary override, and the named definition has not been separately run. 64K capacity report · 128K/256K and stateful continuation summary.
Practical notes
Use Fast for short replies and allow more output for longer code or reasoning. Low can consume a 2,048-token allowance before producing a final answer, as observed in two explanation requests. A valid schema does not guarantee a correct answer. Applications should validate results and manage tool execution, retries, and duplicate actions. The 64K profile expands the tested configuration, not the model's underlying architecture.
Scope & limitations
- What Sovereign adds. This release packages upstream weights with conversion, runtime integration, profiles, and optional editing tools. It does not claim a new foundation model or newly fine-tuned release weights. Compact Repair runs outside the model and must be invoked separately.
- What “uncensored” means. The name identifies the ARA abliteration lineage. It is not a promise of zero refusals, factual accuracy, or unrestricted behaviour on every request.
- What the numbers cover. Reported results come from the documented hardware, prompts, settings, and small evaluation suites. They are not a frontier-model comparison, a family-wide ranking, or a guarantee of sustained speed. Synthetic tasks and supplied tests do not establish general repository-level coding quality.
- What the checks guarantee. Passing syntax, import, or supplied functional tests establishes only those checks. Review generated code and validate tool arguments before granting access or executing actions.
- Compatibility and provenance. Native Splash and MLX are separate formats. The 64K, 128K and 256K load definitions use the same Splash Q4 weights. The 256K tests used a temporary override; the newly packaged 256K definition itself has not been separately tested. Qwen, Heretic, Inco, LM Studio, and Apple are credited components or platforms; no affiliation or endorsement is implied. Their applicable licences remain in force.
The yellow-and-black identity draws on Rothbardian themes of self-ownership and voluntary exchange. It describes the project's visual identity, not a claim about the model's beliefs or political fine-tuning.
Built on, packaged by
Qwen created the base model. Heretic supplied the ARA target. Inco AI supplied Splash and the stock DFlash2 draft. Proofofvitalik prepared the native conversion, packaging, LM Studio integration, presets, and evaluations.
- Target and vision: heretic-org/Qwen3.8-27B-heretic-ara, revision
2dc9b364104881cbb85e390f00195ba6b9d745e9. - Stock packed draft: incoai/Qwen3.8-27B-Splash, revision
cac1885f0f7bf90e10e1c57b7e4af0433b3f1195. - Engine: Inco AI / Splash. Runtime distributed separately.
- Integrity: manifest, converters, and 78/78 runtime artifact hashes verified in the recorded installation audit.
Apache-2.0 model package; component and runtime terms remain in their respective notices. License · Notices.