Call for testers: Janas, a CPU-only engine that runs MoE models larger than RAM — one command to test your machine

Hi everyone,

I’m looking for people willing to spend half an hour (or a quiet evening) running a test of Janas-LLM on their own Linux machine.

What it is. Janas-LLM is an inference engine written in C, GPL-3.0, built for ordinary machines: no GPU needed, and Mixture-of-Experts models larger than the RAM, with the experts streamed from the SSD as they are needed. On my laptop (Core Ultra 9 185H, 32 GB, NVMe) Qwen3-Next-80B-A3B — a 48 GB file — writes about 23 tokens/s, 30 with its multi-token prediction block, and it has read a full 262,144-token context and still found three facts hidden in it. It runs the Qwen3 family (dense and MoE, Qwen3.5/3.6, Qwen3-Next), has a terminal chat, an OpenAI-compatible HTTP server (chat, Responses, embeddings, tools) and MCP support.

Why I need you. The engine tunes itself from what it measures — threads, GPU, how much of each expert to read — but it has only ever measured one machine. Every other CPU, disk and amount of memory is a guess until someone runs it. Most wanted: AMD CPUs, Intel without AVX-VNNI, machines with 16 GB or less, SATA SSDs, distributions other than Debian.

How: two commands.

git clone https://github.com/prabanta-dev/janas && cd janas
./tools/janas-try.sh

You need Linux on x86-64, a C compiler (gcc or clang) and curl — nothing else to install, and no root. The script:

  1. builds Janas and runs its tests;
  2. downloads a model from its official Hugging Face repository and checks its SHA-256;
  3. converts it to Janas’s own format and checks the result against published fingerprints (the conversion is deterministic, so any difference is a bug found);
  4. checks that the model’s answer at temperature 0 is the very same text every other machine gets (the arithmetic is bit-identical across thread counts and instruction sets — your machine confirms it or finds where it is not);
  5. measures the speed, after checking that the machine is idle;
  6. tries the HTTP server;
  7. writes a report and shows it to you; only if you say yes does it open a GitHub issue (with gh, or a pre-filled link).

It asks before every long step, and at the end deletes only what it downloaded, if you want. The report holds the CPU, memory, disk model, distribution, compiler and the results — no host name, user name, paths, or anything about what else runs on your machine.

Three levels, and it offers only those your machine can hold:

Level Model Download Disk while it works
quick Qwen3-4B 2.5 GB ~5 GB
medium Qwen3.6-35B-A3B 22 GB ~45 GB
full Qwen3-Next-80B-A3B + MTP ~52 GB ~100 GB

On a machine like mine (the laptop above, on a fast connection) quick takes about fifteen minutes; the others take longer, mostly downloading, and a slower CPU, disk or connection stretches every step.

Please measure on an idle machine: close the browser, builds and other models first. The script checks, and the report says how idle the CPU was — speeds taken on a busy machine aren’t that machine’s.

Here is what a report looks like: https://github.com/prabanta-dev/janas/issues/8

The models are the Qwen team’s, under Apache 2.0; Janas converts them locally and never redistributes them (the details are in MODELS.md).

Every report will be read, and what it shows goes back into the engine: a failure on your machine is exactly what I’m looking for. Questions and problems are welcome here or as issues on the repository.

Thank you!

Maurizio “camauri” Cammalleri
https://github.com/prabanta-dev/janas

I don’t have a physical Linux machine, so I tried it on Colab for now:


I ran the quick test on a CPU-only Colab VM, and it ended up covering several of the cases you mentioned: Intel without AVX-VNNI, less than 16 GB of RAM, and a distribution other than Debian.

The short version is that, on Janas commit 02c325ed144cd09c916837a8cead6a0a10243952:

  • ./build.sh release test passed
  • the conversion fingerprints matched
  • the deterministic answer fingerprint matched: c08f0f936eeba5cb
  • the HTTP server checks passed 2/2 (chat + tool call)
  • an independent O_DIRECT smoke test succeeded
  • io_uring_setup() succeeded

So at least for this particular cloud/KVM environment, I did not find a portability or deterministic-correctness failure in the quick path.

I would not treat this as a useful consumer-hardware performance benchmark, though. It was only a 2-vCPU Colab allocation, and quick uses Qwen3-4B, which fit in memory here. So this does not test the interesting “MoE larger than RAM” case yet.

Environment and exact checks

The run was:

Janas:
  commit 02c325ed144cd09c916837a8cead6a0a10243952

Environment:
  Google Colab CPU runtime
  Ubuntu 24.04.5 LTS
  Linux 6.6.122+
  KVM virtual machine

CPU:
  Intel Xeon @ 2.20 GHz
  2 logical CPUs
  AVX2: yes
  AVX-VNNI: no

Memory:
  12.7 GiB
  no swap

Compiler:
  gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0

Workspace filesystem:
  overlayfs

GPU:
  not used

I used the VM-local filesystem rather than Google Drive.

The build/test step:

./build.sh release test

completed successfully. Janas’s expert-cache tests also exercised the relevant asynchronous paths successfully.

I separately checked a few Linux I/O primitives because a cloud VM seemed like a potentially interesting boundary for this project:

statx(STATX_DIOALIGN):
  memory alignment = 512
  offset alignment = 512

O_DIRECT:
  aligned 4096-byte pread: success

io_uring_setup():
  success

I am only reading those results narrowly: the basic direct-I/O and io_uring primitives worked on this particular Colab VM. They are not evidence about physical SATA/NVMe performance, since Colab does not expose the underlying storage to me as a normal local-drive benchmark target.

The model path was the current quick level from tools/janas-try.sh:

Qwen3-4B-Q4_K_M.gguf
    ↓
janas-get / gguf2jns / jns_planes
    ↓
qwen3-4b-q4km.jns

Results:

GGUF SHA-256 check:          PASS
gguf2jns fingerprint:        PASS
jns_planes fingerprint:      PASS

deterministic answer:
  expected: c08f0f936eeba5cb
  observed: c08f0f936eeba5cb
  result:   PASS

HTTP server:
  chat completion: PASS
  tool call:       PASS
  result:          2/2

The deterministic-answer check seems particularly useful here: a successful build and conversion alone would not tell us that the resulting inference path is behaving identically.

One connection I noticed: Ubuntu / GCC 13 / VM

This run also happens to overlap with issue #4, the Ubuntu/GCC 13 build failure.

That report was from Ubuntu with GCC 13.3 in a virtual machine. After fixing the GCC 13 #pragma GCC unroll / -Werror interaction, you reproduced and checked the fix using Debian’s GCC 13.3, but noted that you could not directly test Ubuntu itself or the reporter’s VM.

This Colab run is obviously not the same VM, so I would not call it a reproduction of that environment. But it does give one additional current-commit datapoint for:

Ubuntu 24.04.5
Ubuntu GCC 13.3
KVM

with the full release test passing.

Diagnostic benchmark — probably not useful as a headline number

For completeness, I also ran a small diagnostic benchmark.

configuration     prefill      decode
2 threads         4.6 tok/s    2.7 tok/s
1 thread          3.6 tok/s    2.4 tok/s

The expert cache was reported as 100%.

I would not use these numbers to compare Janas with results from actual PCs. This was only a 2-vCPU cloud VM, and the CPU was about 79% idle immediately before the measurement rather than the >=85% condition the normal tester tries to establish.

The useful part to me is just that the benchmark/tuning path ran and made a sensible local choice; the absolute throughput is mostly a property of this Colab allocation.

What this run does and does not establish

What I think this run actually establishes:

  • the current Janas release builds and passes its tests on this particular Ubuntu 24.04.5 / GCC 13.3 / KVM environment;
  • the non-AVX-VNNI x86-64 path worked;
  • the test ran with only 12.7 GiB of RAM;
  • basic O_DIRECT and io_uring operation worked on the VM-local filesystem;
  • Qwen3-4B conversion reproduced the published fingerprints;
  • the fixed greedy answer reproduced the expected fingerprint;
  • the HTTP chat and tool-call smoke tests worked.

What it does not establish:

  • performance on a physical CPU;
  • SATA or NVMe performance;
  • that overlayfs/cloud storage behaves like a local SSD;
  • the medium or full MoE paths;
  • correctness or performance when the model is actually larger than available RAM;
  • GPU/Vulkan behavior;
  • that other Colab allocations will behave identically.

In particular, quick fit comfortably enough that this was mostly a portability + correctness test, not a test of Janas’s main SSD-streaming claim.

For anyone trying to reproduce this later, I think the exact Janas commit is worth keeping with the result because the project is changing quickly.

The main references I used were the Janas repository, the current janas-try.sh, and the earlier Ubuntu/GCC 13 issue.

I kept the generated report and raw logs as well, in case a particular part of the Colab run turns out to be useful for comparison.