Hi everyone,
I’m looking for people willing to spend half an hour (or a quiet evening) running a test of Janas-LLM on their own Linux machine.
What it is. Janas-LLM is an inference engine written in C, GPL-3.0, built for ordinary machines: no GPU needed, and Mixture-of-Experts models larger than the RAM, with the experts streamed from the SSD as they are needed. On my laptop (Core Ultra 9 185H, 32 GB, NVMe) Qwen3-Next-80B-A3B — a 48 GB file — writes about 23 tokens/s, 30 with its multi-token prediction block, and it has read a full 262,144-token context and still found three facts hidden in it. It runs the Qwen3 family (dense and MoE, Qwen3.5/3.6, Qwen3-Next), has a terminal chat, an OpenAI-compatible HTTP server (chat, Responses, embeddings, tools) and MCP support.
Why I need you. The engine tunes itself from what it measures — threads, GPU, how much of each expert to read — but it has only ever measured one machine. Every other CPU, disk and amount of memory is a guess until someone runs it. Most wanted: AMD CPUs, Intel without AVX-VNNI, machines with 16 GB or less, SATA SSDs, distributions other than Debian.
How: two commands.
git clone https://github.com/prabanta-dev/janas && cd janas
./tools/janas-try.sh
You need Linux on x86-64, a C compiler (gcc or clang) and curl — nothing else to install, and no root. The script:
- builds Janas and runs its tests;
- downloads a model from its official Hugging Face repository and checks its SHA-256;
- converts it to Janas’s own format and checks the result against published fingerprints (the conversion is deterministic, so any difference is a bug found);
- checks that the model’s answer at temperature 0 is the very same text every other machine gets (the arithmetic is bit-identical across thread counts and instruction sets — your machine confirms it or finds where it is not);
- measures the speed, after checking that the machine is idle;
- tries the HTTP server;
- writes a report and shows it to you; only if you say yes does it open a GitHub issue (with
gh, or a pre-filled link).
It asks before every long step, and at the end deletes only what it downloaded, if you want. The report holds the CPU, memory, disk model, distribution, compiler and the results — no host name, user name, paths, or anything about what else runs on your machine.
Three levels, and it offers only those your machine can hold:
| Level | Model | Download | Disk while it works |
|---|---|---|---|
| quick | Qwen3-4B | 2.5 GB | ~5 GB |
| medium | Qwen3.6-35B-A3B | 22 GB | ~45 GB |
| full | Qwen3-Next-80B-A3B + MTP | ~52 GB | ~100 GB |
On a machine like mine (the laptop above, on a fast connection) quick takes about fifteen minutes; the others take longer, mostly downloading, and a slower CPU, disk or connection stretches every step.
Please measure on an idle machine: close the browser, builds and other models first. The script checks, and the report says how idle the CPU was — speeds taken on a busy machine aren’t that machine’s.
Here is what a report looks like: https://github.com/prabanta-dev/janas/issues/8
The models are the Qwen team’s, under Apache 2.0; Janas converts them locally and never redistributes them (the details are in MODELS.md).
Every report will be read, and what it shows goes back into the engine: a failure on your machine is exactly what I’m looking for. Questions and problems are welcome here or as issues on the repository.
Thank you!
Maurizio “camauri” Cammalleri
https://github.com/prabanta-dev/janas