Parallax 8B lambda: mid-training checkpoints

These are research checkpoints from lambda, a training run of Parallax that is still in progress. Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several countries. The hosts have no direct connections to each other and synchronize over the public internet.

  • Base model. It is pretrained on a mixture of web, document, math, code, knowledge-focused and book text. It is not instruction-tuned, chat-tuned or safety-tuned, and it will continue text rather than follow instructions.
  • Mid-training. Each export is a snapshot of a run that has not finished. Later exports are expected to differ, and quality between exports is not monotonic.
  • Research artifact. It is published so the training can be followed and inspected. It is not intended for production use.

Live training dashboard: parallax.chutes.ai.

Model

Parameters ~7.77B total, ~1.20B active per token
Layers 64: 32 mixers (18 EDA recurrent, 8 sliding-window attention, 6 global attention) interleaved with 32 mixture-of-experts layers
Recurrent mixers EDA recurrence (a gated delta rule) with per-key-channel decay bounded below by e^-5 per token; erase rank 16; output normalization
Sliding-window attention First and last mixers use 2048 tokens; the six interior layers share a per-step draw of 128 or 2048 tokens with equal probability throughout training; evaluation and inference use 2048
Global attention sparse-capable global attention (MSA) with 256-dimensional latent KV, 16 query heads, 4 KV heads and RoPE; dense causal attention throughout this run, with sparse top-k selection inactive
Experts 4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token
Expert weights ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight
Width Model width 1152; expert intermediate width 2304, latent width 384; ReLUยฒ expert activation
Tokenizer Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings
Training context 4096 tokens
Document masking In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document)
Output logit scale learnable, capped at 3.0

Training

  • Data: seven stationary sources totalling about 1.09T tokens: two general pretraining mixtures of web, document, math and code text, an additional web set, two math sets, a knowledge-focused set and a books set. At about 899.6B tokens the data switches to two further sources with no learning-rate event. Planned training totals about 972.9B tokens.
  • Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Routed experts are trained through low-rank adapters on the GPUs that use them, and the adapter updates are folded into full-precision masters that publish new ternary versions. The tech report describes the details.
  • Learning rate: linear warm-up from 0 to 6e-4 over the first 19.66B tokens, then 6e-4 until 400B tokens, 3e-4 until 700B and 1.5e-4 through the end of training. There is no final decay.
  • Batch: 10 gradient-accumulation micro-batches per step at launch (about 8.5M tokens per fleet step with 26 machines), switched to 20 at about 100B tokens on 2026-10-04 at 22:08 UTC. This gives about 17M tokens per step with all 26 machines; tokens per step vary with fleet size.
  • Fleet: 26 machines and 208 GPUs at launch. Three machines left on 2026-10-05 and their roles were reassigned automatically while training continued; the fleet has since comprised 23 machines and 184 GPUs.
  • Document masking: In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document).
  • Output logits: the learnable output logit scale is capped at 3.0 in training and inference (see Files and format).

Exports

One folder per export under exports/, named by training tokens (rounded down to whole billions) and fleet step. A new export is added about every 30 minutes while the run continues. Published exports are never modified or removed.

Export Tokens Step Size Exported (UTC)
468B-tokens_step37937 (latest) 468.7B 37937 2.43 GB 2026-10-07 02:48
465B-tokens_step37674 465.2B 37674 2.43 GB 2026-10-07 02:20
462B-tokens_step37429 462.1B 37429 2.43 GB 2026-10-07 01:45
458B-tokens_step37164 458.5B 37164 2.43 GB 2026-10-07 01:17
455B-tokens_step36914 455.4B 36914 2.43 GB 2026-10-07 00:45
451B-tokens_step36659 451.9B 36659 2.43 GB 2026-10-07 00:19
448B-tokens_step36416 448.7B 36416 2.43 GB 2026-10-06 23:48
445B-tokens_step36150 445.2B 36150 2.43 GB 2026-10-06 23:17
442B-tokens_step35905 442.0B 35905 2.43 GB 2026-10-06 22:45
438B-tokens_step35646 438.5B 35646 2.43 GB 2026-10-06 22:17
435B-tokens_step35402 435.3B 35402 2.43 GB 2026-10-06 21:46
431B-tokens_step35140 431.9B 35140 2.43 GB 2026-10-06 21:18
428B-tokens_step34890 428.7B 34890 2.43 GB 2026-10-06 20:46
425B-tokens_step34634 425.2B 34634 2.43 GB 2026-10-06 20:21
422B-tokens_step34390 422.0B 34390 2.43 GB 2026-10-06 19:46
418B-tokens_step34134 418.5B 34134 2.43 GB 2026-10-06 19:17
415B-tokens_step33886 415.2B 33886 2.43 GB 2026-10-06 18:45
411B-tokens_step33618 411.8B 33618 2.43 GB 2026-10-06 18:20
408B-tokens_step33381 408.7B 33381 2.43 GB 2026-10-06 17:47
405B-tokens_step33116 405.1B 33116 2.43 GB 2026-10-06 17:17
402B-tokens_step32861 402.0B 32861 2.43 GB 2026-10-06 16:45
398B-tokens_step32598 398.6B 32598 2.43 GB 2026-10-06 16:16
395B-tokens_step32353 395.3B 32353 2.43 GB 2026-10-06 15:44
391B-tokens_step32098 391.8B 32098 2.43 GB 2026-10-06 15:18
388B-tokens_step31854 388.6B 31854 2.43 GB 2026-10-06 14:45
385B-tokens_step31191 385.1B 31191 2.43 GB 2026-10-06 14:18
381B-tokens_step30942 381.8B 30942 2.43 GB 2026-10-06 13:45
378B-tokens_step30696 378.2B 30696 2.43 GB 2026-10-06 13:18
374B-tokens_step30455 374.8B 30455 2.43 GB 2026-10-06 12:44
371B-tokens_step30200 371.2B 30200 2.43 GB 2026-10-06 12:14
367B-tokens_step29959 367.8B 29959 2.43 GB 2026-10-06 11:45
364B-tokens_step29699 364.3B 29699 2.43 GB 2026-10-06 11:17
360B-tokens_step29461 360.9B 29461 2.43 GB 2026-10-06 10:47
357B-tokens_step29209 357.3B 29209 2.43 GB 2026-10-06 10:17
353B-tokens_step28971 354.0B 28971 2.43 GB 2026-10-06 09:46
350B-tokens_step28721 350.4B 28721 2.43 GB 2026-10-06 09:19
347B-tokens_step28477 347.0B 28477 2.43 GB 2026-10-06 08:46
343B-tokens_step28219 343.4B 28219 2.43 GB 2026-10-06 08:17
340B-tokens_step27983 340.1B 27983 2.43 GB 2026-10-06 07:47
336B-tokens_step27720 336.4B 27720 2.43 GB 2026-10-06 07:18
333B-tokens_step27484 333.0B 27484 2.43 GB 2026-10-06 06:46
329B-tokens_step27229 329.5B 27229 2.43 GB 2026-10-06 06:20
326B-tokens_step26986 326.1B 26986 2.43 GB 2026-10-06 05:43
322B-tokens_step26728 322.5B 26728 2.43 GB 2026-10-06 05:17
319B-tokens_step26493 319.2B 26493 2.43 GB 2026-10-06 04:44
315B-tokens_step26240 315.6B 26240 2.43 GB 2026-10-06 04:17
312B-tokens_step25999 312.2B 25999 2.43 GB 2026-10-06 03:44
308B-tokens_step25748 308.6B 25748 2.43 GB 2026-10-06 03:18
305B-tokens_step25502 305.2B 25502 2.43 GB 2026-10-06 02:45
301B-tokens_step25253 301.7B 25253 2.43 GB 2026-10-06 02:19
298B-tokens_step25009 298.3B 25009 2.43 GB 2026-10-06 01:43
294B-tokens_step24762 294.7B 24762 2.43 GB 2026-10-06 01:18
291B-tokens_step24519 291.3B 24519 2.43 GB 2026-10-06 00:46
287B-tokens_step24258 287.7B 24258 2.43 GB 2026-10-06 00:15
284B-tokens_step24014 284.4B 24014 2.43 GB 2026-10-05 23:44
280B-tokens_step23768 280.7B 23768 2.43 GB 2026-10-05 23:15
277B-tokens_step23530 277.4B 23530 2.43 GB 2026-10-05 22:46
273B-tokens_step23276 273.8B 23276 2.44 GB 2026-10-05 22:16
270B-tokens_step23035 270.4B 23035 2.44 GB 2026-10-05 21:44
266B-tokens_step22784 266.9B 22784 2.44 GB 2026-10-05 21:19
263B-tokens_step22539 263.5B 22539 2.44 GB 2026-10-05 20:47
259B-tokens_step22282 259.9B 22282 2.44 GB 2026-10-05 20:17
256B-tokens_step22052 256.6B 22052 2.44 GB 2026-10-05 19:47
252B-tokens_step21791 252.9B 21791 2.44 GB 2026-10-05 19:17
249B-tokens_step21550 249.6B 21550 2.44 GB 2026-10-05 18:45
245B-tokens_step21250 245.3B 21250 2.45 GB 2026-10-05 18:09
242B-tokens_step21067 242.4B 21067 2.45 GB 2026-10-05 17:45
238B-tokens_step20776 238.0B 20776 2.45 GB 2026-10-05 17:14
235B-tokens_step20575 235.2B 20575 2.46 GB 2026-10-05 16:46
230B-tokens_step20280 230.7B 20280 2.46 GB 2026-10-05 16:12
227B-tokens_step20081 227.9B 20081 2.46 GB 2026-10-05 15:45
223B-tokens_step19786 223.4B 19786 2.45 GB 2026-10-05 15:11
220B-tokens_step19602 220.6B 19602 2.45 GB 2026-10-05 14:47
216B-tokens_step19301 216.2B 19301 2.45 GB 2026-10-05 14:11
213B-tokens_step19100 213.3B 19100 2.45 GB 2026-10-05 13:45
208B-tokens_step18807 208.9B 18807 2.45 GB 2026-10-05 13:15
206B-tokens_step18619 206.0B 18619 2.45 GB 2026-10-05 12:45
201B-tokens_step18313 201.6B 18313 2.45 GB 2026-10-05 12:12
198B-tokens_step18118 198.7B 18118 2.45 GB 2026-10-05 11:44
194B-tokens_step17828 194.3B 17828 2.45 GB 2026-10-05 11:11
191B-tokens_step17631 191.5B 17631 2.45 GB 2026-10-05 10:44
187B-tokens_step17334 187.0B 17334 2.45 GB 2026-10-05 10:12
184B-tokens_step17145 184.2B 17145 2.45 GB 2026-10-05 09:45
179B-tokens_step16851 179.8B 16851 2.45 GB 2026-10-05 09:11
176B-tokens_step16645 176.9B 16645 2.45 GB 2026-10-05 08:45
172B-tokens_step16359 172.6B 16359 2.45 GB 2026-10-05 08:12
169B-tokens_step16164 169.7B 16164 2.45 GB 2026-10-05 07:45
165B-tokens_step15861 165.2B 15861 2.45 GB 2026-10-05 07:10
162B-tokens_step15667 162.4B 15667 2.46 GB 2026-10-05 06:45
157B-tokens_step15384 157.9B 15384 2.46 GB 2026-10-05 06:13
155B-tokens_step15179 155.1B 15179 2.46 GB 2026-10-05 05:46
150B-tokens_step14880 150.6B 14880 2.46 GB 2026-10-05 05:12
147B-tokens_step14693 147.8B 14693 2.46 GB 2026-10-05 04:45
143B-tokens_step14118 143.4B 14118 2.47 GB 2026-10-05 04:13
140B-tokens_step13940 140.5B 13940 2.47 GB 2026-10-05 03:44
135B-tokens_step13714 135.8B 13714 2.47 GB 2026-10-05 03:14
132B-tokens_step13524 132.5B 13524 2.47 GB 2026-10-05 02:48
127B-tokens_step13242 127.7B 13242 2.48 GB 2026-10-05 02:12
123B-tokens_step13016 123.6B 13016 2.48 GB 2026-10-05 01:43
119B-tokens_step12790 119.5B 12790 2.48 GB 2026-10-05 01:13
115B-tokens_step12532 115.4B 12532 2.48 GB 2026-10-05 00:38
111B-tokens_step12302 111.4B 12302 2.48 GB 2026-10-05 00:10
107B-tokens_step12081 107.2B 12081 2.48 GB 2026-10-04 23:41
103B-tokens_step11842 103.2B 11842 2.48 GB 2026-10-04 23:14
99B-tokens_step11533 99.2B 11533 2.48 GB 2026-10-04 22:39
95B-tokens_step11160 96.0B 11160 2.47 GB 2026-10-04 22:12
92B-tokens_step10766 92.6B 10766 2.47 GB 2026-10-04 21:43
89B-tokens_step10388 89.3B 10388 2.47 GB 2026-10-04 21:12
85B-tokens_step9986 86.0B 9986 2.47 GB 2026-10-04 20:41
82B-tokens_step9591 82.7B 9591 2.47 GB 2026-10-04 20:09
79B-tokens_step9195 79.3B 9195 2.47 GB 2026-10-04 19:39
76B-tokens_step8835 76.1B 8835 2.48 GB 2026-10-04 19:13
72B-tokens_step8443 72.8B 8443 2.48 GB 2026-10-04 18:40
69B-tokens_step8039 69.5B 8039 2.48 GB 2026-10-04 18:11
66B-tokens_step7652 66.1B 7652 2.48 GB 2026-10-04 17:42
62B-tokens_step7275 62.9B 7275 2.48 GB 2026-10-04 17:10
59B-tokens_step6711 59.6B 6711 2.48 GB 2026-10-04 16:40
56B-tokens_step6444 56.3B 6444 2.48 GB 2026-10-04 16:09
53B-tokens_step6096 53.0B 6096 2.48 GB 2026-10-04 15:40
49B-tokens_step5723 49.8B 5723 2.48 GB 2026-10-04 15:12
46B-tokens_step5320 46.4B 5320 2.48 GB 2026-10-04 14:40
43B-tokens_step4946 43.1B 4946 2.48 GB 2026-10-04 14:10
39B-tokens_step4549 39.8B 4549 2.48 GB 2026-10-04 13:38
36B-tokens_step4166 36.5B 4166 2.48 GB 2026-10-04 13:15
33B-tokens_step3768 33.2B 3768 2.48 GB 2026-10-04 12:40
29B-tokens_step3391 29.8B 3391 2.49 GB 2026-10-04 12:16
26B-tokens_step3002 26.5B 3002 2.49 GB 2026-10-04 11:43
23B-tokens_step2604 23.1B 2604 2.50 GB 2026-10-04 11:15
19B-tokens_step2198 19.9B 2198 2.50 GB 2026-10-04 10:40
16B-tokens_step1817 16.7B 1817 2.50 GB 2026-10-04 10:11
13B-tokens_step1419 13.3B 1419 2.50 GB 2026-10-04 09:42
6B-tokens_step635 6.6B 635 2.50 GB 2026-10-04 08:41
  • This table lists the exports only. Validation loss and benchmark results for this run are shown on the live training dashboard.

Latest export

exports/468B-tokens_step37937: 468.7B training tokens, step 37937.

hf download chutesai/parallax-8b-lambda --include "exports/468B-tokens_step37937/*" --local-dir parallax-8b-lambda
cd parallax-8b-lambda/exports/468B-tokens_step37937
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS   # every file of the original export, byte for byte

Files and format

The files are in Parallax's native compact export format (inference only, no optimizer state). They are byte-identical to the export the training system produced:

File Contents
manifest.json export manifest: tensor inventory, per-file sha256 digests, token clock
model_config.json model configuration
coverage.json, layouts.json tensor coverage and expert frame layouts
indexer.bundle sparse-attention indexer weights
relay_pack/ trunk (non-expert) weights in bf16, with their own manifest
packed_experts.tar the 4096 routed experts (packed_experts/*.t24p, packed ternary codes and scales)
SHA256SUMS sha256 of every file of the original export
export_info.json step, tokens, time, sizes and digests of this export

The only change from the original export is packaging. The 4096 expert files are stored in one uncompressed tar to keep the repository's file count manageable. Extract it and check SHA256SUMS as shown above. Every upload was checked against the training system's own digests before and after it was published.

Logit-scale bound. The model's output logit scale is bounded: the forward pass uses exp(min(logit_scale_log, log 3.0)). The stored trunk tensor is the raw training parameter (it can sit slightly above the bound, e.g. from bf16 rounding), and each export records the bound in manifest.json (logit_scale: max, raw, effective) and in coverage.json (inference_policy.logit_scale_max). A loader must apply the recorded bound; the files themselves are left byte-identical.

Running it

This is a base model only. It is not chat- or instruction-tuned, so it does plain text completion: give it the start of a text and it continues it. It will not follow instructions or hold a conversation.

Standard transformers cannot load this format. Use the parallax-lambda branch of our llama.cpp fork, https://github.com/chutesai/llama.cpp. It adds a dedicated runtime, llama-parallax, for CPU (x86-64, ARM64) and Apple GPUs (Metal). Lambda's recurrent mixers (EDA) are implemented only on this branch. The fork's master branch runs the earlier kappa checkpoints and cannot run lambda.

A vLLM serving path for GPU servers is in preparation.

Ready-made GGUF

parallax-8b-lambda-431B-tokens.gguf at the repository root (2,133,006,592 bytes, 2.13 GB, sha256 687296380f21e2c133b6d00668e09a70e1fc5488186b27a4e4ca20cdf4ddd694, also in the .sha256 file next to it) was converted from exports/431B-tokens_step35140 (431.9B training tokens, step 35140). It keeps the export's bf16 trunk and stores the routed experts' exact ternary codes and per-row scales; no weight is requantized. Each GGUF is a snapshot of one export and is not updated as new exports are added. To use a later export, convert it yourself (see tools/parallax/README.md on the branch).

The earlier snapshot, parallax-8b-lambda-381B-tokens.gguf, from exports/381B-tokens_step30942, remains at the repository root (sha256 86312194ffc72b7b8d87f37389fe0fdfe7c510e2f6066f0c424cc0eaf592aedf).

hf download chutesai/parallax-8b-lambda parallax-8b-lambda-431B-tokens.gguf --local-dir .
hf download NousResearch/Meta-Llama-3-8B tokenizer.json --local-dir .   # Llama 3 tokenizer

Build and run

Requires CMake, a C++17 compiler, and Python 3 with tokenizers for the text front end.

git clone -b parallax-lambda https://github.com/chutesai/llama.cpp && cd llama.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON   # -DGGML_METAL=OFF for CPU only
cmake --build build --target llama-parallax -j 8
build/bin/llama-parallax --backend metal --self-test               # or --backend cpu
pip install tokenizers

# Apple GPU
python tools/parallax/run.py --binary build/bin/llama-parallax \
  --model ../parallax-8b-lambda-431B-tokens.gguf --tokenizer ../tokenizer.json \
  --backend metal --experts lut9 --prompt 'The capital of France is' --predict 64

# CPU (x86-64 or ARM64)
python tools/parallax/run.py --binary build/bin/llama-parallax \
  --model ../parallax-8b-lambda-431B-tokens.gguf --tokenizer ../tokenizer.json \
  --backend cpu --experts lut8 --threads 8 --prompt 'The capital of France is' --predict 64

Generation is greedy by default. build/bin/llama-parallax --help lists sampling, context size (-c, default 4096), prompt scoring and benchmark options.

Speed and memory

Measured on an Apple M5 Max on AC power (batch 1, 128 generated tokens) with the bf16-trunk GGUF of the earlier 270B-tokens_step23035 export, which has the same architecture, using the same runtime code; speed does not depend on the weight values. Prompt times used --prefill-batch 128:

Backend Experts Tokens/s, short context Tokens/s after 4,096 tokens 2,048-token prompt
Metal lut9 192 172 2.2 s
CPU, 12 threads lut8 157 112 4.5 s
CPU, 12 threads lut (exact) 90 74

Peak resident memory was about 2.8 GB on Metal and 3.4 GB on CPU at the default 4,096-token context. Other Apple and ARM hardware will differ. ARM64 CPUs use NEON kernels. x86-64 CPUs use portable C++ that gives the same results but is not tuned for speed: on a shared, loaded 64-core AMD EPYC 9534 server it generated about 22 tokens/s with 16 threads.

Fidelity

The 431B GGUF was checked on CPU with lut against the training code's own model loaded from the same export, with FP32 activations in both, on the same 2,555 held-out positions from public text and code passages as before. Mean KL was 1.7e-6 nats/token, the top-1 next token agreed at 2,554 of 2,555 positions, and three 48-token greedy continuations matched exactly.

The small differences trace to two tokens where the router's expert-selection scores for the 12th and 13th expert were tied or within a few FP32 rounding steps (in one case exactly tied in the reference). The two implementations picked different experts there, and the difference carried forward in that passage. When the reference was given the runtime's choice at those two tokens, mean KL was about 6e-10 nats/token and the top-1 next token agreed at every position.

The training code itself runs in bf16; against that bf16 forward the top-1 agreement was 94.9% (mean KL 0.0027, measured on the earlier 270B-tokens_step23035 export), which reflects the bf16 rounding of the reference, not an error of the runtime.

On the 431B GGUF, Metal and CPU had per-layer relative error within 2.2e-5 at the final position of a 512-token text, and identical 48-token greedy continuations on three prompts. lut8 (CPU) quantizes expert lookup tables to 8 bits and is slightly approximate; lut and lut9 are exact.

Limitations of the runtime

  • One sequence per process, greedy or simple sampled generation. No chat template, batching or server.
  • Not integrated into llama-cli, llama-server or the standard llama.cpp model loader.
  • Context is limited by -c. Inference uses a 2,048-token window for all sliding-window layers and dense global attention, the settings the run is evaluated with. Lambda was trained on 4,096-token sequences; longer inputs run but were not qualified.

Tech report

The Parallax tech report: https://parallax.chutes.ai/tech-report.pdf. It is AI-generated from the project's measurements and logs, and it is a living document that changes as the run progresses. The report currently published there describes an earlier run of the same system; lambda's architecture and training recipe are as described on this card.

Limitations

This is an early base model. It can produce incorrect, biased or nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.

License

MIT.

Table updated 2026-10-07 02:58 UTC.

Downloads last month
-
GGUF
Model size
2B params
Architecture
parallax
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support