Compiler and runtime
Flux runs on two engines: a graph interpreter, which serves the editor, the live preview and the debugger, and a compiled WebAssembly module, which serves the run. They are not an approximation of each other. They produce the same bytes, and that equality is checked at every single compilation, on hostile data, before anything ships.
That one invariant is the spine of everything else on this page. It is what lets the engine be swapped underneath a running chart with no visual glitch, what lets an optimizer be aggressive without being trusted, and what lets a server re-execute a client’s work and detect a lie.
Together the two engines are one abstract machine — the FVM, the Flux Virtual Machine: a total, sandboxed, deterministic dataflow machine whose instruction set is the kernels and the graph operations, whose memory is the static linear layout of the memory model, and whose arithmetic is pinned to the byte. Neither the interpreter nor the WebAssembly module is the FVM; each is a way of running it, and I7 (below) is the guarantee that the two ways cannot disagree. It is a virtual machine in the exact sense the JVM is — a single semantics with more than one conforming implementation — not a bytecode loop: the “instructions” are a typed graph of kernels, and the two implementations lower it differently but must land on the same bits.
New here? Start with Guide §11 — Determinism, replay and trust → — the accessible version of the two-engine contract, before the machine-level rules below.
The pipeline
source
│ parse Lezer — incremental, total: an arbitrary input yields one tree or a clean error
▼
│ resolve + inline names and `def` bodies (the call graph is acyclic, so inlining terminates)
▼
typed DAG kinds inferred bottom-up; presentation derived from the kinds
│ causality check every feedback cycle must cross a unit delay
▼
│ optimize common-subexpression elimination · dead code · constant folding · fusion · kernel selection
▼
│ memory plan liveness intervals → slots → an exact footprint
▼
emit ──┬─→ interpreter closures (edit · preview · debug · THE ORACLE)
└─→ WebAssembly (Binaryen) (run · distribute)
Figure — one graph, two back-ends, one gate between them.
The interpreter pays once, when the graph is built. After that the hot loop is native kernels over pre-allocated columns.
The unit of compilation is (graph, resolved parameters, bar capacity). Parameters are
resolved at compile time — changing a knob recompiles and re-gates. That is a deliberate
trade: it buys a fully static state layout, an exact min = max memory, and a gate that
validates the very bytes and the very instance that will serve. There is no gap between
what was verified and what runs.
Two engines, one contract
| Interpreter | WebAssembly module | |
|---|---|---|
| Serves | editing, live preview, the dataflow debugger | the run, and distribution |
| Feedback | instant — no compile step in the loop | compiled and gated, then swapped in |
| Role in the contract | the oracle | the candidate |
I6 — a leaf is byte-identical to its kernel
A node that maps to a native kernel produces exactly the bytes that kernel produces, warm-up
included. Flux does not impose its own na-until-N convention: a Flux indicator on a clock is
the same citizen as a built-in, from the first bar. That is what makes “rewrite the catalogue in
Flux” a safe proposition rather than a rewrite of every golden.
I7 — the interpreter and WASM agree, byte for byte
The gate runs at every compilation and it blocks:
- The oracle is the interpreter, evaluating the graph directly.
- The candidate is the instantiated WebAssembly module.
- They are compared byte-wise on every sink column, over a hostile, deterministic, adaptive corpus — series longer than the program’s resolved maximum lookback, seeded with holes, ±infinity, negative zero, raw non-canonical NaN patterns, exact half-integers, flat runs, monotone runs, magnitudes at 1e±9 — first in batch, then live (bar by bar, through the incremental step), then on the real data if any is supplied.
- Any divergence blocks the compilation. The module is not shipped; the interpreter keeps serving.
Figure — the gate compares the two engines on adversarial data, in batch and live, before a single byte is allowed out.
Why the live path is in the gate too. Batch equality is the easy half. The live path steps one bar at a time through a different code path in the module, and it is exactly where a subtle state bug hides. In the emitted module, one shared body serves both — the batch range and the single-bar advance are the same code, called with different bounds — so live ≡ batch by construction; the gate then checks the one seam that remains (column bases versus the live scratch) rather than trusting it.
What it buys, concretely
- Engine swaps are invisible. Serving interpreter results now and WASM results a moment later produces the same bytes. No migration state, no reconciliation, no repaint, no golden churn.
- The optimizer needs no trust. Any value change it introduces is caught by the blocking gate at the compilation that produced it.
- A distributed module is verifiable. Anyone holding the graph can re-run the validation locally: identity is checkable, not promised.
The browser path, as built
The main thread never imports the compiler. It orchestrates:
- A registered script serves immediately, on the interpreter — the proven path, with no compile step between typing and seeing.
- The service compiles in a worker, where compilation and the I7 gate are atomic: divergent bytes are never handed back, because they are never handed back at all.
- When the gated module exists, the service upgrades silently. By I7 the bytes are identical, so the swap requires no re-render and produces no visible event.
The compiler bundle is fetched lazily, on editor intent — never on the chart path, and excluded from the offline precache. Compiled modules are cached in tiers (attached instances by script and configuration; modules by a canonical key, in a bounded cache), because the expensive part is compilation, not instantiation.
Determinism: the pinned-routine discipline
Byte-identity does not survive first contact with a standard library. Two engines can disagree about the last bit of a logarithm, the sign of a zero, or the rounding of a half-integer, and every one of those disagreements is enough to break replay. So Flux pins the routines — the same code, on both sides, with no delegation to the platform.
| Routine | Pinned to | Why the obvious choice is wrong |
|---|---|---|
log exp sin cos tan atan atan2 pow |
a pinned WebAssembly libm, used by both engines | two JavaScript engines differ by ≥ 1 unit in the last place |
% |
the pinned library’s remainder, linked directly module-to-module | the remainder itself is exactly specified — what is pinned is the binding: routing it through the host instead would reshape a produced NaN, and the canonicalization has to happen in one place |
round |
f64.nearest — ties to even |
half-up rounding disagrees on every half-integer |
min / max |
the exact na-absorbing selection chain |
the native instructions propagate NaN and return −0 for min(−0, +0) — both disagree with the language’s rule |
| NaN constants | emitted as an integer bit pattern, then reinterpreted | a NaN marshalled through a host API has no guaranteed bits |
na at rest |
the canonical quiet NaN 0x7FF8000000000000 |
WebAssembly leaves a produced NaN’s sign and payload undetermined |
decimal |
one shared multi-limb integer routine | a floating-point engine has no native 128-bit integer; any emulation would diverge |
string, fmt.* |
pinned Unicode tables and one canonical formatter | platform string length is in UTF-16 units; platform number formatting differs in the last digit |
| calendar | a pinned epoch ↔ civil routine, pinned time-zone data | two correct implementations still disagree on DST gaps and end-of-month clamping |
rand(seed) |
one pinned counter-based integer generator; in v1, rand()/rand(seed) live on the presentation plane only, and the deterministic-domain form lands after |
integer arithmetic is bit-identical for free; floating-point mixing is not |
ordering of na in sorts |
one pinned total order (absent values last, stable by index) | a partial comparator leaves the order to the platform’s sort |
| the memory plan | a deterministic function of the graph | the value oracle is blind to layout, so nothing else would catch a divergence |
Why the interpreter does not call the platform. It is written in TypeScript and runs on a JavaScript engine, so
Math,Number,String.prototype,DateandIntlare right there. Using them would make the interpreter agree with itself and disagree with the module — and the disagreement would be invisible to a program (nothing observes a NaN payload; nothing observes the last bit of a logarithm) and visible only to a byte-level oracle. The discipline is: the interpreter runs the same pinned routine the module runs.
The toolchain is pinned too
The lowering to WebAssembly goes through Binaryen, and it must be a deterministic function:
- the optimization pipeline is fixed (an
O3pipeline, no shrink bias, fast-math off, a fixed feature set) and re-validated by the gate at every compile; - the Binaryen version is an input to the build hash, not just a note in a changelog;
- iteration is independent of pointer or address order; function and local ordering is canonical; the name and producer sections are omitted or pinned; no timestamp, path or build id enters the artifact.
Optimization is mandatory, not optional — an unoptimized module is not a “safer” module, it is a different module, and the whole point is that there is only one.
The emitted bytes are a pure function of (program, toolchain), so the toolchain has a single identity — the compiler version, the Binaryen pin, the identity of the pinned math routines, and the enabled feature set — and every cache key for a compiled artifact carries the whole of it. Without that, bumping an emission rule would keep serving stale bytes that a recompilation no longer produces: a cache that is keyed on less than the toolchain is a silent divergence with a long fuse.
Minification and obfuscation, where used, are deterministic and applied after the gate, on the distribution artifact. The provenance hash and the rebuild gate are defined on the artifact that actually ships.
The runtime surface
The compiled module exposes the same shape the native engine already uses:
| Surface | Meaning |
|---|---|
make |
instantiate and reset |
advance(bar) |
the live step |
run(n) |
the batch range |
snapshot / restore |
a copy of the state region — a checkpoint is a memcpy, and resumption is byte-exact |
One body serves run and advance, so the live and batch paths cannot drift apart, and a
checkpoint is a contiguous copy rather than a bespoke serializer.
Why WebAssembly
Not for speed. The honest measurement: fusing kernels with glue gains nothing in batch (the
boundary is amortized over a long history), SIMD is excluded by byte-identity (a horizontal
reduction reassociates floating-point and changes the bits), and scalar f64 compute in a
modern JavaScript engine is already within a small factor. Real speed comes from algorithmic
work and native kernels, not from the execution language.
WebAssembly is the target for four structural reasons:
- One execution artifact from day one. The interpreter and the module must agree bit for bit; freezing that equality at the start is far cheaper than retrofitting it.
- Cross-machine determinism. WebAssembly’s floating-point semantics are strictly specified — which is what makes server-side re-execution meaningful rather than “same runtime, probably”.
- Application panes. One execution and distribution format for indicators, representations, drawings, scenes, transitions and application logic.
- An opaque, sandboxed distribution artifact. A shared or purchased script ships as WASM, never as source: the intellectual property is protected, and the consumer’s trust boundary is a binary plus a sealed manifest.
The honest frontier. WebAssembly computes. It never paints: it produces geometry, a draw-list, and a view tree in linear memory, and the host — JavaScript, or the GPU — does the painting. There is no Web API reachable from the module. Rendering runs entirely on WebGPU: scene graphics — the geometry and the draw-list — and text alike, the text as an SDF glyph atlas that stays crisp at any zoom and DPR. A software rasterizer inside WebAssembly was considered and rejected: it loses that sharpness and is slower than the GPU. The render targets are specified in display; what the module guarantees is narrower and exact — a deterministic scene, the same numbers and the same draw-list on every machine, never identical pixels across GPUs.
eval and dynamic code generation remain forbidden. WebAssembly is not “eval with extra steps” —
it is a sandbox with linear memory, no DOM, and host access only through vetted imports, admitted
by its own dedicated policy token.
Budgets — counted at compile time, never timed at run time
A budget that is enforced by a stopwatch is not a budget: the same script would be accepted on one machine and killed on another, and replay would die of it. Every ceiling in Flux is therefore a counter, evaluated on the graph, before anything runs.
| Ceiling | Value | What it bounds |
|---|---|---|
N_max |
10 000 in the browser; 100 000 on the server and in backtests | the const length of a window or a vec. Beyond it: [ErrTotal] |
maxNodes |
3 072 per script | the size of the graph |
maxBricksPerBar |
1 000 | how many re-binned units one bar may produce (Renko, P&F). Beyond the cap the host aggregates rather than blow the budget |
N_active |
16 co-active scripts per chart | how many scripts merge into one shared DAG; the aggregate ceiling is N_active × maxNodes |
| memory per instance | the structural bound maxNodes × N_max |
the legal worst case. The declared footprint is the exact liveness plan, which is far smaller |
The numbers are the least interesting part of that table. What matters is how they are enforced.
The graph is the authority, not the text. maxNodes is judged on the DAG after inlining,
common-subexpression elimination and dead-code elimination — the graph that will actually run. The
front-end guards (AST size, inlining expansion) sit far above it and exist only to keep a hostile
input from exhausting the compiler; they never pronounce a budget verdict. nMax likewise never
touches a computed byte: browser and server differ in what they accept, never in what they
produce.
The memory ceiling is not a check — it is the allocation. The module declares its linear
memory with min = max, sized from the liveness plan (§ memory model). Growth
is not forbidden at run time; it is impossible. The static part is verified at compile time and
the remaining lengths at instantiation — never while a bar is being stepped.
There is no per-bar timeout. Runtime cost is statically bounded by
maxNodes × N_max × maxBricksPerBar: totality gives termination, and the ceilings give the
practical bound. An over-budget graph is rejected at compilation, not killed mid-frame — the
runtime guard below is defence in depth, not the enforcement path. A build does carry a wall-clock
timeout, but it is an interactive cancellation in the editor, never a verdict: accept and
reject stay a pure function of the source.
When the aggregate frame budget is nevertheless exceeded — many co-active scripts on one chart — the host applies a deterministic degradation policy, throttling or pausing low-priority scripts in an explicit declared order. The order is never data-dependent, so the degradation replays like everything else.
Where the work runs is a counted decision too. The service estimates a graph’s cost in cost units and dispatches to a worker fleet only above a calibrated threshold; below it, a merged single-threaded pass finishes before a fleet would have started. The estimate is a pure function of the graph, so the routing is reproducible — and by 1 ≡ N the choice cannot change a byte either way.
Fault isolation
A script that fails at runtime — an over-budget graph, a genuine error on a valid graph — is quarantined: its node is marked and removed from the active graph, without killing the worker and without disturbing its neighbours, which share the same instance map. Restarting the worker is the last resort, and the neighbours resume from their last checkpoint.
A NaN is not a fault. It is na, and it is displayed as a gap.
Verification and reproducible builds
I7 is one clause of a larger contract. Around it sits a verification harness — a first-class deliverable, not a folder of tests — whose sub-suites each declare an oracle, a corpus, and whether they block the ship: deterministic goldens, the formal properties (principality, confluence, totality, causality, the peak-≤-sum memory plan), a total-parser fuzzer driven by a well-typed generator, the three-way differential oracle (interpreter ↔ WASM ↔ native kernel, covering I6, optimized ≡ reference and I7), the enumerated metamorphic relations, the 1 ≡ N stress suite, lattice enumeration, and a model-checked capability monitor. Above them, reproducible builds make the emitted module a pure function of source, lockfile, compiler version, pinned routines and the canonical memory plan, sealed by a rebuild gate that recompiles the same inputs on different machines and asserts byte-identity — the condition server-side replay silently depends on.
One subtlety carries over into every byte-identity claim on this page: the three-way oracle calls the same pinned routine on all three sides, so it is blind to a bug inside a pinned routine. That is why each pinned routine also ships a second, independent reference implementation, compared bit for bit on fuzzed input — the oracle catches disagreement, only a second implementation catches a shared mistake.
Canonical: the full suite table with its blocking verdicts, and the build-hash inputs and rebuild-gate procedure, live in Verification and reproducible builds. This section is a summary, not a second copy.
Performance — measured, and honest
Why WebAssembly argued that the target was not chosen for speed. The
certification benchmark settles what the speed actually is — measured on one Apple M4, over a
hundred thousand bars, with the TypeScript, interpreter and WASM legs proven byte-identical before a
single timing is taken. So every row below compares the same algorithm: the TypeScript leg is the
interpreter’s own kernels, called directly on a plain Float64Array, not a naive rewrite. And
against that, in batch — the path every chart takes — the compiled module is faster everywhere, by a
margin that grows with how much intermediate data the indicator moves:
| Workload — the same algorithm on both sides | How much faster than TypeScript |
|---|---|
Trivial O(1) kernels — rsi, change |
≈ parity — V8 compiles a tight f64 loop nearly as well |
Weighted scans and deques — wma, alma, highest |
×1.3–2 |
| Realistic charts — a classic eleven-plot chart runs ≈ 21.5 ns/bar | ×1.6–2.5 |
Order-p scans and reductions — percentrank, kama |
×2–3 |
Multi-stage composites — fisher, connorsRsi |
×3.7–5 |
Heavy multi-output composites — stdErrorBands |
×9–14 |
Figure — the same algorithm on both sides: the batch gap grows with how much intermediate data the workload moves, from parity on trivial kernels to ×9–14 on heavy composites.
Why — the datapath, not the engine. Decompose the gap and the execution engine is the small
part of it. Wasm is machine code with single-instruction f64 ops and no dynamic bounds-checks on a
linear memory proven in range at compile time — but V8’s TurboFan compiles a hot, monomorphic
typed-array loop nearly as well, so the engine difference alone is only ×1.1–2, which is exactly
why the trivial kernels land at parity. The lever is the datapath. A multi-stage indicator
hand-written in TypeScript materialises a full Float64Array per stage, writing then re-reading
its values each time; the flux compiler fuses every stage into one pass, holding the
intermediates in locals and rings. stdErrorBands’ TypeScript leg allocates eight arrays; the module
allocates none — and every materialised intermediate is two full memory passes of pure overhead.
That, not the instruction set, is where the ×2-to-×14 comes from. (Compilation itself is tens of
milliseconds, almost all of it Binaryen lowering the graph to bytes — an editor-time cost, paid once,
never per bar.)
And beneath it, the algorithms keep the bytes. A rolling maximum reads a monotone deque instead
of rescanning its window; a reset touches only the live geometry, about 28.5 µs down to 0.8 µs; a
bounded length knob shrinks a script’s state by roughly twenty times; an sma and a sum that
coincide share one ring. Every one is an O(n) change with the same output bits — a win the
optimizer and the native kernels deliver in any language, and one more reason the
figure above is not a story about the execution engine.
The one case TypeScript wins — and why it is not a headline. Stream a single O(1) kernel that has
a native f32 stepper — an ema or rsi advanced bar by bar — and the native library wins, ×5–8:
flux pays a call across the WASM boundary on every bar where the library takes a single step. But that
leg runs in f32, not f64 — a different-precision result, recorded as a documented divergence and
never quoted as a speed headline. And it exists only where a native stepper does: a windowed wma or
a deque highest has none, so that path recomputes in batch, where the module takes it all back.
Why SIMD is off the table — by choice. The one lever that would move scalar compute is SIMD, and
byte-identity forecloses it: f32x4 packs four lanes into one, and a horizontal reduction
reassociates floating-point and drops precision, so the bits change. Same-bits-everywhere is what
no-repaint, replay, the goldens and the server’s re-execution all stand on; trading it for a fraction
of a factor would spend the property that makes these numbers mean something. It is a decision,
recorded and enforced, not a feature still to come — the same line the optimizer draws one level down
when it forbids reassociation.
So what WebAssembly is for. The speed is real, but read the decomposition again: it comes from the fused dataflow the compilation model gives you — not from WebAssembly as an execution engine, which is only the ×1.1–2. The reasons to target WebAssembly specifically are the four structural ones in Why WebAssembly: one execution artifact, so the interpreter and the module can be held equal from the start; floating-point specified across machines, so a server can re-run a client’s work and catch a lie; one format that carries an indicator, a scene and a whole application pane alike; and a sandboxed, opaque binary to distribute, source withheld. Flux is fast because it compiles a fused dataflow graph — and it compiles that graph to WebAssembly for correctness, determinism and distribution. A benchmark you can reproduce is a better argument than a superlative.
See also
- Verification and reproducible builds — the full harness table and the rebuild gate this page summarizes.
- Memory model — the layout the plan produces, and why the plan itself is pinned.
- Optimizer — the correctness law, the tiers, and translation validation.
- Concurrency — the scheduler, and the 1 ≡ N proof the stress suite exercises.
- Packages — content addressing, the lockfile, and what the rebuild gate seals.
- Guarantees — the same properties, stated for the reader who must trust them.
- Inference — why deterministic inference is a prerequisite for all of this.