◆ Flux

Compiler and runtime

Flux runs on two engines: a graph interpreter, which serves the editor, the live preview and the debugger, and a compiled WebAssembly module, which serves the run. They are not an approximation of each other. They produce the same bytes, and that equality is checked at every single compilation, on hostile data, before anything ships.

That one invariant is the spine of everything else on this page. It is what lets the engine be swapped underneath a running chart with no visual glitch, what lets an optimizer be aggressive without being trusted, and what lets a server re-execute a client’s work and detect a lie.

Together the two engines are one abstract machine — the FVM, the Flux Virtual Machine: a total, sandboxed, deterministic dataflow machine whose instruction set is the kernels and the graph operations, whose memory is the static linear layout of the memory model, and whose arithmetic is pinned to the byte. Neither the interpreter nor the WebAssembly module is the FVM; each is a way of running it, and I7 (below) is the guarantee that the two ways cannot disagree. It is a virtual machine in the exact sense the JVM is — a single semantics with more than one conforming implementation — not a bytecode loop: the “instructions” are a typed graph of kernels, and the two implementations lower it differently but must land on the same bits.

New here? Start with Guide §11 — Determinism, replay and trust → — the accessible version of the two-engine contract, before the machine-level rules below.

The pipeline

source
  │  parse            Lezer — incremental, total: an arbitrary input yields one tree or a clean error
  ▼
  │  resolve + inline names and `def` bodies  (the call graph is acyclic, so inlining terminates)
  ▼
typed DAG            kinds inferred bottom-up; presentation derived from the kinds
  │  causality check  every feedback cycle must cross a unit delay
  ▼
  │  optimize         common-subexpression elimination · dead code · constant folding · fusion · kernel selection
  ▼
  │  memory plan      liveness intervals → slots → an exact footprint
  ▼
emit ──┬─→ interpreter closures     (edit · preview · debug · THE ORACLE)
       └─→ WebAssembly (Binaryen)   (run · distribute)

The compilation pipeline Figure — one graph, two back-ends, one gate between them.

The interpreter pays once, when the graph is built. After that the hot loop is native kernels over pre-allocated columns.

The unit of compilation is (graph, resolved parameters, bar capacity). Parameters are resolved at compile time — changing a knob recompiles and re-gates. That is a deliberate trade: it buys a fully static state layout, an exact min = max memory, and a gate that validates the very bytes and the very instance that will serve. There is no gap between what was verified and what runs.

Two engines, one contract

Interpreter WebAssembly module
Serves editing, live preview, the dataflow debugger the run, and distribution
Feedback instant — no compile step in the loop compiled and gated, then swapped in
Role in the contract the oracle the candidate

I6 — a leaf is byte-identical to its kernel

A node that maps to a native kernel produces exactly the bytes that kernel produces, warm-up included. Flux does not impose its own na-until-N convention: a Flux indicator on a clock is the same citizen as a built-in, from the first bar. That is what makes “rewrite the catalogue in Flux” a safe proposition rather than a rewrite of every golden.

I7 — the interpreter and WASM agree, byte for byte

The gate runs at every compilation and it blocks:

The I7 oracle Figure — the gate compares the two engines on adversarial data, in batch and live, before a single byte is allowed out.

Why the live path is in the gate too. Batch equality is the easy half. The live path steps one bar at a time through a different code path in the module, and it is exactly where a subtle state bug hides. In the emitted module, one shared body serves both — the batch range and the single-bar advance are the same code, called with different bounds — so live ≡ batch by construction; the gate then checks the one seam that remains (column bases versus the live scratch) rather than trusting it.

What it buys, concretely

The browser path, as built

The main thread never imports the compiler. It orchestrates:

  1. A registered script serves immediately, on the interpreter — the proven path, with no compile step between typing and seeing.
  2. The service compiles in a worker, where compilation and the I7 gate are atomic: divergent bytes are never handed back, because they are never handed back at all.
  3. When the gated module exists, the service upgrades silently. By I7 the bytes are identical, so the swap requires no re-render and produces no visible event.

The compiler bundle is fetched lazily, on editor intent — never on the chart path, and excluded from the offline precache. Compiled modules are cached in tiers (attached instances by script and configuration; modules by a canonical key, in a bounded cache), because the expensive part is compilation, not instantiation.

Determinism: the pinned-routine discipline

Byte-identity does not survive first contact with a standard library. Two engines can disagree about the last bit of a logarithm, the sign of a zero, or the rounding of a half-integer, and every one of those disagreements is enough to break replay. So Flux pins the routines — the same code, on both sides, with no delegation to the platform.

Routine Pinned to Why the obvious choice is wrong
log exp sin cos tan atan atan2 pow a pinned WebAssembly libm, used by both engines two JavaScript engines differ by ≥ 1 unit in the last place
% the pinned library’s remainder, linked directly module-to-module the remainder itself is exactly specified — what is pinned is the binding: routing it through the host instead would reshape a produced NaN, and the canonicalization has to happen in one place
round f64.nearest — ties to even half-up rounding disagrees on every half-integer
min / max the exact na-absorbing selection chain the native instructions propagate NaN and return −0 for min(−0, +0) — both disagree with the language’s rule
NaN constants emitted as an integer bit pattern, then reinterpreted a NaN marshalled through a host API has no guaranteed bits
na at rest the canonical quiet NaN 0x7FF8000000000000 WebAssembly leaves a produced NaN’s sign and payload undetermined
decimal one shared multi-limb integer routine a floating-point engine has no native 128-bit integer; any emulation would diverge
string, fmt.* pinned Unicode tables and one canonical formatter platform string length is in UTF-16 units; platform number formatting differs in the last digit
calendar a pinned epoch ↔ civil routine, pinned time-zone data two correct implementations still disagree on DST gaps and end-of-month clamping
rand(seed) one pinned counter-based integer generator; in v1, rand()/rand(seed) live on the presentation plane only, and the deterministic-domain form lands after integer arithmetic is bit-identical for free; floating-point mixing is not
ordering of na in sorts one pinned total order (absent values last, stable by index) a partial comparator leaves the order to the platform’s sort
the memory plan a deterministic function of the graph the value oracle is blind to layout, so nothing else would catch a divergence

Why the interpreter does not call the platform. It is written in TypeScript and runs on a JavaScript engine, so Math, Number, String.prototype, Date and Intl are right there. Using them would make the interpreter agree with itself and disagree with the module — and the disagreement would be invisible to a program (nothing observes a NaN payload; nothing observes the last bit of a logarithm) and visible only to a byte-level oracle. The discipline is: the interpreter runs the same pinned routine the module runs.

The toolchain is pinned too

The lowering to WebAssembly goes through Binaryen, and it must be a deterministic function:

Optimization is mandatory, not optional — an unoptimized module is not a “safer” module, it is a different module, and the whole point is that there is only one.

The emitted bytes are a pure function of (program, toolchain), so the toolchain has a single identity — the compiler version, the Binaryen pin, the identity of the pinned math routines, and the enabled feature set — and every cache key for a compiled artifact carries the whole of it. Without that, bumping an emission rule would keep serving stale bytes that a recompilation no longer produces: a cache that is keyed on less than the toolchain is a silent divergence with a long fuse.

Minification and obfuscation, where used, are deterministic and applied after the gate, on the distribution artifact. The provenance hash and the rebuild gate are defined on the artifact that actually ships.

The runtime surface

The compiled module exposes the same shape the native engine already uses:

Surface Meaning
make instantiate and reset
advance(bar) the live step
run(n) the batch range
snapshot / restore a copy of the state region — a checkpoint is a memcpy, and resumption is byte-exact

One body serves run and advance, so the live and batch paths cannot drift apart, and a checkpoint is a contiguous copy rather than a bespoke serializer.

Why WebAssembly

Not for speed. The honest measurement: fusing kernels with glue gains nothing in batch (the boundary is amortized over a long history), SIMD is excluded by byte-identity (a horizontal reduction reassociates floating-point and changes the bits), and scalar f64 compute in a modern JavaScript engine is already within a small factor. Real speed comes from algorithmic work and native kernels, not from the execution language.

WebAssembly is the target for four structural reasons:

  1. One execution artifact from day one. The interpreter and the module must agree bit for bit; freezing that equality at the start is far cheaper than retrofitting it.
  2. Cross-machine determinism. WebAssembly’s floating-point semantics are strictly specified — which is what makes server-side re-execution meaningful rather than “same runtime, probably”.
  3. Application panes. One execution and distribution format for indicators, representations, drawings, scenes, transitions and application logic.
  4. An opaque, sandboxed distribution artifact. A shared or purchased script ships as WASM, never as source: the intellectual property is protected, and the consumer’s trust boundary is a binary plus a sealed manifest.

The honest frontier. WebAssembly computes. It never paints: it produces geometry, a draw-list, and a view tree in linear memory, and the host — JavaScript, or the GPU — does the painting. There is no Web API reachable from the module. Rendering runs entirely on WebGPU: scene graphics — the geometry and the draw-list — and text alike, the text as an SDF glyph atlas that stays crisp at any zoom and DPR. A software rasterizer inside WebAssembly was considered and rejected: it loses that sharpness and is slower than the GPU. The render targets are specified in display; what the module guarantees is narrower and exact — a deterministic scene, the same numbers and the same draw-list on every machine, never identical pixels across GPUs.

eval and dynamic code generation remain forbidden. WebAssembly is not “eval with extra steps” — it is a sandbox with linear memory, no DOM, and host access only through vetted imports, admitted by its own dedicated policy token.

Budgets — counted at compile time, never timed at run time

A budget that is enforced by a stopwatch is not a budget: the same script would be accepted on one machine and killed on another, and replay would die of it. Every ceiling in Flux is therefore a counter, evaluated on the graph, before anything runs.

Ceiling Value What it bounds
N_max 10 000 in the browser; 100 000 on the server and in backtests the const length of a window or a vec. Beyond it: [ErrTotal]
maxNodes 3 072 per script the size of the graph
maxBricksPerBar 1 000 how many re-binned units one bar may produce (Renko, P&F). Beyond the cap the host aggregates rather than blow the budget
N_active 16 co-active scripts per chart how many scripts merge into one shared DAG; the aggregate ceiling is N_active × maxNodes
memory per instance the structural bound maxNodes × N_max the legal worst case. The declared footprint is the exact liveness plan, which is far smaller

The numbers are the least interesting part of that table. What matters is how they are enforced.

The graph is the authority, not the text. maxNodes is judged on the DAG after inlining, common-subexpression elimination and dead-code elimination — the graph that will actually run. The front-end guards (AST size, inlining expansion) sit far above it and exist only to keep a hostile input from exhausting the compiler; they never pronounce a budget verdict. nMax likewise never touches a computed byte: browser and server differ in what they accept, never in what they produce.

The memory ceiling is not a check — it is the allocation. The module declares its linear memory with min = max, sized from the liveness plan (§ memory model). Growth is not forbidden at run time; it is impossible. The static part is verified at compile time and the remaining lengths at instantiation — never while a bar is being stepped.

There is no per-bar timeout. Runtime cost is statically bounded by maxNodes × N_max × maxBricksPerBar: totality gives termination, and the ceilings give the practical bound. An over-budget graph is rejected at compilation, not killed mid-frame — the runtime guard below is defence in depth, not the enforcement path. A build does carry a wall-clock timeout, but it is an interactive cancellation in the editor, never a verdict: accept and reject stay a pure function of the source.

When the aggregate frame budget is nevertheless exceeded — many co-active scripts on one chart — the host applies a deterministic degradation policy, throttling or pausing low-priority scripts in an explicit declared order. The order is never data-dependent, so the degradation replays like everything else.

Where the work runs is a counted decision too. The service estimates a graph’s cost in cost units and dispatches to a worker fleet only above a calibrated threshold; below it, a merged single-threaded pass finishes before a fleet would have started. The estimate is a pure function of the graph, so the routing is reproducible — and by 1 ≡ N the choice cannot change a byte either way.

Fault isolation

A script that fails at runtime — an over-budget graph, a genuine error on a valid graph — is quarantined: its node is marked and removed from the active graph, without killing the worker and without disturbing its neighbours, which share the same instance map. Restarting the worker is the last resort, and the neighbours resume from their last checkpoint.

A NaN is not a fault. It is na, and it is displayed as a gap.

Verification and reproducible builds

I7 is one clause of a larger contract. Around it sits a verification harness — a first-class deliverable, not a folder of tests — whose sub-suites each declare an oracle, a corpus, and whether they block the ship: deterministic goldens, the formal properties (principality, confluence, totality, causality, the peak-≤-sum memory plan), a total-parser fuzzer driven by a well-typed generator, the three-way differential oracle (interpreter ↔ WASM ↔ native kernel, covering I6, optimized ≡ reference and I7), the enumerated metamorphic relations, the 1 ≡ N stress suite, lattice enumeration, and a model-checked capability monitor. Above them, reproducible builds make the emitted module a pure function of source, lockfile, compiler version, pinned routines and the canonical memory plan, sealed by a rebuild gate that recompiles the same inputs on different machines and asserts byte-identity — the condition server-side replay silently depends on.

One subtlety carries over into every byte-identity claim on this page: the three-way oracle calls the same pinned routine on all three sides, so it is blind to a bug inside a pinned routine. That is why each pinned routine also ships a second, independent reference implementation, compared bit for bit on fuzzed input — the oracle catches disagreement, only a second implementation catches a shared mistake.

Canonical: the full suite table with its blocking verdicts, and the build-hash inputs and rebuild-gate procedure, live in Verification and reproducible builds. This section is a summary, not a second copy.

Performance — measured, and honest

Why WebAssembly argued that the target was not chosen for speed. The certification benchmark settles what the speed actually is — measured on one Apple M4, over a hundred thousand bars, with the TypeScript, interpreter and WASM legs proven byte-identical before a single timing is taken. So every row below compares the same algorithm: the TypeScript leg is the interpreter’s own kernels, called directly on a plain Float64Array, not a naive rewrite. And against that, in batch — the path every chart takes — the compiled module is faster everywhere, by a margin that grows with how much intermediate data the indicator moves:

Workload — the same algorithm on both sides How much faster than TypeScript
Trivial O(1) kernels — rsi, change ≈ parity — V8 compiles a tight f64 loop nearly as well
Weighted scans and deques — wma, alma, highest ×1.3–2
Realistic charts — a classic eleven-plot chart runs ≈ 21.5 ns/bar ×1.6–2.5
Order-p scans and reductions — percentrank, kama ×2–3
Multi-stage composites — fisher, connorsRsi ×3.7–5
Heavy multi-output composites — stdErrorBands ×9–14

How much faster the module runs than the same algorithm in TypeScript Figure — the same algorithm on both sides: the batch gap grows with how much intermediate data the workload moves, from parity on trivial kernels to ×9–14 on heavy composites.

Why — the datapath, not the engine. Decompose the gap and the execution engine is the small part of it. Wasm is machine code with single-instruction f64 ops and no dynamic bounds-checks on a linear memory proven in range at compile time — but V8’s TurboFan compiles a hot, monomorphic typed-array loop nearly as well, so the engine difference alone is only ×1.1–2, which is exactly why the trivial kernels land at parity. The lever is the datapath. A multi-stage indicator hand-written in TypeScript materialises a full Float64Array per stage, writing then re-reading its values each time; the flux compiler fuses every stage into one pass, holding the intermediates in locals and rings. stdErrorBands’ TypeScript leg allocates eight arrays; the module allocates none — and every materialised intermediate is two full memory passes of pure overhead. That, not the instruction set, is where the ×2-to-×14 comes from. (Compilation itself is tens of milliseconds, almost all of it Binaryen lowering the graph to bytes — an editor-time cost, paid once, never per bar.)

And beneath it, the algorithms keep the bytes. A rolling maximum reads a monotone deque instead of rescanning its window; a reset touches only the live geometry, about 28.5 µs down to 0.8 µs; a bounded length knob shrinks a script’s state by roughly twenty times; an sma and a sum that coincide share one ring. Every one is an O(n) change with the same output bits — a win the optimizer and the native kernels deliver in any language, and one more reason the figure above is not a story about the execution engine.

The one case TypeScript wins — and why it is not a headline. Stream a single O(1) kernel that has a native f32 stepper — an ema or rsi advanced bar by bar — and the native library wins, ×5–8: flux pays a call across the WASM boundary on every bar where the library takes a single step. But that leg runs in f32, not f64 — a different-precision result, recorded as a documented divergence and never quoted as a speed headline. And it exists only where a native stepper does: a windowed wma or a deque highest has none, so that path recomputes in batch, where the module takes it all back.

Why SIMD is off the table — by choice. The one lever that would move scalar compute is SIMD, and byte-identity forecloses it: f32x4 packs four lanes into one, and a horizontal reduction reassociates floating-point and drops precision, so the bits change. Same-bits-everywhere is what no-repaint, replay, the goldens and the server’s re-execution all stand on; trading it for a fraction of a factor would spend the property that makes these numbers mean something. It is a decision, recorded and enforced, not a feature still to come — the same line the optimizer draws one level down when it forbids reassociation.

So what WebAssembly is for. The speed is real, but read the decomposition again: it comes from the fused dataflow the compilation model gives you — not from WebAssembly as an execution engine, which is only the ×1.1–2. The reasons to target WebAssembly specifically are the four structural ones in Why WebAssembly: one execution artifact, so the interpreter and the module can be held equal from the start; floating-point specified across machines, so a server can re-run a client’s work and catch a lie; one format that carries an indicator, a scene and a whole application pane alike; and a sandboxed, opaque binary to distribute, source withheld. Flux is fast because it compiles a fused dataflow graph — and it compiles that graph to WebAssembly for correctness, determinism and distribution. A benchmark you can reproduce is a better argument than a superlative.

See also