Memory model
Flux has no garbage collector, no runtime allocator, and no way to run out of memory while running. That is not an optimization — it is a consequence of the language. Every buffer has a const-folded size, every lifetime is statically exact, and the compiler therefore does not check a memory ceiling: it computes the allocation. A program that would exceed its budget is rejected before it runs, never killed while running.
This page describes how values are represented, how the data path is laid out, how the liveness plan turns a graph into a memory map, and which bounds are enforced where. It distinguishes what is implemented today in the analysis-plane backend from what the sealed design specifies beyond it, and marks each of the latter where it appears.
New here? Start with Guide §2 — Why Flux is built this way → — the static-budget pillar in plain terms, before the exact layout.
Value representation
At runtime, every scalar is an IEEE-754 double. Kinds are a compile-time discipline: once
the dimension has been checked, a price and a level are both an f64. Nothing about the
kind survives into the data path — which is precisely why a dimensional type system costs
nothing at runtime.
na and its canonical bit pattern
na is a NaN. That is convenient — arithmetic propagates it for free — and it is a trap,
because the WebAssembly specification leaves the sign bit and payload of a produced NaN
undetermined (0/0, sqrt(-1), ∞ − ∞). Two engines could therefore store different bytes
for the same absent value, and byte-identity would break silently: no program can observe a
NaN payload (is_na is a x ≠ x test, payload-insensitive), so only a byte-level oracle would
ever see it.
The rule, implemented on both sides:
- Per-bar values live in machine registers (WASM locals, the interpreter’s stack). There, a NaN may carry any payload — no operation in the instruction set observes one.
- At every observable boundary — a sink column write, a snapshot, a hash, a serialization —
the value is forced to the single canonical quiet NaN
0x7FF8000000000000. In the WASM backend this is ani64.storeof the bit-exact twin of the interpreter’s canonicalizing store.
So na behaves like an ordinary absent value in the language, and like one exact bit pattern
everywhere it can be compared.
Decimals
decimal(scale) is a scaled integer, not a float. The backing width follows the declared
precision — i64 for up to ~18 digits (a native WebAssembly type, eight bytes, the fast path),
i128 by default, i256 for the largest crypto magnitudes.
Storage width is not compute width. A product promotes its intermediate (i64 × i64 → i128)
and then re-quantizes to the declared precision of the destination, so an overflow is
impossible to hide. Exceeding the declared bound yields na plus a diagnostic — never a
wraparound, never undefined behaviour.
Why the threshold follows the declaration. The numeric digits of a result do not depend on the backing width; only the point at which “this is out of domain” fires does. Declaring
decimal(18, s)means “beyond this, it is a domain error, not a bigger number”. Because that threshold decides when annaappears, the declared precision is part of the script’s hash: two parties replaying the same program must agree on when absence begins.
Strings
A string is immutable, UTF-8, and bounded by a declared cap. Its unit is the Unicode
scalar — never a byte, never a UTF-16 code unit — so that indexing and slicing agree across
engines, and truncation always cuts on a scalar boundary.
Most strings in practice are short (a label, a formatted price), so they live inline in the value — a small-string optimization, zero allocation. Longer ones go into a bump arena reset once per evaluation tick (per bar, or per frame): no GC, because purity plus bounded lifetimes make the reset always safe.
One rule keeps that safe under replay: a string that survives its tick — captured by a
scan, stored in a Model field, written into a checkpoint — is materialized out of the
arena (copied), never left as a view into memory that is about to be overwritten. Without
that, scrubbing backwards in the debugger would read an arena rewritten a thousand times since.
Aggregates
| Kind | Representation |
|---|---|
vec(κ, N) |
a contiguous span of N elements of κ. N is a capacity, so a shorter vector inhabits a longer one with an na tail. |
record{…} |
a flat struct — in the data path, a group of parallel columns, one per field |
variant{…} |
a tag plus its payload |
Map, Set, Deque, Tree |
bounded arena structures, ordered, no hashing — see collections |
The data path is columnar
The engine evaluates a graph over bars, and it does so column by column, not row by row.
Each node produces one Float64Array of length equal to the bar capacity; sinks carry their own
columns. Records are struct-of-arrays: a bollinger node is three columns, not an array of
three-field objects.
Figure — nodes produce columns; a windowed kernel rings its history, and most rings derive their position from the bar counter while two families persist a cursor.
Two consequences worth naming:
- The hot loop allocates nothing. Every buffer is allocated once, from the graph, before the first bar. There is no allocation, no free, and no fragmentation while stepping.
- Most ring positions are derived; two families keep a cursor. A windowed kernel holds its
history in a ring of its period
p. For the cursorless kernels — the sliding sums (sma,sum,stdev,bollinger), the pure ring-scans (wma,cci,change,stderr,percentrank) and the weighted windows (winw,alma) — the slot for barkisk mod p, read straight off the bar counter, and there is no write index to persist because there is none to keep. Two families are genuinely different, and the difference is exactly the state a checkpoint has to carry:rsirings the variations, not the bars. It pushes one only on a bar whose predecessor is finite, so its write cursor counts pushes, not bars — a hole in the data advanceskand leaves the cursor where it was. It is therefore not a function ofk, and it is persisted, alongside the running averages.- The monotone deques —
highest,lowest,aroonup,aroondown— persist a head and a tail. The ring positions are derived from those two counters; the counters themselves are state, written back on every bar.
The state plan counts each of them: a derived position costs nothing, a cursor costs a cell. And
since a checkpoint is a memcpy of the whole state region, those cells travel with it — the ring
of an rsi and the head and tail of a highest are in the snapshot, not reconstructed from the
bar counter on the far side.
The linear memory map
The compiled module owns one linear memory, internal and exported, with this layout:
Figure — the plan is the allocation: min = max pages, and growth is impossible by construction.
[ header: bar count · write index ][ state cells ][ 6 bar columns ][ sink columns ]The header is two i32 fields, and the second earns its place. The first is the bar count.
The second, four bytes further in, is the write index of the columns — how many bars have
actually been written. In a batch run, and in a live run over history, the two are equal and
nothing observable separates them. Under a window they part company: the write index is the
count of window bars written, and every column read follows it, not the bar count. They share
one header cell, so a snapshot and a reset cover both at once.
The sink columns are the observable output, and their set is closed: plot, mark, fill,
colorBars, alert, assert.
Everything in it is sized at compile time, from the declared bar capacity and the resolved
parameters. The memory is declared with min = max = ceil(layout / 64 KiB) pages, so growth
is not merely unused — it is impossible. The ceiling is not a runtime check; it is the
allocation.
This is what makes the budget honest. A pane’s footprint is known before it opens: the bounded Model, plus the view arena, plus the graph’s own plan.
Why parameters are resolved at compile time. The unit of compilation is (graph, resolved parameters, bar capacity). Changing a knob recompiles and re-gates. The reward is a 100 % static state layout, an exact
min = maxmemory, and — the load-bearing part — a byte-identity gate that validates the very bytes and the very instance that will serve. There is no gap between what was validated and what runs.
The liveness plan
The graph is pure, total, causal and free of aliasing, and every buffer has a const-folded size. Therefore every buffer’s lifetime is statically exact: its first use and its last use can be read off the schedule. The compiler exploits that with a mandatory liveness pass.
Figure — two buffers whose intervals do not overlap share one slot; the reported footprint is the peak of what is simultaneously live, never the sum of everything allocated.
The pass has three steps:
- Compute each buffer’s interval
[first use, last use]on the canonical topological order — the single linearization obtained by breaking every tie between ready nodes with the pinned lexical node identity (the same identity that anchors hashing and the random generator’s draw index, invariant under inlining, dead-code elimination, common-subexpression elimination and recompilation). - Classify. A buffer is live-out — kept for the whole instance — if it feeds an
observable output, or if it survives its tick (captured by a
scan, a Model field, a checkpoint). Otherwise it is transient and recyclable the moment its last use passes. - Colour the intervals. Transients are assigned slots by greedy interval colouring within a size class: two buffers whose intervals are disjoint share the same slot, in an assignment order pinned to the canonical rank. Never a first-seen index, never a hash iteration order, never an address.
The reported footprint is then the peak of simultaneously-live bytes, not the sum of every
buffer that ever existed. (The maxNodes × N_max product remains a sum bound — it guarantees
termination and compile-time rejection, and it is deliberately not the same number as the peak.)
In the WebAssembly backend the plan does double duty: one f64 local per liveness slot, so
the plan is literally the local-allocation map, and the peak is far below the node count —
especially once several scripts are merged into one graph.
In-place donation (sealed design — the shipped planner today shares slots only between
disjoint lifetimes; donation is specified but not yet emitted). A functional node (vec.setAt, a
column derivation, a record update) with a single consumer at its last use would write in
place into the slot of its dying input instead of copying — opt-in, and gated by the translation
validation so that an unsafe donation (two consumers, or a source still live) is a build
failure, never a silent overwrite.
Why the plan itself must be deterministic. The value oracle compares outputs; it is blind to layout. Slot sharing is value-invariant (a slot is only reused after its occupant’s last use, so no read ever sees an overwritten value) — so two different plans would produce identical outputs and the oracle would notice nothing. But the plan is baked into the emitted module. If it were not a deterministic function of the graph, two compilations of the same program would differ in bytes. The plan therefore joins the pinned routines — the pinned mathematics, the decimal routine, the Unicode tables, the calendar conversion, the random generator — as something that must be identical on every engine and every machine.
Bounds and budgets
| Bound | Value | Nature |
|---|---|---|
N_MAX |
10 000 (browser target) | the maximum window, period or delay of a kernel |
N_MAX_SERVER |
100 000 | the same bound, server/backtest target |
MAX_NODES |
3 072 | the size of the graph — judged on the graph after inlining, elimination and common-subexpression elimination |
N_ACTIVE_MAX |
16 | co-active scripts merged into one global graph |
maxBricksPerBar |
1 000 | the cap on how many boxes one time bar may cross in a price-driven representation |
Three properties of these numbers matter more than the numbers themselves.
The window bound is a validation bound, not a compute parameter. It decides acceptance; it
never enters a computation. A program that compiles under two different values of N_MAX
produces identical bytes, and its memory is sized by the periods it actually uses. That is
what lets the browser and the server carry different ceilings without forking the language: the
compiler records the program’s real maximum length, and the loader gates
environment ≥ program. A server pack with a 50 000-bar period is refused outright in a
browser — cleanly, at load — instead of failing at instantiation.
The graph bound is judged on the graph. The front-end guards (on the syntax tree, on the inlining expansion) sit at 64× the graph budget: they are a compiler anti-abuse measure, not a verdict on your program.
The compile verdict is deterministic counters only. Acceptance or rejection is a pure function of the source. There is an interactive build timeout in the editor — around two seconds — but it is a cancellation, never a verdict. A wall-clock verdict would mean the same program is accepted on one machine and rejected on another, which would break replay and anti-cheat at the root.
There is likewise no per-bar timeout. Runtime cost is statically bounded by
maxNodes × N_max × maxBricksPerBar; totality gives termination, and those ceilings give the
practical bound. Exceeding the budget is a compile-time rejection — a program is never killed
mid-bar.
Isolation: panes, workers, arenas
The worker layer described here is shipped: the graph is partitioned, scheduled and run across a pool today, and the 1≡N proof in Concurrency is the reason that is safe. Two things in this area are sealed design rather than shipped runtime — the application realm (the per-pane Model and its view arena) and the shared compute module. Both are called out where they appear below, and neither describes the current runtime.
One module instance per task, with its own memory. This is the shipped isolation, and it is
the strong form of it. A task is handed the compiled bytes and instantiated on the worker that
will run it. The module declares its own internal linear memory, min = max pages, and it is
instantiated against function imports only — the pinned transcendentals, nothing else. It
imports no memory, and there is no memory for it to import: not one linear memory in the runtime
is declared shared.
Shared, read-only: the six input columns. The host writes time, open, high, low, close and volume once, into a single shared array buffer, and every worker reads its bars from there. That is the whole of the sharing.
Owned, write: everything a task produces. A sink column is computed in the module’s private
memory and handed back by buffer transfer — zero-copy, and the sender loses the buffer as it
gives it away. Never by sharing. Which is what lets the next rule be absolute: no atomic ever
touches data. Atomics serve only the scheduler’s claimed and done counters, in a small
buffer of their own.
Arenas are per worker. The scratch arena for transient buffers is private to each worker, sized to the peak of that worker’s sub-graph. Under the rule above it could hardly be anything else — nothing writable is shared, so there is no shared arena to fall into.
Why per-worker arenas are load-bearing. The liveness plan reads disjointness on the sequential canonical order. Two buffers may legitimately share a slot because their intervals do not overlap there — and yet, under a dynamic scheduler, the nodes that own them can execute at the same time on two workers. A shared arena would give them the same address: a write-write race, and a byte-identity failure that no value oracle would catch. Private arenas make the aliasing impossible by construction rather than by scheduling discipline.
A shared compute module. The eventual design has a different shape, and it is worth stating precisely so that nobody reads it back into the runtime. One shared compute module — the kernels, the engine core — instantiated per worker against one shared linear memory, each worker addressing its own region by offset; each application module then imports it instead of recompiling the kernels into itself, and so carries only its own logic: small, fast to instantiate, independently invalidated. None of that ships. Today a kernel that carries state between bars is inlined per node into the module that uses it, and the pure ring-scans that remain shared functions are shared within one module, never across two. The rewrite would trade zero writable sharing for region-disciplined sharing and reopen the snapshot, the windowing and the verification machinery to do it — for a future the current runtime does not need yet. The reasoning is set out in Concurrency.
One instance per pane. An application pane is a WebAssembly module instantiated for that pane — its own linear memory, no DOM, host access only through vetted imports. Closing the pane releases the instance and everything in it. Two levels, deliberately distinct: the application realm (a Model, a view arena) is isolated per pane, while the compute pool runs the graph. Confusing the two is the classic mistake here.
Checkpoints and snapshots
Because all state lives in one contiguous region with a static layout, a checkpoint is a memcpy of that region — and a restore is the same copy back. Resumption is byte-exact, which is what makes debugging time-travel, live-preview scrubbing and server-side replay the same mechanism rather than three approximations of one.
At the serialization boundary the canonical na rule applies, so a snapshot taken by one engine
is byte-identical to one taken by another.
The APP plane: bounded models and slotmaps
An application’s Model admits only bounded kinds — so its footprint is computable at compile
time, exactly like a graph’s. Its variable collections use the slotmap pattern: a bounded
vector, tombstones instead of compaction, a live mask, and a free list held in a parallel
index vector so that the tombstone itself is never overwritten. Nothing is ever moved, so no
handle is ever invalidated, and the memory plan stays flat.
See App plane.
What does not exist
No garbage collector. No runtime allocator. No fragmentation. No out-of-memory at runtime. No “it was fine on my machine”. A program either fits its declared budget — and then it fits it on every machine, byte for byte — or it does not compile.
See also
- Compiler and runtime — the pipeline, the byte-identity gate, and the pinned routines.
- Optimizer — the passes that shape the graph the plan is computed from.
- Concurrency — the scheduler, shared memory, and the 1 ≡ N proof.
- Kinds — where
na,decimal,stringandveccapacities come from. - collections — the bounded arena structures.
- App plane — bounded Models, slotmaps, journals and checkpoints.