◆ Flux

Determinism, replay and trust

There is a moment every serious analyst reaches: a backtest looks brilliant, and you have to decide whether to believe it. Did the signal really fire on that step (the bar, in charting), or did it quietly read a price that had not printed yet? Would it compute the same thing on your colleague’s laptop, on the server, tomorrow? And the harder question underneath: could you run someone else’s strategy — one you did not write and cannot fully read — without handing it the keys to your machine?

This chapter is about the property that answers all four at once. Flux is deterministic to the byte, and almost everything you can trust about a Flux program falls out of that single fact: history that cannot be rewritten, replay that reconstructs a run exactly, tests that mean what they say, and a sandbox strong enough that installing a stranger’s code is informed consent to a short list rather than an act of faith. We will build the idea up from the bottom, and — because a trust page that lists only strengths is a sales page — name the limits with the same confidence as the strengths.

The replay guarantee Figure — the same source produces the same bytes on every engine, which is what lets a run be replayed, golden-tested, and re-checked by a server.

Same numbers on every machine

Start with the load-bearing property. The same program on the same data produces the same bytes — between the interpreter that runs while you type and the compiled WASM that ships, between two runs, between an ARM laptop and an x86 server.

FLUX
r = math.log(close / close[1])        // ratio in, dimensionless out — via the pinned libm, same bits everywhere
dot { at: (bar.i, close), glow: 8 * rand() }   // per-frame randomness — canvas-only, never in the replayable numbers

The comment on the first line is doing real work. “Roughly equal” is not a property you can build on; byte-equality is. Floating point is deterministic if and only if every source of variance is pinned, so Flux pins all of them as language policy rather than author burden: scalar f64 with no SIMD and no reassociation in the deterministic core, a fixed reduction order, one pinned libm build for every transcendental (never the host engine’s Math, which legitimately differs by an ULP between browsers), one shared routine for decimal, Unicode, calendar and the seeded PRNG, and a single canonical bit pattern for na so two engines never disagree even on the bytes of “nothing”. The second line matters too: rand is a presentation generator — per-frame and non-replayable — so it exists only on the canvas side of the firewall, where determinism is deliberately not promised; analysis code that reads it is rejected with [ErrFirewall], and the numbers stay replayable precisely because the one non-replayable thing cannot reach them.

None of this is assumed. The equivalence is verified at every compilation — the interpreter and the compiled module are run on real data and asserted bit-for-bit equal (this is invariant I7), and a divergence blocks the artifact. What you get in exchange is worth the discipline: replay-exact debugging where you can step backwards as reliably as forwards, golden tests with no tolerance to tune and no flakiness to excuse, and cross-machine agreement — which is exactly what lets a verifying server re-run a client’s work and get the same answer.

Note the precise claim, because it is easy to overstate. Determinism means same numbers, same scene — not same pixels. The compute is bit-identical and the draw-list the compiler emits is deterministic, but the picture on the glass is not promised identical across GPUs, and does not try to be. That is not a leak; it is the design of the next section but one.

History that cannot be rewritten

Byte-determinism gives you agreement across space. Causality gives you agreement across time — and it is the guarantee traders feel most directly. In any system that re-evaluates over growing history, the deadliest defect is the value that quietly changes retroactively: an analytic that looked prophetic on the chart because, at each past step, it had silently read data that did not exist yet. Both the live run and the backtest are “correct” — for different definitions of time — so no test catches the divergence. Flux removes the defect by making it inexpressible.

FLUX
prev  = close[1]                 // yesterday's close — legal, and na on the first bar
gain  = math.max(close - prev, 0)
peek  = close[-1]                // ✗ [ErrCausal] — a negative delay reads the future

Three rules produce the property. Delays reach backwards only — x[n] with a constant n ≥ 0; a negative index does not parse into a meaning, it raises [ErrCausal]. A resample reads the last closed unit of a coarser clock, never the one still forming. And every feedback cycle must cross a unit delay, so today’s output may depend on yesterday’s output but never on itself. By induction over the graph, output[t] = f(inputs[0..t]) — a past value has nothing left to depend on, so nothing can move it. This is no-repaint: a value, once produced for a step, never changes, and it holds for every accepted program, not a careful subset. Live and historical evaluation are the same function computed over the same inputs, so they produce the same bytes — the honest backbone of any evaluation.

There is exactly one exception, and it comes with a wall. live(e) re-evaluates an analysis sub-graph including the step in formation — genuinely useful for watching an indicator move within the current step — and it may flow only to display sinks. Feed it into an alert, an assertion or a calculation and you get [ErrFirewall]. A script that uses it is flagged non-replayable, visibly, so you always see the trade-off you made. The firewall that enforces this is the subject of Guide §8; here it is enough to know that the one value that can still change is kept out of every number a decision rests on.

Bounded, so it is safe to run

Determinism assumes the program finishes. Flux guarantees it: every program terminates, and its cost per step is known at compile time. There are no unbounded loops, no unbounded recursion, no unbounded collections. Iteration exists — window with map/fold for counted loops, scan for running state, loop(max, …) for “iterate until done” with a declared ceiling — but every bound is a compile-time constant under a global cap. A program that cannot state its bound, or that exceeds the budget, is rejected before it runs with [ErrTotal], the offending bound named. It is never killed mid-run by a timeout, because there is nothing to time out.

One subtlety is worth pausing on, because it is where determinism has to start. The accept/reject verdict is itself a pure function of the source, decided by counters alone — never by a clock. The editor’s build timeout, on the order of two seconds, is an interactive cancellation to keep the UI responsive; it is never a verdict. A wall-clock verdict would be machine-dependent — the same script accepted on a fast machine and rejected on a slow one — and then two users would not be running the same language, and replay, which assumes that what compiled there compiles here, would break. Determinism has to hold at the compiler’s answer, or it holds nowhere downstream.

Flux is, deliberately, not Turing-complete — and it is worth being honest about what that costs and what it does not. Totality is the synchronous-dataflow trade, in the lineage of Lustre and SCADE: the language is total and bounded per step, while remaining unbounded over time — it runs step after step forever, it just cannot spin without limit within one. In Flux’s client-side, reactive domain that ceiling is invisible, because an unbounded inner loop there would not be a clever program — it would be a hung tab, i.e. a bug. What totality buys is the foundation everything else stands on: a static budget the host can honor, and static checks (kinds, causality, exhaustiveness) that are decidable because the language is total and its bounds are constants.

Replay comes for free

The properties so far are about analysis — pure series over closed steps. Applications add the one thing analysis cannot: state that persists between events and decides what is displayed — a score, a document, a selection, a blotter of positions. You might expect that to be where determinism frays. It is where it pays off most.

An application’s state never changes except through one deterministic, journaled reducer, so the model is nothing more than fold(init, [the journal of every message so far]). Undo rewinds the journal one step and folds again; redo re-extends it. Recall the counter from Guide §10 — its entire undo history is those two moves:

FLUX
variant Msg { Bump | Undo | Redo }

app tally {
  capabilities: [ journal ]

  init(p)        = { doc: { n: 0 }, ui: { pending: na } }
  update(m, msg) = match msg {
                     Bump -> { model: m with { doc: m.doc with { n: m.doc.n + 1 } }, cmds: [] }
                     Undo -> { model: m, cmds: [ Journal(UndoToMark) ] }
                     Redo -> { model: m, cmds: [ Journal(RedoToMark) ] }
                   }
  view(m)        = row {
                     button("undo", Undo)
                     text("{m.doc.n}")
                     button("+1", Bump)
                     button("redo", Redo)
                   }
  subs(m)        = []
}

Read the two undo arms carefully: they return the model unchanged and hand the host an inert command. The host — which is where the journal lives — truncates it to the target and re-folds into the new model. Nothing in the application code knows how to reverse an edit; it only knows how to move forward. Because the fold is a pure function of the journal, the state you land on is exactly the state you left — reproduced, not reconstructed. And because Flux is deterministic to the byte, “exactly” means bit-for-bit, including the answers that came back from the network: an async result entered the model as a message, so it is already in the journal, replayed verbatim rather than re-fetched. Rewinding never fires the request a second time; it reads the reply that already happened.

One mechanism wears two faces. For the user it is undo/redo that cannot silently forget a field or resurrect a stale selection — because the Model is split into a doc (the business state history owns) and a ui (the cursor, the draft, the in-flight request) that history does not, so undo rewinds doc and leaves ui alone. For the developer the same journal is a time-travel debugger: scrub to any past message, step backwards, set a data breakpoint — stop at the first message where score crosses 100 — and put the run beside a reference run to see where, if ever, they diverge. It is undo over an application’s execution, the events it processed, never over its source code. The full architecture — the doc/ui partition, checkpoints, schema evolution — is Guide §10.

What replay proves — and what it does not

Guide §10 first drew this limit; here it faces an adversary head-on, and the guide owes you a precise statement rather than a slogan. A verdict is a pure function of (init, [msg]), so a server can recompute it on the same bytes, and a claimed score that was not earned is a journal that does not fold to the claimed result. That property was not designed; it was inherited from determinism. But it proves one specific thing:

Replay proves the COHERENCE of a journal, not its NON-FALSIFIABILITY.

“Diverges ⇒ tampered” is complete for a score derived from the seed and the elapsed time, because the server owns both — the seed is derived server-side from (runId, level, qIndex) and never accepted from the client, and the elapsed time of a ranked run is host-stamped and substituted at re-fold, so a forged Tick count buys nothing. It is not complete for a third class, and the gap is open by name. A score fed by a host-pushed outcome — an OnReveal message carrying a result the host computed — enters the journal as data, and a re-fold replays data verbatim. The seed re-derives the messages that came from randomness; this one did not. So a forged outcome re-folds to the claimed result without a whisper of divergence: the check passes, and the claim is still a lie. The same class covers a pixel — a bounded pixel reading derives from the client’s own pan and zoom, and a server with no viewport cannot re-derive it. Hence the standing, language-level rule: a pixel value never feeds a ranked verdict. It is a readout, or it is cosmetic.

An outcome-fed run therefore has exactly two honest destinations and no third: the host re-derives the outcome server-side by re-running the kernel itself, or the run is local-score-only, excluded from the shared leaderboard. It is never accepted on the strength of the client’s journal alone.

Server re-execution is designed, and rests on exactly the determinism above — a server re-running a client’s work to catch a forged result. In v1 the native/server leg is verified client-side: re-execution lands with the server port of the grader.

Running code you did not write

Now the question underneath the others. Flux is the language of a platform where code is shared, sold and run by people who did not write it — an ordinary Flux target is untrusted code, and a program a language model emits is a clean instance of exactly that. That is only tenable if safety is a property of the language and host, not of a review process, and Flux gets there in layers.

The language has no primitive that does I/O — no fetch, no DOM access, no eval, no file handle — so a script’s only channel to the world is data it hands the host. Scripts hold no ambient authority; there is no global object in scope at all. An effect is an inert Cmd that carries values — a sound’s name, a storage key, a score — never a token, socket, URL or handle. The host holds every resource and runs the command only if the corresponding capability was declared in the app’s manifest and granted by the user. Emitting a command the manifest does not grant is rejected at compile time with [ErrCapDenied] — it never becomes a runtime event to be caught. And the manifest is not a promise the author writes: the compiler derives it from the emit Cap sites, as the union of capability needs over the whole transitive dependency closure, intersected with the user’s grant — so a dependency’s appetite for the network surfaces before install, no dependency can exceed what you granted to the whole, and no dependency holds a capability object it could re-delegate or amplify.

The consequence is the one that makes user-generated content routine rather than risky: security is the grant, not the code. The first-party interface and a stranger’s marketplace application run the same language in the same sandbox; they differ only in what has been granted. This is the third pillar as a runtime fact: a stranger’s island renders beside yours, and touches none of the data your own logic depends on. A trusted tier grants effects — it never loosens causality, no-repaint, totality or the firewall. Repaint is inexpressible for everyone. And trust is decided by the host from a server-side provenance record keyed by the binary’s content hash: a module cannot declare itself trusted, and no embedded metadata is believed.

Name the honest limit here too. The language is safe by construction; the capability monitor and the view sanitizer are ordinary code, and they are the residual attack surface — which is precisely why they are the components that earn a model check, and why the view primitives are a closed, typed set rather than a string a hostile view could smuggle markup through.

The optimizer you don’t have to trust

One more actor could break determinism: the optimizer. A Flux program is a pure, typed, total, causal DAG, which makes classic optimizations — sharing, pruning, reordering, specializing — safe by construction, and its small bounded size makes whole-graph searches affordable. The distinctive part is the trust model. The optimizer obeys one law: the optimized program must be bit-identical to the reference semantics — the unoptimized graph’s canonical evaluation. That law is enforced by translation validation: every compilation runs the reference and optimized artifacts on hostile data and asserts byte-equality. A miscompilation cannot ship; it can only fail loudly at the gate, where the compile falls back to the unoptimized path and turns the test suite red. The optimizer can therefore be aggressive precisely because no one has to believe in it.

What you get is free global common-subexpression elimination — write ema(close, 26) four times in a script and it is computed once — dead code elimination, constant folding and element-wise fusion, with the editor’s cost gutter showing the optimized graph, so what you read is what you pay. The default is bit-exact, always.

Sharing a subgraph across four co-active scripts — one merged graph for the whole chart — is a later optimizer tier, and relaxed floating-point (reassociation, fused multiply-add) is designed as an explicit, labelled opt-in, never a silent default and never bit-exact.

What the compile actually earned you

A guarantee whose only enforcement is a promise in a document is not a guarantee, so every property in this chapter is checked by a machine, and the editor tells you which ones your program earned. After a compile, the guarantees panel states it plainly:

✓ No-repaint     ✓ No look-ahead     ✓ Deterministic
✓ Bounded memory ✓ Byte-identical    ⚠ contains live() → non-replayable

That last line is not decoration: a guarantee you traded away should be visible at the moment you traded it. Behind the panel is a verification harness whose sub-suites each declare an oracle and a corpus — goldens, a differential oracle across interpreter, WASM and native kernel, a metamorphic suite of semantics-preserving relations, a 1 ≡ N worker stress, a capability monitor that asserts “no command outside the manifest is ever executed”, and more — most of them blocking. Reproducible builds sit under all of it: the build hash is a pure function of the source, the dependency closure, the compiler version, the pinned routines and the memory plan, and a rebuild gate recompiles the same inputs on different machines with different thread counts and asserts a byte-identical module. Server-side replay depends on this, because a value-level oracle cannot see the bytes emitted across two compilations.

The limits, stated plainly

Because a trust page that lists only strengths is a sales page, here is what Flux does not guarantee:

See also

The formal rules → This chapter narrated determinism, replay and the trust model. Their precise statements live in the Spec, FVM and FDK: