Determinism, replay and trust
There is a moment every serious analyst reaches: a backtest looks brilliant, and you have to decide whether to believe it. Did the signal really fire on that step (the bar, in charting), or did it quietly read a price that had not printed yet? Would it compute the same thing on your colleague’s laptop, on the server, tomorrow? And the harder question underneath: could you run someone else’s strategy — one you did not write and cannot fully read — without handing it the keys to your machine?
This chapter is about the property that answers all four at once. Flux is deterministic to the byte, and almost everything you can trust about a Flux program falls out of that single fact: history that cannot be rewritten, replay that reconstructs a run exactly, tests that mean what they say, and a sandbox strong enough that installing a stranger’s code is informed consent to a short list rather than an act of faith. We will build the idea up from the bottom, and — because a trust page that lists only strengths is a sales page — name the limits with the same confidence as the strengths.
Figure — the same source produces the same bytes on every engine, which is what lets a run be replayed, golden-tested, and re-checked by a server.
Same numbers on every machine
Start with the load-bearing property. The same program on the same data produces the same bytes — between the interpreter that runs while you type and the compiled WASM that ships, between two runs, between an ARM laptop and an x86 server.
r = math.log(close / close[1]) // ratio in, dimensionless out — via the pinned libm, same bits everywhere
dot { at: (bar.i, close), glow: 8 * rand() } // per-frame randomness — canvas-only, never in the replayable numbersThe comment on the first line is doing real work. “Roughly equal” is not a property you can build
on; byte-equality is. Floating point is deterministic if and only if every source of variance is
pinned, so Flux pins all of them as language policy rather than author burden: scalar f64 with no
SIMD and no reassociation in the deterministic core, a fixed reduction order, one pinned libm
build for every transcendental (never the host engine’s Math, which legitimately differs by an
ULP between browsers), one shared routine for decimal, Unicode, calendar and the seeded PRNG, and a
single canonical bit pattern for na so two engines never disagree even on the bytes of “nothing”.
The second line matters too: rand is a presentation generator — per-frame and non-replayable —
so it exists only on the canvas side of the firewall, where determinism is deliberately not
promised; analysis code that reads it is rejected with [ErrFirewall], and the numbers stay
replayable precisely because the one non-replayable thing cannot reach them.
None of this is assumed. The equivalence is verified at every compilation — the interpreter and the compiled module are run on real data and asserted bit-for-bit equal (this is invariant I7), and a divergence blocks the artifact. What you get in exchange is worth the discipline: replay-exact debugging where you can step backwards as reliably as forwards, golden tests with no tolerance to tune and no flakiness to excuse, and cross-machine agreement — which is exactly what lets a verifying server re-run a client’s work and get the same answer.
Note the precise claim, because it is easy to overstate. Determinism means same numbers, same scene — not same pixels. The compute is bit-identical and the draw-list the compiler emits is deterministic, but the picture on the glass is not promised identical across GPUs, and does not try to be. That is not a leak; it is the design of the next section but one.
History that cannot be rewritten
Byte-determinism gives you agreement across space. Causality gives you agreement across time — and it is the guarantee traders feel most directly. In any system that re-evaluates over growing history, the deadliest defect is the value that quietly changes retroactively: an analytic that looked prophetic on the chart because, at each past step, it had silently read data that did not exist yet. Both the live run and the backtest are “correct” — for different definitions of time — so no test catches the divergence. Flux removes the defect by making it inexpressible.
prev = close[1] // yesterday's close — legal, and na on the first bar
gain = math.max(close - prev, 0)
peek = close[-1] // ✗ [ErrCausal] — a negative delay reads the futureThree rules produce the property. Delays reach backwards only — x[n] with a constant n ≥ 0;
a negative index does not parse into a meaning, it raises [ErrCausal]. A resample reads the
last closed unit of a coarser clock, never the one still forming. And every feedback cycle must
cross a unit delay, so today’s output may depend on yesterday’s output but never on itself. By
induction over the graph, output[t] = f(inputs[0..t]) — a past value has nothing left to depend
on, so nothing can move it. This is no-repaint: a value, once produced for a step, never changes,
and it holds for every accepted program, not a careful subset. Live and historical evaluation are
the same function computed over the same inputs, so they produce the same bytes — the honest
backbone of any evaluation.
There is exactly one exception, and it comes with a wall. live(e) re-evaluates an analysis
sub-graph including the step in formation — genuinely useful for watching an indicator move
within the current step — and it may flow only to display sinks. Feed it into an alert, an
assertion or a calculation and you get [ErrFirewall]. A script that uses it is flagged
non-replayable, visibly, so you always see the trade-off you made. The firewall that enforces
this is the subject of Guide §8; here it is enough to know that the one
value that can still change is kept out of every number a decision rests on.
Bounded, so it is safe to run
Determinism assumes the program finishes. Flux guarantees it: every program terminates, and its
cost per step is known at compile time. There are no unbounded loops, no unbounded recursion, no
unbounded collections. Iteration exists — window with map/fold for counted loops, scan for
running state, loop(max, …) for “iterate until done” with a declared ceiling — but every bound is
a compile-time constant under a global cap. A program that cannot state its bound, or that exceeds
the budget, is rejected before it runs with [ErrTotal], the offending bound named. It is
never killed mid-run by a timeout, because there is nothing to time out.
One subtlety is worth pausing on, because it is where determinism has to start. The accept/reject verdict is itself a pure function of the source, decided by counters alone — never by a clock. The editor’s build timeout, on the order of two seconds, is an interactive cancellation to keep the UI responsive; it is never a verdict. A wall-clock verdict would be machine-dependent — the same script accepted on a fast machine and rejected on a slow one — and then two users would not be running the same language, and replay, which assumes that what compiled there compiles here, would break. Determinism has to hold at the compiler’s answer, or it holds nowhere downstream.
Flux is, deliberately, not Turing-complete — and it is worth being honest about what that costs and what it does not. Totality is the synchronous-dataflow trade, in the lineage of Lustre and SCADE: the language is total and bounded per step, while remaining unbounded over time — it runs step after step forever, it just cannot spin without limit within one. In Flux’s client-side, reactive domain that ceiling is invisible, because an unbounded inner loop there would not be a clever program — it would be a hung tab, i.e. a bug. What totality buys is the foundation everything else stands on: a static budget the host can honor, and static checks (kinds, causality, exhaustiveness) that are decidable because the language is total and its bounds are constants.
Replay comes for free
The properties so far are about analysis — pure series over closed steps. Applications add the one thing analysis cannot: state that persists between events and decides what is displayed — a score, a document, a selection, a blotter of positions. You might expect that to be where determinism frays. It is where it pays off most.
An application’s state never changes except through one deterministic, journaled reducer, so the
model is nothing more than fold(init, [the journal of every message so far]). Undo rewinds the journal one
step and folds again; redo re-extends it. Recall the counter from Guide §10 — its entire undo
history is those two moves:
variant Msg { Bump | Undo | Redo }
app tally {
capabilities: [ journal ]
init(p) = { doc: { n: 0 }, ui: { pending: na } }
update(m, msg) = match msg {
Bump -> { model: m with { doc: m.doc with { n: m.doc.n + 1 } }, cmds: [] }
Undo -> { model: m, cmds: [ Journal(UndoToMark) ] }
Redo -> { model: m, cmds: [ Journal(RedoToMark) ] }
}
view(m) = row {
button("undo", Undo)
text("{m.doc.n}")
button("+1", Bump)
button("redo", Redo)
}
subs(m) = []
}Read the two undo arms carefully: they return the model unchanged and hand the host an inert command. The host — which is where the journal lives — truncates it to the target and re-folds into the new model. Nothing in the application code knows how to reverse an edit; it only knows how to move forward. Because the fold is a pure function of the journal, the state you land on is exactly the state you left — reproduced, not reconstructed. And because Flux is deterministic to the byte, “exactly” means bit-for-bit, including the answers that came back from the network: an async result entered the model as a message, so it is already in the journal, replayed verbatim rather than re-fetched. Rewinding never fires the request a second time; it reads the reply that already happened.
One mechanism wears two faces. For the user it is undo/redo that cannot silently forget a field
or resurrect a stale selection — because the Model is split into a doc (the business state history
owns) and a ui (the cursor, the draft, the in-flight request) that history does not, so undo
rewinds doc and leaves ui alone. For the developer the same journal is a time-travel debugger:
scrub to any past message, step backwards, set a data breakpoint — stop at the first message where
score crosses 100 — and put the run beside a reference run to see where, if ever, they diverge.
It is undo over an application’s execution, the events it processed, never over its source code.
The full architecture — the doc/ui partition, checkpoints, schema evolution — is
Guide §10.
What replay proves — and what it does not
Guide §10 first drew this limit; here it faces an adversary head-on, and the guide owes you a
precise statement rather than a slogan. A verdict is a pure function of (init, [msg]), so a server can recompute it on
the same bytes, and a claimed score that was not earned is a journal that does not fold to the
claimed result. That property was not designed; it was inherited from determinism. But it proves
one specific thing:
Replay proves the COHERENCE of a journal, not its NON-FALSIFIABILITY.
“Diverges ⇒ tampered” is complete for a score derived from the seed and the elapsed time,
because the server owns both — the seed is derived server-side from (runId, level, qIndex) and
never accepted from the client, and the elapsed time of a ranked run is host-stamped and substituted
at re-fold, so a forged Tick count buys nothing. It is not complete for a third class, and the
gap is open by name. A score fed by a host-pushed outcome — an OnReveal message carrying a
result the host computed — enters the journal as data, and a re-fold replays data verbatim. The
seed re-derives the messages that came from randomness; this one did not. So a forged outcome
re-folds to the claimed result without a whisper of divergence: the check passes, and the claim is
still a lie. The same class covers a pixel — a bounded pixel reading derives from the client’s
own pan and zoom, and a server with no viewport cannot re-derive it. Hence the standing,
language-level rule: a pixel value never feeds a ranked verdict. It is a readout, or it is
cosmetic.
An outcome-fed run therefore has exactly two honest destinations and no third: the host re-derives the outcome server-side by re-running the kernel itself, or the run is local-score-only, excluded from the shared leaderboard. It is never accepted on the strength of the client’s journal alone.
Server re-execution is designed, and rests on exactly the determinism above — a server re-running a client’s work to catch a forged result. In v1 the native/server leg is verified client-side: re-execution lands with the server port of the grader.
Running code you did not write
Now the question underneath the others. Flux is the language of a platform where code is shared, sold and run by people who did not write it — an ordinary Flux target is untrusted code, and a program a language model emits is a clean instance of exactly that. That is only tenable if safety is a property of the language and host, not of a review process, and Flux gets there in layers.
The language has no primitive that does I/O — no fetch, no DOM access, no eval, no file handle
— so a script’s only channel to the world is data it hands the host. Scripts hold no ambient
authority; there is no global object in scope at all. An effect is an inert Cmd that carries
values — a sound’s name, a storage key, a score — never a token, socket, URL or handle. The host
holds every resource and runs the command only if the corresponding capability was declared in the
app’s manifest and granted by the user. Emitting a command the manifest does not grant is rejected
at compile time with [ErrCapDenied] — it never becomes a runtime event to be caught. And the
manifest is not a promise the author writes: the compiler derives it from the emit Cap sites, as
the union of capability needs over the whole transitive dependency closure, intersected with the
user’s grant — so a dependency’s appetite for the network surfaces before install, no dependency
can exceed what you granted to the whole, and no dependency holds a capability object it could
re-delegate or amplify.
The consequence is the one that makes user-generated content routine rather than risky: security is the grant, not the code. The first-party interface and a stranger’s marketplace application run the same language in the same sandbox; they differ only in what has been granted. This is the third pillar as a runtime fact: a stranger’s island renders beside yours, and touches none of the data your own logic depends on. A trusted tier grants effects — it never loosens causality, no-repaint, totality or the firewall. Repaint is inexpressible for everyone. And trust is decided by the host from a server-side provenance record keyed by the binary’s content hash: a module cannot declare itself trusted, and no embedded metadata is believed.
Name the honest limit here too. The language is safe by construction; the capability monitor and the view sanitizer are ordinary code, and they are the residual attack surface — which is precisely why they are the components that earn a model check, and why the view primitives are a closed, typed set rather than a string a hostile view could smuggle markup through.
The optimizer you don’t have to trust
One more actor could break determinism: the optimizer. A Flux program is a pure, typed, total, causal DAG, which makes classic optimizations — sharing, pruning, reordering, specializing — safe by construction, and its small bounded size makes whole-graph searches affordable. The distinctive part is the trust model. The optimizer obeys one law: the optimized program must be bit-identical to the reference semantics — the unoptimized graph’s canonical evaluation. That law is enforced by translation validation: every compilation runs the reference and optimized artifacts on hostile data and asserts byte-equality. A miscompilation cannot ship; it can only fail loudly at the gate, where the compile falls back to the unoptimized path and turns the test suite red. The optimizer can therefore be aggressive precisely because no one has to believe in it.
What you get is free global common-subexpression elimination — write ema(close, 26) four times in
a script and it is computed once — dead code elimination, constant folding and element-wise fusion,
with the editor’s cost gutter showing the optimized graph, so what you read is what you pay. The
default is bit-exact, always.
Sharing a subgraph across four co-active scripts — one merged graph for the whole chart — is a later optimizer tier, and relaxed floating-point (reassociation, fused multiply-add) is designed as an explicit, labelled opt-in, never a silent default and never bit-exact.
What the compile actually earned you
A guarantee whose only enforcement is a promise in a document is not a guarantee, so every property in this chapter is checked by a machine, and the editor tells you which ones your program earned. After a compile, the guarantees panel states it plainly:
✓ No-repaint ✓ No look-ahead ✓ Deterministic
✓ Bounded memory ✓ Byte-identical ⚠ contains live() → non-replayableThat last line is not decoration: a guarantee you traded away should be visible at the moment you
traded it. Behind the panel is a verification harness whose sub-suites each declare an oracle and a
corpus — goldens, a differential oracle across interpreter, WASM and native kernel, a metamorphic
suite of semantics-preserving relations, a 1 ≡ N worker stress, a capability monitor that asserts
“no command outside the manifest is ever executed”, and more — most of them blocking. Reproducible
builds sit under all of it: the build hash is a pure function of the source, the dependency closure,
the compiler version, the pinned routines and the memory plan, and a rebuild gate recompiles the
same inputs on different machines with different thread counts and asserts a byte-identical module.
Server-side replay depends on this, because a value-level oracle cannot see the bytes emitted across
two compilations.
The limits, stated plainly
Because a trust page that lists only strengths is a sales page, here is what Flux does not guarantee:
- Presentation is not deterministic, and does not try to be. The scene graphics and text both render on the GPU — text as SDF glyphs, crisp at any zoom, sidestepping DOM layout cost — and the GPU, the compositor and unseeded randomness are outside the oracle by design. Same numbers, same scene; not same pixels. The firewall contains them, it does not eliminate them.
- Replay proves coherence, not truthfulness. A host-pushed outcome journaled as data is re-folded verbatim, so a score depending on it needs the server to re-derive it, or must be excluded from a shared leaderboard. This is a named open problem, not a hidden one.
- Server re-execution is designed, not shipped in v1 — the determinism that would let a server catch a forged result is real and verified, but in v1 it is verified on the client.
- Bounds are claims, not invariants.
osc(0,100)says what a value conventionally is, not what it provably is; onlyclampmakes a bound real. Flux has no solver, and says so. - The sandbox rests on two pieces of ordinary code — the capability monitor and the view sanitizer. Everything else is safe by construction; those two are safe by review, model checking and fuzzing.
See also
- The four planes — the firewall that keeps the wall clock and the forming unit out of analysis.
- Building an application — the journal, the
doc/uipartition, and the reducer replay rests on. - Streams, delay and running state — delays,
scan, and causality in practice. - Guarantees — each promise, its exact scope, and the machine check that enforces it.
- The App plane — capabilities, journals and replay, in full.
- The optimizer — the correctness law and the rewrites that look sound and are not.
The formal rules → This chapter narrated determinism, replay and the trust model. Their precise statements live in the Spec, FVM and FDK:
- Byte-determinism, the pinned routines and the I7 gate — Guarantees › Byte-determinism and Verification & reproducible builds.
- Causality, delays, closed-unit resampling and
live()— Time & state.- Totality — the bound,
[ErrTotal], and the deterministic verdict — Guarantees › Totality.- Replay, the journal, undo bounds and checkpoints — The App plane › Totality, determinism, replay.
- What replay proves — and the open anti-cheat vector — The App plane › What replay proves and server, which carries the closing argument.
- Capability security — the catalogue, the two trust tiers, transitive manifests — The App plane › Capabilities.
- Verified optimization — translation validation and the correctness law — The optimizer.
- The verification harness and reproducible builds — Verification & reproducible builds.