compute — numbers, dataframes and domain libraries
The compute pillar is what makes Flux an analysis language for arbitrary tabular data rather than a chart-scripting one. It is a columnar, bounded, kind-typed dataframe algebra — pure, total, fused, dimensional — with a numeric layer (matrices, linear algebra, statistics) and domain libraries above it. A market series is not the foundation here; it is a special case of a table.
Everything in it is built out of the frozen core: a table is a record of columns, a column is a
bounded vector, a query is a pure graph. No new sort, no new grammar, no new substrate. The
dimensional system every column carries lives in Kinds; the bounded structures
the tables rest on are collections; the step that turns a Table into a scene
is display. This page specifies the table algebra, the numeric layer above it, and
the domain libraries above that.
New here? Start with Guide §5 — Kinds: types that carry meaning →
The dataframe layer, the numeric layer and the domain libraries are sealed in design; the scalar
namespaces (math, stat, vec, decimal, time, ta) are the v1 surface the analysis plane
already uses. This page describes the whole pillar in the present tense; its maturity sits in the
status ledger.
Six properties, each a theorem rather than a wish
- Bounded and total. Every collection is a
vec(κ, N)with a const-folded capacity. The output capacity of every relational operator is statically derivable from its inputs, and no operation ever shrinks. The memory a query needs is computable at compile time; an over-budget query is rejected, never killed mid-run. - Columnar. A table is a record whose fields are columns — struct-of-arrays, the same layout the engine already uses for bars.
- Kind-typed and dimensional. Every column carries its dimension, its numeric representation
and its asset tag. An entire class of bugs — adding a price to a volume, a BTC price to an ETH
price, an
f64to adecimal, joining on keys from different instruments — becomes a compile error. - Pure. A table touches neither the network nor application state, so it lives on any plane without a firewall question.
- Lazy and fused. A query is an inert graph, fused into one columnar pass. Projection pushdown is dead-code elimination; predicate pushdown is fusing the mask into the producing scan; nothing intermediate is ever materialized.
- Byte-identical. Scalar
f64, no SIMD, no floating-point reassociation, every reduction order pinned, every order statistic through the pinnedna-ordering routine, every decimal through the pinned integer routine.
The scalar namespaces
These are the everyday surface — available in every plane, and dimensional throughout.
| Namespace | Contents |
|---|---|
math.* |
abs sign min max clamp floor ceil round (dimension-preserving) · sqrt (halves the exponents, so stdev = sqrt(variance) types) · pow(x, n:lit) · log exp sin cos tan atan atan2 (demand dimensionless — log(price) is [ErrDim], with a quick-fix) · lerp norm |
stat.* |
mean stdev variance skew kurtosis median percentile rank correl covar zscore linreg · the exponentially-weighted family ewmVar ewmStd ewmCov ewmCorr — all bounded window reducers |
vec.* |
map fold scan zip sum avg product min max reverse take drop window · fill(N, x) · range(N) · setAt(v, i, x) · where / mask (length-preserving) · sortBy / topK · count any all |
decimal.* |
div round — division names its target scale (half-even), round quantizes to a scale; with bare toDecimal / toFloat as the f64 bridge — exact fixed-point money |
time.* |
the calendar: years months weeks days (the only producers of a period), the accessors, epoch* conversions, in_session, barsPerYear |
ta.* |
the indicator catalogue |
enc.* crypto.* id.* bits.* |
encoding, pinned hashes and signature verification, deterministic ids, bit and byte-buffer work |
Three rules run through all of them and are worth internalizing:
Maths is unit-correct. sqrt halves exponents, pow scales them, the transcendentals
demand dimensionless input. This is not pedantry — it is what makes sqrt(variance) → level
type-check while log(price) does not.
Order statistics are pinned. median, percentile and rank sort internally, and they use
the one pinned total order over na (absent values last, stable by index) — never the
platform’s sort. percentile names one interpolation method; median on an even count is the
midpoint; rank names one tie policy. A window containing a hole yields na, and there is a
golden for exactly that.
Higher moments and EW statistics are stable by construction. Skew and kurtosis extend the
pinned Welford recurrence; the exponentially-weighted family is a single bounded scan with a
pinned λ-recurrence — never the numerically unstable one-pass form.
Tables
A table is a record of columns, plus a presence mask and a live count:
Table<{ c₁:κ₁, …, cₘ:κₘ }>[N] ≡ record{
c₁: vec(κ₁, N), …, cₘ: vec(κₘ, N), // the columns — struct-of-arrays, one shared row cap
live: vec(signal, N), // 1 = a real row, 0 = a tombstone
count: num // live rows ≤ N
}
Figure — the slotmap of the APP plane, transposed from array-of-structs to struct-of-arrays.
Because it is a record of vectors, it inherits the lattice laws, deep na-aware equality, and
projection (t.close) for free — and it compiles straight onto the existing columnar engine.
Col(κ, N) is a named alias of a bounded vector; Mat(κ, R, C) adds a second const axis with a
rectangularity guarantee.
A Series — the time-indexed data of the analysis plane — is the special case, and the
bridge is an identity rather than a conversion.
Totality, verb by verb
The hardest problem this pillar solves is stated plainly: mainstream dataframes get their power from operations whose output cardinality depends on the data — a filter shrinks, a group-by returns one row per distinct key, a join can explode. Flux forbids that. The five rules that replace it:
| Rule | |
|---|---|
| T1 | The output cap is a static function of the input caps. N for select/filter/sort/window · min(N,k) for a slice · a declared G for a group-by · Ng for an as-of join · Ng × F for an equi-join with a declared fan-out · Na + Nb for a concat. Never data-dependent — and a derived cap that would exceed N_max is [ErrTotal] at compile time. |
| T2 | Selection is a MASK, never a shrink. A “filter” writes na where the predicate is false and clears the live bit. The length does not move; count does. |
| T3 | Iteration is na-aware, so nothing needs compacting. Aggregations reduce the living rows: tombstones are excluded before the reduction, not absorbed after it. |
| T4 | Overflow is a named policy, never growth. Past the declared cap: Reject (drop the surplus item, with a diagnostic), Truncate (keep the canonical survivors), or Bucket (an “other” bucket). A missing cap is a compile error. |
| T5 | Everything is functional, and lowers to the frozen vec.* primitives — which is why the fusion into one columnar pass comes for free. |
The verbs themselves are ordinary, and they chain:
def liquidBars(n) =
let bars = series("BTC-USD").toTable(n).derive(range, high - low) in // ADD a column — a fresh record
let hot = bars.where(bars.volume > sma(bars.volume, 20) and bars.range > atr(14)) in
hot.select(time, close, volume, range) // projection
def dailyVwap(n) =
let bars = series("BTC-USD").toTable(n) in
bars.groupBy((r) -> { y: year(r.time), d: dayOfYear(r.time) }, maxGroups: 366)
.agg((g) -> { vwap: (g.close * g.volume).sum() / g.volume.sum(), // pv ÷ volume → price
hi: g.high.highest() })Three details in that snippet carry the whole design.
where takes a column mask, not a row lambda — it never removes a row, it blanks one. The mask
is an ordinary column expression over the receiver, which is why the receiver has a name: there
is no implicit row variable in a Flux query, so an intermediate table is bound with let … in and
its columns are projected off it (bars.volume). A chain that never names its intermediate cannot
speak about its columns.
groupBy declares its cap (maxGroups: 366): the memory a group-by needs is therefore known
before it runs, which is exactly what an unbounded hash-aggregate cannot promise.
derive adds a column (it builds a fresh record); with { … } redefines an existing one.
The distinction is not stylistic — with is shape-preserving by law, so adding a field through it
would be [ErrField].
Joins
As-of is the primitive, not an afterthought — because time-series data is what most joins in this domain are about, and because it is the join whose output cap is trivially static (one row per left row):
def joins(quotes, refs) =
let bars = series("BTC-USD").toTable(256) in
{ aligned: bars.asofJoin(quotes, key: time, take: [bid, ask]), // the most recent quote at or before each row
paired: bars.join(refs, key: symbol, kind: Inner, maxFanout: 4) } // equi-join with a DECLARED maximum fan-outAn equi-join without a declared maxFanout does not compile. That is the price of never having a
query explode in production, and it is a price worth paying.
The execution model
A query is a DAG of pure nodes, and it is the same DAG the engine already optimizes. So:
- Projection pushdown is dead-code elimination — a column nobody reads is never materialized.
- Predicate pushdown is fusing the mask into the producing scan.
- Common-subexpression elimination is global, across a whole query, and across scripts.
- Nothing intermediate is materialized. A chain of ten verbs is one pass over the columns.
Execution follows the morsel → sink model: the columns are cut into chunks, chunks flow through the fused pipeline, and the results land in the sink. Parallelism applies to independent reductions (distinct groups, distinct cells) — never to the interior of a single reduction, because that would reassociate floating-point and break byte-identity.
Why the reduction order is pinned even when it costs speed. A parallel sum is faster and gives a different last bit. Two engines would then disagree, replay would drift, and a server could not verify a client’s work — the re-derivation argument server sets out, what replay proves, rests on this byte-identity holding through every reduction. The reduction order is therefore fixed, and the parallelism is found where it does not change a single bit.
The numeric layer
Mat(κ, R, C) with two const axes, and above it:
- Linear algebra — decompositions, solves, eigenvalues for symmetric matrices — each with a frozen reduction order, so the last bit is the same on every machine.
- Statistics — the moment family, order statistics, correlation, the OLS/QR regression path.
- Signal processing — the FFT, bounded and deterministic.
- Randomness — the pinned counter-based integer generator, so a simulation replays exactly.
- Optimization — bounded solvers, with an iteration cap that is part of the type.
Domain libraries
| Function | What it computes |
|---|---|
sharpe(returns, rf, barsPerYear(clock)) |
the annualized ratio, with the base made explicit |
rollOls(y, x, 60) |
a rolling regression — bounded window |
drawdown(equity) |
the running peak-to-trough |
impliedVol(price, strike, t, r) |
an option surface point |
ytm(cashflows), duration(bond) |
fixed income — a yield, and a time-sensitivity |
laspeyres(p0, p1, q0), paasche(p0, p1, q1) |
index numbers |
Every one of them is a bounded composition of the primitives above — which means each of them is
also readable, and each of them type-checks dimensionally. sharpe cannot silently annualize
with the wrong base, because the base is an argument whose kind says what it is.
geom.* is the 2-D geometry family — distances, point-in-shape, intersections — used by drawing
tools and custom layout. It is screen-space by design: data-space projection belongs to the
host, which is what keeps the firewall intact.
The spreadsheet
A bounded, reactive dataframe with a cell grammar (SUM, AVERAGE, VLOOKUP, SUMIF, COUNTIF,
…), evaluated by the same graph engine with early cut-off — a sheet is a query that recomputes only
what changed. It is one of the domain libraries above: included here because it falls out of the
algebra rather than being bolted onto it, a sheet being a bounded table plus a dependency graph, and
Flux already has both.
Reserved seams
metric[id] — a non-price series (an economic indicator, an on-chain metric, an analytics stream)
entering the analysis plane as a first-class causal stream, tagged with its identity so that adding
a CPI series to a hashrate series is [ErrDim]. The seam is designed and held inert; nothing
non-price enters analysis until it is armed.
See also
- Kinds — the dimensional system every column carries.
- collections — the bounded structures under the tables.
- display —
viz.*, which turns aTableinto a scene. - units —
meas[u], for quantities outside the market domain. - Memory model — the columnar layout and the liveness plan.
- Optimizer — the fusion and the pushdowns, and why they are safe.