◆ Flux

compute — numbers, dataframes and domain libraries

The compute pillar is what makes Flux an analysis language for arbitrary tabular data rather than a chart-scripting one. It is a columnar, bounded, kind-typed dataframe algebra — pure, total, fused, dimensional — with a numeric layer (matrices, linear algebra, statistics) and domain libraries above it. A market series is not the foundation here; it is a special case of a table.

Everything in it is built out of the frozen core: a table is a record of columns, a column is a bounded vector, a query is a pure graph. No new sort, no new grammar, no new substrate. The dimensional system every column carries lives in Kinds; the bounded structures the tables rest on are collections; the step that turns a Table into a scene is display. This page specifies the table algebra, the numeric layer above it, and the domain libraries above that.

New here? Start with Guide §5 — Kinds: types that carry meaning →

The dataframe layer, the numeric layer and the domain libraries are sealed in design; the scalar namespaces (math, stat, vec, decimal, time, ta) are the v1 surface the analysis plane already uses. This page describes the whole pillar in the present tense; its maturity sits in the status ledger.

Six properties, each a theorem rather than a wish

The scalar namespaces

These are the everyday surface — available in every plane, and dimensional throughout.

Namespace Contents
math.* abs sign min max clamp floor ceil round (dimension-preserving) · sqrt (halves the exponents, so stdev = sqrt(variance) types) · pow(x, n:lit) · log exp sin cos tan atan atan2 (demand dimensionless — log(price) is [ErrDim], with a quick-fix) · lerp norm
stat.* mean stdev variance skew kurtosis median percentile rank correl covar zscore linreg · the exponentially-weighted family ewmVar ewmStd ewmCov ewmCorr — all bounded window reducers
vec.* map fold scan zip sum avg product min max reverse take drop window · fill(N, x) · range(N) · setAt(v, i, x) · where / mask (length-preserving) · sortBy / topK · count any all
decimal.* div round — division names its target scale (half-even), round quantizes to a scale; with bare toDecimal / toFloat as the f64 bridge — exact fixed-point money
time.* the calendar: years months weeks days (the only producers of a period), the accessors, epoch* conversions, in_session, barsPerYear
ta.* the indicator catalogue
enc.* crypto.* id.* bits.* encoding, pinned hashes and signature verification, deterministic ids, bit and byte-buffer work

Three rules run through all of them and are worth internalizing:

Maths is unit-correct. sqrt halves exponents, pow scales them, the transcendentals demand dimensionless input. This is not pedantry — it is what makes sqrt(variance) → level type-check while log(price) does not.

Order statistics are pinned. median, percentile and rank sort internally, and they use the one pinned total order over na (absent values last, stable by index) — never the platform’s sort. percentile names one interpolation method; median on an even count is the midpoint; rank names one tie policy. A window containing a hole yields na, and there is a golden for exactly that.

Higher moments and EW statistics are stable by construction. Skew and kurtosis extend the pinned Welford recurrence; the exponentially-weighted family is a single bounded scan with a pinned λ-recurrence — never the numerically unstable one-pass form.

Tables

A table is a record of columns, plus a presence mask and a live count:

Table<{ c₁:κ₁, …, cₘ:κₘ }>[N]  ≡  record{
    c₁: vec(κ₁, N), …, cₘ: vec(κₘ, N),   // the columns — struct-of-arrays, one shared row cap
    live:  vec(signal, N),                // 1 = a real row, 0 = a tombstone
    count: num                            // live rows ≤ N
}

A table is a record of columns Figure — the slotmap of the APP plane, transposed from array-of-structs to struct-of-arrays.

Because it is a record of vectors, it inherits the lattice laws, deep na-aware equality, and projection (t.close) for free — and it compiles straight onto the existing columnar engine. Col(κ, N) is a named alias of a bounded vector; Mat(κ, R, C) adds a second const axis with a rectangularity guarantee.

A Series — the time-indexed data of the analysis plane — is the special case, and the bridge is an identity rather than a conversion.

Totality, verb by verb

The hardest problem this pillar solves is stated plainly: mainstream dataframes get their power from operations whose output cardinality depends on the data — a filter shrinks, a group-by returns one row per distinct key, a join can explode. Flux forbids that. The five rules that replace it:

Rule
T1 The output cap is a static function of the input caps. N for select/filter/sort/window · min(N,k) for a slice · a declared G for a group-by · Ng for an as-of join · Ng × F for an equi-join with a declared fan-out · Na + Nb for a concat. Never data-dependent — and a derived cap that would exceed N_max is [ErrTotal] at compile time.
T2 Selection is a MASK, never a shrink. A “filter” writes na where the predicate is false and clears the live bit. The length does not move; count does.
T3 Iteration is na-aware, so nothing needs compacting. Aggregations reduce the living rows: tombstones are excluded before the reduction, not absorbed after it.
T4 Overflow is a named policy, never growth. Past the declared cap: Reject (drop the surplus item, with a diagnostic), Truncate (keep the canonical survivors), or Bucket (an “other” bucket). A missing cap is a compile error.
T5 Everything is functional, and lowers to the frozen vec.* primitives — which is why the fusion into one columnar pass comes for free.

The verbs themselves are ordinary, and they chain:

FLUX
def liquidBars(n) =
  let bars = series("BTC-USD").toTable(n).derive(range, high - low) in   // ADD a column — a fresh record
  let hot  = bars.where(bars.volume > sma(bars.volume, 20) and bars.range > atr(14)) in
  hot.select(time, close, volume, range)                                 // projection

def dailyVwap(n) =
  let bars = series("BTC-USD").toTable(n) in
  bars.groupBy((r) -> { y: year(r.time), d: dayOfYear(r.time) }, maxGroups: 366)
      .agg((g) -> { vwap: (g.close * g.volume).sum() / g.volume.sum(),   // pv ÷ volume → price
                    hi:   g.high.highest() })

Three details in that snippet carry the whole design.

where takes a column mask, not a row lambda — it never removes a row, it blanks one. The mask is an ordinary column expression over the receiver, which is why the receiver has a name: there is no implicit row variable in a Flux query, so an intermediate table is bound with let … in and its columns are projected off it (bars.volume). A chain that never names its intermediate cannot speak about its columns.

groupBy declares its cap (maxGroups: 366): the memory a group-by needs is therefore known before it runs, which is exactly what an unbounded hash-aggregate cannot promise.

derive adds a column (it builds a fresh record); with { … } redefines an existing one. The distinction is not stylistic — with is shape-preserving by law, so adding a field through it would be [ErrField].

Joins

As-of is the primitive, not an afterthought — because time-series data is what most joins in this domain are about, and because it is the join whose output cap is trivially static (one row per left row):

FLUX
def joins(quotes, refs) =
  let bars = series("BTC-USD").toTable(256) in
  { aligned: bars.asofJoin(quotes, key: time, take: [bid, ask]),      // the most recent quote at or before each row
    paired:  bars.join(refs, key: symbol, kind: Inner, maxFanout: 4) } // equi-join with a DECLARED maximum fan-out

An equi-join without a declared maxFanout does not compile. That is the price of never having a query explode in production, and it is a price worth paying.

The execution model

A query is a DAG of pure nodes, and it is the same DAG the engine already optimizes. So:

Execution follows the morsel → sink model: the columns are cut into chunks, chunks flow through the fused pipeline, and the results land in the sink. Parallelism applies to independent reductions (distinct groups, distinct cells) — never to the interior of a single reduction, because that would reassociate floating-point and break byte-identity.

Why the reduction order is pinned even when it costs speed. A parallel sum is faster and gives a different last bit. Two engines would then disagree, replay would drift, and a server could not verify a client’s work — the re-derivation argument server sets out, what replay proves, rests on this byte-identity holding through every reduction. The reduction order is therefore fixed, and the parallelism is found where it does not change a single bit.

The numeric layer

Mat(κ, R, C) with two const axes, and above it:

Domain libraries

Function What it computes
sharpe(returns, rf, barsPerYear(clock)) the annualized ratio, with the base made explicit
rollOls(y, x, 60) a rolling regression — bounded window
drawdown(equity) the running peak-to-trough
impliedVol(price, strike, t, r) an option surface point
ytm(cashflows), duration(bond) fixed income — a yield, and a time-sensitivity
laspeyres(p0, p1, q0), paasche(p0, p1, q1) index numbers

Every one of them is a bounded composition of the primitives above — which means each of them is also readable, and each of them type-checks dimensionally. sharpe cannot silently annualize with the wrong base, because the base is an argument whose kind says what it is.

geom.* is the 2-D geometry family — distances, point-in-shape, intersections — used by drawing tools and custom layout. It is screen-space by design: data-space projection belongs to the host, which is what keeps the firewall intact.

The spreadsheet

A bounded, reactive dataframe with a cell grammar (SUM, AVERAGE, VLOOKUP, SUMIF, COUNTIF, …), evaluated by the same graph engine with early cut-off — a sheet is a query that recomputes only what changed. It is one of the domain libraries above: included here because it falls out of the algebra rather than being bolted onto it, a sheet being a bounded table plus a dependency graph, and Flux already has both.

Reserved seams

metric[id] — a non-price series (an economic indicator, an on-chain metric, an analytics stream) entering the analysis plane as a first-class causal stream, tagged with its identity so that adding a CPI series to a hashrate series is [ErrDim]. The seam is designed and held inert; nothing non-price enters analysis until it is armed.

See also