◆ Flux

Lexical structure

This page defines how a Flux source text becomes a stream of tokens: the source model, comments, the significant newline, the complete token catalogue (identifiers, the numeric family and its glued suffixes, strings and interpolation, capability references, operators and punctuation), and the keyword model — which words are reserved, where, and why most of them are only reserved contextually. Everything downstream — the grammar, kind inference, the editor — consumes exactly the token stream specified here.

Two properties frame the whole page. First, the lexer is total: any input text produces a token stream (malformed input surfaces as precise diagnostics, never as a crash). Second, the lexer is linear and incremental: every context-dependent device below (the significant newline, interpolation, capability references, the numeric munch) is bounded by counters the lexer already maintains, so retokenizing after an edit touches only the edited neighborhood — the property that keeps live preview under its frame budget.

New here? Start with Guide §3 — Your first session → for the same tokens met in program order, then return here for the exact rules.

Source model

A Flux program is a Unicode text. The fixed alphabet of the language itself — identifiers, keywords, operators, punctuation — is ASCII; arbitrary text (any Unicode) lives inside string literals and comments. Between tokens, spaces and tabs are insignificant and may appear in any number; newlines are significant (see TERM below).

One rule shapes many diagnostics on this page: Flux has no juxtaposition. Two adjacent primary expressions with no operator between them never form a term:

FLUX
x = 1.50 d      // ✗ syntax error — NUMBER then IDENT: two adjacent primaries
y = 1.50d       // decimal(2) — the glued suffix makes this ONE token
d = 2           // `d` alone is an ordinary identifier — a licit binding
z = 1.50 * d    // valid — the detached `d` is just a name here

Because adjacency is never meaningful, the lexer can afford glued suffix tokens (1.50d, 4px) without ambiguity: either the suffix touches the number and the pair is one token, or it does not and the program is ill-formed unless an operator intervenes.

A script may span several .flux source files; the package visibility modifier (see grammar — modules) is scoped to exactly that set of files.

Comments

Form Token Extent Role
// … LINE to end of line ordinary comment, skipped
/// … DOC to end of line doc-comment, attaches to the next def
FLUX
/// z-score of a series over n bars
def zscore(x, n=20) = (x - sma(x, n)) / stdev(x, n)

plot zscore(close, 20) as z   // an ordinary comment

A doc-comment is lexically a comment (the parser skips it) but it is not thrown away: the documentation pipeline collects the /// block preceding a def and publishes it as the definition’s hover documentation and doc-as-data entry (see Guide §12 — Working in the editor). There are no block comments; a comment always ends at the newline. Comments are transparent to the significant-newline rules below — a line that ends in a trailing // … comment is classified by the last token before the comment, and a comment-only line neither ends nor continues a statement.

The significant newline (TERM)

Flux has no mandatory statement terminator. A statement ends at the end of its line — the lexer emits a TERM token at the newline — unless the line is visibly unfinished. Semicolons are optional separators (never required); writing several statements on one line with ; is legal but idiomatic Flux is one statement per line.

The three-clause newline decision: open paren/bracket depth, a pending line, a continuator next — otherwise TERM Figure — the three-clause newline decision: open ( / [ depth, a pending line end, or a continuator next each continue the statement; when none holds, the lexer emits TERM and the statement ends.

A newline does not emit TERM (the statement continues) when any of the three clauses holds:

  1. (a) Open parenthesis or bracket. The ( / [ nesting depth is greater than zero. Inside parentheses and square brackets, newlines are pure whitespace — which is how a long call wraps across lines. Braces do not reset that depth: a brace body written inside a call is still at depth > 0, so its items need an explicit , or ; separator. At depth zero, a newline separates items, which is why a top-level multi-line record or match needs no commas at all:

    FLUX
    def f(r) = r.a + r.b
    m = { a: 1                      // depth 0 — the newline separates
          b: 2 }
    n = f({ a: 1 ; b: 2 })          // inside a call — the `;` is REQUIRED
  2. (b) The line is pending. The last significant token of the line is an operator, an opening delimiter, a separator, or a keyword that grammatically awaits an operand — if, then, else, let, in, not, and, or, as, from, with, over, def, match, variant, record, app, import, type, representation, tool, on, every, when, tween, spawn, burst, emit, rate, set, morph, replay, focus, cross_up, cross_down, pub, private, package, scene. A contextual keyword only counts here in its keyword role: in x = p.on the trailing on is a field name and the statement ends. (One lexical subtlety: a trailing % glued to a digit is a percent literal, not a pending modulo — 50% ends the line.)

  3. © The next token is a continuator. The first token of the following line can only extend an expression, never begin a new item.

Clause © is not an ad-hoc list. Normative definition: a token is a continuator iff it belongs to no FIRST set of any list item — formally, iff it is outside ⋃ FIRST(item) taken over every repeated-list production of the grammar (statements, view children, variant constructors, representation and tool hooks, match arms, record fields, capability references, app members). The set is computed from the grammar, not maintained by hand, which makes the rule machine-checkable and keeps it in lockstep with the grammar forever.

Illustratively, the continuators are the purely infix or postfix tokens — . (member access), .., @, +, *, /, %, <, >, <=, >=, ==, !=, cross_up, cross_down, and, or, ->, with, ?, ??, ?., :, ,, ;, the closers ) ] }, the variant separator | — and the purely medial keywords then, else, in, as, from.

Three tokens are deliberately not continuators even though they can extend an expression, because they can also begin a new item: - (unary minus), [ (a list literal) and ( (a parenthesized form). A leading . is a continuator only when not followed by a digit — .upper continues, but .5 is a number and begins an item. To continue across a line on one of these ambiguous tokens, end the previous line pending (clause b) instead:

FLUX
u = bollinger(close, 20)
  .upper                      // `.` continues the postfix chain — clause (c)

total = close
  + open -                    // `+` continues (c); trailing `-` leaves the line pending (b)
  low

cond = if close > open
  then 1
  else 0                      // medial keywords are continuators — clause (c)

variant Wide { A | B
  | C }                       // `|` is a continuator inside the declaration

multi = sma(
  close,
  20)                         // clause (a): inside ( ) newlines are whitespace

m = { a: 1 }
updated = m
  with { a: 2 }               // `with` is a continuator — postfix record update

The clause-(b) idiom for the ambiguous heads: a -⏎ b (line ends on the operator), v[⏎ i] (line ends on the opener). Similarly, over is not a continuator — as a contextual keyword it may begin an item (for example as a property key) — so a multi-line morph … over d breaks after over, which is a pending word under clause (b).

Why this rule exists. A terminator-free surface reads like the notation authors actually sketch, but it must never become whitespace guesswork. The three clauses are decidable from at most one token of lookahead and the counters the lexer already carries, so newline classification is O(1) per line, linear over the file, and incremental under edits. And because the continuator set is defined as the complement of the computed item-head set, there is no hand-curated list to drift out of sync when the grammar grows: adding a statement head automatically removes it from the continuators.

Token catalogue

The complete token inventory. Each class is detailed in the sections that follow.

Class Tokens Notes
Identifier IDENT [A-Za-z_][A-Za-z0-9_]*, not reserved at its position
Numbers NUMBER integer, decimal, leading-dot, exponent forms
Suffixed literals DUR PCT PX DEC SPAN RATE glued suffixes; SPAN allows one space
Booleans / absence true false na typed literals (signal; na inhabits every kind)
Strings STRING "…" or '…', single line
Interpolation STR_HEAD STR_MID STR_TAIL fragment tokens of "… {e} …"
Capability ref CAPREF namespace:verb, only inside capability lists
Range / arrow .. -> one token each
Comparison < > <= >= == != cross_up cross_down non-associative level
Additive / multiplicative + - · * / % - also unary
The ? family ?? ?. ? maximal munch, in that order
Punctuation @ . , : ; = ( ) [ ] { } ~ | | separates variant constructors only
Comments LINE DOC //, ///
Layout TERM significant newline

Identifiers

IDENT = [A-Za-z_] [A-Za-z0-9_]*

Identifiers are ASCII: a letter or underscore, then letters, digits and underscores. An identifier is a name for a binding, parameter, field, kind, argument label, module or declaration. The lone underscore _ is lexically an ordinary identifier with two blessed roles given to it by the grammar and the elaborator: the wildcard pattern in match arms, and the implicit single-parameter placeholder in expression position (vec.map(_ * 1.1), see operators).

Whether a given identifier is available depends on the keyword model below — most keywords in Flux are contextual, so words like render, view or color remain usable as field names, parameters and kind names.

Numbers

NUMBER = ( digits "." digits | "." digits | digits ) [ ("e"|"E") ["+"|"-"] digits ]

42, 2.5, .5 and 1.5e3 are all NUMBER tokens. A bare NUMBER has kind lit — the const-folded literal that is dimension-polymorphic (close + 10 is a price; see kinds).

Maximal munch, bounded by the dot rule. The number scanner never consumes a . that is followed by another . or by a non-digit. This single rule makes ranges and member access compose with numbers without separators:

FLUX
len   = input(14, 2..200)      // NUMBER RANGE NUMBER — the dot is never eaten before `..`
half  = .5                     // a leading-dot NUMBER
band  = (2.5..3.5)             // NUMBER RANGE NUMBER — a range lives in its own slots
u     = bollinger(close, 20).upper   // `.upper` is member access, not a malformed number
sci   = 1.5e3                  // exponent form
neg   = -0.5                   // unary minus applied to NUMBER (the sign is not part of the token)

Suffixed literals

Six literal classes carry their unit as a suffix glued to the number (no space); SPAN alone also accepts exactly one space. Each token is typed at the expression position by rule [LitTyped] of the kind system:

Token Form Examples Kind
DUR NUMBER s | ms 2s, 300ms, 1.5s duration
PCT NUMBER % 50%, 0.5% ratio
PX NUMBER px 4px, 12px num tagged px (screen space)
DEC NUMBER d 1.50d, 1.5e3d decimal(scale)
SPAN NUMBER [ ] bar | bars 200 bars, 1 bars, 3bar barspan
RATE NUMBER /s | /min 40/s, 3/min num·T⁻¹ (per-time rate)
FLUX
b  = 300s + 20ms               // duration arithmetic
c  = 50%                       // ratio
d  = 4px                       // screen-space size (CANVAS styling)
r  = 40/s                      // emission rate
sp = life(200 bars)            // barspan — unifies every(n bars) and life: n bars
a  = 1.50d                     // decimal, scale 2 — exact money arithmetic

The glued-d rule (representative of the whole family). The d suffix must touch the number, and the character after the suffix must not extend an identifier:

FLUX
a = 1.50d       // DEC(1.50, scale 2) — one token
b = 1.5e3d      // DEC — the suffix composes with scientific notation
c = 1.50 d      // ✗ syntax error — NUMBER IDENT, two adjacent primaries (no juxtaposition)
d = 2           // a detached `d` is an ordinary identifier — a licit binding
d2 = 1.50 * d   // valid — 1.50 times the bound `d`

The same boundary check protects every suffix: 2se is not a DUR (the trailing e would extend an identifier, so the lexer yields NUMBER(2) then IDENT(se), which the parser rejects as juxtaposition), and 40/sec is not a RATE (it is 40 / sec, a division by the identifier sec). The DEC scale is the number of written fraction digits: 1.50d is decimal(2), 2d is decimal(0).

Why glued suffixes. Units on literals could have been separate tokens (1.50 d) or constructor calls (dur(300)), but both spellings put a parse between the number and its unit, and the detached word would collide with ordinary identifiers — d, s and bars are all reasonable binding names. Gluing makes the unit part of the token, so the decision is made by the lexer with zero grammar impact: 1.50d can never be misread, d alone is never stolen from the author, and each literal arrives at inference already carrying its kind. The one relaxation — 200 bars with a single space — is accepted because bars is a reserved head there and the span form reads as prose in every(…) and life: positions.

The plan keeps SPAN, PX, RATE and the binary % operator in v1 and flags them for final ratification; they are documented here as kept.

Strings

STRING = '"' frag '"' | "'" frag "'"

Both delimiters are equivalent — Flux has no character type, so the single quote is free to be a string delimiter. A string literal has kind string and must close on the line it opened (a raw newline inside a string is a syntax error). Inside a string, a backslash makes the next character literal: \" inside a "…" string, \{ and \} for literal braces (see interpolation below).

FLUX
plain  = "no interpolation"
single = 'quote style'
alert close cross_up open "crossed"       // strings feed the text channels (labels, messages)

Strings are values of the categorical kind string — bounded, immutable text for labels, prompts and messages. They are never plotted as a series, and + on two strings concatenates (the one categorical overload of +; see operators).

String interpolation

A { inside a string opens an interpolation hole holding a full Flux expression. Lexically the literal is split into fragment tokens:

STR_HEAD = '"' frag '{'          (opening fragment — also with ' delimiter)
STR_MID  = '}' frag '{'          (middle fragment)
STR_TAIL = '}' frag '"'          (closing fragment)
interpStr = STR_HEAD expr { STR_MID expr } STR_TAIL

A string containing no unescaped { stays one atomic STRING token. \{ and \} denote literal braces inside a fragment. The interpolation tokenizer matches its closing delimiter to the opening one ("…" or '…'), and the { of a fragment opens a lexer mode that is bounded by the brace counter the lexer already keeps — so holes nest to any depth the program itself can nest, and tokenization stays linear and incremental.

FLUX
m = { a: 0 }
x = close
mark close > open "close {close} above {str(open)}"
label  = "nested {(m with { a: 1 }).a} brace"    // a full expression, inner braces counted
single = 'quote {x} style'                 // interpolation works in '…' too
lit    = "a literal \{brace\}"             // escaped — no hole opened

An interpolated literal has kind string. Its AST is a concatenation of the fragments with each hole’s expression formatted by the canonical formatter for its kind (fmt.* — pinned, identical across every execution target), so a mark or alert label can be dynamic without any formatting boilerplate. See text for the formatting rules.

Capability references (CAPREF)

CAPREF = IDENT ":" IDENT        (only inside a capability list)

An APP-plane descriptor declares what it may touch as a list of capability references — namespace:verb pairs like chart:read, storage:own, levels:write, or a bare IDENT for single-token capabilities (sfx):

FLUX
app structureGame {
  capabilities: [chart:read, storage:own, levels:write, sfx, app:launch]
  // …
}

CAPREF is deliberately not a greedy lexer rule. It is recognized only in one grammatical state — element of a capability list — so its : never competes with the other colons of the language (record fields f: v, properties at: (x, y), when …: children, the ternary’s :). Everywhere else, a:b is an identifier, a colon and an identifier with their usual meanings.

Inside that one state the lexer reads the first segment literally, even when it spells a reserved word: app:launch is a legal capability reference although app is a hard keyword everywhere else. The relaxation is scoped to exactly the namespace position of a capability list; it does not extend to any other position in the language.

This page owns only the token — the namespace:verb shape and where it is recognized. The catalogue of legal capabilities and what each verb grants is defined by FDK — host services, and the app descriptor that carries the list is specified in the app plane.

Operators and punctuation

Tokens Role
.. range — 2..200, fill a..b, lo..hi properties; never an arithmetic operator
-> THE arrow — one token, one grammar production, five contextual readings (grammar)
< > <= >= == != cross_up cross_down comparisons (one non-associative level)
+ - · * / % additive / multiplicative; - is also the unary minus
?? ?. ? null-coalescing · safe navigation · ternary head
@ clock suffix — close@"1d", sma(close, 9)@tf("4h")
. member access / UFCS call (and ..'s shorter sibling)
= : , ; binding, key/value and label colon, separators
( ) [ ] { } grouping and call · index and list · blocks, records, bodies
~ approximate-cadence marker in every(~ d) (CANVAS; see canvas)
| separator of variant constructors — its only role

Three lexical facts here are load-bearing:

Why a single arrow token. Arrow-like syntax appears in five places (lambdas, event wiring, tween pairs, match arms, comprehensions), and languages that grow separately spelled arrows for those roles force readers to memorize which arrow belongs where. Flux decrees one symbol: -> is one token and one grammar production, and the five readings are selected by the context — the guarding head (on, match, for … in) or, for the unguarded uses, by kind inference. The lexer’s contribution to that decree is minimal and strict: exactly one arrow spelling exists. The full disambiguation story is on the grammar page.

The keyword model

Flux keeps the hard keyword set as small as the grammar allows, in three tiers plus a deliberate non-tier.

Tier 0 — hard-reserved everywhere

The core binders, connectors, operators and literals are reserved at every position:

def  let  in  if  then  else  for  match  with  variant  record  app
and  or  not  cross_up  cross_down  true  false  na

plus the two purely medial connectors as and from. These words can never be an identifier — they are binders and separators whose contextual release would buy nothing and cost lookahead. The single scoped exception, described above, is the namespace position of a capability reference, where app:launch reads app literally.

Tier 1 — contextual heads

Every other keyword of the language is contextual: it is a keyword at its head position (the first token of its production, or a medial connector like over in morph … over d) and an ordinary IDENT in the five binding slots:

# Binding slot Example with a Tier-1 word
a field name (declaration or literal) record Stop { color: color }
b parameter name def zone(rect) = rect.w * rect.h
c member access (.name) gpu.dot, x = p.on
d kind name c: color, state: variant { … }
e argument label focus(view, over: 600ms, pad: 5%)

The Tier-1 words, by family:

The visibility modifiers pub, private and package form a distinct contextual subclass: they are not production heads but declaration prefixes, decided by one token of lookahead (followed by a declaration head or a binding, they are modifiers; followed by =, :, . or (, they are plain identifiers).

A Tier-1 word is a keyword only where the parser can shift it as one — at its head or medial position. Everywhere else the lexer hands back an ordinary identifier, so def f(render) = render * 2 is legal, and so is the binding render = 3: a bind’s left-hand side shifts an identifier, not a head. What you cannot do is use the word where its production expects it and mean something else.

FLUX
record Stop { color: color }               // (a) field name + (d) kind name — both `color`
representation pnf(box, rev) {
  transform: rebin(close, box, rev)
  render: column { at: (bar.i, hi), h: hi - lo, w: 1 }
}                                          // `render` is a keyword here — hook-head position
focus(view, over: 600ms)                   // (e) `over` as an argument label

Why contextual reservation. Reserving every head outright would force authors into distorted names — focusRef, bounds, vdot — precisely where focus, rect and dot are the intended vocabulary of the domain. Contextual reservation keeps the readability that keyword-headed statements buy (every statement is committed by its first token) while returning the words to authors in the five slots where no head can ever occur. The two sides are provably disjoint: the binding slots sit in parser states where only an IDENT is expected, so freeing a word there introduces no ambiguity anywhere — a claim the grammar build re-verifies mechanically on every change.

Reserved ahead of need

Flux reserves a word before shipping its production whenever the word is destined for the surface, so that no existing program can shadow it in the meantime (see additivity). The APP-plane words (app, match, capabilities, init, update, subs, contributes, view, emit, variant) and the module words (import, pub, private, package) were reserved this way from the first version and have since received their productions.

test is the current instance: a reserved word in v1 with no production. The test "name" { … } block is a tooling convenience whose production ships later, additively, together with its golden tests.

Deliberately NOT reserved — built-ins

The built-in value names are ordinary identifiers, not keywords:

close  open  high  low  volume  time  hl2  hlc3  ohlc4
bar  clock  screen  pane  ratio  depth  z  up  down  self  range

A program may bind them — and a style lint immediately flags the shadowing:

FLUX
close = 42            // legal — the shadowing lint flags: `close` hides the built-in series
plot close            // now plots 42 at every bar

Why not reserve them. These names are vocabulary, not structure. Reserving open, range or z would poison huge swaths of ordinary naming (every rectangle has a range, every record an open something), for zero parsing benefit — no grammar production is anchored on them. A lint gives the author the warning that matters (accidental shadowing of a data source) while keeping the keyword set minimal, which in turn keeps the additivity story clean: the fewer words the language owns, the fewer future collisions it must manage.

Lexical errors

Because both lexer and parser are total, malformed input yields diagnostics, never crashes. The characteristic lexical-level rejections:

FLUX
s = "unterminated          // ✗ syntax error — a string must close on its own line
x = 1.50 d                 // ✗ syntax error — juxtaposition (NUMBER then IDENT)
t = 2se                    // ✗ syntax error — `se` does not complete a duration suffix

Everything else that looks lexical — a < b < c, a stray ->, a { a, b } in expression position — is rejected one stage later, by the grammar; those diagnostics are catalogued on the grammar page.

See also