Lexical structure
This page defines how a Flux source text becomes a stream of tokens: the source model, comments, the significant newline, the complete token catalogue (identifiers, the numeric family and its glued suffixes, strings and interpolation, capability references, operators and punctuation), and the keyword model — which words are reserved, where, and why most of them are only reserved contextually. Everything downstream — the grammar, kind inference, the editor — consumes exactly the token stream specified here.
Two properties frame the whole page. First, the lexer is total: any input text produces a token stream (malformed input surfaces as precise diagnostics, never as a crash). Second, the lexer is linear and incremental: every context-dependent device below (the significant newline, interpolation, capability references, the numeric munch) is bounded by counters the lexer already maintains, so retokenizing after an edit touches only the edited neighborhood — the property that keeps live preview under its frame budget.
New here? Start with Guide §3 — Your first session → for the same tokens met in program order, then return here for the exact rules.
Source model
A Flux program is a Unicode text. The fixed alphabet of the language itself — identifiers, keywords, operators, punctuation — is ASCII; arbitrary text (any Unicode) lives inside string literals and comments. Between tokens, spaces and tabs are insignificant and may appear in any number; newlines are significant (see TERM below).
One rule shapes many diagnostics on this page: Flux has no juxtaposition. Two adjacent primary expressions with no operator between them never form a term:
x = 1.50 d // ✗ syntax error — NUMBER then IDENT: two adjacent primaries
y = 1.50d // decimal(2) — the glued suffix makes this ONE token
d = 2 // `d` alone is an ordinary identifier — a licit binding
z = 1.50 * d // valid — the detached `d` is just a name hereBecause adjacency is never meaningful, the lexer can afford glued suffix tokens (1.50d,
4px) without ambiguity: either the suffix touches the number and the pair is one token, or it
does not and the program is ill-formed unless an operator intervenes.
A script may span several .flux source files; the package visibility modifier (see
grammar — modules) is scoped to exactly that set of files.
Comments
| Form | Token | Extent | Role |
|---|---|---|---|
// … |
LINE | to end of line | ordinary comment, skipped |
/// … |
DOC | to end of line | doc-comment, attaches to the next def |
/// z-score of a series over n bars
def zscore(x, n=20) = (x - sma(x, n)) / stdev(x, n)
plot zscore(close, 20) as z // an ordinary commentA doc-comment is lexically a comment (the parser skips it) but it is not thrown away: the
documentation pipeline collects the /// block preceding a def and publishes it as the
definition’s hover documentation and doc-as-data entry (see Guide §12 — Working in the
editor).
There are no block comments; a comment always ends at the newline. Comments are transparent to
the significant-newline rules below — a line that ends in a trailing // … comment is classified
by the last token before the comment, and a comment-only line neither ends nor continues a
statement.
The significant newline (TERM)
Flux has no mandatory statement terminator. A statement ends at the end of its line — the lexer
emits a TERM token at the newline — unless the line is visibly unfinished. Semicolons are
optional separators (never required); writing several statements on one line with ; is legal
but idiomatic Flux is one statement per line.
Figure — the three-clause newline decision: open
( / [ depth, a pending line end, or a continuator next each continue the statement; when none holds, the lexer emits TERM and the statement ends.
A newline does not emit TERM (the statement continues) when any of the three clauses holds:
-
(a) Open parenthesis or bracket. The
(/[nesting depth is greater than zero. Inside parentheses and square brackets, newlines are pure whitespace — which is how a long call wraps across lines. Braces do not reset that depth: a brace body written inside a call is still at depth > 0, so its items need an explicit,or;separator. At depth zero, a newline separates items, which is why a top-level multi-line record ormatchneeds no commas at all:FLUXdef f(r) = r.a + r.b m = { a: 1 // depth 0 — the newline separates b: 2 } n = f({ a: 1 ; b: 2 }) // inside a call — the `;` is REQUIRED -
(b) The line is pending. The last significant token of the line is an operator, an opening delimiter, a separator, or a keyword that grammatically awaits an operand —
if,then,else,let,in,not,and,or,as,from,with,over,def,match,variant,record,app,import,type,representation,tool,on,every,when,tween,spawn,burst,emit,rate,set,morph,replay,focus,cross_up,cross_down,pub,private,package,scene. A contextual keyword only counts here in its keyword role: inx = p.onthe trailingonis a field name and the statement ends. (One lexical subtlety: a trailing%glued to a digit is a percent literal, not a pending modulo —50%ends the line.) -
© The next token is a continuator. The first token of the following line can only extend an expression, never begin a new item.
Clause © is not an ad-hoc list. Normative definition: a token is a continuator iff it
belongs to no FIRST set of any list item — formally, iff it is outside ⋃ FIRST(item) taken
over every repeated-list production of the grammar (statements, view children, variant
constructors, representation and tool hooks, match arms, record fields, capability references,
app members). The set is computed from the grammar, not maintained by hand, which makes the
rule machine-checkable and keeps it in lockstep with the grammar forever.
Illustratively, the continuators are the purely infix or postfix tokens — . (member access),
.., @, +, *, /, %, <, >, <=, >=, ==, !=, cross_up, cross_down, and,
or, ->, with, ?, ??, ?., :, ,, ;, the closers ) ] }, the variant separator
| — and the purely medial keywords then, else, in, as, from.
Three tokens are deliberately not continuators even though they can extend an expression,
because they can also begin a new item: - (unary minus), [ (a list literal) and ( (a
parenthesized form). A leading . is a continuator only when not followed by a digit — .upper
continues, but .5 is a number and begins an item. To continue across a line on one of these
ambiguous tokens, end the previous line pending (clause b) instead:
u = bollinger(close, 20)
.upper // `.` continues the postfix chain — clause (c)
total = close
+ open - // `+` continues (c); trailing `-` leaves the line pending (b)
low
cond = if close > open
then 1
else 0 // medial keywords are continuators — clause (c)
variant Wide { A | B
| C } // `|` is a continuator inside the declaration
multi = sma(
close,
20) // clause (a): inside ( ) newlines are whitespace
m = { a: 1 }
updated = m
with { a: 2 } // `with` is a continuator — postfix record updateThe clause-(b) idiom for the ambiguous heads: a -⏎ b (line ends on the operator), v[⏎ i]
(line ends on the opener). Similarly, over is not a continuator — as a contextual keyword it
may begin an item (for example as a property key) — so a multi-line morph … over d breaks
after over, which is a pending word under clause (b).
Why this rule exists. A terminator-free surface reads like the notation authors actually sketch, but it must never become whitespace guesswork. The three clauses are decidable from at most one token of lookahead and the counters the lexer already carries, so newline classification is O(1) per line, linear over the file, and incremental under edits. And because the continuator set is defined as the complement of the computed item-head set, there is no hand-curated list to drift out of sync when the grammar grows: adding a statement head automatically removes it from the continuators.
Token catalogue
The complete token inventory. Each class is detailed in the sections that follow.
| Class | Tokens | Notes |
|---|---|---|
| Identifier | IDENT |
[A-Za-z_][A-Za-z0-9_]*, not reserved at its position |
| Numbers | NUMBER |
integer, decimal, leading-dot, exponent forms |
| Suffixed literals | DUR PCT PX DEC SPAN RATE |
glued suffixes; SPAN allows one space |
| Booleans / absence | true false na |
typed literals (signal; na inhabits every kind) |
| Strings | STRING |
"…" or '…', single line |
| Interpolation | STR_HEAD STR_MID STR_TAIL |
fragment tokens of "… {e} …" |
| Capability ref | CAPREF |
namespace:verb, only inside capability lists |
| Range / arrow | .. -> |
one token each |
| Comparison | < > <= >= == != cross_up cross_down |
non-associative level |
| Additive / multiplicative | + - · * / % |
- also unary |
The ? family |
?? ?. ? |
maximal munch, in that order |
| Punctuation | @ . , : ; = ( ) [ ] { } ~ | |
| separates variant constructors only |
| Comments | LINE DOC |
//, /// |
| Layout | TERM |
significant newline |
Identifiers
IDENT = [A-Za-z_] [A-Za-z0-9_]*Identifiers are ASCII: a letter or underscore, then letters, digits and underscores. An
identifier is a name for a binding, parameter, field, kind, argument label, module or
declaration. The lone underscore _ is lexically an ordinary identifier with two blessed roles
given to it by the grammar and the elaborator: the wildcard pattern in match arms, and the
implicit single-parameter placeholder in expression position (vec.map(_ * 1.1), see
operators).
Whether a given identifier is available depends on the keyword model below — most keywords in
Flux are contextual, so words like render, view or color remain usable as field names,
parameters and kind names.
Numbers
NUMBER = ( digits "." digits | "." digits | digits ) [ ("e"|"E") ["+"|"-"] digits ]42, 2.5, .5 and 1.5e3 are all NUMBER tokens. A bare NUMBER has kind lit — the
const-folded literal that is dimension-polymorphic (close + 10 is a price; see
kinds).
Maximal munch, bounded by the dot rule. The number scanner never consumes a . that is
followed by another . or by a non-digit. This single rule makes ranges and member access
compose with numbers without separators:
len = input(14, 2..200) // NUMBER RANGE NUMBER — the dot is never eaten before `..`
half = .5 // a leading-dot NUMBER
band = (2.5..3.5) // NUMBER RANGE NUMBER — a range lives in its own slots
u = bollinger(close, 20).upper // `.upper` is member access, not a malformed number
sci = 1.5e3 // exponent form
neg = -0.5 // unary minus applied to NUMBER (the sign is not part of the token)Suffixed literals
Six literal classes carry their unit as a suffix glued to the number (no space); SPAN
alone also accepts exactly one space. Each token is typed at the expression position by rule
[LitTyped] of the kind system:
| Token | Form | Examples | Kind |
|---|---|---|---|
DUR |
NUMBER s | ms |
2s, 300ms, 1.5s |
duration |
PCT |
NUMBER % |
50%, 0.5% |
ratio |
PX |
NUMBER px |
4px, 12px |
num tagged px (screen space) |
DEC |
NUMBER d |
1.50d, 1.5e3d |
decimal(scale) |
SPAN |
NUMBER [ ] bar | bars |
200 bars, 1 bars, 3bar |
barspan |
RATE |
NUMBER /s | /min |
40/s, 3/min |
num·T⁻¹ (per-time rate) |
b = 300s + 20ms // duration arithmetic
c = 50% // ratio
d = 4px // screen-space size (CANVAS styling)
r = 40/s // emission rate
sp = life(200 bars) // barspan — unifies every(n bars) and life: n bars
a = 1.50d // decimal, scale 2 — exact money arithmeticThe glued-d rule (representative of the whole family). The d suffix must touch the
number, and the character after the suffix must not extend an identifier:
a = 1.50d // DEC(1.50, scale 2) — one token
b = 1.5e3d // DEC — the suffix composes with scientific notation
c = 1.50 d // ✗ syntax error — NUMBER IDENT, two adjacent primaries (no juxtaposition)
d = 2 // a detached `d` is an ordinary identifier — a licit binding
d2 = 1.50 * d // valid — 1.50 times the bound `d`The same boundary check protects every suffix: 2se is not a DUR (the trailing e would
extend an identifier, so the lexer yields NUMBER(2) then IDENT(se), which the parser
rejects as juxtaposition), and 40/sec is not a RATE (it is 40 / sec, a division by the
identifier sec). The DEC scale is the number of written fraction digits: 1.50d is
decimal(2), 2d is decimal(0).
Why glued suffixes. Units on literals could have been separate tokens (
1.50 d) or constructor calls (dur(300)), but both spellings put a parse between the number and its unit, and the detached word would collide with ordinary identifiers —d,sandbarsare all reasonable binding names. Gluing makes the unit part of the token, so the decision is made by the lexer with zero grammar impact:1.50dcan never be misread,dalone is never stolen from the author, and each literal arrives at inference already carrying its kind. The one relaxation —200 barswith a single space — is accepted becausebarsis a reserved head there and the span form reads as prose inevery(…)andlife:positions.
The plan keeps SPAN, PX, RATE and the binary % operator in v1 and flags them for final
ratification; they are documented here as kept.
Strings
STRING = '"' frag '"' | "'" frag "'"Both delimiters are equivalent — Flux has no character type, so the single quote is free to be
a string delimiter. A string literal has kind string and must close on the line it opened
(a raw newline inside a string is a syntax error). Inside a string, a backslash makes the next
character literal: \" inside a "…" string, \{ and \} for literal braces (see
interpolation below).
plain = "no interpolation"
single = 'quote style'
alert close cross_up open "crossed" // strings feed the text channels (labels, messages)Strings are values of the categorical kind string — bounded, immutable text for labels,
prompts and messages. They are never plotted as a series, and + on two strings concatenates
(the one categorical overload of +; see operators).
String interpolation
A { inside a string opens an interpolation hole holding a full Flux expression. Lexically
the literal is split into fragment tokens:
STR_HEAD = '"' frag '{' (opening fragment — also with ' delimiter)
STR_MID = '}' frag '{' (middle fragment)
STR_TAIL = '}' frag '"' (closing fragment)
interpStr = STR_HEAD expr { STR_MID expr } STR_TAILA string containing no unescaped { stays one atomic STRING token. \{ and \} denote
literal braces inside a fragment. The interpolation tokenizer matches its closing delimiter to
the opening one ("…" or '…'), and the { of a fragment opens a lexer mode that is bounded
by the brace counter the lexer already keeps — so holes nest to any depth the program itself
can nest, and tokenization stays linear and incremental.
m = { a: 0 }
x = close
mark close > open "close {close} above {str(open)}"
label = "nested {(m with { a: 1 }).a} brace" // a full expression, inner braces counted
single = 'quote {x} style' // interpolation works in '…' too
lit = "a literal \{brace\}" // escaped — no hole openedAn interpolated literal has kind string. Its AST is a concatenation of the fragments with each
hole’s expression formatted by the canonical formatter for its kind (fmt.* — pinned, identical
across every execution target), so a mark or alert label can be dynamic without any formatting
boilerplate. See text for the formatting rules.
Capability references (CAPREF)
CAPREF = IDENT ":" IDENT (only inside a capability list)An APP-plane descriptor declares what it may touch as a list of capability references —
namespace:verb pairs like chart:read, storage:own, levels:write, or a bare IDENT for
single-token capabilities (sfx):
app structureGame {
capabilities: [chart:read, storage:own, levels:write, sfx, app:launch]
// …
}CAPREF is deliberately not a greedy lexer rule. It is recognized only in one grammatical
state — element of a capability list — so its : never competes with the other colons of the
language (record fields f: v, properties at: (x, y), when …: children, the ternary’s :).
Everywhere else, a:b is an identifier, a colon and an identifier with their usual meanings.
Inside that one state the lexer reads the first segment literally, even when it spells a
reserved word: app:launch is a legal capability reference although app is a hard keyword
everywhere else. The relaxation is scoped to exactly the namespace position of a capability
list; it does not extend to any other position in the language.
This page owns only the token — the namespace:verb shape and where it is recognized. The
catalogue of legal capabilities and what each verb grants is defined by
FDK — host services, and the app descriptor that carries the list is
specified in the app plane.
Operators and punctuation
| Tokens | Role |
|---|---|
.. |
range — 2..200, fill a..b, lo..hi properties; never an arithmetic operator |
-> |
THE arrow — one token, one grammar production, five contextual readings (grammar) |
< > <= >= == != cross_up cross_down |
comparisons (one non-associative level) |
+ - · * / % |
additive / multiplicative; - is also the unary minus |
?? ?. ? |
null-coalescing · safe navigation · ternary head |
@ |
clock suffix — close@"1d", sma(close, 9)@tf("4h") |
. |
member access / UFCS call (and ..'s shorter sibling) |
= : , ; |
binding, key/value and label colon, separators |
( ) [ ] { } |
grouping and call · index and list · blocks, records, bodies |
~ |
approximate-cadence marker in every(~ d) (CANVAS; see canvas) |
| |
separator of variant constructors — its only role |
Three lexical facts here are load-bearing:
- Maximal munch in the
?family: the lexer tries??, then?., then?. Soa??bis a coalescing,a?.ba safe navigation, anda ? b : ca ternary — no spacing tricks needed, exactly as..wins over.. ..and->are single tokens. Neither is ever two characters glued at parse time, sobb.upper..bb.lowercannot be misread (..never appears inside an expression’s postfix chain) and no second arrow spelling exists anywhere in the language.|has no operator role. It appears only between the constructors of avariantdeclaration (variant Phase { ask | suspense | revealed }). Boolean disjunction is the wordor; there is no bitwise-or operator (bitwise operations are the named functions ofbits.*). This keeps|collision-free with every expression level.
Why a single arrow token. Arrow-like syntax appears in five places (lambdas, event wiring, tween pairs, match arms, comprehensions), and languages that grow separately spelled arrows for those roles force readers to memorize which arrow belongs where. Flux decrees one symbol:
->is one token and one grammar production, and the five readings are selected by the context — the guarding head (on,match,for … in) or, for the unguarded uses, by kind inference. The lexer’s contribution to that decree is minimal and strict: exactly one arrow spelling exists. The full disambiguation story is on the grammar page.
The keyword model
Flux keeps the hard keyword set as small as the grammar allows, in three tiers plus a deliberate non-tier.
Tier 0 — hard-reserved everywhere
The core binders, connectors, operators and literals are reserved at every position:
def let in if then else for match with variant record app
and or not cross_up cross_down true false naplus the two purely medial connectors as and from. These words can never be an identifier —
they are binders and separators whose contextual release would buy nothing and cost lookahead.
The single scoped exception, described above, is the namespace position of a capability
reference, where app:launch reads app literally.
Tier 1 — contextual heads
Every other keyword of the language is contextual: it is a keyword at its head position
(the first token of its production, or a medial connector like over in morph … over d) and
an ordinary IDENT in the five binding slots:
| # | Binding slot | Example with a Tier-1 word |
|---|---|---|
| a | field name (declaration or literal) | record Stop { color: color } |
| b | parameter name | def zone(rect) = rect.w * rect.h |
| c | member access (.name) |
gpu.dot, x = p.on |
| d | kind name | c: color, state: variant { … } |
| e | argument label | focus(view, over: 600ms, pad: 5%) |
The Tier-1 words, by family:
- statement heads (ANALYSIS):
plot mark fill alert assert bars input - CANVAS heads and event verbs:
on group repeat spawn burst emit rate tween flash bounce pulse shake set when every hover click drag enter exit move wheel - CANVAS primitives:
dot circle ring rect square triangle poly line path text image svg sparkline backdrop column - TRANSITION heads and connectors:
morph over focus replay scene - representation heads and hooks:
representation transform render reduce liveReduce updateLastUnit persistKey - drawing-tool heads and hooks:
tool barExtent priceExtent - APP-plane member heads:
capabilities init update subs contributes view - tooling and modules:
test import type - clause connectors reserved contextually at their clause:
at(inassert … at),color(incolor bars:)
The visibility modifiers pub, private and package form a distinct contextual subclass:
they are not production heads but declaration prefixes, decided by one token of lookahead
(followed by a declaration head or a binding, they are modifiers; followed by =, :, . or
(, they are plain identifiers).
A Tier-1 word is a keyword only where the parser can shift it as one — at its head or medial
position. Everywhere else the lexer hands back an ordinary identifier, so def f(render) = render * 2
is legal, and so is the binding render = 3: a bind’s left-hand side shifts an identifier, not a
head. What you cannot do is use the word where its production expects it and mean something else.
record Stop { color: color } // (a) field name + (d) kind name — both `color`
representation pnf(box, rev) {
transform: rebin(close, box, rev)
render: column { at: (bar.i, hi), h: hi - lo, w: 1 }
} // `render` is a keyword here — hook-head position
focus(view, over: 600ms) // (e) `over` as an argument labelWhy contextual reservation. Reserving every head outright would force authors into distorted names —
focusRef,bounds,vdot— precisely wherefocus,rectanddotare the intended vocabulary of the domain. Contextual reservation keeps the readability that keyword-headed statements buy (every statement is committed by its first token) while returning the words to authors in the five slots where no head can ever occur. The two sides are provably disjoint: the binding slots sit in parser states where only anIDENTis expected, so freeing a word there introduces no ambiguity anywhere — a claim the grammar build re-verifies mechanically on every change.
Reserved ahead of need
Flux reserves a word before shipping its production whenever the word is destined for the
surface, so that no existing program can shadow it in the meantime (see
additivity). The APP-plane words (app, match,
capabilities, init, update, subs, contributes, view, emit, variant) and the
module words (import, pub, private, package) were reserved this way from the first
version and have since received their productions.
test is the current instance: a reserved word in v1 with no production. The test "name" { … }
block is a tooling convenience whose production ships later, additively, together with its golden
tests.
Deliberately NOT reserved — built-ins
The built-in value names are ordinary identifiers, not keywords:
close open high low volume time hl2 hlc3 ohlc4
bar clock screen pane ratio depth z up down self rangeA program may bind them — and a style lint immediately flags the shadowing:
close = 42 // legal — the shadowing lint flags: `close` hides the built-in series
plot close // now plots 42 at every barWhy not reserve them. These names are vocabulary, not structure. Reserving
open,rangeorzwould poison huge swaths of ordinary naming (every rectangle has arange, every record anopensomething), for zero parsing benefit — no grammar production is anchored on them. A lint gives the author the warning that matters (accidental shadowing of a data source) while keeping the keyword set minimal, which in turn keeps the additivity story clean: the fewer words the language owns, the fewer future collisions it must manage.
Lexical errors
Because both lexer and parser are total, malformed input yields diagnostics, never crashes. The characteristic lexical-level rejections:
s = "unterminated // ✗ syntax error — a string must close on its own line
x = 1.50 d // ✗ syntax error — juxtaposition (NUMBER then IDENT)
t = 2se // ✗ syntax error — `se` does not complete a duration suffixEverything else that looks lexical — a < b < c, a stray ->, a { a, b } in expression
position — is rejected one stage later, by the grammar; those diagnostics are catalogued on the
grammar page.
See also
- Grammar — the normative grammar these tokens feed; the single arrow; precedence.
- Kinds — the dimensional kinds that typed literals (
DUR,PCT,SPAN,DEC…) carry. - Operators — the dimensional algebra behind
+ - * / %, the?family, UFCS. - Time and state — what the
@clock suffix means and why it is restricted. - Text — the
stringkind, formatting (fmt.*) and the interpolation pipeline. - Host services — the capability catalogue behind
CAPREFverbs. - Guide §3 — Your first session — the same tokens met in program order.