◆ Flux

text — strings, structured text, and editing

The string kind and the str.* / fmt.* surface belong to the sealed language core. The structured-text codec, the sanitized render, the editing protocol, highlighting, segmentation, diff, search and the validator catalogue above them are a sealed additive design, each landing as a pinned version with its own golden the way a table bump does. The page states the whole pillar in the present tense; its maturity sits in the status ledger.

Text is the substrate a general application language cannot avoid. A forum has posts, a course has lessons, an editor has a document, a form has a field somebody typed into, and a chart has a label under a mark. Flux carries all of it — under one wall and one unlock. The wall is rule A12: no regular expressions, no arbitrary parsing. The unlock is the move that already carried the feed and form codecs, the calendar tables and the Unicode tables — a fixed grammar compiled to a declared kind, bounded and pinned. The Unicode tables are the precedent that matters most here, because this pillar leans on them for the next four hundred lines: case, normalization, and the segmentation that tells a caret where it may stand. Nothing on this page relaxes A12; this pillar is what populates it, and the last two sections show why the ban is satisfiable rather than merely restrictive.

New here? Start with Guide §5 — Kinds →

Reading the examples. A line marked ✗ is often a bare expression fragment — Flux has no expression-statements, so it illustrates a kind rule rather than a program. Positive samples are always legal statements.

The string kind

string is bounded, immutable UTF-8 text: labels, prompts, messages, keys, a document’s source. It is a categorical sort — flat, like color — and four consequences follow.

It is outside numeric arithmetic, with exactly one overload. + on two strings is concatenation: the single categorical line in the + rule table. Nothing else in the numeric algebra touches a string.

It has no ordering. There is no string < string, because a defensible answer would need a locale, and a locale read out of the air is a value that differs between two readers. Equality is admitted — bit-equality on the UTF-8 bytes, which every engine computes identically. A locale-aware order exists, but only as a named, explicit, pinned combinator: see i18n.

It is never a series. A string is consumed by the channels that expect text — a mark’s label, an alert’s message, a ui node’s content — and it is not plottable.

FLUX
sym   = "BTC-USD"
label = sym + " — " + fmt.price(close)     // string + string → string
alert close cross_up ema(close, 50) "{sym} crossed its 50"

same = "a" == "b"          // ✓ signal — bit-equality
"a" < "b"                  // ✗ [ErrDim] — `string` is never an ordered kind
plot label                 // ✗ [ErrPlot] — a string feeds text channels; it is not traced

It is bounded. Every string has a length cap — declared at the boundary that produces it (a text input’s maxLen, a codec’s maxTextLen) or inherited from the global cap. Overflow is a deterministic truncation, and the truncation cuts at a scalar boundary, never in the middle of a scalar (which would leave invalid UTF-8). This is the vec(κ, N) discipline applied to text: the memory an instance can occupy is computable before it runs.

The unit is the Unicode scalar

len, slice, indexOf and split count and index in Unicode scalars (code points). Never a byte. Never a UTF-16 code unit.

Why this rule exists. The interpreter runs on the platform, and the compiled module runs on its own UTF-8 memory. If the interpreter delegated len to the platform’s string length, it would be counting UTF-16 code units, and the compiled module would be counting scalars. The two agree on ASCII and diverge on the first character outside it: a label carrying a currency symbol, an accented name, a CJK title or an emoji would have one length here and another there. Byte-identity (invariant I7) is the property that lets a script be re-executed and trusted; it cannot survive a string API that means two different things. So str.* never delegates to the platform’s string methods — it runs the same pinned routine on both sides.

Three units exist in this pillar, they are not interchangeable, and the difference between them is not academic:

Layer Unit Where you meet it
Storage UTF-8 byte invisible to the program; the small-string threshold is a storage detail
The string kind Unicode scalar len, slice, indexOf, split, rep; truncation at the cap
Editing and display grapheme cluster Caret.off, Sel, truncate, Tok.start/len, diff lengths, highlight spans
Text Bytes Scalars Graphemes
abc 3 3 3
é as one code point 2 1 1
é as e plus a combining acute 3 2 1
a zero-width-joiner family emoji 25 7 1

The scalar is the canonical unit of the kind — the one two engines must agree on. The grapheme is the unit of the user — what a person calls “a character”, and therefore what a caret must step by: stepping by scalars would split that family emoji into four people and three joiners, and let an insertion point land between a letter and its accent. Both are pinned; neither is the platform’s.

str.* and fmt.*

Function Notes
len, slice, indexOf, contains, startsWith, endsWith scalar-indexed; slice is bounded
split(s, sep, maxParts) the part count is declared, so the result length is const-folded
trim, pad, padStart, padEnd, rep(s, n) rep’s count is a literal — the output length is known at compile time
replace(s, from, to) literal replacement, bounded — not a pattern
upper, lower locale-invariant, through the pinned Unicode case table
normalize(s, form) NFC / NFD, riding the same sealed table
truncate(s, n) grapheme-safe — it never cuts a cluster in half
graphemes(s), graphemeAt, graphemeSlice the segmentation surface (below)
fmt.num, fmt.price, fmt.pct, fmt.time the pinned canonical formatter
fmt.cat the concatenation an interpolated literal desugars to

The pinned formatter

fmt.num/price/pct/time are one canonical routine, shared byte-for-byte by the interpreter, the compiled module and the server. They are never the platform’s number-to-string.

Why this rule exists. Two engines disagree about numbers-as-text in ways nobody notices until a golden fails: how many decimals a f64 renders by default, which way the last digit rounds, and at what magnitude the output flips into scientific notation. A single divergent digit in a label is a divergent output, and the byte-identity oracle would then fail on every script that prints a number — which is nearly all of them. Text formatting is held to the same standard as the transcendental functions: one pinned routine, one golden, no exceptions.

upper and lower are locale-invariant for the same reason. That is a deliberate limit, not an oversight: locale-aware case belongs to i18n, where the locale is an explicit argument and the tables are versioned.

Interpolation

A { inside a string literal opens a hole holding a full Flux expression. The literal lexes into fragment tokens and the parser interleaves the expressions; the AST is a fmt.cat of the fragments, with each hole desugared through fmt.* according to its kind. So a label is dynamic without any string-building API:

FLUX
sym   = "BTC-USD"
stamp = fmt.time(time, "HH:mm")
mark close cross_up ema(close, 50) "{sym} crossed at {fmt.price(close)}"

Literal braces escape as \{ and \}, and both delimiters — "…" and '…' — behave identically; the token-level details are in lexical structure.

Memory: small strings, the arena, and promotion

Nearly every string in an application is short — a label, a formatted price, a key, a prompt. Those live inline in the value itself (the small-string optimization) and cost zero allocations. Longer ones go into a bump arena that is reset once per evaluation tick — per bar, per frame. There is no garbage collector: Flux is pure and its lifetimes are bounded, so the arena’s reset is the deallocation.

Concatenation fuses. The line below does not build three intermediates: it compiles to one length computation and one arena write — the string-builder pattern, made invisible by purity and common-subexpression elimination.

FLUX
def tag(c) = "px " + fmt.price(c) + " @ " + fmt.time(time, "HH:mm")

record Model { last: string ; n: num }      // `last` outlives the tick → promoted out of the arena

Promotion. A string that survives its tick is materialized out of the arena: copied into node-lifetime or Model-lifetime memory, never left as a view. Three things trigger it — capture by a scan or a stateful node, a field of a Model, and a checkpoint.

Why promotion is not an optimization detail. A checkpoint that stored a slice-view into a per-tick arena would, after that arena had been rewritten a thousand times, restore whatever happened to be sitting at those offsets. Scrubbing backwards through a session would produce different text on every attempt, and the replay would not be bit-exact. Copying on promotion is what reconciles “garbage-less” with “replayable” — two properties this language refuses to trade against each other.

Structured text — the Md codec

Markdown-class documents enter through a codec, not a parser. What the grammar admits (closed and versioned as md-v1, a strict CommonMark subset): ATX headings 1–6 · paragraphs · emphasis and strong · inline code · fenced and indented code blocks · blockquotes with bounded nesting · ordered and unordered lists with bounded nesting · thematic breaks · links · images · tables with bounded columns · hard breaks.

Permanently excluded: raw markup passthrough. There is no production for it, and the sanitizer below would not accept it if there were.

Footnotes and definition lists are not permanent exclusions — they are named for a later, additive grammar version (md-v2), which arrives as a new pinned version with its own golden, exactly as a table bump does.

The output is a bounded node-pool tree, not a recursive kind:

FLUX
variant MdNode {
  Doc | Heading(level: num) | Para | Em | Strong | Code | CodeBlock(lang: string)
  | Quote | List(ordered: signal) | Item | Link(href: string) | Image(ref: string)
  | Table | Row | Cell | Text(s: string) | Break | Rule
}

The document is a Tree(MdNode, N) — the node pool from collections. The caps are declared at the decode site, and overflow is a bounded truncation with a diagnostic, the same discipline as vec(κ, N):

FLUX
MD_CAPS = { maxNodes: 2000, maxDepth: 8, maxTextLen: 4000 }

def article(body) = md.parse(body, MD_CAPS)     // → Tree(MdNode, N)

Two entry points, one pinned routine. Md in the net codec catalogue decodes a fetched body straight to the tree at the boundary; md.parse(s, caps) does the same in the script, which an editor needs in order to preview a draft living in the Model. They are the same routine — one source of truth, interpreter ≡ compiled module, with a golden per grammar version.

Why this is a codec and not a parser. A parser is a program that runs on data, and its cost is a function of the data. A codec is a projection into a declared kind: the grammar is fixed before the program runs, the depth and the node count are capped at the call site, and the work is therefore bounded by numbers the compiler can read. The distinction is exactly what makes totality survive contact with text. It is also why decoding a payload never appears in your code as parsing — see the last section.

A link is data. Link(href) carries a string, and a string is not an authority. An href becomes navigation, and a ref becomes a load, only at the render boundary, under the host’s policy. A tree that arrived from the network cannot reach anything by itself.

The sanitized render

The text pipeline Figure — bytes become a bounded tree, the host holds every authority, and only committed edits enter the journal.

richText(ast) -> ui is a host-rendered primitive in the closed ui catalogue: the host walks the tree and renders vetted constructs only.

Node What the host does
text runs inserted as text content, never as markup; typography from tokens
Link(href) a host-vetted anchor: internal routes resolve through the navigation allowlist; an external href gets the host’s external-link affordance and opens through the host
Image(ref) resolved only through asset:load (allowlist plus quota), or dropped with a placeholder and a diagnostic
CodeBlock(lang) highlighted (below)
a malformed node rejected

Why images go through the asset policy. A fetched document that could hotlink a pixel would be a tracking beacon, and the reader would have no way to know. Routing every image through the allowlist means a document loads only what the application’s asset policy already admits — the network cannot introduce a new origin by writing one into a link.

There is no “unknown node class” in transit: MdNode is a closed variant produced by a pinned routine, so the sanitizer never guesses at a foreign tag — it judges only the malformed, which is a far smaller and far more decidable job. A prose container then wraps long-form output with a reader-width measure and vertical rhythm tokens; course pages, documentation bodies and forum posts are its consumers.

Relative hrefs inside a fetched document have two admissible treatments — resolve them against the feed’s origin, or forbid them outright; the plan names both and fixes the choice at the first consumer that fetches documents.

The editing protocol

A rich-text editor is a widget in the ui catalogue, but the interesting part is not the widget — it is the protocol underneath it, designed once, here, so that editing is replay-exact no matter what the host’s input stack does. Positions are data, and they are grapheme-safe:

FLUX
record Caret { node: num ; off: num }       // `off` counts GRAPHEME CLUSTERS
record Sel   { anchor: Caret ; focus: Caret }

Edits are messages. The host widget delivers them through constructors the application declared — the same OnX(args, C) carve-out every host event uses:

FLUX
record Caret { node: num ; off: num }
record Sel   { anchor: Caret ; focus: Caret }
variant EditOp  { Insert(at: Caret, s: string) | Delete(r: Sel) | Replace(r: Sel, s: string)
                | SetSel(r: Sel) | SetMark(r: Sel, m: MarkKind) }
variant MarkKind { Em | Strong | Code | Link(href: string) }

One reducer applies them. text.apply(doc, op) -> doc is a single pinned, total function — and totality holds at the cap, not below it: no input, and no sequence of inputs, makes it fail to return a document.

Situation Behaviour
a position outside the valid range clamped to the range
an operation wholly outside the document no-op plus a diagnostic (the vec.setAt precedent)
an edit spanning node boundaries the node pool splits and merges deterministically
Replace(r, s) exactly Delete(r) then Insert(…, s) — so it inherits both disciplines
SetMark over part of a text run the run splits at the selection edges
an edit that would overflow the pool’s cap N no-op plus a diagnostic — never an overrun

IME composition never enters the journal. While a composition is in flight, its intermediate states are presentation-local — continuous-class input, the same class as a drag’s in-flight position or a scroll offset. Only the committed text lands, as an Insert or a Replace message.

Why the journal only sees commitments. Input-method engines differ — the same keystrokes produce different intermediate candidate strings on different platforms and different versions. If those intermediates were journaled, a session recorded on one machine would not re-fold on another, and undo would step through candidate states no user ever chose. Journaling the commitment makes replay byte-exact across IME engines, and makes undo mean what a writer expects it to mean.

Undo is the application journal — its bounds and its coalescing, nothing else. There is no second undo stack inside the editor, which is why undo cannot resurrect a stale selection: the Model’s doc sub-record is versioned by history, and its ui sub-record is not.

FLUX
variant Msg { Edit(op: EditOp) | Move(r: Sel) | Undo }

app notes {
  capabilities: [ storage:own, journal ]

  init(p)        = { doc: md.parse(p.seed, MD_CAPS), ui: { sel: p.sel } }   // MD_CAPS: above
  update(m, msg) = match msg {
                     Edit(op) -> { model: m with { doc: text.apply(m.doc, op) }, cmds: [] }
                     Move(r)  -> { model: m with { ui:  m.ui with { sel: r } },  cmds: [] }
                     Undo     -> { model: m, cmds: [ Journal(UndoToMark) ] }
                   }
  view(m)        = prose { richText(m.doc) }
  subs(m)        = []
}

Move writes only into ui, so a caret movement is not an undoable step; Edit writes into doc, so it is. The partition is the undo semantics. A plain multi-line textarea is the same protocol minus SetMark.

Syntax highlighting

FLUX
record Tok { start: num ; len: num ; class: TokClass }     // start/len in GRAPHEME clusters
variant TokClass { Kw | Ident | Num | Str | Comment | Op | Punct | Plain }

TXT_CAPS = { maxTextLen: 4000 }
def toksOf(src) = hl.tokens("flux", src, TXT_CAPS)         // vec(Tok, N)

The grammars are a closed catalogue, exactly like the codecs: bounded single-pass tokenizers, pinned and versioned per language. The initial set is flux, json, js, html-escaped, md. codeBlock(lang, text) -> ui renders the classes through the theme’s tokens.

An unknown lang renders as plain text with a diagnostic — never a guess. Detection by heuristic is a non-goal: it would make a document’s rendering depend on a classifier, and a classifier is exactly the kind of thing that changes its mind between two versions.

Which languages the catalogue takes on after the committed set is a growth decision, not a grammar one: the plan lists candidates, and each is added the same way — a pinned, versioned tokenizer with a golden.

Unicode segmentation

The tables are pinned and versioned, and they are the foundation the editing model stands on:

Table Gives you
Grapheme clusters (UAX #29) str.graphemes, graphemeAt, graphemeSlice; caret arithmetic; truncate
Word and line break (UAX #29 / #14) word-jump for the caret; wrap hints for the renderer
Case and normalization upper, lower, normalize (NFC / NFD) — one sealed table, two uses

The routine is shared between the interpreter and the compiled module, and it is never the platform’s segmenter. This is the formatter’s argument again: a platform table is a moving target that ships on the platform’s schedule, and two engines on two versions would then disagree about where a caret may stand. A pinned table with a version number is a table you can put in a golden.

Grapheme-safe truncation supersedes scalar truncation for anything a person reads; scalar truncation remains the rule at the kind’s cap, because that boundary is about storage validity — never split a scalar, never emit invalid UTF-8 — rather than about what a reader sees.

Diff and patch

FLUX
variant Edit { Keep(len: num) | Ins(s: string) | Del(len: num) }   // lengths in GRAPHEME clusters

edits = txt.diff(prev, next, 400)          // vec(Edit, K) — bounded by the declared maxD
back  = txt.patch(prev, edits)             // total; `back == next` when the diff was exact

The algorithm is Myers, bounded by a declared maxD. Beyond that bound it does not fail and it does not run longer: it falls back to a coarse result in the same variant — one Del of the old text, one Ins of the new — with a diagnostic. txt.diff is therefore total by construction, and every consumer handles one shape. txt.patch is total. txt.diffLines(a, b, maxD) shares the routine at line granularity.

The tie-break inside the longest-common-subsequence search is one canonical choice, pinned, with a golden — because two equally good diffs are two different byte outputs, and byte-identity does not accept “equally good”. Consumers are the ordinary ones: revision history, the “changes” gutter, and optimistic-UI reconciliation when the server’s answer arrives.

The search stack

Search composes on collections; this pillar supplies the text-side pieces, all bounded and all pinned:

Piece Signature Notes
Tokenizer search.tokens(s, caps) -> vec(string, N) UAX #29 word boundaries plus pinned stopword tables
Fuzzy match search.fuzzy(q, s, maxDist: lit) bounded Levenshtein → record{ hit: signal ; dist: num }
Subsequence search.subseq(q, s) the command-palette match → record{ hit: signal ; score: num }
Prefix lookup m.range(lo, hi, k: lit) a range scan on the ordered Map — no trie kind exists, and none is needed
Ranking search.bm25(postings, stats, q, k: lit) deterministic scoring, stable tie-break by document id
Highlight spans search.spans(q, s, caps) grapheme-indexed spans, feeding the render

The indexes are ordinary bounded collections, and an index is an ordinary record holding them — three fields, one per question you can ask of it:

Field Kind Answers
terms Map(Token, Set(DocId, D), T) membership, and — because the Map is ordered — prefix
postings Map(Token, vec(record{ doc: DocId ; tf: num }, N), T) which documents, and how often
stats record{ n: num ; avgdl: num ; docLen: Map(DocId, num, D) } the corpus norms the ranker divides by

Both bounded calls take their result cap as the named argument k, exactly as the Map’s own range does in collections — the count is a declared ceiling on the answer, not one more positional number to miscount:

FLUX
def query(idx, q) = {
  hits: search.bm25(idx.postings, idx.stats, q, k: 20),   // vec(record{ doc, score }, 20)
  head: idx.terms.range("flu", "flv", k: 10)              // the typeahead window, O(log N + k)
}

Stopword tables are per-locale, with en as the base — the same uniform treatment i18n gives every table. A stemmer exists as an explicit, pinned variant, off by default: stemming changes what a query means, and that should be a decision rather than a default.

Two smaller boundaries stay unsettled by design until a consumer needs them fixed: which stemmer languages follow the first two, and whether the stopword tables are owned by this pillar or by i18n.

Validators — the catalogue that replaces patterns

Every validator is a predicate over a fixed grammar, and the set of them is a closed catalogue — never a pattern the caller supplies.

Validator Grammar
valid.isEmail(s) the WHATWG email grammar
valid.isUrl(s) the URL grammar (the same one the URL codec uses)
valid.isPhone(s) the E.164 shape
valid.luhn(s) the checksum
valid.isSlug(s) the slug shape
valid.inRange(x, lo, hi) a numeric bound
valid.matches(s, fmt) a named format, drawn from a closed variant
FLUX
variant DateFmt { Iso8601 | Rfc3339 | Ymd | Dmy | Mdy }
variant NamedFormat { Date(f: DateFmt) | Hex | Base64 | Uuid | Iban }
variant Rule { Format(NamedFormat) | Email | Url | Phone | Luhn | Slug | InRange(lo: num, hi: num) | Required }

DateFmt is itself a closed enumeration of pinned date shapes — never a user-supplied pattern string, which would be a pattern language smuggled in through a parameter. The catalogue grows by addition: a new entry is a new pinned routine with a golden. It never grows a runtime grammar.

FLUX
valid.matches(s, "^[a-z]+$")    // ✗ [ErrArg] — `matches` takes a NamedFormat, not a pattern
re.match("(a+)+b", s)           // ✗ [ErrUnbound] — no such name: there is no regex engine

Forms are validated field-wise. valid.form(form, rules) takes a record whose fields parallel the form’s, each carrying a Rule, and returns a record that parallels them again, each field an Ok or an Err. Required is a presence check — the field is non-na and non-empty — and it is evaluated before the value predicate, so a missing field reports “missing” rather than “malformed”. Every other arm names one of the predicates above. A Rule is data, not a function value: the descriptor is closed, which keeps the no-arrow discipline intact and lets the editor show the rule set as a table.

FLUX
variant Check { Ok | Err(reason: string) }

def errText(v) = match v { Ok -> "" ; Err(r) -> r }       // renders one field's verdict

def validate(form) =
  let rules = { email: Email, age: InRange(13, 120) } in   // one Rule per field
  valid.form(form, rules)                                  // { email: Ok, age: Err(reason) }

note = errText(Err("too young"))                           // the message a field renders under itself

The per-field record is exactly what a text input or a form widget consumes to render its own error — which is why validation never needs a side channel.

What is excluded, and why the ban is satisfiable

There are no regular expressions. Not “discouraged” — no name in the language evaluates one.

Why they are excluded. A regular expression is an unbounded computation described by data. Its cost is not a function of the input’s declared cap but of a pattern that arrives at runtime, and a backtracking engine’s worst case is catastrophic on inputs that look ordinary. A language whose central promise is that every program terminates within a budget the compiler can state cannot admit a construct whose budget is written by whoever supplies the pattern. This is the same reason there is no filter that shrinks a vector and no unbounded queue: the exclusions are one exclusion, applied consistently.

A ban is only honest if the work it forbids can still be done. This one is replaced from two sides at once:

  1. Validation is the named, bounded, deterministic catalogue above. You do not write a pattern for an email address, a URL, a UUID or an IBAN — you name the format, and the grammar behind the name is fixed, pinned and golden-tested. If a format is missing, the answer is a catalogue entry, not a pattern language.
  2. Structure never needs parsing, because it arrives already typed. A payload is decoded against the schema the application declared (Json(Trade), Md, Csv, a URL form), and a broken required field surfaces as DecodeError(field, reason) — never a silent na that poisons a computation three hops later. See net.

Between them, the cases that usually reach for a pattern — “is this a valid address”, “pull the fields out of this body”, “check the shape of this identifier” — are covered by constructs whose cost is a number written in the source.

Locale-dependent case and collation are also excluded from the core. upper and lower are locale-invariant, and there is no < on strings. That work exists — with an explicit locale and pinned tables — in i18n, so a computed value never depends on who is reading it.

One door is left ajar, and the plan describes it without walking through it: a pattern that is a compile-time literal could be compiled at build into a pinned automaton (linear time, no backtracking, capped size) — at which point it is not a regular expression in the A12 sense but sugar over the fixed-grammar discipline, because the constant pattern is a fixed grammar and the automaton is the pinned routine. If a named consumer ever justifies it, it would arrive as re.match(litPattern, s) and re.find, literal-only forever. A pattern that arrives at runtime, or through data, stays excluded permanently. That is the actual wall.

See also