text — strings, structured text, and editing
The string kind and the str.* / fmt.* surface belong to the sealed language core. The
structured-text codec, the sanitized render, the editing protocol, highlighting, segmentation, diff,
search and the validator catalogue above them are a sealed additive design, each landing as a
pinned version with its own golden the way a table bump does. The page states the whole pillar in
the present tense; its maturity sits in the status ledger.
Text is the substrate a general application language cannot avoid. A forum has posts, a course has lessons, an editor has a document, a form has a field somebody typed into, and a chart has a label under a mark. Flux carries all of it — under one wall and one unlock. The wall is rule A12: no regular expressions, no arbitrary parsing. The unlock is the move that already carried the feed and form codecs, the calendar tables and the Unicode tables — a fixed grammar compiled to a declared kind, bounded and pinned. The Unicode tables are the precedent that matters most here, because this pillar leans on them for the next four hundred lines: case, normalization, and the segmentation that tells a caret where it may stand. Nothing on this page relaxes A12; this pillar is what populates it, and the last two sections show why the ban is satisfiable rather than merely restrictive.
New here? Start with Guide §5 — Kinds →
Reading the examples. A line marked ✗ is often a bare expression fragment — Flux has no expression-statements, so it illustrates a kind rule rather than a program. Positive samples are always legal statements.
The string kind
string is bounded, immutable UTF-8 text: labels, prompts, messages, keys, a document’s
source. It is a categorical sort — flat, like color — and four consequences follow.
It is outside numeric arithmetic, with exactly one overload. + on two strings is
concatenation: the single categorical line in the + rule table. Nothing else in the numeric
algebra touches a string.
It has no ordering. There is no string < string, because a defensible answer would need a
locale, and a locale read out of the air is a value that differs between two readers. Equality is
admitted — bit-equality on the UTF-8 bytes, which every engine computes identically. A locale-aware
order exists, but only as a named, explicit, pinned combinator: see i18n.
It is never a series. A string is consumed by the channels that expect text — a mark’s
label, an alert’s message, a ui node’s content — and it is not plottable.
sym = "BTC-USD"
label = sym + " — " + fmt.price(close) // string + string → string
alert close cross_up ema(close, 50) "{sym} crossed its 50"
same = "a" == "b" // ✓ signal — bit-equality
"a" < "b" // ✗ [ErrDim] — `string` is never an ordered kind
plot label // ✗ [ErrPlot] — a string feeds text channels; it is not tracedIt is bounded. Every string has a length cap — declared at the boundary that produces it (a
text input’s maxLen, a codec’s maxTextLen) or inherited from the global cap. Overflow is a
deterministic truncation, and the truncation cuts at a scalar boundary, never in the middle of
a scalar (which would leave invalid UTF-8). This is the vec(κ, N) discipline applied to text: the
memory an instance can occupy is computable before it runs.
The unit is the Unicode scalar
len, slice, indexOf and split count and index in Unicode scalars (code points). Never a
byte. Never a UTF-16 code unit.
Why this rule exists. The interpreter runs on the platform, and the compiled module runs on its own UTF-8 memory. If the interpreter delegated
lento the platform’s string length, it would be counting UTF-16 code units, and the compiled module would be counting scalars. The two agree on ASCII and diverge on the first character outside it: a label carrying a currency symbol, an accented name, a CJK title or an emoji would have one length here and another there. Byte-identity (invariant I7) is the property that lets a script be re-executed and trusted; it cannot survive astringAPI that means two different things. Sostr.*never delegates to the platform’s string methods — it runs the same pinned routine on both sides.
Three units exist in this pillar, they are not interchangeable, and the difference between them is not academic:
| Layer | Unit | Where you meet it |
|---|---|---|
| Storage | UTF-8 byte | invisible to the program; the small-string threshold is a storage detail |
The string kind |
Unicode scalar | len, slice, indexOf, split, rep; truncation at the cap |
| Editing and display | grapheme cluster | Caret.off, Sel, truncate, Tok.start/len, diff lengths, highlight spans |
| Text | Bytes | Scalars | Graphemes |
|---|---|---|---|
abc |
3 | 3 | 3 |
é as one code point |
2 | 1 | 1 |
é as e plus a combining acute |
3 | 2 | 1 |
| a zero-width-joiner family emoji | 25 | 7 | 1 |
The scalar is the canonical unit of the kind — the one two engines must agree on. The grapheme is the unit of the user — what a person calls “a character”, and therefore what a caret must step by: stepping by scalars would split that family emoji into four people and three joiners, and let an insertion point land between a letter and its accent. Both are pinned; neither is the platform’s.
str.* and fmt.*
| Function | Notes |
|---|---|
len, slice, indexOf, contains, startsWith, endsWith |
scalar-indexed; slice is bounded |
split(s, sep, maxParts) |
the part count is declared, so the result length is const-folded |
trim, pad, padStart, padEnd, rep(s, n) |
rep’s count is a literal — the output length is known at compile time |
replace(s, from, to) |
literal replacement, bounded — not a pattern |
upper, lower |
locale-invariant, through the pinned Unicode case table |
normalize(s, form) |
NFC / NFD, riding the same sealed table |
truncate(s, n) |
grapheme-safe — it never cuts a cluster in half |
graphemes(s), graphemeAt, graphemeSlice |
the segmentation surface (below) |
fmt.num, fmt.price, fmt.pct, fmt.time |
the pinned canonical formatter |
fmt.cat |
the concatenation an interpolated literal desugars to |
The pinned formatter
fmt.num/price/pct/time are one canonical routine, shared byte-for-byte by the interpreter, the
compiled module and the server. They are never the platform’s number-to-string.
Why this rule exists. Two engines disagree about numbers-as-text in ways nobody notices until a golden fails: how many decimals a
f64renders by default, which way the last digit rounds, and at what magnitude the output flips into scientific notation. A single divergent digit in a label is a divergent output, and the byte-identity oracle would then fail on every script that prints a number — which is nearly all of them. Text formatting is held to the same standard as the transcendental functions: one pinned routine, one golden, no exceptions.
upper and lower are locale-invariant for the same reason. That is a deliberate limit, not an
oversight: locale-aware case belongs to i18n, where the locale is an explicit argument
and the tables are versioned.
Interpolation
A { inside a string literal opens a hole holding a full Flux expression. The literal lexes
into fragment tokens and the parser interleaves the expressions; the AST is a fmt.cat of the
fragments, with each hole desugared through fmt.* according to its kind. So a label is dynamic
without any string-building API:
sym = "BTC-USD"
stamp = fmt.time(time, "HH:mm")
mark close cross_up ema(close, 50) "{sym} crossed at {fmt.price(close)}"Literal braces escape as \{ and \}, and both delimiters — "…" and '…' — behave identically;
the token-level details are in lexical structure.
Memory: small strings, the arena, and promotion
Nearly every string in an application is short — a label, a formatted price, a key, a prompt. Those live inline in the value itself (the small-string optimization) and cost zero allocations. Longer ones go into a bump arena that is reset once per evaluation tick — per bar, per frame. There is no garbage collector: Flux is pure and its lifetimes are bounded, so the arena’s reset is the deallocation.
Concatenation fuses. The line below does not build three intermediates: it compiles to one length computation and one arena write — the string-builder pattern, made invisible by purity and common-subexpression elimination.
def tag(c) = "px " + fmt.price(c) + " @ " + fmt.time(time, "HH:mm")
record Model { last: string ; n: num } // `last` outlives the tick → promoted out of the arenaPromotion. A string that survives its tick is materialized out of the arena: copied into
node-lifetime or Model-lifetime memory, never left as a view. Three things trigger it — capture by
a scan or a stateful node, a field of a Model, and a checkpoint.
Why promotion is not an optimization detail. A checkpoint that stored a slice-view into a per-tick arena would, after that arena had been rewritten a thousand times, restore whatever happened to be sitting at those offsets. Scrubbing backwards through a session would produce different text on every attempt, and the replay would not be bit-exact. Copying on promotion is what reconciles “garbage-less” with “replayable” — two properties this language refuses to trade against each other.
Structured text — the Md codec
Markdown-class documents enter through a codec, not a parser. What the grammar admits
(closed and versioned as md-v1, a strict CommonMark subset): ATX headings 1–6 · paragraphs ·
emphasis and strong · inline code · fenced and indented code blocks · blockquotes with bounded
nesting · ordered and unordered lists with bounded nesting · thematic breaks · links · images ·
tables with bounded columns · hard breaks.
Permanently excluded: raw markup passthrough. There is no production for it, and the sanitizer below would not accept it if there were.
Footnotes and definition lists are not permanent exclusions — they are named for a later, additive
grammar version (md-v2), which arrives as a new pinned version with its own golden, exactly as a
table bump does.
The output is a bounded node-pool tree, not a recursive kind:
variant MdNode {
Doc | Heading(level: num) | Para | Em | Strong | Code | CodeBlock(lang: string)
| Quote | List(ordered: signal) | Item | Link(href: string) | Image(ref: string)
| Table | Row | Cell | Text(s: string) | Break | Rule
}The document is a Tree(MdNode, N) — the node pool from collections. The caps
are declared at the decode site, and overflow is a bounded truncation with a diagnostic, the
same discipline as vec(κ, N):
MD_CAPS = { maxNodes: 2000, maxDepth: 8, maxTextLen: 4000 }
def article(body) = md.parse(body, MD_CAPS) // → Tree(MdNode, N)Two entry points, one pinned routine. Md in the net codec catalogue decodes a
fetched body straight to the tree at the boundary; md.parse(s, caps) does the same in the script,
which an editor needs in order to preview a draft living in the Model. They are the same routine —
one source of truth, interpreter ≡ compiled module, with a golden per grammar version.
Why this is a codec and not a parser. A parser is a program that runs on data, and its cost is a function of the data. A codec is a projection into a declared kind: the grammar is fixed before the program runs, the depth and the node count are capped at the call site, and the work is therefore bounded by numbers the compiler can read. The distinction is exactly what makes totality survive contact with text. It is also why decoding a payload never appears in your code as parsing — see the last section.
A link is data. Link(href) carries a string, and a string is not an authority. An href
becomes navigation, and a ref becomes a load, only at the render boundary, under the host’s
policy. A tree that arrived from the network cannot reach anything by itself.
The sanitized render
Figure — bytes become a bounded tree, the host holds every authority, and only committed edits enter the journal.
richText(ast) -> ui is a host-rendered primitive in the closed ui catalogue: the host walks the
tree and renders vetted constructs only.
| Node | What the host does |
|---|---|
| text runs | inserted as text content, never as markup; typography from tokens |
Link(href) |
a host-vetted anchor: internal routes resolve through the navigation allowlist; an external href gets the host’s external-link affordance and opens through the host |
Image(ref) |
resolved only through asset:load (allowlist plus quota), or dropped with a placeholder and a diagnostic |
CodeBlock(lang) |
highlighted (below) |
| a malformed node | rejected |
Why images go through the asset policy. A fetched document that could hotlink a pixel would be a tracking beacon, and the reader would have no way to know. Routing every image through the allowlist means a document loads only what the application’s asset policy already admits — the network cannot introduce a new origin by writing one into a link.
There is no “unknown node class” in transit: MdNode is a closed variant produced by a pinned
routine, so the sanitizer never guesses at a foreign tag — it judges only the malformed, which is
a far smaller and far more decidable job. A prose container then wraps long-form output with a
reader-width measure and vertical rhythm tokens; course pages, documentation bodies and forum posts
are its consumers.
Relative hrefs inside a fetched document have two admissible treatments — resolve them against the feed’s origin, or forbid them outright; the plan names both and fixes the choice at the first consumer that fetches documents.
The editing protocol
A rich-text editor is a widget in the ui catalogue, but the interesting part is not the widget —
it is the protocol underneath it, designed once, here, so that editing is replay-exact no matter
what the host’s input stack does. Positions are data, and they are grapheme-safe:
record Caret { node: num ; off: num } // `off` counts GRAPHEME CLUSTERS
record Sel { anchor: Caret ; focus: Caret }Edits are messages. The host widget delivers them through constructors the application
declared — the same OnX(args, C) carve-out every host event uses:
record Caret { node: num ; off: num }
record Sel { anchor: Caret ; focus: Caret }
variant EditOp { Insert(at: Caret, s: string) | Delete(r: Sel) | Replace(r: Sel, s: string)
| SetSel(r: Sel) | SetMark(r: Sel, m: MarkKind) }
variant MarkKind { Em | Strong | Code | Link(href: string) }One reducer applies them. text.apply(doc, op) -> doc is a single pinned, total function —
and totality holds at the cap, not below it: no input, and no sequence of inputs, makes it fail to
return a document.
| Situation | Behaviour |
|---|---|
| a position outside the valid range | clamped to the range |
| an operation wholly outside the document | no-op plus a diagnostic (the vec.setAt precedent) |
| an edit spanning node boundaries | the node pool splits and merges deterministically |
Replace(r, s) |
exactly Delete(r) then Insert(…, s) — so it inherits both disciplines |
SetMark over part of a text run |
the run splits at the selection edges |
an edit that would overflow the pool’s cap N |
no-op plus a diagnostic — never an overrun |
IME composition never enters the journal. While a composition is in flight, its intermediate
states are presentation-local — continuous-class input, the same class as a drag’s in-flight
position or a scroll offset. Only the committed text lands, as an Insert or a Replace
message.
Why the journal only sees commitments. Input-method engines differ — the same keystrokes produce different intermediate candidate strings on different platforms and different versions. If those intermediates were journaled, a session recorded on one machine would not re-fold on another, and undo would step through candidate states no user ever chose. Journaling the commitment makes replay byte-exact across IME engines, and makes undo mean what a writer expects it to mean.
Undo is the application journal — its bounds and its coalescing, nothing else. There is no second
undo stack inside the editor, which is why undo cannot resurrect a stale selection: the Model’s doc
sub-record is versioned by history, and its ui sub-record is not.
variant Msg { Edit(op: EditOp) | Move(r: Sel) | Undo }
app notes {
capabilities: [ storage:own, journal ]
init(p) = { doc: md.parse(p.seed, MD_CAPS), ui: { sel: p.sel } } // MD_CAPS: above
update(m, msg) = match msg {
Edit(op) -> { model: m with { doc: text.apply(m.doc, op) }, cmds: [] }
Move(r) -> { model: m with { ui: m.ui with { sel: r } }, cmds: [] }
Undo -> { model: m, cmds: [ Journal(UndoToMark) ] }
}
view(m) = prose { richText(m.doc) }
subs(m) = []
}Move writes only into ui, so a caret movement is not an undoable step; Edit writes into doc,
so it is. The partition is the undo semantics. A plain multi-line textarea is the same protocol
minus SetMark.
Syntax highlighting
record Tok { start: num ; len: num ; class: TokClass } // start/len in GRAPHEME clusters
variant TokClass { Kw | Ident | Num | Str | Comment | Op | Punct | Plain }
TXT_CAPS = { maxTextLen: 4000 }
def toksOf(src) = hl.tokens("flux", src, TXT_CAPS) // vec(Tok, N)The grammars are a closed catalogue, exactly like the codecs: bounded single-pass tokenizers,
pinned and versioned per language. The initial set is flux, json, js, html-escaped, md.
codeBlock(lang, text) -> ui renders the classes through the theme’s tokens.
An unknown lang renders as plain text with a diagnostic — never a guess. Detection by heuristic
is a non-goal: it would make a document’s rendering depend on a classifier, and a classifier is
exactly the kind of thing that changes its mind between two versions.
Which languages the catalogue takes on after the committed set is a growth decision, not a grammar one: the plan lists candidates, and each is added the same way — a pinned, versioned tokenizer with a golden.
Unicode segmentation
The tables are pinned and versioned, and they are the foundation the editing model stands on:
| Table | Gives you |
|---|---|
| Grapheme clusters (UAX #29) | str.graphemes, graphemeAt, graphemeSlice; caret arithmetic; truncate |
| Word and line break (UAX #29 / #14) | word-jump for the caret; wrap hints for the renderer |
| Case and normalization | upper, lower, normalize (NFC / NFD) — one sealed table, two uses |
The routine is shared between the interpreter and the compiled module, and it is never the platform’s segmenter. This is the formatter’s argument again: a platform table is a moving target that ships on the platform’s schedule, and two engines on two versions would then disagree about where a caret may stand. A pinned table with a version number is a table you can put in a golden.
Grapheme-safe truncation supersedes scalar truncation for anything a person reads; scalar truncation remains the rule at the kind’s cap, because that boundary is about storage validity — never split a scalar, never emit invalid UTF-8 — rather than about what a reader sees.
Diff and patch
variant Edit { Keep(len: num) | Ins(s: string) | Del(len: num) } // lengths in GRAPHEME clusters
edits = txt.diff(prev, next, 400) // vec(Edit, K) — bounded by the declared maxD
back = txt.patch(prev, edits) // total; `back == next` when the diff was exactThe algorithm is Myers, bounded by a declared maxD. Beyond that bound it does not fail and it
does not run longer: it falls back to a coarse result in the same variant — one Del of the old
text, one Ins of the new — with a diagnostic. txt.diff is therefore total by construction, and
every consumer handles one shape. txt.patch is total. txt.diffLines(a, b, maxD) shares the
routine at line granularity.
The tie-break inside the longest-common-subsequence search is one canonical choice, pinned, with a golden — because two equally good diffs are two different byte outputs, and byte-identity does not accept “equally good”. Consumers are the ordinary ones: revision history, the “changes” gutter, and optimistic-UI reconciliation when the server’s answer arrives.
The search stack
Search composes on collections; this pillar supplies the text-side pieces, all bounded and all pinned:
| Piece | Signature | Notes |
|---|---|---|
| Tokenizer | search.tokens(s, caps) -> vec(string, N) |
UAX #29 word boundaries plus pinned stopword tables |
| Fuzzy match | search.fuzzy(q, s, maxDist: lit) |
bounded Levenshtein → record{ hit: signal ; dist: num } |
| Subsequence | search.subseq(q, s) |
the command-palette match → record{ hit: signal ; score: num } |
| Prefix lookup | m.range(lo, hi, k: lit) |
a range scan on the ordered Map — no trie kind exists, and none is needed |
| Ranking | search.bm25(postings, stats, q, k: lit) |
deterministic scoring, stable tie-break by document id |
| Highlight spans | search.spans(q, s, caps) |
grapheme-indexed spans, feeding the render |
The indexes are ordinary bounded collections, and an index is an ordinary record holding them — three fields, one per question you can ask of it:
| Field | Kind | Answers |
|---|---|---|
terms |
Map(Token, Set(DocId, D), T) |
membership, and — because the Map is ordered — prefix |
postings |
Map(Token, vec(record{ doc: DocId ; tf: num }, N), T) |
which documents, and how often |
stats |
record{ n: num ; avgdl: num ; docLen: Map(DocId, num, D) } |
the corpus norms the ranker divides by |
Both bounded calls take their result cap as the named argument k, exactly as the Map’s own
range does in collections — the count is a declared ceiling on the answer, not
one more positional number to miscount:
def query(idx, q) = {
hits: search.bm25(idx.postings, idx.stats, q, k: 20), // vec(record{ doc, score }, 20)
head: idx.terms.range("flu", "flv", k: 10) // the typeahead window, O(log N + k)
}Stopword tables are per-locale, with en as the base — the same uniform treatment
i18n gives every table. A stemmer exists as an explicit, pinned variant, off by
default: stemming changes what a query means, and that should be a decision rather than a default.
Two smaller boundaries stay unsettled by design until a consumer needs them fixed: which stemmer languages follow the first two, and whether the stopword tables are owned by this pillar or by i18n.
Validators — the catalogue that replaces patterns
Every validator is a predicate over a fixed grammar, and the set of them is a closed catalogue — never a pattern the caller supplies.
| Validator | Grammar |
|---|---|
valid.isEmail(s) |
the WHATWG email grammar |
valid.isUrl(s) |
the URL grammar (the same one the URL codec uses) |
valid.isPhone(s) |
the E.164 shape |
valid.luhn(s) |
the checksum |
valid.isSlug(s) |
the slug shape |
valid.inRange(x, lo, hi) |
a numeric bound |
valid.matches(s, fmt) |
a named format, drawn from a closed variant |
variant DateFmt { Iso8601 | Rfc3339 | Ymd | Dmy | Mdy }
variant NamedFormat { Date(f: DateFmt) | Hex | Base64 | Uuid | Iban }
variant Rule { Format(NamedFormat) | Email | Url | Phone | Luhn | Slug | InRange(lo: num, hi: num) | Required }DateFmt is itself a closed enumeration of pinned date shapes — never a user-supplied pattern
string, which would be a pattern language smuggled in through a parameter. The catalogue grows by
addition: a new entry is a new pinned routine with a golden. It never grows a runtime grammar.
valid.matches(s, "^[a-z]+$") // ✗ [ErrArg] — `matches` takes a NamedFormat, not a pattern
re.match("(a+)+b", s) // ✗ [ErrUnbound] — no such name: there is no regex engineForms are validated field-wise. valid.form(form, rules) takes a record whose fields parallel
the form’s, each carrying a Rule, and returns a record that parallels them again, each field an
Ok or an Err. Required is a presence check — the field is non-na and non-empty — and it
is evaluated before the value predicate, so a missing field reports “missing” rather than
“malformed”. Every other arm names one of the predicates above. A Rule is data, not a function
value: the descriptor is closed, which keeps the no-arrow discipline intact and lets the editor show
the rule set as a table.
variant Check { Ok | Err(reason: string) }
def errText(v) = match v { Ok -> "" ; Err(r) -> r } // renders one field's verdict
def validate(form) =
let rules = { email: Email, age: InRange(13, 120) } in // one Rule per field
valid.form(form, rules) // { email: Ok, age: Err(reason) }
note = errText(Err("too young")) // the message a field renders under itselfThe per-field record is exactly what a text input or a form widget consumes to render its own error — which is why validation never needs a side channel.
What is excluded, and why the ban is satisfiable
There are no regular expressions. Not “discouraged” — no name in the language evaluates one.
Why they are excluded. A regular expression is an unbounded computation described by data. Its cost is not a function of the input’s declared cap but of a pattern that arrives at runtime, and a backtracking engine’s worst case is catastrophic on inputs that look ordinary. A language whose central promise is that every program terminates within a budget the compiler can state cannot admit a construct whose budget is written by whoever supplies the pattern. This is the same reason there is no
filterthat shrinks a vector and no unbounded queue: the exclusions are one exclusion, applied consistently.
A ban is only honest if the work it forbids can still be done. This one is replaced from two sides at once:
- Validation is the named, bounded, deterministic catalogue above. You do not write a pattern for an email address, a URL, a UUID or an IBAN — you name the format, and the grammar behind the name is fixed, pinned and golden-tested. If a format is missing, the answer is a catalogue entry, not a pattern language.
- Structure never needs parsing, because it arrives already typed. A payload is decoded
against the schema the application declared (
Json(Trade),Md,Csv, a URL form), and a broken required field surfaces asDecodeError(field, reason)— never a silentnathat poisons a computation three hops later. See net.
Between them, the cases that usually reach for a pattern — “is this a valid address”, “pull the fields out of this body”, “check the shape of this identifier” — are covered by constructs whose cost is a number written in the source.
Locale-dependent case and collation are also excluded from the core. upper and lower are
locale-invariant, and there is no < on strings. That work exists — with an explicit locale and
pinned tables — in i18n, so a computed value never depends on who is reading it.
One door is left ajar, and the plan describes it without walking through it: a
pattern that is a compile-time literal could be compiled at build into a pinned automaton
(linear time, no backtracking, capped size) — at which point it is not a regular expression in the
A12 sense but sugar over the fixed-grammar discipline, because the constant pattern is a fixed
grammar and the automaton is the pinned routine. If a named consumer ever justifies it, it would
arrive as re.match(litPattern, s) and re.find, literal-only forever. A pattern that arrives
at runtime, or through data, stays excluded permanently. That is the actual wall.
See also
- Kinds — the
stringsort, bit-equality, and why there is no ordering. - Lexical structure — string literals and the interpolation tokens.
- collections —
Tree,MapandSet: the node pool and the indexes this pillar composes on. - i18n — locale-aware rendering, plural selection, and the only string ordering that exists.
- net — the codec catalogue, and schema-typed decoding at the boundary.