Strings
Text and bytes: the immutable String and ByteBuffer, their mutable twins, and the views and iterators over them.
Generated by
bin/build_library_doc.pyfrom the kernel sources. Do not edit by hand: change the generator, or the doc comments inkernel/src/, and re-run it.
Text
CLASS String
IMPLEMENTS Countable, Cloneable, Hashable, Equatable
Immutable readable/printable text.
Storage is char32[] (UTF-32 code points), one entry per code point.
The compiler lowers char32[] to ENVZN::_ValueArray<char32_t> so
every positional method is O(1) by code-point index. No UTF-8
decode lives in String — those conversions live in Converter and
Codec classes.
Storage rule: STRING IS IMMUTABLE. The .data field is populated
once in INIT and never mutated. Every transformation (substring,
toUpper, trim, concat, …) returns a fresh String.
Case fold: ASCII only. Code points >= 0x80 pass through unchanged
in toUpper / toLower / case-insensitive find / contains /
count. Full Unicode case folding waits for V2 Unicode tables.
kernel/src/String.ev:27
Constructors
INIT()
Empty String.
INIT(REFERENCE char32[] origin)
Take ownership of a pre-built code-point array. The literal-
lowering path in the emitter and every internal transformation
(substring, toUpper, concat, trim*, …) construct their result
this way: build a fresh char32[], hand it to a new String.
From a raw code-point array, COPYING it. REFERENCE is now explicit:
it always lowered to const EvArray<char32_t>&, so this states the
borrow the signature already had rather than changing it.
INIT(char32[] origin, int64 length)
THE LITERAL CONSTRUCTOR. The emitter's, not a developer's.
It could not be hidden. INTERNAL INIT is E1141 and PRIVATE INIT is a
parse error: a constructor carries no visibility modifier in Envzn,
because "INIT and CLEANUP are lifecycle, not module surface". So this is
public, and what keeps it honest is its own signature: the array
is a plain, owning parameter — a developer who calls it sees that it is taken, and
the analyzer poisons the source so a later read is E3010. That is the
language's normal contract rather than a hidden trapdoor.
It exists to spend ONE allocation on a string literal instead of two.
String s = "bob" lowers to a _ValueArray<char32_t>(U"bob", 3)
handed to a String constructor. Through INIT(REFERENCE char32[]) above
that array is built, then COPIED into .data and destroyed — two heap
allocations and two memcpys for every literal in the program, of which
there are ~2,800. Taking it as a plain (owning) parameter makes it a by-value
EvArray<char32_t>, so C++17's guaranteed copy elision constructs the
caller's prvalue DIRECTLY into it — nothing is materialised twice — and
:= then transfers the buffer pointer into .data. Measured 2 -> 1.
length IS REDUNDANT AND MUST NOT BE REMOVED. origin.length already
carries the count; this parameter exists solely to give this INIT a
different ARITY from INIT(REFERENCE char32[]). Owning and REFERENCE of
the SAME type at the SAME arity both lower to a one-argument
constructor over EvArray<char32_t> — by value and by const reference —
and clang rejects the pair as ambiguous. Differing arity is the
established way round it; StringIterator carries both forms for
exactly this reason. Delete this parameter and the kernel stops building.
It is String-only on purpose: a character-sequence literal has exactly
one representation in the language, and it is String. The other five
text/binary classes do not get it, which
run_text_binary_consistency.sh records as an allow-listed asymmetry
rather than drift.
INIT(REFERENCE String other)
Copy-construct from another String. ByteBuffer has had
INIT(REFERENCE ByteBuffer) all along; the text axis had no equivalent,
so CREATE String(other) was rejected outright until Phase 1 added this.
This does NOT collide with the array form above: the two differ in
parameter TYPE (String against char32[]), which C++ overload
resolution separates cleanly. What it cannot separate is the same type
owned versus REFERENCE — both lower to one-argument
constructors taking EvArray<char32_t> by value and by const reference,
and clang calls that ambiguous. That is why there is no owning companion.
Methods
METHOD clone() RETURNS String
Cloneable contract — deep copy. Returns a fresh String whose
.data is an independent code-point array with the same values.
METHOD length() RETURNS int64
Code-point count. O(1).
METHOD size() RETURNS int64
UTF-8 byte count of the text — distinct from length() (the code-point count). Storage is UTF-32 code points, so size() sums each code point's UTF-8 width per RFC 3629 (the same widths UTFCodec.encodeToUTF8 emits): U+0000–007F → 1 byte, U+0080–07FF → 2, U+0800–FFFF → 3, ≥ U+10000 → 4. Mirrors DynamicString.size().
METHOD isEmpty() RETURNS boolean
True iff the String has zero code points.
METHOD __op_index__(int64 i) RETURNS char32
Code point at index i. Declared as operator [] so the
only access syntax is s[i] — there is no public at()
method on String. Throws IndexOutOfBoundsError if i is out
of range — bounds check lives in _ValueArray<char32_t>::operator[].
METHOD equals(REFERENCE String other) RETURNS boolean
Code-point-by-code-point equality. UTF-32 storage means no
normalisation surprises; two Strings are equal iff their code-
point sequences match exactly.
Value equality, delegated to the value array's own equals. That is not
merely shorter: for an element type with a unique object representation
the array compares with a single length-checked memcmp (one SIMD
compare) where this method used to walk element by element. Measured
2026-08-27 at 6x faster on 64 elements and 11-14x from 1K up, with
identical results across an 8-cell differential matrix — including the
float guard, since the array deliberately does NOT memcmp an arithmetic
element type (-0.0 == +0.0, NaN != NaN would come out wrong).
METHOD isLessThan(REFERENCE String other) RETURNS boolean
Use the lexicographic comparision below
METHOD compareTo(REFERENCE String other) RETURNS int32
Lexicographic code-point comparison. Returns -1 / 0 / +1 in the usual three-way-comparison sense.
METHOD hash() RETURNS uint64
FNV-1a 64-bit hash over the code-point array. Satisfies the Hashable contract. Deterministic.
METHOD hash(uint64 seed) RETURNS uint64
FNV-1a 64-bit hash with caller-supplied seed. Each code point is mixed in as a 4-byte little-endian sequence: XOR then multiply per byte. Wrap on uint64 overflow is the algorithm.
METHOD substring(int64 start, int64 len) RETURNS String
Substring by code-point index. start is inclusive; length is
the count of code points. A bad range is a precondition violation
(programming error) — it THROWS, the same category as operator[].
Recoverable runtime outcomes (search misses) live on find/count
as the pipe-XOR (... | STATUS); a bad slice range is not one.
Also the lowering target of the s[start:length] slice sugar.
METHOD view(int64 start, int64 length) RETURNS StringView
Non-owning bounded view starting at start for length code
points. Suitable for searching / scanning without allocating a
copy. StringView's INIT validates the bounds; length == 0
yields an empty view; negative or out-of-range bounds throw.
METHOD find(REFERENCE String needle) RETURNS | int64
First occurrence of needle in this String (case-sensitive) — the
Searchable contract. Pipe-XOR: SUCCESS yields the absolute index
where the match begins; FAILURE means no match. To search a
sub-range, view(start, length) then search the window.
METHOD find(REFERENCE String needle, boolean caseMatch) RETURNS | int64
As find, with ASCII case-fold when caseMatch == FALSE. Forwards to
the one canonical char32 scan in StringView: a full-span view over this
String's codepoints searches for needle's codepoints.
METHOD contains(REFERENCE String needle) RETURNS boolean
True iff needle occurs at least once (case-sensitive).
METHOD contains(REFERENCE String needle, boolean caseMatch) RETURNS boolean
METHOD startsWith(REFERENCE String prefix) RETURNS boolean
True iff this String begins with prefix. Forwards to the one canonical
char32 scan in StringView, the same way find does — a full-span view
over this String's code points tests prefix's code points. An empty
prefix is a prefix of everything.
METHOD endsWith(REFERENCE String suffix) RETURNS boolean
True iff this String ends with suffix. Same delegation as startsWith.
METHOD lastIndexOf(REFERENCE String needle) RETURNS | int64
Index of the LAST occurrence of needle, pipe-XOR: SUCCESS populates
index, FAILURE when the needle does not occur. The forward search is
find; this is its mirror, and like find it forwards to StringView so
there is exactly one char32 scan in the kernel.
METHOD count(REFERENCE String needle) RETURNS int64
Non-overlapping occurrence count of needle (case-sensitive).
Empty needle returns 0.
METHOD count(REFERENCE String needle, boolean caseMatch) RETURNS int64
METHOD find_first_of(REFERENCE String set) RETURNS | int64
B1: index of the first code point that IS a member of set (the code
points of set, treated as a set). Pipe-XOR: SUCCESS = index, FAILURE =
none present. Forwards through a full-span view to the one canonical
char32 scan in StringView.
METHOD find_first_not_of(REFERENCE String set) RETURNS | int64
B1: index of the first code point that is NOT a member of set.
METHOD split(REFERENCE String delim) RETURNS String[]
Split on delim into an array of Strings. Adjacent delimiters
produce empty pieces; a trailing delimiter produces an empty
trailing piece. Empty receiver returns a single-element array
containing one empty String. Empty delim returns a single-
element array containing this String unchanged.
METHOD splitLines() RETURNS String[]
Split on any of the Unicode line-break code points: LF, CR, CRLF (one boundary), VT, FF, NEL (U+0085), LS (U+2028), PS (U+2029). The line terminators are NOT included in the pieces. A trailing terminator does NOT yield a trailing empty piece.
METHOD toUpper() RETURNS String
ASCII uppercase fold. Code points >= 0x80 unchanged.
METHOD toLower() RETURNS String
ASCII lowercase fold. Code points >= 0x80 unchanged.
METHOD trim() RETURNS String
TRIM (ASCII WHITESPACE: space, tab, LF, CR, VT, FF) -----------
METHOD trimStart() RETURNS String
METHOD trimEnd() RETURNS String
METHOD format(opaque args...) RETURNS String
Format this String as a template — substitute $1..$9 with
the corresponding argument and $$ with a literal $. The
everyday ("text $1")->format(a) form. stdio-spec.md §4.3 —
the String is the receiver of the formatting work; a Formatter
only supplies configuration, so format never instantiates one.
V1 scope: arguments must already be text — the compiler rewrite
at stdio-spec.md §4.3 redirects a .format(...) call whose args
include a non-String type to Formatter().format(...) so this
body only sees text args. A bare opaque holding a non-String
renders as ? (defensive fallback).
METHOD formatWith(BinaryMode mode, REFERENCE opaque[] args) RETURNS String
Template-substitution engine. mode is carried for the V2
binary-rendering cascade; V1 substitution is text-only and
does not consult it. Public so Formatter can delegate here.
Builds the result using only String methods (substring /
concat) so the kernel .hpp topology stays a clean DAG —
String never depends on another kernel class.
METHOD concat(REFERENCE String other) RETURNS String
String concatenation. Returns a fresh String; both receivers are unchanged.
METHOD iterator() RETURNS ValueIterator[char32]
ITERATOR — value-form (char32 is by-value). Wraps the bounded
walk in a sibling StringIterator class so the surface
matches the standard ValueIterator pipe-XOR shape (nextValue /
peekValue / skipValue). _ValueArray's raw iterator has a
different signature (no STATUS pipe-XOR), so the wrap is
necessary.
CLASS DynamicString
IMPLEMENTS Countable, Cloneable, Hashable, Equatable
Stack-resident mutable text builder.
DynamicString is one of four sibling classes in the Envzn text/binary type system:
- String — heap, immutable, text (sibling)
- DynamicString — heap, mutable, text (this file)
- ByteBuffer — heap, immutable, opaque binary (sibling)
- DynamicByteBuffer — heap, mutable, opaque binary (sibling)
PURPOSE The text builder. Mutate in place; snapshot to immutable String with toString() when done.
API SHAPE The non-modifying surface mirrors String exactly: length / isEmpty / at / equals / compare / hash / hash(seed) / find / contains / count / split / splitLines / concat / iterator / clone / format / formatWith / toString.
The modifying surface is DynamicString's reason for existing: MODIFY toUpper / toLower / trim / trimStart / trimEnd / substring (in-place slice) / append (String / DynamicString / char32) / prepend (String / DynamicString / char32) / replace / padStart / padEnd. Mutating methods return STATUS.
STRING ↔ DynamicString BRIDGE
Both classes now carry UTF-32 codepoint storage. Cross-class
operations (INIT(String), equals(String), find(String, …),
toString()) are direct codepoint copies — no UTF-8 encode/decode.
kernel/src/DynamicString.ev:52
Constructors
INIT()
INIT(int64 reserve)
Pre-size the builder to reserve code points, so a build of known
length performs ONE allocation instead of walking the growth schedule.
This was a NO-OP: it accepted the number and discarded it, on the
grounds that "pre-spilling to capacity N would need an storage reserve
method". reserveExact is that method — the substrate's resize already
takes INLINE and handles the spill, so the storage was never the
obstacle. Every caller that believed it was pre-sizing has been growing
from zero.
reserveExact, not reserve: this is a STATED size. reserve is the
incremental hint append uses and deliberately leaves 79% headroom.
INIT(REFERENCE String s)
Bridge ctor — start with the contents of an immutable String. Both storages are UTF-32 code points, so this is a direct copy with no encode step.
REFERENCE: the source is borrowed and survives construction. It used
to be taken by value, which handed the constructor a whole String it
only ever read from.
INIT(REFERENCE DynamicString other)
Copy-construct from another DynamicString. Every owning class in the text/binary family now has a same-type copy constructor: String, ByteBuffer, DynamicString and DynamicByteBuffer.
REFERENCE, so the source is borrowed and this instance allocates its
own storage. It does NOT collide with the other constructors here — they
differ in parameter TYPE, which C++ overload resolution separates. What
it cannot separate is one type owned against REFERENCE; that is
why there is no owning companion anywhere in this family.
INIT(REFERENCE char32[] origin)
Construct from a raw code-point array, COPYING it. REFERENCE, so the
source is borrowed and the builder allocates its own storage — the same
rule the other three owning classes in this family follow.
Methods
METHOD clone() RETURNS DynamicString
Deep copy. Returns a fresh DynamicString with the same codepoints. Implements the Cloneable contract.
METHOD isEmpty() RETURNS boolean
O(1) emptiness test.
METHOD length() RETURNS int64
Code-point count, O(1). Storage is one codepoint per slot.
METHOD size() RETURNS int64
UTF-8 byte count of the stored text — distinct from length() (the code-point count). Storage is UTF-32 code points (Phase 8.7), so size() sums each code point's UTF-8 width per RFC 3629 (the same widths UTFCodec.encodeToUTF8 emits): U+0000–007F → 1 byte, U+0080–07FF → 2, U+0800–FFFF → 3, ≥ U+10000 → 4.
METHOD __op_index__(int64 i) RETURNS char32
Code point at index i. Declared as operator [] so the
only access syntax is s[i] — there is no public at()
method on DynamicString. Throws IndexOutOfBoundsError if i is out
of range — bounds check lives in _ValueArray<char32_t>::operator[].
METHOD equals(REFERENCE DynamicString other) RETURNS boolean
Compare codepoint content against another DynamicString.
Value equality, delegated to the value array's own equals. That is not
merely shorter: for an element type with a unique object representation
the array compares with a single length-checked memcmp (one SIMD
compare) where this method used to walk element by element. Measured
2026-08-27 at 6x faster on 64 elements and 11-14x from 1K up, with
identical results across an 8-cell differential matrix — including the
float guard, since the array deliberately does NOT memcmp an arithmetic
element type (-0.0 == +0.0, NaN != NaN would come out wrong).
METHOD equals(REFERENCE String other) RETURNS boolean
Compare codepoint content against a String. Direct compare — both storages are UTF-32 codepoints.
METHOD isLessThan(REFERENCE DynamicString other) RETURNS boolean
Lexicographic codepoint comparison against another DynamicString.
METHOD compareTo(REFERENCE String other) RETURNS int32
Lexicographic codepoint comparison against a String. Returns -1/0/+1.
METHOD compareTo(REFERENCE char32[] otherData) RETURNS int32
Lexicographic codepoint comparison against a String. Returns -1/0/+1.
METHOD hash() RETURNS uint64
FNV-1a 64-bit hash over the UTF-32 codepoints. Matches String.hash() for identical text content (both fold over the same codepoint sequence).
METHOD hash(uint64 seed) RETURNS uint64
FNV-1a 64-bit hash with caller-supplied seed.
METHOD find(REFERENCE DynamicString needle) RETURNS | int64
First occurrence of needle in this DynamicString. Pipe-XOR:
SUCCESS yields the codepoint index of the match; FAILURE means no
match. This is the Searchable contract (case-sensitive). The scan is
the one canonical char32 scan in StringView — DynamicString builds a
full-span view over its own codepoints and forwards. A String-needle
overload is provided for convenience.
METHOD find(REFERENCE DynamicString needle, boolean caseMatch) RETURNS | int64
As find, with ASCII case-fold when caseMatch == FALSE. Views both
sides in place and runs the one canonical char32 scan in StringView.
This method used to snapshot BOTH sides to immutable Strings — two full
copies per call — on the belief that a char32[128+] field could not
bind to a REFERENCE char32[] parameter. E3 (2026-08-27) disproved it:
since SBO was removed both spellings are the same EvArray<char32_t>.
METHOD find(REFERENCE String needle) RETURNS | int64
Convenience: search for a String needle.
METHOD find(REFERENCE String needle, boolean caseMatch) RETURNS | int64
As find, with ASCII case-fold when caseMatch == FALSE. Views the
storage in place — no snapshot — and forwards to the canonical scan,
which owns the native fast path and its bounds re-derivation (gh #197).
METHOD contains(REFERENCE String needle) RETURNS boolean
True iff needle occurs at least once (case-sensitive).
METHOD contains(REFERENCE String needle, boolean caseMatch) RETURNS boolean
METHOD count(REFERENCE String needle) RETURNS int64
Non-overlapping occurrence count of needle (case-sensitive).
Empty needle returns 0.
METHOD count(REFERENCE String needle, boolean caseMatch) RETURNS int64
METHOD view(int64 start, int64 len) RETURNS StringView
Viewable — a non-owning, bounded window over this builder's code points
without copying. The mutable half of the family had no view() until
Phase 5, because a view into a GROWABLE buffer is only safe once the
analyzer refuses to mutate the source while the view is live.
That lock now exists: a view() on a named variable arms the same
per-source mutation lock an iterator() does (E5003), so appending to
or truncating this builder while a view over it is in scope is a
compile-time error rather than a stale window.
What the lock does NOT cover is the view outliving this object — that
is gh #248, it predates this method, and it already applies to
String.view().
METHOD startsWith(REFERENCE String prefix) RETURNS boolean
True iff this builder begins with prefix.
Builds a StringView directly over .data — no snapshot. E3 (2026-08-27)
established that a char32[128+] field binds to a REFERENCE char32[]
parameter: both lower to EvArray<char32_t>, and have since SBO was
removed. The view is a local, consumed before this method returns, and
nothing mutates the builder in between — so it needs none of the
mutation-lock machinery that a PUBLIC view() will require, and it
costs no copy. find above uses the same shape.
METHOD endsWith(REFERENCE String suffix) RETURNS boolean
True iff this builder ends with suffix.
METHOD lastIndexOf(REFERENCE String needle) RETURNS | int64
Index of the LAST occurrence of needle, pipe-XOR.
METHOD find_first_of(REFERENCE String set) RETURNS | int64
Index of the first code point that IS a member of set, pipe-XOR.
METHOD find_first_not_of(REFERENCE String set) RETURNS | int64
Index of the first code point that is NOT a member of set, pipe-XOR.
METHOD split(REFERENCE String delim) RETURNS String[]
Split on delim into an array of Strings. Empty delim
returns a single-element array containing a snapshot of this
DynamicString. Adjacent delimiters produce empty pieces; a
trailing delimiter produces an empty trailing piece. Receiver
is not mutated.
METHOD splitLines() RETURNS String[]
Split into lines on any of the Unicode line boundaries (same set as String.splitLines): LF, VT, FF (0x0A-0x0C), FS, GS, RS (0x1C-0x1E), CR and CR-LF (one boundary), NEL (U+0085), LS (U+2028), PS (U+2029). Terminators are NOT included; a trailing terminator does not yield a trailing empty piece. Receiver not mutated.
METHOD concat(REFERENCE String other) RETURNS String
Build a fresh String from this DynamicString's content followed
by other's. Non-modifying — receiver and other are both
unchanged. (To append in place, use the modifying append(String)
below.)
METHOD iterator() RETURNS ValueIterator[char32]
Snapshot to a String first, then return its iterator. The snapshot is an O(n) codepoint copy; subsequent iteration is O(1) per codepoint against the contiguous char32[] storage.
METHOD format(opaque args...) RETURNS String
Snapshot to a String first, then delegate to String.format — keeps the format-engine source-of-truth on String and avoids duplicating template-walking logic across the two classes.
METHOD formatWith(BinaryMode mode, REFERENCE opaque[] args) RETURNS String
METHOD toString() RETURNS String
BRIDGE OUT — explicit snapshot to immutable String ----------- Direct codepoint copy into a fresh String. No decode step.
MODIFY METHOD toUpper() RETURNS STATUS
─── MODIFYING METHODS ───────────────────────────────────────── Mutate receiver in place; return STATUS. ASCII uppercase fold, in place. Codepoints >= 0x80 pass through. Full Unicode case mapping is V2.
MODIFY METHOD toLower() RETURNS STATUS
ASCII lowercase fold, in place.
MODIFY METHOD toTitleCase() RETURNS STATUS
ASCII title-case fold, in place: the first character of each word is
uppercased and the remainder of the word lowercased. Word boundaries are
ASCII whitespace, -, and ' (see isTitleBoundary). Codepoints >= 0x80
pass through uncased but do not break a word. Full Unicode case mapping is
V2 — the same limit toUpper / toLower carry.
MODIFY METHOD trim() RETURNS STATUS
Strip ASCII whitespace at both ends, in place.
MODIFY METHOD trimStart() RETURNS STATUS
MODIFY METHOD trimEnd() RETURNS STATUS
METHOD substring(int64 start, int64 len) RETURNS String
A NEW String over [start, start+length). PURE — the receiver is
untouched. Matches String.substring, which throws on a bad range
rather than returning a STATUS, so this does too.
MODIFY METHOD append(REFERENCE String s) RETURNS STATUS
BUILDER MUTATIONS — append ------------------------------------
MODIFY METHOD append(DynamicString d) RETURNS STATUS
MODIFY METHOD append(char32 c) RETURNS STATUS
Append one Unicode code point. Tier-1 text check rejects a noncharacter / surrogate / control code point with FAILURE.
MODIFY METHOD setAt(int64 i, char32 c) RETURNS STATUS
Assign the code point at i. Pipe-XOR on the index, and gated by the
same Tier-1 text check append(char32) applies — a DynamicString may
not be mutated into holding a surrogate, noncharacter or control.
int64, not the uint64 DynamicByteBuffer.setAt still takes; Phase 2
normalises that split and a new method should not join the wrong side.
MODIFY METHOD clear() RETURNS STATUS
Empty the builder, keeping its capacity. The mutable text class had no way to be cleared at all, while its byte peer did.
MODIFY METHOD truncate(int64 n) RETURNS STATUS
Drop code points from the END until length() == n. FAILURE if n
exceeds the current length or is negative — this never grows.
MODIFY METHOD dropFirst(int64 n) RETURNS STATUS
Drop the first n code points. O(length - n) shift left.
MODIFY METHOD prepend(REFERENCE String s) RETURNS STATUS
BUILDER MUTATIONS — prepend -----------------------------------
MODIFY METHOD prepend(DynamicString d) RETURNS STATUS
MODIFY METHOD prepend(char32 c) RETURNS STATUS
MODIFY METHOD replace(REFERENCE String needle, REFERENCE String replacement) RETURNS STATUS
BUILDER MUTATIONS — replace -----------------------------------
Replace every occurrence of needle with replacement, in
place. Empty needle returns FAILURE (avoids infinite-substitution).
MODIFY METHOD padStart(int64 targetLength, char32 padChar) RETURNS STATUS
BUILDER MUTATIONS — padStart / padEnd -------------------------
Left-pad with padChar until the receiver is targetLength
CODE POINTS long. Already-longer receivers are left alone.
MODIFY METHOD padEnd(int64 targetLength, char32 padChar) RETURNS STATUS
Right-pad with padChar until the receiver is targetLength
CODE POINTS long.
CLASS StringView
IMPLEMENTS View, Viewable
Non-owning, bounded slice over a code-point array.
A StringView is a short-lived locals-only descriptor. It carries a
REFERENCE char32[] source plus start + length (code-point offsets
into the source array), and exposes a narrow read-only surface — find /
contains / count, length / operator[] / iterator(), plus view() (narrow
to a sub-window — composition) and copy() (materialise the window into a
fresh char32[]; wrap with CREATE String(view->copy()) for an owned
String). It IMPLEMENTS View and Viewable.
STRING-FREE BY DESIGN. StringView references only char32[] and itself —
never String. This breaks what would otherwise be a String↔StringView
dependency cycle, so String can implement Viewable with a covariant
view() returning a (complete, emitted-first) StringView. Consequences:
the search needle is a StringView (wrap a needle String with
needleStr->view(0, needleStr->length())), copy() yields char32[],
and bounds violations PANIC IndexOutOfBoundsError(...) directly (bodies
emit out-of-line).
The REFERENCE field is safe by the same reference-checker rules the rest
of the kernel relies on: the source's lifetime must enclose the view's;
the source must not be re-bound underneath a live view; the view is
non-storable into another class. An empty view (length == 0) is permitted
— find reports not-found and iterator() is immediately exhausted — so
String can route search / iteration uniformly through a full-span view.
kernel/src/StringView.ev:45
Constructors
INIT(REFERENCE char32[] source, int64 start, int64 length)
Construct a bounded view over [start, start+length) in source.
Throws (via the C-string helper) if start < 0 or length < 0 or
start + length > source.length(). A zero length is allowed.
Methods
METHOD length() RETURNS int64
Code-point count of this view.
METHOD size() RETURNS int64
Countable — View EXTENDS Countable, so the view owes size() and
isEmpty() as well as its own length(). All three are the same
number here: the count of code points the window spans.
METHOD isEmpty() RETURNS boolean
METHOD __op_index__(int64 i) RETURNS char32
Code point at index i within this view. sv[i] is the only read
syntax. Throws (C-string helper) if i is out of [0, view.length).
METHOD find(REFERENCE char32[] needle) RETURNS | int64
First occurrence of needle (itself a view) within this view's whole
code-point range. Pipe-XOR: SUCCESS yields the view-relative index of
the match; FAILURE means no match. Case-sensitive — to search a
sub-range, narrow with view(start, length) first.
METHOD find(REFERENCE char32[] needle, boolean caseMatch) RETURNS | int64
As find, with ASCII case-fold when caseMatch == FALSE.
METHOD startsWith(REFERENCE char32[] needle) RETURNS boolean
True iff this view begins with needle (case-sensitive). An empty
needle returns TRUE; a needle longer than the view returns FALSE.
METHOD endsWith(REFERENCE char32[] needle) RETURNS boolean
True iff this view ends with needle (case-sensitive). An empty
needle returns TRUE; a needle longer than the view returns FALSE.
METHOD lastIndexOf(REFERENCE char32[] needle) RETURNS | int64
Last occurrence of needle within this view (case-sensitive). Pipe-XOR:
SUCCESS yields the view-relative index of the final match; FAILURE means
no match. An empty needle yields the view length. Scans backward from the
last candidate window, so the first hit found is the last occurrence.
METHOD contains(REFERENCE char32[] needle) RETURNS boolean
True iff needle occurs at least once in this view (case-sensitive).
METHOD contains(REFERENCE char32[] needle, boolean caseMatch) RETURNS boolean
METHOD count(REFERENCE char32[] needle) RETURNS int64
Non-overlapping occurrence count of needle in this view
(case-sensitive). Empty needle returns 0.
METHOD count(REFERENCE char32[] needle, boolean caseMatch) RETURNS int64
METHOD find_first_of(REFERENCE char32[] set) RETURNS | int64
B1: index of the first code point that IS a member of set. Pipe-XOR:
SUCCESS = the index; FAILURE = no code point in this view is in set.
METHOD find_first_not_of(REFERENCE char32[] set) RETURNS | int64
B1: index of the first code point that is NOT a member of set.
METHOD view(int64 start, int64 length) RETURNS StringView
Narrow to a sub-window [start, start+length) relative to THIS view
(composition — a sub-window of a window is a window). Bounds are
checked against this view's length, so a sub-view can never escape
its parent. Satisfies Viewable.
METHOD copy() RETURNS char32[]
Materialise this window's code points into a fresh char32[] (the
"keep this slice" escape hatch — the view itself never escapes).
Wrap with CREATE String(view->copy()) for an owned String.
DEFECT FIXED 2026-08-27. This used .source->copyTo(piece, length),
which copies from the SOURCE'S index 0 and ignores .start — so any
view whose window did not begin at 0 returned the wrong code points.
_ValueArray::copyTo(other, n) takes no source offset, so it cannot
express a windowed copy at all; the element loop below (which was here,
commented out) is the only correct form. ByteBufferView.copy() never had
the bug — it always used this shape.
It stayed latent because every caller passed a FULL-SPAN view
(view(0, length())), where .start is 0 and the two agree. The first
caller with a non-zero start was DynamicString.substring, added in
Phase 2, which is what surfaced it.
METHOD iterator() RETURNS ValueIterator[char32]
Walk the view's code points in order. The iterator snapshots the
bounded slice into its own storage at construction, so it is
independent of this view and its source char32[] for its entire
lifetime (matching ByteBufferIterator's design). The snapshot is
materialised inline because the emit does not currently auto-deref a
REFERENCE self-field when passed to a value-typed ctor param.
Bytes
CLASS ByteBuffer
IMPLEMENTS Countable, Cloneable, Hashable, Equatable
Immutable opaque binary data container.
ByteBuffer is one of four sibling classes in the Envzn text/binary type system:
- String — heap, immutable, text (sibling)
- DynamicString — heap, mutable, text (sibling)
- ByteBuffer — heap, immutable, opaque binary (this file)
- DynamicByteBuffer — heap, mutable, binary (sibling)
PURPOSE A handle to a chunk of opaque binary data — bytes whose meaning the language does not interpret. Once constructed the contents never change; transformations like subBuffer or concat return a new ByteBuffer. Display in human-readable form goes through HexCodec.toHex(...) (default separator " ", e.g. "3F 40 6A").
STORAGE
A single INTERNAL binary[] field. At the C++ layer this lowers
to _ValueArray<uint8_t> — a contiguous, growable byte buffer
governed by the value-array storage shorthand of
ENVZN_CONSTITUTION.md §10.1. All positional accessors read via
.data[i]; appends are high-water-mark assignments at
.data[.data.length].
INDEXING SEMANTICS — BYTES, NOT CODE POINTS
Every positional method indexes in bytes. length() and size()
both exist and return the same number, because a byte IS the
storage unit; they differ only on the text axis. buf[i] returns a
binary primitive, not a char. 'binary' is an opaque char, aka
'unsigned char'.
ERROR-RETURN POLICY
One policy, shared with the text axis: [] throws
IndexOutOfBoundsError, the find family is pipe-XOR
(int64 index | STATUS) with no -1 sentinel, and everything else
is total.
This header used to say "No method throws", and that sentence was
the origin of most of the drift between the two axes: it is what
produced a second element accessor at(), a pipe-XOR subBuffer
against a throwing substring, and the -1 sentinel. It was also
already false about this class's own [], nine lines below where
it was written — [] has always delegated to
_ValueArray<uint8_t>::operator[], which throws. Phase 3 kept the
behaviour the code had and retired the sentence.
kernel/src/ByteBuffer.ev:62
Constructors
INIT()
CONSTRUCTORS / CLEANUP
INIT(REFERENCE binary[] origin)
Construct from a pre-built byte array, COPYING it. Every internal
transformation (subBuffer, concat, clone) constructs its result this
way: build a fresh binary[], hand it to a new ByteBuffer.
REFERENCE, not owning: the class allocates its own storage rather
than adopting the caller's. This was the last owning constructor in the
text/binary family and the only one that adopted caller memory — every
other kernel class allocates its own, so the outlier moved rather than
the pattern. It costs one copy per construction, measured at 8-21%
(E2, 2026-08-27), and buys a uniform rule: handing an array to a
constructor never takes it away from you.
It also cannot coexist with an owning companion. Both lower to a
one-argument constructor over the same EvArray<uint8_t> — one by
value, one by const reference — which clang rejects as ambiguous.
INIT(binary[] origin, int64 length)
Similar to String, we have an owning binary[] for performance
length IS REDUNDANT AND MUST NOT BE REMOVED.
INIT(REFERENCE ByteBuffer other)
Copy-construct from another ByteBuffer. One allocation, one bulk copy — the element loop this replaced wrote at the high-water mark, so the storage reallocated as it filled.
Methods
METHOD clone() RETURNS ByteBuffer
CLONE — hand-written. Hands the storage straight to the copying constructor, which reserves once and bulk-copies.
It used to build an intermediate binary[] byte by byte and hand THAT
over, which cost two full copies and an extra allocation: once into the
intermediate, once out of it. That shape made sense when the constructor
took the array by ownership and adopted the array; under REFERENCE the intermediate is
pure waste, because the constructor was going to copy anyway.
METHOD isEmpty() RETURNS boolean
INSPECTION
METHOD size() RETURNS int64
METHOD length() RETURNS int64
Byte count — the same number size() returns. On the byte axis the two
coincide, because a byte IS the storage unit; on the text axis they
differ (String.length() counts code points, String.size() counts
UTF-8 bytes). Both names exist on all six text/binary classes so a
caller never has to remember which axis they are holding.
METHOD __op_index__(int64 i) RETURNS binary
Byte at index i — buf[i] yields the raw binary. THE element
accessor: a pipe-XOR at() stood beside it until Phase 3, which left
the byte axis with two ways to read one byte and the text axis with one.
The bounds check lives in _ValueArray<uint8_t>::operator[], which
throws IndexOutOfBoundsError out of range, mirroring String's code-point
operator []. Delegating rather than range-checking here is deliberate:
the array's throw carries the offending index and the bound in its
message, and this class sits upstream of Formatter, so a PANIC written
here could only say "out of bounds" with no numbers in it.
METHOD equals(REFERENCE ByteBuffer other) RETURNS boolean
COMPARISON
Value equality, delegated to the value array's own equals. That is not
merely shorter: for an element type with a unique object representation
the array compares with a single length-checked memcmp (one SIMD
compare) where this method used to walk element by element. Measured
2026-08-27 at 6x faster on 64 elements and 11-14x from 1K up, with
identical results across an 8-cell differential matrix — including the
float guard, since the array deliberately does NOT memcmp an arithmetic
element type (-0.0 == +0.0, NaN != NaN would come out wrong).
METHOD isLessThan(REFERENCE ByteBuffer other) RETURNS boolean
Use the lexicographic comparision below
METHOD compareTo(REFERENCE ByteBuffer other) RETURNS int32
Lexicographic byte comparison: shorter buffer orders before a longer one when prefixes match. Returns -1 / 0 / 1.
METHOD find(REFERENCE ByteBuffer needle) RETURNS | int64
Searchable — first index at which needle's bytes occur in this
buffer, else FAILURE. Forwards to the one canonical binary scan in
ByteBufferView: a full-span view over this buffer searches for
needle's bytes.
METHOD lastIndexOf(REFERENCE ByteBuffer needle) RETURNS | int64
Index of the LAST occurrence of needle, pipe-XOR. The mirror of
find, forwarded through a full-span view the same way, so there is
exactly one byte scan in the kernel.
METHOD find_first_of(REFERENCE ByteBuffer set) RETURNS | int64
B1: index of the first byte that IS a member of set (the bytes of
set, treated as a set). Forwards through a full-span view.
METHOD find_first_not_of(REFERENCE ByteBuffer set) RETURNS | int64
B1: index of the first byte that is NOT a member of set.
METHOD view(int64 start, int64 length) RETURNS ByteBufferView
Viewable — a non-owning, bounded window over this buffer's bytes
without copying. The byte-side analogue of String.view.
METHOD hash() RETURNS uint64
HASHING — same FNV-1a as String, deterministic across runs FNV-1a 64-bit hash. Matches String.hash() and DynamicString.hash() byte-for-byte for identical content, so the three classes hash to the same uint64 when carrying the same bytes.
METHOD hash(uint64 seed) RETURNS uint64
Seeded FNV-1a, the peer of String.hash(uint64 seed). Chaining a seed
lets a composite key hash its parts in sequence without materialising a
concatenation.
METHOD contains(REFERENCE ByteBuffer needle) RETURNS boolean
CONTENT LOOKUP Empty needle returns true (matches String.contains).
METHOD startsWith(REFERENCE ByteBuffer prefix) RETURNS boolean
METHOD endsWith(REFERENCE ByteBuffer suffix) RETURNS boolean
METHOD find(REFERENCE ByteBuffer needle, int64 fromIndex) RETURNS | int64
First-occurrence index at or after fromIndex. Pipe-XOR:
FAILURE on out-of-range fromIndex or when needle does not
occur from that point onward.
METHOD count(REFERENCE ByteBuffer needle) RETURNS int64
Non-overlapping occurrence count. Empty needle returns 0 (avoids infinite-match scenario).
METHOD subBuffer(int64 start, int64 length) RETURNS ByteBuffer
SLICING / CONCATENATION
Copy a contiguous slice into a fresh ByteBuffer. start is inclusive,
length is a byte count. TOTAL, and the exact peer of
String.substring: a bad range is a precondition violation — a
programming error — so it THROWS, the same category as [].
length == 0 yields an empty buffer and is not an error.
This returned pipe-XOR (ByteBuffer | STATUS) until Phase 3, which made
a caller slicing a buffer write an IF where the same caller slicing a
String writes none. A recoverable outcome is a search that misses; a
caller asking for bytes that were never there is a bug in the caller.
METHOD concat(REFERENCE ByteBuffer other) RETURNS ByteBuffer
Concatenation. Returns a fresh ByteBuffer; receivers unchanged. operator + pending — see Bug #57.
METHOD iterator() RETURNS ByteBufferIterator
Yield a value-iterator over the bytes. The iterator copies the bytes into its own storage at construction — it is fully independent of this buffer for its whole lifetime.
METHOD data() RETURNS REFERENCE binary[]
Non-owning REFERENCE to the underlying binary[] storage. The
receiver retains ownership; the returned reference is valid
for the receiver's lifetime per §11.6 lifetime elision. The
canonical use is at a FOREIGN BIND call site that needs to
expose the byte storage to a C function without copying.
CLASS DynamicByteBuffer
IMPLEMENTS Countable, Cloneable, Hashable, Equatable
Stack-resident mutable binary builder.
DynamicByteBuffer is one of four sibling classes in the Envzn text/binary type system:
- String — heap, immutable, text (sibling)
- DynamicString — heap, mutable, text (sibling)
- ByteBuffer — heap, immutable, opaque binary (sibling)
- DynamicByteBuffer — heap, mutable, binary (this file)
PURPOSE The binary builder. Mutate in place; snapshot to immutable ByteBuffer with toByteBuffer() when done. Mutating methods return STATUS only — no chaining.
ACCESS IDIOMS
Subscript read .data[i] returns binary. Subscript write at HWM
(.data[.data.length] = v) extends; in-range write replaces.
.data.length returns int32 (current logical length). .data.clear()
resets length to 0.
kernel/src/DynamicByteBuffer.ev:39
Constructors
INIT()
CONSTRUCTORS / CLEANUP
INIT(int64 reserve)
Pre-size the builder to reserve bytes — one allocation instead of the
growth schedule. that is
reserveExact, and the substrate's resize already handles the INLINE
spill. reserveExact and not reserve because this is a STATED size,
not the incremental hint append uses.
INIT(REFERENCE binary[] origin)
Construct from a raw byte array, COPYING it. DynamicString has had a raw-array constructor all along; the byte builder had none, so the only way in was append-in-a-loop.
REFERENCE, not owning: the class allocates its own storage rather than
adopting the caller's. That is what every kernel class does — the last
owning constructor in this family, ByteBuffer's, was retired in Phase 2.
INIT(REFERENCE DynamicByteBuffer other)
Copy-construct from another DynamicByteBuffer. Every owning class in the text/binary family now has a same-type copy constructor: String, ByteBuffer, DynamicString and DynamicByteBuffer.
REFERENCE, so the source is borrowed and this instance allocates its
own storage. It does NOT collide with the other constructors here — they
differ in parameter TYPE, which C++ overload resolution separates. What
it cannot separate is one type owned against REFERENCE; that is
why there is no owning companion anywhere in this family.
INIT(REFERENCE ByteBuffer other)
Methods
METHOD clone() RETURNS DynamicByteBuffer
CLONE — hand-written. Constructs a fresh DynamicByteBuffer and re-appends every byte through the public append() method, so we stay on the same write path as ordinary growth (no cross- instance private-field access). Mirrors DynamicString.clone().
METHOD isEmpty() RETURNS boolean
INSPECTION
METHOD size() RETURNS int64
METHOD length() RETURNS int64
Byte count — the same number size() returns. On the byte axis the two
coincide, because a byte IS the storage unit; on the text axis they
differ (String.length() counts code points, String.size() counts
UTF-8 bytes). Both names exist on all six text/binary classes so a
caller never has to remember which axis they are holding.
METHOD data() RETURNS REFERENCE binary[]
Non-owning REFERENCE to the underlying binary[512+]
storage. Mirror of ByteBuffer.data() (§11.6 lifetime
elision applies); canonical use is at a FOREIGN BIND call
site that needs to expose the byte storage to a C function
without copying.
METHOD iterator() RETURNS ByteBufferIterator
Snapshot iterator over the current bytes. Returns the shared ByteBufferIterator (the value-form ValueIterator OF binary used by ByteBuffer too): it copies the bytes in at construction, so the walk stays stable even if this mutable buffer is appended to afterward.
METHOD equals(REFERENCE DynamicByteBuffer other) RETURNS boolean
COMPARISON
Value equality, delegated to the value array's own equals. That is not
merely shorter: for an element type with a unique object representation
the array compares with a single length-checked memcmp (one SIMD
compare) where this method used to walk element by element. Measured
2026-08-27 at 6x faster on 64 elements and 11-14x from 1K up, with
identical results across an 8-cell differential matrix — including the
float guard, since the array deliberately does NOT memcmp an arithmetic
element type (-0.0 == +0.0, NaN != NaN would come out wrong).
METHOD isLessThan(REFERENCE DynamicByteBuffer other) RETURNS boolean
Use the lexicographic comparison below.
METHOD compareTo(REFERENCE DynamicByteBuffer other) RETURNS int32
Lexicographic byte comparison. Shorter buffer orders before longer when prefixes match. Returns -1 / 0 / 1.
METHOD find(REFERENCE DynamicByteBuffer needle) RETURNS | int64
Searchable — first index at which needle's bytes occur in this
buffer, else FAILURE. Snapshots both sides to immutable ByteBuffers
and forwards to ByteBuffer.find — the one canonical byte scan.
METHOD hash() RETURNS uint64
HASHING FNV-1a 64-bit hash. Matches String.hash() / DynamicString.hash() / ByteBuffer.hash() byte-for-byte for identical content.
METHOD hash(uint64 seed) RETURNS uint64
METHOD contains(REFERENCE ByteBuffer needle) RETURNS boolean
CONTENT LOOKUP
METHOD startsWith(REFERENCE ByteBuffer prefix) RETURNS boolean
METHOD endsWith(REFERENCE ByteBuffer suffix) RETURNS boolean
METHOD find(REFERENCE ByteBuffer needle, int64 fromIndex) RETURNS | int64
First-occurrence index at or after fromIndex. Pipe-XOR:
FAILURE on out-of-range fromIndex or when needle does not
occur from that point onward.
METHOD count(REFERENCE ByteBuffer needle) RETURNS int64
Non-overlapping occurrence count. Empty needle returns 0 (avoids infinite-match scenario).
MODIFY METHOD append(binary b) RETURNS STATUS
MUTATIONS — return STATUS, mutate receiver in place
MODIFY METHOD append(REFERENCE ByteBuffer bb) RETURNS STATUS
MODIFY METHOD append(DynamicByteBuffer dbb) RETURNS STATUS
MODIFY METHOD prepend(binary b) RETURNS STATUS
MODIFY METHOD prepend(REFERENCE ByteBuffer bb) RETURNS STATUS
MODIFY METHOD setAt(int64 i, binary b) RETURNS STATUS
MODIFY METHOD clear() RETURNS STATUS
MODIFY METHOD truncate(int64 n) RETURNS STATUS
Drop bytes from the END until size() == n. FAILURE if n exceeds current size or n < 0 (no-grow contract).
MODIFY METHOD dropFirst(int64 n) RETURNS STATUS
Drop the first n bytes. O(size - n) shift left.
METHOD subBuffer(int64 start, int64 length) RETURNS ByteBuffer
A NEW ByteBuffer over [start, start+length). PURE — the receiver is
untouched. Same name, same meaning and same shape as
ByteBuffer.subBuffer, which is what the vertical rule requires — TOTAL
since Phase 3, throwing on a bad range like String.substring does.
METHOD __op_index__(int64 i) RETURNS binary
Byte at i. Throws IndexOutOfBoundsError out of range — the bounds
check lives in the value array's own subscript, exactly as ByteBuffer's
and String's do. The mutable class had no index operator at all, so
dbb[i] did not compile while bb[i] did.
METHOD view(int64 start, int64 length) RETURNS ByteBufferView
Viewable — a non-owning, bounded window over this builder's bytes
without copying. The mutable half of the family had no view() until
Phase 5, because a view into a GROWABLE buffer is only safe once the
analyzer refuses to mutate the source while the view is live.
That lock now exists: a view() on a named variable arms the same
per-source mutation lock an iterator() does (E5003), so appending to
or truncating this builder while a view over it is in scope is a
compile-time error rather than a stale window.
What the lock does NOT cover is the view outliving this object — that
is gh #248, it predates this method, and it already applies to
ByteBuffer.view().
METHOD lastIndexOf(REFERENCE ByteBuffer needle) RETURNS | int64
Index of the LAST occurrence of needle, pipe-XOR. Scans in place over
the backing rather than snapshotting to a ByteBuffer first — the same
choice find already makes here, and the one DynamicString.find still
does not. Once view() reaches the mutable classes this can forward to
ByteBufferView.lastIndexOf like the immutable peer does.
METHOD find_first_of(REFERENCE ByteBuffer set) RETURNS | int64
Index of the first byte that IS a member of set, pipe-XOR.
METHOD find_first_not_of(REFERENCE ByteBuffer set) RETURNS | int64
Index of the first byte that is NOT a member of set, pipe-XOR.
METHOD concat(REFERENCE ByteBuffer other) RETURNS ByteBuffer
A NEW ByteBuffer holding this buffer's bytes followed by other's.
Pure: this buffer is unchanged. The mutating sibling is append.
MODIFY METHOD padStart(int64 targetLength, binary padByte) RETURNS STATUS
Pad the FRONT with padByte until the buffer is targetLength long.
Already at or above the target is SUCCESS and a no-op. int64 rather
than the uint64 DynamicString still uses — the sign split across these
two classes is normalised in Phase 2, and a new method should not be
born on the wrong side of it.
MODIFY METHOD padEnd(int64 targetLength, binary padByte) RETURNS STATUS
Pad the END with padByte until the buffer is targetLength long.
MODIFY METHOD replace(REFERENCE ByteBuffer needle, REFERENCE ByteBuffer replacement) RETURNS STATUS
Replace every occurrence of needle with replacement. Empty
needle returns FAILURE (avoids infinite-substitution).
METHOD toByteBuffer() RETURNS ByteBuffer
BRIDGE OUT — explicit snapshot to immutable ByteBuffer
CLASS ByteBufferView
IMPLEMENTS View, Viewable
A non-owning, bounded window over a ByteBuffer's binary[] storage.
The byte-side mirror of StringView: it holds the one canonical binary
scan (find / contains / count) that ByteBuffer — and, via a ByteBuffer
snapshot, DynamicByteBuffer — forward into. Like StringView it is
ByteBuffer-free: it knows only binary[], so it introduces no
dependency cycle with the owning buffer types.
Bytes carry no case, so there is no case-fold variant — the scans are exact-byte only (the one structural difference from StringView).
kernel/src/ByteBufferView.ev:25
Constructors
INIT(REFERENCE binary[] source, int64 start, int64 length)
Construct a bounded view over [start, start+length) in source.
Throws (via the C-string helper) if start < 0 or length < 0 or
start + length > source.length(). A zero length is allowed.
Methods
METHOD length() RETURNS int64
Byte count of this view.
METHOD size() RETURNS int64
Countable — View EXTENDS Countable, so the view owes size() and
isEmpty() as well as its own length(). All three are the same
number here: the count of octets the window spans.
METHOD isEmpty() RETURNS boolean
METHOD __op_index__(int64 i) RETURNS binary
Byte at index i within this view. bv[i] is the only read syntax.
Throws (C-string helper) if i is out of [0, view.length).
METHOD find(REFERENCE binary[] needle) RETURNS | int64
First occurrence of needle's bytes within this view's whole range.
Pipe-XOR: SUCCESS yields the view-relative index of the match;
FAILURE means no match. To search a sub-range, narrow with
view(start, length) first.
METHOD startsWith(REFERENCE binary[] needle) RETURNS boolean
True iff this view begins with needle. An empty needle returns TRUE;
a needle longer than the view returns FALSE.
METHOD endsWith(REFERENCE binary[] needle) RETURNS boolean
True iff this view ends with needle. An empty needle returns TRUE;
a needle longer than the view returns FALSE.
METHOD lastIndexOf(REFERENCE binary[] needle) RETURNS | int64
Last occurrence of needle within this view. Pipe-XOR: SUCCESS yields
the view-relative index of the final match; FAILURE means no match. An
empty needle yields the view length. Scans backward from the last
candidate window, so the first hit found is the last occurrence.
METHOD contains(REFERENCE binary[] needle) RETURNS boolean
True iff needle occurs at least once in this view.
METHOD count(REFERENCE binary[] needle) RETURNS int64
Non-overlapping occurrence count of needle. Empty needle returns 0.
METHOD find_first_of(REFERENCE binary[] set) RETURNS | int64
B1: index of the first byte that IS a member of set. Pipe-XOR:
SUCCESS = the index; FAILURE = no byte in this view is in set.
METHOD find_first_not_of(REFERENCE binary[] set) RETURNS | int64
B1: index of the first byte that is NOT a member of set.
METHOD view(int64 start, int64 length) RETURNS ByteBufferView
Narrow to a sub-window [start, start+length) relative to THIS view
(composition — a sub-window of a window is a window). Bounds are
checked against this view's length, so a sub-view can never escape
its parent. Satisfies Viewable.
METHOD copy() RETURNS binary[]
Materialise this window's bytes into a fresh binary[] (the "keep this
slice" escape hatch — the view itself never escapes). Wrap with
CREATE ByteBuffer(view->copy()) for an owned ByteBuffer.
METHOD iterator() RETURNS ValueIterator[binary]
Walk this window's bytes in order. Mirrors StringView.iterator(),
which the byte axis had no counterpart to until now.
The window is materialised and GIVEN to the iterator. It cannot be
lent: .source is a REFERENCE field — a pointer — and does not bind to
a REFERENCE binary[] parameter, so the only array this view can offer
is a local snapshot, and lending a local returns an iterator over freed
storage (gh #248). The owning constructor takes the array, so the
referent lives exactly as long as the iterator does.
CLASS ByteBufferIterator
IMPLEMENTS ValueIterator
Value-form iterator over a ByteBuffer.
ByteBufferIterator is the value-iterator counterpart to the collection iterators (ArrayIterator and friends), and mirrors StringIterator on the text axis.
It has TWO constructors, and they differ in what they do with the
bytes. The BORROWING form takes a window [start, start+length) of
an array the caller keeps owning, and aliases it — ByteBuffer and
DynamicByteBuffer pass their own field, which outlives the
iterator. The OWNING form takes the array outright and lends itself
the borrow, which is what a view must use: a view's .source is a
REFERENCE field, so the only array it can offer is a local snapshot,
and lending a local is the gh #248 dangle.
The header used to say the bytes were "copied in — the iterator never aliases", which the single borrowing constructor had never done. Now one form aliases, one owns, and each says which.
IMPLEMENTS ValueIterator OF binary — a ByteBuffer's storage yields bytes by value, not by reference, so the iterator advances with nextValue() / peekValue() / skipValue() rather than the by-reference forms ReferenceIterator (the collection iterators) carries.
STORAGE
Five fields: snapshot storage used only by the owning constructor,
the borrowed payload, the window offset, a cached length, and a
cursor. Every read indexes .data[.start + .cursor], so a window
that does not begin at zero walks the right bytes. The compiler
emits this CLASS as a C++ struct (the value-class carve-out), so
allocation and teardown are deterministic and RAII.
kernel/src/ByteBufferIterator.ev:52
Constructors
INIT(REFERENCE binary[] source, int64 start, int64 length)
BORROWING form — built by ByteBuffer.iterator() and
DynamicByteBuffer.iterator(), which pass their own binary[] FIELD
plus the window [start, start+length). The iterator ALIASES that
storage, so the argument must outlive the iterator. Both owners pass the
whole array (0, .data.length).
INIT(binary[] snapshot)
OWNING form — the iterator takes the array and lends itself the borrow.
.data =@ .owned is SELF-rooted, so the referent lives exactly as long
as the iterator does and the gh #248 dangle cannot arise. This is the
form ByteBufferView.iterator() uses.
It is separable from the borrowing form ONLY because their arities
differ. Two one-argument constructors over the same EvArray<uint8_t>,
one owning and one REFERENCE, lower to by-value and const-reference
and clang rejects the pair as ambiguous — which is what blocked
ByteBufferView.iterator() through Phase 1.
Methods
METHOD hasNext() RETURNS boolean
MODIFY METHOD nextValue() RETURNS | binary
Value-form advance — copy out the byte at the cursor, then step the cursor forward. FAILURE at end-of-iteration.
METHOD peekValue() RETURNS | binary
The byte at the cursor without advancing. FAILURE at end.
MODIFY METHOD skipValue(uint64 n) RETURNS | binary
Move the cursor forward by n and return the byte there.
FAILURE if that would pass the end. skipValue(0) == peekValue().
MODIFY METHOD nextBytes(int64 numberOfBytes) RETURNS | binary[]
Block read: copy up to numberOfBytes bytes from the cursor into
a fresh binary[] and advance the cursor past them. Returns fewer
than requested when the buffer is exhausted first (the caller
checks .length); FAILURE only when the cursor is already at the
end, or numberOfBytes is not positive. The block-oriented
counterpart to byte-at-a-time nextValue(), mirroring the
(binary[] | STATUS) shape of File.read().