# Mach 6.10.1: complete reference > Mach is a statically typed, compiled, self-hosted systems language with no hidden control flow, no hidden allocation, and no type inference. This file is served by https://machlang.org, which rebuilds it from each new Mach release. It covers the current release, Mach 6.10.1: the guide at https://machlang.org/docs/, then the language reference in https://github.com/briar-systems/mach/tree/v6.10.1/doc/language. The canonical repository is https://github.com/briar-systems/mach. Where the two parts differ, the language reference is authoritative. The short index is https://machlang.org/llms.txt. ## Rules models most often get wrong Mach is not C, Rust, Zig, or Go. Code written from those habits does not compile. Hold to these rules: - **No type inference.** Every binding states its type: `val n: i64 = 42;` and `var i: i64 = 0;`. There is no `let`, no `:=`, and no `auto`. `val` is immutable and `var` is mutable. - **`or`, not `else`.** A conditional chain is `if (a) { ... } or (b) { ... } or { ... }`. There is no `else` and no `else if`, and bodies always take braces. - **`for` is the only loop.** `for (cond) { ... }` loops while the condition holds, and a bare `for { ... }` loops until a `brk` or `ret` leaves it. There is no `while`, no `loop`, and no for-each or range loop. - **`ret`, `brk`, and `cnt`.** Return is `ret expr;` or `ret;`. Loop exit is `brk;` and next iteration is `cnt;`. There is no `return`, `break`, or `continue`. - **`?x` and `@p` for pointers.** `?x` takes the address of a place and `@p` dereferences a pointer, for reads and writes alike: `var p: *i64 = ?x; @p = 11;`. There is no `&x` or `*p` expression. `*T` appears only in types. - **No methods.** Records hold fields only. There is no `impl`, no `self`, and no member function. Write a free function that takes the record, or a pointer to it. In `print.println(...)`, `print` is a module alias, not a receiver. - **`sel` and guards for tagged unions.** `sel p.case` tests which case a `tag` holds. A payload `p.case` may be read only inside a guard: an `if` or `or` arm whose condition is exactly `sel p.case`, or the rest of a block after a chain whose every arm exits. There is no `match`, no `switch`, and no `==` on tags. - **Explicit generic instantiation.** Type arguments are always written in brackets at the use: `identity[i64](42)` and `Pair[i64, u8]`. They are never inferred from the arguments. - **Executables need `use std.runtime;` and `#[symbol("main")]`.** The runtime supplies `_start`, and the entry point is whichever function exports the `main` symbol, with the signature `fun main(argc: i64, argv: **u8) i64`: ```mach use std.runtime; use print: std.print; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.println("Hello, World!"); ret 0; } ``` ## Guide Source: https://machlang.org/docs/index.html ### mach documentation A small, explicit systems language. These pages are a pamphlet, not a bible: short, example-first, and focused on what the language can do. - [Install](https://machlang.org/docs/install.html): One line to the latest release. No external toolchain to set up. - [Hello world](https://machlang.org/docs/hello-world.html): Name your entry symbol and import the runtime yourself. - [Tags & sel](https://machlang.org/docs/records-unions.html#tags): Tagged values, payload guards, and compile-time case reflection. - [Comptime](https://machlang.org/docs/variadics.html): Monomorphized generics, variadic packs, and compile-time control. - [Manifest](https://machlang.org/docs/manifest.html): Declare artifacts, targets, and dependencies in mach.toml. - [Inline assembly](https://machlang.org/docs/asm.html): One ISA-tagged form, with clobbers inferred from the opcodes. - [Secrecy](https://machlang.org/docs/secrecy.html): The ^ qualifier and the constant-time guarantee. Experimental. - [GPU shaders](https://machlang.org/docs/shaders.html): Pipeline stages, descriptor bindings, and textures in the same language. #### Where to go next New to the language? Start with [Hello world](https://machlang.org/docs/hello-world.html) and the [language tour](https://machlang.org/docs/types.html). Want to see tagged values? Read [Tags & sel](https://machlang.org/docs/records-unions.html#tags). Want to see the compile-time system in depth? Read [Variadic packs](https://machlang.org/docs/variadics.html). Working close to the machine? [Inline assembly](https://machlang.org/docs/asm.html) and [Secrecy](https://machlang.org/docs/secrecy.html) cover the two subsystems that reach furthest down. Source: https://machlang.org/docs/install.html ### Install Install the latest mach release with a single command, grab a precompiled binary, or build the self-hosted compiler from source. There are no external toolchain dependencies to install first. #### Quick install The one-line installer fetches the latest release for your platform. On Unix and macOS: ``` curl -fsSL https://machlang.org/install.sh | sh ``` On Windows (PowerShell): ``` irm https://machlang.org/install.ps1 | iex ``` > **Note:** The docs read like a pamphlet, not a bible, and assume you know basic programming concepts from other languages. Skim the language reference alongside this guide as you go. #### What the installer does Each script verifies the download against the release `SHA256SUMS` before extracting it, and a mismatch or a missing entry installs nothing. It then drops the `mach` binary into a per-user location. At a terminal it asks for the install directory, offering the default below. `MACH_INSTALL_DIR` skips the prompt with a directory of your own, and `MACH_VERSION` pins a release (for example `6.8.0`) instead of the latest. | Platform | Install location | | --- | --- | | Unix / macOS | `~/.local/bin` | | Windows | `%LOCALAPPDATA%\mach\bin` | #### Precompiled binaries Precompiled binaries for each release are available directly on the [releases page](https://github.com/briar-systems/mach/releases). Download the archive for your target and place the `mach` binary on your `PATH` if you prefer not to use the installer. #### Build from source Mach builds itself, so compiling from source needs an existing `mach` installation. Install a release first, then clone, pull dependencies, and build. ``` git clone https://github.com/briar-systems/mach cd mach mach dep pull mach build . ``` The compiler is written to `out//bin/mach`, where `` is the selected target name. > **Restriction:** Because mach is self-hosted, there is no bootstrap path without a working `mach` binary. Use the quick installer or a precompiled binary to get the first compiler, then build from source. #### See also - [Hello world](https://machlang.org/docs/hello-world.html) - write, build, and run your first program - [Project layout](https://machlang.org/docs/project-layout.html) - scaffold a project with `mach init` - [CLI](https://machlang.org/docs/cli.html) - the `mach` command-line reference - [Dependencies](https://machlang.org/docs/dependencies.html) - what `mach dep pull` vendors Source: https://machlang.org/docs/hello-world.html ### Hello world A complete mach program in a handful of lines: a couple of imports, an entry point, and one printed line. Every piece is explicit, so nothing runs that you did not write. #### The program This is the whole thing, and it is exactly the `src/root.mach` that `mach init` writes. It imports the runtime and the print module, declares an entry point, prints a line, and returns a zero exit code. ```mach use std.runtime; use print: std.print; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.println("Hello, World!"); ret 0; } ``` > **Note:** The program relies on the standard library as a dependency. `mach init` scaffolds a project with std already declared and fetched, or clone the standalone [Mach Sieve](https://github.com/briar-systems/mach-sieve) starting point. #### Imports Modules are named by full, project-rooted dotted paths. The first line pulls in the standard runtime that every program builds against; it takes no alias because nothing here calls it directly. ```mach use std.runtime; # bring the runtime into the build use print: std.print; # bind std.print to the short name print ``` The `use alias: path` form binds a short handle to a path. With `print` bound, the module's symbols are reached through it as `print.println`. See [Modules](https://machlang.org/docs/modules.html) for path resolution and aliases. #### The entry point Execution starts at a single function. Its name in source can be anything; what makes it the entry point is the exported symbol and its signature. ##### The symbol decorator `#[symbol("main")]` exports the function under the symbol `main`, which is the name the program is entered through. Decorators attach to the declaration that follows them. ```mach #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { ... } ``` > **Note:** Decorators use the `#[...]` syntax. See [Decorators](https://machlang.org/docs/decorators.html) for the full set. ##### Arguments and exit code The entry function takes the argument count `argc: i64` and the argument vector `argv: **u8` (a pointer to an array of C-style strings). It returns an `i64` process exit code, so `ret 0;` reports success. #### Printing `print.println` writes its argument followed by a newline. For formatted output, `print.printf` fills each `{}` hole in the format string with the next trailing argument, in order, and writes exactly what you give it. ```mach print.printf("x = {}\n", 42::i64); # x = 42 ``` The `::i64` casts the literal to a concrete type. Because `printf` writes the string verbatim, include the trailing `\n` yourself. #### Build and run The same toolchain scaffolds, fetches dependencies, builds and links: ``` mach init hello # scaffold hello/ with src/root.mach and std cd hello mach build . # compile and link mach run . # build and run: Hello, World! ``` The binary lands at `out///bin/hello`, for example `out/linux-x86_64/debug/bin/hello`. `mach init` already fetches std, so there is no separate fetch step. See the [CLI](https://machlang.org/docs/cli.html) reference for the full command set. #### See also - [Functions](https://machlang.org/docs/functions.html) - declarations, parameters, and generics - [Modules](https://machlang.org/docs/modules.html) - dotted paths, use, and aliases - [CLI](https://machlang.org/docs/cli.html) - build, dependency, and test commands - [Project layout](https://machlang.org/docs/project-layout.html) - how a project tree maps to modules Source: https://machlang.org/docs/project-layout.html ### Project layout A mach project is a tree of `.mach` files under a source directory, anchored by a `mach.toml` manifest at the root. The project's `id` is the root of every module path the project exposes. #### Source files Source files use the `.mach` extension. `mach init` writes one of two conventional entry files under `src/`: - `root.mach` - executable entry, written by `mach init`. - `lib.mach` - library entry, written by `mach init --lib`. Exposes code for other projects to import and produces no executable. These names are convention only, and the compiler never looks for a function called `main`. An entry is simply the source file an `[artifact.*]` names with its `entry` key; what marks a function as the program's entry point is its exported symbol and signature, covered in [Hello world](https://machlang.org/docs/hello-world.html). The artifact's `kind` fixes the role - `"bin"` links an executable, `"static"` produces an archive. A project may carry both files. #### The project root Every project has a `mach.toml` at its root, and every command takes the project explicitly as a directory or a manifest path. `mach init` writes a complete one. Abridged, a binary scaffold for a project named `myproj` on an x86_64 host looks like this: ``` [project] id = "myproj" # root of every module path version = "0.1.0" mach = "^6.0" # the compiler versions this project builds with src = "src" # source dir out = "out/{target.name}/{profile.name}" [target.linux-x86_64] isa = "x86_64" os = "linux" abi = "sysv64" # ... windows and darwin targets [artifact.myproj] kind = "bin" entry = "root.mach" # executable entry, relative to src out = "bin/myproj{artifact.suffix}" targets = ["*"] link = ["kernel32"] need = [] # ... [link.kernel32] and the debug and release profiles [dep.std] git = "https://github.com/briar-systems/mach-std" version = "^9.0" # newest std release that admits this compiler ``` `[project]` requires `id`, `version`, `src`, and `out`, and a root manifest also requires `mach`, the [compiler range](https://machlang.org/docs/manifest.html#compiler-range). A root manifest without it is refused. `src` is where module paths resolve, and vendored dependencies live under `dep/`. The full schema - targets, profiles, artifacts, dependencies, output templates - lives on the [Manifest](https://machlang.org/docs/manifest.html) page, and version ranges on [Dependencies](https://machlang.org/docs/dependencies.html#version-ranges). #### Files map to module paths The path separator is `.`. A file at `/foo/bar.mach` in a project with `id = "myproj"` is reachable as `myproj.foo.bar`. ``` myproj/ mach.toml src/ root.mach # myproj.root foo/ bar.mach # myproj.foo.bar ``` There is no `this.` self-prefix. Within a project, modules always reference each other by their full project-rooted path, so every reference is syntactically uniform no matter where it appears. #### Surface and split modules A file `foo.mach` may co-exist with a directory `foo/`. The file is the **surface** module - the public face of `foo`. The directory holds **split** implementations that the surface loads and re-exports. ``` src/ foo.mach # surface: myproj.foo foo/ a.mach # split: myproj.foo.a b.mach # split: myproj.foo.b ``` The surface loads each split with `use` and re-exports its public symbols with `fwd`. Consumers `use myproj.foo` and reach symbols through the surface; they never name the split files directly. ```mach use myproj.foo.a; fwd a.X; # re-export a's public symbol through the surface ``` ##### Topical splits Organize a large module by topic. Every split is forwarded unconditionally, keeping one logical module spread across readable files. ##### Multiplatform splits Keep one implementation per target and select it with `$if` on `$mach.build.os` or `$mach.build.arch`, so the surface forwards only the split that matches the build. #### Importing a project by id A one-segment `use`/`fwd` path equal to a resolvable project id - a dependency's id or the current project's own - resolves to that project's declared `[project].module`. A library that sets `module = "glfw.mach"` is imported by its id instead of by the full path to its surface file. ``` # in glfw's manifest [project] id = "glfw" module = "glfw.mach" # the surface a bare use glfw binds ``` ```mach # in a consumer use glfw; # binds glfw's [project].module ``` > **Note:** A bare import of a project that declares no `module` is a resolution error naming the fix: import a full path, or add a `module` key. A declared `module` that names no file is a manifest error at build start. Longer paths are unaffected, so this is purely additive. #### See also - [Modules](https://machlang.org/docs/modules.html) - how files map to dotted module paths - [Manifest](https://machlang.org/docs/manifest.html) - the full `mach.toml` schema - [Dependencies](https://machlang.org/docs/dependencies.html) - vendored deps under the `dep` dir - [Visibility](https://machlang.org/docs/visibility.html) - `pub` and `fwd` re-exports Source: https://machlang.org/docs/types.html ### Types Mach ships a small set of compiler-seeded primitive types plus a uniform grammar for building pointers, arrays, and function types out of them. There are no compiler-known aliases: names like `bool`, `usize`, and `str` are stdlib `def`s. #### Primitive scalars Eleven names make up the complete set of primitive types the compiler seeds. Everything else is built on top of them. | Family | Members | | --- | --- | | Unsigned int | `u8`, `u16`, `u32`, `u64` | | Signed int | `i8`, `i16`, `i32`, `i64` | | Float | `f32`, `f64` | | Untyped pointer | `ptr` | > **Note:** There is no compiler `bool`. It is a stdlib alias `def bool: u8;`, with `true` and `false` as stdlib `val`s (`1` and `0`). #### SIMD vectors A vector type is a **form**, not a fixed list: any primitive numeric element followed by `x` and a lane count. ```mach f32x4 # 4 lanes of f32 i32x8 # 8 lanes of i32 u8x16 # 16 lanes of u8 ``` The spelling is `x` with a **single** `x`. A name like `f32x4x4` is not a vector type and resolves as an ordinary identifier: a matrix is an algorithm over vectors and belongs in a library, not in the language. `ptr` is not a lane element, since its width is target-defined rather than a scalar bit count. Two rules bound the form, and **neither depends on the target**: at least 2 lanes (`f32x1` is refused - a one-lane vector is just its scalar), and at most 65535 (a compiler limit; the lane count rides a 16-bit field through codegen). Everything else is legal at every width. `f32x3` is 96 bits, `f32x8` is 256, and both compile on every target - including one with no vector unit at all. **Width is a realization question, not a legality one.** How a shape is realized does depend on the target, and there are three answers: one packed instruction where the shape fits a vector register and the operation has a packed form; a value placed in memory and worked one lane at a time where it does not; and per-lane scalar code on a target with no vector unit (riscv64 today). All three compute identical lanes - the expansion is a fixed unroll, never a reassociation - so only performance varies. A target that gains wider vector registers therefore gets **better code**, not new spellings. The `simd` profile key reports or refuses the scalar cases when a project cannot afford them; see [Manifest](https://machlang.org/docs/manifest.html#profiles). A vector spelling is recognized only in **type** position, so a *value* may still be named `f32x4`. A *type* may not: `rec`, `uni`, `tag`, and `def` reject a name spelled as a vector form, because such a type would be silently unreachable - every use in type position resolves to the vector instead. ```mach val f32x4: i64 = 7; # fine: values are a different position rec f32x3 { x: f32; } # error: `f32x3` is spelled as a vector type ``` ##### Size and alignment `$size_of` is **lane-derived** - lanes x element size, packed, with no padding at any width. | Type | `$size_of` | `$align_of` | | --- | --- | --- | | `f32x2` | 8 | 4 | | `f32x3` | 12 | 4 | | `f32x4` | 16 | 16 | | `i16x4` | 8 | 2 | | `f32x8` | 32 | 16 | `$align_of` has two rungs, each for its own reason. A vector *narrower* than the vector register is a packed aggregate that loads piecewise, so it aligns to a single lane - which is what makes `[N]f32x3` a usable packed vertex buffer, since padding `f32x3` to 16 bytes would make it indistinguishable from `f32x4` in memory. One that *fills* the register aligns to its whole size, because the machine's vector load requires it. One *wider* than the register is placed as several register-width pieces and so aligns to the register width - 16 for `f32x8`, not 32. ##### Literals and lanes **Literals are full-arity**: one initializer per lane, mirroring array literals. Too few or too many lanes is a compile error. An uninitialized vector local default-initializes to all-zero lanes. ```mach val v: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; var z: i32x4; # every lane is 0 ``` `v[i]` reads or writes a single lane. The index must be a comptime constant in `[0, lanes)`, checked the same way an array index is: `v[4]` on a `u32x4` reports ``index 4 is out of bounds for `u32x4` of length 4``. A dynamic lane index is not supported in this increment. ```mach var v: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; val x: f32 = v[0]; # read lane 0 v[3] = 9.0; # write lane 3 ``` Arrays of vectors (`[4]f32x4`) and pointers to vectors (`*f32x4`) are ordinary composite types over a vector element. There are no scalar-to-vector casts in this increment: neither an implicit conversion nor a `1.0::f32x4` reinterpret is legal. #### Handles A **handle** is a type whose representation is not the program's: the owning target mints it and the pipeline binds it. A shader reads a texture through one. The language knows only the machinery. A handle is a **bodyless `def`** carrying `#[handle(target, constructor, operands...)]`, and which constructors exist, what operands each takes, and what an operand means all belong to the named target. ```mach #[handle("spirv", "image", TEXEL_F32, DIM_2D, NO_DEPTH, NONARRAYED, SINGLE_SAMPLED, SAMPLED)] pub def Texture2D; #[handle("spirv", "sampled_image", Texture2D)] pub def Sampler2D; #[handle("spirv", "sampler")] pub def Sampler; ``` `def` is the carrier because it already means "this name denotes a type" and promises no fields and no storage, which is exactly what a handle is. There is no body because the target supplies the definition. Operands are ordinary comptime constants, with one exception: a constructor composing over another handle takes a **type name**, so `Sampler2D` names the image it wraps rather than restating that image's operands and cannot disagree with it. Every rule a handle carries follows from the one fact the directive states, and the set is fixed and closed rather than varied per declaration: no fields, no indexing, no construction; it cannot be a local binding, a record or union field, or sit behind a pointer or inside an array; it reaches an operation only by being passed to one, bound as a descriptor; and its extent is declared by the owning target. `$size_of` a handle is the target's pointer size - it is a name for a resource, and a pointer is the shape every target already has for that. On a target that mints no such type the declaration is **inert**: it still denotes a type and still sizes, and an operation over it is an undefined symbol at link. A target refuses an operand combination its constructor spells but it cannot emit, naming the operand rather than the declaration. See [GPU shaders](https://machlang.org/docs/shaders.html) for binding a handle and sampling through one. #### Pointers `*T` is a pointer to a value of type `T`. Take an address with `?` and read through it with `@`. ```mach var x: i64; var p: *i64 = ?x; # address-of yields a pointer val v: i64 = @p; # dereference reads through it ``` #### Arrays `[N]T` is an array of exactly `N` values of type `T`. Arrays nest as `[N][M]T`. ```mach val a: [4]i64 = [4]i64{1, 2, 3, 4}; val g: [2][2]i64 = [2][2]i64{ [2]i64{1, 2}, [2]i64{3, 4} }; ``` **Constant indices are bounds-checked at compile time.** `N` is part of the type, so an index the compiler can fold must land in `[0, N)`. ```mach var xs: [4]i32; val a: i32 = xs[3]; # ok val b: i32 = xs[4]; # error: index 4 is out of bounds for `[4]i32` of length 4 ``` The rule is keyed on the length the type carries, not on how the array was spelled, so a `def` alias, an array field of a generic instance, a nested array, and a `^`-qualified array are all checked the same way. It is exactly the length `$length_of` reports, and a vector's lane count takes the identical rule. Only a **constant** index is checked: a runtime index is not, and a pointer is not indexed against any length at all, since `*T` carries none. #### Function types `fun(T1, T2) R` is a first-class function-pointer type, usable as a value's type just like any scalar. ```mach def BinOp: fun(i64, i64) i64; val op: BinOp = add; val r: i64 = op(2, 3); ``` #### Named types and aliases `rec`, `uni`, and `tag` declarations produce named types. A `def` introduces an alias for any type, including the constructed forms above. ```mach def Bytes: [16]u8; # Bytes now names an array type def bool: u8; # the stdlib's own bool alias ``` #### See also - [Values and variables](https://machlang.org/docs/values.html) - `def` aliases, `val` constants, and how `bool` is built - [Records, unions, and tags](https://machlang.org/docs/records-unions.html) - declaring named `rec`, `uni`, and `tag` types - [Expressions](https://machlang.org/docs/expressions.html) - which operations work on each type - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$size_of`, `$align_of`, and `$length_of` over any type - [GPU shaders](https://machlang.org/docs/shaders.html) - what a handle is bound to, and the vectors a stage computes over - [Secrecy](https://machlang.org/docs/secrecy.html) - the `^` qualifier over any of these types Source: https://machlang.org/docs/values.html ### Values and variables Bindings introduce named values. `val` is immutable, `var` is mutable, and both require an explicit type - mach has no type inference. #### Declaring bindings A binding names a value with a declared type. `val` is constant once set; `var` can be reassigned. Every form pins the type in the declaration. ```mach val NAME: TYPE = EXPR; # immutable; initializer required var NAME: TYPE; # mutable; default-initialized var NAME: TYPE = EXPR; # mutable; explicit initializer ``` #### Mutability and defaults A `val` always carries an initializer and cannot be reassigned. A `var` may omit its initializer, in which case it is default-initialized (zeroed), and it can be written to afterward. ```mach val pi: f64 = 3.14159; val n: i64 = 42; var counter: i64 = 0; var buf: [256]u8; # default-initialized to zero counter = counter + 1; # var is reassignable ``` #### Scope and exports Both forms work at module top level and inside function bodies. Inside a function, a binding is local to its enclosing block. At module top level, `pub` exports it: - `pub val NAME` exports the constant. - `pub var NAME` exports the variable as a writable global. #### No type inference Every binding declares its type. An untyped numeric literal is checked *against* the binding's declared type; it never participates in inferring that type. ```mach val n: i64 = 42; # ok - 42 conforms to i64 val x = 42; # ERROR - no type to check against ``` > **Note:** When the surrounding context does not constrain a literal's type, give it a typed suffix such as `42i64`. #### See also - [Types](https://machlang.org/docs/types.html) - the type grammar used in the annotation - [Visibility](https://machlang.org/docs/visibility.html) - how `pub` exports bindings - [Functions](https://machlang.org/docs/functions.html) - parameters and return values Source: https://machlang.org/docs/functions.html ### Functions A function takes typed arguments, optionally returns a typed value, and runs a body of statements. Beyond the plain form, `fun` also carries generic type parameters, compile-time value parameters, and a trailing variadic pack. #### Declaring a function Every function starts with `fun`, a name, a parenthesized parameter list, an optional return type, and a brace-delimited body. The return type sits between the parameter list and the body; omit it for a function that returns nothing. ```mach fun NAME(args) RET { ... } # function with return type fun NAME(args) { ... } # no return type fun NAME[T](args) RET { ... } # generic over type parameters fun NAME($p: T, args) RET { ... } # comptime value parameter fun NAME(fixed, va: ...) RET { ... } # variadic pack parameter ``` Parameters are written `name: Type`. A value is returned with `ret`. ```mach pub fun add(a: i64, b: i64) i64 { ret a + b; } pub fun bump() { # no return value counter = counter + 1; } ``` #### Generic type parameters Generic functions take type parameters in brackets, `[T]`. There are no constraints: any type may be substituted. The compiler monomorphizes a separate instance per unique type instantiation. ```mach pub fun identity[T](value: T) T { ret value; } pub fun make_pair[T, U](a: T, b: U) Pair[T, U] { var p: Pair[T, U]; p.left = a; p.right = b; ret p; } ``` Call sites supply the types explicitly in brackets: ```mach val x: i64 = identity[i64](42); val p: Pair[i64, u8] = make_pair[i64, u8](1, 2u8); ``` ##### An instance is a value `f[T, ..]` with no call after it denotes the monomorphized **instance itself** and types as the instantiated signature, so a generic can be stored in a binding, passed as a callback, returned, placed in a table, and addressed. ```mach val fp: fun(i64) i64 = identity[i64]; # the instance as a value val ap: *fun(i64) i64 = ?identity[i64]; # its address ``` The instance is reached through the same path a call site uses, so the two spellings name one instance under one linkage name. A **bare** `f` is still not a value and has no address: a generic is a template, not code, and only an instance denotes a function. #### Comptime value parameters A parameter marked `$name: T` must be supplied with a value the compiler can resolve at compile time. The body can branch on it with `$if`, producing different code per call-site instantiation. ```mach pub fun pick_op($mode: Mode, a: i64, b: i64) i64 { $if (mode == MODE_FAST) { ret a + b; } $or (mode == MODE_SAFE) { # extra logic here ret a + b; } } ``` > **Restriction:** Comptime value parameters apply to function parameters only - not record fields, and not other contexts. #### Variadic packs A trailing named pack parameter, `va: ...`, accepts a variable number of trailing arguments. Any number of leading fixed parameters may precede it. The compiler monomorphizes the function once per distinct call-site type-list; the pack is consumed by `$each a in va` at compile time. There is no runtime `va_list`. ```mach pub fun sum(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a; } ret t; } # leading fixed params are allowed before the pack pub fun bias(base: i64, va: ...) i64 { var t: i64 = base; $each a in va { t = t + a; } ret t; } ``` `va.len` folds to the element count, and `g(va...)` forwards the whole pack to another pack-tailed callee. See the variadic packs reference for the full rules. #### See also - [Variadic packs](https://machlang.org/docs/variadics.html) - the full pack parameter reference - [Control flow](https://machlang.org/docs/comptime-control.html) - `$if` and `$each` inside function bodies - [Expressions](https://machlang.org/docs/expressions.html) - function calls and generic instantiation - [Types](https://machlang.org/docs/types.html) - the types parameters and returns are written in Source: https://machlang.org/docs/expressions.html ### Expressions Expressions evaluate to values. They appear on the right of a binding, as conditions, and as call arguments. #### Names A bare identifier references a name in scope. Module-qualified names use the dot path. ``` counter # local or module-level binding core.add # symbol from module core ``` #### Literals The primitive literal forms are numeric, char, string, and `nil`. ``` 42 # numeric literal 'a' # char literal "hello" # string literal nil # the nil value ``` #### Composite literals A type name followed by a brace-delimited initializer builds a record, array, union, or tag value. For generics, the type arguments appear in brackets before the body. ```mach val p: Point = Point{ x: 1, y: 2 }; val a: [3]i64 = [3]i64{ 10, 20, 30 }; val u: Number = Number{ i: 99 }; val pair: Pair[i64, u8] = Pair[i64, u8]{ left: 5, right: 6u8 }; ``` A tag value specifies the case name before the braces. A payload case supplies its value inside the braces, while a unit case uses empty braces. ```mach val s: Shape = Shape.circle{4.5}; val p: Shape = Shape.point{}; ``` A vector literal takes the same shape and is **full-arity**: one initializer per lane, with too few or too many a compile error. See [Types](https://machlang.org/docs/types.html#vectors). ```mach val v: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; ``` #### Field, index, and tag access A field is read with the dot, an array element with a bracketed index. ```mach val x: i64 = p.x; # record field val first: i64 = a[0]; # array index ``` A tag payload is read with `place.case` under a dominating guard that proves the case. In comptime reflection over `$cases(T)`, a tag case can also be accessed dynamically through `place.[case]`. #### Function calls A call applies a function to a parenthesized argument list. ```mach add(2, 3) ``` ##### Generic calls Type arguments for a generic function appear in brackets before the call parentheses. ```mach identity[i64](42) # generic call: type args in [ ] ``` ##### Variadic calls A variadic call passes extra trailing arguments past the fixed parameters. ```mach sum(3, 10i64, 20i64, 30i64) ``` > **Not yet implemented:** Variadic call sites parse, but the callee-side `va_list` machinery is not yet implemented. See [Functions](https://machlang.org/docs/functions.html). For the compile-time alternative, see [Variadic packs](https://machlang.org/docs/variadics.html). ##### Comptime arguments A comptime argument is passed positionally, just like a runtime argument. The function signature decides whether a given argument must be comptime-knowable. ```mach checked_add(MODE_FAST, 1, 2) # MODE_FAST is comptime-knowable ``` #### Operators Operators combine expressions into larger expressions. Precedence follows the usual C-family conventions. ```mach val n: i64 = a + b * c; # * binds tighter than + ``` ##### Lane-wise operators Vector types carry lane-wise operators at **every** lane count, not only the 128-bit shapes, and **which operators are legal is target-independent**: an `f32x8` add is as legal as an `f32x4` one, and the two differ only in how they are realized. | Lane family | `+` `-` | `*` | `/` | `%` | `& \| ^ ~` | `<< >>` | comparisons | | --- | --- | --- | --- | --- | --- | --- | --- | | float | yes | yes | yes | no | - | no | same-shape unsigned mask | | integer | yes | yes | no | no | yes | no | same-shape unsigned mask | Both operands must be the **same** vector shape: there is no implicit scalar-to-vector mixing and no cross-shape widening. Anything the table marks `no` is a compile error, not a silent fallback. A comparison produces the same-shape **unsigned mask** vector - one lane per input lane, all-ones bits for true and all-zeros for false, exactly what the hardware compare yields. There is no vector-bool type, and select is not an operator: it is the library idiom `(mask & a) | (~mask & b)` over matching integer lanes. ```mach val a: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; val b: f32x4 = f32x4{4.0, 3.0, 2.0, 1.0}; val sum: f32x4 = a + b; # lane-wise -> {5.0, 5.0, 5.0, 5.0} val mask: u32x4 = a < b; # -> {0xFFFFFFFF, 0xFFFFFFFF, 0, 0} ``` ##### Case selection: sel The `sel place.case` expression tests the active case of a tag, returning a `bool`. In comptime reflection over `$cases(T)`, `sel place.[case]` tests the reflected case. ```mach if (sel r.ok) { val value: i64 = r.ok; } ``` ##### Declassification: :>T The declassify operator `x:>T` converts a secret value of type `^T` into a public value of type `T`. The public destination type must always be written explicitly. See [Secrecy](https://machlang.org/docs/secrecy.html). ```mach val diff: ^u64 = ct_diff(secret, given); val ok: bool = (diff:>u64) == 0; ``` #### See also - [Statements](https://machlang.org/docs/statements.html) - how expressions appear inside statements - [Functions](https://machlang.org/docs/functions.html) - declarations, generic and comptime parameters - [Records, unions, and tags](https://machlang.org/docs/records-unions.html) - tag declarations and lexical payload guards - [Secrecy](https://machlang.org/docs/secrecy.html) - secret types and declassification rules - [Types](https://machlang.org/docs/types.html) - record, array, union, tag, and vector types - [Values and variables](https://machlang.org/docs/values.html) - bindings and literal forms Source: https://machlang.org/docs/statements.html ### Statements Statements are the runtime steps that fill function bodies and blocks. Every statement ends with `;`, except the ones that end with a block `{...}`. #### Conditionals: `if` / `or` An `if` head opens a branch chain. Each `or (cond) { ... }` adds another conditional branch, and a trailing `or { ... }` with no condition is the catch-all. Every body is a block; there is no brace-less single-statement form. ```mach if (cond) { ... } or (cond) { ... } or { ... # final else (no condition) } ``` ##### Lexical payload guards An `if (sel p.c)` condition opens a compile-time payload guard for the case `p.c` inside its block. Reading `p.c` is only permitted within such a dominating guard. Guards also propagate across logical `&&`: the right side of `sel p.c && expr` is guarded by the check on the left. In addition, when an `if` branch exits the enclosing block using `ret`, `brk`, or `cnt`, the compiler knows the exited case cannot reach subsequent statements. > **Guard limits:** An `or` branch is not guarded because the tested condition failed. Logical `||` and `!` do not establish guards. Keep tests to one condition per check when unwrapping payloads. #### Loops: `for` mach has a single condition-driven loop. The body repeats while the condition holds. There is no for-each form. ```mach var i: i64 = 0; for (i < 10) { i = i + 1; } ``` #### Returning: `ret` `ret expr;` returns a value; bare `ret;` returns from a void function. ```mach ret expr; # return a value ret; # return from a void function ``` #### Loop control: `brk` / `cnt` `brk` exits the enclosing `for`; `cnt` skips to the next iteration. ```mach for (i < 10) { if (i == 3) { cnt; } if (i == 8) { brk; } i = i + 1; } ``` > **Note:** Both are operand-less, so they are keywords only in their bare `brk;` / `cnt;` form. The same word followed by anything else is an ordinary identifier: a variable named `cnt` reads and assigns normally (`cnt = x;`), even inside a loop that also uses bare `cnt;` for control flow. #### Deferred cleanup: `fin` `fin` schedules a statement or block to run when the enclosing scope exits, in reverse order of declaration. It is for cleanup that should happen regardless of how the scope is left. ```mach { fin counter = counter - 1; fin { counter = counter * 2; } # ... code ... } # at scope exit: fin block runs first, then fin counter = counter - 1 ``` `fin` takes either a single statement (`fin stmt;`) or a block (`fin { ... }`). There is no bare-expression form. #### Blocks `{ ... }` introduces a new lexical scope. Statements inside run in order, and a block can stand alone to bound the lifetime of locals. ```mach { val tmp: i64 = compute(); use_tmp(tmp); } ``` #### See also - [Expressions](https://machlang.org/docs/expressions.html) - what goes on the right side of `=` and inside conditions - [Values and variables](https://machlang.org/docs/values.html) - `var` and `val` declarations - [Records, unions, and tags](https://machlang.org/docs/records-unions.html) - tag declarations and payload safety - [Control flow](https://machlang.org/docs/comptime-control.html) - the comptime counterparts `$if` and `$each` Source: https://machlang.org/docs/records-unions.html ### Records, unions, and tags A `rec` is a named collection of typed fields, each with its own storage. A `uni` overlaps its fields in the same memory. A `tag` defines a discriminated value with compiler-enforced lexical guards. #### Records A `rec` lays out its fields side by side. Each field has independent storage, and the record's size is the sum of the field sizes plus any padding the compiler inserts for alignment. A record may be generic over type parameters. ```mach rec NAME { field1: type; field2: type; ... } rec NAME[T, U] { ... } # generic over type parameters ``` ```mach pub rec Point { x: i64; y: i64; } pub rec Pair[T, U] { left: T; right: U; } ``` #### Construction and access A record literal names the type and supplies each field by name. Read a field with `.`. ```mach val p: Point = Point{ x: 1, y: 2 }; val q: Pair[i64, u8] = Pair[i64, u8]{ left: 5, right: 6u8 }; val n: i64 = p.x; # field access via . ``` #### Memory layout By default the compiler may insert padding between fields to satisfy each field's alignment, following the natural C-style rule. The `#[align(N)]` decorator on a record raises its minimum type alignment to `N` bytes, where `N` is a power of two. ```mach #[align(16)] pub rec Aligned { a: u8; b: i64; } ``` `#[packed]` is the inverse: every field sits immediately after the previous one, there is no tail padding, and the type takes no alignment from its fields. That is what describes a shape whose layout is not mach's to choose - a C struct, a file header, a wire frame. The two compose, each owning one question: `packed` decides padding, `align` decides the record's own alignment. ```mach #[packed] rec Header { magic: u8; # offset 0 version: u16; # offset 1 length: u32; # offset 3 checksum: u64; # offset 7 } # $size_of == 15, $align_of == 1 ``` Packing is **not transitive**: a record a packed record contains keeps its own internal padding and is merely placed without any. Taking the address of a packed field is refused, since the resulting pointer type would state an alignment its storage does not have; `?r` on the whole record stays legal. See [Decorators](https://machlang.org/docs/decorators.html#packed) for the full rule set. #### Unions A `uni` is a collection of named fields that share the same memory. Writing one field overwrites whatever bytes the others held. A union's size is the size of its largest field plus any alignment padding. Unions may be generic. ```mach pub uni Number { i: i64; f: f64; } pub uni Maybe[T] { some: T; none: u8; } ``` > **Restriction:** The compiler does not track which field of a union is live. Reading a field other than the one last written reinterprets raw bytes; keeping the active field straight is the programmer's responsibility. #### Tags A `tag` defines a discriminated value with compiler-enforced payload safety. Unlike an untracked union, a tag stores a discriminator alongside an overlapping payload area, and the compiler prevents reading any payload unless the active case is proven by a dominating guard. The discriminator type is written after the tag name and must be an unsigned integer type from `u8` through `u64`. Cases may carry a single typed payload or remain payload-free. ```mach pub tag Shape: u8 { circle: f64; rect: Rect; point; } pub tag Option[T]: u8 { none; some: T; } ``` Values are constructed using the type name, case name, and curly braces. A payload case supplies its value inside the braces, while a unit case uses empty braces. ```mach val c: Shape = Shape.circle{3.14}; val p: Shape = Shape.point{}; ``` #### Payload guards and sel The `sel` keyword tests whether a tag instance currently holds a specific case. The expression `sel place.case` yields a `bool`. Reading a case payload is permitted only where a lexical guard guarantees that case is active. Accessing a payload without a dominating guard produces a compile error. ```mach fun area(s: Shape) f64 { if (sel s.circle) { ret 3.141592653589793 * s.circle * s.circle; } if (sel s.rect) { ret s.rect.w * s.rect.h; } ret 0.0; } ``` Guards obey strict lexical scope rules: - The body of an `if (sel p.c)` block guards `p.c` for the duration of that block. - The right-hand side of a logical `&&` expression is guarded when `sel p.c` appears on the left. - When an `if (sel p.c)` arm exits through a return, break, or continue, later code in the enclosing block is not guarded for `p.c`. However, exiting when all other possibilities are handled allows narrowing. > **Guard restrictions:** Guards must be unambiguous. An `or` branch is not guarded because the condition evaluated to false. Logical `||` and negation with `!` do not open a payload guard. Test one condition at a time. #### See also - [Decorators](https://machlang.org/docs/decorators.html) - `#[deprecated]` on declarations and cases - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$cases`, `$is_tag`, and `$discriminant_of` for tag reflection - [Statements](https://machlang.org/docs/statements.html) - control flow and conditional guards - [Types](https://machlang.org/docs/types.html) - aggregate and primitive type definitions Source: https://machlang.org/docs/modules.html ### Modules A Mach project is a tree of modules rooted at the project's id. Each file is reachable by a dotted path, where `use` imports a name privately and `fwd` re-exports it on a module's public surface. #### Module paths Every `.mach` file under a project's source directory is a module, reachable by a dotted path from the project root. The path separator is `.`, and each segment mirrors the directory and file name. | Project id | File | Module path | | --- | --- | --- | | `myproj` | `src/foo/bar.mach` | `myproj.foo.bar` | There is no `this.` self-prefix. Within a project, modules always reference each other by their full project-rooted path, so every reference reads the same regardless of where it appears. #### Imports with `use` `use` brings an external symbol or module into the current scope under a local name. It is a **private** import: the imported name is not exposed to consumers of this module. ```mach use PATH; # binds the leaf component use ALIAS: PATH; # binds ALIAS ``` - The alias defaults to the path's leaf component when omitted. - One name per line. There is no splat (`use foo.*` does not exist) and no combined form (`use foo.{a, b, c}` does not exist). ##### Module vs symbol binding The resolver binds whatever the path points to. A path ending at a **module** binds the module, and you reach its members with qualified `module.member` access. A path ending at a **symbol** binds the symbol for bare use. Importing a module does not pull its members into scope unqualified: to use `usize` bare, import the symbol, not its module. ```mach use std.types.size; # binds module 'size'; use as size.usize use sz: std.types.size; # binds module under 'sz'; use as sz.usize use std.types.size.usize; # binds symbol 'usize'; use bare as usize use my_usize: std.types.size.usize; # binds symbol under 'my_usize' ``` > **Note:** A Mach module imports every dependency it directly names, including any reached only through a re-export. There is no shortcut for "my surface uses these transitively, just give me the leaf." The dependency graph stays visible at the top of every file. #### Re-exports with `fwd` `fwd` re-exports a symbol or module from another module under this module's public surface. It is the public counterpart to `use`, and it mirrors `use`'s grammar exactly. ```mach fwd PATH; # re-export under the path leaf fwd ALIAS: PATH; # re-export with rename ``` - `fwd` always publishes, so there is no `pub fwd` form. - One name per line. No splat. ```mach use impl: full.core.data; fwd impl.Point; # re-exports as 'Point' fwd Pt: impl.Point; # re-exports as 'Pt' ``` A `fwd` path that ends at a **module** re-exports the whole module as a public module alias, mirroring `use`'s module binding. A consumer reaches the alias's members with qualified access, chaining through any depth of re-export, including a `fwd` of another library's `fwd`. ```mach fwd demo.alpha; # re-exports module 'alpha' fwd deep: demo.deep.beta; # re-exports module under 'deep' use demo.lib; # lib.mach contains: fwd demo.alpha; lib.alpha.answer(); # resolves through the module re-export ``` As with `use`, a module alias is not a value; only its members can be named. #### The shadow-module pattern A file `foo.mach` may co-exist with a directory `foo/`. The file is the **surface** module, the public face of `foo`. The directory's files are **split** implementations that the surface loads and re-exports. ``` myproj/ foo.mach # surface foo/ a.mach # split: myproj.foo.a b.mach # split: myproj.foo.b ``` The surface loads each split with `use` and re-exports its public symbols with `fwd`. Consumers `use myproj.foo` and reach symbols through the surface, never naming the split files directly. ```mach # myproj/foo.mach (surface) use myproj.foo.a; use myproj.foo.b; fwd a.X; fwd b.Y; ``` Two common uses: - **Topical splits** - organize a large module by topic, with all splits forwarded unconditionally. - **Multiplatform splits** - one impl per target, selected by `$if` on `$mach.build.os` or `$mach.build.arch`, then aliased under a stable name for consumers. #### Bare project-id imports A one-segment `use`/`fwd` path equal to a resolvable project id (a dependency's id or the current project's own id) binds that project's public module. For a dependency, that is the `entry` shared by its library artifacts (`static` or `shared`) marked `default = true`. ```mach use mylib; # binds the entry of mylib's default library artifact ``` > **Note:** A bare import of a dependency that declares no default library artifact is an error, directing you to import a full path or mark a library artifact `default = true` in its manifest. Longer paths are unaffected. #### See also - [Visibility](https://machlang.org/docs/visibility.html) - what `pub` exposes for `use` and `fwd` to reach - [Manifest](https://machlang.org/docs/manifest.html) - declaring library artifacts with `default = true` - [Project layout](https://machlang.org/docs/project-layout.html) - how the source tree maps to module paths Source: https://machlang.org/docs/visibility.html ### Visibility Two declaration modifiers control how a symbol is seen: `pub` decides what crosses a module boundary, and `ext` declares a body-less function resolved at link time. #### The `pub` modifier `pub` marks a declaration as part of its module's public surface. Other modules that `use` this module can reference `pub`-marked symbols by name. A declaration without `pub` is file-private: only code in the same file can see it. ```mach pub fun add(a: i64, b: i64) i64 { ret a + b; } fun helper() { ... } # private: only callable inside this file pub rec Point { x: i64; y: i64; } pub tag Status { Ready, Failed(u32) } pub val MAX: i64 = 100; ``` The modifier applies to declarations: `fun`, `tag`, `rec`, `uni`, `def`, `val`, `var`, and `ext fun`. > **Note:** `fwd` always publishes and does not take an explicit `pub` modifier. #### External functions with `ext` `ext` declares a function with the C ABI as a forward reference: it has no body, and the linker resolves the symbol at link time. Only functions can be `ext`. ```mach #[symbol("write")] pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ``` - `ext fun` declarations have no body. - The C ABI is the contract: argument and return types must be representable in C. ##### Overriding the linker name By default the linker looks up the declared name. The `#[symbol("real_name")]` decorator overrides it, binding the declaration to a different external symbol. > **Restriction:** There are no body-less functions outside of `ext fun`. Regular forward declarations do not exist. #### See also - [Functions](https://machlang.org/docs/functions.html) - regular function declarations - [Modules](https://machlang.org/docs/modules.html) - how `use` reaches public symbols - [Decorators](https://machlang.org/docs/decorators.html) - `#[symbol]` and other attributes Source: https://machlang.org/docs/ext-fun.html ### External functions mach's foreign function interface (FFI) is a single declaration form: `ext fun`. It declares a body-less function with the C ABI - a forward reference the linker resolves at link time. Declare the foreign symbol, call it like any other function, and supply its definition (a C object, archive, or shared library) when you build. This is the only body-less function form mach allows. #### Grammar An `ext fun` is a function signature with no body. The declaration ends with a semicolon instead of a block. ```mach ext fun NAME(args) RET; ``` - No body block - the declaration ends with a semicolon. - Argument and return types must be representable in C. - `pub ext fun` exposes the import to other modules; without `pub` it is file-private. ```mach pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ext fun strlen(s: *u8) i64; # private, file-local ``` #### The C ABI boundary The C ABI is the contract at the boundary: every argument and return type must be representable in C, and values are passed exactly as the target's C calling convention dictates. mach's fixed-width scalars and pointers correspond directly to their C counterparts. - `i8` through `i64` and `u8` through `u64` are the sized C integers. - `f32` and `f64` are C `float` and `double`. - `*T` and `**T` are C pointers; a C-style string is a `*u8`. The exact calling convention is the one declared for the target (for example `sysv64` on Linux or `win64` on Windows). See [Manifest](https://machlang.org/docs/manifest.html#targets) for how a target's `abi` is set. #### Renaming the linker symbol By default the linker looks up the declared name. The `#[symbol("real_name")]` decorator overrides it, binding the declaration to a different external symbol. ```mach #[symbol("write")] pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ``` Two common reasons to rename: - The C name (such as `write`) would shadow other things in your namespace. - The target decorates symbol names in its ABI and you need to match the decorated form. #### Library attribution On a format that attributes each import to a specific dependency - the PE (Windows) import directory - the `#[library("...")]` decorator pins an `ext` import to the DLL that exports it. ```mach #[library("ws2_32.dll")] ext fun WSAStartup(ver: u16, data: *u8) i32; ``` - The named value must resolve to a `[link.X]` requirement selected for this build cell. Pinning to an absent library is a link error, never a silent fallback. - An import with no `library` is unattributed; on PE it binds to the first declared dependency. Pin every import that belongs to a different DLL. - On ELF (Linux) the loader resolves imports by global search, so `library` has no effect on the emitted binary; the value is still validated against the dependency set. `library` composes with `symbol`: the rename sets the imported symbol's name, `library` sets the DLL it is imported from. ```mach # imported as `socket` from ws2_32.dll, called as `ws2_socket` in mach. #[library("ws2_32.dll")] #[symbol("socket")] ext fun ws2_socket(af: i32, kind: i32, proto: i32) i64; ``` #### Linking the definition An `ext fun` is only a forward reference. Its definition is supplied at link time, either **statically** by a precompiled object or archive or **dynamically** by a shared library bound at load time. Provide those inputs to `mach build` on the command line or through the manifest. An undefined `ext` symbol that no input resolves is a link error, so a typo never silently drops a dependency. ``` mach build . path/to/libfoo.a # static archive (every member is pulled) mach build . -L build/libs -l foo # search dir + name mach build . -l c # link libc dynamically ``` The same inputs live in the manifest's `libs` overlay, merged with the command-line inputs: ```mach [target.linux] libs = ["build/libs/libfoo.a", "c"] ``` A loose `.o` object or static `.a` archive is a static input merged into the binary; a shared `.so` is a dynamic dependency whose undefined `ext` symbols bind at load time through an emitted PLT. A static definition always wins over a same-named dynamic import. Dynamic linking is implemented for the ELF (Linux) and PE (Windows) targets; the Mach-O (Darwin) import path is not yet implemented. > **Note:** The link-input resolution rules - explicit paths, `-l`/`-L` search order, and static-vs-dynamic selection - are documented once on the [CLI](https://machlang.org/docs/cli.html#link-inputs) page, and the manifest overlay on the [Manifest](https://machlang.org/docs/manifest.html#link-inputs) page. #### A worked example This program imports two libc functions, calls them, and prints a line to standard output. `strlen` measures the message and `libc_write` (the renamed `write`) writes it to file descriptor `1`. ```mach use std.runtime; #[symbol("write")] pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ext fun strlen(s: *u8) i64; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val msg: *u8 = "hello from ffi\n"; val n: i64 = strlen(msg); libc_write(1, msg, n); ret 0; } ``` libc supplies both symbols, so link it dynamically. The result is a dynamically-linked binary that binds `write` and `strlen` at load time. ``` mach build . -l c # resolves the ext imports against libc.so ``` #### See also - [Visibility](https://machlang.org/docs/visibility.html#ext) - `pub` and `ext` modifiers - [Decorators](https://machlang.org/docs/decorators.html#symbol) - the `symbol` and `library` directives - [CLI](https://machlang.org/docs/cli.html#link-inputs) - external link inputs and resolution - [Manifest](https://machlang.org/docs/manifest.html#link-inputs) - the `libs` link overlay - [Functions](https://machlang.org/docs/functions.html) - regular function declarations Source: https://machlang.org/docs/secrecy.html ### Secrecy A program can give away a password it never prints. Writing `^` in front of a type marks a value as secret, and the compiler then refuses to build code that would leak it - and keeps refusing all the way through optimization. > **Experimental preview:** The constant-time support is incomplete and unaudited. Read [Assurance](https://machlang.org/docs/secrecy.html#assurance) before relying on any of this. **Do not build production cryptography on it at this version.** #### What this is for Say you check a password by comparing it one character at a time, and stop at the first character that does not match. That is the obvious way to write it, and it is correct - it never prints the password, never logs it, and returns nothing but yes or no. It still gives the password away. An attacker who can *time* the check learns something from every attempt: a guess starting with the right character takes a hair longer to reject than one starting with the wrong character, because the comparison got one step further before giving up. Guess the first character - only one of them is measurably slower. Keep it, and guess the second. A password that would take longer than the universe to brute-force all at once falls in a few thousand tries, one character at a time. Timing is not the only such channel. Three things about a running program are observable without reading its memory: - **Which way it branched.** Taking one path rather than another takes a different amount of time, and leaves different traces in the processor. - **Which memory it touched.** Reading `table[secret]` pulls exactly that entry into the cache. An attacker sharing the machine can often work out which one. - **How long an instruction took.** Some instructions - division especially - finish faster or slower depending on the values fed to them. The standard defence is to write the code so none of those depend on the secret: compare every character whether or not an earlier one already failed, and combine the results arithmetically instead of branching. This is called **constant-time** code, and it is notoriously easy to get wrong. ##### Why the compiler has to be involved Even when you get it right by hand, the compiler can undo it. An optimizer's whole job is to notice that work is unnecessary and remove it - and branch-free code that carefully does the same work every time looks exactly like work worth eliminating. A shortcut you deliberately avoided writing gets helpfully added back, and nothing warns you. The source is still constant-time; the binary is not. So the guarantee cannot live in a coding convention or a library. It has to be something the compiler itself knows about and is obliged to preserve. That is what `^` is. ```mach # `^u64` is a secret. the compiler tracks where it flows. #[oblivious] fun ct_diff(a: ^u64, b: ^u64) ^u64 { ret a ^ b; # one xor, always the same work } fun verify(given: ^u64, secret: ^u64) bool { if (given == secret) { ret true; } # refused - see below ret (ct_diff(given, secret):>u64) == 0; } ``` That branch does not compile: ``` error: secret value used as a branch condition --> ./src/root.mach:16:5 | 16 | if (given == secret) { ret true; } | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ | = note: a secret may not steer control flow; branching on it leaks through timing. compute branch-free, or `:>T` to a public value when disclosure is intended ``` Three things are worth drawing out of that. The mistake is caught **when you compile**, not by an audit or a timing measurement after the fact. Secrecy is **part of the type**, so it travels with the value through every function it is passed to - you cannot lose track of which variable holds the sensitive one. And disclosure is possible but never accidental: `:>T` is the one way to turn a secret back into an ordinary value, so every deliberate release is a visible mark in the source that a reviewer can search for. The rest of this page is the precise version of all that. #### The secrecy lattice There are two secrecy levels in a two-point lattice: public is the bottom, secret the top. `^` lifts a type to secret and binds to the type immediately to its right, so it nests with `*` and `[N]` in any order: `^u32`, `*^u8`, `^*u8`, `[N]^u8`, `^MyRec`. Doubling collapses: `^^T` is `^T`. A public value coerces *up* to secret wherever a secret is expected, with no syntax. The reverse never happens implicitly. ```mach fun up(p: u32) ^u32 { ret p; } # public u32 flows into a secret slot ``` A **literal is public by construction** and stays public through that coercion: its value sits in the instruction stream, so classifying it as secret would protect nothing. That does not weaken the join below - a value *computed* with a secret is secret however its other operand is spelled - so `v << 3` on a secret `v` still yields a secret, while the constant `3` is not mistaken for a secret shift count. #### Join Any operation with a secret operand yields a secret result. Taint joins across arithmetic, bitwise, shift, and comparison operators, and through a value read out of a secret container. ```mach fun mix(a: ^u32, b: u32) ^u32 { ret a + b; } # ^u32 + u32 -> ^u32 rec Key { d: ^[32]u8; } fun first(k: Key) ^u8 { ret k.d[0]; } # element of a secret array is ^u8 ``` Taking the address of a `^T` value with `?` gives the public pointer `*^T` - the address is public, the pointee secret - and dereferencing it with `@` recovers the secret `^T`. #### Gates A secret may not reach a position the leakage model observes. Each is a compile error decided by operand type: - a secret branch or loop condition (`if`, `for`) - a secret left operand of a short-circuiting `&&` / `||` - it is the branch the operator keys on; a secret right operand only taints the result - a secret memory index (`table[i]` with `i` secret) - a secret memory address - an access through a secret *pointer*, whether by `@p`, `p[i]`, or the auto-deref in `p.x` - a secret operand of the always-variable-latency `/` or `%` ```mach fun leak(a: ^u32, t: *u8, p: ^*u8) u8 { if (a) { ret 1; } # error: secret value used as a branch condition ret t[a]; # error: secret value used as a memory index ret @p; # error: secret value used as a memory address } ``` The index and the address are the two halves of one effective address, so both are gated. Only a secret *pointer* is an address: a `^[N]T` or `^Rec` is a secret value living at a public address, and a `*^T` is a public address to secret storage. Three more gates decide against the target's constant-time capabilities rather than the source alone, so they are reported at lowering: a secret operand of a **floating-point** operation (always variable-latency, gated on every target); a secret operand of an **integer multiply** on a target without a trusted data-independent-timing mode, conservatively every ISA today; and a secret **variable shift count** on a target without a barrel shifter. A secret passed to a variadic pack is also rejected, including one wrapped inside an aggregate. > **Note:** The gates are checked against the types of the **instance**, not of the template. A generic's body is re-checked per instantiation under its concrete type arguments, so a `T` that instantiates to a secret is gated exactly as the secret spelled out in full would be. #### Asking about secrecy at comptime `$is_secret(T)` folds true when `T` is `^`-qualified **at the outermost level**. It is a type predicate like `$is_record` / `$is_union` / `$is_pointer`: comptime-only, valid as a `$if` / `$or` gate condition, and answered per instantiation inside a generic. The full reference is in [Intrinsics](https://machlang.org/docs/intrinsics.html#is-secret). It exists because secrecy was otherwise invisible to a library. Every other predicate asks about the shape *under* the `^` and so answers false for every secret, which makes `$is_record(^u64)` and `$is_record(u64)` the same answer - a reflection walk could only ever meet a secret field as a fallthrough it had to refuse. `$is_secret` is the positive question, and it is what lets a derive **decide** rather than refuse: a formatter redacts a secret field, a hash refuses one (a data-dependent fold is a leak in the shape of a digest), and an equality picks the constant-time comparison instead of the early-out whose timing *is* the secret. ```mach rec Session { id: u64; key: ^[32]u8; } $each f in $fields(Session) { $if ($is_secret(f.type)) { } # redact: no read of `key` is emitted $or { render(s.[f]); } } ``` **Outermost only**, the same line the rest of the family draws. `^*u8` is secret - the pointer *is* the secret, which is the shape the welded-storage rules exist for. `*^u8` is not: it is a public address to secret storage, and `$is_secret($pointee_of(f.type))` asks about the pointee. `[N]^u8` is not, and neither is a record with a secret field - the **field** is, and that is where a walk meets the question. There is deliberately no transitive "contains a secret anywhere" query: the per-field question is the one a walk actually has, and answering the transitive one in its place would make the common case wrong. > **Note:** A walk that skips a secret field is pinned by the flow rules rather than by convention - reading one into a public accumulator does not compile, so a walk that gates wrongly is a compile error, not a silent disclosure. #### Downgrade with `:>T` `:>T` is the only way to remove `^`. It produces a new public value of type `T` and never reinterprets storage in place. The destination type `T` must always be named explicitly. ```mach fun publish(a: ^u32) u32 { ret a:>u32; } ``` `:>T` peels exactly the outer qualifier, so it can never launder a welded pointee - `*^T` stays `*^T`. `::` and `:~` may neither add nor drop `^`. #### Welded-storage pointers Secrecy is fixed at declaration and is non-launderable, which makes the public/secret aliasing leak unconstructable with no alias analysis at all: - a `^T` is stored only through a `*^T`, never a `*T` - a secret-welded pointer cannot be erased to the untyped `ptr` - a `uni`'s overlapping variants must agree on secrecy ```mach fun erase(p: *^u8) ptr { ret p; } # error: cannot erase a secret pointer to ptr uni Bad { a: ^u32; b: u32; } # error: variants disagree on secrecy ``` The union rule is a property of the union **type**, not of the syntax that declared it, so it holds for an inline `uni { ... }` and at every *instance* of a generic union. At the declaration a variant typed by a generic parameter says nothing about secrecy, so the check happens where each instance is formed. ```mach uni U[T] { a: T; b: u32; } rec Box[T] { u: U[T]; } var s: U[^u32]; # error: this instantiation makes the variants disagree var b: Box[^u32]; # same error: the instance need not be spelled ``` Welding extends to **nested aggregates**: a `*Pair` whose `x` field is `^u64` has a welded field, so `?p.x` yields a `*^u64` and inherits the refusal. #### `#[oblivious]` - the codegen contract The flow typing constrains the *source*; `#[oblivious]` carries the obligation through *codegen*. Inside a function carrying it, the backend must not introduce a secret-dependent branch or select a variable-latency instruction on a secret operand. ```mach #[oblivious] fun ct_select(mask: ^u32, a: ^u32, b: ^u32) ^u32 { ret (a & mask) | (b & ~mask); } ``` A function instance that **computes** on a secret must carry it; one that only moves, stores, or declassifies secrets is transparent and stays annotation-free. The check runs per monomorphized instance. A secret-taint bit is threaded from sema's flow typing through IR and MIR to the emitted instruction stream, preserved across every value replacement, inline clone, instruction selection, and register-allocator copy. The one place taint stops is the declassify barrier a `:>T` cast lowers to. Secret-free code carries no taint and compiles byte-identically. A **translation validator** then re-derives the taint over the lowered MIR as a monotone dataflow fixpoint and independently re-checks the leakage conditions. The stages after that MIR are trusted, not re-checked: see [Where the check runs](https://machlang.org/docs/secrecy.html#assurance-scope). ##### Inline assembly inside an oblivious function Inline `asm` inside such a function is **validated**, not rejected. The block is parsed into instructions and walked for the same three leaks the compiler checks everywhere else. Taint enters through the block's `{name}` bindings, whose secrecy is stamped from the local's declared type. What the walk cannot model, it refuses: | Construct | Why it is refused | | --- | --- | | a body that does not parse | nothing to analyze | | a data directive (`.byte`, `.word`, `.long`, `.quad`) | its payload can encode any instruction | | a mnemonic the target has not classified | its timing behaviour is unknown | | a flags-conditioned branch (x86-64 `jcc`, aarch64 `b.`) | its condition rides the flags register, which the inline-asm effect model does not represent | That last row is a per-target asymmetry worth stating precisely. A branch whose condition is a **register operand** is visible to the walk and is checked: aarch64's `cbz` / `cbnz`, and every riscv64 branch, which compares two registers - RISC-V has no flags register at all. A branch whose condition rides the **flags register** cannot be checked, because a `cmp` of a secret before it would be invisible, so those are refused. `#[oblivious]` remains a **per-function** contract. A call out to a non-oblivious function is not validated - that is the boundary the decorator draws, not a hole in it. #### The zeroizing-write guarantee Wiping a secret is only useful if the wipe survives to run. That guarantee exists, but it is **not** provided by `#[oblivious]`, and it is scoped more broadly than the decorator is. A store into secret storage is tainted at lowering, keyed on the storage rather than on any decorator: ```mach # no decorator: the wipe is protected anyway fun clear(p: *^u8, n: usize) { var i: usize = 0; for (i < n) { p[i] = 0; i = i + 1; } } ``` **What it covers.** Memory reached through a pointer that escapes the function - the shape a `zeroize` helper has. Such a store cannot be promoted out of memory, and the taint is present for any future pass to read. **What it does not cover.** A value the compiler keeps in a **register**. Writing to a promoted local is not a memory write, so wiping one is not preserved: ```mach var x: ^u8 = k; x = 0; # NOT guaranteed: `x` may never have been in memory ``` Adding `#[oblivious]` does not change this. To wipe reliably, write through a pointer whose target is memory the compiler cannot promote away. #### Trusted base The only secret-to-public crossings are the explicit `:>T` cast and inline `asm` blocks. Everything else is enforced. A proof is always relative to a leakage model, and its fidelity to real silicon is empirical. The contract is only offered where mach emits the instructions that execute. A target whose back half hands a module to a downstream compiler instead - the SPIR-V backend - **rejects `#[oblivious]`**: neither that translation nor the device's timing behaviour is covered by the leakage model. Compile constant-time code for a machine target, and pass such a target only public data. #### Assurance > **Read this first:** The constant-time guarantee is incomplete. This support is an experimental preview and has not been audited. **Do not build production cryptography on it at this version.** What holds today: the type system checks that the source respects the leakage model, `#[oblivious]` carries the obligation through codegen, and the translation validator independently re-checks the lowered MIR. All three are static. Two host-executed measurements sit under them, and each only means anything on the machine it runs on, which is why both are unit tests rather than build checks. ##### The timing harness is a tool, not a gate `mach.lang.ct.probe` is a dudect-style harness in the compiler's tree. `mach test` runs it at deliberately tiny sample counts and prints its table. It asserts **nothing about the numbers**, not a threshold and not positivity, because even "a mean is positive" proved to be a property of the host's clock resolution rather than of the harness ([mach#3092](https://github.com/briar-systems/mach/issues/3092)). The run can only fail if the harness itself crashes. **No property of a measured time gates anything**, here or anywhere else in the suite. That is deliberate and it is not a gap. A dudect score is a statistic over wall-clock time on hardware nobody controls, so any threshold over it has a false-failure rate that belongs to the machine rather than to the code. A required check that can fail nondeterministically is worse than no check: it teaches a reader to re-run a red constant-time result until it turns green. An earlier form of this harness did assert on such thresholds, and failed a release gate on `x86_64-darwin` and then passed a re-run of the identical commit ([mach#3070](https://github.com/briar-systems/mach/issues/3070)). So the claims below are **measurements, taken deliberately**, not properties the test suite enforces. They are re-taken by raising the counts in `mach.lang.ct.probe` and reading the printed table, on a quiet machine, by a person who then reads the numbers. Performance and timing work belongs in [mach-bench](https://github.com/briar-systems/mach-bench), which is built for it, rather than in a correctness suite that must be deterministic. What such a measurement can establish is bounded. It works by refutation, because that is all a timing measurement can do: a clean score is consistent with a leak the instrument cannot see. So each mode carries a *planted* leak that must be detected, and a constant-time reference that must not separate the way the planted one does. Read the separation between the two class means, never `|t|`. Welch's statistic divides by the sample variance, and concurrent load inflates the variance without moving the means, so `|t|` collapses under load while the leak is plainly still there. Measured at load average 27, six consecutive runs of one binary gave `|t|` between 7.51 and 11.24 against a threshold of 10, while the mean ratio never fell below 3.4. A third probe leaks nothing at all and says whether a run counts. Any separation *it* shows is the machine rather than the code, so a run in which it is not flat has no discriminating power, and its verdict means nothing in either direction. A flat null is necessary and not sufficient: it must also be sampled at a comparable cost to the probe it is bounding, or a quiet control at one magnitude certifies nothing about noise at another. ##### The x86-64 flags table is measured, not inferred The inline-asm model `#[oblivious]` rests on classifies eighteen x86-64 instructions by which ones *define* ZF and CF, which merely write them, and which read them. That table is re-derived on x86-64 hosts by `mach.lang.target.isa.x64.probe`, which runs each instruction twice with RFLAGS preset all-set and all-clear and compares what the CPU did against the transcription. Defining the flags is the only fact that clears a taint, so a row that drifts from the silicon is a permission rather than a refusal. Three rows are exempt and classified by reasoning instead. `popfq`, `iretq` and `syscall` pass the writer probe cleanly and are still not definers, because their flags come from the stack, the interrupt frame, or an existing value masked through `IA32_FMASK`. Where a value came *from* is structural rather than measurable. A host that is not x86-64 declines the probe by name. ##### What a timing harness can and cannot assure The leakage model has three channels, and no single sampling regime covers them ([mach#2363](https://github.com/briar-systems/mach/issues/2363)): | Channel | Assured by | | --- | --- | | control-flow trace | the source-level branch gate. A latency-mode measurement can corroborate. | | variable-latency operands | the sema and lowering gates. A latency-mode measurement can corroborate. | | memory-address trace | the source-level **secret-index and secret-address gates**. Only an address-mode measurement can corroborate. | The two harness modes need opposite sampling and neither substitutes for the other. Latency mode times a large batch of calls per sample, which is what lifts a running-time difference above clock resolution. That same batching *hides* an address-trace leak: every call in a batch is handed the same input, so after the first call both input classes are reading a warm cache line and the single cache miss carrying the signal is averaged away. Measured on one function, a secret-indexed read over a table larger than the last-level cache, latency mode scored `|t|` of about 1 to 15 across runs, straddling its own threshold so it neither confirmed nor denied, while address mode, at one call per sample, scored in the hundreds. Two consequences worth stating plainly: - A clean latency-mode number for a table lookup is **not** evidence of address-trace safety. It is the wrong instrument for that channel. - An address-trace leak is only *measurable* when the table exceeds the last-level cache. A cache-resident table leaks its index just as truly and no timing harness will see it. For small tables the property rests entirely on the secret-index gate and on reading the emitted code. ##### Where the check runs The constant-time contract is checked at the **IR level**. The validator runs over the lowered, target-independent MIR before width legalization, instruction selection, register allocation, spilling, frame insertion, and encoding, and trusts those stages to be timing-preserving. What it refuses, it refuses closed: an operation it does not know, an out-of-range register reference, and inline assembly without a complete effect declaration are rejected, never defaulted to public. A proof over the final allocated machine program, after selection, allocation, spills and frame insertion, over physical registers and flags, is planned additive work ([mach#3591](https://github.com/briar-systems/mach/issues/3591)) and not something this version claims. Where mach does not own the later stages at all, as with a whole-module emitter such as SPIR-V, the contract is refused rather than assumed. #### See also - [Types](https://machlang.org/docs/types.html) - the compound type grammar `^` qualifies - [Decorators](https://machlang.org/docs/decorators.html#oblivious) - the `#[oblivious]` reference - [Inline assembly](https://machlang.org/docs/asm.html) - the blocks an oblivious function may contain - [Expressions](https://machlang.org/docs/expressions.html) - the `::` / `:~` casts that preserve secrecy Source: https://machlang.org/docs/asm.html ### Inline assembly Mach has one inline-assembly form: an ISA-tagged block of raw instructions with local-variable substitution. The compiler parses the instruction stream and infers operand direction and clobbers from the opcode semantics - no `in` / `out` declarations, no clobber list. #### Grammar ```mach asm { # raw instructions, one per line, # for comments mov rcx, {ptr} mov rax, [rcx] mov {result}, rax } ``` - The ISA tag is mandatory. Bare `asm { ... }` does not exist. - The tag comes from a closed set: `x86_64`, `aarch64`, `riscv64`, `riscv32`. Each has a working assembler that emits native bytes; the first three run in CI, riscv64 under qemu. - **The RISC-V tag names the machine, not the family.** An `asm riscv64` block is refused on a riscv32 target and the other way round, so a body reaching the assembler was written for the register width it is being assembled for. Gate the two with `$if ($mach.build.arch == $mach.arch.riscv32)` when a routine needs both. An RV64-only spelling - `ld`, `sd`, the `*W` group, the doubleword atomics - is refused under a `riscv32` tag, naming the instruction and the machine, because each instruction states the register widths it exists at and the assembler reads that rather than a second list of its own. The shift-amount field goes the same way: `slli a0, a1, 32` is refused on rv32, where its sixth bit would land in `funct7` and decode as a different instruction. - Each line is an instruction in the ISA's native syntax. - `#` introduces a line comment: everything from `#` to the end of the line is ignored, whatever it contains. The three ISAs share one statement grammar, one effect model, and one numeric-local-label scope; each supplies only its mnemonic table and encoder. #### Operand substitution `{name}` substitutes a local in scope. The compiler resolves the reference to a memory or register operand based on liveness and the instruction's expected operand class. ```mach pub fun add_via_asm(a: i64, b: i64) i64 { var result: i64 = 0; asm x86_64 { mov rax, {a} add rax, {b} mov {result}, rax } ret result; } ``` > **Note:** In practice a `{name}` binds the local's storage - typically a stack slot - so a pointer local's pointee is reached by staging the pointer through a scratch register first (`mov rcx, {ptr}` then `mov rax, [rcx]`), never by a direct `[{ptr}]` indirection. Only an identifier inside braces names a local. Any other braced text belongs to the ISA's own syntax, such as the aarch64 register list `{v0.16b}`, and reaches its grammar untouched. ##### A body that binds `{name}` may not move the stack pointer A `{name}` becomes a fixed displacement off a base register, measured once when the block is assembled. On aarch64 and riscv64 that base is the **stack pointer**, the only base whose displacement stays inside those ISAs' immediate forms however deep the frame is. A statement that moves the stack pointer would move every `{name}` in the block out from under its own address, so the compiler refuses the block rather than assembling a wrong one. ```mach var x: i64 = 0; asm aarch64 { ldr x9, {x} stp x1, x2, [sp, -16]! # refused: this body binds {x} } ``` The refusal covers the whole block, not just the statements after the write, because a backward branch reaches an earlier `{name}` again with the pointer already moved. Push and pop around the block instead, or drop the `{name}` and stage the address into a register yourself. A body that binds no `{name}` is unaffected, which is what a `#[naked]` function's hand-written prologue relies on. #### What the compiler infers - **Operand direction.** Position within an instruction determines whether an operand is read or written. - **Clobber set.** The compiler reads each instruction, knows what registers and flags it touches, and adds them to the surrounding function's clobber set. - **Memory clobber.** Every `asm` block is conservatively assumed to modify arbitrary memory. - **Vector clobbers.** A write to a vector register lands in the block's vector clobber set, kept apart from the general-purpose one. See [Vector registers](https://machlang.org/docs/asm.html#vector). - **Raw encodings.** A data directive is a stream the parser cannot read, so it clobbers every register in every bank unless it [declares what it writes](https://machlang.org/docs/asm.html#raw-effects). #### Calls and jumps (x86-64) `call` and `jmp` take the same three shapes, and which one a statement means is read off the operand. ```mach asm x86_64 { call some_symbol # direct: E8 rel32, relocated against the symbol call rax # indirect through a register call [0x100018] # indirect through an absolute address call [rax + 8] # indirect through a computed address jmp some_symbol # direct: E9 rel32 jmp rax # indirect through a register jmp [rax + 8] # ... and the same memory forms } ``` The absolute form exists for a fixed-address ABI - one whose entry points are addresses rather than symbols. Its displacement is sign-extended to 64 bits, so an address outside signed 32-bit range is refused rather than silently truncated. `call [symbol]` and `jmp [symbol]` are refused too: the rip-relative form would mean "transfer to the pointer *stored* at the symbol", which is not what the direct form beside it means. Both indirect operands are fixed 64-bit in long mode, so a narrower register (`jmp eax`) is refused rather than widened. An indirect call clobbers exactly as a direct one does. The register or memory holding the target is **read**, not written. #### Operand sizes (x86-64) A register operand states its own width, so `mov eax, [rcx]` is a four-byte load and needs nothing else. A memory operand states none, and where the instruction does not settle it either, the width is written out in nasm's spelling. ```mach asm x86_64 { movzx eax, word [rcx] # a two-byte load, zero-extended into eax movsx rax, dword [rcx] # a four-byte load, sign-extended (movsxd) mov dword [rcx], 1 # a four-byte store, not the machine word neg qword [rcx] # an eight-byte read-modify-write } ``` `byte`, `word`, `dword` and `qword` are accepted before a memory operand and nowhere else, so `mov qword rax, rcx` is refused rather than ignored. A vector instruction's memory operand takes `xmmword` instead (see [below](https://machlang.org/docs/asm.html#vector-x86-64)). **A prefix that contradicts the instruction is a build error, not a dropped token**, and what counts as a contradiction is per mnemonic: | Shape | Rule | | --- | --- | | most instructions | every operand shares one width, so a prefix must agree with any register operand, and with no register operand it *sets* the width | | `movzx` / `movsx` | the source is narrower by design, so a memory source **must** be sized, and strictly narrower than the destination | | `push` / `pop`, indirect `call` / `jmp` | fixed 64-bit in long mode, so any narrower prefix names no instruction | | `lidt` | its pseudo-descriptor is ten bytes, which no keyword names | So `mov eax, word [rcx]` is refused (two widths for one access), and so is `movzx eax, [rcx]`: an unsized source names no width, and reading it as a same-width move would silently assemble a plain `mov` where a zero-extending load was written. #### Vector registers The x86-64 and aarch64 grammars take vector registers as operands. A vector register belongs to the floating-point and vector bank, so writing one adds it to the block's **vector clobber set**, and a live vector value crossing the block is kept the same way a general-purpose one is. The allocator and the call conventions both see those writes. - A memory base or index is always a general-purpose register. A vector register there is refused by name. - A vector register in a general-purpose form is refused by name rather than encoded as something else. ##### x86-64 The registers are `xmm0` to `xmm15`. A memory operand of a vector instruction is a whole 128-bit vector, written bare or as `xmmword [...]`. A narrower width prefix is refused. `[symbol]` addresses RIP-relative data, and the relocation is correct after a trailing immediate. | Form | Mnemonics | | --- | --- | | move, either direction between a register and memory | `movdqa` `movdqu` `movaps` `movups` | | `xmm, xmm/m128` | integer arithmetic: `paddb` `paddw` `paddd` `paddq` `psubb` `psubw` `psubd` `psubq` `psubusb` `psubusw` `pmullw` logic: `pand` `por` `pxor` compares: `pcmpeqb` `pcmpeqw` `pcmpeqd` `pcmpgtb` `pcmpgtw` `pcmpgtd` interleave and saturate: `punpcklbw` `punpcklwd` `punpckldq` `packsswb` `packssdw` float arithmetic: `addps` `subps` `mulps` `divps` `addpd` `subpd` `mulpd` `divpd` conversions: `cvtdq2ps` `cvttps2dq` `cvtdq2pd` `cvttpd2dq` `cvtps2pd` `cvtpd2ps` | | `xmm, xmm/m128, imm8` | `pshufd` `cmpps` `cmppd` | ```mach asm x86_64 { movdqu xmm0, [rsi] # load 16 bytes movdqu xmm1, xmmword [rdx] # the same, with the width written out pxor xmm0, xmm1 pshufd xmm0, xmm0, 0x1b # reverse the four 32-bit lanes movdqu [rdi], xmm0 # store } ``` ##### aarch64 The registers are spelled `vN.16b`, `vN.8h`, `vN.4s` or `vN.2d`. The suffix is the lane arrangement, and every operand of one instruction shares it. `add`, `sub`, `and`, `orr`, `eor` and `mov` select their vector form when their operands are vector registers. | Form | Mnemonics | | --- | --- | | three registers | `add` `sub` `mul`, `and` `orr` `eor`, `cmeq` `cmgt` `cmge` `cmhi` `cmhs`, `fadd` `fsub` `fmul` `fdiv`, `fcmeq` `fcmgt` `fcmge` | | two registers | `mov` `mvn` | | one element structure | `ld1 {vT.}, [Xn]`, `st1 {vT.}, [Xn]` | Not every member exists at every arrangement: | Members | Arrangements | | --- | --- | | `and` `orr` `eor` `mov` `mvn` | `.16b` only | | the float members (`fadd` through `fcmge`) | `.4s` and `.2d` only | | `mul` | every arrangement except `.2d` | `ld1` and `st1` post-index their base, either by the structure size (`, 16`) or by an X register (`, x9`). Either form writes the base register. ```mach asm aarch64 { ld1 {v0.16b}, [x1], 16 # load 16 bytes, then x1 += 16 ld1 {v1.16b}, [x1], 16 eor v0.16b, v0.16b, v1.16b st1 {v0.16b}, [x2], x9 # store, then x2 += x9 } ``` The braces in `{v0.16b}` are aarch64's register-list syntax, not a substitution: only an identifier inside braces names a local. ##### Constant-time checking The multiply and float members are variable-latency operations, so the constant-time check in an [`#[oblivious]`](https://machlang.org/docs/secrecy.html#oblivious-asm) function treats them as it treats their scalar counterparts. The check tracks secrets per register bank, so a secret in `xmm3` is never mistaken for a secret in `rbx`, which shares its index. ##### Instruction-set availability The vector tables above hold only each architecture's baseline set: SSE2 on x86-64 and AdvSIMD on aarch64. An `asm` block may use any listed form on any target of its architecture, with no declaration. #### Privileged and system instructions The privileged and system instruction families are reachable from inline assembly on every target: the x86-64 systems instructions below, aarch64 `mrs` / `msr` and the exception conduits, and the RISC-V CSR family. ##### Systems instructions (x86-64) ```mach asm x86_64 { cli # clear the interrupt flag sti # ... and set it cld # clear DF before entering a program hlt # park the core in al, dx # port i/o, through dx or by immediate port out dx, al rdtsc # the time-stamp counter rdmsr # the model-specific registers wrmsr lidt [rax] # install an interrupt descriptor table pushfq # save RFLAGS popfq # ... and restore it swapgs # per-CPU state on a syscall entry iretq # return from an interrupt handler mov rax, cr2 # the faulting address in a page-fault handler mov cr3, rax # switch page tables mov eax, cs # the live selector, for programming STAR } ``` Control registers are `cr0`, `cr2`, `cr3`, `cr4` and `cr8`, and a control-register move takes a 64-bit general-purpose register on its other side. Segment registers (`es`, `cs`, `ss`, `ds`, `fs`, `gs`) can be **read** into a general-purpose register but not written through `mov`. None of these writes a register the allocator tracks except `rdtsc` and `rdmsr`, which land their result in EDX:EAX. `iretq` does not fall through, but the effect model cannot say so: statements after it are unreachable without the compiler reporting it. ##### System registers (aarch64) `mrs` and `msr` name a system register by its architectural name, in either case. ```mach asm aarch64 { mrs x0, cntvct_el0 # the virtual counter mrs x1, CNTFRQ_EL0 # ... and its frequency, capitalized as ARM spells it msr vbar_el1, x2 # install an exception vector base msr daifset, 0xf # mask every interrupt } ``` The named set covers what freestanding code reaches for. It is deliberately not exhaustive: **any** system register is also nameable by its encoding, exactly as ARM and GNU as spell it, which is what makes the surface complete rather than a list that always lags the architecture. ```mach asm aarch64 { mrs x0, s3_3_c14_c0_2 # the same register as `mrs x0, cntvct_el0` } ``` A field the architecture cannot hold is refused rather than truncated, because a truncated selector would name a *different* register than the text does. `msr , #imm` writes a PSTATE field (`daifset`, `daifclr`, `spsel`, `pan`, `uao`, `ssbs`, `dit`, `tco`); the architecture spells these by name only, so there is no numeric escape for that form. ##### Exception conduits and waits (aarch64) ```mach asm aarch64 { svc 0 # the kernel, at EL1 hvc 0 # the hypervisor, at EL2 smc 0 # the secure monitor, at EL3 wfi # wait for an interrupt: the correct idle loop wfe # wait for an event } ``` `hvc` and `smc` are how PSCI is reached, which is the only way to power off or restart a `virt` board. Both follow the SMC Calling Convention, so the compiler declares them as destroying **x0 to x17**. x18 to x30 and SP survive, which makes a conduit cheaper than an ordinary `bl`. The waits write nothing, so an idle loop holds every live value across them. `yield` is the weaker hint - it may do nothing at all. ##### Control-and-status registers (riscv64) The Zicsr extension's six instructions - read-write, read-set, and read-clear, each taking its source from a register or a five-bit immediate - reach a CSR by name. ```mach asm riscv64 { csrrw a0, mstatus, a1 # read mstatus into a0, write a1 into it csrr a0, mtvec # the read-only pseudo csrw stvec, a1 # install a trap vector rdtime a0 # the unprivileged counters } ``` The privileged spec defines several hundred addresses across three privilege levels, so **any** CSR is also reachable by its numeric address. RISC-V spells no separate escape syntax for this - a CSR operand simply parses as the ordinary integer literal it looks like, bounded to the twelve bits a CSR address occupies. ```mach asm riscv64 { csrr a0, 0xc01 # the same register as `csrr a0, time` } ``` > **Restriction:** Access permission is not checked, on either target: whether a register is readable or writable depends on the exception or privilege level the code runs at, which the compiler does not know. Accessing one the current level cannot reach traps at run time, as the architecture defines. #### Raw encodings Four data directives emit their values verbatim, for an encoding the ISA's mnemonic table does not name. They work on every target. ```mach asm x86_64 { .byte 0x0f, 0x01, 0xd0 # xgetbv } asm aarch64 { .word 0xd53be040 # mrs x0, cntvct_el0 } ``` The widths are GNU as's, per target - `.word` is the one that differs: | Directive | x86-64 | aarch64 / riscv64 | | --- | --- | --- | | `.byte` | 1 byte | 1 byte | | `.word` | **2 bytes** | **4 bytes** | | `.long` | 4 bytes | 4 bytes | | `.quad` | 8 bytes | 8 bytes | Values are written in the target's byte order, so `.word 0xd503201f` is the aarch64 `nop` as its manual prints it. A non-negative value is read unsigned (`.quad 0xFFFFFFFFFFFFFFFF` is a legal address), and a negative one is its two's complement at the directive's width. A value the width cannot hold is refused rather than truncated, and one directive carries a whole sequence - up to 256 payload bytes - not four. On aarch64 and riscv64 a statement must emit a whole number of instruction words (`.byte 0x1f, 0x20, 0x03` is refused there), and x86-64 has no such constraint. > **Cost:** A raw encoding is an instruction stream the parser cannot read, so by default the block's clobber set becomes **every register in every bank**, and an [`#[oblivious]`](https://machlang.org/docs/secrecy.html#oblivious-asm) function may not contain one at all. A real mnemonic is always preferable where one exists. ##### Declaring a raw encoding's effects That default is correct and expensive: with nothing held live across the statement, the allocator spills every value in a callee-saved register and the prologue saves every callee-saved register the function could reach. A raw encoding may instead state what it writes, with a `::` clause on the directive. ```mach asm x86_64 { mov dx, {port} .byte 0xee :: writes() # out dx, al - reads dx and al, writes nothing .byte 0x0f, 0x31 :: writes(rax, rdx) # rdtsc - lands its result in EDX:EAX } asm aarch64 { .word 0xd4000002 :: writes(x0, x1, x2, x3) # hvc #0, returning per SMCCC } asm riscv64 { .word 0xc0102573 :: writes(a0) # csrr a0, time } ``` `writes()` with an empty list is a **declaration** that the encoding writes no register. Leaving the clause off keeps the conservative default. Registers are named in the target's own vocabulary, in either bank: | Target | General-purpose | Float / vector | | --- | --- | --- | | x86-64 | `rax` to `r15` | `xmm0` to `xmm15` | | aarch64 | `x0` to `x30`, `sp`, `xzr` | `v0` to `v31` | | riscv64 | `x0` to `x31`, psABI aliases (`a0`, `t0`, `sp`, ...) | `f0` to `f31`, psABI aliases (`fa0`, `ft0`, ...) | A width is not a register: `eax`, `w0` and `v2.4s` are refused, because the allocator tracks each register's bits as a unit. A malformed clause, an unknown register, and an unknown attribute all fail the build. A clause may only follow a data directive, so it can only narrow the conservative default and never contradict an effect the compiler derived from a mnemonic. `writes(...)` is the only attribute today. > **Not checked:** The compiler does not decode the payload, so it cannot verify that the bytes write only what the clause says. **Under-declaring is a miscompile**: the allocator will keep a value in a register the encoding overwrites. Declare everything the instruction's manual says it writes, including implicit destinations. A declaration buys allocation quality, not verification, and an `#[oblivious]` function still refuses a data directive whether or not it is declared. #### Multi-arch dispatch Different architectures use different mnemonics, registers, and calling conventions. There is no nested arch-block construct inside `asm`; instead, wrap each block in `$if` on `$mach.build.arch`. ```mach $if ($mach.build.arch == $mach.arch.x86_64) { asm x86_64 { ... } } $or ($mach.build.arch == $mach.arch.aarch64) { asm aarch64 { ... } } ``` The discarded branches do not compile, so each `asm` block only needs to be valid for its tagged ISA. #### When to use it - Truly target-specific operations: syscalls, register reads, stack-frame surgery. - Anything that does not have a one-to-one standard-library wrapper. For operations that exist as named standard-library functions - atomics, fences, traps, the SIMD long tail - use the library API. Those wrappers already contain the arch-dispatched `asm`. #### See also - [Decorators](https://machlang.org/docs/decorators.html#naked) - `#[naked]`, whose body may hold only `asm` - [Secrecy](https://machlang.org/docs/secrecy.html#oblivious-asm) - how an `#[oblivious]` function validates an `asm` block - [Comptime control flow](https://machlang.org/docs/comptime-control.html) - `$if` over `$mach.build.arch` Source: https://machlang.org/docs/shaders.html ### GPU shaders A `spirv` target compiles ordinary mach into a finished SPIR-V module. A shader is written in the same language as the rest of the project - the same records, the same vectors, the same comptime - and a small set of decorators say which functions are pipeline stages and which module-scope variables the pipeline binds. > **Note:** Every directive on this page is accepted on **every** target, because which target a module is built for is not a property of its source. Only a target that forms pipeline stages acts on them: on a machine target a staged function is compiled normally and the interface variables are ordinary globals. #### The target A `spirv` target's object output is a complete, self-contained module rather than a link input, so the build delivers a **module tree** and runs no link phase. ```toml [target.gpu] isa = "spirv" os = "freestanding" abi = "spirv" # no `of`: the finished-module format resolves on its own ``` ``` mach build . --target gpu # writes out/gpu//obj/.spv ``` A `static` or `shared` artifact kind, and `mach test`, are refused by name: there is no archive, shared object, or executable form for a module. See [Manifest](https://machlang.org/docs/manifest.html#finished-modules). #### `stage(str)`: a pipeline stage Marks a function as the entry point of a graphics or compute pipeline stage. The value set is closed - `"vertex"`, `"fragment"`, `"compute"` - and an unrecognized value is a compile error, not a module that quietly forms no stage. ```mach #[stage("vertex")] fun vertex_main() { } #[stage("fragment")] fun fragment_main() { } ``` A staged function **takes no parameters and returns nothing**. A pipeline stage has no caller: its inputs arrive through input interface variables and its results leave through output ones, so there is no argument list or return value to carry them. A staged function with either is rejected. A module that declares any stage is a **shader module**, and that changes the whole artifact rather than just the one function. A shader module carries entry points and no external linkage at all; a module with no stage is a **library module**, which publishes each function as a linkage export so a consumer can find it. The two are exclusive - a Vulkan consumer refuses a module carrying linkage - so adding the first `#[stage(...)]` to a module stops it exporting its functions. The entry point's name, as a pipeline-creation call looks it up, is the function's **bare source name**. A shader module has no linker symbols to mangle. ##### `workgroup(x, y, z)` Sizes the workgroup of a `#[stage("compute")]` function. It requires a stage on the same function - without one it would silently mean nothing - and applies only to the compute stage. ```mach #[stage("compute")] #[workgroup(64, 1, 1)] fun compute_main() { } ``` Omitted, a compute stage takes the single-invocation default `(1, 1, 1)`. The dimensions are always declared in the emitted module, since a compute stage that does not state its workgroup size is not one a consumer can dispatch. #### The interface A stage reads and writes **module-scope variables** that the pipeline binds. These directives say which kind each variable is. They apply only to module-level `val` / `var` bindings, and a variable carries **exactly one** of them - they are mutually exclusive. ```mach #[input(0)] var in_position: f32x4; #[output(0)] var out_colour: f32x4; #[builtin("position")] var position: f32x4; rec Camera { view: f32x4; proj: f32x4; } #[uniform(0, 0)] var camera: Camera; rec Particles { pos: [64]f32x4; } #[storage(0, 1)] var particles: Particles; #[sampler(1, 0)] var albedo: Sampler2D; ``` `input` and `output` number a **varying** with a location, which is how one stage's outputs line up with the next stage's inputs: the producer's `#[output(0)]` feeds the consumer's `#[input(0)]`. ##### `builtin(str)` Names a value the pipeline supplies or consumes instead of one a location carries. The set is closed. | Value | Meaning | Type | Direction | | --- | --- | --- | --- | | `"position"` | clip-space vertex position | `f32x4` | written | | `"point_size"` | rasterized point size | `f32` | written | | `"vertex_index"` | index of the current vertex | `u32` | read | | `"instance_index"` | index of the current instance | `u32` | read | | `"frag_coord"` | fragment window coordinate | `f32x4` | read | | `"global_invocation"` | compute global invocation id | `u32x3` | read | | `"local_invocation"` | compute local invocation id | `u32x3` | read | | `"workgroup_id"` | compute workgroup id | `u32x3` | read | The direction is a property of the built-in, not something you restate - a stage writes its position and reads what the pipeline hands it - so there is no input/output marker to pair with `builtin`, and none that could disagree with it. The **type** is a property of the built-in too, and it is a requirement rather than a suggestion: the pipeline binds the variable itself, so a wider or narrower one is an invalid module rather than a wasteful one. Declaring a built-in at any other type is a compile error naming both the declared type and the required one. The two integer rows accept `i32` as well as `u32`, because the compiler carries an integer's width and not its sign and the emitted type is sign-less either way. ##### `uniform` and `storage` blocks `uniform` binds a read-only block by descriptor set and binding; `storage` binds a **read-write** buffer the same way. Both types **must be a `rec`**: a block has a host-visible layout, and a bare scalar or vector has no block layout for a pipeline to bind. Wrap a single value in a one-field record. Each is emitted with its `Block` decoration and an explicit byte offset on every member, taken from the same layout the rest of the compiler uses, so what the shader reads is what the host wrote. They differ in one place: their **layout rules**. A uniform block follows std140-shaped rules, under which an array's stride is rounded up to 16 - which mach's own layout does not do, so an array of anything narrower than 16 bytes is refused rather than silently repacked. A storage buffer follows std430-shaped rules, which use the element's natural stride, and that *is* mach's layout. So `[8]f32` is fine in a `storage` block and rejected in a `uniform` one. A compute stage's data path is `storage`: Vulkan forbids the `Output` storage class in a compute execution model, so a compute shader reads and writes buffers rather than varyings. ##### The `"readonly"` qualifier `storage` takes memory qualifiers after the descriptor pair. There is one, and it says that nothing writes the binding. ```mach rec Palette { columns: [512]f32x4; } #[storage(0, 3, "readonly")] var palette: Palette; ``` A store through a `"readonly"` binding is a compile error on every target, naming the line that wrote it. That is what the qualifier buys over what the compiler works out on its own: a buffer no body in the module stores through is emitted with the SPIR-V `NonWritable` decoration whether or not it is marked, and Vulkan reads that decoration to decide whether a stage needs `vertexPipelineStoresAndAtomics`. So an accidental write does not produce a wrong module, it produces a **correct one that quietly costs a hardware feature**. Marking the binding turns that into a diagnostic instead. The inference is one-sided on purpose: anything the compiler cannot follow, such as the binding's address handed to a function, counts as a write, so a missing decoration is possible and a wrong one is not. #### Textures and samplers `sampler` binds a [handle](https://machlang.org/docs/types.html#handles) by descriptor set and binding, at the same descriptor addressing `uniform` and `storage` use, so a host binds one the way it binds the others. Its type must be a handle type - a bodyless `def` carrying `#[handle]` - and a handle type must carry this decorator: a handle names a descriptor rather than an object with storage, so one with no descriptor address is reachable from no stage. Sampling is an `#[op(...)]` declaration rather than a language form, because a sample *is* one SPIR-V instruction, exactly as `sqrt` and `dot` are. ```mach #[op("spirv", "core", "OpImageSampleImplicitLod")] fun sample(s: Sampler2D, uv: f32x2) f32x4; #[stage("fragment")] fun frag_main() { out_colour = sample(albedo, in_uv); } ``` The separately-bound form works the same way, with the instruction that combines an image and a sampler declared alongside it. ```mach #[op("spirv", "core", "OpSampledImage")] fun combine(t: Texture2D, s: Sampler) Sampler2D; #[sampler(1, 0)] var base_tex: Texture2D; #[sampler(1, 1)] var base_smp: Sampler; #[stage("fragment")] fun frag_sep() { out_colour = sample(combine(base_tex, base_smp), in_uv); } ``` The combined value is handed straight to the sample rather than named: SPIR-V requires an `OpSampledImage` result be consumed in the block that produced it, which is the same rule that makes a handle-typed local a compile error. #### `op(target, set, name)`: a function that *is* an instruction A shader needs `sqrt`, `normalize`, `dot`, and `mix`. None of them is an operator, and none of them is a call SPIR-V can make: each is one instruction. This directive says which one a function is, so that on a `spirv` target a call to it becomes that instruction, inline, rather than a call. ```mach #[op("spirv", "GLSL.std.450", "Sqrt")] pub fun sqrt(x: f32) f32; #[op("spirv", "GLSL.std.450", "Normalize")] pub fun normalize(v: f32x4) f32x4; #[op("spirv", "core", "OpDot")] pub fun dot(a: f32x4, b: f32x4) f32; ``` The first argument names the **target** - the ISA name the manifest selects with - the second the instruction set, and the third the instruction within it. All three value sets are closed and checked at compile time on **every** target: the directive is legal everywhere, so a typo caught only where it is acted on would go unreported on a CPU build. The parameter count is checked against the instruction's own operand count, which is not uniform across a family that looks it: `Reflect` takes two operands where `Refract` takes three, and `FMin` two where `FClamp` takes three. | Set | Meaning | | --- | --- | | `"core"` | the core opcode space; needs no import | | `"GLSL.std.450"` | the standard extended set; imported once per module, on use | The substitution is uniform: the emitted instruction's **result type is the function's declared return type** and its **operands are the function's parameters in declaration order**. That is what lets `dot` and `length` return a scalar from vectors, and `refract` mix a scalar operand with vector ones, without any of them being a special case. > **Check the specification, not GLSL:** `dot` is **core `OpDot`**, not a GLSL.std.450 instruction, even though GLSL spells it beside `normalize` and `length`. Check each function against the SPIR-V specification rather than against GLSL's surface. On every target other than `spirv` the directive is inert and a decorated function is an ordinary function. A **bodyless** one - which is what the shader-side maths library uses - is then an undefined symbol, so a CPU build that calls it fails at link naming the symbol. That is the library's design choice, not a property of the directive: a decorated function may have a body, and if it does, that body is what every non-`spirv` target runs while `spirv` substitutes the instruction. #### Shaders across modules A shader can call a function defined in another module and can use an **interface variable** one declares, so a shared library may own a whole descriptor set. A drawn-in module contributes globals by **reachability, not membership**: a shared module may declare a whole set, and a shader importing it for one helper does not inherit the rest. The root is exempt - its interface list is the module's pipeline contract, declared whole. A repeated `(set, binding)` pair is checked over the set one entry point reaches, **across** the descriptor roles rather than within one, since a `#[uniform(0, 0)]` and a `#[storage(0, 0)]` are two descriptor types claiming one slot. Within one module a repeat is a typo its author can see; across modules it is neither, and the shader importing both is where the conflict first exists. `spirv-val` does not catch it, because two variables at one pair are well-formed SPIR-V and the conflict is with the pipeline layout, which is not in the module. The diagnostic names the declaring **module** when it is not the one being compiled. #### What a stage may not do - **Recursion** and a signature past 15 parameters are refused rather than emitted. - **`#[oblivious]`** is rejected outright on this target: the back half emits a module for a downstream compiler rather than the executed instructions, so the constant-time contract cannot be validated or upheld there. See [Secrecy](https://machlang.org/docs/secrecy.html). - **`#[packed]`** cannot apply to a `#[uniform]` or `#[storage]` block, whose member offsets the layout rules fix. - A **handle** cannot be a local binding, a record or union field, or sit behind a pointer or inside an array. #### See also - [Types](https://machlang.org/docs/types.html#handles) - handle types, and the SIMD vectors a stage computes over - [Decorators](https://machlang.org/docs/decorators.html) - the full directive set and its applicability matrix - [Manifest](https://machlang.org/docs/manifest.html#finished-modules) - declaring a `spirv` target and what its build delivers Source: https://machlang.org/docs/comptime.html ### The comptime channel The `$` prefix opens the comptime channel, the compiler-owned namespace a program reads at compile time. It is read-only: `$` selects and expands, it never executes or mutates. Conditional compilation, intrinsics, and target queries all ride one of its shapes. #### The shapes The parser tells the shapes apart by structure, so a name never has to be looked up to know what kind of read it is. | Shape | Meaning | | --- | --- | | `$mach.*` / `$project.*` / `$bin.*` | rooted compiler-owned read (compiler to developer) | | `$sym(args)` | comptime function call (an intrinsic) | | `$if`, `$or` | comptime control flow | - `$.` reads into a compiler-owned tree. The roots `mach`, `project`, and `bin` are reserved at the top of `$`; user symbols cannot collide with them. - `$ident(args)` is a comptime call; the closed compiler-intrinsic set lives here. - `$if` / `$or` are comptime branches, structurally distinct from runtime `if` / `or`. > **Note:** Per-declaration codegen attributes (symbol rename, library pin, inline, align, section) are written as `#[...]` decorators, not `$`-comptime shapes. See [Decorators](https://machlang.org/docs/decorators.html). #### Compiler-owned roots Three rooted trees carry the read-only facts a build exposes. `$mach.*` (the active build and compiler identity) has its own section below; `$project.*` and `$bin.*` read project metadata and the current artifact, one member at a time: | Read | Value | | --- | --- | | `$project.id` | project id, from `[project]` | | `$project.version` | whole version string, e.g. `"2.0.0"` | | `$project.version.major` | major version component (folded integer) | | `$project.version.minor` | minor version component | | `$project.version.patch` | patch version component | | `$project.target.os` | selected target os string, e.g. `"linux"` | | `$project.target.arch` | selected target arch string, e.g. `"x86_64"` | | `$project.target.abi` | selected target abi string, e.g. `"sysv64"` | | `$bin.name` | the artifact (build unit) being built | The `$project.target.*` reads carry the manifest declared string spellings, distinct from the numeric tags used for `$mach.build.*` comparison. The flat `$project.version` string and its structured `major`, `minor`, and `patch` components are all available. ```mach val ver: str = $project.version; # "2.0.0", from [project].version $if ($project.target.os == "windows") { ... } # the declared os string ``` > **Note:** A path the root does not carry, such as `$project.name` or `$project.description`, is rejected as an unknown `$project.*` path. #### The `$mach.*` namespace The `$mach.*` subtree is the compiler view of the world: the resolved build context, the compiler identity, and source position. Every read is a comptime constant. > **Implementation status:** Live today: `$mach.build.os`, `$mach.build.arch`, `$mach.build.abi`, `$mach.build.pointer_width`, `$mach.build.mode`, `$mach.build.pie`, `$mach.build.platform`, the tag tables, `$mach.version` (with its components), `$mach.compiler.name`, and `$mach.compiler.version`. Reserved stubs: `$mach.build.timestamp`, `$mach.build.host`, `$mach.build.git.*`, `$mach.project.*`, and `$mach.source.*`. Reading a stub is a compile error ("not yet available"). ##### `$mach.build.*` - what we build for The resolved active build facts. `os`, `arch`, `abi`, and `mode` share the numeric tag space of `$mach.{os,arch,abi,mode}.*`, so a comparison is a plain integer compare. ```mach $mach.build.os # compared against $mach.os.* tags $mach.build.arch # compared against $mach.arch.* tags $mach.build.abi # compared against $mach.abi.* tags $mach.build.pointer_width # integer count of bytes $mach.build.mode # compared against $mach.mode.* tags $mach.build.pie # 1 when building position-independent, else 0 $mach.build.platform # target platform tag as a string, "" when unset ``` The members above are the whole subtree, and no manifest key adds custom defines. An unknown `$mach.build.` is a compile error at the use site. Project configuration constants are ordinary `val`s selected with `$if` over these build facts. ##### Version and compiler identity ```mach $mach.version # the version string, e.g. "2.0.0" $mach.version.major # integer component (minor, patch too) $mach.compiler.name # compiler identity $mach.compiler.version # same value as $mach.version ``` ##### Tag tables The `$mach.{os,arch,abi,mode}.*` tables exist for path-value comparison against the resolved build. The lists are closed and grow only when backend support lands; an unrecognized tag name is a compile error, never a silent fold. ```mach $mach.os.linux $mach.os.darwin $mach.os.windows $mach.os.freestanding $mach.arch.x86_64 $mach.arch.aarch64 $mach.arch.riscv64 $mach.arch.riscv32 $mach.abi.sysv64 $mach.abi.win64 $mach.abi.aapcs64 $mach.abi.lp64 $mach.abi.lp64f $mach.abi.lp64d $mach.abi.ilp32 $mach.abi.ilp32f $mach.abi.ilp32d $mach.mode.debug $mach.mode.release ``` #### Comparing tags Tag comparisons are path-value, with no `.id` suffix or unwrapping. Both sides share one numeric space, so the comparison is an ordinary integer compare. ```mach $if ($mach.build.os == $mach.os.linux) { ... } $if ($mach.build.arch == $mach.arch.x86_64) { ... } ``` #### Folding into runtime values A `$mach.*` read can initialize a runtime binding; the compiler folds the right-hand side at compile time. ```mach pub val IS_LINUX: u8 = $mach.build.os == $mach.os.linux; pub val COMPILER: *u8 = $mach.compiler.name; ``` #### The bare `$ident` is rejected A standalone bare `$ident` such as `$mode` or `$foo` is none of the shapes above. A comptime parameter is referenced by its bare name (no `$`); every comptime path is rooted. The rule applies identically in a `$if` gate and in value position, with sema and lowering deferring to the evaluator's single verdict. > **Diagnostic:** comptime parameters are referenced without `$`; comptime paths are rooted: `$mach`, `$project`, `$bin` #### Not in the channel - No reflection via a `$.*` subtree. Types are not first-class comptime values. - No decl-attached prefix sugar; `$inline pub fun ...` does not exist. Use `#[...]` decorators instead. - No comptime function definitions, and no comptime loops. - No bare `$ident`. #### See also - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$size_of`, `$assert`, and the comptime call set - [Control flow](https://machlang.org/docs/comptime-control.html) - `$if` / `$or` gates over these reads - [Decorators](https://machlang.org/docs/decorators.html) - `#[...]` codegen attributes - [Manifest](https://machlang.org/docs/manifest.html) - the `[project]` and target stanzas feeding `$project.*` Source: https://machlang.org/docs/intrinsics.html ### Intrinsics Intrinsics are compiler-shipped comptime functions. They share the ordinary `$name(args)` call shape, but their names are reserved and their bodies live in the compiler. The set is closed: adding one requires a compiler change. #### Value intrinsics These fold to a comptime constant `u64`. Mach has no implicit widening or narrowing, so a binding of another integer type takes an explicit `::` cast. ```mach $size_of(T) # byte size of type T $length_of(T) # ELEMENT count of type T $align_of(T) # byte alignment of type T $offset_of(T, field) # byte offset of T's field ``` ```mach pub val POINT_SIZE: u64 = $size_of(Point); pub val POINT_X: i64 = $offset_of(Point, x)::i64; ``` `T` is a **type**, written with the ordinary type grammar - not just a bare name. A generic instance, a pointer, an array, a `^` secret, and a qualified `module.Type` are all valid, including inside the generic that owns the parameter. `$offset_of`'s second argument is the exception: a bare field name, resolved against the record's layout, never a type or a value. ##### `$length_of`: elements, not bytes `$size_of` counts bytes and `$length_of` counts elements. For `[N]u8` the two answers are equal; for everything else they are not, and the difference is deliberately in the surface rather than in the caller's head - making a caller divide by an element size is exactly the silent-arithmetic error the `u8` case hides during development. ```mach $length_of(PIXELS) # 400 - elements, for a [400]f32x4 $size_of(PIXELS) # 6400 - bytes $length_of(f32x4) # 4 - a vector's lane count ``` **Only a fixed array and a vector have an answer**, and everything else is refused rather than guessed: a pointer (`str` included) has a length the compiler does not know, and a record has a field count rather than an element count. The refusal names the type it was handed. ##### The operand may be a binding Every intrinsic taking a type operand also accepts a **value binding** there, denoting that binding's type. A binding's type has no spelling - `val LOGO: [_]u8` really is a concrete `[7194]u8` that cannot be written - so without this a program could index an embed and pass it around and never learn its length. ```mach #[embed("assets/logo.qoi")] val LOGO: [_]u8; $length_of(LOGO) # 7194 - elements $size_of(LOGO) # 7194 - bytes ``` This is the operand slot's rule, not a per-intrinsic one, and it is **only** that slot. A name written where a type is expected and resolving to a value is still a mistake in every other position. ##### Where a layout intrinsic folds `$size_of`, `$length_of`, and `$align_of` fold in every **type** position, including ones resolved before layout would otherwise be known: the measured type's layout is established on demand when the measurement asks for it, so where the type is *declared* relative to where it is measured makes no difference. | Position | `$size_of` / `$length_of` / `$align_of` | | --- | --- | | `val` / `var` initializer | yes | | global `align` | yes | | record / union type `align` | yes | | array length `[N]T` | yes | | `$if` / `$or` condition, in a function body | yes | | `$if` / `$or` condition, in declaration scope | only when no arm of the chain declares anything | ```mach rec Pair { a: u64; b: u64; } #[align($align_of(Pair))] # a type's alignment rec Over { x: u8; } rec Holder { buf: [$size_of(Pair)]u8; } # an array length, inside a field type ``` A layout relationship is the thing most worth asserting and it now has a home: `$if ($size_of(A) != $size_of(B)) { $error("..."); }` fails the build that introduces the divergence, rather than a runtime test that fires only if someone runs the suite. `$offset_of` is the exception among the four: it folds in a value position but not in a type one, because a field offset is settled during lowering rather than by the front end. Which times a declaration-scope gate may measure at is covered in [Control flow](https://machlang.org/docs/comptime-control.html#decl-scope). **Cycles are refused, not resolved.** A measurement whose answer is one of its own inputs - `#[align($size_of(Self))]`, or two types each aligned to the other's size - is reported as a layout cycle naming the type that closes it. A pointer field does not create one: it stores an address of fixed width, so `rec Node { next: *Node; }` measures normally. #### Type intrinsic `$type_of(expr)` produces a comptime type value: the resolved type of its argument. Type values have no runtime representation; they are meaningful only as operands in comptime type comparisons. ```mach $type_of(expr) # comptime type value of expr ``` Compare type values with `==` or `!=` inside a `$if` condition. A bare type name (`i64`, `str`, `Point`) is the other valid operand. The comparison selects one branch at compile time per monomorphization instance - useful for per-element type dispatch inside `$each` bodies. ```mach $if ($type_of(arg) == i64) { write_i64(w, arg); } $or ($type_of(arg) == str) { write_str(w, arg); } $or { $error("unsupported type"); } ``` > **Note:** Provably-dead arms are pruned before type-checking, so each arm uses `arg` at its own concrete type with no per-arm cast: the `str` arm above is never checked against a `u64` element. Only the selected arm is type-checked and emitted. #### Type predicates Where `$type_of` asks what a type *is*, the predicates ask about its **shape**. Each takes one type operand and folds in a `$if` / `$or` gate. ```mach $is_record(T) # T is a record (or an instance of one) $is_union(T) # T is a union (or an instance of one) $is_tag(T) # T is a tag (or an instance of one) $is_pointer(T) # T is a reference: the raw `ptr` or a typed `*U` $is_secret(T) # T is `^`-qualified at the outermost level ``` They are **comptime-only**. A gate condition selects an arm, and there is no runtime boolean for one to become, so using a predicate as a value is an error. **`^` is a constructor, and a predicate answers about the outermost one.** `^Pair` is a secret, not a record, so the three shape predicates answer false and a reflection walk refuses it instead of descending into secret storage. That is what keeps a predicate and `$fields` in agreement: `$fields(^Pair)` refuses, so a gate that called `^Pair` a record would send a walk into an operand the intrinsic then rejects. Outermost means outermost. `^*u8` is a secret pointer and `$is_pointer` answers false; `*^u8` is a **public** pointer to secret storage and is still a pointer, since the address is public. A generic instance answers as the declaration it instantiates, so `Box[i64]` is a record - which is the type a reflection loop actually meets - and `Box[^u64]` is a record too, because the instance is not itself secret; its field is, and the field is where a walk meets the question. Inside a generic, a predicate is answered **per instantiation**: `$is_record(T)` in a `fun f[T]()` body is not decided against the template's placeholder, so `f[SomeRecord]` and `f[u64]` take different arms from one template. ##### `$is_secret` `$is_secret` is the family's fourth member and its one exception, and the exception is coherent rather than special-cased: the other three ask about the shape *under* the wrapper, this one asks about the wrapper itself. It is what makes the other three's false readable - `$is_record(^Pair)` and `$is_record(u64)` are otherwise the same answer, so without it a library could only ever meet a secret as a fallthrough it had to refuse. ```mach $each f in $fields(T) { $if ($is_secret(f.type)) { ... } # redact, refuse, or compare in constant time $or { ... } # an ordinary public field } ``` | Operand | `$is_secret` | Why | | --- | --- | --- | | `^u64`, `^Pair`, `^[4]u8` | true | `^` is outermost | | `^*u8` | true | a secret **pointer**: the address is the secret | | `^^T` | true | `^^T` collapses to `^T` | | `*^u8` | false | a **public** pointer to secret storage | | `[4]^u8` | false | a public array of secret elements | | `rec S { k: ^u64; }` | false | the record is public, its **field** is secret | | `u64`, `Pair`, `ptr` | false | no `^` anywhere | **It is not transitive, deliberately.** "Does this contain a secret anywhere" is a different question, and folding the two together would make the common case answer wrong: a `fmt` derive gating on a transitive answer would redact a whole record over one field, and could not tell which field to redact. Where the transitive question is genuinely wanted through a reference it is *composed*: `$is_secret($pointee_of(f.type))` says whether a pointer field points at secret storage. #### `$pointee_of(T)`: descend through a reference `$is_pointer` tells a walk that a field is a reference. `$pointee_of` says what it refers to, which is what makes the reference traversable rather than merely detectable. ```mach $pointee_of(*U) # U $pointee_of(**U) # *U - one level, not all of them ``` It is a type **constructor**, in the same family as `*`, `[N]`, and `^`, not a call that returns a value. So it is written wherever a type is written, including nested inside another intrinsic's operand and inside a generic argument list. ```mach rec Node { value: i64; next: *Inner; } $each f in $fields(Node) { $if ($is_pointer(f.type)) { $each g in $fields($pointee_of(f.type)) { # gate, then descend total = total + (@(n.[f])).[g]; } } $or { total = total + n.[f]; } } ``` Because `str` is `def str: *char`, a `str` field is a reference field, and `$pointee_of(str)` is `u8` - which is what a formatter rendering a `str` field needs. **Everything that is not a typed reference is refused, and the refusal names what it was handed**, because a plausible wrong type here flows into a `$fields` walk that then reports about the wrong record. `ptr` is refused (the raw pointer is untyped and carries no pointee) and so is `^*U` (a `^` secret is not a reference). `^` is **not** stripped, which puts `$pointee_of` with the predicates rather than with `$size_of`: `$is_pointer(^*U)` answers false, so descending through `^*U` would hand a walk the secret storage the gate refused it. > **Termination:** Following references does not terminate structurally. A record cannot contain itself by value, so a `$is_record` descent reaches a finite set of types; a reference graph has no such property. `rec Grow[T] { p: *Grow[*T]; n: i64; }` is legal and has unboundedly many instances. The compiler's generic-instantiation guard turns that into a diagnostic naming the derivation chain rather than a hang, but that is a backstop, not a termination story - a library that walks references owes its callers one of its own. #### Where `^` is stripped One rule covers the whole surface: **`^` is stripped only where the question is about storage.** | Asks about | Strips `^` | | --- | --- | | `$size_of` / `$length_of` / `$align_of` / `$offset_of` / `$discriminant_of` | yes - a secret occupies its base type's storage | | `$is_record` / `$is_union` / `$is_tag` / `$is_pointer` | no - `^T` is a secret, not a `T` | | `$is_secret` | no - and it is the one query *about* the `^` | | `$pointee_of` | no - `^*U` is a secret, and is refused rather than followed | | `$type_name` | no - the spelling is `^T` | | `$fields` / `$cases` | no - a secret aggregate is refused, not walked | | type comparison (`f.type == u64`) | no - `^u64` is not `u64` | #### Recursive reflection A `$fields` walk can ask each field whether to descend into it. ```mach rec Inner { x: u64; y: u64; } rec Outer { i: Inner; n: u64; } $each f in $fields(Outer) { $if ($is_record(f.type)) { $each g in $fields(f.type) { ... } # descend } $or { ... } # a scalar field } ``` #### `$type_name(T)` A type's spelling, as a NUL-terminated string. ```mach $type_name(T) # the type's spelling, as a *u8 ``` Unlike the predicates this **is** a value (`*u8`), usable anywhere one is. The spelling is the same one diagnostics print, so a name a program reads and a name an error reports cannot drift. Composites spell compositely - `$type_name(*Pair)` is `"*Pair"` - and `^` spells too: `$type_name(^Pair)` is `"^Pair"`. Stripping it would be a drift on the one qualifier where a drift matters most, since a diagnostic about that type prints `^Pair`. #### The type operand `$size_of`, `$length_of`, `$align_of`, `$offset_of`, `$fields`, and the queries above all take a **type** in argument 0, written with the ordinary type grammar - plus one extra form: a field descriptor's `f.type` inside a `$each` body. `$pointee_of` is part of that grammar rather than one of its consumers, so it composes with every one of them. ```mach $each f in $fields(T) { val n: u64 = $size_of(f.type); # the field's own size $each g in $fields(f.type) { ... } # its own fields } ``` Inside the loop a field's type has no spelling, only the descriptor. A path that is genuinely a qualified type name (`mod.Type`) still reads as one. The same form is valid in a **generic argument list**, which is what makes a walk recursive rather than merely descending. ```mach fun eq[T](a: *T, b: *T) bool { $each f in $fields(T) { $if ($is_record(f.type)) { if (!eq[f.type](?a.[f], ?b.[f])) { ret false; } # re-enter at the field's type } $or { if (a.[f] != b.[f]) { ret false; } } } ret true; } ``` Without it, `$fields(f.type)` gives one level of descent per `$each` someone wrote, so a walk reaches only as deep as its author hand-unrolled. With it the walk is written once and reaches any depth. **Termination is structural and needs no depth limit**: each descent instantiates at a field's own type, a record's fields are finite, and a record cannot contain itself by value - a self-reference must go through a pointer, which `$is_record` does not select. #### Field intrinsic and projection `$fields(T)` produces a comptime sequence of field descriptors for record type `T`, written with the full type grammar exactly as the layout intrinsics take it (`$fields(Box[T])`, `$fields(mod.Rec)`). A **union is refused**: its variants overlap in storage, so a member walk over them would report distinct fields at distinct offsets that do not exist. Each descriptor carries three readable properties. | Property | Type | Value | | --- | --- | --- | | `f.name` | `*u8` | field name as a NUL-terminated string | | `f.type` | type value | comptime type value of the field's type | | `f.offset` | integer | byte offset of the field in `T`'s layout | The sequence is consumed by `$each f in $fields(T)`. Inside the body, `v.[f]` projects the concrete field off an instance `v`. It is an lvalue: readable and writable, including through a pointer receiver. ```mach $fields(T) # comptime field sequence for record T v.[f] # comptime field projection: access the field f on v ``` ```mach rec Pair { x: i64; y: i64; } fun sum(p: Pair) i64 { var total: i64 = 0; $each f in $fields(Pair) { total = total + p.[f]; # p.x on iteration 1, p.y on iteration 2 } ret total; } ``` `$each f in $fields(Empty)` expands to nothing when `T` has no fields. ##### Heterogeneous fields Because each iteration re-types `v.[f]` to the concrete field type, heterogeneous records work naturally - cast each field as you fold it. ```mach rec Mixed { a: i64; b: u8; } fun total(m: Mixed) i64 { var t: i64 = 0; $each f in $fields(Mixed) { t = t + m.[f]::i64; # m.a (i64) on iter 1, m.b (u8) cast to i64 on iter 2 } ret t; } ``` ##### Descriptor reads A descriptor's properties can be read inside the loop body - `f.offset` for layout math, `f.type` for type comparisons. ```mach fun offsum(m: Mixed) i64 { var s: i64 = 0; $each f in $fields(Mixed) { s = s + f.offset::i64; # 0 + 8 = 8 for Mixed { a: i64; b: u8; } } ret s; } fun count_i64(m: Mixed) i64 { var n: i64 = 0; $each f in $fields(Mixed) { $if (f.type == i64) { n = n + 1; } $or { } } ret n; } ``` > **Note:** A field literally named `type` is unaffected: ordinary `v.type` access still works. The `v.[f]` projection uses the `$each` loop variable, which is always a field descriptor, never a regular member. ##### Nested $each `$each` can be nested to walk a record's fields against another's. ```mach fun cross(p: Pair, q: Pair) i64 { var t: i64 = 0; $each f in $fields(Pair) { $each g in $fields(Pair) { t = t + p.[f] * q.[g]; } } ret t; } ``` #### Tag reflection: `$cases` and `$discriminant_of` `$cases(T)` produces a comptime sequence of case descriptors for tag type `T`. Each descriptor carries five readable properties. | Property | Type | Value | | --- | --- | --- | | `c.name` | `*u8` | case name as a NUL-terminated string | | `c.has_payload` | `bool` | whether the case takes a payload | | `c.type` | type value | comptime type value of the payload when present | | `c.offset` | integer | byte offset of the payload in `T`'s layout | | `c.code` | integer | discriminant value assigned to this case | The sequence is consumed by `$each c in $cases(T)`. Inside the body, `sel v.[c]` tests whether an instance holds that case, `v.[c]` accesses the payload under an active guard, and `T.[c]{...}` constructs an instance for case `c`. ```mach fun code_of[T](v: T) u64 { $each c in $cases(T) { if (sel v.[c]) { ret c.code; } } ret 0; } ``` ##### `$discriminant_of(T)` `$discriminant_of(T)` is a comptime type constructor that returns the unsigned integer type declared for tag `T`'s discriminator. It can be used anywhere a type is expected. ```mach val disc_size: u64 = $size_of($discriminant_of(Shape)); ``` #### $each: compile-time unroll `$each` is a statement form that splices its body once per element of a comptime sequence. There are four sequence forms. ```mach $each f in $fields(T) { ... } # one iteration per field of T $each c in $cases(T) { ... } # one iteration per case of tag T $each a in va { ... } # one iteration per element of pack va $each x in ARR { ... } # one iteration per element of a constant array val ``` It is valid only in statement scope, inside a function body. It is not a loop: the body is duplicated at compile time, not iterated at runtime. Enclosing runtime variables (an index, an accumulator) are shared across all unrolled copies. See [Variadic packs](https://machlang.org/docs/variadics.html) for the pack form. ##### $each over a comptime-constant array `$each x in ARR` unrolls the body once per element of `ARR`, binding `x` to that element's compile-time constant. Unlike the pack and `$fields` forms, every element shares one type - the array's element type - so the loop variable is an ordinary constant value: it reads as a value, casts, dispatches a per-element `$if`, and for a record element projects fields with `x.field`. ```mach val PRIMES: [4]i64 = [4]i64{2, 3, 5, 7}; fun sum() i64 { var total: i64 = 0; $each x in PRIMES { total = total + x; # x is 2, then 3, then 5, then 7 } ret total; # 17 } ``` A per-element `$if` selects its arm from the element's constant, so heterogeneous handling falls out of the unroll - including a function-pointer field, which folds to the element's function. ```mach rec Rule { tag: i64; fn: fun(i64) i64; } val RULES: [3]Rule = [3]Rule{ Rule{tag: 1, fn: inc}, Rule{tag: 2, fn: dbl}, Rule{tag: 3, fn: neg}, }; fun run(n: i64) { $each r in RULES { $if (r.tag == 2) { use_double(r.fn(n)); } $or { use_other(r.tag, r.fn(n)); } } } ``` **Eligibility.** `ARR` must name an immutable `val` (never a `var`) declared in the current module, whose type is a fixed-size array `[N]E` fully initialized by an array literal of exactly `N` elements. `E` must be a scalar or record type; nested-array element types are not supported. An empty array unrolls to nothing, and each violation is reported with a teaching diagnostic. Projection is one level deep (`x.field`); `x` itself is a constant and has no address, so `?x` is rejected. #### Diagnostic intrinsics `$error("msg")` fails compilation with `msg` when it is reached on a live path: an unconditional position, or a `$if` / `$or` arm the compiler selects. A `$error` in a discarded arm never fires, so it is the natural total-coverage fallback for a `$type_of` dispatch: the unhandled-type `$or {}` arm fails the build at compile time instead of falling through to a runtime error. It is valid in both declaration and statement scope and takes one string-literal message. ```mach $error("msg") # fails compilation when reached $if (!supported) { $error("this target is not supported"); } $if ($type_of(arg) == i64) { write_i64(w, arg); } $or ($type_of(arg) == str) { write_str(w, arg); } $or { $error("no writer for this argument type"); } # compile error on an unhandled type ``` #### Not provided as intrinsics Code intrinsics - runtime-instruction emitters like `trap`, `fence`, and `pause` - are not in the compiler-shipped set. They belong in the standard library as functions with per-arch `asm` bodies. **`$assert` is not an intrinsic either**, and is not planned as one: `$if` and `$error` already compose to it exactly, so a dedicated directive would add spelling without adding capability. Write the composition directly. ```mach # instead of $assert(cond, "msg") $if (!cond) { $error("msg"); } $if (!($mach.build.arch == $mach.arch.x86_64)) { $error("expected x86_64"); } ``` The composition inherits `$if`'s condition rules, which is the point: the same conditions fold there as in any other gate, and the ones that do not refuse with their own cause rather than through a second surface that could describe them differently. A chain written this way declares nothing, so in declaration scope it is decided during type checking and can measure a type. #### See also - [The comptime channel](https://machlang.org/docs/comptime.html) - what runs at compile time and how it folds - [Control flow](https://machlang.org/docs/comptime-control.html) - `$if` and `$or` branch selection - [Variadic packs](https://machlang.org/docs/variadics.html) - `$each a in va` over argument sequences Source: https://machlang.org/docs/comptime-control.html ### Control flow `$if` and `$or` branch on comptime-evaluable conditions. Only the taken arm is compiled; the discarded arms are never resolved, type-checked, or emitted - unlike runtime `if` / `or`, which always generates a branch. #### The `$if` / `$or` form A `$if` chain mirrors runtime `if` / `or` in shape, but each gate is a comptime condition and only the selected branch reaches the binary. `$or` with a condition is an else-if; `$or` with no condition is the comptime else. ```mach $if (cond) { ... } $or (cond) { ... } $or { ... # comptime else } ``` Target gating is the most common use: select the right import or backend per build, and leave the rest out of the binary entirely. ```mach $if ($mach.build.os == $mach.os.linux) { use full.os.linux; } $or ($mach.build.os == $mach.os.windows) { use full.os.windows; } $or { $error("unsupported OS"); } ``` #### Comptime conditions The gate must be a comptime expression. Common shapes: - `$mach.*` reads for target and build conditions - comparisons of comptime constants (`pub val` declarations) - comparisons of a comptime function parameter - see below A comptime comparison or arithmetic relates the mathematical values of its operands, exactly as the runtime operators do. A constant in the range 2^63 to 2^64 - 1 carries its true unsigned magnitude, so `$if (0xFFFFFFFFFFFFFFFF > 0)` is taken, and any cross-sign comparison agrees with the runtime form: `$if (X < Y)` never selects a branch that `if (X < Y)` would not. > **Note:** Comptime arithmetic that overflows the value's range is a compile error, not a silent wrap. #### When a declaration-scope `$if` is decided A `$if` chain written in declaration scope runs at one of two times, and what its arms *contain* picks which. - **Some arm declares something.** The chain is decided while names are being resolved, because what it decides is which declarations exist and every later stage reads the resulting set. Nothing has a type at that point, so the gate cannot ask a type question: a layout intrinsic, a type predicate, or a `$type_of` comparison there is rejected, with a message naming the reason. - **No arm declares anything.** The chain contributes no name and no type whichever arm is taken, so nothing depends on deciding it early. It is decided during type checking instead, where its gate may measure a type (`$size_of`, `$align_of`, `$length_of`), query one (`$is_record` and friends), or compare one. The question is answered from the **syntax**, over every arm - `$if`, every `$or`, and a trailing `$or {}` - before any gate is evaluated. One declaring arm anywhere keeps the whole chain at the earlier time. Per-arm answers are not possible: which stage runs the gate would then depend on which arm the gate selects, and the stage that would have to know that is the one being chosen. A `use` is a declaration, so a conditional import is always decided while names are resolved. That is what makes the common target-gating form above work. ```mach rec MeshUniforms { model: [16]f32; } # no arm declares: decided during type checking, so the gate may measure $if ($size_of(MeshUniforms) != 64) { $error("MeshUniforms must be 64 bytes"); } # the second arm declares, so the whole chain is decided while names are # resolved - and the gate is rejected there $if ($size_of(MeshUniforms) != 64) { $error("MeshUniforms must be 64 bytes"); } $or { val PADDING: u32 = 0; } ``` The one visible consequence is ordering: a `$error` reached under a chain that declares nothing is reported during type checking, so an unrelated name-resolution error elsewhere in the same module is reported before it rather than after. A `$if` inside a **function body** is always decided during type checking - it selects statements rather than declarations, so the question does not arise and its gate may always ask about a type. #### Branching on a comptime parameter A comptime function parameter (`$mode: u8`) is a compile-time-known argument, fixed per call site. `$if` / `$or` may branch on it: the compiler monomorphizes the body once per distinct comptime-argument value, and each instance compiles only the arm its value selects. ```mach val MODE_DOUBLE: u8 = 0; val MODE_SQUARE: u8 = 1; fun apply($mode: u8, n: i64) i64 { $if (mode == MODE_DOUBLE) { ret n + n; } $or (mode == MODE_SQUARE) { ret n * n; } ret 0; } # apply(MODE_DOUBLE, ..) and apply(MODE_SQUARE, ..) emit two distinct bodies, # each carrying only its selected arm. ``` ##### Rules - The argument bound to a `$`-parameter must be a compile-time constant at the call site - a literal, a `pub val`, or another comptime parameter. A runtime value is rejected with `comptime argument is not a compile-time constant`. Cross-module constants work, whether imported by bare name or as a qualified member (`alias.CONST`). - Each arm gate must itself be comptime-foldable: its identifiers must all be comptime (parameters or constants). A gate that references a runtime local or parameter is rejected. - A comptime parameter has no storage, so its address cannot be taken: `?$mode` is rejected with `cannot take the address of a comptime parameter`. - A comptime-parameter function is a template, not a value - it can only be called, never assigned, passed, or compared. `val fp = apply;` is rejected with `cannot reference a comptime-parameter function as a value`. - Comptime parameters carry no runtime cost: they are stripped from the lowered signature and ABI, so only the runtime parameters are passed. They may be mixed freely with runtime parameters in any order. - An instance is emitted against its declaring module and folds its gates against that module's own comptime constants, so a library can export a comptime-parameter function gated on its own `pub val`s. - Combining a comptime parameter with a generic function (`fun f[T]($mode: u8, ...)`) is planned - it is reported today with a clear diagnostic. Because each instance compiles only its taken arm, a comptime parameter can safely gate per-target `asm` blocks: a register one backend does not recognize lives only in the arm the other backend never compiles. ```mach pub fun load($order: Order, ptr: *i64) i64 { var result: i64 = 0; $if ($mach.build.arch == $mach.arch.aarch64) { $if (order == RELAXED) { asm aarch64 { ldr {result}, [{ptr}] } } $or (order == ACQUIRE) { asm aarch64 { ldar {result}, [{ptr}] } } } ret result; } ``` #### Discarded branches An untaken `$if` branch is absent from the compiled output, but what "not taken" means for name resolution depends on what the gate reads. - `$mach.*` or constant gates - the untaken arms are entirely absent and the compiler does not even resolve names inside them. This is what makes per-target `asm` blocks safe when one arm references registers the other backend has never heard of. - comptime-parameter gates - the one exception. Because arm selection happens per call site at monomorphization, name resolution and type checking run over all arms structurally, and only the selected arm is emitted into each instance. Every arm must therefore be independently resolvable and type-checkable. - `$type_of` type gates - not an exception. At monomorphization the operand's concrete type is known, so provably-dead arms are pruned and only the selected arm is type-checked and emitted. Each arm may use its value at its own concrete type with no per-arm cast. #### See also - [The comptime channel](https://machlang.org/docs/comptime.html) - `$mach.*` reads for target and build conditions - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$type_of` dispatch and the comptime toolkit - [Statements](https://machlang.org/docs/statements.html) - the runtime `if` / `or` counterpart - [Variadic packs](https://machlang.org/docs/variadics.html) - `$if` / `$or` inside `$each` bodies Source: https://machlang.org/docs/variadics.html ### Variadic packs A trailing `va: ...` parameter collects a variable number of call-site arguments into a compile-time sequence. The compiler monomorphizes the function once per distinct argument type-list. There is no runtime structure, no `va_list`, and no `any`. #### Declaring a pack A pack parameter is a named parameter whose type is `...`. It must be last; any number of fixed, typed parameters may precede it. ```mach fun name(va: ...) RetType { ... } fun name(fixed: T, va: ...) RetType { ... } # leading fixed params are allowed ``` #### Iterating with `$each` `$each a in va` unrolls the body once per element, with `a` bound to the element and re-typed to that element's concrete type per instantiation. This is the only way to consume a pack. ```mach fun sum(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a; } ret t; } sum(1, 2, 3) # 6 sum() # 0, the body never runs on an empty pack ``` Because each element has its own concrete type at monomorphization, the body can handle a heterogeneous pack by casting each element: ```mach $each a in va { t = t + a::i64; # cast each element's concrete type to i64 } ``` > **Note:** A `$each` body is a normal statement block. Runtime variables in the enclosing scope (a cursor, an accumulator) thread across every unrolled iteration: each one reads where the previous left off. #### Counting with `va.len` `va.len` folds to the instance's element count at compile time. ```mach fun count(va: ...) i64 { ret va.len::i64; } count(1, 2, 3) # 3 ``` #### Forwarding a whole pack Inside a pack instance, `g(va...)` forwards the whole pack to another pack-tailed function, which is monomorphized for the forwarded type-list. ```mach fun outer(va: ...) i64 { ret sum(va...); } ``` > **Restriction:** `va...` is valid only as the sole trailing argument of a pack-tailed callee. Spreading into a callee with no pack parameter, spreading with other arguments after it, or spreading only part of a pack are all rejected. #### Monomorphization and ABI A pack-tailed function is compiled once per distinct argument type-list at each call site. Different arities, or the same arity with different types, produce separate instances. | Call | Instance | | --- | --- | | `sum(1, 2, 3)` | `(i64, i64, i64)` | | `sum(10, 20)` | `(i64, i64)` | | `sum(5::u8, 1::u32)` | `(u8, u32)` | > **Not a stable ABI:** A pack-tailed function has no single entry point, so it is not a stable-ABI symbol. It cannot be the target of an `ext fun` or a function pointer shared across compilation units. It is source-level only. #### Example: vformat The standard library's `vformat` is a pack-tailed function. Each `$each` iteration handles one format argument in order, with `$type_of` dispatch selecting the right writer per element type. ```mach $each arg in va { $if ($type_of(arg) == str) { write_str(w, arg); } $or ($type_of(arg) == i64) { write_i64(w, arg); } $or { $error("no writer for this argument type"); } } ``` The runtime format cursor threads across all unrolled iterations; the argument sequence is consumed entirely at compile time. #### See also - [Functions](https://machlang.org/docs/functions.html) - declarations, generic and comptime parameters - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$type_of`, `$fields`, `$each` - [Control flow](https://machlang.org/docs/comptime-control.html) - `$if` and `$or` inside pack bodies Source: https://machlang.org/docs/decorators.html ### Decorators A decorator attaches metadata to a declaration. It can carry a source-use notice, confine a declaration to tests, or control how the compiler emits the symbol - its linker name, alignment, section placement, inlining, dynamic import attribution, constant-time obligations, or exclusion from auto-vectorization. > **Note:** Visibility (`pub` / `ext`) is a separate concern and is never controlled by a decorator. #### Surface A decorator is written as an attribute clause. A bare flag takes no arguments; a directive takes comptime-expression arguments. ```mach #[name] # bare flag (e.g. inline) #[name(args)] # directive with comptime-expr arguments ``` > **Restriction:** A line comment that begins `#[` with no space opens an attribute. Write such a comment with a separating space: `# [...]`. #### Placement Decorators appear before the declaration they target, one per line or space-separated on the same line. They attach to the immediately following declaration only and do not bleed across declarations. ```mach #[inline] #[symbol("big")] fun big(a: i64, b: i64) i64 { ... } #[align(64)] #[symbol("g_lit64")] pub var g_lit64: u8 = 7; ``` #### Directives The directive set is closed; new directives require a compiler change. ##### `deprecated` / `deprecated(str)` - deprecation notice Marks a declaration or tag case as deprecated. An optional string argument supplies guidance for consumers. When code references a deprecated symbol or case, the compiler emits a warning naming the declaration and message. ```mach #[deprecated("use new_parser instead")] pub fun old_parser() { ... } pub tag Status: u8 { ready; #[deprecated("use ready")] pending; } ``` ##### `testing` - test-only declaration Marks a declaration as existing only for tests, giving a helper or fixture the semantics a [`test`](https://machlang.org/docs/testing.html) block already has. It takes no arguments and may appear once. - **Checked in every build.** A `#[testing]` declaration is resolved and type-checked in every build, so it cannot rot, and a type error in one fails `mach build` as well as `mach test`. - **Emitted only for tests.** It is omitted from IR and object files in ordinary builds, emitted under `mach test`, and left out of `mach doc`. - **Referenced only from tests.** A reference is legal only from a `test` body or from another `#[testing]` declaration, including its decorators. Any other reference is an error at the use site. Every kind of reference is checked - value references, calls, address-of, type references (a field, parameter or return type of a production declaration), generic instantiation, comptime evaluation and decorator arguments - and there is no same-module exemption. ```mach # file: src/queue.mach #[testing] pub fun filled(n: u32) Queue { ... } test queue__drains_in_order { var q: Queue = filled(3); # a test body may use the fixture ... } #[testing] fun drained() Queue { ret filled(0); } # so may another testing declaration pub fun reset() Queue { ret filled(0); } # error: `filled` is a `#[testing]` declaration: only a test body or another # `#[testing]` declaration may reference it ``` The mark follows imports. A plain `use` of a testing declaration is itself testing, so every reference through it is checked, and a `#[testing] use` confines its alias even when the target is an ordinary declaration. A `fwd` of a testing declaration must itself be marked `#[testing]`, because an unmarked `fwd` puts the name on a production surface. `pub #[testing]` is valid, and a dependency's testing helper is usable from a dependent's tests. - It applies to `fun`, `rec`, `uni`, `tag`, `def`, `val`, `var`, `use` and `fwd` declarations at module scope. `test` blocks reject it because they are already test-only, and comptime directives reject it too. - It cannot combine with `ext`, `symbol`, `section`, `stage`, or the shader interface decorators (`input`, `output`, `builtin`, `uniform`, `storage`, `sampler`). Each names a consumer outside mach source that the check cannot see. - Inline `asm` `{name}` operands bind only locals, so they never reference a declaration. - The mark is not a cycle escape: `mach test` builds with testing declarations present, so a `use` cycle that only tests need is still an error. ##### `symbol(str)` - linker name Overrides the emitted or imported symbol name. Applies to functions and globals. Without it the compiler mangles the mach name; `symbol` gives the linker the exact name. The mangled name is the source FQN *as the source spells it* - `std.types.string.str_len`, with generic arguments after a `$` - so a profile, a crash report, and a disassembly all read the name you wrote. A test's symbol is its [qualified name](https://machlang.org/docs/testing.html#qualified-names), `#`. There is no prefix: a mangled name always contains a `.` or a `#`, and a C identifier never can. `ext fun` foreign symbols and `#[symbol("...")]` names are literal and unaffected. ```mach #[symbol("main")] fun entry(argc: i64, argv: **u8) i64 { ... } #[symbol("write")] ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ``` ##### `library(str)` - dynamic import attribution Pins an `ext` import to a specific dependency in the link set. Applies to `ext` functions only, and composes with `symbol` - the import is emitted under the renamed symbol within the named dependency. ```mach #[library("ws2_32.dll")] #[symbol("WSAStartup")] ext fun wsa_startup(ver: u16, data: *u8) i32; ``` - The value normally names a `[link.X]` requirement's stable logical identity: its `library` value, or `X` when that key is omitted. A bare command-line `-l name` also exposes `name`. Exact canonical loader names remain accepted. Pinning to an absent dependency is a link error, never a silent fallback. - PE and Mach-O use two-level namespaces, so every dynamic import on those targets needs a `library` attribution. - On ELF (Linux) the loader resolves imports by global search, so `library` has no effect on the emitted binary; the value is still validated against the link's dependency set. ##### `inline` - force inlining Marks a function for inlining at every call site, overriding the compiler's size- and use-count heuristics. Applies to functions only and takes no arguments. ```mach #[inline] fun fast_path(x: i64) i64 { ret x * 2; } ``` Release builds inline small helpers and `#[inline]` functions across modules. Marking a function `#[inline]` declares that the body is part of its public surface so an importer can inline it directly. The copy is emitted weak under the origin's mangled name so the defining module still provides the single canonical symbol at link time. ##### `noinline` - forbid inlining The inverse of `inline`: forbids inlining a function into any caller, overriding the heuristics that would otherwise fold it in. Applies to functions only and takes no arguments. ```mach #[noinline] fun cold_path(code: i64) i64 { panic("unreachable state"); } ``` Use it to keep a function's frame and symbol real - for a profiler or stack sampler to attribute its cost correctly, to keep a cold path from bloating a hot caller's instruction cache, or to hold code size down on a constrained target. - `inline` and `noinline` on the same function is a direct contradiction and is rejected in sema; neither wins silently. - `scalar` already declines inlining as a side effect, so pairing it with `noinline` is legal but redundant. - Purely a hint to the inliner; it does not otherwise change codegen. The debug pipeline runs no inlining pass at all, so `noinline` is inert (and unnecessary) there. ##### `align(expr)` - alignment override Sets the alignment of a global variable or a record/union type. `expr` must be a comptime integer - a literal, or a comptime expression such as `$size_of(T)` or `$align_of(T)`. ```mach #[align(64)] pub var cache_line: u8 = 0; #[align($size_of(Pair))] pub var g_cmp: u8 = 0; #[align(32)] rec Over { a: u8; } ``` - On a `var` / `val`, it sets the global's section and address alignment. - On a `rec` or `uni`, it sets the type's own alignment, inherited by any global of that type. - It does not apply to `def` aliases, which are transparent and have no layout of their own. ##### `packed` - no padding Lays a `rec` or `uni` out with no padding: every field sits immediately after the previous one, there is no padding at the tail, and the type takes no alignment from its fields. Takes no arguments. `align` only ever raises alignment. `packed` is the inverse, and it exists for the case where the layout is not mach's to choose - a C struct, a file header, a wire frame, a vertex whose stride a buffer fixes. Without it such a shape cannot be described as a record at all. ```mach #[packed] rec Header { magic: u8; # offset 0 version: u16; # offset 1 length: u32; # offset 3 checksum: u64; # offset 7 } # $size_of == 15, $align_of == 1 ``` Naturally the same shape is 24 bytes. `$size_of`, `$align_of`, and `$offset_of` all report the packed layout, and so does the code that reads and writes the fields - there is one layout, not a declared one and an emitted one. **It composes with `align` rather than conflicting**, and each owns one question: `packed` decides padding (none between fields, none at the tail) and `align(N)` decides the record's own alignment, rounding its size up to a multiple of `N`. ```mach #[packed] #[align(8)] rec Frame { a: u8; b: u32; } # fields at 0 and 1; $align_of == 8, $size_of == 8 ``` **Packing is not transitive.** A packed record packs *its own* fields; a record it contains keeps its own internal padding and is merely *placed* without padding. This matches C, and it is the rule that composes - an inner type's layout does not change depending on who holds it. A transitive rule would silently change the inner record's meaning inside its holder. If the inner record must be packed too, write `#[packed]` on it as well. ```mach rec Point { x: u8; y: u32; } # natural: y at 4, size 8 #[packed] rec Msg { tag: u8; p: Point; } # p at offset 1, still 8 bytes; $size_of(Msg) == 9 ``` ###### What is refused - **The address of a packed field.** `?r.b` would yield a `*u32`, and a `*u32` states alignment 4 to everything downstream of it while the storage it names has none. The access through such a pointer is correct on the targets mach supports today; the *pointer type* is what is untrue, and it travels. The refusal covers the whole access chain, so `?r.arr[0]`, `?r.inner.x`, and `?p.b` through a `*Packed` go the same way. `?r` on the **whole** record stays legal - a `*Packed` describes an align-1 pointee correctly. To work with a field's value, copy it into a local. This is fail-closed on purpose: refusing can be relaxed later once alignment can ride in a pointer type; permitting cannot be tightened. - **Atomics on a packed field**, by that same rule rather than one of their own. `std.sync.atomic` is ordinary functions over `*i64`, so a pointer is the only route an atomic has to a field, and there is no pointer to hand it. - **Vector fields**, including one reached through an array or a nested record. The reason is evidence rather than arithmetic: an unaligned *scalar* access is measured on real hardware, and that measurement is what `#[packed]` rests on. This is a sequencing decision and is expected to be lifted. - **Interface blocks.** `packed` cannot apply to a `#[uniform]` or `#[storage]` block: its member offsets are fixed by the std140 / std430 layout rules, which packing would contradict. > **Target note: riscv64:** Unaligned access is permitted-but-may-trap on RV64, and where the hardware does not do it Linux emulates the access in the kernel. A packed field access there is expected to be **correct and pathologically slow** - a trap-and-emulate round trip per access rather than a load. x86-64 and aarch64 do unaligned scalar access in hardware. ##### `section(str)` - object section placement Places a function or global variable in a named section instead of the default `.text` / `.data`. The section is created if absent; cross-section calls and accesses use ordinary relocations. ```mach #[section(".hottext")] #[symbol("f_hot")] fun f_hot(x: i64) i64 { ret x + 1; } #[section(".machsec")] #[symbol("g_sec")] pub var g_sec: u64 = 100; ``` ##### `oblivious` - constant-time boundary Marks a function as a constant-time boundary. Applies to functions only and takes no arguments. Inside it the backend must not introduce a secret-dependent branch or select a variable-latency instruction on a secret operand. A translation validator re-derives the secret taint over lowered MIR, and a second validation pass verifies the final machine instructions to reject any leaks. ```mach #[oblivious] fun ct_eq(a: ^[8]u8, b: ^[8]u8) u8 { ... } ``` - Inline `asm` inside such a function is **validated rather than rejected**: the block is parsed and walked for the same leaks, and refused only where a leak is found or where the construct cannot be modelled. - The decorator is purely subtractive - on a secret-free function it is a no-op. A function instance that *computes* on a `^` secret is required to carry it; an instance that only moves, stores, or declassifies secrets stays annotation-free. - The zeroizing-write guarantee is *not* one of the decorator's obligations. A write into secret storage carries a taint keyed on the storage's secrecy rather than on any decorator, so a zeroizing wipe is protected in a function carrying no `#[oblivious]` at all. - Rejected outright for a target whose back half emits a module for a downstream compiler rather than the executed instructions (the SPIR-V backend): the contract cannot be validated or upheld there. > **Experimental preview:** The constant-time guarantee is not complete and has not been audited end to end. Do not build production cryptography on it at this version. See [Assurance](https://machlang.org/docs/secrecy.html#assurance) for what the guarantee covers, where the check runs, and what is still open. ##### `scalar` - opt out of auto-vectorization Excludes a function from loop auto-vectorization, so its loops compile to scalar code even in the release pipeline on a vector-capable target. Applies to functions only and takes no arguments. ```mach #[scalar] fun reference_sum(a: *i64, n: usize) i64 { ... } ``` A `#[scalar]` function is also declined by the inliner, so the opt-out survives inlining - it cannot be lost by the body moving into an unflagged caller. Use it for a scalar reference twin in a differential test, or where vectorized codegen is undesirable for a specific function. The project-wide equivalent is the `vectorize` profile key. ##### `naked` - no prologue, no epilogue, body as written Emits the function's body exactly as written and nothing else: no frame-pointer record, no stack allocation, no callee-save stores, no argument moves, and no return. Applies to functions only and takes no arguments. ```mach #[naked] #[symbol("_start")] fun start() { $if ($mach.build.arch == $mach.arch.x86_64) { asm x86_64 { mov rdi, [rsp] # argc, straight off the kernel-supplied stack lea rsi, [rsp+8] # argv call main } } $or { asm aarch64 { ... } } } ``` The programmer owns the frame, the stack alignment, the link register, and the return. That is the whole point: a reset vector, an interrupt handler that must return with `iret` / `rti` rather than `ret`, a syscall or context-switch stub, or a thread entry point whose register state at entry *is* the interface. - **The body may contain only inline `asm`** - plus the `$if $mach.build.arch` chain that is how mach spells per-ISA assembly. Any other statement is rejected, because it would lower to code assuming a frame the function does not have and would run and return a wrong answer rather than fail. - **No return is generated.** If the asm falls off the end, control runs into whatever the linker placed next. Write the return the ABI (or the interrupt controller) actually calls for. - **Parameters and the return type are still checked** at every call site, so a naked function is called like any other. No moves are emitted for them: the arguments arrive in the ABI's registers and the body reads them there. - **Mutually exclusive with `inline`** - there is no coherent winner between a body spliced into a caller and one that owns its own frame - and with `oblivious`, which already forbids inline asm. Both combinations are rejected in sema. `noinline` is redundant: the inliner declines a naked function unconditionally. Frame *elision* is a separate, automatic thing: the compiler already omits the prologue for a leaf that provably never touches its frame. `naked` is the declared form, and it is unconditional. Merely containing an `asm` block does **not** suppress a frame: a function that also makes a call gets one, since an unaligned call boundary or a clobbered link register is not something the author asked for by writing assembly. ##### `embed(str)` - compile-time file embedding Sources a `val`'s bytes from a file at compile time: the file's content **is** the initializer. Applies to `val` only - not `var` (the storage is read-only data) and not an `ext` data import (which has no storage here). Takes one string-literal argument. ```mach #[embed("assets/logo.qoi")] val LOGO: [_]u8; # length taken from the file's byte count #[embed("boot/sector.bin")] val SECTOR: [512]u8; # length pinned; a size change fails the build ``` - The declaration carries no initializer of its own; writing one alongside `embed` is rejected. This is a second exemption to `val`'s requires-an-initializer rule, alongside `ext`. - The path resolves relative to the **declaring source file's** directory and must remain within the project root. Any path escaping the project root is refused and the file is not read. - The annotation must be `[_]u8` or `[N]u8`. `[_]` is an inferred array length, legal **only** on an `#[embed]` declaration. A `[_]u8` embed can be asked for its own length: `$length_of(LOGO)` is its element count and `$size_of(LOGO)` its byte count, both folded at compile time. The explicit `[N]u8` form is for pinning a size by contract, not for recovering one. - Two `#[embed]` globals whose files hold byte-identical content and whose final section name, kind, and alignment match share **one** read-only data placement within a module, so their addresses compare equal. This is specific to embedded data; an ordinary global is never merged this way. - An explicit `[N]u8` whose `N` disagrees with the file is rejected, naming both counts. This is how a declaration pins a fixed-size asset - a boot sector, a ROM image - so the build fails the moment it stops being that size. - Bytes are placed in read-only data exactly like any other constant byte array: no runtime I/O, no copy. Works for every artifact kind and target, freestanding included. - The embedded file is a build input: its content digest feeds the embedding module's incremental cutoff, so editing the asset invalidates that module and an untouched asset stays a cache hit. #### Shader directives Nine directives belong to the GPU pipeline surface: `stage`, `workgroup`, `input`, `output`, `builtin`, `uniform`, `storage`, `sampler`, and the two target-owned declarations `handle` and `op`. They are accepted on **every** target, because which target a module is built for is not a property of its source; only a target that forms pipeline stages acts on them. Each is documented with the surface it belongs to, on [GPU shaders](https://machlang.org/docs/shaders.html). #### Applicability Each directive targets a fixed set of declaration kinds. The full matrix: | Directive | `fun` | `ext fun` | `val` / `var` | `rec` / `uni` / `tag` | | --- | --- | --- | --- | --- | | `deprecated` | yes | yes | yes | yes | | `testing` | yes | no | yes | yes | | `symbol` | yes | yes | yes | no | | `library` | no | yes | no | no | | `inline` | yes | no | no | no | | `noinline` | yes | no | no | no | | `align` | no | no | yes | yes | | `packed` | no | no | no | yes | | `section` | yes | yes | yes | no | | `oblivious` | yes | no | no | no | | `scalar` | yes | no | no | no | | `naked` | yes | no | no | no | | `embed` | no | no | yes | no | | `stage` | yes | no | no | no | | `workgroup` | yes | no | no | no | | `input` / `output` | no | no | yes | no | | `builtin` | no | no | yes | no | | `uniform` / `storage` | no | no | yes | no | | `sampler` | no | no | yes | no | | `op` | yes | no | no | no | | `handle` | a bodyless `def` only | | | | The `val` / `var` column is shared, but `embed` accepts only `val` - a `var` is refused. `deprecated` also applies to `def`, `use`, `fwd`, and individual tag cases. `testing` also applies to `def`, `use` and `fwd`, but not to a tag case. #### See also - [Testing](https://machlang.org/docs/testing.html) - `test` blocks, whose semantics `testing` gives a declaration - [Visibility](https://machlang.org/docs/visibility.html) - `pub` / `ext`, the separate concern decorators never touch - [Intrinsics](https://machlang.org/docs/intrinsics.html) - `$size_of` and `$align_of` as `align` arguments - [Secrecy](https://machlang.org/docs/secrecy.html) - the `^` qualifier `oblivious` is the codegen contract for - [Inline assembly](https://machlang.org/docs/asm.html) - what a `naked` function's body may contain - [Manifest](https://machlang.org/docs/manifest.html) - the `[link.X]` requirements a `library` pin resolves against Source: https://machlang.org/docs/cli.html ### CLI The `mach` binary is the compiler driver. It dispatches on the first argument to a small set of commands - `build`, `run`, `test`, and more - that compile, execute, and manage a project rooted at a `mach.toml`. #### Invocation ``` mach [options] ``` The compiler dispatches on `argv[1]`. With no command, or an unknown one, it prints usage and exits non-zero. The project commands - `build`, `run`, `test`, and `doc` - take the project as a **required** positional: a directory, whose `mach.toml` is read, or a manifest file directly. The manifest is never guessed from the working directory, and a bare invocation with no path is a user error. Flags are matched exactly: `--flag value` (the value follows in the next argument) or a bare `--flag` toggle. > **Restriction:** Bundled short flags are **not** recognized, and a value is its own argument. The one exception is `--emit-ir`, whose optional form attaches to the same argument (`--emit-ir=listing`) and never consumes the next one. A lone `-` is a positional operand, not a flag. #### Commands | Command | Summary | | --- | --- | | `build` | compile the project to objects and linked outputs | | `check` | parse, resolve, and type-check without compiling or linking | | `run` | execute an existing built binary without rebuilding | | `test` | build the test dispatcher and run the collected tests | | `fmt` | format source files to the canonical layout | | `clean` | remove the project's build output | | `dep` | manage project dependencies: resolve, realize, verify, and update | | `init` | scaffold a new project, with std added at its newest compatible release | | `doc` | generate Markdown reference docs from source doc-comments | | `info` | print the compiler version and the host target it resolves, or every buildable target | | `help` | print usage; `mach help ` for detail | #### Global flags Read by `build`, `check`, `run`, `test`, and `doc`, which share one config parser. Passing a verbosity flag and `--quiet` together is a parse error. | Flag | Value | Effect | | --- | --- | --- | | `-v` | - | verbose readout: build phase timing, plus per-test lines under `mach test` | | `-vv` | - | `-v` plus per-module / per-file (or passing-test) detail | | `--quiet`, `-q` | - | suppress non-error output | | `-a`, `--artifact ` | name or glob | select `[artifact.]` entries; repeatable. Absent, the default selection: the artifacts marked `default = true` for each selected target, or all of them when none is marked | | `-t`, `--target ` | name or glob | select `[target.]` entries; repeatable. Absent, the [native](https://machlang.org/docs/manifest.html#native) target, or the one a named artifact settles | | `-p`, `--profile ` | name or glob | select `[profile.]` entries; repeatable. Absent, the sole declared profile or the one marked `default = true` | | `--all` | - | fill every axis no `-a`, `-t` or `-p` narrows with `*`: every artifact on every target it supports, in every profile | | `-o ` | path | override the linked-binary path; accepted only when the selection resolves to a single build cell | | `-g` | - | emit debug info for this build, overriding the profile's `debug` key | | `--pie` | - | emit a position-independent (`ET_DYN`) executable | | `--subsystem ` | console, gui | the environment a windows executable declares; overrides the artifact's `subsystem` key, inert off windows | | `--emit-asm` | - | emit per-module `.s` assembly text beside each object | | `--emit-ir[=
]` | debug, listing | emit per-module `.ir` text beside each object. Bare (or `=debug`) writes the ir-debug dump. `=listing` writes a readable listing with language-spelled types and `line:column` positions, whose format is a stable contract. | | `--verify-ir` | - | run the IR verifier after each optimisation pass | A selection pattern is an exact name, which must be declared, or a glob in which `*` matches any run of characters and `?` any one, which must match at least one entry. The axis takes every entry any of its patterns names, in declaration order. Quote a glob so the shell leaves it alone: `-a '*'`. A glob skips the cells an artifact's `targets` list does not declare, while naming such a pair exactly is refused. `mach test . --all -p debug` selects every artifact and target in `debug` only. `check` does not take `-o`, `run` does not take `--all`, and `doc` selects with `-a` and `-t` only. > **Note:** IR and assembly emission is a CLI concern only - there is no profile key for it. There are no `--no-emit-*` inverses, no `--release` flag (use `-O2` or `--profile`), no `--color` flag, and no `--verbose` long form. > **Note:** `mach dep` and `mach init` do not use the shared config parser; they read only their own flags. #### build ``` mach build [options] ``` Compiles the project rooted at `` (for example `mach build .`). With no `-a`, it builds the default selection for the default target and profile, and `--all` builds every artifact on every target it supports in every profile. A selection that spans several profiles builds one profile after another. Every reachable module is driven through sema, lower, optimise, and codegen to one relocatable object under the manifest's `obj` template. For a `kind = "bin"` artifact the objects are linked into the resolved `out` path; for a `static` artifact they are archived into an `ar` archive there, and with `--emit obj` the objects are the deliverable and nothing is linked. | Flag | Value | Effect | | --- | --- | --- | | `-O0` | - | force the debug pipeline, overriding the profile | | `-O2` | - | select the release pipeline, overriding the profile | | `--emit ` | obj, exe | `obj` stops at the objects; `exe` (default) links a binary | | `--jobs ` | count | codegen worker threads (default: host CPUs; `1` serialises) | | `--plan` | - | print the effective build plan and exit without building | | `--no-cache` | - | force an uncached build: no `obj/` reuse and no build-step reuse (default: the object cache is on) | | `-L ` | dir | add a search directory for `-l`-resolved inputs; repeatable | | `-l ` | name | link a named object, archive, or shared library; repeatable | Plus the global flags. `-O0` or `-O2` overrides the selected profile's `opt` for this invocation; absent one, the profile decides. `-O1` was removed and is refused (`-O1 was removed; use -O0 or -O2`). The build reads `obj/` as its object cache by default: a module whose object there carries the build's key is reused instead of lowered and generated again, and an edit rebuilds only that module and the modules whose view of it changed. A bare `.o` / `.a` / `.so` / `.dylib` / `.dll` argument, or any `/`-bearing path, is linked verbatim. ##### Link inputs `ext fun` declarations are forward references resolved at link time by external precompiled code - a loose `.o` object, a static `.a` archive, or a shared `.so` library. Inputs come from the command line and from the manifest's merged `libs` overlay; both sets are linked. An input that resolves to no existing file is a hard error, so a typo never silently drops a dependency. - **Explicit input path** - a bare (non-flag) argument that contains a `/`, ends in `.o` or `.a`, or names a `.so` is treated as an input path. The first non-flag positional after `build` is the project root and is skipped; the rest are link inputs. A relative path is tried verbatim against the working directory first, then rooted at the project root. - **`-l `** - each `-L ` is searched for `lib.o`, `.o`, `lib.a`, then `.a`; if none hit, the same candidates are tried against the working directory. Only if no static object or archive is found does resolution fall back to a shared `lib.so` (the `-L` dirs, then the target OS's default and common system library directories). - **`-L `** - adds a search directory for the resolution above. Both `-L` and `-l` may be repeated. How an input resolves decides static vs dynamic linking. A loose `.o` or static `.a` is a **static** input merged into the executable (an archive contributes every member object). A shared `.so` is a **dynamic** dependency: its `DT_SONAME` is recorded and undefined `ext` symbols are bound at load time through an emitted PLT. A static definition always wins over a dynamic import of the same name. `-l` prefers a static candidate, so the `.so` fallback applies only when none exists (the common case for system libraries like libc). Manifest `libs` resolve before CLI inputs, giving a deterministic link order. > **Note:** Dynamic linking is implemented for the ELF (Linux) and PE (Windows) targets. The Mach-O (Darwin) import path is not yet implemented. #### check ``` mach check [options] ``` Parses, resolves, and type-checks the source reachable from the selected artifacts through the compiler frontend. No code is generated, compiled, linked, or written to disk. It serves as a fast verification loop during development. #### run ``` mach run [options] [-- args...] ``` Executes the binary `mach build` already produced for `` without rebuilding. The selection options (`-a`, `-t`, `-p`) must resolve to exactly one cell, and `--all` is not accepted. With no `-a`, it runs the sole `bin` the target builds. Arguments after a `--` separator are forwarded to the child as its `argv`, and the child's exit code becomes this command's exit code. > **Note:** `mach run` does not build. Run `mach build` first to produce or refresh the executable. | Flag | Value | Effect | | --- | --- | --- | | `--runner ` | command | execute the binary as ` ` instead of directly | | `--timeout ` | duration | terminate the run and its process group after the duration (`30ms`, `30s`, `5m`, `1h`), exiting `3` (default unbounded) | `--runner` names a host-side launcher for binaries the host cannot exec directly, for example `mach run . -t windows --runner wine`. `` is a single command name or path (no shell-style word splitting); a bare name is resolved on `PATH`. Without the flag the binary is exec'd directly, and a launch failure reports exit `127` when `execve` rejects the binary, with no auto-detection. #### test ``` mach test [options] ``` Builds the artifact's objects as `build` does, adds a test object `obj//.test.o` beside each module that declares tests, links **one dispatcher executable** of only the selected tests under `test//`, and runs each test as its own process (` `), reporting a per-module roll-up that expands failures. A test is named by its qualified name, `#`. A crashing test reports its signal and the run continues. Only the current project's own tests run by default. A test build always links executables, even for a library target. | Flag | Value | Effect | | --- | --- | --- | | `--filter ` | substring | run only tests whose qualified name contains `` | | `--include-deps` | - | also run tests declared in dependency modules | | `--list` | - | print each collected test's qualified name and test object path, and exit | | `--format ` | human, json | `human` (default), or a `json` NDJSON event stream for tooling, schema `2` | | `--jobs ` | count | run up to `n` test processes at once, and size the build's codegen workers | | `--runner ` | command | launch every test as ` ` instead of exec'ing it directly | | `--timeout ` | duration | terminate a test process and its process group after the duration (`30ms`, `30s`, `5m`, `1h`); a test that exceeds it is reported as timed out (default unbounded) | | `--no-cache` | - | as on `build` | Plus the `build` and global flags. `--runner` has the same semantics as on `run`. Without it, a test the host cannot launch reports a per-test failure - `FAIL(exit 127)` when `execve` rejects the binary, `FAIL(spawn)` when the spawn itself fails. #### fmt ``` mach fmt ||- [--check] ``` Rewrites source in the one canonical layout. There is no configuration. The operand decides what is formatted: - **a directory or `mach.toml`** - rewrites that project's own source files in place, walking only `[project].src`, so no `[profile.*]` is needed. Dependencies are never visited or fetched. - **any other file** - rewrites that one source file, with no manifest read. - **`-`** - reads source from stdin and writes its canonical layout to stdout, touching nothing on disk. Diagnostics go to stderr. With `--check`, nothing is written and every input whose layout differs is named on stdout (the stream as ``). A file that does not parse is reported and left unchanged. Exit codes: `0` formatted or already canonical, `1` an input differs under `--check`, is malformed, or the invocation is a user error, `2` internal failure, `3` filesystem failure. ``` mach fmt . # the whole project mach fmt src/root.mach # one file mach fmt - < draft.mach # stdin to stdout, for editors mach fmt . --check # ci: fail if anything is unformatted ``` #### clean ``` mach clean ``` Removes the project's build output for every declared target and profile: `obj/`, `ir/`, `asm/`, `.cache/`, `.stage/`, `test/`, and `dep/` under the output directory, along with every artifact output (see [the output directory](https://machlang.org/docs/manifest.html#output-directory)), so the build after it is a cold one that reuses nothing. `` is a project directory or a manifest file. It is idempotent and takes no options. Exit codes are its own: `0` ok, `1` no project or a bad manifest, `2` an IO failure. #### dep ``` mach dep [args] [options] ``` Manages the project's dependency tree. A dependency is either a **git** URL, realized as a submodule under `dep/` and selected by a `version` range or an exact `ref`, or a **path** to another project tree, copied into `dep/`. Every action accepts `--quiet` / `-q`. | Action | Args | Effect | | --- | --- | --- | | `list` | `` | list declared dependencies with their source, ref, staged pin, and state (`realized` or `missing`). | | `add` | ` (--git [--ref \| --version ] \| --path )` | declare a dependency and realize the tree. With `--git` alone, writes `version = "^X.Y.Z"` of the highest release every range and the running compiler accept. | | `remove` | ` [--purge]` | remove the declaration. `--purge` also deletes the dependency directory. | | `update` | ` ( \| --all) [--lowest] [--offline]` | refresh path copies, advance branch refs, and resolve version ranges to the releases they select. `` may be any identity in the closure, and the others stay put while they still fit. `--lowest` selects the lowest release every range accepts. `--offline` resolves from releases already fetched. | | `pull` | `` | realize declared dependencies without advancing revisions or refreshing existing path copies. Idempotent. | | `outdated` | ` [--offline]` | report each version-selected dependency's pinned, highest compatible, and latest release. | | `verify` | ` [--release]` | check the realized closure against the manifests without changing anything. `--release` also requires every dependency of the project itself to be selected by `version` or an exact `tag/`. | `add` takes exactly one of `--git` or `--path`. `--ref` and `--version` are valid only with `--git` and cannot be combined. `update` takes a name or `--all`, not both. Resolution runs only in `add`, `update`, and `outdated`. `mach build` never resolves and never needs the network: it verifies the realized tree offline, including that each version-selected pin is a release inside every requirer's range. Two different selections of one identity are a hard error that names both chains and the root declaration that would decide. See [Dependencies](https://machlang.org/docs/dependencies.html) for the range grammar and the full rules. ##### Dependency pins Dependencies are pinned by the gitlinks the root repository commits under `dep/`. There is no lock file. Each dependency in the transitive closure lives directly under `dep//`, where the manifest key, directory name, and project id must match. #### init ``` mach init [dir] [options] ``` Scaffolds a new project in `[dir]` (default: the current directory). An absent directory is created, and an existing one is filled in. It writes a complete `mach.toml` - a `[project]` block whose `mach` key is the running compiler's caret range (`mach = "^6.0"`), `[target.-]` entries for linux, windows, and darwin on the host architecture, one artifact, and debug and release profiles - plus a starter source file. It then initializes a Git repository and adds std as `mach dep add std --git https://github.com/briar-systems/mach-std` would: at the caret range of the newest std release whose `[project].mach` admits the running compiler, checked out at that release. Every refusal happens before the first write, so a refused init changes nothing. | Flag | Value | Effect | | --- | --- | --- | | `--name ` | id | project id (default: the directory name). Letters, digits, `_`, and `-` only. | | `--force` | - | overwrite an existing `mach.toml` and entry file | | `--lib` | - | library layout: write `src/lib.mach` and a `kind = "static"` artifact marked `default = true`, instead of `src/root.mach` and a `kind = "bin"` artifact | | `--no-deps` | - | resolve and declare std without checking it out. Resolving still needs the network. | | `--no-git` | - | skip Git initialization and realize dependencies as plain checkouts | | `--quiet`, `-q` | - | suppress non-error output | The first non-flag argument is the target directory. A default binary scaffold links and runs from `mach build .` without further manifest edits. Offline, init fails to add std, says so, prints the `mach dep add` command to run once connected, and writes no `[dep.std]` table. Exit codes: `0` success, `1` an invalid id or refused overwrite, `2` internal failure, `3` the std dependency could not be added. #### doc, info, and help **`mach doc `** loads the module graph and generates Markdown reference docs from source doc-comments - one page per module plus an index. Each `pub` declaration is paired with the run of `#` comment lines immediately preceding it. `--out ` sets the output directory (default `doc/api`); `--target ` selects a target for module discovery. The hand-written language material is never touched. **`mach info`** prints the compiler's version and the host target it resolves for the machine it runs on, one dimension per line. It needs no project. Every value is taken from the target the host request resolves to, so on an x86_64 Linux machine it prints: ``` mach host: linux-x86_64 isa: x86_64 os: linux abi: sysv64 object: elf ``` **`mach info targets`** prints every target the registries compose end to end, one full tuple per line. A name appears once for each ABI and object format it supports, so it never advertises a tuple that would fail to resolve. Abridged: ``` linux-x86_64 isa=x86_64 os=linux abi=sysv64 object=elf linux-aarch64 isa=aarch64 os=linux abi=aapcs64 object=elf linux-riscv64 isa=rv64gc os=linux abi=lp64d object=elf darwin-aarch64 isa=aarch64 os=darwin abi=aapcs64 object=macho windows-x86_64 isa=x86_64 os=windows abi=win64 object=coff freestanding-riscv32 isa=rv32imac os=freestanding abi=ilp32 object=raw freestanding-spirv isa=spirv os=freestanding abi=spirv object=spv ``` `mach info --version` prints the version alone on one line, for tooling. Combining `--version` with `targets`, or passing any other verb, is a usage error. The exit status is `0` on success, `1` for a usage error, and `2` when target registration or host resolution fails. **`mach help [command]`** prints the top-level usage summary, or - with a known command - that command's detail page. An unknown command prints usage and exits non-zero. #### Exit codes The commands share a stable convention for scripting: | Code | Meaning | | --- | --- | | `0` | success (for `test`, all tests passed) | | `1` | user error: missing project path, no `mach.toml`, unknown target, compile errors, an unresolvable link input (for `test`, any test failed) | | `2` | internal error | | `3` | environment failure, where a command documents it: a filesystem failure in `fmt`, an unaddable dependency in `init`, an expired `--timeout` in `run` | `mach run` returns the child process's exit code directly. #### See also - [Manifest](https://machlang.org/docs/manifest.html) - the `mach.toml` schema and keys - [Dependencies](https://machlang.org/docs/dependencies.html) - version ranges, resolution, and git and path dependencies in detail - [Testing](https://machlang.org/docs/testing.html) - writing tests and the `mach test` runner Source: https://machlang.org/docs/manifest.html ### Manifest `mach.toml` declares a project. It separates the orthogonal axes of a build - *what* is produced (`[artifact.*]`), *where* it runs (`[target.*]`), *how* it is compiled (`[profile.*]`), and what it links or must run first (`[link.*]`, `[step.*]`) - and the build engine takes their product. Nothing is inferred from another key. > **Totality:** A **root** manifest is parsed strictly and totally: every required key must be present, and an unknown key is an error rather than a silent no-op. A **dependency's** manifest is read permissively - only the keys the consumer needs are consulted - so a library cannot impose its build choices on you. #### The schema ``` [project] id = "demo" # required: identifier; root of every module path version = "0.1.0" # required src = "src" # required: source dir, project-root-relative out = "out/{target.name}/{profile.name}" # required: output-path template root mach = "^6.0" # required in a root manifest: the compiler range [target.linux] # a platform: a fully-spelled tuple default = true isa = "x86_64" os = "linux" abi = "sysv64" [profile.debug] # a build variant: requires all five keys default = true opt = 0 # 0 (debug pipeline) | 1 | 2 (release pipeline) debug = true # emit debug info for this profile simd = "scalarize" # SIMD lever: "scalarize" | "require" vectorize = true # auto-vectorization lever float_reassoc = false # associative float transforms [artifact.demo] # a produced artifact default = true kind = "bin" # "bin" | "static" | "shared" entry = "root.mach" # entry source, relative to src out = "bin/demo" # output path, relative to the project out targets = ["*"] # which declared targets build it ("*" = all) link = [] # [link.X] names this artifact links need = [] # category-qualified names demanded directly [dep.std] # a dependency: directory, manifest key, and id match git = "https://github.com/briar-systems/mach-std" version = "^9.0" # a release range, or ref = "tag/..." for one exact selection ``` #### `[project]` | Key | Type | Meaning | | --- | --- | --- | | `id` | string | Root segment of every module path the project exposes: a file at `/foo/bar.mach` is reachable as `.foo.bar`. Must be a plain identifier - letters, digits, `_`, `-` - since it names the dependency store and keys step stamp files. Read by `$project.id`. | | `version` | string | Project version. Read by `$project.version` and `$project.version.{major,minor,patch}`, and stamped into a Windows executable's version resource. A release tag `vX.Y.Z` is a [release](https://machlang.org/docs/dependencies.html#resolution) only when this matches it. | | `src` | string | Source root, project-root-relative. Module paths resolve under it. | | `out` | string | The output-path template root, referenced as `{project.out}` by artifact `out`, step paths, and `cmd`s. | | `mach` | string | The compiler versions this project builds with, as a [version range](https://machlang.org/docs/dependencies.html#version-ranges) such as `"^6.0"`. Required in a root manifest. See [Compiler range](https://machlang.org/docs/manifest.html#compiler-range). | > **Restriction:** `[project]` permits only `id`, `version`, `src`, `out`, and `mach`. Any other key, `name` and `description` included, is an unknown-key error. ##### The output directory Everything a build and its cache write lands under the expanded `out`, in one layout: | Path | Holds | | --- | --- | | `obj/` | one object per module, each carrying its cache key; it is the object cache, and is read by developers and tooling too | | `ir/`, `asm/` | the human-readable views `--emit-ir` and `--emit-asm` write | | `.cache/` | compiler-only state, read and written by nothing but the compiler | | `.cache/steps/` | one fingerprint stamp per [build step](https://machlang.org/docs/manifest.html#steps) | | `.stage/` | build step scratch space, one directory per step, reset before the step runs | | `test//dispatch.o` | the test dispatcher object of a tested artifact | | `test//` | the test dispatcher executable | | `test//log/` | a failing test's captured output | | `dep//` | a dependency's artifact outputs | Test objects sit in `obj/` beside the module objects, as `obj//.test.o`. Artifact outputs go wherever their own `out` names under the directory. `mach build` and `mach test` read `obj/` as the object cache by default. Each module has its own key, so an edit rebuilds only that module and the modules whose view of it changed, and a warm build links the same binary a cold one does. An object that is missing or carries another key is rebuilt, never linked stale. `--no-cache` forces an uncached build. `mach clean` removes `obj/`, `ir/`, `asm/`, `.cache/`, `.stage/`, `test/`, and `dep/` along with every artifact output, for every declared target and profile, so the build after it is a cold one that reuses nothing. A step output may not name a path inside `obj//`, `.cache/`, or `.stage/`. ##### Compiler range `mach` states which compilers a project builds with. Every command that reads the dependency closure - `build`, `test`, `check`, `mach dep verify`, and the language server - checks it for the root and for every realized dependency. A compiler outside any of those ranges is refused once, with every unmet requirement and the chain that states it: ``` error: this is mach 5.2.1, and the dependency closure does not accept it: app (mach.toml) requires mach ^5.3 app -> gfx -> glfw requires mach >=5.4, <6 ``` A root manifest must state `mach`. One without it is refused, with the line to add. A dependency without it states no constraint. `mach init` writes the oldest release of the running compiler's major that reads the key: `mach = "^6.0"` for every 6.x compiler. Dependency resolution also reads it: a release whose `mach` range excludes the running compiler is never selected. The compiler's version is the last release it was built from. A build from an unreleased tree reports that release, so a project cannot require an unreleased feature by version. #### Targets Each `` in `[target.]` is a selector you pass to `-t `. A target is a fully-spelled platform tuple; nothing is inferred from another key. | Key | Req | Meaning | | --- | --- | --- | | `isa` | yes | Instruction-set architecture. Read by `$project.target.arch`. | | `os` | yes | Operating system. Read by `$project.target.os`. | | `abi` | yes | Application binary interface. Read by `$project.target.abi`. | | `of` | no | Object-format override; defers to the os's format when omitted. | | `base` | no | Load-address override (integer). Overrides the os's default base virtual address; defers to it (`0` for `freestanding`) when omitted. | | `platform` | no | Open platform tag (string), surfaced to comptime as `$mach.build.platform` (empty when unset). A support library keys its backend on it; the compiler treats it as opaque. | > **Restriction:** `native` is a reserved name - declaring `[target.native]` is an error, because `native` resolves to whichever *declared* target matches the host. ##### Accepted tuple values | Axis | Values | | --- | --- | | `isa` | `x86_64`, `aarch64`, `riscv64`, `riscv32`, `spirv` | | `os` | `linux`, `windows`, `darwin`, `freestanding` | | `abi` | `sysv64`, `win64`, `aapcs64`, `lp64`, `lp64f`, `lp64d`, `ilp32`, `ilp32f`, `ilp32d`, `spirv` | | `of` | `elf`, `coff`, `macho`, `raw`, `spv` | A value outside its axis set is a strict-parse error, so a typo is caught rather than silently never matching. `mach info targets` prints the tuples this binary can actually build. It is derived from the same declarations composition reads, so it never advertises a tuple that would fail to resolve. > **RISC-V is a family of ABIs:** `lp64` means **soft float**: arguments ride integer registers even where the hardware has an FPU. `lp64f` passes single-precision floats in float registers and `lp64d` passes both single and double, which is what a Linux riscv64 toolchain means by its default. `ilp32`, `ilp32f`, and `ilp32d` are the same three on 32-bit RISC-V. Passing `lp64` where `lp64d` was meant is an ABI mismatch against every C object on the system, not a performance choice. `riscv64` defaults to `rv64gc`, while `riscv32` defaults to `rv32imac` and `ilp32` (specify `rv32imafdc` when RV32 hardware float is wanted). Each os has a default object format: `linux` to `elf`, `windows` to `coff`, `darwin` to `macho`, `freestanding` to `raw`, and `of` names a different one. An os accepts only the formats it can load, so an override the os cannot enter is refused. The default is a function of the whole tuple, not the os alone: a `spirv` target resolves to `spv` regardless of the os it names. A `freestanding` target can select `of = "elf"` on every architecture, and gets an image with no loader construct in it: no header segment, since nothing parses a program header at run time on bare metal. ``` [target.metal] isa = "x86_64" os = "freestanding" # os default object format is "raw" abi = "sysv64" of = "elf" # override: emit an ELF image instead ``` `--pie`, dynamic linking, and a shared library are each refused by name on a loaderless os: a position-independent executable exists to be relocated by a program loader, a dynamically-linked one to have its imports resolved by one, and `os = "freestanding"` has none. ##### Finished-module targets A `spirv` target's object output is a complete, self-contained module rather than a link input. The build delivers the **module tree** - one `/obj/.spv` per module - and runs no link phase, so a default build and `--emit obj` produce the same files. ``` [target.gpu] isa = "spirv" os = "freestanding" abi = "spirv" # no `of`: the finished-module format resolves on its own ``` The artifact's `out` template and `-o` name a linked binary, which such a target has none of; the module tree is delivered instead. A `static` or `shared` artifact kind, and `mach test`, are refused by name. ##### Platform targets (bare metal) A bare-metal platform is not its own `os`. It is `os = "freestanding"` plus two optional keys: a `base` load-address override, and an open `platform` tag a support library keys its backend on. ``` [target.bmos] isa = "x86_64" os = "freestanding" abi = "sysv64" of = "raw" base = 0xFFFF800000000000 platform = "bmos" ``` #### Artifacts Every artifact is declared explicitly and named by its table key. `$bin.name` reads the selected artifact name. | Key | Req | Meaning | | --- | --- | --- | | `kind` | yes | `"bin"`, `"static"`, or `"shared"`. | | `entry` | yes | Entry source, relative to the project `src` dir. The entry module FQN is `.`, with `/` turned into `.`. | | `out` | yes | This artifact output path, **relative to the expanded project `out`** and rooted there automatically, such as `bin/demo` or `bin/demo{artifact.suffix}`. | | `targets` | yes | Array of declared target names this artifact builds for, where `["*"]` means every declared target. | | `link` | yes | Array of `[link.X]` names this artifact links. `[]` for none. | | `need` | yes | Array of category-qualified requirements such as `step.generate`, `artifact.support`, and `artifact.shader-*`. Each glob matches only its named category. `[]` for none. | | `subsystem` | no | `"console"` (default) or `"gui"`, the environment a Windows executable declares it runs under. Refused on targets whose image format has no subsystem. | | `icon` | no | Project-root-relative `.ico` path embedded in a Windows executable PE resources. `bin` artifacts only. | | `manifest` | no | Project-root-relative application-manifest path embedded byte-for-byte in a Windows executable PE resources. `bin` artifacts only. | | `default` | no | `true` marks the artifact chosen when a command needs one artifact (such as `mach test` or `mach run`) and several declared artifacts support the selected target. Exactly one candidate may carry it. | - **`bin`** links an executable at the resolved `out` path. - **`static`** materialises a real `ar` archive at the resolved `out` path: the per-module objects with an archive symbol index, the deliverable a consumer links as a `.a`. - **`shared`** links a shared object at the resolved `out` path, on ELF targets only. Mach-O and PE targets are refused with `link: object format cannot write shared libraries`: dylib and DLL output is not implemented yet and is tracked in [mach#3588](https://github.com/briar-systems/mach/issues/3588). The library exports the root project's `pub` declarations and the names its modules re-export with `fwd`. A dependency's `pub` surface is not part of its ABI, and `#[symbol]` chooses a name, not a visibility. Every other definition still resolves inside the link but is not exported: it is a local symbol in the library's `.symtab` and absent from `.dynsym`. A shared artifact that exports nothing is refused. A freestanding target is refused at artifact naming (`this object format has no shared-library form`), or, with `of = "elf"`, at link time because it has no loader. Per-target extension or per-target entry is not a per-cell exception table. It is a second artifact stanza, so the condition stays visible like everything else. ##### Windows subsystem and resources A PE executable records in its optional header which environment it wants, and the Windows loader honours it: `"console"` gets a console window attached to the process, `"gui"` does not. A graphical application sets `"gui"` to stop an empty console from opening behind it on launch. `--subsystem console|gui` overrides the key for one invocation. ``` [artifact.game] kind = "bin" entry = "root.mach" out = "bin/game.exe" targets = ["*"] link = [] need = [] subsystem = "gui" icon = "assets/game.ico" manifest = "assets/game.manifest" ``` Only a PE image carries the subsystem field. A key written on an artifact planned for a target whose format has none (ELF, Mach-O, raw) is refused as unsupported, naming the key, the target, and the format. An omitted key is the console default and is never refused, so an artifact declaring no subsystem still builds everywhere. Resource keys like `icon` and `manifest` remain valid in a multi-target artifact, but are inert off Windows where Mach does not read either path. #### Profiles A `[profile.]` is a build variant. The optimization level, debug toggle, and SIMD levers live here because they are variant concerns. | Key | Req | Meaning | | --- | --- | --- | | `opt` | yes | Optimization level: `0` selects the debug pipeline (the always-on passes only), `1` and `2` select the release pipeline. Any other integer is a manifest error. | | `debug` | yes | Emit debug info (DWARF, including in `.debug_*` sections of COFF objects and PE images). Gates emission only, never the optimizer, so a `release` profile can keep symbols with `debug = true`. | | `simd` | yes | SIMD scalarization lever. `"scalarize"` builds for a target without hardware SIMD by emitting a defined unrolled scalar expansion of each vector operator; `"require"` makes a build for an incapable target a hard error naming the offending operator. | | `vectorize` | yes | Auto-vectorization lever. When `true`, the release pipeline rewrites provably safe counted loops to 128-bit SIMD on a target with hardware vectors. `false` skips the pass, changing performance but never semantics. | | `float_reassoc` | yes | Permission to treat floating-point addition and multiplication as **associative**. It lets the vectorizer reduce an `f32` / `f64` accumulator through lane-count partial sums, which changes the result. The only profile key that can change a program computed answer. | | `default` | no | `true` marks the profile a build uses when several are declared and `-p` is omitted. Exactly one profile may carry it. | All five keys (`opt`, `debug`, `simd`, `vectorize`, `float_reassoc`) are required in a declared root profile. Emission of the human-readable IR and assembly side-artifacts is not a profile concern. It is controlled only by the `--emit-ir` and `--emit-asm` CLI flags. > **Note:** The `simd`, `vectorize`, and `float_reassoc` levers are always the **consumer's**. A dependency `[profile.*]` is parsed permissively and never read to build the consumer, so a library's values are inert. There is no ecosystem fork and no dual API. #### Link requirements A `[link.]` is a named external link requirement. Artifacts reference entries by name in their `link = [...]`, and an entry whose filters do not match the build cell is skipped. An entry with `export = true` also applies to any project that links this project's modules, so a platform link requirement lives once in the manifest that needs it and cascades to consumers. | Key | Req | Meaning | | --- | --- | --- | | `source` | yes | `"system"` (a system library resolved by name), `"framework"` (a macOS framework), or `"local"` (a file on disk). | | `name` | shape | Library / framework name: required for `"system"` and `"framework"`, forbidden for `"local"`. | | `path` | shape | File path: required for `"local"`, forbidden otherwise. A template. | | `library` | no | Stable logical name used by `#[library("...")]`; defaults to the table name. | | `symbols` | no | Array of symbol names this dependency provides, attributing imports that have no `ext` declaration to decorate. | | `os` / `isa` / `abi` | yes | Filter axes: a canonical value, `"*"` (any), an array of values, or `[]` (none). An entry applies to a cell when all three match. | | `export` | yes | `true` cascades this entry to consumers; `false` keeps it to this project's own builds. | ``` [link.kernel32] source = "system" name = "kernel32.dll" library = "kernel32" symbols = ["Sleep", "CreateFileW", "CloseHandle"] os = "windows" isa = "*" abi = "*" export = true ``` `library` decouples source attribution from platform loader spelling. Give mutually exclusive platform entries the same logical value when they provide the same API; one unconditional `#[library("glfw")]` can then bind against `libglfw.so.3` on Linux, an install name on Darwin, and `glfw3.dll` on Windows. `symbols` names the symbols the dependency provides. On a two-level-namespace format (PE, Mach-O) every import must identify its provider, and `#[library]` can only attribute a symbol your Mach source declares. A vendored static archive leaves its own undefined references with no declaration to decorate, so the entry that provides them claims them. A symbol may be claimed only once per link. Whether an input links **statically** or **dynamically** follows the resolved file: a loose `.o` / `.obj` or static `.a` / `.lib` links statically, while ELF `.so`, Mach-O `.dylib`, and PE `.dll` inputs are recorded using their format canonical loader name. #### Build steps A `[step.]` is a command, make-recipe style, that produces files a build consumes, typically a `local` link input such as a vendored-C object. | Key | Req | Meaning | | --- | --- | --- | | `argv` | yes | Nonempty array of strings, spawned directly with no shell. `argv[0]` names the executable, by path or resolved on the planner `PATH`. Templates expand in every element. To invoke a shell, write `["sh", "-c", "..."]`. | | `env` | no | Table of string values added to the step process environment. | | `in` | yes | Declared input file list. Accepts globs (`*`, `**`), expanded sorted for a stable fingerprint. A glob that matches nothing is an error. | | `out` | yes | Declared output file list. Concrete paths only: a glob here is an error, since demand matching and caching expand `out` verbatim. | | `need` | yes | Array of `step.` requirements or `step.` globs this step must run after. Steps may require only steps. Cycles error. `[]` for none. | | `timeout` | no | Duration string (`"30ms"`, `"30s"`, `"5m"`, `"1h"`) after which the step process group is terminated and the build fails. Omit for an unbounded step. | Steps carry **no filters** and **never run automatically**. A step runs only when demanded: by a selected `[link.X]` whose `local` `path` matches the step's `out`, by another step's `need`, or by an artifact's `need`. Because a step has no filter of its own, the condition for running it lives in the link entry that demands it. A step is cached by content: its declared inputs, resolved executable, expanded arguments, and effective environment contribute to its fingerprint. An unchanged step whose outputs still exist is skipped. The fingerprint of the last successful run is kept as a stamp in `{project.out}/.cache/steps/`, and a step's outputs under `{project.out}` are written into scratch space in `{project.out}/.stage//` and published only once the step succeeds. A declared `out` inside `.cache/` or `.stage/` fails at manifest load. Every step process inherits the active build cell target tuple as `MACH_TARGET_ISA`, `MACH_TARGET_OS`, and `MACH_TARGET_ABI`. #### Dependencies Each `[dep.]` names a dependency realized under `dep//`. The key, the directory, and the dependency's own `[project].id` are one name. The build resolves a dependency purely by vendor layout: it reads that directory's own `mach.toml` for its `[project].id` and `src`. A module path whose head matches a dep `id` resolves into that dep tree. A stanza declares exactly one source, `git` or `path`, and a `git` stanza exactly one selector, `version` or `ref`. | Key | Meaning | | --- | --- | | `git` | Git URL. The dependency is a git submodule at `dep//`, pinned by the root repository's gitlink. Requires `version` or `ref`. | | `version` | A [version range](https://machlang.org/docs/dependencies.html#version-ranges) over the dependency's `vX.Y.Z` release tags, such as `"^5.3.0"`. Valid only with `git`. Resolved by `mach dep add` and `update`, never by a build. | | `ref` | Exact selector (with `git`): `branch/`, `tag/`, or `commit/`. Any other spelling is rejected. Cannot be combined with `version`. | | `path` | Local project tree, resolved relative to this manifest and never fetched. `mach dep` copies its files into `dep//`. Forbids `version` and `ref`. | Pins are the gitlinks committed in the root repository, and there is no lock file. `mach dep pull ` realizes the whole transitive closure flat at `dep//`, one level deep. Builds never fetch or resolve. They check offline that every checkout is clean at its pin, that each version-selected pin is a release inside every requirer's range, and that every [`mach` range](https://machlang.org/docs/manifest.html#compiler-range) in the closure accepts the compiler. > **Note:** The range grammar, how resolution picks releases, root narrowing and overriding, and everything a build verifies are covered in [Dependencies](https://machlang.org/docs/dependencies.html). #### Bare project-id imports A module path whose head segment matches a declared dependency's `[project].id` resolves into that dependency's source tree. The directory under `dep/` and the `id` match what source code writes. #### Path templates Paths and `argv` arguments expand over a closed, final set of eight variables: | Variable | Expands to | | --- | --- | | `{project.out}` | the **root** project expanded `[project].out`, in every manifest of the closure | | `{target.name}` | the resolved target name (never the literal `native`) | | `{target.isa}` | the resolved target `isa`, e.g. `x86_64` | | `{target.os}` | the resolved target `os`, e.g. `linux` | | `{target.abi}` | the resolved target `abi`, e.g. `sysv64` | | `{profile.name}` | the selected profile name | | `{artifact.suffix}` | the conventional filename suffix of the artifact being named, for its kind on the selected target (such as `.exe`, `.a`, `.so`). Available only in an artifact's own `out`. | | `{artifact..out}` | the output path of a required artifact, relative to the project root exactly as `{project.out}` is | An artifact's `out` is relative to the expanded project `out` and is rooted there automatically. Step `out` lists and local link `path`s are **not** auto-rooted: they name `{project.out}` explicitly, which is what homes a dependency's build products into the *consumer's* output tree rather than the dependency's checkout. > **Restriction:** There are no `{name}` / `{ext}` or bare `{target}` / `{profile}` aliases. An unresolvable `{...}` reference, or an unterminated `{`, is a strict-parse error. `{project.out}` is not available inside `[project].out` itself, and `{artifact..out}` is not available inside an artifact's own `out`. #### The build matrix A build cell is one artifact times one target times one profile. - Cells are selected with `-a/--artifact`, `-t/--target` and `-p/--profile`, one per axis. Each takes an exact name, which must be declared, or a glob with `*` and `?`, which must match, and repeats. `--all` fills every axis no option names with `*`: `mach build . --all` builds every artifact on every target it supports in every profile, and `mach test . --all -p debug` does the same in `debug` only. - An axis no option names, without `--all`, takes the manifest's default: the [native](https://machlang.org/docs/manifest.html#native) target (or the one a named artifact settles), the sole profile or the one marked `default = true`, and the default selection of artifacts, those marked `default = true` for the target when any is marked and every one of them when none is. No default is chosen by table order. - `mach build ` and `mach check ` take the whole default selection. `mach test ` takes it when it holds one artifact, and `mach run ` takes the sole `bin` the target builds; with several candidates and none marked, they refuse, naming every candidate. - A cell whose artifact does not list the cell's target is a cell the manifest never declared, so a glob that reaches it skips it. *Naming* that pair exactly (`-a kernel -t linux-x86_64`) is a different act and is refused by name, because you asked for a cell that does not exist. - Every cell is attempted; a failure does not abandon the ones after it. Each cell's diagnostics are reported under its own heading, and every cell that succeeded leaves its artifact on disk. - `-o` is accepted exactly when the selection resolves to a single build cell, and refused otherwise, naming the cells it resolved to. Artifacts cannot share an output path: a manifest whose expanded `out` templates collide is rejected before the build starts. ##### native target resolution `native` resolves the host's `(isa, os)` against the **declared** targets only, never a synthesized tuple. Exactly one host match is chosen; several matching tuples is an ambiguity error naming the candidates. With no match `native` is an error, raised before any step runs: a declared target that does not match the host is built only when `-t` names it, so a cross-only project with hosted targets selects its target with `-t`. #### Annotated example A project that cross-compiles to linux and windows, links a system library on each, and vendors a C object through a build step. Note where the conditions live: `[step.miniz]` carries no filter of its own, so it runs only because `[link.miniz]` demands it, and that entry is gated to `os = "linux"` since the host `cc` emits an ELF object. On the windows cell the entry filters out, the step is never demanded, and it never runs. To vendor C for the windows cell too, add a second step whose `argv` cross-compiles (it can branch on `MACH_TARGET_OS`) and a second link entry gated to `os = "windows"`. ``` [project] id = "demo" version = "0.1.0" mach = "^6.0" src = "src" out = "out/{target.name}/{profile.name}" [target.linux] isa = "x86_64" os = "linux" abi = "sysv64" default = true [target.windows] isa = "x86_64" os = "windows" abi = "win64" [profile.debug] opt = 0 debug = true simd = "scalarize" vectorize = false float_reassoc = false default = true [profile.release] opt = 2 debug = false simd = "scalarize" vectorize = true float_reassoc = false [step.miniz] # builds the vendored c object the link entry below demands argv = ["cc", "-c", "vendor/miniz.c", "-o", "{project.out}/miniz.o"] in = ["vendor/miniz.c"] out = ["{project.out}/miniz.o"] need = [] [link.miniz] source = "local" path = "{project.out}/miniz.o" # matches step.miniz's out, so it demands that step os = "linux" # host cc emits an ELF object, so this entry is linux-only isa = "*" abi = "*" export = false [link.kernel32] source = "system" name = "kernel32.dll" os = "windows" # skipped on the linux cell, never a per-platform file isa = "*" abi = "*" export = true [artifact.demo] kind = "bin" entry = "root.mach" out = "bin/demo{artifact.suffix}" targets = ["*"] link = ["miniz", "kernel32"] need = [] [dep.std] git = "https://github.com/briar-systems/mach-std" version = "^9.0" ``` #### See also - [CLI](https://machlang.org/docs/cli.html) - the flags that select a target, profile, and artifact - [Dependencies](https://machlang.org/docs/dependencies.html) - version ranges, release resolution, `mach dep`, and gitlink pins - [Project layout](https://machlang.org/docs/project-layout.html) - how `src` maps to module paths - [Decorators](https://machlang.org/docs/decorators.html#library) - `#[library]`, which resolves against a `[link.X]` identity Source: https://machlang.org/docs/dependencies.html ### Dependencies A mach project declares its dependencies in `mach.toml` and vendors them into a flat `dep/` tree. A git dependency selects releases by version range, or one exact commit, tag, or branch by ref. `mach dep` resolves and realizes that tree, and `mach build` then verifies it offline, with no network. #### Declaring a dependency A dependency is named by its **project id**, and that one name appears in three places: the manifest key `[dep.]`, the directory `dep//`, and the head segment of every module path the dependency exposes (`use .x;`). The compiler checks that all three agree, so `dep//mach.toml` must declare `id = ""`. A dependency whose id is the declaring project's own is refused where it is declared. ``` # in mach.toml [dep.std] git = "https://github.com/briar-systems/mach-std" version = "^5.3.0" # any release from 5.3.0 up to, not including, 6.0.0 [dep.gfx] git = "https://example.com/gfx" ref = "tag/v1.4.0" # one exact tag [dep.util] path = "../util" # a local project tree, never fetched ``` | Key | Type | Meaning | | --- | --- | --- | | `git` | string | Git URL. The dependency is a git submodule at `dep//`, pinned by the gitlink the root repository commits. Requires `version` or `ref`. | | `version` | string | A [version range](https://machlang.org/docs/dependencies.html#version-ranges) over the dependency's releases. Valid only with `git`. | | `ref` | string | Exact selector for `git`: `branch/`, `tag/`, or `commit/`. | | `path` | string | Local project tree, resolved relative to this manifest's directory. Never fetched. Forbids `version` and `ref`. | A stanza names exactly one source, `git` or `path`. A `git` stanza also names exactly one selector, `version` or `ref`, and naming both is an error (`[dep.std] names both 'ref' and 'version'; keep one`). `version` selects among releases. `ref` selects one exact tag or commit, or follows a branch. Any other `ref` spelling is rejected: there are no bare names, short hashes, or empty refs. A **path** dependency is copied into `dep//`, without the source's own `dep/`, its Git metadata, or its build output. When the source lives in a git work tree, only what git keeps is copied: tracked files and the untracked files its ignore rules keep. Copied files are not staged automatically. #### Version ranges `[dep.].version` and the compiler range [`[project].mach`](https://machlang.org/docs/manifest.html#compiler-range) share one range grammar: ``` range = clause *( "," clause ) ; the intersection of every clause clause = op partial op = "^" / "~" / ">=" / ">" / "<=" / "<" / "=" partial = major [ "." minor [ "." patch [ "-" pre ] ] ] ``` A version satisfies a range when it satisfies every clause. Each clause names its operator, so a bare `1.2` is refused (`a clause needs an operator, such as ^1.2 or >=1.2`). Whitespace is allowed around `,` and between an operator and its version, and nowhere else. A missing component is 0. | Clause | Means | | --- | --- | | `^1.2.3` | `>=1.2.3, <2.0.0` | | `^1.2` | `>=1.2.0, <2.0.0` | | `^1` | `>=1.0.0, <2.0.0` | | `~1.2.3` | `>=1.2.3, <1.3.0` | | `~1.2` | `>=1.2.0, <1.3.0` | | `~1` | `>=1.0.0, <2.0.0` | | `>=1.2`, `>1.2`, `<=1.2`, `<2` | the bound with missing components as 0, so `>1.2` is `>1.2.0` | | `=1.2.3` | exactly `1.2.3`. `=` needs all three components. | Below 1.0 a minor release is breaking, so caret fixes everything up to the first nonzero component: | Clause | Means | | --- | --- | | `^0.4.2` | `>=0.4.2, <0.5.0` | | `^0.4` | `>=0.4.0, <0.5.0` | | `^0.0.3` | `=0.0.3` | | `^0.0` | `>=0.0.0, <0.1.0` | | `^0` | `>=0.0.0, <1.0.0` | A pre-release version (`1.3.0-rc.1`) satisfies a range only when one of its clauses names a pre-release of that same release. So `^1.2` never selects `1.3.0-rc.1`, and `>=1.3.0-rc.1, <2` does. Build metadata (`+...`) is refused in a range and ignored in a release's version. There is no `*` and no `||`. #### Releases and resolution A **release** of a git dependency is a tag `vX.Y.Z` (or `vX.Y.Z-pre`) together with the `mach.toml` at that tag. A tag whose manifest's `[project].version` differs from the tag name is not a candidate. Resolution runs in exactly three places: `mach dep add`, `mach dep update`, and `mach dep outdated`. **Builds never resolve.** They [verify](https://machlang.org/docs/dependencies.html#verify), offline. For every identity in the closure that some manifest selects by `version`, resolution picks one release such that: 1. every requirer's range contains it, 2. its own `[project].mach` contains the running compiler, and 3. the closure its own manifest implies also resolves. Among the choices that satisfy all three, it takes the highest release of each identity, and writes the result as gitlinks like any other pin. `mach dep update ` keeps every other identity at its current release while that release still fits, so an update moves as little as it can. `--all` resolves from scratch. When nothing fits, the error lists every requirement that took part, including a release that needs a newer compiler, and names the identity the root can settle: ``` error: no set of releases satisfies every requirement: root requires b ^1.2 a 1.0.0 requires b ^2.0 the root decides by declaring the identity itself, for example: [dep.b] version = "" ``` Resolution never silently settles for a lower release than the ranges allow. - **What `add` writes.** With `--git ` alone, `version = "^X.Y.Z"` of the release resolution picked, so the lower bound is the release actually tested when the dependency was added. With `--version `, that range. With `--ref `, that selector. `mach init` adds std the same way. - **`--offline`.** Resolution normally reads candidates from each dependency's repository: one `git ls-remote --tags` per URL, and the manifest at a release through a shallow fetch of its tag. `update --offline` and `outdated --offline` use only the tags already present in the realized checkouts, and say so. A resolution that needs a candidate it doesn't have fails, naming the identity. - **`--lowest`.** `mach dep update --all --lowest` picks the lowest release every range accepts. A library's CI runs it in a scratch checkout and then builds and tests, which proves its declared lower bounds are honest. Run it in a release or manually dispatched job, never a scheduled one. - **`mach dep outdated `** prints, for each version-selected identity, the pinned release, the highest release resolution would pick now, and the highest release published. A newer release held back by a range or by the compiler is marked as such. - **No yanking.** A bad release is fixed forward with a new one. A consumer that must avoid one raises its range's lower bound (`^3.2.1`). - **Forks.** Identity is the project id, not the URL. A root that declares a fork's URL makes that fork the candidate source, so its `vX.Y.Z` tags compete under the same ranges. A fork that wants to stay distinguishable tags pre-releases (`v1.4.3-fork.1`), and a consumer opts in by naming the pre-release in its range. ##### Root declarations: narrowing and overriding A root `version` for an identity **narrows**: it is intersected with every requirer's range, and resolution and verification hold the pin to all of them. A root `ref` or `path` **overrides**: the requirers' ranges and selectors for that identity no longer apply, and `add` and `update` print each one they override. A range is never widened silently, and the escape hatch is one visible line in the root manifest. > **Overrides are unchecked:** Because the closure is flat, a requirer's `use b.*` binds to whatever the root selected, even a major that requirer was never built or tested against. A passing build only shows that the code it reached compiled. `mach dep verify` prints a note for every edge an override replaced, without failing. Treat each note as a claim to confirm. ##### A release selects only releases A dependency reached through a **release** (a `version` range or an exact `tag/`) may itself select dependencies only by `version` or `tag/`, so a release is reproducible from its tag all the way down. A dependency reached through `branch/` or `commit/` is in development, and its manifest may use any selector. A release that breaks the rule is refused wherever it is reached, naming the chain and the offending line, unless the root declares the offending dependency itself. `mach dep verify --release` holds the project itself to the same rule, so a library's release workflow catches the mistake before it tags. #### The mach dep command `mach dep ` manages the tree under `dep/`. Every action takes the project path first and accepts `--quiet` / `-q`. | Action | Effect | | --- | --- | | `add ` | Declare a dependency and realize the tree. Takes `--git ` with an optional `--version ` or `--ref `, or `--path `. Without a selector it writes the caret range of the highest compatible release. | | `update ( \| --all)` | The only command that moves a pin. Refreshes path copies, advances branch refs to their remote tips, moves exact refs to the commit they name, and resolves version ranges. Takes `--lowest` and `--offline`. Never edits the manifest. | | `outdated ` | Report each version-selected dependency's pinned, highest compatible, and latest release. Takes `--offline`. | | `pull ` | Realize the declared closure without advancing revisions or refreshing existing path copies. Idempotent. | | `verify ` | Check the realized closure against the manifests without changing anything, and print `ok`. `--release` also applies the [release rule](https://machlang.org/docs/dependencies.html#release-rule) to the project itself. | | `remove ` | Drop the entry from `mach.toml`. `--purge` also deletes `dep//`. | | `list ` | Print each declared entry with its source, ref, staged pin, and whether it is `realized` or `missing`. | ``` # add std at the caret range of its newest compatible release mach dep add . std --git https://github.com/briar-systems/mach-std # or state the range, or pin one exact tag mach dep add . std --git https://github.com/briar-systems/mach-std --version "^9.0" mach dep add . std --git https://github.com/briar-systems/mach-std --ref tag/v9.0.0 # see what newer releases exist, then move to them mach dep outdated . mach dep update . std ``` `pull` looks at what a git dependency's `dep/` holds together with its record (the staged gitlink, its `.gitmodules` entry, and any module directory Git retained) and takes the one step that brings it to what a build verifies. A staged gitlink with nothing checked out is initialized in place, a checkout at the wrong commit is checked out at its gitlink, a clean unregistered checkout is registered, and with neither the submodule is added at its selector. A symlink, a file, a directory that is not a checkout of its own, and a dirty checkout are refused and left alone. A version-selected dependency that has no pin yet is refused, with the `mach dep update` command to run. `update ` looks the name up in the whole closure, so it can re-pin a transitive dependency the root does not declare. For an identity reached by more than one path, the root's selector wins if the root declares it. Otherwise the requirers must agree, or the command stops, prints both chains, and names the root declaration that would decide. #### Pins are gitlinks The record of which commit a dependency is at is the **gitlink** committed in the root repository, generated into `.gitmodules` by `mach dep`. Nothing else records a pin. There is no lock file, and a `mach.lock` in the project root is an unrelated file no command reads. The verifier reads the git index, so a freshly realized dependency is verifiable before it is committed. A project does not need its own Git repository. A project root is identified by its own `mach.toml`. In a repository root, git dependencies use the staged gitlinks as their pins. In a plain directory, or a project nested inside an unrelated repository, they are plain clones whose own checkout commits are verified. Path dependencies are verified from their copies. #### The flat dep tree The root's `dep/` holds every identity in its **transitive** closure, one directory each, one level deep. The root manifest declares only what the root uses directly, plus any override. A dependency's own dependencies reach the root's `dep/` through closure computation and need no declaration in the consumer. With a root that declares `a`, and `a` that declares `b`, the layout is `dep/a/` and `dep/b/`, and `a`'s `use b.*` resolves against the root's `dep/b/`. Git materializes `a`'s own gitlink as an empty `dep/a/dep/b/` that is neither realized, verified, nor descended into. `dep` itself must be a physical directory, not a symlink. The build resolves a dependency purely by vendor layout: it reads that directory's own `mach.toml` for its `[project].id` and `[project].src`, and routes any module path whose head segment matches a dep's `id` into that tree. ##### One identity, one commit One identity resolves to exactly one commit per build, with no exception for majors. Two different selections of one identity is a **clash**, and the diagnostic prints both chains and the root declaration that would resolve it: ``` error: dependency conflict: project id 'b' is reached with two different selections: root -> a -> b requires git @ tag/v1.0.0 root -> c -> b requires git @ tag/v2.0.0 the root decides by declaring the identity itself, for example: [dep.b] git = "" ref = "tag/v2.0.0" ``` The root's declaration may point at upstream or at a fork carrying the same id, with no change to any consumer. Two unrelated packages claiming one id is a collision and is rejected. URL disagreement is a mirror, not a conflict: the root's declared URL wins, else the first declaring path's, and verification compares commits, never URLs. ##### Removed forms `pull`, `verify`, and every build refuse two shapes of a realized closure: - a **key that is not the project id**, such as `[dep.foo]` whose realized project declares `id = "std"`. Rename the table to `[dep.std]` and the directory to `dep/std`. - a **nested realization**, a `dep//dep//mach.toml`. The root's `dep/` owns the flat closure, so delete `dep//dep`. The empty directory git materializes for a consumed dependency's own gitlink is not a realization and passes. #### What a build verifies Builds never fetch and never write under `dep/`. Every build, and `mach dep verify`, checks offline that: 1. every git dependency is a clean checkout at its applicable pin, and every path dependency is a contained tree without repository metadata, 2. its project id equals its directory name, 3. the closure computed from the realized manifests equals the set of directories under `dep/`, with nothing missing and nothing extra, 4. there are no cycles, 5. every realized manifest's [`[project].mach`](https://machlang.org/docs/manifest.html#compiler-range) accepts the running compiler, and 6. for every identity selected by `version`, the pinned commit carries a release tag that is inside every requirer's range, the root's included, and every release in the closure [selects only releases](https://machlang.org/docs/dependencies.html#release-rule). A pin outside a range names the requirer chain, the range, the pinned release, and the `mach dep update` command that re-pins it. For an identity the root does not declare, every requirer's exact `tag/` or `commit/` selector must match the realized commit. A root `ref` or `path` is the override, so the root's gitlink is its pin. A `branch/` selector is an input to `update`, never something a build checks. > **Note:** `mach build` never requires the network: a project whose dep tree is present builds on a bare machine. Only `add`, `update`, `outdated`, and `pull` reach a remote, through git discovered on `PATH`. A failed git command names the command, its exit status, and git's own message. #### Importing a dependency Once vendored, a dependency's modules are imported by their full path: the head segment is the dep's `[project].id`, and the rest mirrors its source tree. ```mach use std.print; use std.types.result.res; ``` A one-segment `use` path equal to a dependency's id binds that project's public module. For a dependency, that is the `entry` shared by its library artifacts (`static` or `shared`) marked `default = true`. ``` # in the dependency manifest: [artifact.lib] kind = "static" entry = "mylib.mach" out = "lib/libmylib{artifact.suffix}" targets = ["*"] link = [] need = [] default = true ``` ```mach # in a consumer: use mylib; # binds mylib.mach from the default library artifact ``` A dependency that declares no default library artifact, or whose default library artifacts name different entries, has no public module. A bare import of it is an error naming the missing declaration, and callers import by full module path instead. #### Cascading link requirements A dependency declares its own external link requirements as `[link.X]` entries. An entry marked `export = true` also applies to any project that links this dependency's modules, so a platform link requirement lives once in the manifest that needs it and out of every consumer. A standalone build and a consumed build read the same entries, so nothing behaves differently as a dependency. ``` # in the dependency's manifest: [link.kernel32] source = "system" name = "kernel32.dll" os = "windows" # filtered out on any non-windows cell isa = "*" abi = "*" export = true # and cascades to every consumer ``` An entry applies to a build cell when all three of its `os` / `isa` / `abi` axes match. Target names are local to each manifest and are never matched on. Across dependencies the order is topological, then declaration order, and the union is deduplicated. #### See also - [Manifest](https://machlang.org/docs/manifest.html) - the full `mach.toml` reference, including `[project].mach`, `[dep.*]`, and `[link.*]` - [CLI](https://machlang.org/docs/cli.html) - every `mach` command, including `mach dep` actions - [Modules](https://machlang.org/docs/modules.html) - how source files map to importable module paths - [Project layout](https://machlang.org/docs/project-layout.html) - where `dep/` and the manifest live in a project Source: https://machlang.org/docs/testing.html ### Testing A `test` is a named block of statements the runner can execute on its own. Tests live inline with the code they exercise, and `mach test` collects every one in the artifact under test, builds it, and runs it. #### Declaring a test A test is a declaration at module scope - the same level as `fun`, `rec`, and `val`. It takes no parameters and is not callable from ordinary code; it exists only for the runner to invoke. ```mach test { ... } ``` The name is an identifier, following the ordinary identifier rules. A string in its place (`test "label" { ... }`) is a compile error that names the identifier form. The body is an ordinary block that may use anything in scope in the enclosing module, just like a function body. ```mach test is_leap_year__centuries { if (!is_leap_year(2000)) { ret 1; } if (is_leap_year(1900)) { ret 1; } if (is_leap_year(2023)) { ret 1; } ret 0; } test log__nil_message_does_not_crash { debug(nil); info(nil); ret 0; } ``` Tests live in their own namespace in each module. A test's name is never in scope in code, so `test str_len { ... }` and `fun str_len` in one module do not conflict, and no code can name, call, or reference a test. Two tests with the same name in one module are a compile error. Related tests group under a common subject as `subject__case` (`str_len__empty`, `str_len__multibyte`), and a regression test is named `regression__*`. Both are conventions the compiler does not check. Every build resolves and type-checks each test body in the modules its artifact reaches, and ordinary builds omit test bodies from IR and object files. > **Note:** A `pub` modifier is syntactically accepted before `test` but carries no meaning - a test is never part of a module's public surface. ##### Qualified names A test's **qualified name** is its module path, `#`, and its name: `std.types.string#str_len__empty`. `#` appears in no identifier or module path, so a qualified name never collides with another symbol. It is the test's symbol, and it is the name `--list`, `--filter`, the readout, and `--format json` show. A debugger takes it unquoted: `break std.types.string#str_len__empty` in gdb. #### Test-only helpers A helper or fixture that exists only for tests is marked [`#[testing]`](https://machlang.org/docs/decorators.html#testing). It gets the same treatment as a test body: checked in every build, omitted from ordinary ones, and emitted under `mach test`. Only a test body or another `#[testing]` declaration may reference it. ```mach #[testing] fun sample_year() i64 { ret 2000; } test is_leap_year__sample_year { if (!is_leap_year(sample_year())) { ret 1; } ret 0; } ``` #### Reporting pass and fail A test body is checked against an `i32` return type, and each `test` lowers to a zero-parameter function whose result becomes the test process's exit status, in `0..255`. A return value of `0` indicates success, while any value in `1..255` signals failure and is reported as `(exit N)`. A `ret` whose value is a literal (or a literal-shaped expression) outside `0..255` is a compile error. A value computed at run time that lands outside the range is folded to `255` and always fails, so a result of `256` can never read as a pass. ```mach ret 0; # pass ret 1; # fail - any non-zero ordinal indicates failure ``` Falling off the end of a body returns `0`, treating an empty or fallthrough body as a pass. Use positive integers to distinguish which assertion or condition failed in your test. #### Where tests live Tests are not tied to a single file. Write `test name { }` declarations directly in the relevant `src/` module, or in modules of their own. `mach test` builds exactly what `mach build` builds for the same artifact - the closure its entry reaches through `use` and `fwd` - and runs the tests declared there. A module no selected artifact reaches is not loaded under test either, so its tests do not run. A module that exists only for tests, such as a suite that exercises several modules together, is reached by no artifact on its own. Give such modules a **test artifact**: an ordinary library artifact whose entry `use`s each of them, tested by name. ``` # the test-only modules no other artifact reaches [artifact.tests] kind = "static" entry = "test/all.mach" out = "lib/tests" targets = ["*"] link = [] need = [] ``` ```mach # src/test/all.mach use std.runtime; use app.test.parser; use app.test.roundtrip; ``` `mach test .` then runs the tests the default artifact reaches, and `mach test . -a tests` the test-only suites. The entry reaches the runtime's startup (`use std.runtime;`) because a library artifact's closure is all the test dispatcher links. Only the current project's own tests are collected by default, so a dependency's suite never runs (or fails) as part of your `mach test`. `--include-deps` collects dependency tests as well, which is useful when working on a dependency in-tree. #### Running with mach test ``` mach test [options] ``` `mach test` is `mach build` with a different goal. It builds the artifact's objects exactly as `mach build` does, adds a test object beside each module that declares tests, links one test **dispatcher** executable covering the selected tests, then runs each of them as its own process and reports the results. A crashing test reports its signal and the run continues. A test build always links an executable, even for a library artifact. The readout is per module: a module whose tests all pass collapses to one roll-up line, and a module with a failure expands to show the failing test's location, its exit code, signal, or timeout, its captured output, and the exact `rerun:` command. The run closes with a summary that re-lists every failure: ``` failures: app.main#fails_on_purpose src/main.mach:11 (exit 3) 1 passed, 1 failed, 2 total (1ms) ``` ##### Test objects and the dispatcher A module that declares tests or `#[testing]` declarations gets a **test object**, `obj//.test.o`, beside its normal object. The test object holds the module's tests under their qualified names and only what the module's object lacks, so the module's object never changes for tests. After `mach build`, `mach test` recompiles no module object, and editing a test recompiles only its module's test object and relinks. Each run links a small dispatcher of only the selected tests, `test//` under the output directory, and spawns it once per test. Each test's stdout and stderr are captured to a file under `test//log/`. A passing test's file is removed on the spot, and a failing test's file stays. ##### Selecting and listing | Flag | Effect | | --- | --- | | `--filter ` | run only tests whose qualified name contains `` | | `--include-deps` | also run tests declared in dependency modules | | `--list` | list the collected tests and exit, running nothing | | `--jobs ` | run up to `n` test processes at once (default: host CPUs) | | `--timeout ` | terminate a test and its process group after the duration (default unbounded) | | `--format ` | `human` (default), or a `json` NDJSON event stream | | `--runner ` | launch every test through `` instead of exec'ing it directly | ``` mach test . # build and run every test mach test . --filter date # only tests whose qualified name contains date mach test . --list # enumerate tests, run nothing mach test . --timeout 30s # bound each test to 30 seconds ``` `--filter` selects before the dispatcher links, so the dispatcher holds only the selected tests and what they reach, and changing the filter relinks without recompiling. `--list` prints each selected test's qualified name and the test object that holds it, and exits without linking or running anything: ``` app.parser#rejects_trailing_comma ./out/linux-x86_64/debug/obj/app/parser.test.o ``` The `build` and global flags apply too, and `-v` reports build phase timing as `mach build -v` does. `--runner` names a host-side launcher for foreign-target tests the host cannot exec directly, e.g. `mach test . -t windows --runner wine`, and needs the selection to resolve to one cell. Without it, only a target whose `os` and `isa` are the host's runs its tests: any other is built and reported on a `skip` line, whatever emulation the host has. A run in which nothing was runnable exits `1`, so a green run always ran something. ##### Timeouts `--timeout ` bounds each spawned test process independently, from its own spawn, on its whole process group, so a process the test started dies with it. A test that exceeds the bound is the distinct outcome **timed out**: it renders as `(timed out after )`, is counted separately on the summary line, and still fails the run. ``` failures: app.main#spins src/main.mach:4 (timed out after 1s) 0 passed, 1 failed (1 timed out), 1 total (1.0s) ``` `` is a positive integer followed by a unit: `ms`, `s`, `m`, or `h` (`30ms`, `30s`, `5m`, `1h`). A bare number, a fraction, zero, or any other unit is a usage error naming the accepted forms. There is no default: omitting the flag leaves every test unbounded. ##### JSON output `--format json` replaces the readout with one JSON object per line on stdout: `run_start`, one `test` per result, and `summary`, or one `case` per test under `--list`. Build diagnostics stay on stderr. A `test` or `case` event names its test by qualified name in `name`, beside its `module`, `file`, `line`, test `object`, and dispatcher `index`. A timed-out test reports `"kind":"timeout"` with its bound in nanoseconds in `timeout_ns`. The schema is versioned, `"schema":2` on every event. ##### Exit codes - `0` - every test that ran passed. - `1` - at least one test failed, was killed by a signal, or timed out; also a user error such as an unknown flag. - `2` - a build or internal error before the tests could run, or a test that failed for an infrastructure reason. #### See also - [CLI](https://machlang.org/docs/cli.html) - the full `mach` command surface, including `mach test` - [Decorators](https://machlang.org/docs/decorators.html#testing) - `#[testing]` for declarations that exist only for tests - [Functions](https://machlang.org/docs/functions.html) - a test body is checked like a function body - [Statements](https://machlang.org/docs/statements.html) - `if`, `ret`, and the blocks a test body uses - [Project layout](https://machlang.org/docs/project-layout.html) - where `src/` modules and tests live ## Language reference Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/README.md ### Mach language reference Per-element reference docs. Each file covers one language component; read the index below or follow `see also` links to navigate. This directory is the authoritative reference for the Mach 5 dialect. Each file is a focused doc with grammar, examples, and neighboring links; start from the index below. #### Files and structure - [files.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/files.md) — extensions, `lib.mach` / `main.mach` conventions - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) — module tree, path separator, shadow-module pattern - [use.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/use.md) — imports - [fwd.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fwd.md) — re-exports #### Declarations - [visibility.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/visibility.md) — `pub` and `ext` modifiers - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — codegen decorators, `#[name]` (`symbol`, `library`, `inline`, `align`, `section`, `embed`) - [def.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/def.md) — type alias - [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md) - record - [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md) - raw union - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - tagged value - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) - function - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) - external function - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) - immutable and mutable bindings - [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md) - test declaration and the mach test workflow - [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) - variadic packs (va: ..., $each, va.len, va...) #### Values and types - [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md) - numeric, char, string - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) - primitive grammar, compound types, tag types and the std failure tags - [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) - the ^ secret qualifier, flow typing, gates, :>T - [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) - arithmetic, bitwise, comparison, logical, pointer, cast - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) - construction, access, calls, generic instantiation, `sel` #### Control flow - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) - if/or, for, ret, brk, cnt, fin, blocks, payload guards #### Comptime channel - [comptime.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime.md) - channel overview - [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md) - $mach.* compiler-owned namespace - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) - codegen decorators, #[name] (replaces the removed $sym.attr setters) - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) - $size_of, $length_of, $align_of, $offset_of, $type_of, $fields, $cases, $is_tag, $discriminant_of, $pointee_of, $is_record, $is_union, $is_pointer, $is_integer, $is_float, $is_secret, $holds_secret, $type_name, $each, $error - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) - $if / $or #### Low-level - [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md) — inline assembly - [policy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md) — compiler vs stdlib boundary #### Diagnostics - [diagnostics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md) — diagnostic keys, the code registry and its never-reused rule - [diagnostics-json.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md) — `--diagnostics=json`, the versioned NDJSON record schema #### Formal grammar - [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md) — full EBNF grammar of the implemented dialect, derived from the lexer and parser #### Conventions - [documentation.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/documentation.md) — docstring style for functions, types, modules, and values - The ```` ```mach ```` blocks on these pages are compiled by `test/run.sh --docs`. A plain block compiles, and runs when it declares a main. ```` ```mach fragment ```` marks a block that is not a whole program, and ```` ```mach error ```` a block whose compile fails with `` in the output. A block showing several files starts each with `# file: src/.mach`, in a project whose id is `example`. See [test/README.md](https://github.com/briar-systems/mach/blob/v6.10.1/test/README.md#doc-blocks). #### Build system - [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md) — the `mach.toml` manifest reference - [ir-output.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ir-output.md) — `--emit-ir` and its two forms, and which one is stable for tooling - `mach --help` and `mach help ` — the command-line reference The supported, source-stable surface of the compiler is the editor API (`mach.lang.editor`, documented by its docstrings and rendered by `mach doc`), the command line, and the manifest schema. Everything else under `src/` is internal. Source API authors can mark deprecated declarations with [`#[deprecated]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#deprecated--deprecatedstr--source-use-notice), which warns on external use. Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/files.md ### Files Source files use the `.mach` extension. A project has a `mach.toml` at its root, a source directory the manifest names (`src` by convention), and one entry module per artifact, named by the artifact's `entry` key. The compiler attaches no meaning to any file name: `mach init` scaffolds `src/main.mach` for a binary and `src/lib.mach` for a library, and either basename is arbitrary. An artifact's build is rooted at its entry module and follows `use` and `fwd` edges from there; a file under `src` that no entry reaches is not part of that artifact, under `mach build`, `mach check` and `mach test` alike. Tests that live in modules no artifact reaches belong to a test artifact of their own (see [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#which-tests-run)). #### Executable entry The compiler does not special-case a `main` function. An executable's entry is whichever function **exports the linker symbol** `main`, tagged with [`#[symbol("main")]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) and matching the runtime's expected signature: ```mach use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { ret 0; } ``` `std.runtime` provides the platform-specific `_start` symbol the linker uses as the true process entrypoint; `_start` decodes `argc`/`argv` and calls whatever function exports `main`, then terminates the process with its returned exit code. `use std.runtime;` is **required** to link `_start` into the binary, even though no code references it by name. Only the exported `main` symbol binds the entry — the Mach-level function name is irrelevant, so `fun entry(...)` tagged `#[symbol("main")]` works identically. #### mach.toml The project manifest. It declares the project's identity, its targets, its profiles, its artifacts, and its dependencies; every table is required to be complete. The full reference is [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md). A minimal binary project, as `mach init` writes it for one target: ```toml [project] id = "myproj" version = "0.1.0" mach = "^5.3" src = "src" out = "out/{target.name}/{profile.name}" [target.linux-x86_64] isa = "x86_64" os = "linux" abi = "sysv64" [artifact.myproj] kind = "bin" entry = "main.mach" out = "bin/myproj{artifact.suffix}" targets = ["*"] link = [] need = [] [profile.debug] opt = 0 debug = true simd = "scalarize" vectorize = false float_reassoc = false [dep.std] git = "https://github.com/briar-systems/mach-std" ref = "branch/main" ``` The `id` is the root of every module path the project exposes. A file at `src/foo/bar.mach` is reachable as `myproj.foo.bar`. The dependency `std` is the standard library, realized as a git submodule at `dep/std/` and addressed as `std.*` in source. #### See also - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) — how files map to module paths - [use.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/use.md) — referencing modules from other modules - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — `#[symbol("main")]` and the linker-name override Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md ### Modules A Mach project is a tree of modules rooted at the project's `id`. Each `.mach` file under the project's source directory is a module reachable by a dotted path from the project root. #### Path structure The path separator is `.`. A file at `src/foo/bar.mach` in a project with `id = "myproj"` is reachable as `myproj.foo.bar`. There is no `this.` self-prefix. Within a project, modules always reference each other by their full project-rooted path. Every reference is syntactically uniform regardless of where it appears. #### Bare project-id imports A one-segment `use`/`fwd` path equal to a resolvable project id — a dependency's id or the current project's own id — binds that project's **public module**. For a dependency that is the `entry` shared by its library artifacts marked `default = true`: a library that declares one `static` artifact with `default = true` and `entry = "lib/libstd.mach"` gives `use std;` the module `std.lib.libstd`. Several default `static`/`shared` artifacts may share that entry; a `bin` never publishes one. For the current project it is the selected artifact's entry. A dependency with no default library artifact, or whose default library artifacts name different entries, has no public module, and a bare import of it is an error (`project 'x' has no public module: a bare import binds the entry shared by its library artifacts marked `default = true`; import a full path, or mark one static or shared [artifact.*] table (or several sharing one entry) default = true in its manifest`). Longer paths are unaffected: `use std.print;` needs no default artifact. A dependency that declares no artifact at all has no public module either, and the refusal says so (`project 'x' declares no artifact, so it has no public module: import a full path, or declare a static or shared [artifact.*] table marked default = true in its manifest`), so a library declares its artifact. #### Shadow-module pattern A file `foo.mach` may co-exist with a directory `foo/`. The file is the **surface** module — the public face of `foo`. The directory's files are **split** implementations that the surface loads and re-exports. ``` myproj/ ├── foo.mach # surface └── foo/ ├── a.mach # split: myproj.foo.a └── b.mach # split: myproj.foo.b ``` The surface loads each split with `use myproj.foo.a;` and re-exports its public symbols with `fwd a.X;`. Consumers `use myproj.foo;` and access symbols through the surface — they never name the split files directly. ```mach # file: src/foo/a.mach pub fun one() i64 { ret 1; } # file: src/foo/b.mach pub fun two() i64 { ret 2; } # file: src/foo.mach use example.foo.a; use example.foo.b; fwd a.one; fwd b.two; # file: src/main.mach use example.foo; use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", foo.one() + foo.two()); ret 0; } ``` Two common uses: - **Topical splits** — organize a large module by topic; all splits forwarded unconditionally. - **Multiplatform splits** — one impl per target, selected by `$if` on `$mach.build.os` or `$mach.build.arch`. #### See also - [files.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/files.md) — how `mach.toml` declares the project root - [use.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/use.md) — loading a module into another module's scope - [fwd.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fwd.md) — re-exporting from a surface module Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/use.md ### `use` — imports `use` brings an external symbol or module into the current scope under a local name. It is a private import; the imported name is not exposed to consumers of this module. #### Grammar ```mach fragment use PATH; # binds the leaf component use ALIAS: PATH; # binds ALIAS ``` - The alias defaults to the path's leaf component when not given. - One name per line. No splat (`use foo.*` does not exist) and no combined forms (`use foo.{a, b, c}` does not exist). The resolver binds whatever the path points to. A path ending at a **module** binds the module — access its members with qualified `module.member`. A path ending at a **symbol** binds the symbol for bare use. Importing a module does **not** pull its members into scope unqualified; to use `usize` bare, import the symbol, not its module. #### Examples ```mach use std.types.size; # binds module 'size'; use as `size.usize` use sz: std.types.size; # binds module under 'sz'; use as `sz.usize` use std.types.size.usize; # binds symbol 'usize'; use bare as `usize` use my_usize: std.types.size.usize; # binds symbol under 'my_usize' val a: size.usize = 1; val b: sz.usize = 2; val c: usize = 3; val d: my_usize = 4; ``` ```mach fragment use mylib; # bare project id: binds mylib's public module ``` A one-segment path equal to a resolvable project id binds that project's public module, the `entry` shared by its library artifacts marked `default = true` — see [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md#bare-project-id-imports). A dependency that marks none has no public module and is imported by full path only; std 2.0.0 is one (`use std;` is refused naming the rule, while `use std.print;` needs no default artifact). #### Design rule A Mach module imports every dependency it directly names — including any dependency reached only through a re-export. There is no shortcut for "my surface uses these transitively; just give me the leaf." The dependency graph is visible at the top of every file. #### See also - [fwd.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fwd.md) — the public-re-export counterpart - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) — how paths form module identifiers Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fwd.md ### `fwd` — re-exports `fwd` re-exports a symbol or module from another module under this module's public surface. It is the public counterpart to `use`. #### Grammar ```mach fragment fwd PATH; # re-export under the path leaf fwd ALIAS: PATH; # re-export with rename ``` - `fwd` always publishes; there is no `pub fwd` form. - One name per line. No splat. Mirrors `use` grammar exactly. #### Examples ```mach fragment use impl: full.core.data; fwd impl.Point; # re-exports as 'Point' fwd Pt: impl.Point; # re-exports as 'Pt' ``` The surface file in a shadow-module pattern is typically a long list of `fwd` lines — one per symbol the surface exposes. #### Module re-exports A `fwd` path that ends at a **module** re-exports the whole module as a public module alias, mirroring `use`'s module binding: ```mach fragment fwd demo.alpha; # re-exports module 'alpha' fwd deep: demo.deep.beta; # re-exports module under 'deep' ``` A consumer reaches the alias's members with qualified access, chaining through any depth of re-export — including a `fwd` of another library's `fwd`. In a project whose `[project] id` is `example`: ```mach # file: src/alpha.mach pub fun answer() i64 { ret 42; } # file: src/lib.mach fwd example.alpha; # re-exports module 'alpha' # file: src/main.mach use example.lib; # lib.mach contains `fwd example.alpha;` use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", lib.alpha.answer()); # resolves through the module re-export ret 0; } ``` As with `use`, a module alias is not a value; only its members can be named. #### When to use - Composing a public surface from split implementation files. - Aliasing platform-specific impls under a stable name for consumers. #### See also - [use.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/use.md) — the private-import counterpart - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) — shadow-module pattern Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/visibility.md ### Visibility — `pub` and `ext` Two declaration modifiers control how a symbol is seen. #### `pub` Marks a declaration as part of its module's public surface. Other modules that `use` this module can reference `pub`-marked symbols by name; symbols without `pub` are file-private. ```mach error no symbol `helper` exported by `example.lib` # file: src/lib.mach pub fun add(a: i64, b: i64) i64 { ret a + b; } # private: only callable inside this file fun helper() i64 { ret 1; } pub rec Point { x: i64; y: i64; } pub val MAX: i64 = 100; # file: src/main.mach use example.lib; fun sum() i64 { ret lib.add(lib.MAX, 1); # fine: pub } fun peek() i64 { ret lib.helper(); # error: helper is not exported } ``` Applies to: `fun`, `rec`, `uni`, `def`, `val`, `var`, `ext fun`, `ext val`, `ext var`. `fwd` always publishes and does not take an explicit `pub` modifier. #### `ext` Declares a function or a data binding whose definition lives in another object, as a forward reference the linker resolves. An `ext fun` has no body and follows the C ABI; an `ext val` / `ext var` has no initializer and no storage of its own. ```mach #[symbol("write")] pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ext var errno: i32; ``` - The C ABI is the contract; argument and return types must be representable in C. - Use the `#[symbol("real_name")]` decorator to override the linker name. There are no body-less functions outside of `ext fun`. Regular forward declarations do not exist. #### What a shared library exports A `kind = "shared"` artifact exports the **root project's** `pub` declarations and nothing else. A dependency's `pub` surface is that dependency's ABI, not this library's, so a `pub fun` in a dependency is reachable from your code and absent from your library's export table. Every other definition — anything without `pub`, and every definition the compiler synthesizes, such as a generic or pack instance — is hidden: other modules in the same link resolve it normally, and no consumer of the linked library can bind to it. `#[symbol("name")]` chooses the linker name, not the visibility. A non-`pub` declaration with a `#[symbol]` name is still hidden, so a C program cannot link against it; a function a C caller links against is `pub`. A `fwd` re-export is a declaration of surface, so what a root-project module re-exports is exported, wherever it is defined: ```mach fragment # file: src/lib.mach fwd impl.helper; # exported: this module published it fwd other.module; # exported: that module's whole public surface ``` A `fwd` of a module exports that module's whole public surface, following its own `fwd`s in turn. A `fwd` of a generic, comptime-parameter or pack declaration exports nothing, because such a declaration has instances rather than one definition and each consumer instantiates its own. A shared library build warns at such a `fwd`, naming the declaration; the way to export one instantiation is a `pub` non-generic wrapper around it. A `fwd` of a module that holds generics warns for none of them, since it asked for the module's exportable surface, and the same `fwd` in an executable or static library build says nothing. A `fwd` of a type or of an `ext` import exports nothing either, since neither defines a symbol in the image. A shared library that exports nothing is refused: it would be callable by nobody, and dead-code elimination would leave it empty. An executable exports nothing at all, so `pub` makes no difference to one. #### See also - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — regular function declarations - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) — full reference for `ext fun` - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md#ext--foreign-data-imports) — `ext val` / `ext var` - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — `#[symbol]` and the other decorators Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md ### Decorators A decorator attaches metadata to a declaration. It can provide source-use notices or influence how the compiler emits the symbol: its linker name, alignment, section placement, inlining, dynamic import attribution, constant-time obligations, or exclusion from auto-vectorization. Visibility (`pub` / `ext`) is separate and unaffected by decorators. #### Surface A decorator is written as an attribute: ``` #[name] # bare flag (e.g. inline) #[name(args)] # directive with comptime-expr arguments ``` > `#[...]` is the only decorator surface. A backtick is not a token: one > anywhere in source is a lexer error. > One caveat the attribute form introduces: a line comment that begins `#[` > (with no space) opens an attribute. Write such a comment with a separating > space — `# [...]`. #### Grammar ``` #[deprecated] # external uses warn #[deprecated("msg")] # external uses warn with this message #[testing] # exists only for tests; omitted from ordinary builds #[expect("key", ..)] # acknowledge the named warnings inside this declaration #[symbol("name")] # linker name override #[library("dep")] # dynamic import attribution (ext only) #[inline] # force inlining (no arguments) #[noinline] # forbid inlining (no arguments) #[align(expr)] # alignment; expr is a comptime integer #[packed] # lay a rec / uni out with no padding (no arguments) #[volatile] # every access to a rec / uni / tag is a volatile access (no arguments) #[section(".name")] # place in a named object section #[oblivious] # constant-time boundary (no arguments) #[scalar] # opt out of auto-vectorization (no arguments) #[naked] # no prologue or epilogue; body as written (no arguments) #[extensions(a, b)] # the function may use these instruction-set extensions #[embed("path")] # compile-time file embedding (val only) #[stage("name")] # GPU pipeline stage; makes the function an entry point #[workgroup(x,y,z)] # compute workgroup dimensions, constants or #[spec] vars (with #[stage("compute")]) #[input(n)] # shader interface input at location n (global only) #[output(n)] # shader interface output at location n (global only) #[builtin("name")] # pipeline built-in variable (global only) #[uniform(set, bnd)] # descriptor-bound uniform block, read-only (global only) #[storage(set, bnd)] # descriptor-bound storage buffer, read-write (global only) #[storage(set, bnd, "readonly")] # the same buffer, with writes refused #[storage(set, bnd, "writeonly")] # the same buffer, with reads refused #[storage(set, bnd, "coherent")] # the same buffer, its writes visible across workgroups #[sampler(set, bnd)] # descriptor-bound image / sampler handle (global only) #[push] # push-constant block, read-only (global only) #[spec(id)] # specialization constant the host supplies (module var only) #[shared] # compute workgroup memory, zero when a stage starts (global only) #[op(tgt,set,name)] # the target instruction this function is (bodyless fun only) #[handle(tgt,ctor,..)] # the target type this declares (bodyless def only) #[abi_type("name")] # a C type whose layout the target declares (bodyless def only) ``` Decorators appear **before** the declaration they target, one per line or space-separated on the same line. They attach to the immediately following declaration only and do not bleed across declarations. ```mach fragment #[inline] #[symbol("big")] fun big(a: i64, b: i64) i64 { ... } #[align(64)] #[symbol("g_lit64")] pub var g_lit64: u8 = 7; ``` Each directive is wrapped in its own clause: `#[name]` for a bare flag or `#[name(args)]` for a directive that takes arguments. Arguments are comptime expressions. ##### Constant arguments Where a directive takes a string or an integer, the argument is a constant expression of that type, evaluated at compile time in the declaring module. A literal is one. So is a `val`, one an `$if` arm selects, or one imported from another module. A directive's argument never depends on the build that evaluates it, only on the constants in scope. `#[symbol]` stays exactly the name you give it: the compiler adds no platform decoration of its own, and a `val` gated by comptime is how one declaration names its symbol per target. ```mach $if ($mach.build.os == $mach.os.darwin) { val SPIN: *u8 = "_spin"; } $or { val SPIN: *u8 = "spin"; } #[symbol(SPIN)] pub fun spin() i64 { ret 0; } ``` An argument that is not a constant expression, such as a `var` or a call, is refused with `decorator.not_constant`, naming the directive. A constant of the wrong type is refused with `decorator.argument`. #### Directives ##### `deprecated` / `deprecated(str)` — source-use notice Marks a declaration deprecated. The optional argument is one constant string carrying a message. Repeating the attribute, giving it more than one argument, or giving it an argument that is not a constant string is an error. The attribute changes nothing about visibility, type identity, ABI or codegen. A use of the deprecated declaration from another source module warns at the identifier, carrying the message. Value references, calls and type references are covered, including through imports, re-exports and generic instantiation, and each source site warns once even when a generic body is instantiated more than once. The declaring module does not warn on its own uses, and an unused import alone produces no warning. ```mach # file: src/legacy.mach #[deprecated("use replacement")] pub fun old() i32 { ret replacement(); } pub fun replacement() i32 { ret 1; } # file: src/main.mach use example.legacy; fun caller() i32 { ret legacy.old(); # warning: `old` is deprecated: use replacement } ``` It applies to `fun`, `rec`, `uni`, `tag`, `def`, `val`, `var`, `use` and `fwd` declarations, and to a tag case, where it is the only decorator a case accepts: ```mach # file: src/reply.mach #[deprecated("the whole tag")] pub tag Old: u8 { empty; } pub tag Reply: u8 { empty; #[deprecated("use fresh")] value: i64; fresh: i64; } # file: src/main.mach use example.reply; fun read(r: reply.Reply) i64 { if (sel r.value) { ret r.value; } # both sites warn: tag case `value` is deprecated: use fresh ret 0; } fun make() reply.Reply { ret reply.Reply.value{1}; # construction warns too } fun stale() reply.Old { ret reply.Old.empty{}; # warning: `Old` is deprecated: the whole tag } ``` A deprecated case warns at every external use that names it: `Reply.value{...}` construction, the `sel place.value` test and the `place.value` payload place. Descriptor forms that name no case in source (`Reply.[c]{...}`, `sel v.[c]`, `v.[c]`) warn nowhere. Notices follow imported symbols and re-exports. A `#[deprecated]` on a `use` alias or a `fwd` re-export belongs to the forwarding module and replaces any inherited notice for that exported name; a clean alias of the same canonical definition keeps no notice. `test` blocks and comptime directives reject the attribute because they declare no externally usable name. ##### `testing` — test-only declaration Marks a declaration as existing only for tests, giving a fixture the semantics a [`test`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md) block already has. It takes no arguments and may appear once. A `#[testing]` declaration is resolved and type-checked in every build, so it cannot rot, and a type error in one fails `mach build` as well as `mach test`. It is omitted from IR and object files in ordinary builds and emitted under `mach test`. A reference to a `#[testing]` declaration is legal only from a `test` body or from another `#[testing]` declaration, including its decorators. Any other reference is an error at the use site. Every reference that names the declaration is checked: value references, calls, address-of, type references (a field, parameter or return type of a production declaration), generic instantiation, comptime evaluation and decorator arguments. The check has no same-module exemption. ```mach fragment # file: src/queue.mach #[testing] pub fun filled(n: u32) Queue { ... } test queue__drains_in_order { var q: Queue = filled(3); # a test body may use the fixture ... } #[testing] fun drained() Queue { ret filled(0); } # so may another testing declaration pub fun reset() Queue { ret filled(0); } # error[testing.use]: `filled` is a `#[testing]` declaration: only a test body or another # `#[testing]` declaration may reference it ``` The mark follows imports. A plain `use` of a testing declaration is itself testing, so every reference through it is checked. A `#[testing] use` confines its alias even when the target is an ordinary declaration. A `fwd` of a testing declaration must itself be marked `#[testing]`, because an unmarked `fwd` puts the name on a production surface. `pub #[testing]` is valid, and a dependency's testing helper is usable from a dependent's tests. It applies to `fun`, `rec`, `uni`, `tag`, `def`, `val`, `var`, `use` and `fwd` declarations at module scope. `test` blocks reject it because they are already test-only, and comptime directives reject it too. It cannot combine with `ext` or with `symbol`, `section`, `stage`, `input`, `output`, `builtin`, `uniform`, `storage`, `sampler` or `push`, because each names a consumer outside Mach source that the check cannot see. Inline `asm` `{name}` operands bind only locals, so they never reference a declaration. The mark is not a cycle escape: `mach test` builds with testing declarations present, so a `use` cycle that only tests need is still an error. `mach doc` omits testing declarations. ##### `expect(key)` — acknowledge a warning Acknowledges warnings a declaration raises on purpose. Each argument is a constant string naming a warning key, or a family of keys, from the [diagnostic key table](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings). A warning of a named kind raised inside the declaration, its doc comment and body included, is neither printed nor counted. Warnings elsewhere, and warnings of other kinds, still report. ```mach fragment #[expect("vector.scalarize")] #[testing] fun divide_i32x4(a: i32x4, b: i32x4) i32x4 { ret a / b; # scalarizes on x86_64 and aarch64, by design } #[expect("float")] test float__rounding { val x: f32 = 3.14159265358979; # float.inexact, acknowledged by its family ... } ``` A warning is matched to the declaration whose source contains its site. That holds for a warning raised late in the build, such as `vector.scalarize` from lowering, and for a generic instance, whose warnings belong to the generic's declaration. A warning replayed from the build cache is matched exactly as a fresh one, and a warning `#[expect]` covers counts as acknowledged even when a profile's `allow` silences its key too. An acknowledgement cannot rot silently. When a key the source alone decides (`import.unused`, `decl.deprecated`, `doc.lint`, `float.inexact`) names no warning raised inside the declaration, that is itself a warning, `expect.unfulfilled`. A key whose warning depends on the target or on what the build compiles, such as `vector.scalarize`, may be quiet in a given build, so its expectation is not reported. A build with errors judges no expectation. It applies to `fun`, `rec`, `uni`, `tag`, `val`, `var` and `test` declarations and may appear once, naming every key it acknowledges. There is no module-wide form: a profile's [`allow`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings) silences a key across the build. An unknown key, a key that covers only errors, and a key named twice are refused, since an error is never acknowledged: ``` error[expect.error_key]: `secret.not_oblivious` names an error, and an error is never acknowledged ``` ##### `symbol(str)` — linker name Overrides the emitted or imported symbol name. Applies to functions and globals. ```mach fragment #[symbol("main")] fun entry(argc: i64, argv: **u8) i64 { ... } #[symbol("write")] ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ``` Without `symbol`, the compiler mangles the Mach name — except on an `ext` declaration, which names a C declaration and takes the target's C symbol name for it (Darwin's underscore prefix, nothing elsewhere; see [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md)). `symbol` gives the exact name the linker sees, with no platform prefix applied to it. A mangled name is the source FQN, dotted, with generic arguments after a `$`: `std.types.string.str_len`, `std.collections.vector.push$ptr`. Each argument is introduced by a run of `$` whose length is its nesting depth, so a nested argument closes without a bracket — `f[Map[Vec[i64], str], u8]` is `m.f$m.Map$$m.Vec$$$i64$$str$u8`. `p$u8` is `*u8`, `sec$u32` is `^u32`, `arr4$u8` is `[4]u8`, `fn$$i64$$u8` is `fun(u8) i64`, a record is its own dotted origin FQN, a comptime value is its literal, and a variadic-pack instance carries a `pack` marker before its element list. A test's symbol is its qualified name, `#` (see [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#grammar)). There is no prefix: a mangled name always contains a `.` or a `#`, and a C identifier never can. Both `.` and `$` are legal in an inline-asm symbol, so any emitted symbol can be named from `asm` — but the spelling is not a stability promise, and binding to one from C is not a supported use. ##### `library(str)` — dynamic import attribution Pins an `ext` import to a specific dependency in the link set. Applies to `ext` functions only. ```mach #[library("ws2_32.dll")] #[symbol("WSAStartup")] ext fun wsa_startup(ver: u16, data: *u8) i32; ``` - The value normally names a `[link.X]` requirement's stable logical identity: its `library` value, or `X` when that key is omitted. A bare command-line `-l name` also exposes `name`. Exact canonical loader names remain accepted. Pinning to an absent dependency is a link error, never a silent fallback. A logical identity may not equal a different dependency's loader name. - PE and Mach-O use two-level namespaces, so every dynamic import on those targets needs a `library` attribution. - On ELF (Linux) the loader resolves imports by global search, so `library` has no effect on the emitted binary; the value is still validated against the link's dependency set. - `library` composes with `symbol`: the import is emitted under the renamed symbol within the named dependency. ##### `inline` — force inlining Marks a function for inlining at every direct call site, overriding the compiler's size and use-count heuristics and exempt from the caller's expansion budget. Applies to functions only and takes no arguments. The optimization pipeline must enable inlining. Indirect calls and recursive call cycles are not expanded by this attribute. Taking a function's address retains its callable identity even when direct calls are inlined. Release optimization makes small ordinary helper bodies available across source modules without emitting extra definitions. A helper is small when its body has fewer than 25 live instructions after promotion, debug annotations excluded, so `-g` never moves the decision. Extraction, import and per-caller heuristic expansion each have a limit of 1024 copied IR instructions and 256 KiB of owned payload. An `inline` function is expanded at every direct call site in the same module without charging that budget, so the outcome never depends on what else the caller expanded first; across modules it is imported within the extraction and import limits like any other body. An `oblivious` function is expanded only into another `oblivious` function, so its instructions never leave a constant-time validated body; a `naked` or `noinline` function, a recursive cycle and an indirect call are never expanded. A remaining call or taken address still names the original defining function. Generic, comptime and pack specializations keep their existing shared weak linkage. When several modules materialize that same specialization, body import uses an already available definition or the first acquired provider of that linkage, and tracks that provider as a query dependency. Helpers referencing compiler-local literal pools retain their calls because those objects have module-local identity. Named globals keep their original symbols, and copied instructions preserve effects, assembly bindings and debug locations. ```mach #[inline] fun fast_path(x: i64) i64 { ret x * 2; } ``` ##### `noinline` — forbid inlining The inverse of `inline`: forbids inlining a function into any caller, overriding the compiler's size- and use-count heuristics that would otherwise fold it in. Applies to functions only; takes no arguments. ```mach #[noinline] fun cold_path(code: i64) i64 { ret code * 100; } ``` Use it to keep a function's frame and symbol real — for a profiler or stack sampler to attribute its cost correctly, to keep a cold path from bloating a hot caller's instruction cache, or to hold code size down on a constrained target. - `inline` and `noinline` on the same function is a direct contradiction and is rejected in sema; neither wins silently. - `scalar` already declines inlining as a side effect (#2141), so pairing it with `noinline` is legal but redundant. - Purely a hint to the inliner; it does not otherwise change codegen. It binds at every optimization level — the debug pipeline runs no inlining pass at all, so `noinline` is inert (and unnecessary) there — and it will bind identically when a callee body is available from another module. Recursive peeling also respects `noinline` and `scalar`. ##### `align(expr)` — alignment override Sets the alignment of a global variable, a record/union type, or a function's entry. `expr` must be a comptime integer — either a literal or a comptime expression such as `$size_of(T)` or `$align_of(T)`, in both positions. A type's alignment is settled during type resolution, before layouts are otherwise known; the measured type's layout is established on demand when the intrinsic asks for it, so the answer does not depend on whether `T` is declared above or below. ```mach rec Pair { a: u64; b: u64; } #[align(64)] pub var cache_line: u8 = 0; #[align($size_of(Pair))] pub var g_cmp: u8 = 0; #[align($align_of(Pair))] rec Over { a: u8; } #[align(64)] fun hot(n: i64) i64 { ret n + 1; } ``` A type aligned to a measurement of itself — `#[align($size_of(Self))]`, or two types each aligned to the other's size — is a layout cycle and is reported as one, naming the type that closes it. - On a `var` / `val`, sets the global's section alignment and address alignment. - On a `rec` or `uni`, sets the type's own alignment, which is then inherited by any global of that type. - On a `fun`, sets the alignment of the function's entry address. Without it a function still starts at the target's own entry alignment: 16 bytes on x86-64 and aarch64, as gcc and llvm align them, and no padding on riscv. The bytes before an entry are an instruction that traps. `align` only raises the target's alignment, so a value below it changes nothing. - `align` does not apply to `def` aliases (transparent, no layout of their own). The alignment holds for **every** object of the type, including a local on the stack. An ABI only promises the stack 16 bytes at a call boundary, so a function holding a local that asks for more gets a prologue that masks the stack pointer down to the largest alignment its frame contains, and addresses its locals and spills from there. The cost falls on those functions alone: one masking instruction and up to `N - 16` bytes of frame. A function with nothing over-aligned emits exactly the prologue it always did. The frame pointer stays where the ABI put it, so incoming stack arguments, the frame record and a stack walk through it are unaffected. ##### `packed` — no padding Lays a `rec` or `uni` out with no padding: every field sits immediately after the previous one, there is no padding at the tail, and the type takes no alignment from its fields. Takes no arguments. `align` only ever raises alignment. `packed` is the inverse, and it exists for the case where the layout is not mach's to choose — a C struct, a file header, a wire frame, a vertex whose stride a buffer fixes. Without it such a shape cannot be described as a record at all. ```mach use std.print; use std.runtime; #[packed] rec Header { magic: u8; # offset 0 version: u16; # offset 1 length: u32; # offset 3 checksum: u64; # offset 7 } # $size_of == 15, $align_of == 1 #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {} {}", $size_of(Header), $align_of(Header), $offset_of(Header, checksum)); ret 0; } ``` Naturally the same shape is 24 bytes. `$size_of`, `$align_of` and `$offset_of` all report the packed layout, and so does the code that reads and writes the fields — there is one layout, not a declared one and an emitted one. ###### Composition with `align` The two compose rather than conflict, and each owns one question: - `packed` decides **padding** — none between fields, none at the tail. - `align(N)` decides the **record's own alignment**, and rounds its size up to a multiple of `N`. ```mach use std.print; use std.runtime; #[packed] #[align(8)] rec Frame { a: u8; b: u32; } # fields at 0 and 1; $align_of == 8, $size_of == 8 #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {} {}", $size_of(Frame), $align_of(Frame), $offset_of(Frame, b)); ret 0; } ``` ###### Packing is not transitive A `packed` record packs **its own** fields. A record it contains keeps its own internal padding and is merely *placed* without padding. This matches C, and it is the rule that composes: an inner type's layout does not change depending on who holds it. ```mach use std.print; use std.runtime; rec Point { x: u8; y: u32; } # natural: y at 4, size 8 #[packed] rec Msg { tag: u8; p: Point; } # p at offset 1, still 8 bytes; $size_of(Msg) == 9 #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {} {} {}", $offset_of(Point, y), $size_of(Point), $offset_of(Msg, p), $size_of(Msg)); ret 0; } ``` A transitive rule would make `Msg` 6 bytes and silently change `Point`'s meaning inside it. If the inner record must be packed too, write `#[packed]` on it as well. ###### What is refused **The address of a packed field.** `?r.b` on a packed record would yield a `*u32`, and a `*u32` states alignment 4 to everything downstream of it while the storage it names has none. The access through such a pointer is correct on the targets mach supports today; the **pointer type** is what is untrue, and it travels — an atomic, which every ISA requires naturally aligned, is exactly what a caller does with a pointer it was handed. The refusal covers the whole access chain, so `?r.arr[0]`, `?r.inner.x` and `?p.b` through a `*Packed` are refused for the same reason. `?r` on the **whole** record stays legal: a `*Packed` describes an align-1 pointee correctly, and nothing is lost by handing it out. To work with a field's value, copy it into a local. This is fail-closed on purpose. Refusing can be relaxed later, once alignment can ride in a pointer type; permitting cannot be tightened later without breaking programs that came to depend on it. Rust refuses; C permits, and it is a standing source of faults. **Atomics on a packed field** are refused by that same rule, not by one of their own. `std.sync.atomic` is ordinary functions over `*i64`, so a pointer is the only route an atomic has to a field, and there is no pointer to hand it. **Vector fields.** A vector in a packed record is refused for now, including one reached through an array or a nested record. The reason is evidence rather than arithmetic: an unaligned **scalar** access is measured on real hardware, and that measurement is what `#[packed]` rests on. The vector measurement now exists too. The codegen corpus's `vec/vec_mem` case writes `f32x3` and `f32x5` — the two shapes whose size and alignment disagree — into packed buffers and records where every write has a live neighbour, folding the neighbour after the write so a store too wide by a lane changes the checksum. It runs at both pipelines on every target with an execution engine, against a C reference the host's own compiler built, so a dropped lane or a disturbed neighbouring byte is a differing number rather than a passing run. The row that matters is `aarch64-linux` on `ubuntu-24.04-arm`, because aarch64 has 128-bit forms with alignment requirements; `x86_64-linux` and `x86_64-windows` carry it too. `riscv64-linux` also passes and is not evidence: it runs under qemu-user, and riscv64 declares no 128-bit vector support, so the access there is a scalar expansion rather than a vector access. What the refusal still waits on is the other half, [#2687](https://github.com/briar-systems/mach/issues/2687) — a lane-dependent vector footprint through aggregate layout and ABI classification. This is a sequencing decision and is expected to be lifted, not a permanent rule. **Interface blocks.** `packed` cannot apply to a `#[uniform]`, `#[storage]` or `#[push]` block: its member offsets are fixed by the std140 / std430 layout rules and emitted as explicit SPIR-V `Offset` decorations, which packing would contradict. ###### Target note: riscv64 Unaligned access is permitted-but-may-trap on RV64. Where the hardware does not do it, Linux emulates the access in the kernel, so a packed field access there is expected to be **correct and pathologically slow** — a trap-and-emulate round trip per access rather than a load. Correctness tests on that target will pass and prove nothing about usability, so treat a green riscv64 leg as evidence about correctness only. This has not been measured on riscv64 hardware. It cannot be: qemu-user emulates a misaligned guest load directly and never takes the kernel path, so a qemu measurement shows no cost whether or not real silicon would. If the cost turns out to matter, the answer is byte-wise lowering of packed field access on faulting targets, which is codegen work and not part of `#[packed]` as it stands. x86-64 and aarch64 do unaligned scalar access in hardware. The aarch64 answer is measured on real hardware rather than assumed — `int`'s `linux-arm64` leg runs natively. ##### `volatile` — every access to the type is a volatile access `#[volatile]` on a `rec`, `uni` or `tag` declaration makes every load and store of that type's storage volatile: the optimizer keeps each one, in program order, at the width written. A volatile access is never elided as dead or redundant, never hoisted out of a loop, never merged with its neighbour into a wider or unaligned access, and never promoted to a register (`mem2reg` leaves a function with one alone; `licm` refuses to hoist one; the bulk-memory combiner skips one). It takes no arguments and applies to nothing else: `#[volatile]` on a function or a variable is refused with `` `volatile` decorator applies only to records, unions, and tags``. Volatility is a property of a **declared type**, so every volatile access in a program traces back to a declaration. There is no variable-level decorator and no pointer qualifier (see [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#pointer)); a memory-mapped device is expressed by declaring its register block as a volatile record and casting its address to a pointer to it: ```mach #[volatile] rec Fb { status: u32; px: [1024]u32; } val fb: *Fb = 0xB8000::*Fb; fun fill(c: u32) { var i: u64 = 0; for (i < 1024) { fb.px[i] = c; # one volatile store per iteration i = i + 1; } fb.status = 1; # a volatile store, ordered after the loop } ``` The access decides by the storage it reaches, not by its form. A member (`fb.status`), a projection, an index (`fb.px[i]`), a dereference (`@fb`) and a whole-record copy are volatile alike, and so is a field of a plain record stored inside a volatile one (`blk.ctl.bits` when `Blk` is volatile and `Ctl` is not), because the storage is the volatile block. An indirection ends the walk: a pointer field of a volatile record is itself read volatile, but what it points at is ordinary storage unless its own type says otherwise. A raw scalar pointer (`@p` with `p: *u32`) is never volatile, since it traces to no declaration. A volatile record copied by value is a volatile copy in both directions, and a `#[volatile]` type is compatible with `#[packed]` and `#[align(N)]`, which only shape the layout. `rec.md`, `uni.md` and `tag.md` link here. ##### `section(str)` — object section placement Places a function or global variable in a named section instead of the default `.text` / `.data`. ```mach #[section(".hottext")] #[symbol("f_hot")] fun f_hot(x: i64) i64 { ret x + 1; } #[section(".machsec")] #[symbol("g_sec")] pub var g_sec: u64 = 100; ``` The named section is created if absent. Cross-section calls and accesses use ordinary relocations. ##### `oblivious` — constant-time boundary Marks a function as a constant-time boundary. Applies to functions only; takes no arguments. Inside it the backend must not introduce a secret-dependent branch or select a variable-latency instruction on a secret operand; a translation validator re-derives the secret taint over the lowered MIR and rejects any such leak. Inline `asm` inside such a function is **validated rather than rejected**: the block is parsed and walked for the same leaks, and refused only where a leak is found or where the construct cannot be modelled. See [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#oblivious--the-codegen-contract) for what is checked and what is refused, including the x86-64 conditional-branch limitation. The **zeroizing-write** guarantee is *not* one of the decorator's obligations, and describing it as one understates it. A write into secret storage carries a taint applied at lowering, keyed on the storage's secrecy rather than on any decorator, so a zeroizing wipe is protected in a function carrying no `#[oblivious]` at all. See [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#the-zeroizing-write-guarantee) for what that covers and what it does not. ```mach fragment #[oblivious] fun ct_eq(a: ^[8]u8, b: ^[8]u8) u8 { ... } ``` The decorator is purely subtractive — on a secret-free function it is a no-op. A function instance that *computes* on a `^` secret (arithmetic, bitwise, shift, comparison, negation) is **required** to carry it; an instance that only moves, stores, or declassifies secrets stays annotation-free. It is rejected outright for a target whose back half emits a module for a downstream compiler rather than the executed instructions (the experimental SPIR-V backend): the contract cannot be validated or upheld there. > **Experimental preview.** The constant-time guarantee is not complete and has > not been audited — see [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#assurance) for the known open > holes. Do not build production cryptography on it at this version. ##### `scalar` — opt out of auto-vectorization Excludes a function from loop auto-vectorization, so its loops compile to scalar code even in the release pipeline on a vector-capable target. Applies to functions only; takes no arguments. ```mach fragment #[scalar] fun reference_sum(a: *i64, n: usize) i64 { ... } ``` A `#[scalar]` function is also declined by the inliner, so the opt-out survives inlining — it cannot be lost by the body moving into an unflagged caller. Use it for a scalar reference twin in a differential test, or where vectorized codegen is undesirable for a specific function. The project-wide equivalent is the `vectorize` profile key (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#profilename)). ##### `naked` — no prologue, no epilogue, body as written Emits the function's body exactly as written and nothing else: no frame-pointer record, no stack allocation, no callee-save stores, no argument moves, and no return. Applies to functions only; takes no arguments. ```mach fragment #[naked] #[symbol("_start")] fun start() { $if ($mach.build.arch == $mach.arch.x86_64) { asm x86_64 { mov rdi, [rsp] # argc, straight off the kernel-supplied stack lea rsi, [rsp+8] # argv call main } } $or { asm aarch64 { ... } } } ``` The programmer owns the frame, the stack alignment, the link register, and the return. That is the whole point: a reset vector, an interrupt handler that must return with `iret`/`rti` rather than `ret`, a syscall or context-switch stub, or a thread entry point whose register state at entry *is* the interface. - **The body may contain only inline `asm`** — plus the `$if $mach.build.arch` chain that is how mach spells per-ISA assembly. Any other statement is rejected. A local, an expression, or a `ret` lowers to code that assumes a frame the function does not have, and the result would run and return a wrong answer rather than fail. - **No return is generated.** If the asm falls off the end, control runs into whatever the linker placed next. Write the return the ABI (or the interrupt controller) actually calls for. - **Parameters and the return type are still checked** at every call site, so a naked function is called like any other. No moves are emitted for them: the arguments arrive in the ABI's registers and the body reads them there. - **Mutually exclusive with `inline`** — there is no coherent winner between a body spliced into a caller and one that owns its own frame — and with `oblivious`, which already forbids inline asm because a type system cannot check it. Both combinations are rejected in sema. `noinline` is redundant: the inliner declines a naked function unconditionally. - **Debug info carries no `DW_AT_frame_base`** for a naked subprogram. Every other function has a frame base the compiler established and can name; this one does not, so it declares none rather than pointing at a register the asm may have moved. Frame *elision* is a separate, automatic thing: the compiler already omits the prologue for a leaf that provably never touches its frame. `naked` is the declared form, and it is unconditional — it suppresses the frame whether or not the compiler could prove it safe, because the proof obligation is the author's. Merely containing an `asm` block does **not** suppress a frame: a function that also makes a call gets one, since an unaligned call boundary (x86-64) or a clobbered link register (aarch64, riscv64) is not something the author asked for by writing assembly. ##### `extensions(names)` — an outlier function Lets one function use instruction-set extensions the target does not select. Applies to a function with a body; takes one or more bare extension names. ```mach fragment #[extensions(sha, ssse3)] fun compress_sha_ni(state: *[8]u32, block: *[64]u8) { asm x86_64 { # sha256msg1, sha256rnds2, pshufb, ... } } fun compress(state: *[8]u32, block: *[64]u8) { if (cpu_has_sha_ni()) { compress_sha_ni(state, block); } or { compress_portable(state, block); } } ``` This is the same contract as Rust's `#[target_feature(enable = "...")]` and gcc's `__attribute__((target("...")))`. Inside the function, inline `asm` may use every instruction the named extensions admit, in addition to what the target selects. **The caller owns the run-time check.** Calling an outlier on a processor that lacks one of its extensions is undefined behaviour, typically an illegal-instruction trap. The compiler neither inserts the check nor verifies that one precedes the call, because only the program knows how it detects the processor's features and when that answer holds. - **Names come from every instruction set.** A name no instruction set declares is refused, and the refusal lists the known names. A name another instruction set declares admits nothing on this target and is not an error, so one declaration serves a multi-arch source: `#[extensions(sse41, sha2)]` admits `sse41` on x86_64 and `sha2` on aarch64, and an `asm aarch64 {}` block in an x86_64 build still refuses `sha256h` because the tag, not the decorator, picks the isa. Each name may appear once. The names are those the manifest's [`extensions`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions) key takes, closed over what they imply (`sse41` admits `pshufb`), except the rows only a target selects (riscv `i`, `c`, `f`, `d`, `zkt`), which are refused with the reason. - **One predicate, everywhere.** A function's admitted set is the target's selection plus what the decorator names. An instruction requiring an extension is emitted only into a function whose admitted set holds it (see [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md#extension-instructions)); the encoder checks every row against it, and the inliner checks it before moving a body, so an outlier the target already selects may still inline into a baseline caller. - **Only the instructions change.** The decorator admits instructions in the function's inline `asm`. It does not change how the compiler generates the rest of the body, which stays within the target's selection. The function is called through the ordinary ABI, and its address is an ordinary `fun(...)` value that dispatch through a pointer calls like any other. - **It is never inlined into a caller that admits less.** The inliner declines to move an outlier's body into a function whose admitted set does not hold every extension the outlier's does, so the extension instructions stay behind the call the run-time check guards. `#[inline]` does not override this. A caller carrying a superset of the outlier's extensions may still inline it, and any function may be inlined *into* an outlier. A generic function's instances carry the decorator's set. - **Nothing else changes.** The decorator composes with `inline`, `noinline`, `symbol` and `section`. `#[oblivious]` already refuses the instructions a constant-time check cannot model, and an extension row it can model is checked like any other. ##### `embed(str)` — compile-time file embedding Sources a `val`'s bytes from a file at compile time: the file's content **is** the initializer. Applies to `val` only — not `var` (the storage is read-only data) and not an `ext` data import (which has no storage here). Takes one constant string argument. ```mach fragment #[embed("assets/logo.qoi")] val LOGO: [_]u8; # length taken from the file's byte count #[embed("boot/sector.bin")] val SECTOR: [512]u8; # length pinned; a size change fails the build ``` - The declaration carries no initializer of its own; writing one alongside `embed` is rejected. This is a second exemption to `val`'s requires-an-initializer rule, alongside `ext` (see [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md#ext--foreign-data-imports)). - Exactly one argument, a constant string, decoded like any string: a literal's escapes are its value, as they are for `symbol` and `section`. - The path resolves relative to the **declaring source file's** directory. An absolute path is taken as written. The resolved file must lie inside the project root: an embed that escapes it (`../../outside.txt` from `src/`) is refused at the decorator (`` `embed` path escapes the project root; an embedded file must live inside the project and the file is not read ``) and the file outside is never read. Keep assets under the project. - A path holding `{artifact..out}` names the output of a required artifact and resolves against the **root project's** directory rather than the declaring file's directory; the required artifact is built first. The name is read in the manifest that owns the declaring module: the root's own module names an artifact the built artifact requires, and a dependency's module names an artifact that dependency's `default = true` library artifact requires, which the consumer's build produces for it. No other template variable may appear in an `embed` path. See [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-requirements). - The annotation must be `[_]u8` or `[N]u8`; the element type must be `u8`. `[_]` is an inferred array length, legal **only** on an `#[embed]` declaration — written anywhere else it is rejected (see [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#types)). A `[_]u8` embed can be asked for its own length: `$length_of(LOGO)` is its element count and `$size_of(LOGO)` its byte count, both folded at compile time (see [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md)). The explicit `[N]u8` form is for pinning a size by contract, not for recovering one. - An explicit `[N]u8` whose `N` disagrees with the file is rejected, naming both counts. This is how a declaration pins a fixed-size asset — a boot sector, a ROM image — so the build fails the moment it stops being that size. - Bytes are placed in read-only data exactly like any other constant byte array: no runtime I/O, no copy. Works for every artifact kind and target, freestanding included. - Two `#[embed]` globals whose files hold byte-identical content and whose final section name, kind, and alignment match share **one** read-only data placement within a module, so their addresses compare equal. This is specific to embedded data — an ordinary global is never merged this way, and a named object's address is otherwise its own. - A missing file, a directory where a file is required, an unreadable file, and a file larger than the 4,294,967,295-byte array/section limit each report once, naming the declaration and the resolved path. - The embedded file is a build input: its content digest feeds the embedding module's incremental cutoff, so editing the asset invalidates that module and an untouched asset stays a cache hit — see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps) for the equivalent guarantee on `[step]` `in` entries. ##### `stage(str)` — GPU pipeline stage Marks a function as the entry point of a graphics or compute pipeline stage. The argument names the stage and the set is closed: | Value | Stage | |--------------|---------------------------| | `"vertex"` | vertex shader | | `"fragment"` | fragment (pixel) shader | | `"compute"` | compute shader | An unrecognized value is a compile error, not a module that quietly forms no stage. ```mach #[stage("vertex")] fun vertex_main() {} #[stage("fragment")] fun fragment_main() {} ``` A staged function **takes no parameters and returns nothing**. A pipeline stage does not have a caller: its inputs arrive through input interface variables and its results leave through output ones, so there is no argument list or return value to carry them. A staged function with either is rejected. The decorator is accepted on every target, because which target a module is built for is not a property of its source. Only a target that has pipeline stages acts on it: on `spirv` a staged function becomes an `OpEntryPoint` with the matching execution model, and on a machine target the stage is ignored and the function is compiled normally. A module that declares any stage is a **shader module**, and that changes the whole artifact rather than just the one function. A shader module carries entry points and no external linkage at all; a module with no stage is a **library module**, which publishes each function as a linkage export so a consumer can find it. The two are exclusive — a Vulkan consumer refuses a module carrying linkage — so adding the first `#[stage(...)]` to a module stops it exporting its functions. The entry point's name, as a pipeline-creation call looks it up, is the function's **bare source name** (`vertex_main` above), not a mangled linker symbol. A shader module has no linker symbols to mangle. ##### `workgroup(x, y, z)` — compute workgroup dimensions Sizes the workgroup of a `#[stage("compute")]` function. The three arguments are comptime integers giving the x, y and z dimensions. ```mach #[stage("compute")] #[workgroup(64, 1, 1)] fun compute_main() {} ``` It requires a stage on the same function — without one it would silently mean nothing — and it applies only to the compute stage. When it is omitted, a compute stage takes the single-invocation default `(1, 1, 1)`; the dimensions are always declared in the emitted module, since a compute stage that does not state its workgroup size is not one a consumer can dispatch. Any dimension may instead name a `#[spec]` var (see [`spec`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#specid--specialization-constants)), which the host sets when it creates the pipeline. The var must be a `u32` or `i32`, and any other variable is refused. ```mach fragment #[spec(0)] var tile: u32 = 64; #[stage("compute")] #[workgroup(tile, 1, 1)] fun compute_main() {} ``` On `spirv` how a specialized workgroup is emitted depends on the environment. Where it allows it, the stage carries `LocalSizeId` over the spec constant. That needs SPIR-V 1.2 and, on Vulkan, the `maintenance4` feature, which only Vulkan 1.3 requires of every device, so the form is used with no `env` and with `vulkan1.3`. Every other environment gets a `WorkgroupSize` built-in over an `OpSpecConstantComposite`. That built-in sizes **every** compute stage in the module, so a module in such an environment whose compute stages size their workgroups differently, and at least one of them through a `#[spec]` var, is refused. Give the stages one workgroup or put them in separate modules. ##### `input(n)` / `output(n)` / `builtin(str)` / `uniform(set, binding)` / `storage(set, binding)` / `sampler(set, binding)` / `push` / `spec(id)` / `shared` — shader interface A pipeline stage does not receive its inputs or return its results through a call. It reads and writes **module-scope variables** that the pipeline binds, and these directives say which kind each variable is. They apply only to module-level `val` / `var` bindings, and a variable carries **exactly one** of them — they are mutually exclusive. `spec`, described in its own section below, is one of them too. ```mach fragment #[input(0)] var in_position: f32x4; #[output(0)] var out_colour: f32x4; #[builtin("position")] var position: f32x4; rec Camera { view: f32x4; proj: f32x4; } #[uniform(0, 0)] var camera: Camera; rec Particles { pos: [64]f32x4; } #[storage(0, 1)] var particles: Particles; #[sampler(1, 0)] var albedo: Sampler2D; rec Params { tint: f32x4; count: u32; } #[push] var params: Params; ``` `input` and `output` number a **varying** with a location, which is how one stage's outputs line up with the next stage's inputs: the producer's `#[output(0)]` feeds the consumer's `#[input(0)]`. `builtin` names a value the pipeline supplies or consumes instead of one a location carries. The accepted set is closed: | Value | Meaning | Type | Direction | Stage | |----------------------------|----------------------------------------|---------|-----------|----------| | `"position"` | clip-space vertex position | `f32x4` | written | vertex | | `"point_size"` | rasterized point size | `f32` | written | vertex | | `"vertex_index"` | index of the current vertex | `u32` | read | vertex | | `"instance_index"` | index of the current instance | `u32` | read | vertex | | `"frag_coord"` | fragment window coordinate | `f32x4` | read | fragment | | `"global_invocation"` | compute global invocation id | `u32x3` | read | compute | | `"local_invocation"` | compute local invocation id | `u32x3` | read | compute | | `"workgroup_id"` | compute workgroup id | `u32x3` | read | compute | | `"num_workgroups"` | compute workgroup count of a dispatch | `u32x3` | read | compute | | `"local_invocation_index"` | compute local invocation id, flattened | `u32` | read | compute | | `"subgroup_size"` | invocations in a subgroup | `u32` | read | every | | `"subgroup_invocation"` | the invocation's index in its subgroup | `u32` | read | every | | `"subgroup_id"` | the subgroup's index in its workgroup | `u32` | read | compute | | `"num_subgroups"` | subgroups in the workgroup | `u32` | read | compute | The direction is a property of the built-in, not something you restate — a stage writes its position and reads what the pipeline hands it — so there is no input/output marker to pair with `builtin`, and none that could disagree with it. The **type** is a property of the built-in too, and it is a requirement rather than a suggestion: the pipeline binds the variable itself, so a wider or narrower one is an invalid module rather than a wasteful one. Declaring a built-in at any other type is a compile error naming both the declared type and the required one. The scalar integer rows accept `i32` as well as `u32`, because the compiler carries an integer's width and not its sign and the emitted type is sign-less either way. The **stage** is a property of the built-in as well. Each row exists in the stages SPIR-V defines it in, in its direction: the pipeline supplies an input built-in to those execution models alone and consumes an output one from them alone, so `position` and `point_size` are a vertex stage's outputs, `frag_coord` is a fragment stage's input, the seven compute rows are the `GLCompute` execution model's, and `subgroup_size` and `subgroup_invocation` are every stage's, decorated `Flat` as a fragment stage's integer inputs. A stage that uses a built-in outside its row, directly or through a function it calls, is a compile error naming the built-in and both stages. The four subgroup built-ins need SPIR-V 1.3, and outside a compute stage the `subgroup_graphics_stages` feature, as the subgroup operations do. `uniform` binds a read-only block by descriptor set and binding. Its type **must be a `rec`**: a uniform is a block with a host-visible layout, and a bare scalar or vector has no block layout for a pipeline to bind. The record is emitted with its `Block` decoration and an explicit byte offset on every member, taken from the same layout the rest of the compiler uses, so what the shader reads is what the host wrote. Wrap a single value in a one-field record. `storage` binds a **read-write** buffer by descriptor set and binding, where `uniform` binds a read-only one. Both must be a `rec` for the same reason, and both are emitted as a `Block`-decorated struct with an explicit offset on every member. They differ in one place: their **layout rules**. A uniform block follows std140-shaped rules, under which an array's stride is rounded up to 16 — which mach's own layout does not do, so an array of anything narrower than 16 bytes is refused rather than silently repacked. A storage buffer follows std430-shaped rules, which use the element's natural stride, and that *is* mach's layout, so `[8]f32` is fine in a `storage` block and rejected in a `uniform` one. A compute stage's data path is `storage`: Vulkan forbids the `Output` storage class in a compute execution model, so a compute shader reads and writes buffers rather than varyings. `storage` takes **memory qualifiers** after the descriptor pair, any number of them in any order: `"readonly"`, `"writeonly"` and `"coherent"`. `"readonly"` says that nothing writes the binding: ```mach rec Palette { columns: [512]f32x4; } #[storage(0, 3, "readonly")] var palette: Palette; ``` A store through a `"readonly"` binding is a compile error on every target, naming the line that wrote it. That is what the qualifier buys over what the compiler works out on its own: a buffer no body in the module stores through is emitted with the SPIR-V `NonWritable` decoration whether or not it is marked, and Vulkan reads that decoration to decide whether a stage needs `vertexPipelineStoresAndAtomics`. So an accidental write does not produce a wrong module, it produces a **correct one that quietly costs a hardware feature**. Marking the binding turns that into a diagnostic instead. The inference is one-sided on purpose. Anything the compiler cannot follow, such as the binding's address handed to a function, counts as a write, so a missing decoration is possible and a wrong one is not. `"writeonly"` is the mirror: it says that nothing reads the binding, and the buffer is emitted `NonReadable`. ```mach rec Frame { texels: [4096]f32x4; } #[storage(0, 4, "writeonly")] var frame: Frame; ``` A read of a `"writeonly"` binding is a compile error on every target, naming the expression that read it. Storing into a field or an element reads nothing, and neither does taking the binding's address, so `frame.texels[i] = c` and `?frame.texels[i]` are accepted. Writing **one lane** of a vector inside it is refused, because a lane write loads the whole vector and stores it back: store the whole vector instead. On SPIR-V, a load the compiler follows through the binding's address is refused as well, with the same caution as `"readonly"`: anything the compiler cannot follow counts as a read. `"readonly"` and `"writeonly"` together are refused, since that binding would be neither read nor written. `"coherent"` makes a write one invocation makes visible to invocations in other workgroups, which is what atomics and flags shared across workgroups rely on. It combines with either of the other two. How depends on the memory model the target selects (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)): - Under the GLSL450 model the buffer is decorated `Coherent`. - Under the Vulkan model, which has no such decoration, each access through the buffer carries its own availability or visibility: a store `MakePointerAvailable`, a load `MakePointerVisible`, each with `NonPrivatePointer`, at the `QueueFamily` scope, every invocation of the dispatch and of later work on that queue family. That is glslang's reading of `coherent`, and it needs no device-scope feature. A storage image's `OpImageRead` and `OpImageWrite` carry `MakeTexelVisible` or `MakeTexelAvailable` with `NonPrivateTexel` the same way. A storage image handed to a function as a parameter carries no qualifier into it, so every image read and write in such a function is coherent in a module that binds a coherent image. ```mach rec Counters { done: u32; } #[storage(0, 5, "coherent")] var counters: Counters; ``` `sampler` binds a **handle** by descriptor set and binding, at the same descriptor addressing `uniform` and `storage` use, so a host binds one the way it binds the others. Its type must be a **handle type**, a bodyless `def` carrying `#[handle]` (see [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md)), and a handle type must carry this decorator: a handle names a descriptor rather than an object with storage, so one with no descriptor address is reachable from no stage. A handle cannot sit behind a pointer, inside an array, or in a local binding, and each of those is a compile error naming why. A **storage image** is the exception: an `image` handle whose `Sampled` operand is `2` is read and written directly rather than sampled, which is Vulkan's `STORAGE_IMAGE` descriptor, so it binds through `storage` and not `sampler`. The `"readonly"`, `"writeonly"` and `"coherent"` qualifiers apply to it exactly as they do to a buffer, checked against the instructions that read and write its texels. A storage texel buffer, an image of `Dim` `Buffer` with `Sampled` `2`, binds the same way, and a uniform texel buffer (`Sampled` `1`) keeps `sampler`. Binding a storage image through `sampler`, or any other handle through `storage`, is a compile error. A storage image is read and written by `OpImageRead` and `OpImageWrite`, a uniform texel buffer or a sampled image is fetched texel by texel by `OpImageFetch`, and `OpImageQuerySize`, `OpImageQuerySizeLod`, `OpImageQueryLevels` and `OpImageQuerySamples` read an image's descriptor under the `ImageQuery` capability, which every Vulkan version accepts. `OpImageRead`, `OpImageWrite`, `OpImageFetch`, `OpImageSampleImplicitLod` and `OpImageGather` take an optional `Image Operands` mask leading their tail, so a declaration either stops at the instruction's required operands or passes the mask and the operands its bits bring, each typed by its bit. `Sample` (`0x40`) names one sample of a multisampled image, and a fetch's `Lod` (`0x2`) the level it reads. A multisampled image is read, written and fetched only with `Sample`, and only a multisampled image takes it. `OpImageSampleExplicitLod` always takes the mask, which must set `Lod` or `Grad` (`0x4`), whose two operands are the coordinate's derivatives along x and y, and never both. An implicit-lod sample takes `Bias` (`0x1`), a float added to the level it derives, and is reached only from a fragment stage, the one stage with the coordinate derivatives it derives that level from. A fetch, a sample and a gather take `ConstOffset` (`0x8`), an integer constant added to the coordinate, or `Offset` (`0x10`), the same computed at run time, one component per dimension of the coordinate, and never on a `Cube` image. A gather may take `ConstOffsets` (`0x20`) instead, a constant `[4]i32x2` of the offsets of the four texels it reads. An instruction takes one offset bit at most. `Offset` and `ConstOffsets` need the `image_gather_extended` extension, and `Offset` on a fetch or a sample `maintenance8` as well, since Vulkan admits it outside a gather only under that feature. A sample takes `MinLod` (`0x80`), the least level of detail it reads, which an explicit-lod sample takes only with `Grad` and which needs the `resource_min_lod` extension. `OpImageGather` reads one component, a constant id, of the four texels a sample of a `2D` or `Cube` image would filter. A constant operand is a constant by emission: a literal, a vector of literals directly or through a binding, a module-scope `val`, which a shader reads as the constant it is, or an aggregate passed whole as an array literal of constants or a copy of a `val`, written as one composite constant. An aggregate any run-time value writes, or one written only in part, is a value, and is refused with `op.operand_not_constant`. `OpImageQuerySamples` reads the sample count of a multisampled image only, and `OpImageQuerySizeLod` does not take one. `OpImageQuerySize` reads an image with no level of detail to choose, a multisampled image, a storage image or a texel buffer, so a single-sampled sampled image is queried with `OpImageQuerySizeLod` instead. Each is refused at the call with `op.operand_value`. ```mach fragment #[handle("spirv", "image", TEXEL_F32, DIM_2D, NO_DEPTH, NONARRAYED, MULTISAMPLED, SAMPLED, FORMAT_UNKNOWN)] pub def TextureMS; #[op("spirv", "core", "OpImageFetch")] fun fetch_sample(img: TextureMS, at: i32x2, mask: u32, sample: i32) f32x4; #[op("spirv", "core", "OpImageQuerySamples")] fun sample_count(img: TextureMS) i32; ``` ```mach fragment #[handle("spirv", "image", TEXEL_F32, DIM_2D, NO_DEPTH, NONARRAYED, SINGLE_SAMPLED, STORAGE, FORMAT_RGBA8)] pub def Target2D; #[op("spirv", "core", "OpImageWrite")] fun image_write(img: Target2D, at: i32x2, texel: f32x4); #[storage(0, 6, "writeonly")] var target: Target2D; ``` A **depth comparison** samples a depth image, one whose `Depth` operand is `1` (see [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md)), and compares each texel it reads against a reference value, an `f32` scalar that follows the coordinate. `OpImageSampleDrefImplicitLod` and `OpImageSampleDrefExplicitLod` filter the comparisons into one scalar of the image's texel scalar, and `OpImageDrefGather` returns the comparisons of the four texels a gather reads. Each takes the `Image Operands` its plain form takes, after the reference, under the same rules: the implicit-lod comparison takes `Bias`, an offset and `MinLod` and is reached only from a fragment stage, the explicit-lod one always takes the mask with `Lod` or `Grad`, and the gather takes an offset or `ConstOffsets` and reads only a `2D` or `Cube` image. A comparison of an image whose `Depth` is `0` is refused with `op.operand_value`, and so is one of a `3D` image, which Vulkan never compares (`VUID-StandaloneSpirv-OpImage-04777`). ```mach fragment #[op("spirv", "core", "OpImageSampleDrefImplicitLod")] fun shadow(s: ShadowSampler, uv: f32x2, depth: f32) f32; #[op("spirv", "core", "OpImageSampleDrefExplicitLod")] fun shadow_lod(s: ShadowSampler, uv: f32x2, depth: f32, mask: u32, lod: f32) f32; #[op("spirv", "core", "OpImageDrefGather")] fun shadow_gather(s: ShadowSampler, uv: f32x2, depth: f32, mask: u32, offset: i32x2) f32x4; ``` The projective forms, `OpImageSampleProj*` and `OpImageSampleProjDref*`, have no row. Each divides the coordinate, and a comparison's reference, by the coordinate's last component before sampling, which a shader writes as that division and passes to the plain form: they add no sampling a row here does not already give, HLSL, MSL and WGSL have none, and each brings its own refusals (no `Cube` image, no arrayed one). A declaration naming one is refused as an instruction the target does not define. `push` binds a **push-constant block**, a small `rec` the host supplies with the command that records a dispatch or draw rather than through a descriptor, so it takes no arguments: there is no set or binding to name. Like `uniform` and `storage` it must be a `rec`, and it is emitted as a `Block`-decorated struct with an explicit offset on every member. Its layout rules are std430-shaped, the same ones a `storage` buffer follows and checked the same way, so `[8]f32` is fine in a push block. A push block is **read-only** in the shader. A store to it is a compile error on every target, and on `spirv` so is a store the compiler follows through its address handed to a function. Vulkan admits **one push-constant block per entry point**: two push blocks used by the same stage are refused, naming both, while two stages that each use a different one are accepted. ```mach rec Params { scale: f32; slot: u32; } #[push] var params: Params; ``` Sampling a handle is an `#[op(...)]` declaration rather than a language form, because a sample IS one SPIR-V instruction like `sqrt` and `dot` are: ```mach fragment #[op("spirv", "core", "OpImageSampleImplicitLod")] fun sample(s: Sampler2D, uv: f32x2) f32x4; #[stage("fragment")] fun frag_main() { out_colour = sample(albedo, in_uv); } ``` The separately-bound form works the same way, with the instruction that combines an image and a sampler declared alongside it: ```mach fragment #[op("spirv", "core", "OpSampledImage")] fun combine(t: Texture2D, s: Sampler) Sampler2D; #[sampler(1, 0)] var base_tex: Texture2D; #[sampler(1, 1)] var base_smp: Sampler; #[stage("fragment")] fun frag_sep() { out_colour = sample(combine(base_tex, base_smp), in_uv); } ``` The combined value is handed straight to the sample rather than named: SPIR-V requires an `OpSampledImage` result be consumed by an image instruction in the block that produced it, which is the same rule that makes a handle-typed local a compile error. Returning one, or passing one to a function, is refused with `op.result_flow`, and so is a sample whose later operand branches, as `&&` and `||` do, since its sampled image would then be consumed in another block. Compute such an operand into a binding before the call. `shared` declares **workgroup memory**: one instance per workgroup of a compute stage, which every invocation of that workgroup reads and writes. It applies to a `var` only, since a `val` of workgroup memory could only ever read zero. It takes no arguments, and the variable has no descriptor and no location, because the pipeline never binds it. ```mach fragment #[builtin("local_invocation")] var local_id: u32x3; #[shared] var tile: [256]f32; #[stage("compute")] #[workgroup(64, 1, 1)] fun blur() { tile[local_id[0]] = 1.0; } ``` Workgroup memory exists only in a compute stage, so a `#[shared]` variable used from a vertex or fragment stage, directly or through a function the stage calls, is a compile error naming the stage. A `#[shared]` variable is **zero** when a compute stage starts, as every mach variable is, on every environment. How depends on the environment: - Where workgroup memory is zero-initialized by the consumer, the variable carries an `OpConstantNull` initializer. `vulkan1.3` guarantees that (`shaderZeroInitializeWorkgroupMemory` is core there), and a target that selects the `zero_init_workgroup` extension declares it for an earlier version (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions)). The consumer then has to enable the feature (`VK_KHR_zero_initialize_workgroup_memory`). - Otherwise the compiler zeroes it itself, at the start of each compute stage that uses it. Each invocation stores zero to its own slice, the elements of an array its local invocation index reaches in steps of the workgroup size and the whole of any other type for invocation 0, and then the stage executes one workgroup `OpControlBarrier`. The barrier precedes all of the stage's own code, so every invocation reaches it. Because the value on entry is always zero, a `#[shared]` variable cannot have an initializer. Assign it inside the stage. Workgroup memory is **coherent** among the invocations of a workgroup under either memory model. Under GLSL450 it is so by definition. Under the Vulkan model an access is private unless it says otherwise, and a barrier orders no private access between invocations, so each load and store of a `#[shared]` variable carries `MakePointerVisible` or `MakePointerAvailable` with `NonPrivatePointer` at the `Workgroup` scope, the compiler's own zeroing stores included. A workgroup barrier with acquire-release semantics then orders them as it does under GLSL450. Whether a variable may carry an **initializer** is settled by its role, since the role says who puts the first value in it: | Role | Initializer | Why | |--------------------------------------------------|-------------|-------------------------------------------------------| | `input`, a read built-in | refused | the previous stage or the pipeline supplies the value | | `uniform`, `storage`, `sampler`, `push` | refused | the host binds or supplies the memory | | `shared` | refused | workgroup memory is zero when a stage starts | | `spec` | required | it is the default the pipeline keeps | | `output`, a written built-in | allowed | it is the value the variable starts at | A refused initializer is a compile error, because the value it writes would never be the one the shader sees. An `output` or a written built-in starts at its initializer, and at zero without one, as every mach `var` does: ```mach fragment #[output(0)] var out_colour: f32x4 = f32x4{0.0, 0.0, 0.0, 1.0}; #[output(1)] var out_mask: u32; ``` On `spirv` the Output `OpVariable` carries that value as its initializer: the constant the initializer spells, or `OpConstantNull` where it is zero or absent. SPIR-V and Vulkan both admit an initializer on an Output variable. As with `#[stage(...)]`, these are accepted on every target and acted on only by a target that forms pipeline stages. On `spirv` each becomes an `OpVariable` in the matching storage class, carrying the matching decoration, and the Input and Output variables are named in every entry point's interface list. A `sampler` binding becomes an `OpVariable` in the `UniformConstant` class — the one class Vulkan permits an image, sampler or sampled-image variable in — carrying `DescriptorSet` and `Binding` exactly as a `uniform` does. A `push` block becomes an `OpVariable` in the `PushConstant` class, with no `DescriptorSet` or `Binding`. A `shared` variable becomes an `OpVariable` in the `Workgroup` class, named in the interface of each entry point that uses it from SPIR-V 1.4. ##### `spec(id)` — specialization constants A specialization constant is a value the host supplies when it creates the pipeline, after the shader has been compiled. It is declared as a module-level `var` carrying the constant's id, and its initializer is the default the pipeline keeps when the host supplies nothing for that id. The initializer is required: ```mach #[spec(0)] var tile_size: u32 = 64; #[spec(1)] var gain: f32 = 0.5; ``` It is a `var` like every other value the host supplies, and that settles how the compiler treats it: - It is **never a compile-time value**. It cannot be an array length or a comptime operand, since what it holds is decided after the build. A `#[spec]` on a `val` is refused for the same reason. - The optimizer **never folds it to its initializer**, even when nothing in the module writes it. A mutable global is never replaced by its initial value, and that is exactly what keeps the host's value live. - A **store to it is refused**, naming the line that wrote it. On the GPU it is a constant once the pipeline exists, so there is nothing to write. Copy it into a local to change the value. Handing its address to a function counts as a store. The type must be a scalar integer or float. There is no boolean specialization constant, because mach has no boolean type the compiler knows: `bool` is an alias of `u8`. Write a flag as an integer spec var, which the host sets with the same 4 bytes as a `VkBool32`: ```mach #[spec(3)] var use_fog: u32 = 1; ``` A narrower integer such as `u8` works too, but it needs the capability for its width like any other `u8` in a shader. On `spirv` each one becomes an `OpSpecConstant` whose literal is the initializer, decorated with `SpecId`, and a read uses that constant directly with no load. Two `#[spec]` vars with one id in the same module are refused, since the host names the constant by its id. On a machine target the decorator has no effect and the var is an ordinary global. A `#[spec]` var may also size a compute workgroup (see [`workgroup`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#workgroupx-y-z--compute-workgroup-dimensions)). ##### `handle(target, constructor, operands...)` — a type the target mints A bodyless `def` carrying this directive declares a type whose representation is **not the program's**: the owning target mints it and the pipeline binds it. ```mach fragment #[handle("spirv", "image", TEXEL_F32, DIM_2D, NO_DEPTH, NONARRAYED, SINGLE_SAMPLED, SAMPLED, FORMAT_UNKNOWN)] pub def Texture2D; #[handle("spirv", "sampled_image", Texture2D)] pub def Sampler2D; ``` The first argument names the target and the second the type constructor within it. Both are constant strings matched against the target's own definition table. Everything after them is **operands to that constructor**, never rule knobs: the rules a handle carries are fixed and closed (see [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md)) and never vary per declaration. An operand is an ordinary comptime constant, with one exception. A constructor that composes over another handle takes a **type name**, and that is the only place a decorator argument is read as a type rather than as a value. It exists so a composing declaration names what it wraps instead of restating it, which is what keeps the two from disagreeing. The named type must be a handle the same target mints, with the constructor that position requires. A declaration addressed to a target this build did not select is **inert**: it still denotes a type and still sizes at the target's pointer width, so a library of handles compiles on a machine target. A constructor name the selected target does not define, an operand count that disagrees with the constructor's, or an operand combination the target cannot emit is a compile error at the declaration. ##### `abi_type(name)` — a C type whose layout the target declares A bodyless `def` carrying this directive declares a C type whose size and alignment come from the **selected target** and whose contents the program never reaches. ```mach #[abi_type("va_list")] pub def VaList; ``` `va_list` is the only name, and the set is closed in the front end for the reason `op`'s instruction names are: which target a module is built for is not a property of the source, so a typo checked only where it is acted on would go unreported on every other build of the same library. A target that declares no layout for the named type refuses the declaration rather than substituting a default. The type is an opaque aggregate of the declared extent. A value of one may be received as a parameter and passed on, and nothing else: a local binding, a record or union field, a global, a return position, a pointer or array of one, and a cast in either direction are each refused where they are written. See [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) for the recipe and for why forwarding is the whole scope. ##### `op(target, set, name)` — a function that *is* a target instruction A shader needs `sqrt`, `normalize`, `dot` and `mix`. None of them is an operator, and none of them is a call SPIR-V can make: each is one instruction. This directive says which one a function is, so that on a `spirv` target a call to it becomes that instruction, inline, rather than a call. ```mach #[op("spirv", "GLSL.std.450", "Sqrt")] pub fun sqrt(x: f32) f32; #[op("spirv", "GLSL.std.450", "Normalize")] pub fun normalize(v: f32x4) f32x4; #[op("spirv", "core", "OpDot")] pub fun dot(a: f32x4, b: f32x4) f32; ``` The first argument names the **target**, the second the instruction set, and the third the instruction within it. The target is the ISA name the manifest selects with, so nothing about this directive is specific to one back end. The arguments must be strings on every target, but the instruction set and name are checked only when the named target is the one selected: a declaration for any other target is inert, and that target's table is not consulted. When it is checked, the parameter count is held to the instruction's own operand count, which is not uniform across a family that looks it. | Set | Meaning | |------------------|-------------------------------------------------------------| | `"core"` | the core opcode space; needs no import | | `"GLSL.std.450"` | the standard extended set; imported once per module, on use | The substitution is uniform: the emitted instruction's **result type is the function's declared return type** and its **operands are the function's parameters in declaration order**. That is what lets `dot` and `length` return a scalar from vectors, and `refract` mix a scalar operand with vector ones, without any of them being a special case. Each instruction's row in the target's table also says **how each operand is written** and **whether the instruction has a result**, and when the target is selected the declaration's types are checked against both: | Kind | The operand is | Parameter type | |-----------------|-------------------------------------------------------------|----------------| | value | an ordinary id, the argument's value | not a pointer | | constant id | an id that must be an integer constant by emission, such as a `Scope` or `MemorySemantics` | an integer | | literal | a constant written inline as a literal word, such as an image-operands mask | an integer | | pointer read | the argument's address, only read through | a pointer | | pointer write | the argument's address, only stored through | a pointer | | pointer update | the argument's address, read and written (read-modify-write) | a pointer | | handle read | a handle whose memory the instruction reads, such as a storage image's texels | a handle | | handle write | a handle whose memory the instruction writes | a handle | | truth value | a predicate the instruction takes as SPIR-V's boolean, true where the argument is nonzero | an integer | A **pointer operand takes its storage class from the call site**: the argument's own access chain decides whether it points into a storage buffer, workgroup memory, a function-local object or an image, since an `op` has no body and so no boundary at which its parameter's pointer could be given one. Any access chain is accepted, a member or element as well as a whole object. The kinds are also what the `"readonly"` and `"writeonly"` qualifiers of a `storage` binding are checked against: an atomic load through a `readonly` binding is accepted and an atomic add on it is refused, and an atomic store into a `writeonly` binding is accepted and an atomic load from it is refused. A handle passed as an ordinary value names its descriptor and touches none of its memory, and a handle read or write is checked the same way, so `OpImageWrite` into a `"readonly"` storage image and `OpImageRead` from a `"writeonly"` one are refused. A non-constant argument to a constant id or a literal is refused at the call, naming the operand. A row **without a result** is declared with no return type, and a row with one must return it. A row may also **return a pointer** into a storage class the row itself declares, as `OpImageTexelPointer` returns an `Image` pointer, and that result is accepted as a later instruction's pointer operand. A row whose result is a **truth value**, such as `OpGroupNonUniformElect`, is declared returning an integer, which receives 1 or 0. A `bool` return is an 8-bit integer, which a target without `int8` carries at 32 bits like any other 8-bit local ([manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)), so the module needs no Int8 for it. ```mach #[op("spirv", "core", "OpControlBarrier")] pub fun barrier(execution: u32, memory: u32, semantics: u32); #[op("spirv", "core", "OpMemoryBarrier")] pub fun memory_barrier(memory: u32, semantics: u32); #[op("spirv", "core", "OpAtomicIAdd")] pub fun atomic_add(p: *u32, scope: u32, semantics: u32, v: u32) u32; ``` On the target that owns the instruction, a decorated function **is the instruction and never its body**, so a call to it is never inlined away or deleted, and the optimizer treats it as reading and writing all memory. No load or store is moved across a barrier or an atomic, at any optimization level. `OpControlBarrier` takes an execution scope, a memory scope and memory semantics, and `OpMemoryBarrier` a memory scope and semantics, each an integer constant. A memory scope and memory semantics are held to the module's memory model: under the Vulkan model the `Device` scope needs the `vulkan_memory_model_device_scope` extension, and under GLSL450 the `QueueFamily` scope and the `MakeAvailable`, `MakeVisible` and `Volatile` semantics need `vulkan_memory_model` (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). An atomic's scope and semantics are held the same way. The execution scope is not, since a Vulkan barrier executes at `Workgroup` or `Subgroup` only. `SequentiallyConsistent` semantics are refused under either model, since Vulkan defines no sequentially consistent order (VUID-StandaloneSpirv-MemorySemantics-10866): use `Acquire`, `Release` or `AcquireRelease`. A control barrier must be reached in **uniform control flow**: every invocation of its execution scope executes it, or none does. That is the program's obligation, as it is in GLSL and WGSL, because whether a branch is uniform is not statically decidable in general. The compiler does not check it, and a barrier inside a branch or loop that some invocations of the scope skip is undefined behavior on the device. A row may also carry **requirements**: a capability and the extensions of the target's vocabulary that every use of it needs. A **literal operand can be enumerated**, so that its value is one of a closed set the row names (or, for a mask, a union of that set's bits), and each value brings a requirement of its own and, where the instruction grows with it, operands at the end of the instruction. Such a row has an **optional tail**: a declaration may take its required operands alone or the tail too, and each call must pass exactly the operands its literal's value brings. `OpGroupNonUniformIAdd` is one: ```mach #[op("spirv", "core", "OpGroupNonUniformIAdd")] pub fun subgroup_add(scope: u32, operation: u32, v: u32) u32; #[op("spirv", "core", "OpGroupNonUniformIAdd")] pub fun subgroup_cluster_add(scope: u32, operation: u32, v: u32, cluster_size: u32) u32; ``` Its operation is a `GroupOperation`. `Reduce` (0), `InclusiveScan` (1) and `ExclusiveScan` (2) need the `subgroup_arithmetic` extension and declare `GroupNonUniformArithmetic`, and `ClusteredReduce` (3) needs `subgroup_clustered`, declares `GroupNonUniformClustered` and is followed by the ClusterSize operand, so it is passed only to the four-parameter declaration. Where the specification makes the literal itself optional, as it does an instruction's `Image Operands` or `Memory Operands`, the literal **leads the tail**: a declaration leaves it out with every operand it would bring, or takes it followed by those operands. A mask's set bits bring theirs in ascending bit order, the order the specification writes them in, so the parameters after the mask are declared in that order. Each value types the operands it brings, so `Grad` brings two values and `ConstOffset` one constant wherever they land after the mask, and a parameter receiving one is held to its kind at the call rather than at the declaration. Each is checked at the call, where the literal's value is known: | At the call | Is refused with | |--------------------------------------------------------|------------------------------------| | a value outside the operand's enumeration | `op.operand_value`, naming the values | | a value whose operands the declaration does not pass, or passes without it | `op.operand_value`, naming the count | | a parameter the value's operand kind does not admit | `op.signature`, naming the operand | | a requirement's extension the target does not select | `spirv.capability`, naming the extension | | a capability whose SPIR-V version the environment is below | `spirv.capability`, naming the first `env` that reaches it | A device feature is an extension the target names in its `extensions` once the consumer enables it, since no environment guarantees it: `subgroup_arithmetic` is Vulkan's `VK_SUBGROUP_FEATURE_ARITHMETIC_BIT`, and a target naming no `env` holds every extension. A module declares a capability only when something it emits needs it, with `OpExtension` for a capability a SPIR-V extension defines. A row may also be **typed**: its requirement depends on the type it operates on, read from one operand (a pointer's pointee), and on the storage class that operand's memory lives in. The atomics are typed. A declaration whose type the row admits in no storage class is refused with `op.signature`, and each call is checked where its storage class is known. A load reads its pointer and every other atomic writes it, so a `"readonly"` binding admits only an atomic load. ```mach #[op("spirv", "core", "OpAtomicIAdd")] pub fun atomic_add64(p: *u64, scope: u32, semantics: u32, v: u64) u64; #[op("spirv", "core", "OpAtomicFAddEXT")] pub fun atomic_fadd(p: *f32, scope: u32, semantics: u32, v: f32) f32; ``` A 32-bit integer atomic is core in every storage class. Every other type needs the Vulkan device feature of its storage class, named for its `shaderBuffer*`, `shaderShared*` or `shaderImage*` member: `buffer_*` on buffer memory (`StorageBuffer`, `Uniform` before SPIR-V 1.3, or `PhysicalStorageBuffer` through a [physical pointer](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#pointers-on-spir-v)), `shared_*` on [`#[shared]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#inputn--outputn--builtinstr--uniformset-binding--storageset-binding--samplerset-binding--push--specid--shared--shader-interface) workgroup memory, and `image_*` on a storage image texel (`Image`, through `OpImageTexelPointer`), where Vulkan defines only a 64-bit integer and an `f32`. Any other storage class is refused. An image of 64-bit texels (`R64ui` or `R64i`) declares `Int64ImageEXT` (`SPV_EXT_shader_image_int64`), which `image_int64_atomics` enables, so the image itself needs that feature whatever reaches it. An `f16` atomic also needs `float16`, under which `f16` memory is the `OpTypeFloat 16` the atomic operates on, and the storage buffer holding one needs `storage_buffer_16bit_access` ([manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). | Rows | Type | Extensions | Vulkan feature | Capability (SPIR-V extension) | |------|------|-----------|----------------|-------------------------------| | all 15 integer atomics | `u32`, `i32` | none | core | none | | all 15 integer atomics | `u64`, `i64` | `buffer_int64_atomics`, `shared_int64_atomics`, `image_int64_atomics` | `shaderBufferInt64Atomics`, `shaderSharedInt64Atomics`, `shaderImageInt64Atomics` | `Int64Atomics` | | `OpAtomicLoad`, `OpAtomicStore`, `OpAtomicExchange` | `f16`, `f32`, `f64` | `buffer_float{16,32,64}_atomics`, `shared_float{16,32,64}_atomics`, `image_float32_atomics` | `shader{Buffer,Shared}Float{16,32,64}Atomics`, `shaderImageFloat32Atomics` | none | | `OpAtomicFAddEXT` | `f32`, `f64` | `buffer_float{32,64}_atomic_add`, `shared_float{32,64}_atomic_add`, `image_float32_atomic_add` | `shader{Buffer,Shared}Float{32,64}AtomicAdd`, `shaderImageFloat32AtomicAdd` | `AtomicFloat{32,64}AddEXT` (`SPV_EXT_shader_atomic_float_add`) | | `OpAtomicFAddEXT` | `f16` | `buffer_float16_atomic_add`, `shared_float16_atomic_add` | `shader{Buffer,Shared}Float16AtomicAdd` | `AtomicFloat16AddEXT` (`SPV_EXT_shader_atomic_float16_add`) | | `OpAtomicFMinEXT`, `OpAtomicFMaxEXT` | `f16`, `f32`, `f64` | `buffer_float{16,32,64}_atomic_min_max`, `shared_float{16,32,64}_atomic_min_max`, `image_float32_atomic_min_max` | `shader{Buffer,Shared}Float{16,32,64}AtomicMinMax`, `shaderImageFloat32AtomicMinMax` | `AtomicFloat{16,32,64}MinMaxEXT` (`SPV_EXT_shader_atomic_float_min_max`) | The features come from `VkPhysicalDeviceShaderAtomicInt64Features`, `VkPhysicalDeviceShaderImageAtomicInt64FeaturesEXT`, `VkPhysicalDeviceShaderAtomicFloatFeaturesEXT` and `VkPhysicalDeviceShaderAtomicFloat2FeaturesEXT`. No environment guarantees any of them, so a target names each one its consumer enables, and a target naming no `env` holds them all. The integer atomics are `OpAtomicLoad`, `OpAtomicStore`, `OpAtomicExchange`, `OpAtomicCompareExchange`, `OpAtomicIIncrement`, `OpAtomicIDecrement`, `OpAtomicIAdd`, `OpAtomicISub`, `OpAtomicSMin`, `OpAtomicUMin`, `OpAtomicSMax`, `OpAtomicUMax`, `OpAtomicAnd`, `OpAtomicOr` and `OpAtomicXor`. | At the call | Is refused with | |--------------------------------------------------------|------------------------------------| | a type the row admits only in other storage classes | `spirv.capability`, naming the class | | a type whose feature the target does not select | `spirv.capability`, naming the feature | A row also states how its operands' types and its result's **relate** to the type it operates on, the type of the one operand its typing reads (a pointer's pointee). A declaration that breaks a relation is refused with `op.signature`, naming both the parameter (or the return type) and the operand it relates to, so a mismatch is caught at the declaration rather than as an invalid module. | Rows | Relation | |------|----------| | every atomic | the result and each value operand are the pointer's pointee | | `OpImageRead`, `OpImageFetch`, `OpImageSample*Lod`, `OpImageGather` | the result is a 4-vector of the image's texel scalar, a sampled image's being its image's | | `OpImageSampleDref*Lod` | the result is a scalar of the image's texel scalar, and the reference is an `f32` scalar | | `OpImageDrefGather` | the result is a 4-vector of the image's texel scalar, and the reference is an `f32` scalar | | `OpImageWrite` | the texel is a scalar or vector of the image's texel scalar, with at least as many components as the image's format stores (any for `Unknown`) | | `OpImageTexelPointer` | the result points to the image's texel scalar, into a storage image of `R32ui`, `R32i`, `R32f`, `R64ui` or `R64i` format | | `OpSampledImage` | the result is a sampled image composed over the image operand's own type, and the second operand is a `sampler` | | `OpGroupNonUniformBroadcast*`, `Shuffle*`, `Quad*` and the arithmetic rows | the result is the value operand's type | | the GLSL.std.450 math rows, float and integer | the result and every operand are the first operand's type, except `Refract`'s `eta`, and `Length` and `Distance`, whose result is a scalar | | `Modf` | the result is the value's type, and the out-pointer points to the value's type, where the whole part is stored | | `Frexp` | the result is the value's type, and the out-pointer points to an `i32` scalar or vector with as many components as the value, where the exponent is stored | | `OpDot` | the second vector is the first's type | A relation on a pointer operand, such as the out-pointer `Modf` and `Frexp` store their second part through, holds the type it points to. The pointer may address a local, a `#[shared]` variable or a storage buffer, or any part of one, and the store counts as a write to that binding, so a `readonly` one is refused. ```mach #[op("spirv", "GLSL.std.450", "Frexp")] pub fun frexp(x: f32x4, exp: *i32x4) f32x4; ``` A declaration that returns a handle is refused with `op.signature` unless its row names the operand its result derives from, since a handle holds a binding's descriptor and only that operand says whose. Each math row and subgroup row states the class of number it operates on. The GLSL.std.450 float rows and `OpDot` operate on a scalar or vector of floats, and the subgroup rows on a scalar or vector of integers or floats. The GLSL.std.450 integer rows operate on the signedness their name states: `SAbs`, `SSign`, `SMin`, `SMax`, `SClamp` and `FindSMsb` on signed integers, `UMin`, `UMax`, `UClamp` and `FindUMsb` on unsigned ones, and `FindILsb` on either. A data operand of any other type is refused with `op.signature`: an integer for `FAbs`, an `i32` for `UMin`, a handle, a pointer or an aggregate. GLSL.std.450 removed `IMix`, so it has no row. Each image row names the rule its texel's component count comes from, and the refusal quotes it: the 4-vector read result is Vulkan's `VUID-StandaloneSpirv-Result-04780`, the 4-vector fetch, sample and gather results and the scalar depth-comparison result are SPIR-V's own, and the write's count against the format is Vulkan's `VUID-RuntimeSpirv-OpImageWrite-07112`, which spirv-val cannot check because it sees no `VkFormat`. The **subgroup operations** are the `OpGroupNonUniform*` rows, each taking the Subgroup scope (3) as its first operand. Every one needs SPIR-V 1.3, so `vulkan1.1` or later, and each family needs its capability and the feature Vulkan reports it by: | Family | Rows | Feature | |-------------------|--------------------------------------------------------------|-----------------------------| | basic | `Elect` | none: every vulkan1.1 device | | vote | `All`, `Any`, `AllEqual` | `subgroup_vote` | | arithmetic | `IAdd`, `FAdd`, `IMul`, `FMul`, `SMin`, `UMin`, `FMin`, `SMax`, `UMax`, `FMax`, `BitwiseAnd`, `BitwiseOr`, `BitwiseXor`, `LogicalAnd`, `LogicalOr`, `LogicalXor` with `Reduce` or a scan | `subgroup_arithmetic` | | clustered | the same rows with `ClusteredReduce` and a ClusterSize | `subgroup_clustered` | | ballot | `Ballot`, `InverseBallot`, `BallotBitExtract`, `BallotBitCount`, `BallotFindLSB`, `BallotFindMSB`, `Broadcast`, `BroadcastFirst` | `subgroup_ballot` | | shuffle | `Shuffle`, `ShuffleXor` | `subgroup_shuffle` | | relative shuffle | `ShuffleUp`, `ShuffleDown` | `subgroup_shuffle_relative` | | quad | `QuadBroadcast`, `QuadSwap` | `subgroup_quad` | `BallotBitCount` takes a `GroupOperation` too, which its ballot capability covers and which has no clustered form. `Broadcast`'s lane, `QuadBroadcast`'s index and `QuadSwap`'s direction are constant ids, which every SPIR-V version accepts. Vulkan guarantees subgroup operations only in compute stages, so a use reached from a vertex or fragment stage, an operation or a subgroup built-in alike, also needs `subgroup_graphics_stages`, the device's `subgroupSupportedStages`. ```mach use std.types.bool.bool; #[op("spirv", "core", "OpGroupNonUniformBallot")] pub fun subgroup_ballot(scope: u32, predicate: bool) u32x4; #[op("spirv", "core", "OpGroupNonUniformShuffleXor")] pub fun subgroup_shuffle_xor(scope: u32, v: u32, mask: u32) u32; ``` `OpExtInstImport "GLSL.std.450"` is emitted **once per module and only when that module uses the set**. A module that calls none of these carries no import. Note that `dot` is **core `OpDot`**, not a GLSL.std.450 instruction, even though GLSL spells it beside `normalize` and `length`. Check each function against the specification rather than against GLSL's surface. On every target other than `spirv` the directive is inert, and a decorated function is an ordinary function. A **bodiless** one — which is what the shader-side maths library uses — is then an undefined symbol, so a CPU build that calls it fails at link time naming the symbol. That is a deliberate design choice on the library's part, not a property of the directive: a decorated function may have a body, and if it does, that body is what every non-`spirv` target runs while `spirv` substitutes the instruction. A `spirv` build never emits the body at all. The set of accepted instructions is the table in `src/lang/target/isa/spirv/defs.mach`, where each row carries its operand kinds, its result, its requirements, its typing and the enumerations of its literals. The capabilities, with the SPIR-V version and extension each needs, are the table in `src/lang/target/isa/spirv.mach`. Adding an instruction is a row in it. #### Applicability | Directive | `fun` | `ext fun` | `val` / `var` | `rec` / `uni` | |-------------|:-----:|:---------:|:-------------:|:-------------:| | `deprecated`| yes | yes | yes | yes | | `testing` | yes | no | yes | yes | | `symbol` | yes | yes | yes | no | | `library` | no | yes | no | no | | `inline` | yes | no | no | no | | `noinline` | yes | no | no | no | | `align` | no | no | yes | yes | | `packed` | no | no | no | yes | | `section` | yes | yes | yes | no | | `oblivious` | yes | no | no | no | | `scalar` | yes | no | no | no | | `naked` | yes | no | no | no | | `extensions`| yes | no | no | no | | `embed` | no | no | yes | no | | `stage` | yes | no | no | no | | `workgroup` | yes | no | no | no | | `input` | no | no | yes | no | | `output` | no | no | yes | no | | `builtin` | no | no | yes | no | | `uniform` | no | no | yes | no | | `storage` | no | no | yes | no | | `push` | no | no | yes | no | | `spec` | no | no | yes | no | | `op` | yes | no | no | no | | `handle` | no | no | no | no | | `abi_type` | no | no | no | no | The `val` / `var` column is shared, but `embed` accepts only `val` — a `var` is refused (see [`embed`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#embedstr--compile-time-file-embedding) above). `spec` is the reverse and accepts only `var` (see [`spec`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#specid--specialization-constants)). `deprecated` also applies to `tag`, `def`, `use` and `fwd` declarations and to a tag case, and `testing` to `tag`, `def`, `use` and `fwd` declarations, none of which the table columns cover. The set is closed. New directives require a compiler change. #### See also - [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md) — `test` blocks, whose semantics `testing` gives a declaration - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) — `ext` imports, `library` and `symbol` use cases - [visibility.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/visibility.md) — `pub` / `ext` visibility (not decorator-controlled) - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) — `$size_of` / `$align_of` as `align` arguments - [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) — `^` secret types and the `oblivious` constant-time contract - [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md) — inline `asm`, the only body a `naked` function may have, and the extension instructions `extensions` admits - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) — `val` / `var` bindings, and the `embed` exemption to `val`'s initializer requirement - [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#types) — the `[_]` inferred array length `embed` introduces - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) — the SIMD vector types a shader stage computes over and `op` operates on, and the handle types `handle` declares - [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md) — the `vectorize` profile key `scalar` opts out of, and content-fingerprinted build inputs (`embed`, `[step]` `in`) Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/def.md ### `def` — type alias `def` introduces a new name for an existing type. The alias and the underlying type are interchangeable; there is no nominal distinction. #### Grammar ```mach fragment def NAME: TYPE; ``` #### Examples ```mach pub def Age: i64; # alias for a primitive pub def BinOp: fun(i64, i64) i64; # alias for a function type pub def Anon: rec { x: i64; y: i64; }; # inline record pub def Choice: uni { a: i64; b: f64; }; # inline union ``` Aliases may name any type: primitives, pointers, arrays, function types, records, unions, tags, or other aliases. `def` is a module-scope declaration; there is no function-scope alias. An alias of a tag constructs, tests and copies as the tag: ```mach use std.types.result.res; tag ParseError: u8 { invalid; overflow; } def R: res[i64, ParseError]; fun parse(x: i64) R { if (x < 0) { ret R.err{ParseError.invalid{}}; } ret R.ok{x}; } ``` #### Stdlib aliases Mach has no compiler-known type aliases. Names like `usize` and `str` live in stdlib as ordinary `def`s — `def usize: u64;` (or `u32`, target- conditional via `$if`) and `def str: *u8;`. A module that wants the shorthand imports the appropriate stdlib module. #### See also - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) - the type grammar def references - [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md), [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md), [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - aggregate forms commonly aliased Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md ### `rec` — record A `rec` is a named, structurally-laid-out collection of typed fields. Each field has its own storage; the record's size is the sum of field sizes plus any padding for alignment. #### Grammar ```mach fragment rec NAME { field1: type; field2: type; ... } rec NAME[T, U] { ... } # generic over type parameters ``` #### Examples ```mach pub rec Point { x: i64; y: i64; } pub rec Pair[T, U] { left: T; right: U; } ``` #### Construction and access A record literal names the type and provides each field by name: ```mach use std.print; use std.runtime; rec Point { x: i64; y: i64; } rec Pair[T, U] { left: T; right: U; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val p: Point = Point{x: 1, y: 2}; val q: Pair[i64, u8] = Pair[i64, u8]{left: 5, right: 6u8}; val n: i64 = p.x; # field access via . print.printlnf("{} {}", n, q.right); ret 0; } ``` A literal may leave fields out. **Every omitted field is zero-initialized**: integers and floats read `0`, pointers read `nil`, and a nested record or array is zero in every byte. This holds for secret (`^`) fields and secret-welded pointers (`*^T`) as well, for a literal built at runtime from non-constant values, for one built from constants inside a function, and for a constant literal that initializes a module-level `val`. `T{}` names no field and so is all zero. The guarantee is a contract, not an accident of the stack: a literal never exposes what the storage held before. ```mach rec Key { id: u64; material: ^[32]u8; buf: *^u8; } fun fresh(id: u64) Key { ret Key{id: id}; # material is all zero, buf is nil } val EMPTY: Key = Key{id: 0}; # the same holds for a constant literal ``` #### Layout By default the compiler may insert padding between fields for alignment. The `#[align(N)]` decorator on a record raises its minimum type alignment to `N` bytes (a power of two), and `#[packed]` lays the record out with no padding at all, for a shape whose layout is fixed elsewhere (a C struct, a file header, a wire frame); see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#packed--no-padding) for what a packed record refuses. `#[volatile]` makes every access to the record's storage a volatile access, the way a memory-mapped register block is declared; see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#volatile--every-access-to-the-type-is-a-volatile-access). #### See also - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - tagged value - [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md) - overlapping-memory counterpart - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) - #[align], #[packed] and #[volatile] - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) - $size_of, $offset_of Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md ### `uni` — raw union A `uni` is a collection of named fields that share the same memory. Writing to one field overwrites whatever bytes were in the others. The compiler does not track which field is "live"; the user is responsible for that. #### Grammar ```mach fragment uni NAME { field1: type; field2: type; ... } uni NAME[T] { ... } # generic ``` #### Examples ```mach pub uni Number { i: i64; f: f64; } pub uni Maybe[T] { some: T; none: u8; } ``` A `uni`'s size is the size of its largest field, plus any alignment padding. A `uni`'s overlapping variants must agree on secrecy: every field is either secret (`^`) or every field is public. A mixed union would let the same storage be read at two secrecy classifications, the aliasing leak the welded-storage rules forbid. See [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md). #### Raw unions versus tagged values A `uni` is unchecked, raw memory. The compiler tracks neither which variant was written nor whether reading a variant is valid. For safe, discriminated values where the selected case is stored and payload access requires a guard, use `tag`. See [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). Low-level systems code can still compose `rec` and `uni` manually when modeling foreign data structures, hardware registers, or wire formats: ```mach fragment rec RawPacket { kind: u8; data: uni { header: Header; raw: [64]u8; }; } ``` With manual composition, the compiler does not verify that `kind` and `data` agree. Keeping them consistent is the programmer's responsibility. In ordinary Mach code, prefer first-class `tag` declarations. A union may carry `#[align(N)]`, `#[packed]` and `#[volatile]` like a record; `#[volatile]` makes every access to the union's storage a volatile access, for a device register that reads and writes as different types. See [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#volatile--every-access-to-the-type-is-a-volatile-access). #### See also - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - checked tagged values with guarded payload access - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) - `#[align(N)]`, `#[packed]`, `#[volatile]` - [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md) - records and sequential aggregate layout - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) - if/or chains for branching - [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) - secrecy agreement across overlapping variants Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md ### `tag`: tagged value A `tag` is a discriminated aggregate value that represents exactly one active case at any moment. Each case has a name and either one explicitly typed payload or no payload at all. This page is the tagged-value contract (#3217): declarations with an explicit discriminator, `Type.case{payload}` construction, `sel` case tests, lexical payload guards, the debug-profile discriminator trap, checked layout and the reflection intrinsics. There is no flow-sensitive analysis, no `try` expression and no compiler-known failure type: `res`, `opt` and `err` are ordinary std tags. #### Grammar ```mach fragment tag NAME: u8 { case1; case2: type; ... } tag NAME[T, E]: u8 { ... } # generic over type parameters ``` A case is declared with an identifier, an optional colon followed by a payload type, and a terminating semicolon. Empty tags and duplicate case names are rejected. Tags use the standard aggregate declaration and generic parameter rules. Tag declarations reject names that match SIMD vector spellings such as `f32x4`, because vector spellings resolve as vector types in type positions. #### Examples ```mach pub tag Reply: u8 { empty; value: i64; } pub tag ParseError: u8 { invalid; overflow; } pub tag Tree[T]: u8 { leaf: T; empty; } ``` When multiple values must accompany a case, use an ordinary record payload: ```mach use std.types.size.usize; use std.types.string.str; pub tag Entry: u8 { none; pair: rec { key: str; count: usize; }; } ``` An empty tag, a duplicate case name and a discriminator too narrow for the case count are rejected: ```mach error duplicate tag case name tag Twice: u8 { one; one; } ``` #### Construction and initialization A tag value is constructed by naming the type, the case and the payload: ```mach tag Reply: u8 { empty; value: i64; } val empty_reply: Reply = Reply.empty{}; val num_reply: Reply = Reply.value{42}; ``` The payload is positional because a case carries exactly one, and a payloadless case takes empty braces. There is no other construction form: - Omitting a payload on a payload-bearing case (`Reply.value{}`) is a compile error - Supplying a payload to a payloadless case (`Reply.empty{1}`) is a compile error - Supplying more than one payload (`Reply.value{1, 2}`) is a compile error - Naming the payload (`Reply.value{value: 1}`) is a compile error - The record-literal form (`Reply{value: 1}` or `Reply{empty}`) is a compile error - A case selector alone (`Reply.value`) is not a value ```mach error this tag case requires a payload tag Reply: u8 { empty; value: i64; } val missing: Reply = Reply.value{}; # the case declares a payload ``` Whole-value assignment replaces the selected case and payload together. ##### Default initialization Zero initialization selects the first declared case and zero-initializes its payload if one exists. ```mach use std.print; use std.runtime; tag Reply: u8 { empty; value: i64; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var reply: Reply; # selects Reply.empty if (sel reply.empty) { print.println("empty"); } ret 0; } ``` #### The std failure tags The canonical failure types are ordinary std tags, declared in `std.types.result`, `std.types.option` and `std.types.error` with fixed generic arities and no compiler knowledge of their names. They use the same mechanisms as every user tag: the same construction form, `sel`, guards, layout and reflection. A module imports the ones it spells: ```mach use std.types.error.err; use std.types.option.opt; use std.types.result.res; ``` std declares them as: ```mach pub tag res[T, E]: u8 { err: E; ok: T; } pub tag opt[T]: u8 { none; some: T; } pub tag err[E]: u8 { err: E; ok; } ``` - `res[T, E]` is either an error of type `E` or a value of type `T`; `err: E` is the first declared case and `ok: T` the second. - `opt[T]` is either absence or a value of type `T`; payloadless `none` is first and `some: T` second. - `err[E]` is either an error of type `E` or payloadless success; `err: E` is first and payloadless `ok` second. It is distinct from `opt[E]`, not an alias. There is no defaulted type argument, general type inference, dummy success type, unit value, constructor function or automatic error conversion. ```mach use std.types.error.err; use std.types.option.opt; use std.types.result.res; tag ParseError: u8 { invalid; overflow; } val good: res[i64, ParseError] = res[i64, ParseError].ok{42}; val bad: res[i64, ParseError] = res[i64, ParseError].err{ParseError.invalid{}}; val present: opt[i64] = opt[i64].some{42}; val absent: opt[i64] = opt[i64].none{}; val finished: err[ParseError] = err[ParseError].ok{}; val failed: err[ParseError] = err[ParseError].err{ParseError.overflow{}}; ``` Default initialization of `opt[T]` selects `none`. Default initialization of both `res[T, E]` and `err[E]` selects `err` with a zero-initialized error payload. The case names `ok`, `err`, `some` and `none` are members of their tags, not keywords. Their layouts and closed case sets are part of std's SemVer contract; they ship in std 2.0.0 paired with Mach 5.0.0. They are std declarations like any other: the compiler has no knowledge of the three names, they resolve only through an import or a declaration in scope, and a module may declare its own `res`, `opt` or `err` as a tag or as anything else. Declaring one twice in a module is the ordinary duplicate definition. A module that spells `res` without importing it is rejected the way any unresolved type name is: ```mach error unresolved type name `res` tag ParseError: u8 { invalid; } fun parse(x: i64) res[i64, ParseError] { ret res[i64, ParseError].ok{x}; } ``` #### Case tests `sel place.case` is a boolean expression that is true when `place` currently holds `case`. It reads only the discriminator, never a payload, and has no side effects. ```mach tag Reply: u8 { empty; value: i64; } fun describe(reply: Reply) i64 { if (sel reply.value) { ret reply.value; # reply holds value here } or { ret 0; # reply holds empty here } } ``` The operand is a place: a binding, a field, an index or a dereference, followed by exactly one case name of that place's tag type. A pointer to a tag is a place too: `sel p.case` with `p: *Reply` auto-dereferences the pointer once and tests its pointee, exactly as the payload place `p.value` already reads through it. A call or other temporary is not a place. The result is an ordinary `bool`, so it composes with `!`, `&&` and `||`, can initialize a `bool` binding, and can be returned. ```mach use std.types.bool.bool; use std.types.bool.false; use std.types.option.opt; tag Reply: u8 { empty; value: i64; } fun tests(reply: Reply, next: opt[i64], a: opt[i64], b: opt[i64]) bool { val done: bool = sel reply.value; for { if (!sel next.some) { brk; } brk; } if (sel a.some && sel b.some) { ret done; } ret false; } ``` `sel` takes a place, so a call result is refused: ```mach error `sel` tests a place tag Reply: u8 { empty; value: i64; } fun make() Reply { ret Reply.empty{}; } fun test() i64 { if (sel make().empty) { ret 1; } ret 0; } ``` `sel` is a keyword. Comparing a tag with `==`, whole-tag equality, payload equality, ordering, a `.kind` field and a `match` construct do not exist. An outer-secret `^Tag` protects the selected case as well as the payload, so `sel` refuses one. `sel` also evaluates at comptime over a constant tag. A `$if (sel C.case)` gate on a module `val` constructed with `C.case{...}` tests the constant's selected case, guards `C.case` in its arm the same way a runtime chain arm does, and its payload reads fold to the constant payload. A comptime `sel` on anything that is not a constant tag element (a parameter, a local, a module `var`, or a constant of another module) is rejected with a located diagnostic at the condition, and a payload read of a case the constant does not hold is a compile-time error rather than the runtime undefined behavior, because comptime state is never stale. #### Payload places and guards `value.case` is a payload place. Reading it, writing it and taking its address are legal only inside a guard for that place and case. A guard is a lexical region, not a flow fact: a chain arm whose condition is exactly `sel P.c` guards `P.c` inside its block, and a chain whose every arm exits guards the remainder of the enclosing block for the case the chain left untested. ```mach tag Reply: u8 { empty; value: i64; } fun read_value(reply: Reply) i64 { if (sel reply.value) { ret reply.value; # guarded by the arm condition } ret 0; } ``` A payload read outside a guard is a compile error: ```mach error a tag payload requires a `sel` guard tag Reply: u8 { empty; value: i64; } fun read_value(reply: Reply) i64 { ret reply.value; # no guard opens here } ``` Inside a condition, the right operand of `&&` is guarded by a `sel P.c` that is its left operand, because `&&` short-circuits. `||`, `!` and every other operator open no guard. A guard is matched by spelling, so a guarded place must be a path of identifiers, fields, dereferences and indexes by a name or a literal, `sel` rejects anything else, and reassigning an index binding inside the guard is not tracked and leaves the debug-profile discriminator check as the only net. Inside a guard the payload place is ordinary storage: reading it copies under the existing value rules, writing it keeps the selected case, and `?value.case` yields a typed pointer to naturally aligned storage. Whole-value assignment to the guarded place is a compile error; rebind to a new name instead. ```mach error cannot assign to this place inside a guard tag Reply: u8 { empty; value: i64; } fun reset(reply: Reply) i64 { var r: Reply = reply; if (sel r.value) { r = Reply.empty{}; # whole-value assignment inside the guard ret 1; } ret 0; } ``` The check covers the guarded place and every object it is reached through by value (`b.r = ...` under `sel b.r.value` is refused too). A pointer ends that chain: under `sel p.value` with `p: *Reply`, reassigning `p` or writing `@p = ...` is not tracked, because the guard covers the storage the pointer reached when the test ran, and a write through a pointer is the ordinary raw-memory obligation below. A payload read whose case is no longer selected is undefined behavior of the same class as a stale pointer read. A raw pointer to a payload does not pin a case or extend a lifetime. There is no borrow checker, proof analysis or runtime validator. In the debug profile only, each guarded payload access compares the discriminator and traps on mismatch. #### Layout and representation A tag is laid out with its discriminator at byte offset zero in target byte order. The discriminator type is the one the declaration names after the tag name (`tag Name: u8 { ... }`), one of `u8`, `u16`, `u32` or `u64`; a type too narrow to number every declared case is rejected. Declaration order determines case codes, starting at zero. Single-case tags still include a discriminator. Let `D` be the discriminator size. Let `M` be the maximum payload size, and `PAlign` the maximum payload alignment (or 1 when no cases have payloads). The common payload offset is `align_up(D, PAlign)`. Total object alignment is the maximum of discriminator alignment, payload alignment, and any explicit `#[align(N)]` decorator. Total size is the aligned extent of the common payload area. Tags with no payloads contain only the discriminator and trailing padding. Packing with `#[packed]` places the payload immediately after the discriminator at offset `D` with base alignment 1. Explicit alignment raises object alignment and rounds total size without repacking nested payloads. Packed value access uses legal unaligned operations. A typed payload pointer must not promise stronger alignment than its storage provides. Construction captures the active payload before overwriting the destination and zeroes tag-owned gaps, inactive payload suffix bytes, and tail padding. Active payload representation rules remain unchanged, including raw union bytes. Replacing a larger case clears the bytes that become inactive. Representation-changing `::` and `:~` casts are rejected when either by-value representation contains a tag, including through records, arrays, or unions. Same-type identity casts remain valid, and transparent aliases preserve type identity. Ordinary pointer retyping remains an explicit raw-memory operation. Reading through a typed tag pointer requires live, aligned storage with a valid case code and a valid selected payload. A pointer cast does not validate raw data. Stripping secrecy with `:>` removes only outer secrecy from the tag value and preserves the active case and inner payload qualifiers. #### Secrecy and ownership A public tag has a public discriminator, while each payload retains its declared secrecy. Outer `^Tag` also protects the active case. Secret-dependent case tests must obey the constant-time rules. Copies preserve potentially secret storage across all cases and use the fixed public type extent, rather than choosing a copy size from a secret active case. An immutable pointer binding does not make its pointee immutable. Raw payload pointers carry the lifetime and active-case obligations above. Tags introduce no borrow checker, moves, destructors, allocation, unwinding, or automatic resource rollback. Address-bound owners initialize caller-owned final storage and are not returned by value inside results. #### Reflection Tags support comptime reflection through three intrinsics: - `$is_tag(T)` returns true for public tag shapes, including canonical tags - `$cases(T)` yields a comptime sequence of case descriptors - `$discriminant_of(T)` returns the unsigned integer type used for the discriminator Case descriptors expose `name`, `has_payload`, `type`, `offset`, and `code`. Accessing `type` or `offset` on a descriptor where `has_payload` is false is an error. ```mach fragment $each case in $cases(T) { if (sel value.[case]) { $if (case.has_payload) { consume[case.type](value.[case]); } } } ``` `sel value.[case]` is the case test, `value.[case]` is the guarded payload place, and `T.[case]{payload}` constructs through the same single-case rule after specialization. #### Deprecation `#[deprecated]` and `#[deprecated("msg")]` apply to a `tag` declaration and, inside the body, to a single case, where it is the only decorator a case accepts. A deprecated case warns at every external construction, `sel` test and payload place that names it; see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md). `#[volatile]` on a tag makes every access to its storage, discriminator and payload alike, a volatile access; see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#volatile--every-access-to-the-type-is-a-volatile-access). #### See also - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) - `#[deprecated]`, `#[packed]`, `#[align(N)]`, `#[volatile]` - [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md) - records and struct layout - [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md) - raw unions - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) - primitive and compound type reference - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) - reflection intrinsics Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md ### `fun` — function A function takes typed arguments, optionally returns a typed value, and has a body of statements. #### Grammar ```mach fragment fun NAME(args) RET { ... } # function with return type fun NAME(args) { ... } # no return type fun NAME[T](args) RET { ... } # generic over type parameters fun NAME($p: T, args) RET { ... } # comptime value parameter fun NAME(fixed, va: ...) RET { ... } # variadic pack parameter ``` #### Examples ```mach rec Pair[T, U] { left: T; right: U; } var counter: i64 = 0; pub fun add(a: i64, b: i64) i64 { ret a + b; } pub fun bump() { # no return value counter = counter + 1; } pub fun identity[T](value: T) T { ret value; } pub fun make_pair[T, U](a: T, b: U) Pair[T, U] { var p: Pair[T, U]; p.left = a; p.right = b; ret p; } ``` #### Generic type parameters Generic functions take type parameters in brackets `[T]`. There are no constraints; any type may be substituted. The compiler monomorphizes per unique type instantiation. Call sites supply the types explicitly: ```mach fragment val x: i64 = identity[i64](42); val p: Pair[i64, u8] = make_pair[i64, u8](1, 2u8); ``` The same spelling with no call after it **names the instance as a value**. Its type is the instantiated signature, so it can be stored, passed, returned, and addressed: ```mach fragment val f: fun(i64) i64 = identity[i64]; # the i64 instance, as a value val p: ptr = ?identity[i64]; # its address ret apply(identity[i64], 42); # passed as a callback ``` A generic itself is a template rather than code, so a bare `identity` has no address and cannot be a value; only an instance can. ##### A body is checked at each instantiation There are no constraints on `T`, so nothing about a type parameter is decided before an instantiation supplies its argument. An operator, a cast, a `:~` reinterpret, a literal beside a `T`-typed operand and a branch condition are all resolved against the instance's concrete type, so an ordered or numeric algorithm is writable without a comparator parameter: ```mach fun maxof[T](a: T, b: T) T { if (a > b) { ret a; } ret b; } fun sum[T](p: *T, n: u64) T { var acc: T = 0; # the literal takes `T` var i: u64 = 0; for (i < n) { acc = acc + p[i]; i = i + 1; } ret acc; } ``` A type that genuinely does not support the operator is refused at the instantiation that asked for it, naming that type: ``` error[operator.operand_type]: `Point` has no ordering: `<`, `<=`, `>` and `>=` order integers and floats --> src/main.mach:12:5 | 12 | maxof[Point](p, q); | ^^^^^^^^^^^^^^^^^^ | --> src/main.mach:4:9 | 4 | if (a > b) { ret a; } | ----- in this generic body, checked against this instance's type arguments ``` The template itself reports nothing. It types the body so each instance and the lowering have an expression table to read, and every judgement is the instance's, which is what keeps one bad instantiation to one refusal at one site. **A generic nothing instantiates is not checked.** `fun never_used[T](a: T) T { ret a + a; }` compiles, because there is no type to decide `+` against and no site to report at. The way to check a generic is to instantiate it, so a library instantiates its own generics in its own tests. #### Comptime value parameters A parameter marked with `$name: T` must be supplied with a value the compiler can resolve at compile time. The function body can branch on the comptime parameter via `$if`, producing different code per call-site instantiation. ```mach fragment pub fun pick_op($mode: Mode, a: i64, b: i64) i64 { $if (mode == MODE_FAST) { ret a + b; } $or (mode == MODE_SAFE) { # extra logic here ret a + b; } } ``` Comptime value parameters apply to function parameters only — not record fields, not other contexts. #### Variadic packs A function with a trailing named pack parameter (`va: ...`) accepts a variable number of trailing arguments. The compiler monomorphizes the function once per distinct call-site type-list; the pack is consumed by `$each a in va` at compile time — there is no runtime `va_list`. ```mach use std.print; use std.runtime; pub fun sum(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a; } ret t; } # leading fixed params are allowed before the pack pub fun bias(base: i64, va: ...) i64 { var t: i64 = base; $each a in va { t = t + a; } ret t; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {}", sum(10i64, 20i64, 30i64), bias(1, 10i64, 20i64, 30i64)); ret 0; } ``` `va.len` folds to the element count. `g(va...)` forwards the whole pack to another pack-tailed callee. See [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) for the full reference. #### Tail calls The release pipeline (`opt = 2`) treats a call whose result the function returns unchanged, as in `ret f(a, b);`, as a tail call. A call to the function itself becomes a loop: the recursion runs in the stack of one frame at any depth. `ret f(n - 1) + x`, and the same with `*`, `|` or `^` on integers, is a loop too, with the pending operands carried in an accumulator. A call to any other function becomes a jump that hands the callee the caller's frame, so a chain of such calls, mutual recursion included, also runs in constant stack. A call stays a call when the function has a `fin` pending at the `ret`, when the address of one of its locals escapes (to the callee or to any earlier call), when it runs inline assembly or is `#[naked]` or `#[oblivious]`, when a by-value argument is a copy the caller has to make, or when the callee takes more stack arguments than the caller received. The debug pipeline (`opt = 0`) makes no tail calls. `-g` changes none of this: debug information is added to the same code. A debugger or unwinder therefore never sees the frame a tail call left. A backtrace taken inside the callee shows it directly below the caller's own caller, and a self tail call shows as the one frame of the loop, its parameters holding the current iteration's values. A local of the caller that holds the call's result reads as optimized out, since the caller never receives it. #### See also - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) — body-less external functions - [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) — variadic pack parameter reference - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` inside function bodies, and the two regimes a generic body is checked under - [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) — the operator table an instance is checked against - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) — function calls and generic instantiation Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md ### `ext fun` — external function `ext fun` declares a function with C ABI as a forward reference — body-less. The linker resolves the symbol at link time. This is the only body-less function form Mach allows. #### Grammar ```mach fragment ext fun NAME(args) RET; ``` - No body block. The declaration ends with a semicolon. - Argument and return types must be representable in C. - `pub ext fun` exposes the import to other modules; without `pub`, the import is file-private. #### Examples ```mach pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ext fun strlen(s: *u8) i64; # private, file-local ``` #### C-variadic imports A C function declared `int open(const char *, int, ...)` is **C-variadic**: its parameter list ends in `...` and a call may pass further arguments the declaration does not type. Spell it the same way: ```mach ext fun open(path: *u8, flags: i32, ...) i32; ext fun fcntl(fd: i32, cmd: i32, ...) i32; ``` - The `...` is a **bare** trailing marker written after the last fixed parameter. It is not a named parameter and has no type. - **`ext` only.** A mach-bodied `fun` has no C-variadic form — its variadic is the comptime pack `va: ...`, an ordinary *named* parameter with an entirely different meaning (see [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md)). A bare `...` in a non-`ext` parameter list is an error that names the pack. - **At least one fixed parameter** must precede the `...`, as in C. Every psABI defines its variadic rule relative to the named arguments. - Only the **call** side exists. Mach cannot *define* a C-variadic function: there is no `va_list` and no `va_arg`, so such a declaration cannot carry a body. The declared parameters are the *fixed* arity. A call must supply all of them and may supply any number of further arguments, including none. ##### Naming the signature as a type The same shape is spellable as a function **type**, so a C-variadic function can be stored in a variable, passed as a callback, or held in a table: ```mach ext fun printf(fmt: *u8, ...) i32; def Printf: fun(*u8, ...) i32; fun main() i32 { val p: Printf = printf; ret p("%d\n"::*u8, 42); } ``` The `...` is part of the type's **identity**, not a modifier on it. `fun(*u8, ...) i32` and `fun(*u8) i32` are different types and do not unify: a call through the first places its tail by the target's variadic rule and a call through the second does not, so mixing them would reach a C callee's `va_arg` with arguments laid out the ordinary way — a wrong *value* at run time, not a link error. The same two rules as the declaration apply: the `...` is bare and trailing, and at least one fixed parameter must precede it. An indirect call through such a type is checked and placed exactly as a direct one is; the tail rule is read off the type, since no declaration is in view at the call. ##### Arguments in the variadic tail A tail argument has no declared parameter type to check against, so it is checked against the C variadic contract directly. C applies **default argument promotions** there: an integer narrower than `int` arrives as an `int`, and a `float` arrives as a `double`. Mach inserts no implicit conversions anywhere and makes no exception here — a narrow argument is **rejected**, naming the cast: ```mach fragment var c: u8 = 65; var f: f32 = 1.5; log(fmt, c); # error: cast it to `i32` / `u32` log(fmt, f); # error: cast it to `f64` log(fmt, c::i32, f::f64); # correct ``` An argument of 32 bits or wider passes as written: `i32`/`u32`, `i64`/`u64`, `f64`, any pointer, and a record or vector passed by value. A **secret** (`^T`) may not enter a variadic tail: the callee is foreign code that prints or logs the value, so it is an observation boundary — declassify with `:>T` first. ##### Per-target argument passing The fixed parameters always follow the target's ordinary calling convention. The *tail* is where targets differ: | Target | Variadic tail | |--------|---------------| | **Apple arm64** (`darwin` + `aarch64`) | **Every** tail argument is passed on the **stack**, naturally aligned with an 8-byte minimum. There is no register phase at all — not for integers, not for floats. | | **Other AAPCS64** (`linux` + `aarch64`) | Ordinary rule: register-then-stack, exactly as for a named argument. | | **System V x86-64** (`linux` / `darwin` + `x86_64`) | Ordinary rule, plus `AL` set to the number of vector registers the call uses. It is set for **every** call to a C-variadic callee, including one whose tail is empty. | | **Microsoft x64** (`windows`) | Ordinary rule. A tail float rides **both** its XMM register and the integer register of the same positional slot, since a `va_arg` reader walks only the integer slots. Vectors follow the extent-based carrier mapping below. | A tail argument keeps its *form* on every target: a record too large to pass by value is still passed by reference, and only the hidden pointer's location moves. The *fixed* parameters diverge on Apple arm64 too, on a separate axis: a fixed argument that lands on the stack there takes its **natural size and alignment**, where AAPCS64 rounds every stack argument up to an 8-byte slot. Past the eight GP registers, `(…, int i, short j, long long k)` occupies `[sp+0]`, `[sp+4]`, `[sp+8]` on `darwin-aarch64` and `[sp+0]`, `[sp+8]`, `[sp+16]` on `linux-arm64`. A record that is not a homogeneous float or vector aggregate is the exception: Apple arm64 gives it a slot aligned to at least 8 bytes and rounded up to a multiple of 8, as AAPCS64 does. Mach applies each target's own rule, so nothing in a declaration has to say which one is in force. Note only that Apple's two rules genuinely differ from each other: the variadic tail in the table above keeps its 8-byte minimum on the very target where a fixed argument does not. Apple arm64 is the reason this form exists. The obvious workaround — declaring the call at its fixed arity, `ext fun open(path: *u8, flags: i32, mode: u32) i32` — works on Linux and on x86-64 and is **silently wrong** there: the fixed-arity lowering puts `mode` in `x2` while libSystem reads it off the stack, so the callee sees an uninitialized value. That is a wrong value, not a link error and not a crash. Mach cannot detect the mistake, because nothing in a fixed-arity declaration says the callee is variadic — so **declare every C-variadic callee with `...`**, on every target. #### `va_list` parameters A C function may take a `va_list` as an ordinary **fixed** parameter. The common shape is a logging callback: ```c typedef void (*TraceLogCallback)(int logLevel, const char *text, va_list args); void SetTraceLogCallback(TraceLogCallback callback); ``` That is a fixed 3-arity function. It has no `...` in it, and binding it needs nothing from the C-variadic form above — only a parameter type that is ABI-correct in the third slot. Declare that type with `#[abi_type("va_list")]` on a bodyless `def`: ```mach use std.types.size.usize; #[abi_type("va_list")] def VaList; pub def TraceLogCallback: fun(i32, *u8, VaList) i64; #[library("raylib")] ext fun SetTraceLogCallback(cb: TraceLogCallback); ext fun vsnprintf(buf: *u8, n: usize, fmt: *u8, ap: VaList) i32; fun trace(level: i32, text: *u8, args: VaList) i64 { var buf: [512]u8; ret vsnprintf(?buf[0], 512, text, args)::i64; } ``` The declaration carries no size of its own. The selected target declares what a `va_list` is on that platform, and the same source text binds correctly on every one of them. A target that declares none refuses the declaration by name rather than guessing. ##### Mach can forward one and never read one The whole scope is **forward-only**: receive an opaque token from C and pass it to C. There is no `va_start`, no `va_arg`, no `va_end`, and no way to construct one. Each of these is refused where it is written: ```mach fragment val saved: VaList = args; # refused: a local binding rec Held { ap: VaList; } # refused: a record or union field var kept: VaList; # refused: a global fun make() VaList { … } # refused: a return position fun peek(ap: *VaList) { … } # refused: behind a pointer val raw: ptr = args :~ ptr; # refused: a cast in either direction ``` Reading one needs `va_arg`, and mach cannot have `va_arg`: it has no runtime variadics, so it has no callee that walks its own argument tail. That is not a gap waiting to be filled. Constructing a `va_list` is only meaningful for such a callee, and mach never needs to originate one — it would call `printf` rather than `vprintf`. One further rule comes from C rather than from mach: **a forwarded `va_list` is spent.** Handing one to a function that reads it leaves its value indeterminate, exactly as in C, and what actually happens differs per platform — under AAPCS64 the type is a composite the caller copies, so each callee walks its own; under System V x86-64 it is a pointer into shared state that the callee advances. Forward it once. ##### Why the type is declared per target `va_list` is the one C type mach binds whose shape cannot be written down portably: | Target | `va_list` in a parameter slot | |--------|-------------------------------| | **System V x86-64** | `__va_list_tag[1]`, an array type, so a parameter decays to one pointer | | **Apple arm64** | `char *` | | **Microsoft x64** and **Windows arm64** | `char *` | | **RISC-V lp64d** | `void *` | | **AAPCS64 linux** | a 32-byte, 8-aligned composite | On the first four, spelling the parameter `ptr` is already ABI-correct. On AAPCS64-linux it is not, and the way it fails is the reason this type exists rather than a documented caveat: AAPCS64 passes a composite over 16 bytes indirectly, with the **caller** owning the copy, so a `ptr` forwards the received pointer instead of a fresh copy and the next callee advances the caller's list. That usually appears to work, which is worse than failing. #### Aggregate arguments and the caller's copy **The caller owns the copy**, on every convention and on every call edge. An aggregate too large for registers is passed as the address of a copy the caller allocates on AAPCS64, under the RISC-V psABI and under the Microsoft convention, and in the outgoing stack area under System V. The callee treats that memory as its own parameter and may overwrite it; nothing copies again at entry: ```mach rec Wide { a: i64; b: i64; c: i64; d: i64; } ext fun consume(w: Wide) i64; ``` This is the platform convention, and mach has no second one. A Mach→Mach call, a call through a `fun` value, a callback C invokes and an exported symbol reached by `dlsym` are all the same edge, so a C caller and a mach caller hand a mach callee the same thing, and a C callee and a mach callee do the same thing with it. Mach used to copy an incoming aggregate into storage of its own at entry as well, which made two copies of every such argument; [mach#3418][3418] ruled that out and [mach#3416][3416] removed it. What the language guarantees on top of the convention is the capture point, not a second copy: reading an aggregate captures its value where it is read, so a later argument cannot change what an earlier one passed (see [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md)). [3418]: https://github.com/briar-systems/mach/issues/3418 [3416]: https://github.com/briar-systems/mach/issues/3416 #### 128-bit integers and `__int128` `u128` and `i128` are C's `unsigned __int128` and `__int128` at every boundary: 16 bytes, 16-byte aligned, low half first in memory. Each convention passes them the way its C compilers do, and mach follows the convention rather than a rule of its own: | Convention | Argument | Return | |---|---|---| | System V x86-64 | two consecutive integer registers (`rdi:rsi`, ... `r8:r9`); when fewer than two are left the whole value goes to the stack, never split | `rax:rdx` | | AAPCS64 | the register number is rounded up to even and the value takes that pair (`x0:x1`, `x2:x3`, ...); a skipped register is not reused (C.10). When fewer than two are left the remaining registers are given up and the value goes to the stack, 16-byte aligned (C.11) | `x0:x1` | | RISC-V LP64 | two consecutive argument registers; with one left, that register carries the low half and the stack the high half; with none, the stack | `a0:a1` | | Win64 | a pointer to a caller-owned, 16-byte-aligned copy, in the argument's positional slot, like every other 16-byte object | `XMM0`, low half in the low 64 bits | Microsoft's x64 convention has no 128-bit integer (MSVC has no such type). The Win64 row is the convention GCC (mingw-w64) and Clang share: LLVM 18 made `i128` match `__int128` ("Changes to the X86 Backend", LLVM 18.1 release notes), and the Clang change that gave `fp128` the same treatment describes it as "identical to `i128`" and "the same as GCC": passed on the stack by reference and returned in `xmm0` ([llvm/llvm-project#115052][115052]). The corpus verifies each row against a clang `-O0` probe (`test/link/cases`). The mach side of a 128-bit return on Win64 goes through a 16-byte frame slot: the two lanes are stored and `XMM0` is loaded from it (`movups`), because no x86-64 move carries 128 bits between the integer and the vector bank. The caller reverses it. That is the same price a 16-byte vector pays on Win64. [115052]: https://github.com/llvm/llvm-project/pull/115052 #### `f16` and `_Float16` `f16` is C's `_Float16` at every boundary, and a record or array holding `f16` is the C record or array holding `_Float16`. Every convention mach has passes it as the float it is, two bytes wide, the way GCC and clang do: ```mach # _Float16 scale(_Float16 x, _Float16 k); ext fun scale(x: f16, k: f16) f16; pub rec Vertex { pos: [3]f16; u: f16; # a C struct of four _Float16 } ``` | Convention | Argument | Return | As a record member | |---|---|---|---| | System V x86-64 | SSE class: the low 16 bits of the next xmm register, the stack past xmm7 | `xmm0` | a float in its eightbyte, so an eightbyte holding only floats is SSE | | Win64 | the xmm register of its positional slot, the stack past the fourth slot | `xmm0` | no change: a record of 1, 2, 4 or 8 bytes rides its integer slot | | AAPCS64 | the next `h` register, the stack past `v7` | `h0` | a two-byte member of a homogeneous float aggregate, so up to four `f16` ride `h` registers | | RISC-V `lp64d`, `lp64f`, `ilp32d`, `ilp32f` | the next `f` register, NaN-boxed to its width, and by the integer rule once they run out | `fa0` | a float leaf of the two-leaf rule | | RISC-V `lp64`, `ilp32` | the integer rule, as a `u16` | `a0` | an integer leaf | The float registers above the 16 bits belong to nobody on System V, Win64 and AAPCS64: mach writes zeros there and reads only the low 16 bits. The RISC-V psABI boxes a float narrower than its register, so mach sets those bits to ones, as clang does. The rule depends on the ABI, not on Zfh: without Zfh the value still rides an `f` register. The Win64 row is clang's. Mach's windows target is the Microsoft x64 convention, and MSVC has no half type, so clang is the only compiler for that convention that has one. clang places `_Float16` the same way under `x86_64-pc-windows-msvc` and `x86_64-w64-windows-gnu`. mingw-w64 GCC does not. It places a scalar `_Float16` as it would a 2-byte integer: in the integer register of its positional slot, and returned in `AX`. GCC's `function_arg_ms_64` sends only `float` and `double` to an xmm register. So a scalar `f16` crossing to or from code GCC built for Windows disagrees with mach. On that side, declare the parameter or result as a `float` whose low 16 bits are the `f16`, which GCC places where clang places the `_Float16`. `f16` record members are unaffected, since a record of 1, 2, 4 or 8 bytes rides its integer slot under both compilers. SPIR-V has no C boundary: a SPIR-V function takes and returns an `f16` as a value like any other. No convention refuses `f16` in a declared parameter or result, since each of them has a C type to match. The one refusal is the C-variadic tail, where an `f16` is held to the rule for a `float`: the callee reads a `double`, and mach inserts no implicit conversion, so the argument is refused with the cast to write, `::f64` (see [Arguments in the variadic tail](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#arguments-in-the-variadic-tail)): ```mach error is promoted to `f64` in a C-variadic argument ext fun printf(fmt: *u8, ...) i32; fun show(h: f16) i32 { ret printf("%f\n"::*u8, h); # error: `f16` is promoted to `f64`; write the cast } ``` A vector of `f16` lanes is a vector like any other at the boundary and follows [Vector arguments and the C ABI](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#vector-arguments-and-the-c-abi) for its size: on AAPCS64 an `f16x8` is arm_neon.h's `float16x8_t` and an `f16x4` its `float16x4_t`, each in the next `v` register. #### Symbol name The declaration names a C declaration, so its linker symbol is whatever the target's C toolchain emits for that identifier. On Linux and Windows that is the name as written; on Darwin the C ABI prefixes an underscore, so `ext fun mad_open` binds the `_mad_open` a C object there defines. You write the C name either way — the prefix is never spelled in source. Override the symbol outright with the `symbol` decorator: ```mach #[symbol("write")] pub ext fun libc_write(fd: i64, buf: *u8, n: i64) i64; ``` An overridden name is taken **literally**: it is the object symbol, so no platform prefix is applied to it and you spell the exact name the object carries (`_write` on Darwin, `write` elsewhere). That is what lets inline asm and the runtime entry points name a symbol exactly. Common reasons to rename: - The C name (e.g. `write`) would shadow other things in your namespace. - The target's calling convention or ABI decorates symbol names and you need to match the decorated form. #### Library attribution On a two-level dynamic format — PE (Windows) or Mach-O (Darwin) — the `library` decorator pins an `ext` import to the dependency that exports it: ```mach #[library("ws2_32.dll")] ext fun WSAStartup(ver: u16, data: *u8) i32; ``` - The value names a link dependency's stable logical identity. A `[link.X]` requirement supplies it through `library = "..."`, defaulting to `X`; a bare command-line `-l name` uses `name`. The dependency's exact canonical loader name (`libfoo.so.3`, an `LC_ID_DYLIB` install name, or `foo.dll`) is also accepted for compatibility. The named library **must** be among the link's dependencies; pinning to one that is not is a hard link error (`import '' pinned to library '' not among the link's dependencies`), never a silent fallback. A logical identity that equals a different dependency's loader name is rejected as ambiguous. - An `ext` import with no `library` is unattributed. PE and Mach-O require every dynamic import to identify its provider, so an unattributed import is a hard link error on those targets. - A symbol with **no `ext` declaration at all** — one a linked static archive leaves undefined — has nothing to decorate. Its provider is named from the manifest instead, by the `symbols` key on the `[link.X]` entry that supplies it; see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#linkname--link-requirements). The two declarations write the same attribution, so a symbol may be claimed only once. A claim is spelled like the source name, not like the object symbol: both routes key on the link name, and Mach applies the target's C symbol prefix to the manifest name just as it does to a declaration's. - "Only once" is checked across every declaration, and two `#[library]` decorators are no exception. Two `ext` declarations that resolve to the same link name but name different libraries are a hard error identifying both (`import '' is attributed to library '' by declaration '' in module '' and to library '' by declaration ...`), so which library the import table names never depends on module load order. Two declarations naming the *same* library are accepted: they name one provider, which is what a shared binding declared in two modules does. - Declarations under **different arms of one `$if` chain** are exempt, because no target selects both arms, so they never both apply. A target-conditional attribution is written the obvious way and is not a conflict. `library` composes with `symbol`: the rename sets the imported symbol's name, `library` sets the dependency it is imported from. On one decl the import is emitted under the renamed symbol within the named library. ```mach # imported as `socket` from ws2_32.dll, called as `ws2_socket` in Mach #[library("ws2_32.dll")] #[symbol("socket")] ext fun ws2_socket(af: i32, kind: i32, proto: i32) i64; ``` A linked Windows COFF object compiled with C/C++ `dllimport` commonly refers to `__imp_X`, the address of X's Import Address Table cell, instead of calling X directly. The PE linker applies X's normal `#[library]` or manifest `symbols` attribution, emits the undecorated loader import X, and resolves the object reference to that IAT cell. If another object also refers to X directly, both spellings share one import and one IAT entry; do not declare or map `__imp_X` as a separate export. `__imp_X` names that IAT cell, which the import table synthesizes — it is never itself an export, so no library provides a symbol under that name. An import whose loader-facing name still carries the prefix is a hard link error naming the export it denotes, not something the image is allowed to carry: such an image links clean and fails only when Windows loads it, with a missing-entry-point error. Attribute and import the undecorated `X`; every `__imp_X` reference resolves to its address cell from there. On a format whose loader resolves imports by global search (ELF), an import is not bound to a single dependency, so `library` has no effect on the emitted binary — the value is still validated against the link's dependencies, but every dependency is searched at load time regardless. #### Vector arguments and the C ABI A 128-bit vector (`f32x4`, `i32x4`, …) follows the target's **C** vector convention, which is the only one mach has, so a call is bit-compatible with a C `__m128` parameter: - **x86_64-windows (Microsoft x64):** the caller passes each vector argument **by reference** — it stores the vector to a 16-byte-aligned temporary and passes that temporary's address in the parameter's integer register (RCX/RDX/R8/R9) or, once those are exhausted, on the stack. A variadic vector argument is passed the same way. Vector **returns** ride XMM0, which both conventions already agree on. - **Every other target** (System V, AAPCS64, RISC-V lp64d): a vector rides a vector register, which is what the convention says and what a boundary-free call would do anyway. A vector **wider** than the target's vector register (`i32x8`, `f32x16`, …) follows the C rule for a vector of that size on each convention, so a call is bit-compatible with a C `int __attribute__((vector_size(32)))` parameter: - **System V x86-64** (`linux` / `darwin`): the psABI makes a vector wider than the vector registers the target declares **MEMORY** class. As an argument its bytes are placed on the stack at its C alignment, its size (32 for `i32x8`, 64 for `i32x16`), and it takes no register. As a result it is returned through the hidden result pointer, as gcc does. clang (Apple clang included) returns such a vector in `xmm0` and `xmm1` instead, against the psABI, so a C function built by clang that returns one disagrees with mach; declare the C side's result as a record of the same size, which clang returns as MEMORY. Apple clang also places such a vector *argument* on the stack at 16 bytes rather than at its C alignment, so it disagrees with mach wherever the vector's stack offset is not already a multiple of its size; on darwin declare the C side's argument as a record of the same size with `__attribute__((aligned(32)))` (or the vector's size), which Apple clang places as the psABI places the vector. Mach follows the psABI on darwin as on linux. - **System V x86-64 with `avx`** (`x86-64-v3` and up): the vector register is `ymm`, so a 32-byte vector (`i32x8`, `f32x8`, `f64x4`) is SSE class and rides one `ymm` register, as an argument and as a result, which is what gcc and clang do at `-mavx`. Once the vector registers are used up it goes on the stack at its C alignment. A 64-byte vector is still wider than the register and MEMORY class, as above. A vector whose lanes compute at 256 bits there (every lane under `avx2`, the float lanes under `avx` alone) travels whole in its `ymm` register. An integer vector under `avx` without `avx2` computes on 128-bit halves, so it moves between `ymm` and its memory image with `vmovdqu`, and the function runs `vzeroupper` once the value has left the register (after storing incoming `ymm` arguments, and after storing a returned `ymm` value). A function that holds a value in a `ymm` register runs `vzeroupper` before each call and return that passes none in one, so the 128-bit instructions outside it do not pay for a dirty upper half. A call that places an argument above the 16-byte stack alignment realigns the caller's stack pointer to that argument's alignment, as the psABI requires of the argument area. - **x86_64-windows (Microsoft x64):** the carrier table under [Windows vector carriers](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#windows-vector-carriers): a pointer to a caller copy, and caller-provided result storage. - **AAPCS64 and RISC-V lp64d:** C treats the vector as a composite larger than two registers: a pointer to a caller copy for an argument, and the indirect-result register for a return. Mach does the same. The width that decides "wider" is the vector register the target declares under its selected extensions: 16 bytes on every target mach has, and 32 on System V x86-64 once the target selects `avx`. Microsoft x64 passes a 32-byte vector by reference whatever the extensions, so `avx` changes nothing there. A size no C vector type has has no C rule to follow: C's `vector_size` takes a power of two, so `i32x5` (20 bytes) and `f32x6` (24 bytes) have no C counterpart. Mach keeps its own placement for such a vector on every convention: a pointer to its storage for an argument, and the indirect-result pointer for a return. Every call edge marshals the same way, because the C convention is the only one mach has: a direct call to a declared `ext fun`, an ordinary Mach→Mach call, and a call through a function *pointer* all place a vector where the target's C convention puts it. A mach function whose address reaches C is therefore callable from C whatever route it took. #### Linking external objects An `ext fun` is only a forward reference — its definition must be supplied at link time, either **statically** by an external precompiled object/archive or **dynamically** by a shared library bound at load time. Provide those inputs to `mach build` either on the command line or through the project manifest. An undefined `ext` symbol that no static object defines and no shared library can bind is a link error. ##### Command line C-toolchain-style flags, consumed by `mach build`: ```sh # explicit object / archive / shared-library path mach build . path/to/libfoo.o mach build . path/to/libfoo.a mach build . path/to/libfoo.so mach build . path/to/libfoo.dylib mach build . foo.dll # search dir + library name: static candidates first, then the target's .so/.dylib spelling mach build . -L build/libs -l foo # link against a system shared library (libc) dynamically mach build . -l c ``` - A bare argument that contains a `/`, ends in `.o` (object) or `.a` (archive), or names a `.so`, `.dylib`, or `.dll` is treated as an explicit input. A relative path is tried first against the working directory, then against the project root. - `-l ` resolves to an object, archive, or shared library: each `-L ` is searched for `lib.o`, `.o`, `lib.a`, then `.a`; finally the working directory is searched for the same four names (loose objects preferred over archives). Only if no static candidate exists does it fall back to the selected target's shared spelling: `lib.so[.N]` for ELF or `lib[.].dylib` for Mach-O. A resolved `@rpath/` dylib carries its selected directory into the executable as `LC_RPATH`. `-L` and `-l` may each be repeated. ##### Manifest Artifacts name typed `[link.X]` requirements. Use `system` for a named library, `framework` for a Darwin framework, or `local` for an object/archive/shared library path. Filters select the applicable target: ```toml [link.foo-unix] source = "system" name = "foo" library = "foo" os = ["linux", "darwin"] isa = "*" abi = "*" export = false [link.foo-win] source = "system" name = "foo.dll" library = "foo" os = "windows" isa = "*" abi = "*" export = false ``` Both entries expose the logical name `foo`, so `#[library("foo")]` is portable across the mutually exclusive target filters. Manifest and command-line inputs are both included; a name that cannot be resolved is a hard error. ##### Scope Loose `.o` relocatable objects and static `.a` archives are linked **statically**. An archive contributes only the members that define a symbol left undefined by the objects and members before it, selected to a fixed point within that archive, so a vendored archive costs the binary exactly the members it uses (`mach help build` states the resolution rules). A shared `.so`, `.dylib`, framework, or `.dll` is a **dynamic** dependency: its format-specific canonical loader name is recorded and undefined `ext` functions become run-time imports. A shared input is validated before it is recorded: a `.so` that is not a loadable ELF shared object for the selected architecture (a linker script, a foreign-architecture file) is refused (`'' is not a loadable ELF shared object for the selected architecture`). A static definition always wins over a same-named dynamic import. #### See also - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — regular function declarations - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) — `ext val` / `ext var`, the data analogue (foreign data imports) - [visibility.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/visibility.md) — `pub` and `ext` modifiers - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — `symbol`, `library`, and other codegen decorators - [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) — the comptime pack `...`, which is a Mach calling convention and not this one #### Windows vector carriers The `win64` ABI maps a vector's lane bytes to the following C carrier. This mapping applies to arguments, returns, and typed function pointers, including calls between Mach functions. It is independent of the selected C compiler. | Vector extent | C carrier | Arguments | Return | |---|---|---|---| | 2 bytes | `unsigned short` | integer register or stack slot | low 16 bits of `RAX` | | 4 bytes | `unsigned int` | integer register or stack slot | low 32 bits of `RAX` | | 8 bytes | `__m64` | integer register or stack slot | `RAX` | | 16 bytes | `__m128`, `__m128i`, or `__m128d` | pointer to a caller-owned, 16-byte-aligned copy | `XMM0` | | Any other extent `N` | `struct { unsigned char bytes[N]; }` | pointer to a caller-owned, 16-byte-aligned copy | caller-provided result storage | Include `` for `__m64` and `` for the 128-bit intrinsic carriers. The byte-array carrier has exactly `N` bytes and no trailing padding. The result-storage pointer for an aggregate return occupies the first integer argument position, shifting the other arguments by one position. The callee returns that pointer in `RAX`. Lane zero occupies the first bytes. Each lane keeps its little-endian integer or IEEE floating-point bit representation. Carrier conversion copies bits, without numeric conversion, lane widening, or padding between lanes. Signed and floating-point lanes therefore use the same transport as unsigned lanes of the same width. A C function can inspect or construct lane bytes with a union or `memcpy`. Only the vector's `N` bytes belong to the value, even when its argument temporary has additional alignment padding. For example, `i16x4` uses `__m64`, `f32x4` uses a 128-bit intrinsic carrier, and `f32x3` uses `struct { unsigned char bytes[12]; }`. The memory layout of an enclosing Mach record is a separate contract from this function-boundary carrier mapping. A generic C extension such as `short __attribute__((vector_size(8)))` is not a substitute for `__m64`: GCC and Clang assign that extension different Windows argument and return conventions. Declare the carrier above at the boundary. Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md ### `val` and `var` — bindings Bindings introduce named values. `val` is immutable; `var` is mutable. Both require an explicit type — Mach has no type inference. #### Grammar ```mach fragment val NAME: TYPE = EXPR; # immutable; initializer required var NAME: TYPE; # mutable; default-initialized var NAME: TYPE = EXPR; # mutable; explicit initializer ``` #### Examples ```mach val pi: f64 = 3.14159; val n: i64 = 42; var counter: i64 = 0; var buf: [256]u8; # default-initialized to zero fun bump() { counter = counter + 1; # var is reassignable } ``` ```mach error a `val` is immutable val n: i64 = 42; fun change() { n = 43; # ERROR: `n` is a val } ``` #### Immutability Assigning to a `val` is a compile-time error, at module scope and inside a function alike. So is writing a field or an element of one (`v.field = x`, `v[i] = x`): the store lands in the same storage the binding names. A pointer is the exception that proves the rule — `val p: *T` binds the **address** immutably, not the storage it addresses, so `@p = x`, `p.field = x` and `p[i] = x` write the pointee and are legal. The check ends where the address escapes. `?NAME` on a `val` yields a plain `*T` — there is deliberately no read-only pointer type — and writing through any pointer that reaches a `val`'s storage is **undefined behaviour**: a module-level `val` lives in read-only data, so the store typically faults, while a local one may silently appear to work. The compiler does not diagnose it. #### Scope `val` and `var` work at module top level and inside function bodies. At module top level: - `pub val NAME` exports the constant. - `pub var NAME` exports the variable as a writable global. Inside functions, they are local to the enclosing block. A module-level `val` of integer, `bool` or float type whose initializer is a compile-time constant **is** that value at every use site, in its own module and in every module that imports it — however the reference is spelled (bare, module-qualified, or through a `use` alias). No load is emitted, and arithmetic over it folds and strength-reduces like any other literal. Its storage is still emitted, so `?NAME` has an address, and a `val` whose initializer is not a compile-time constant keeps its load. A constant whose value **is** an address — a `str`, whose value is a pointer to its bytes — is never duplicated this way. It keeps one definition and every reference reads it, so two references to one `str` constant compare pointer-equal no matter which module they are in or how they are spelled. #### `ext` — foreign data imports `ext val` / `ext var` declares a binding whose storage lives in another object, imported by name — the data analogue of [`ext fun`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md). It is a forward reference the linker resolves, so it is **storage-less** and carries no initializer: ```mach ext var errno: i32; # imported mutable datum #[symbol("environ")] ext var env: **u8; # renamed import #[library("libfoo.so")] ext val foo_flags: u32; # library-pinned import ``` - No initializer. `ext val x: T = ...;` is an error — the definition, and its value, live in the providing object (mirrors `ext fun`'s absent body). - The symbol name (the target's C spelling of the identifier), the `symbol` and `library` decorators, and the static/dynamic linking inputs all work exactly as for [`ext fun`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md). - On a dynamic target the reference is emitted GOT-indirect so the loader binds it to the runtime definition. ELF uses a dynamic pointer relocation, `GLOB_DAT` on x86-64 and ARM64 or `R_RISCV_64` on RV64; Mach-O binds the `__GOT` slot through dyld, and on arm64 the object carries the reference as the `GOT_LOAD_PAGE21`/`GOT_LOAD_PAGEOFF12` pair. A cell another object of the same image defines is reached through a linker-owned slot instead. An ordinary cross-module reference to a `val`/`var` defined elsewhere in the same artifact stays directly addressed. Executed dynamic-import resolution is proven on the native ELF legs. #### `#[embed(...)]` — the other exemption to "requires an initializer" A `val` carrying `#[embed("path")]` also carries no initializer — the named file's content **is** the initializer, read at compile time. It is the second (and only other) exemption to `val`'s initializer requirement, alongside `ext` above; unlike `ext`, the binding still owns real storage, placed in read-only data. See [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#embedstr--compile-time-file-embedding) for the full rule set. #### No inference Every binding declares its type. An untyped numeric literal is checked against the binding's declared type; it does not participate in inferring that type. ```mach error bindings require an explicit type annotation val n: i64 = 42; # ok — 42 conforms to i64 val x = 42; # ERROR — a binding declares its type val y = 42i64; # ERROR — a suffix is not an annotation ``` The annotation is required whatever the initializer is: a typed suffix gives the literal a type, it does not give the binding one. Suffixes earn their keep where there is no annotation to read from, such as the elements of a pack tail. See [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md#typed-suffixes). #### See also - [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md) — numeric / string / char literal forms - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) — the type grammar for the annotation - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#embedstr--compile-time-file-embedding) — `#[embed]`, the other initializer exemption Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md ### `test` — test declaration A `test` declaration names a block of statements the test runner can execute on its own. Any module may declare them, and `mach test` collects every test in the selected artifact's closure and links the selected ones into one small test binary. What earns a test, and where it sits, is set by the [test policy](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#test-policy). #### Grammar ```mach fragment test { ... } ``` The name is an identifier, following the ordinary identifier rules. A string in its place (`test "label" { ... }`) is a compile error located at the string that names the identifier form. The body is a block of statements. A test takes no parameters and is not callable from ordinary code; it exists only for the runner to invoke. Tests live in their own namespace in each module. A test's name is never in scope in code, so `test str_len { ... }` and `fun str_len` in one module do not conflict, and no code can name, call or reference a test. Two tests with the same name in one module are a compile error located at the second. A test's **qualified name** is its module path, `#`, and its name: `std.types.string#str_len__empty`. `#` appears in no identifier or module path, so a qualified name never collides with another symbol. It is the test's symbol (hidden, like every other symbol that is not exported), and it is the name `--list`, `--filter`, the readout and `--format json` show. A debugger takes it unquoted: `break std.types.string#str_len__empty` in gdb. Related tests group under a common subject as `subject__case` (`str_len__empty`, `str_len__multibyte`), and a regression test is named `regression__*`. Both are conventions the compiler does not check. `test` is a reserved keyword and appears at module (declaration) scope, the same level as `fun`, `rec`, and `val`. Visibility modifiers such as `pub` are syntactically accepted before `test` but carry no meaning — a test is never part of a module's public surface. #### Examples ```mach use std.runtime; use std.types.bool.bool; fun is_leap_year(y: i64) bool { if (y % 400 == 0) { ret 1; } if (y % 100 == 0) { ret 0; } ret y % 4 == 0; } fun debug(msg: *u8) { if (msg == nil) { ret; } } fun info(msg: *u8) { if (msg == nil) { ret; } } test is_leap_year__centuries { if (!is_leap_year(2000)) { ret 1; } if (is_leap_year(1900)) { ret 1; } if (is_leap_year(2023)) { ret 1; } ret 0; } test log__nil_message_does_not_crash { debug(nil); info(nil); ret 0; } ``` A test body may use anything in scope in the enclosing module, just like a function body. #### Semantics Every build resolves and type-checks each `test` body in the modules its artifact entry reaches. `mach test` builds exactly what `mach build` builds for the same artifact: the closure its entry reaches through `use` and `fwd`. A module no selected artifact reaches is not loaded under test either, so it is not checked and its tests do not run (see [Which tests run](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#which-tests-run)). Each test then lowers to a zero-parameter, `i32`-returning function tagged as a test entry point so the runner can iterate it. Its qualified name is the lowered function's name. Ordinary builds omit test bodies from IR and object files. A helper or fixture that exists only for tests is marked [`#[testing]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#testing--test-only-declaration). It gets the same treatment: checked in every build and omitted from ordinary ones. Only a test body or another `#[testing]` declaration may reference it. The body is checked against an `i32` return type. A test reports its result through that return value, treated as a process-style status in the range `0..255`: - `ret 0` — pass. - any `ret N` with `N` in `1..255` — fail, reported as `(exit N)`. - falling off the end of the body returns `0` (the default terminator for a non-void function is a zero return), so a body that never returns explicitly is treated as a pass. The result is the test process's exit status, and a process exit status is eight bits wide on every host (`mach test` reads the same eight bits on windows). The range is therefore part of the protocol, enforced at both ends: - A `ret` in a test body whose value is a literal (or a literal-shaped expression: a negated literal, or an arithmetic expression over literals) outside `0..255` is a compile error located at the `ret`, naming the value: `test result 256 is outside the status range 0..255`. An ordinary function returning the same value is unaffected; only test bodies carry the range. - A result computed at run time that lands outside `0..255` — a bit mask that has grown past eight bits, a negative code — is folded by the dispatcher to `255` before the process exits. It is reported as `(exit 255)` and is always a failure. The low eight bits are never used on their own, so a result of `256` cannot read as a pass. Within the range the compiler attaches no special pass/fail meaning to particular non-zero codes, nor does it provide built-in assertion intrinsics. A test signals failure by returning non-zero — typically by returning early from a failed check, as in the example above. A test that accumulates a bit mask must keep it within eight bits; past that, return the ordinal of the first failing check instead. ##### Collection across modules Tests are not tied to a single file. Every `test` declaration in every module of the artifact under test that belongs to the current project is collected, and `mach test` links one dispatcher executable in place of the artifact's normal entry and runs each test through it in its own process. By default collection is scoped to the current project's own modules: tests declared in dependency modules are excluded, so a library's own suite never runs (or fails) as part of your project's `mach test`. Pass `--include-deps` to collect dependency tests as well — useful when working on a dependency in-tree. `--filter` narrows the run by qualified name in either mode. #### The `mach test` workflow `mach test ` is `mach build` with a different goal: it builds the artifact's objects exactly as `mach build` does, adds a test object beside each module that declares tests, links one test **dispatcher** executable covering the selected tests, then runs each of them as its own process (` `), captures its output, times it, and renders a per-module readout — collapsing all-passing modules to a single roll-up line and expanding any module with a failure to show the failing test's captured output and location. The full flag reference is `mach help test`; the options that select and shape a run are: ``` --jobs run up to n test processes at once (default: host CPUs) --filter select only tests whose qualified name contains the substring --include-deps also run tests declared in dependency modules --list list the collected tests and exit --format the live readout, or an NDJSON event stream --runner launch each test through a host-side command --timeout terminate a test and its process group after the duration ``` A roll-up is ` ok[ FAIL] `. Each expanded failure shows `file:line`, the exit code (`(exit N)`), signal (`(signal N)`) or `(timed out after )`, the child's captured output indented beneath, and the exact `rerun:` command; a passing test stays quiet. The run closes with a summary that re-lists every failure: ``` failures: app.main#fails_on_purpose src/main.mach:11 (exit 3) 1 passed, 1 failed, 2 total (1ms) ``` The exit code of `mach test`: - `0` — every test that ran passed. - `1` — at least one test failed, was killed by a signal, or timed out; also a user error such as an unknown flag. - `2` — a build or internal error before the tests could run, or a test that failed for an infrastructure reason (the harness, not the test). - `3` — an environment failure before the tests could run: a file, directory or process operation the machine refused. These are the codes every `mach` command shares (`mach help ` lists them under `exit:`); `test` only adds what `1` also means. `--list` prints each selected test's qualified name and the test object that holds it, and exits without linking or running anything: ``` app.parser#rejects_trailing_comma ./out/linux-x86_64/debug/obj/app/parser.test.o ``` `--filter ` selects before the dispatcher links, so the dispatcher holds only the selected tests and what they reach, and changing the filter relinks without recompiling. `--emit` is rejected under `mach test` (`--emit is not applicable to 'test'; test always builds its internal test dispatcher`). ##### Timeouts `--timeout ` bounds each spawned test process independently, from its own spawn, on its whole process group, so a process the test started dies with it. A test that exceeds the bound is the distinct outcome **timed out**: it renders as `(timed out after )`, is counted separately on the summary line, and is still a failing test for the exit code, so the suite exits `1`. ``` failures: app.main#spins src/main.mach:4 (timed out after 1s) 0 passed, 1 failed (1 timed out), 1 total (1.0s) ``` `` is a positive integer followed by a unit: `ms`, `s`, `m` or `h` (`30ms`, `30s`, `5m`, `1h`). A bare number, a fraction, zero or any other unit is a usage error naming the accepted forms. There is no default: omitting the flag leaves every test unbounded. ##### JSON output `--format json` replaces the readout with one JSON object per line on stdout (`run_start`, one `test` per result, `summary`; `case` under `--list`), with build diagnostics kept on stderr, where `--diagnostics=json` writes them, a `test` record per result and a closing `summary` as the records of [diagnostics-json.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md). A `test` or `case` event names its test by qualified name in `name`, beside its `module`, `file`, `line`, test `object` and dispatcher `index`. A timed-out test reports `"kind":"timeout"` with its bound in nanoseconds in `timeout_ns`. The schema is versioned (`"schema":2` on every event) and its writer is `mach.cli.cmd.testing`. #### The runner A test build compiles every module's object exactly as `mach build` does and shares it: after `mach build`, `mach test` recompiles no module object. A module that declares tests or `#[testing]` declarations also gets a **test object**, `obj//.test.o`. The compiler (`mach.lang.me.lower.testrunner`) lowers each of its tests to a zero-parameter, `i32`-returning function under its qualified name, a symbol that never collides with, reserves, or rewrites a user symbol. The test object references every symbol the module's object defines, private ones included, so a test reads and writes the same globals and calls the same functions the module's own code does, and every symbol has exactly one definition. It defines only what the module's object lacks: the tests, the `#[testing]` declarations, a private function inlined everywhere or called only from tests, a private global only tests use, and generic instances only tests use. The module's object never changes for tests, and its cache key leaves the module's test declarations out, so editing a test recompiles only that module's test object and relinks. The test object's key is the module object's key and the module's whole source. Each run synthesizes one dispatcher object, `test//dispatch.o`, whose entry selects a test by its index argument, and links it with the test objects and the module objects into `test//`, even for a library artifact. The link keeps only what the selected tests reach. The dispatcher is the program's `main`: an artifact's own `main` yields to it in the link, so its object is linked unchanged. The dispatcher's entry calls the selected test and exits with its result: as is when the result is in `0..255`, and `255` otherwise (see [Semantics](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#semantics)). A missing, malformed, or out-of-range index exits `2`. `mach test` then keeps up to `--jobs` children in flight, each spawned as ` `, captures each child's stdout and stderr to a per-test file under `log/` beside the dispatcher, `test//log/` unless `-o` moves the dispatcher, named by the test's position in the run (a passing test's file is removed on the spot, a failing test's file stays), and reads its exit status. Results render in collection order regardless of completion order. #### Which tests run `mach test` selects its cells with `-a`, `-t` and `-p` as every command does (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#selection-and-the-build-matrix)): each takes an exact name or a glob and repeats, and `--all` fills every axis no option names with `*`. With no `-a`, the sole artifact the selected target builds is chosen, or among several the one marked `default = true` (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifactname)). Each selected artifact's closure, the same module set `mach build` compiles for it, is built into a test dispatcher for each selected target and profile, so `$bin.name` in a test block, and in every module a dispatcher compiles, is the artifact under test. Tests run once per (target, profile). When several artifacts are selected there, the tests they reach are combined and each qualified name runs once. Only a target whose `os` and `isa` are the host's runs its tests; any other is built and reported on a `skip` line, since an emulator the host happens to have is not the target. `--runner ` launches each test as ` ` and needs the selection to resolve to one (artifact, target, profile), foreign or not. A run that executed nothing exits `1`, naming the host, so a passing run always ran tests. An inline `test name { }` declaration in a module the artifact reaches runs with no further wiring. A module that exists only for tests, such as a suite that exercises several modules together, is reached by no artifact and so never runs on its own. Give such modules a **test artifact**: an ordinary library artifact whose entry `use`s each of them, tested by name. ```toml [artifact.app] kind = "bin" default = true entry = "main.mach" out = "bin/app" targets = ["*"] link = [] need = [] # the test-only modules no other artifact reaches [artifact.tests] kind = "static" entry = "test/all.mach" out = "lib/tests" targets = ["*"] link = [] need = [] ``` ```mach fragment # src/test/all.mach use std.runtime; use app.test.parser; use app.test.roundtrip; ``` `mach test .` then runs the tests `app` reaches and `mach test . -a tests` the test-only suites. The entry reaches the runtime's startup (`use std.runtime;`) because a library artifact's closure is all the test dispatcher links. Marking `app` `default = true` keeps `mach build .` and `mach check .` to `app`, since with no selector they take the marked artifacts; `mach build . -a tests` builds the test artifact. #### Test policy A change does not need a test of its own. A unit test exists only if it: - covers a unique surface, duplicating no other test - covers functionality critical to correctness that cannot be allowed to break - is deterministic, never depending on timing or performance - is valuable to check automatically - covers logic that is not blatantly simple - checks correctness - is no more complicated than the code it tests, unless that is unavoidable - does not pin a problem that no longer exists Coverage means branches and known failure points, not volume. `str_len` gets the inputs that exercise each of its branches, not a pile of strings, and a parser's tests cover its surface concisely, not exhaustively. Inline tests are small and sit in their module for convenience or because they need private access. A test lives outside the module it covers to declutter it, because the test is significant, or, most often, because it exercises several modules together (see [Which tests run](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#which-tests-run)). Regression tests are a separate kind, and rare: they are kept only for regressions that are easy to reintroduce, and are named `regression__*`. They may sit next to unit tests. The name is a convention, not a mechanism. The compiler's codegen corpus and link cases get the same scrutiny, scoped to what cannot be tested inside the compiler: the final codegen and link result on disk. #### See also - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — functions; a test body is checked like a function body - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#testing--test-only-declaration) — `#[testing]` for declarations that exist only for tests - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) — `if`/`or`, `ret`, and the other statements a test body uses - [files.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/files.md) — project layout the build (and `mach test`) discovers Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md ### Variadic packs A **variadic pack** (`va: ...`) is a trailing function parameter that collects a variable number of call-site arguments into a compile-time sequence. The compiler monomorphizes the function once per distinct argument type-list at each call site; there is no runtime structure, no `va_list`, and no `any`. This is unrelated to C's varargs, which share only the spelling. A pack is a *named* parameter (`va: ...`) on a mach function; the C form is a *bare* trailing `...` and may be declared only on an `ext fun`, where it describes a foreign callee's ABI — see [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#c-variadic-imports). #### Declaring a pack parameter A pack parameter is written as a named parameter whose type is `...`: ```mach fragment fun name(va: ...) RetType { ... } fun name(fixed: T, va: ...) RetType { ... } # leading fixed params are allowed ``` The pack parameter must be last. Any number of fixed (typed) parameters may precede it. #### Iterating with `$each` `$each a in va` unrolls the body once per element, with `a` bound to the element and re-typed to that element's concrete type per instantiation. This is the only way to consume a pack. ```mach use std.print; use std.runtime; fun sum(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a; } ret t; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {} {}", sum(1, 2, 3), # 6 sum(10, 20), # 30 sum()); # 0 — empty pack, body never runs ret 0; } ``` Because each element has its own concrete type at monomorphization, the body can handle heterogeneous packs: ```mach use std.print; use std.runtime; fun sumc(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a::i64; # cast each element's concrete type to i64 } ret t; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", sumc(1000, 50::i32, 7::u8, 200::i16)); # 1257 ret 0; } ``` A `$each` body is a normal statement block; it may nest arbitrarily, call functions, read outer-scope runtime variables, and write them back. Runtime variables in the enclosing scope (e.g. a cursor) thread across all unrolled iterations — each iteration reads where the previous one left off. #### `va.len` — element count `va.len` folds to the instance's element count at compile time. ```mach use std.print; use std.runtime; fun count(va: ...) i64 { ret va.len::i64; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {}", count(1, 2, 3), count()); # 3 0 ret 0; } ``` #### `va...` — forwarding a whole pack Inside a pack instance, `g(va...)` forwards the whole pack to another pack-tailed function, which is monomorphized for the forwarded type-list. ```mach use std.print; use std.runtime; fun sum(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + a; } ret t; } fun outer(va: ...) i64 { ret sum(va...); } fun fwdpre(base: i64, va: ...) i64 { ret base + sum(va...); } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {}", outer(1, 2, 3), # forwards (1, 2, 3) to sum → 6 fwdpre(100, 1, 2, 3)); # 106 ret 0; } ``` Leading fixed arguments may precede the spread at the call site: ```mach fragment fun fwdpre(base: i64, va: ...) i64 { ret base + sum(va...); } ``` `va...` is valid **only as the sole trailing argument of a pack-tailed callee**. The compiler rejects: - Spreading into a callee with no pack parameter. - Spreading when other arguments follow the spread. - Spreading only part of a pack (partial forward is not supported). #### Monomorphization and ABI A pack-tailed function is compiled once per distinct argument type-list at each call site. Different arities, or the same arity with different types, produce separate instances: ```mach fragment sum(1, 2, 3) # instance: (i64, i64, i64) sum(10, 20) # instance: (i64, i64) sum(5::u8, 1::u32) # instance: (u8, u32) ``` A pack tail composes with generic parameters. The instance key is then the **product** of the type arguments and the element type-list, so `sink[u32](x)` and `sink[^u32](x)` over the same elements are two instances with two bodies, each typed and lowered against its own type arguments: ```mach fragment fun sink[T](t: T, va: ...) u32 { ... } sink[u32](p, 1u32) # instance: [u32] over (u32) sink[u64](p, 1u32) # instance: [u64] over (u32) sink[u32](p, 1u32, 2u64) # instance: [u32] over (u32, u64) ``` The instance's typing covers the **whole body**, not just the `$each`. A `val` / `var` declared around the unroll and named inside it carries the instance type, so it is gated, checked, and folded exactly as if the concrete type had been written out: ```mach fun mk[T](seed: T, va: ...) T { var out: T = seed; # `Box` at mk[Box], not the bare `T` $each a in va { ret out; } ret out; } ``` Because a pack-tailed function has no single entry point, it is **not a stable-ABI symbol** — it cannot be the target of `ext fun` or a function pointer shared across compilation units. It is source-level only. A `^` secret may not be passed to a pack, including a secret nested inside an aggregate argument. A pack instance is re-inferred per call site and commonly feeds formatting or logging, so admitting one would launder a secret straight to an observable sink — see [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md). #### Cross-module packs Pack-tailed functions may call functions in other modules from inside the `$each` body. The compiler re-infers the body per monomorphization instance against the full module set. ```mach fragment fun sumdbl(va: ...) i64 { var t: i64 = 0; $each a in va { t = t + helper.dbl(a); # helper.dbl resolved per element type } ret t; } ``` #### Real-world example: format The standard library's `vformat` is a pack-tailed function. Each `$each` iteration handles one format argument in order, with `$type_of` dispatch selecting the right writer per element type. A reduced version of the same shape, answering `res[usize, FormatError]` the way std's does: ```mach use std.print; use std.runtime; use std.types.result.res; use std.types.size.usize; use std.types.string.str; tag FormatError: u8 { few_holes; } fun write_byte(b: u8) { var one: [2]u8 = [2]u8{b, 0}; print.print(?one[0]); } fun write_str(s: str) { print.print(s); } fun write_i64(n: i64) { print.printf("{}", n); } fun vformat(fmt: str, va: ...) res[usize, FormatError] { var i: usize = 0; $each arg in va { for (fmt[i] != 0 && fmt[i] != '{') { write_byte(fmt[i]); i = i + 1; } if (fmt[i] != '{') { ret res[usize, FormatError].err{FormatError.few_holes{}}; } i = i + 2; $if ($type_of(arg) == str) { write_str(arg); } $or ($type_of(arg) == i64) { write_i64(arg); } $or { $error("no writer for this argument type"); } } for (fmt[i] != 0) { write_byte(fmt[i]); i = i + 1; } ret res[usize, FormatError].ok{i}; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val hi: str = "hi"; val r: res[usize, FormatError] = vformat("n={} s={}\n", 7, hi); if (sel r.err) { ret 1; } ret 0; } ``` The runtime format cursor `i` threads across all unrolled iterations; the arg sequence is consumed entirely at compile time. #### See also - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — function declarations, generic and comptime parameters - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) — `$type_of`, `$fields`, `$each` - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` / `$or` used inside pack bodies - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#c-variadic-imports) — C varargs, the unrelated `ext`-only form Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md ### Literals #### Numeric | Form | Example | |---|---| | Decimal | `42` | | Hex | `0xDEAD` | | Binary | `0b1010` | | Octal | `0o755` | | Underscores | `1_000_000` | | Scientific | `1.5e10` | | Typed suffix | `7i64`, `255u8`, `2.5f64` | Numeric literals without a suffix are untyped until they flow into a binding or expression context that constrains their type. With no such context an integer literal is `i64` and a float literal is `f64`, and an integer literal that takes that default must fit `i64`: one that does not, such as `10000000000000000000::f64`, is refused, and a suffix (`10000000000000000000u64`) or a typed context gives it the type it needs. ##### Typed suffixes A suffix is the spelling of the primitive type the literal has: `u8`, `u16`, `u32`, `u64`, `u128`, `i8`, `i16`, `i32`, `i64`, `i128` on an integer literal, and `f16`, `f32` or `f64` on a float literal. It works with every radix and with digit separators (`0xFFu8`, `0b1010u16`, `1_000_000u32`, `1.5e2f32`, `0.25f16`). A suffixed literal *is* that type, in the same way a variable of that type is. It is not a hint, and it is never silently retyped: ```mach error type mismatch: expected u64, found u32 val a: u32 = 7u32; # fine val b: u64 = 7u32; # error: type mismatch, expected u64, found u32 val c: f64 = 1.5f32; # error: type mismatch, expected f64, found f32 val d: u32 = 7u32 + 1; # fine: the unsuffixed 1 takes u32 from the other operand val e: u32 = 7u32 + 8u64; # error: operands are different types ``` Because a suffixed literal is already typed, it is what types the elements of a pack tail, where nothing else constrains them: ```mach fun sink(va: ...) u32 { var acc: u32 = 0; $each a in va { acc = acc + a; } ret acc; } fun total() u32 { ret sink(7u32, 8u32); } ``` A literal outside the range of the type its suffix declares is rejected, with the range named. A leading `-` is part of the range check, so the most negative value of a signed type is written the way it reads: ```mach error literal 128 is out of range for i8 val a: i8 = -128i8; # fine val b: i8 = 128i8; # error: literal 128 is out of range for i8 (-128..127) val c: u8 = 256u8; # error: literal 256 is out of range for u8 (0..255) val d: u32 = -1u32; # error: literal -1 is out of range for u32 (0..4294967295) ``` An integer literal may be as large as `u128` holds, in any radix and with separators, and the range check reaches the same width; a literal that no integer type holds is an error: ```mach fragment val a: u128 = 340282366920938463463374607431768211455; # u128 max val b: u128 = 0xFFFF_FFFF_FFFF_FFFF_FFFF_FFFF_FFFF_FFFF; # the same val c: i128 = -170141183460469231731687303715884105728; # i128 min val d: u128 = 1u128 << 100; # folded at comptime ``` A float suffix needs a float literal: the fractional part or the exponent is what makes it one, so `1f32` is an error and `1.0f32` and `1e0f32` are the ways to write it. Any other suffix spelling is an error naming the suffixes that exist. ##### Float values A float literal is rounded to its type once that type is known, from its suffix or from its context: to the nearest value the type holds, ties to even, the IEEE default. The decimal is rounded once, straight to that type. An inexact literal warns only when its written digits are not the value stored. The compiler rounds the literal, takes the shortest decimal that reads back as the result, and compares it with the literal's own significant digits, leading and trailing zeros and the exponent's spelling aside. `0.1` in `f32` is quiet, since `0.1` is how that `f32` value is written, while `3.14159265358979` in `f32` warns that the value stored is `3.1415927`. An exactly representable literal never warns, however many digits it has: `2147483648.0` is exact in `f32`. The warning is the named kind `float.inexact`. The rule is the same at every float width, and it matters most for `f16`, whose 11 significant bits hold about three decimal digits. `0.1` in `f16` is quiet, since it stores 0.0999755859375 and `0.1` is how that value is written, while `3.14159` in `f16` warns that the value stored is `3.14`, which is 3.140625. A literal that rounds to infinity in its type, or whose nonzero value rounds to zero, is an error: ```mach error float literal overflows `f32` val a: f32 = 3.4028235e38; # fine: the largest f32 val b: f64 = 1.0e39; # fine: f64 holds it val c: f32 = 1.0e39; # error: float literal overflows `f32`: it rounds to infinity val d: f32 = 1.0e-46; # error: float literal underflows `f32`: its nonzero value rounds to zero ``` `f16` holds much less, so its bounds come early: ```mach error float literal overflows `f16` val a: f16 = 65504.0; # fine: the largest f16 val b: f16 = 0.00000006; # fine: rounds to the smallest subnormal, 2^-24 val c: f16 = 70000.0; # error: float literal overflows `f16`: it rounds to infinity val d: f16 = 1.0e-8; # error: float literal underflows `f16`: its nonzero value rounds to zero ``` Every phase agrees on the type a suffix declares: compile-time evaluation, type checking, constant folding, and the emitted code all read the same type and the same value. #### Char A single character in single quotes, typed as `u8`: ```mach val c: u8 = 'M'; ``` Char escapes: `\n` `\t` `\r` `\\` `\'` `\"` `\0` `\xHH`. This set is deliberately minimal — there is no `\a` `\b` `\f` `\v`. Any other byte, control characters included, is written `\xHH` (e.g. `\x08` for backspace, `\x07` for bell). #### String A sequence of characters in double quotes, producing a `*u8` pointing at null-terminated bytes in the data segment: ```mach val msg: *u8 = "hello, mach\n"; ``` String escapes: the same set. A string literal is a single line. There is no multi-line string syntax — long content uses `\n` escapes or lives in external files. #### `nil` The `nil` keyword is the null-address literal. With no context it types as `*u8`; it coerces to any pointer-like type — a raw `ptr`, a typed `*T`, or a function type `fun(...)` (a function value is a code address, so the null address is a valid null callback). The coercion is uniform across every position a value flows into: globals, locals, record fields, array elements, call arguments, and return slots. ```mach def F: fun(u32); var p: *i64 = nil; # null pointer var cb: fun(u32) = nil; # null function pointer var k: F = nil::F; # the cast spelling works too fun absent() u8 { ret p == nil; # nil compares against any pointer-like value } ``` nil coerces only to pointer-like targets; assigning it to a non-pointer slot (`var x: u32 = nil;`) is a type error. #### Backticks The backtick (`` ` ``) is not a token: it is an unexpected character wherever it appears. See [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#unexpected-characters). #### See also - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) - what these literals are typed as - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) - record, array, union, and tag literals - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - tagged value construction - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) - using literals as binding initializers Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md ### Types Mach has a small set of compiler-shipped concrete types plus a uniform type-construction grammar for pointers, arrays, and function types. There are no compiler-known type aliases — names like `bool`, `usize`, and `str` are stdlib `def`s. #### Primitive scalars | Family | Members | |---|---| | Unsigned int | `u8`, `u16`, `u32`, `u64`, `u128` | | Signed int | `i8`, `i16`, `i32`, `i64`, `i128` | | Float | `f16`, `f32`, `f64` | | Untyped pointer | `ptr` | These fourteen names are the complete set of compiler-seeded primitive types. A **type** may not take one of these names. `rec`, `uni`, `tag`, `def` and a generic parameter named after a primitive are refused with `name.builtin_type`, because every use of the name in type position resolves to the primitive, so the declaration would be unreachable. The refusal holds for every primitive the compiler ships, so a primitive added later refuses an existing type of its name rather than silently changing what the name means: ```mach error is a built-in type pub def f16: u16; # error: `f16` is a built-in type ``` A value, function or field may take the name, since it never stands in type position: ```mach val f16: i64 = 7; # fine: values are a different position ``` ##### Half precision `f16` is IEEE 754 binary16: a sign bit, 5 exponent bits and 10 significand bits, with `$size_of` 2 and `$align_of` 2, which is what C's `_Float16` has on every ABI mach targets. Its largest finite value is 65504, its smallest normal 2^-14 and its smallest subnormal 2^-24. It is a float like the other two: `$is_float(f16)` holds, every float operator applies with the result correctly rounded to binary16, literals take the `f16` suffix ([literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md#typed-suffixes)), and compile-time evaluation computes in binary16 exactly, with the rounding, subnormals and overflow to infinity of the emitted code: ```mach use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val a: f16 = 1.5; val b: f16 = a * 2.0 + 0.25f16; # 3.25, rounded once to binary16 val max: f16 = 65504.0; # the largest finite f16 val inf: f16 = max + max; # overflows to infinity if (b:~u16 != 0x4280 || inf:~u16 != 0x7C00) { ret 1; } if ($size_of(f16) != 2 || $align_of(f16) != 2) { ret 2; } ret 0; } ``` Every target realizes `f16`, including one with no half-precision hardware: where the target has no instruction for an operation, the compiler emits it as inline integer and binary64 code, never a call to a runtime helper, so a freestanding link needs nothing extra. Which targets compute natively, and why the result is the same bits either way, is on [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#half-precision-arithmetic). The conversions are on [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#f16-conversions), the lane form `f16xN` under [SIMD vectors](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#simd-vectors) below, and the calling conventions on [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#f16-and-_float16). ##### 128-bit integers `u128` and `i128` are integers like the others: every arithmetic, bitwise, shift, comparison and conversion operator applies, literals take the `u128` and `i128` suffixes, comptime evaluates them exactly, and `$size_of` is 16 with `$align_of` 16, which is what C's `__int128` has on every ABI mach targets. No target has a 128-bit register. A 128-bit value is **realized** as two 64-bit lanes on the 64-bit targets (x86-64, aarch64, riscv64), the way a vector wider than the vector register is realized as pieces: addition and subtraction carry across the lanes, shifts move bits between them, and a comparison decides on the high lane first. Which targets realize the width is a declaration of the machine model, not a consequence of its name. A target that declares no 128-bit integer (riscv32, spirv) refuses the type where it is used, at the declaration or expression that names it: ```text error[target.int_width]: a 128-bit integer is not realized on riscv32: the target declares no integer of that width ``` Three operations have a shape worth knowing: - **Widening multiply.** `(a::u128) * (b::u128)` with `a` and `b` 64-bit is recognized as the full 64 x 64 product and compiles to the target's one widening instruction, never to a 128 x 128 multiply. Its high half, `(... >> 64)::u64`, is one high-multiply instruction (`mul` on x86-64, `umulh` on aarch64, `mulhu` on riscv64) and its low half `(...)::u64` is the plain multiply. The same holds for `i128` from `i64` operands, with the signed forms. A `u128 * u128` product that is not of that shape is the schoolbook over its lanes, three multiplies and the widening one. - **Division and remainder** at 128 bits have no instruction on any target and are calls into helpers the compiler provides with the program. - **Vectors** have lanes of 8 to 64 bits, so `u128x2` is not a vector type; it is refused with the lane rule named. The secrecy page states which 128-bit operations a secret may reach ([secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md)); the calling conventions for `u128` and `i128` across an `ext fun` boundary are on [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md). There is no compiler `bool`. `bool` is a stdlib `def bool: u8;` with `true` / `false` as stdlib `val`s (`1` / `0`) — see [def.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/def.md). #### SIMD vectors A vector type is a **form**, not a fixed list: any primitive numeric element followed by `x` and a lane count. On a 128-bit target the spellings this currently accepts are: ```mach fragment f16x8 f32x4 f64x2 # float lanes i8x16 i16x8 i32x4 i64x2 # signed integer lanes u8x16 u16x8 u32x4 u64x2 # unsigned integer lanes ``` The spelling is `x` with a **single** `x`. A name like `f32x4x4` is not a vector type (it resolves as an ordinary identifier); a matrix is an algorithm over vectors and belongs in a library over vector elements, not the language. A vector spelling is recognized only in **type position**, so a *value* may still be named `f32x4` without colliding with the type: ```mach val f32x4: i64 = 7; # fine: values are a different position ``` A **type** may not. `rec`, `uni`, `tag`, `def` and a generic parameter reject a name spelled as a vector form, because a type declared with a vector's name would be silently unreachable: every use in type position resolves to the vector instead: ```mach error is spelled as a vector type rec f32x3 { x: f32; } # error: `f32x3` is spelled as a vector type tag f32x4: u8 { empty; } # error: `f32x4` is spelled as a vector type ``` This holds for any well-formed spelling, so the name cannot be claimed by a type and then collide with the vector it spells. Two rules bound the form, and **neither depends on the target**: - **At least 2 lanes.** `f32x1` is refused; a one-lane vector is just its scalar. - **At most 65535 lanes.** A compiler limit, not a machine one: the lane count is carried as a 16-bit field through code generation. Everything else is legal at every width. `f32x3` is 96 bits, `f32x8` is 256, and both compile on every target — including one with no vector unit at all. **Width is a realization question, not a legality one.** How a shape is realized does depend on the target, and there are three answers: one packed instruction when the shape fits a vector register and the operation has a packed form; one packed instruction per register-width piece when the shape is wider than the register; and per-lane scalar code where the operation has no packed form for the lane shape, or on a target with no vector unit (rv64gc today). All three compute identical lanes — the expansion is a fixed unroll, never a reassociation — so only performance varies. The `simd` manifest lever (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md)) reports or refuses the scalar cases if a project cannot afford them. **A shape wider than the register is split into register-width pieces.** On x86-64 and aarch64, whose vector register is 128 bits, an `i32x8` is two `i32x4` pieces, each its own register, and one `+` is two packed adds. The last piece holds whatever lanes are left, so an `i32x9` is two `i32x4` pieces and a one-lane piece. A lane-wise operation (arithmetic, bitwise, compare, select, shift, a conversion that keeps its lanes in place) runs on each piece. One that moves lanes between pieces has its own lowering over the pieces: a range `v[start, count]` reads the pieces it crosses and joins them, a widening conversion or multiply writes each half of a source piece into its own result piece (`i16x8 → i32x8` is `pmullw`/`pmulhw` with `punpcklwd`/`punpckhwd` on x86-64, and `smull` with `smull2` on aarch64), and a narrowing one joins the narrowed lanes of its pieces. A reduction written over the lanes reads each lane from its piece. An operation with no packed form for the lane shape (a 64-bit lane multiply on x86-64) runs lane by lane on each piece, and under `simd = "scalarize"` it warns at its site while `simd = "require"` refuses it. The piece width is the target's **declared** vector width, not its name, so a target that gains wider vector registers gets **better code**, not new spellings: one that declares a 256-bit register holds an `i32x8` whole. x86-64 declares one per lane kind: with `avx` its float lanes (`f16`, `f32`, `f64`) compute at 256 bits, and with `avx2` (`x86-64-v3`) its integer lanes do too, so under `x86-64-v3` an `i32x8`, an `i16x16` or an `f64x4` is one `ymm` register and one VEX.256 instruction per operation, and an `i32x16` is two such pieces. SPIR-V, whose vectors are values rather than registers, splits nothing. A secret vector keeps its secrecy on every piece. `ptr` is not a lane element: its width is target-defined rather than a scalar bit count, so `ptrx2` is not a vector spelling. A lane is an integer of 8 to 64 bits, an `f16`, an `f32` or an `f64`. **`f16` lanes.** `f16xN` is a vector like any other: `+ - * /` and the comparisons apply lane-wise, a comparison gives a `u16xN` mask, `::` converts each lane with the scalar `f16` rule, and `:~` reads the bits (`f16x8:~u16x8`). Every lane is bit for bit what the scalar operation gives on the same target, whichever way the target realizes the shape: | target | `f16` lane arithmetic | `f16` lane comparisons | |---|---|---| | aarch64 with `fp16` | packed `.8h` and `.4h` instructions | packed `.8h` and `.4h` compares | | aarch64 without `fp16` | each lane's scalar `f16` operation | the lanes widened to `f32` (`fcvtl`) and compared there | | x86-64 with `f16c` (`x86-64-v3` and up) | four lanes at a time widened to `f32` lanes (`vcvtph2ps`), computed there and narrowed back (`vcvtps2ph`) | four lanes at a time widened to `f32` lanes and compared there | | x86-64 without `f16c` | each lane's scalar `f16` operation | each lane's scalar comparison | | riscv64, riscv32 | each lane's scalar `f16` operation (no vector unit) | each lane's scalar comparison | | spirv with `float16` | the core float instructions on an `f16` vector | the same | | spirv without `float16` | each lane's scalar `f16` operation | each lane's scalar comparison | A lane's scalar operation is the target's own half instruction where it has one and the inline expansion otherwise (see [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#half-precision-arithmetic)). On aarch64 the conversions between `f16` and `f32` lanes are packed (`fcvtl`, `fcvtn`) with or without `fp16`. A comparison makes no NaN and every widening is exact, so comparing in `f32` lanes gives the scalar mask. x86-64 has no packed `f16` arithmetic below AVX-512 FP16, and the AVX-512 FP16 rows are #4159. ##### Size and alignment `$size_of` is **lane-derived**: `lanes × element size`, packed, with no padding at any width. | type | `$size_of` | `$align_of` | |---|---|---| | `f16x4` | 8 | 2 | | `f16x8` | 16 | 16 | | `f32x2` | 8 | 4 | | `f32x3` | 12 | 4 | | `f32x4` | 16 | 16 | | `i16x4` | 8 | 2 | | `f32x5` | 20 | 16 | | `f32x8` | 32 | 16 (32 where the target's vector register is 256 bits) | `$align_of` has **two rungs, each for its own reason**. A vector *narrower* than the vector register is a packed aggregate that loads piecewise, so it aligns to a single lane. One that *fills* the register aligns to its whole size, because the machine's vector load requires it. One *wider* than the register is split into register-width pieces, each loaded and stored on its own, so it aligns to the register width — 16 for `f32x8`, not 32, because each piece is one 16-byte load and a larger alignment would serve none of them. The register is the widest the target's selected extensions give it. On x86-64 with `avx` it is the 32-byte `ymm`, and a vector aligns to the widest register it fills: 32 for `f32x8`, `i32x8` and every wider vector, 16 for `f32x4` and for a 20- or 24-byte vector, and a lane for one narrower than 16 bytes, as C aligns `__m256` and `__m128`. The first rung is what makes `[N]f32x3` a usable packed vertex buffer: padding `f32x3` to 16 bytes would make it indistinguishable from `f32x4` in memory. Arrays of vectors (`[4]f32x4`) and pointers to vectors (`*f32x4`) are ordinary composite types over a vector element. **Literals** are full-arity — one initializer per lane, mirroring array literals. The lane count must match exactly; too few or too many lanes is a compile error. ```mach fragment val v: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; val w: i32x4 = i32x4{1, 2, 3, 4}; ``` An uninitialized vector local default-initializes to all-zero lanes: ```mach var z: i32x4; # every lane is 0 ``` A literal is for a constant or for lanes assembled from unrelated scalars, and it is built lane by lane whatever its lanes read. Loads, stores and halves are ranges. **A window onto memory is a range converted to a vector.** `p[i, 4]` is the four elements at `p[i]` as a `[4]f32` value, and `::` between `[N]T` and `TxN` moves those elements into the lanes and back. The store is the same range as the assignment target: ```mach fragment val a: f32x4 = p[i, 4]::f32x4; # load the four floats at p[i] p[i, 4] = (a * a)::[4]f32; # store four floats at p[i] ``` The load is one unaligned vector load and the store one vector store (`movups` on x86-64), with no array value in between, so any element offset is valid. Nothing requires `i` to be a multiple of the lane count. `::` between an array and a vector is type-checked like every other `::`: the element type and the count must be the same on both sides, so `[4]i32` does not convert to `f32x4` and `[8]i16` does not convert to `i32x4`. The range and the conversion are described with the index and cast operators in [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#range). **The reinterpret is the bit-level view.** `@((?A[i]):~*f32x4)` reads the sixteen bytes at `A[i]` as an `f32x4` and stays legal. It lowers to the same load. It reads bytes, so the pointer's element type and the vector's lane type must agree in size and meaning: `(?A[i]):~*f32x4` over a `*f32` is the four floats at `i`, over a `*i32` it is their bits. **Half a vector is a range of its lanes.** `v[start, count]` on a vector is the vector of those `count` lanes, so `v[4, 4]` on an `i16x8` is an `i16x4`. Widening a half lane by lane is the range followed by a vector `::`: ```mach fragment val lo: i32x4 = v[0, 4]::i32x4; # widen the low half of an i16x8 val hi: i32x4 = v[4, 4]::i32x4; # widen the high half ``` Where the target packs the lane-halving extension (x86-64, aarch64) each line is that one instruction, signed or unsigned by the source lanes, and elsewhere it is the lane path. A vector range's start is a comptime constant, as a lane index is, and it takes at least 2 lanes (one lane is `v[i]`). **Lane access** `v[i]` reads or writes a single lane. The index must be a comptime constant in `[0, lanes)`; a dynamic (runtime) lane index is not supported in this increment. The bound is the same compile-time rule an array gets, reported the same way — `v[4]` on a `u32x4` is `index 4 is out of bounds for `u32x4` of length 4`. ```mach fragment var v: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; val x: f32 = v[0]; # read lane 0 v[3] = 9.0; # write lane 3 ``` There are no scalar↔vector casts in this increment: neither an implicit scalar-to-vector conversion nor a `1.0::f32x4` reinterpret is legal. Between two vectors with the same lane count, `::` converts each lane with the scalar rule (a lane of `-7` becomes `-7.0` in `i32x4::f32x4`), and `:~` reads the bits of any vector of the same byte size. The lane-wise operators and the comparison-to-mask rule are in [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md), including the shifts, whose two forms are `vec << vec` and `vec >> vec` with a count of the same shape (a uniform count is the same count in every lane of a literal, `v >> u32x4{k, k, k, k}`, never a scalar); what a target without hardware SIMD does with a vector operator is the `simd` profile lever ([manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md), [policy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md)). #### Handles A **handle** is a type whose representation is not the program's: the owning target mints it and the pipeline binds it. A shader reads a texture through one. The language knows only the machinery. A handle is a **bodyless `def`** carrying `#[handle(target, constructor, operands...)]`, and which constructors exist, what operands each takes, and what an operand means all belong to the named target: ```mach fragment #[handle("spirv", "image", TEXEL_F32, DIM_2D, NO_DEPTH, NONARRAYED, SINGLE_SAMPLED, SAMPLED, FORMAT_UNKNOWN)] pub def Texture2D; #[handle("spirv", "sampled_image", Texture2D)] pub def Sampler2D; #[handle("spirv", "sampler")] pub def Sampler; ``` `def` is the carrier because it already means "this name denotes a type" and promises no fields and no storage, which is exactly what a handle is. There is no body because the target supplies the definition. The spirv `image` constructor takes the seven operands of `OpTypeImage` in its order: the texel scalar (`0` `f32`, `1` `i32`, `2` `u32`, `3` `i64`, `4` `u64`), `Dim`, `Depth`, `Arrayed`, `MS`, `Sampled` and the `Image Format`. The format is required, and `0` (`Unknown`) is the usual choice for a sampled image. `Sampled` `1` is an image read through a sampler and `2` a storage image, read and written directly and bound through `#[storage]` (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)). A storage image of a known format needs no device feature to read or write, and a format other than `Rgba32f`, `Rgba16f`, `R32f`, `Rgba8`, `Rgba8Snorm` and the `Rgba32`, `Rgba16`, `Rgba8` and `R32` integer forms declares `StorageImageExtendedFormats`. An `Unknown` storage image is read and written under `StorageImageReadWithoutFormat` and `StorageImageWriteWithoutFormat`, which `vulkan1.3` accepts and an earlier `env` accepts only with the `storage_read_without_format` and `storage_write_without_format` extensions (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions)). A format must match the texel scalar: a float or normalized format is read as `f32`, a signed integer format as `i32` and an unsigned one as `u32`. `R64ui` and `R64i` hold 64-bit texels, the texel scalar `u64` or `i64`, and declare `Int64ImageEXT`, which needs the `image_int64_atomics` extension. `Depth` `1` is a **depth image**, the shadow map a depth comparison samples, and `0` any other. Only a depth image is sampled by the `OpImage*Dref*` comparisons (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)), and it is read without one like any other image. Vulkan places no `Depth` constraint on a storage image, so a storage image may declare `1` too, though no comparison reads it: a comparison samples through a sampler. `2`, which states no indication either way, is refused. ```mach fragment #[handle("spirv", "image", TEXEL_F32, DIM_2D, DEPTH, NONARRAYED, SINGLE_SAMPLED, SAMPLED, FORMAT_UNKNOWN)] pub def ShadowMap; #[handle("spirv", "sampled_image", ShadowMap)] pub def ShadowSampler; ``` `MS` `1` is a **multisampled** image, which only a `2D` image is, arrayed or not: it is refused on `1D`, `3D`, `Cube` and `Buffer`. A multisampled image is never sampled, so a `sampled_image` cannot compose over one, and its texels are read per sample with the `Sample` image operand (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)). A multisampled sampled image needs no capability. A multisampled storage image declares `StorageImageMultisample`, and an arrayed one `ImageMSArray` too, under the `storage_image_multisample` extension. The operands are ordinary comptime constants. One position is not: a constructor that composes over another handle takes a **type name**, so `Sampler2D` names the image it wraps rather than restating that image's operands and cannot disagree with it. That is the only place a decorator argument is read as a type. Every rule a handle carries follows from the one fact the directive states, and the set is fixed and closed rather than varied per declaration: - no fields, no indexing, no construction - it cannot be a local binding - it cannot be a record or union field - it cannot sit behind a pointer or inside an array - it reaches an operation only by being passed to one, bound as a descriptor - its extent is declared by the owning target A generic function takes a handle as a type argument, and each instance is held to these rules as the same function written out would be: the handle is passed by value, behind a pointer to its binding, or returned, and never named as a local or placed in an array. A generic record, union or tag is refused a type argument that would put a handle in one of its fields or payloads, directly or behind a pointer, array or secret, since that instance is a record holding a handle. A pointer to a handle is the address of its binding, and that is the only indirection a handle takes. A pointer to a pointer to a handle, or any deeper chain, is refused in every annotation (a parameter, a result, a local, a field or a type argument), since the inner pointer would be one the program stores and may reassign, so the handle would trace to no binding. `$size_of` a handle is the target's pointer size: it is a name for a resource, and a pointer is the shape every target already has for that. On a target that mints no such type the declaration is **inert**: it still denotes a type and still sizes, and an operation over it is an undefined symbol at link. A target refuses an operand combination its constructor spells but it cannot emit, naming the operand rather than the declaration. The SPIR-V target refuses a `Depth` other than `0` or `1`, an `MS` other than `0` or `1`, a `Dim` past `Cube` other than `Buffer`, and a `Sampled` other than `1` or `2`, each with what SPIR-V would need instead. See [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) for `#[handle]` and `#[sampler(set, binding)]`, and the shader library for the handles a SPIR-V target declares. #### ABI types An **ABI type** is a C type whose size and alignment come from the **selected target** and whose contents the program never reaches. It is a bodyless `def` carrying `#[abi_type(name)]`, and `va_list` is the only name: ```mach #[abi_type("va_list")] pub def VaList; ``` It exists for the one C type that has no correct spelling in mach source. `va_list` is a plain pointer on System V x86-64, Apple arm64, Microsoft x64 and RISC-V lp64d, and a 32-byte 8-aligned composite under AAPCS64 — so `ptr` binds correctly on four platforms and silently corrupts on the fifth, where the psABI passes the composite indirectly with the caller owning the copy. Unlike a handle, an ABI type is an **aggregate** of the declared extent rather than a pointer-shaped name for a resource. That is the whole mechanism: an aggregate over 16 bytes rides AAPCS64's by-reference path, which makes the caller copy and pass an address, and an eight-byte one reduces on every other platform to the single register a pointer would have taken. What a program may do with one is bounded to the opposite end from a handle's: - it may be a **parameter**, and it may be **passed on** to another function - it cannot be read: no fields, no indexing, no cast in either direction - it cannot be constructed, and no function may return one - it cannot be a local binding, a global, or a record or union field - it cannot sit behind a pointer or inside an array Reading one needs `va_arg`, which mach has no callee shape for: mach has no runtime variadics, so nothing walks its own argument tail. `$size_of` and `$align_of` answer with the target's declared numbers, which are published psABI facts rather than anything about a value. A target that declares no layout for the named type **refuses** the declaration, naming the exclusion. There is no inert outcome, because every default available would be a pointer and a pointer is wrong under AAPCS64. See [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) for the binding recipe and [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) for `#[abi_type]`. #### Pointer `*T` — pointer to a value of type `T`. ```mach fragment var x: i64; var p: *i64 = ?x; # address-of yields a pointer val v: i64 = @p; # dereference reads through it ``` A pointer carries no qualifiers: there is no `const` or `volatile` pointer. Immutability is a property of the binding (`val`), and volatility is a property of a declared record, union or tag (`#[volatile]`, see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#volatile--every-access-to-the-type-is-a-volatile-access)), so an access through `*T` is volatile exactly when `T`'s storage is. ##### Pointers on SPIR-V A SPIR-V target has two kinds of pointer, and mach spells both `*T`. The compiler works out which one a pointer is from where it comes from, the same way it works out the storage class of every pointer on that target. - A **logical** pointer names memory the module declares: a local, a shader interface variable, or an element or member of one. `?place` makes one, and so does an access through one. It has no address. It can be handed to a function, but it cannot be stored in memory, returned, compared, ordered, cast to an integer or stepped across whole objects. - A **physical** pointer is an address in a buffer the host passes by its device address (`vkGetBufferDeviceAddress`). A pointer read from memory is one, since memory holds no logical pointer. So is one made from an integer (`addr::*Node`), one a function returns, one passed for a physical pointer parameter, `nil`, and one reached through a physical pointer. ```mach fragment rec Node { next: *Node; value: u32; } rec Head { first: *Node; total: u32; } #[storage(0, 0)] var head: Head; #[stage("compute")] #[workgroup(1, 1, 1)] fun walk() { var p: *Node = head.first; var sum: u32 = 0; for (p != nil) { sum = sum + p.value; p = p.next; } head.total = sum; } ``` A physical pointer is a full `*T`. It is dereferenced, indexed, stepped (`?p[i]`), compared, ordered as an unsigned 64-bit address and cast to and from `u64`, and it can be held in memory, passed and returned. Its pointee is laid out by the std430 rules a `#[storage]` block follows (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)), refused where mach's layout disagrees with them. A record that reaches itself through pointers, like `Node` above, is a recursive type, and so is a set of records that reach each other. A tag cannot sit behind a physical pointer. Every load and store through a physical pointer carries the pointee's alignment, so an address made from an integer must be aligned for the type it points at. An atomic through one takes no memory operands and carries no alignment, so its address is held to the same rule, and the semantics passed to it name its memory `UniformMemory`, as on a storage buffer. A 64-bit or float atomic needs the `buffer_*` feature of its operation (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)). Every variable and parameter holding one is decorated aliased, since mach makes no promise that two pointers do not overlap. Accesses through one are private under the Vulkan memory model: no `"coherent"` qualifier reaches device memory. A register that one path makes logical and another physical is refused, as is every operation a logical pointer has no form for, each naming why. Physical pointers need the `buffer_device_address` device feature (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). A module that holds no physical pointer keeps the Logical addressing model and is unchanged by them. #### Array `[N]T` — array of exactly `N` values of type `T`. Nested: `[N][M]T`. ```mach val a: [4]i64 = [4]i64{1, 2, 3, 4}; val g: [2][2]i64 = [2][2]i64{[2]i64{1, 2}, [2]i64{3, 4}}; ``` **Constant indices are bounds-checked at compile time.** `N` is part of the type, so an index the compiler can fold must land in `[0, N)`: ```mach error index 4 is out of bounds for `[4]i32` of length 4 fun read() i32 { var xs: [4]i32; val a: i32 = xs[3]; # ok val b: i32 = xs[4]; # error: index 4 is out of bounds for `[4]i32` of length 4 val c: i32 = xs[-1]; # error: index -1 is out of bounds ... ret a + b + c; } ``` The rule is keyed on the length the type carries, not on how the array was spelled, so a `def` alias, an array field of a generic instance, an array nested in a record or in another array, and a `^`-qualified array are all checked the same way. It is exactly the length `$length_of` reports ([comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md)), and a vector's lane count takes the identical rule. Only a **constant** index is checked. A runtime index is not, and a pointer is not indexed against any length at all — `*T` carries none. **A range** `x[start, count]` is `count` consecutive elements from `start`, as a `[count]T` value, over an array or through a pointer. `count` is a comptime constant of at least 1, and `start` is any index. A constant `start` is checked the way a constant index is, against the whole range: `start + count` may reach `N` and may not pass it. ```mach error range [3, 2] is out of bounds for `[4]i32` of length 4 fun window() i32 { var xs: [4]i32; val a: [2]i32 = xs[2, 2]; # ok: elements 2 and 3 val b: [2]i32 = xs[3, 2]; # error: range [3, 2] is out of bounds for `[4]i32` of length 4 ret a[0] + b[0]; } ``` A range is also an assignment target, storing `count` elements from `start`: the stored value's type is exactly `[count]T`. A range read is a value, not a view: `?xs[0, 2]` is an error, and writing into a range (`xs[0, 2][1] = 3`) is too. With `::` to a vector of the same shape it is the vector load and store in [SIMD vectors](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#simd-vectors). #### Function type `fun(T1, T2) R` — first-class function-pointer type. ```mach use std.print; use std.runtime; fun add(a: i64, b: i64) i64 { ret a + b; } def BinOp: fun(i64, i64) i64; val op: BinOp = add; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val r: i64 = op(2, 3); print.printlnf("{}", r); ret 0; } ``` #### Record and union types `rec` and `uni` declarations produce named types. See [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md) and [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md). #### Tag types and the std failure tags A `tag` declaration introduces a named tagged value type that holds exactly one active case, selected by an explicitly typed discriminator. A case may be payloadless or carry one explicitly typed payload: ```mach tag Reply: u8 { empty; value: i64; } ``` The failure types every std API answers with are three ordinary std tags with fixed generic arities, declared in `std.types.result`, `std.types.option` and `std.types.error` and imported like any other declaration (`use std.types.result.res;`): - `res[T, E]` is an outcome with error case `err: E` first and success case `ok: T` second - `opt[T]` is presence with payloadless `none` first and `some: T` second - `err[E]` is an outcome with error case `err: E` first and payloadless success `ok` second `err[E]` is not an alias of `opt[E]`. There are no defaulted generic arguments, no general type inference and no dummy success types; the compiler has no knowledge of the three names, so a module that imports none of them cannot spell them, and a module may declare its own. The case names `ok`, `err`, `some`, and `none` are members of their tags, not keywords. Numeric vector spellings such as `f32x4` denote SIMD vector types and require full-lane initialization, whereas a tag value names one case and its payload (`Reply.empty{}` or `res[i64, ParseError].ok{42}`). Construction, the `sel` case test and the lexical payload guards are in [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). #### Type aliases `def NAME: TYPE;` introduces an alias. See [def.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/def.md). #### See also - [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) — the `^` secret qualifier over any of these types - [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) — what operations work on each type - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) — `$size_of`, `$align_of` Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md ### `^` — the secret qualifier `^T` marks a type as carrying secret data for Mach's constant-time guarantee. Sema tracks secrecy as an information-flow discipline: secret data may move and be stored, but it may never reach a position a classical leakage model observes (a branch, a memory address, a variable-latency instruction). This page covers the flow-typing rules and the `#[oblivious]` codegen contract. The grammar lives in [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md); the decorator reference is in [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md). > **Experimental preview.** The constant-time support is incomplete and > unaudited. Read [Assurance](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#assurance) before relying on any of this. > **Do not build production cryptography on it at this version.** #### Secrecy lattice There are two secrecy levels in a two-point lattice: public is the bottom, secret the top. `^` lifts a type to secret and binds to the type immediately to its right, so it nests with `*` and `[N]` in any order: `^u32`, `*^u8`, `^*u8`, `[N]^u8`, `^MyRec`. Doubling collapses: `^^T` is `^T`. A public value coerces *up* to secret wherever a secret is expected, with no syntax: ```mach fun up(p: u32) ^u32 { ret p; # public u32 flows into a secret slot } ``` A **literal is public by construction** and stays public through that coercion: its value sits in the instruction stream, so classifying it as secret would protect nothing. `var s: ^u32 = 5;` types `5` as `u32` and relies on the up-coerce. The join below is unaffected — a value *computed* with a secret is secret however its other operand is spelled — so `v << 3` on a secret `v` still yields a secret, while the constant `3` is not mistaken for a secret shift count. The reverse never happens implicitly. The only downgrade is the explicit `:>T` strip below. #### Join Any operation with a secret operand yields a secret result. Taint joins across arithmetic, bitwise, shift, and comparison operators, and through a value read out of a secret container: ```mach #[oblivious] fun mix(a: ^u32, b: u32) ^u32 { ret a + b; # ^u32 + u32 -> ^u32 } rec Key { d: ^[32]u8; } fun first(k: Key) ^u8 { ret k.d[0]; # element of a secret array is ^u8 } ``` Taking the address of a `^T` value with `?` gives the public pointer `*^T` (the address is public, the pointee secret), and dereferencing it with `@` recovers the secret `^T`. #### Gates A secret may not reach a position the leakage model observes. Each is a compile error decided by operand type: - a secret branch or loop condition (`if`, `for`) - a secret left operand of a short-circuiting `&&` / `||` (it is the branch the operator keys on; a secret right operand only taints the result) - a secret memory index (`table[i]` with `i` secret) - a secret memory address — an access through a secret *pointer*, whether by `@p`, `p[i]`, or the auto-deref in `p.x`. The index and the address are the two halves of one effective address, so both are gated. Only a secret *pointer* is an address: a `^[N]T` or `^Rec` is a secret value living at a public address, and a `*^T` is a public address to secret storage - a secret operand of the always-variable-latency `/` or `%` ```mach error secret value used as a branch condition fun leak(a: ^u32, t: *u8, p: ^*u8) u8 { if (a) { ret 1; } # error: secret value used as a branch condition ret t[a]; # error: secret value used as a memory index ret @p; # error: secret value used as a memory address } ``` Three more gates decide against the target's constant-time capabilities rather than the source alone, so they are reported at lowering: - a secret operand of a **floating-point** operation (always variable-latency, gated on every target) - a secret operand of an **integer multiply**, unless the target declares the exact multiply it emits (the low half, a high half or the widening product, at that operand width) as data-independent-timing under a condition the build meets. The declared cells are the table in [Constant-time multiply by instruction set](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#constant-time-multiply-by-instruction-set): x86-64 admits every scalar cell unconditionally, aarch64 admits its cells only while PSTATE.DIT is on and only on an operating system that guarantees the mode, riscv64 admits its cells only when the selection holds `m` and `zkt`, and riscv32, SPIR-V and every lane multiply are refused. The refusal names the target and the condition it lacks. `$mach.build.ct_mul(op, width)` reads the same decision at comptime (see `comptime-mach.md`). A secret 128-bit `/` or `%` is refused like every secret division: the helper it would call is a loop over the dividend's bits - a secret **variable shift count** on a target without a barrel shifter. Where it is admitted, the saturation of a count at or above the operand width ([operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#bitwise)) is a compare to an all-ones-or-zero mask and a mask on the count and the value, with no branch, so it reveals nothing the shift itself would not A secret value passed to a variadic pack is also rejected, including a secret wrapped inside an aggregate. The gates are checked against the types of the **instance**, not of the template. A generic's body is re-checked per instantiation under its concrete type arguments, and a pack-tailed function's body — the statements around its `$each` as well as the unrolled body itself — is re-typed per monomorphized instance, so a `T` that instantiates to a secret is gated exactly as the secret spelled out in full would be. #### Asking about secrecy at comptime Secrecy is asked about at comptime through two predicates, because a library has two different questions and each needs its own answer: - `$is_secret(T)` is the **per-field** question a reflection walk asks: is this type `^`-qualified at the outermost level? - `$holds_secret(T)` is the **byte-class** question a memory primitive asks: is any byte of this type secret? Both are type predicates like `$is_record` / `$is_union` / `$is_pointer`: comptime-only, valid as a `$if` / `$or` gate condition, answered per instantiation inside a generic. The full reference is in [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md). ##### `$is_secret`: the question a walk asks `$is_secret(T)` folds true when `T` is `^`-qualified **at the outermost level**, and false otherwise. It exists because secrecy was otherwise invisible to a library. Every other predicate asks about the shape *under* the `^` and so answers false for every secret, which makes `$is_record(^u64)` and `$is_record(u64)` the same answer — a reflection walk could only ever meet a secret field as a fallthrough it had to refuse. `$is_secret` is the positive question, and it is what lets `std.derive` decide rather than refuse: a formatter redacts a secret field, a hash refuses one (a data-dependent fold is a leak in the shape of a digest), and an equality picks the constant-time comparison instead of the early-out whose timing *is* the secret. ```mach fragment rec Session { id: u64; key: ^[32]u8; } $each f in $fields(Session) { $if ($is_secret(f.type)) { } # redact: no read of `key` is emitted $or { render(s.[f]); } } ``` **Outermost only**, the same line the rest of the family draws: - `^*u8` is secret — the pointer *is* the secret, which is the shape the welded-storage rules exist for - `*^u8` is **not** — a public address to secret storage. Ask `$is_secret($pointee_of(f.type))` for the pointee - `[N]^u8` is not — a public array whose elements are secret - a record with a secret field is not — the **field** is, and that is where a walk meets the question ##### `$holds_secret`: the question a memory primitive asks `$holds_secret(T)` folds true when any part of `T` is secret: `^u8`, `[32]^u8`, `Session` above, and `Box[^u64]` all answer true, while `u64`, `[32]u8` and a record of public fields answer false. It is the answer the compiler already computes when it refuses to erase a typed pointer to the raw `ptr` (see [Welded-storage pointers](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#welded-storage-pointers)), one implementation rather than two, so a gate on it and the cast check can never disagree. That walk follows typed pointers, so `*^u8` holds a secret too: erasing a `**^u8` to `ptr` is refused for the same reason. That makes it the question a primitive like `std.memory.zero[T]` needs. A word kernel over `ptr` is legal only for a type with no secret byte, a wholly secret type goes to the constant-time kernel, and a type that mixes the two is handled element by element: ```mach fragment pub fun zero[T](p: *T, count: usize) { $if (!$holds_secret(T)) { raw_zero(p::ptr, $size_of(T) * count); } $or ($is_secret(T)) { ct.zeroize(p::*^u8, $size_of(T) * count); } $or { zero_typed[T](p, count); } } ``` ##### Why a walk must not ask the deep question A walk that gated on `$holds_secret` instead of `$is_secret` would be wrong in the common case. A formatter walking a record whose one field is secret would see `$holds_secret` answer true for the record and redact all of it, public `id` included, and it still could not say which field was the secret. The per-field question is the one a walk actually has, and `$is_secret(f.type)` is exactly it. The deep question belongs to code that treats a type as bytes, where one secret byte decides how every byte is moved. Note that a walk which skips a secret field is pinned by the flow rules rather than by convention — reading one into a public accumulator does not compile, so a walk that gates wrongly is a compile error, not a silent disclosure. #### Downgrade with `:>T` `:>T` is the only way to remove `^`. It produces a new public value and never reinterprets storage in place, and it always names the public type it lands on: ```mach fun publish(a: ^u32) u32 { ret a:>u32; } fun publish2(a: ^*u8) *u8 { ret a:>*u8; } ``` `:^` is no operator; the parse stops at the colon: ```mach error expected ';' after 'ret' fun publish(a: ^u32) u32 { ret a:^u32; } ``` `:>T` peels exactly the outer qualifier, so it can never launder a welded pointee (`*^T` stays `*^T`). `::` and `:~` may neither add nor drop `^`. Promotion needs no operator of its own: a public value coerces up to secret implicitly, and a cast to a secret type (`x::^u16`) is an ordinary cast. `:>T` is the only declassification and there is no untyped form: `:^` is not an operator, so `x:^` and `x:^T` are parse errors at the colon. Inside a generic body the operand may be typed by a parameter (`fun show[U](s: U) u32 { ret s:>u32; }`): the template cannot decide whether `u32` is `U` stripped, so it checks only that the target is public and each instance settles the equality under its concrete arguments, the way every other secrecy gate is checked against the instance. #### Welded-storage pointers Secrecy is fixed at declaration and is non-launderable, which makes the public/secret aliasing leak unconstructable with no alias analysis: - a `^T` is stored only through a `*^T`, never a `*T` (the lattice forbids the downgrade) - a secret-welded pointer cannot be erased to the untyped `ptr` - a `uni`'s overlapping variants must agree on secrecy ```mach error union variants must agree on secrecy fun erase(p: *^u8) ptr { ret p; # error: type mismatch: expected ptr, found *^u8 } uni Bad { a: ^u32; b: u32; } # error: variants disagree on secrecy ``` The union rule is a property of the union **type**, not of the syntax that declared it, so it holds for an inline `uni { ... }` with no declaration to hang a check on, and at every **instance** of a generic union. At the declaration a variant typed by a generic parameter says nothing about secrecy, so `uni U[T] { a: T; b: u32; }` agrees there and is decided where each instance is formed: ```mach error this instantiation makes overlapping fields part secret uni U[T] { a: T; b: u32; } rec Box[T] { u: U[T]; } var s: U[^u32]; # error: this instantiation makes the variants disagree var b: Box[^u32]; # same error: the instance need not be spelled var p: U[u32]; # fine, and so is an all-secret instantiation ``` A *partially* concrete template (`U[T, ^u32]` written inside another generic) is a legal annotation with both legal and illegal instantiations, so it is never rejected where it is written — only at the arguments that actually make a pair mixed. The check is **deep** and **fails closed**: a secret nested anywhere inside an aggregate (including a generic instance's lazily-materialized fields) counts as secret at these boundaries, and a placement the checker cannot prove severs no weld — a secrecy difference reachable through a function type, or a shape whose layout it cannot determine — is rejected rather than allowed. Types are compared **by byte extent**, not by field ordinal. Each type lays out as a run of bytes, each byte secret or public, with padding counted as public and a pointer counted by what it reaches. A `::` or `:~` from `*S` to `*T` retypes the storage `S` covers, and a `*S` may address a run of `S` values, so the rule is directional: 1. every byte both `S` and `T` cover has the same class on both sides 2. a narrowing (`T` no larger than `S`) is accepted 3. a widening is accepted only when the extra bytes cannot change class: every byte of `S` has one class and every byte of `T` has that same class 4. beneath a pointer (a pointer field, a pointer to a pointer, a union variant of pointer type) the target is shared storage, so the relation holds both ways: equal extent and the same class at every byte The rule is the same for scalars and aggregates, and the variants of a union overlay its own storage from offset 0, so every pair of variants agrees on each byte both cover. A widening past a secret is refused because the extra bytes may be the next element of a run: ```mach error cannot add or drop the secret qualifier rec Key { x: ^u32; } rec Wide { x: ^u32; y: u32; } # fine: one secret class throughout, as a run of `^u8` is fun words(p: *^u8) *^u64 { ret p::*^u64; } # fine: a narrowing keeps the class of every byte it still covers fun head(p: *Wide) *Key { ret p::*Key; } # refused: `y` of the first element is the next element's secret `x` fun widen(p: *Key) *Wide { ret p::*Wide; } ``` A type that stores no value (an empty `rec`, a `[0]T`, or an aggregate of only those) has zero extent and so no class: anything narrows to it, and it widens to nothing that stores a value. Reading past an object by indexing (`(?q.a)[1]`) is a bounds question no cast rule closes. The comparison reads the same layout the backend emits. Where it cannot determine a layout it declines, which rejects. #### Comparing and ordering addresses `==`, `!=`, `<`, `<=`, `>` and `>=` all accept two pointer-like operands whatever their pointees' secrecy, including `*^T` against `*^T` and `*^T` against a public `*U`. The result is a public `bool` — `u8`, per [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) — so it branches, feeds `&&`, and returns like any other public value: ```mach use std.types.bool.bool; use std.types.size.usize; fun overlap(first: *^u8, first_len: usize, second: *^u8, second_len: usize) bool { ret first <= ?second[second_len - 1] && second <= ?first[first_len - 1]; } ``` **This launders nothing**, and the reason is the line the whole model draws: a `*^T` is a *public address to secret storage*. The address was never secret, so reading it is not a downgrade. What the welding rules protect is the *pointee* — that the bytes at that address can only ever be reached as `^T` — and ordering reaches no bytes at all. It consumes two addresses and produces one `bool`; no integer and no pointer comes out of it, so there is no value to turn back into a `*T` and dereference. Ordering is exactly as revealing as the `==` the language has always permitted, which answers the same question one address at a time. The conversions stay closed. None of these become legal by admitting ordering, and each has a test pinning it: | refused | why | | --- | --- | | `p::usize`, `p:~usize` | `::` and `:~` may neither add nor drop `^` | | `p:>*u8` | a `:>T` target must name the operand's *stripped* type, and `*^u8` stripped is `*^u8` | | `@((?p):~*usize)` | the same, reached through a pointer to the slot | | `val e: ptr = p;` | a secret-welded pointer does not erase to `ptr` | Three deliberate decisions: - **A differing pointee type orders.** `*^u8 < *u32` compiles. Addresses are addresses; both operands are pointer-width and the comparison is over the address space, not over either pointee. This is the pair `==` already accepts, and narrowing it for ordering alone would refuse the common case of relating a byte cursor to a typed buffer. - **`nil` orders.** `p > nil` compiles and compares against the null address as zero. It falls out of `nil` being a pointer-like operand, the same way `p == nil` does. It is rarely what you mean — `p != nil` is — but it is not a leak, and special-casing it would be surface with nothing behind it. - **The comparison is not itself constant-time-gated, and does not need to be.** Both operands are public addresses, so ordering is `CT_OP_NONE`: it lowers to one unsigned integer compare of pointer width, data-independent on every supported target. There is no secret operand for the leakage model to track. Inside a `#[oblivious]` function this needs no exception. The result of ordering two public addresses is public, so it may steer a branch there like any other public condition. The gates are unchanged where the *pointer itself* is the secret: a `^*T` is a secret value, the order of two of them is secret, and that `bool` still cannot be a branch condition. ```mach error secret value used as a branch condition use std.types.bool.bool; #[oblivious] fun before(a: *^u8, b: *^u8) bool { ret a < b; # public addresses, public bool } fun leak(a: ^*u8, b: ^*u8) u8 { if (a < b) { ret 1; } # error: secret value used as a branch condition ret 0; } ``` #### `#[oblivious]` — the codegen contract The flow typing constrains the *source*; `#[oblivious]` carries the obligation through *codegen*. Inside a function carrying it, the backend must not introduce a secret-dependent branch or select a variable-latency instruction on a secret operand. Inline `asm` inside such a function is **validated**, not rejected. The block is parsed into instructions and walked for the same three leaks the compiler checks everywhere else — a secret reaching a branch condition, a memory address, or a variable-latency operation the target cannot do in constant time. Taint enters through the block's `{name}` bindings, whose secrecy is stamped from the local's declared type. A pointer to a secret (`*^u32`) is a public address and a secret load: the register it is staged into may address memory, and what a load through it produces is secret, while a secret pointer (`^*u32`) is a secret address. What the walk cannot model, it refuses: | construct | why | |---|---| | a body that does not parse | nothing to analyze | | a data directive (`.byte`, `.word`, `.long`, `.quad`) | its payload can encode any instruction | | a mnemonic the target has not classified | its timing behaviour is unknown | | **a flags-conditioned branch on a secret** (x86-64 `jcc`, aarch64 `b.`) | the flags it reads were filled from a secret | That last row is a per-target asymmetry worth stating precisely, because getting it wrong was a real hole (#2477). A branch whose condition is a **register operand** is visible to the walk and is checked: aarch64's `cbz`/`cbnz`, and every riscv64 branch, which compares two registers — RISC-V has no flags register at all. A branch whose condition rides the **flags register** is checked against a separate flags-taint state (#2460): a `cmp` of a secret marks the flags secret-derived, and a later `jcc` reading them is refused, while a compare-and-branch over public data is accepted. x86-64's `jcc` family is flags-conditioned, and so is **aarch64's `b.`** — which is easy to miss, because `b.` does not appear in aarch64's mnemonic table at all. It is admitted by the grammar's *decode hook*, carrying the unconditional branch's own opcode with the condition in its flags. The flags model is deliberately more than a single "writes flags" bit, because one bit has no correct value for two families. `inc` and `dec` write every status flag **except** carry, and the shifts write nothing at all when the count is zero — so neither can be taken to have refreshed the flags. A secret `cmp`, a public `inc`, and a `jc` is therefore still refused: the carry the branch reads is the secret one. Only an instruction that unconditionally overwrites every flag this grammar can read (`add`, `sub`, `cmp`, `test`, `and`, `or`, `xor`, `neg`, `cmpxchg`, `xadd`) clears the taint. `popfq`, `iretq` and `syscall` load or mask flags from somewhere the model does not represent, so they taint unconditionally and never clear. A `setcc` **consumes** a flag into a register, so with secret-derived flags its destination is secret too — otherwise a secret could be laundered out of the flags and used as a memory index, which the address check would then miss. On a target whose grammar has no flag reader at all, every one of these facts is vacuous, and that vacuity is asserted per target rather than assumed: adding an aarch64 `cset` later has to state its facts deliberately. The rows whose facts could wrongly *permit* a leak were measured against the CPU rather than asserted from a manual: each instruction was run twice with the flags preset all-set and all-clear, and what it actually defines, preserves and reads was read off the hardware and written into the table. Three rows are exempt and stay a reasoned classification: `popfq`, `iretq` and `syscall` take their flags from the stack, the interrupt frame, or a masked prior RFLAGS, and no experiment of that shape can tell you where a value came *from*. **Nothing re-measures the table today.** It is a recorded classification, and a CPU whose behaviour stops matching it is a silent divergence. Re-establishing it is a job for a unit test that runs the probe on the host it is running on, not for a cross-compilation suite: the experiment only means anything on the ISA it executes. **Memory is the third taint domain**, beside the register set and the flags bit (#2706). A secret spilled to the stack and reloaded comes back **secret**: ``` mov rax, {r} mov [rsp], rax mov rbx, [rsp] mov rcx, [rbx] # error: addresses memory with a secret value ``` Before this domain existed the reload came back in a register the walk believed public, and all three checks downstream of it were defeated at once — a branch condition, a memory address, and a variable-latency operand alike. The direct form (`mov rax, {r}` then `mov rbx, [rax]`) was always refused, which made the gate look sound while an ordinary spill walked straight through it. The domain is **one bit covering all of memory**: either a secret has been stored in this body or it has not. That is coarser than it could be, and deliberately so. The obvious refinement — taint individual `[base + literal]` slots — is not sound here, because `[rsp + k]` and `[rbp + j]` can name the same byte with nothing in the walk relating two base registers, and because a slot key is an address rather than an extent, so an 8-byte store and an overlapping 4-byte load compare as different slots. Both would be laundering paths again. The bit is **monotone**: a later store of a public value does not clear it, since "the same slot" is precisely the question the domain cannot answer. So a body loses by this only if it stores a secret, later reloads some *other* public value, and then branches on, addresses with, or variable-latency operates on what it reloaded. The failure direction is refusal, never acceptance. A store through an address the walk cannot resolve needs no special case: there is no slot to fail to identify. The variable-latency check also bites unevenly: x86-64's and aarch64's asm grammars carry no divide, multiply or float instruction at all, so it reaches only their register-count shifts. riscv64's grammar admits the whole M-extension, so on that target the check is substantive. `#[oblivious]` remains a **per-function** contract. A call out to a non-oblivious function is not validated — that is the boundary the decorator draws, not a hole in it, and it applies to a callee containing `asm` exactly as it applies to any other. The zeroizing-write guarantee is separate and broader; it is described below. A secret-taint bit is threaded from sema's flow typing through IR and MIR to the emitted instruction stream, preserved across every value replacement, inline clone, instruction selection, and register-allocator copy. The one place taint stops is the declassify barrier a `:>T` cast lowers to. Secret-free code carries no taint and compiles byte-identically. A function instance that **computes** on a secret must carry `#[oblivious]`; one that only moves, stores, or declassifies secrets is transparent and stays annotation-free. The check runs per monomorphized instance, so a generic instantiated at a secret type is held to the concrete type's rules. ```mach #[oblivious] fun ct_select(mask: ^u32, a: ^u32, b: ^u32) ^u32 { ret (a & mask) | (b & ~mask); } ``` A **translation validator** re-derives the taint over the lowered, target-independent MIR as a monotone dataflow fixpoint and independently re-checks the leakage conditions, so a secret reaching a branch condition, a memory address, or a forbidden variable-latency op is a compile error naming the function and the offending operation. It is a backstop behind the compile-time gates, not a replacement for them. #### The zeroizing-write guarantee Wiping a secret is only useful if the wipe survives to run. That guarantee exists, but it is **not** provided by `#[oblivious]`, and it is scoped more broadly than the decorator is. A store into secret storage is tainted at lowering, from two composing sources: the stored **value**'s secrecy, and — for a public value written into secret storage, the shape a wipe takes — the **destination**'s secrecy, read from the lvalue's semantic type. Either one marks the store. That taint is the thing an optimization must consult, and it is keyed on the storage, not on any decorator, so: ```mach use std.types.size.usize; # no decorator: the wipe is protected anyway fun clear(p: *^u8, n: usize) { var i: usize = 0; for (i < n) { p[i] = 0; i = i + 1; } } ``` **What it covers.** Memory, reached through a pointer that escapes the function — the shape `crypto.ct.zeroize` has. Such a store cannot be promoted out of memory, and the taint is present for any future pass to read. **What it does not cover.** A value the compiler keeps in a **register**. Writing to a promoted local is not a memory write, so wiping one is not preserved: ```mach fragment var x: ^u8 = k; x = 0; # NOT guaranteed: `x` may never have been in memory ``` Adding `#[oblivious]` does not change this — the wipe is outside the memory-address leakage model rather than an exception within it. To wipe reliably, write through a pointer whose target is memory the compiler cannot promote away, which is what the standard library's `zeroize` does. Settled: the wipe guarantee is memory-scoped; secret register lifetimes are outside it (#2456). **Today the guarantee holds trivially.** mach has no dead-store elimination, so nothing removes a store: an entirely dead fill of a *public* local also survives at release. The guarantee becomes load-bearing only if such a pass is added, and the taint is what would hold that pass to it. `mach.lang.driver:secret_store_taint_survives_lower` pins the taint. The contract is only offered where mach emits the instructions that execute. A target whose back half hands a module to a downstream compiler instead — the experimental SPIR-V backend — **rejects `#[oblivious]`**: neither that translation nor the device's timing behaviour is covered by the leakage model, so the obligation could be neither validated nor upheld. Compile constant-time code for a machine target and pass such a target only public data. #### Constant-time multiply by instruction set `*` on a secret operand means what it means on a public one: the wrapping, same-width product. What changes per target is whether the machine can execute it without a timing leak, and mach decides that from a catalog each instruction set declares, never from a guess. A **row** of the catalog names one cell (a multiply, an operand width) and the **condition** under which that cell has data-independent timing: - **always**: nothing to check; - **PSTATE.DIT on**: the cell is data-independent only while the processor's DIT mode is set, so the row is admitted only on an operating system that declares it guarantees the mode (see [PSTATE.DIT at run time](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#pstatedit-at-run-time)); - **extensions selected**: the cell is data-independent only on a machine that has the named extensions, so the row is admitted only when the target selects every one of them (`extensions` in the manifest, or the isa string). The multiply column is the spelling `$mach.build.ct_mul` takes: `low` is the same-width product, `high_u` and `high_s` the upper half of the unsigned and signed full product, `high_su` the upper half of a signed-by-unsigned product, and `wide_u` and `wide_s` the full double-width product of two operands of the named width as one instruction. The width is the operand's: a product whose result is no wider than the machine's ALU is a `low` cell at the result width (`(a::^u64) * (b::^u64)` over 32-bit `a` and `b` is the 64-bit `low` cell on a 64-bit target), and only a product wider than the ALU is fused into the widening cell (see [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md)). The narrower `wide` and `high` rows classify the inline-asm forms (the one-operand `mul r32`) the same way. The declared rows, one line per multiply: | Instruction set | Multiply | Widths | Condition | |---|---|---|---| | `x86_64` | `low` | 8, 16, 32, 64 | always | | `x86_64` | `high_u` | 8, 16, 32, 64 | always | | `x86_64` | `high_s` | 8, 16, 32, 64 | always | | `x86_64` | `wide_u` | 8, 16, 32, 64 | always | | `x86_64` | `wide_s` | 8, 16, 32, 64 | always | | `aarch64` | `low` | 8, 16, 32, 64 | PSTATE.DIT on | | `aarch64` | `high_u` | 64 | PSTATE.DIT on | | `aarch64` | `high_s` | 64 | PSTATE.DIT on | | `aarch64` | `wide_u` | 32 | PSTATE.DIT on | | `aarch64` | `wide_s` | 32 | PSTATE.DIT on | | `riscv64` | `low` | 8, 16, 32, 64 | `m` and `zkt` selected | | `riscv64` | `high_u` | 64 | `m` and `zkt` selected | | `riscv64` | `high_s` | 64 | `m` and `zkt` selected | | `riscv64` | `high_su` | 64 | `m` and `zkt` selected | | `riscv32` | none | | | | `spirv` | none | | | A test reads this table back against the compiler's catalog, so the two cannot drift. A cell the table does not hold is refused, as is every lane multiply (`i32x4 * i32x4` on a secret vector): the vector rows exist in the contract and no instruction set declares one. **Realized cells.** A cell no instruction executes directly is admitted exactly when every cell of its realization is, under the union of their conditions. On x86-64 the one-operand `mul` and `imul` write the full product to a register pair, so the high halves and the widening products are the same instruction; on aarch64 and riscv64 the full product is the low multiply beside the high one (`mul` + `umulh`/`smulh`, `mul` + `mulhu`/`mulh`), so the 64-bit `wide_u` and `wide_s` are admitted there through the 64-bit `low` and high rows. A **128-bit** multiply is never declared by a target, because no target executes one: `(a::^u128) * (b::^u128)` with 64-bit `a` and `b` is the 64-bit widening product, one instruction, and its high half `(... >> 64)::^u64` is that same instruction; a secret `^u128 * ^u128` low product is the schoolbook over the lanes, three 64-bit low products and the widening one. `$mach.build.ct_mul(low, 128)` and `$mach.build.ct_mul(wide_u, 64)` answer for the realized cells the same way the gate does, so on aarch64 they need DIT like the 64-bit high half. **Why each row holds.** Every row cites its source beside the declaration in the instruction set's registration, and the kind of each claim is marked: - **x86-64, always, Intel and AMD (#3508).** On Intel it is a vendor guarantee: the "Data Operand Independent Timing" guidance states that processors which do not enumerate DOITM may be assumed to behave as if it were enabled for the listed instructions, and `mul` (F6, F7), `imul` (69, 6B, 0F AF, F6, F7) and `mulx` (F6) are on that list, so every operand size of the one- and two-operand forms. BearSSL's `ctmul` table, measured per core, records a constant-time multiply since the first Pentium. On AMD the evidence is by table rather than by vendor document: AMD publishes no equivalent list and no mode control (absence, verified 2026-09-18), the same BearSSL table records a constant-time multiply at every width for every AMD core from K7 through Zen, no AMD x86-64 core has ever been documented with a data-dependent multiplier, and AMD's SB-1039 advises constant-time algorithms, which presupposes constant-time instructions. The compiler emits `imul` for the low half and the one-operand `mul`/`imul` for the high halves and widening products. - **aarch64, PSTATE.DIT on.** The Arm ARM (DDI 0487, "About PSTATE.DIT") and the DIT register page of DDI 0601 list the data-processing (3 source) instructions `madd` (`mul`), `msub`, `smaddl` (`smull`), `smsubl`, `smulh`, `umaddl` (`umull`), `umsubl` and `umulh` as data-independent while DIT is set. The 8- and 16-bit low halves are the 32-bit `madd`, the 64-bit high halves are `umulh` and `smulh`, and the 32-bit widening rows are `umaddl` and `smaddl`. Whether a process can set and hold the mode is the operating system's declaration, and on an OS that declares nothing (windows, freestanding) the refusal names the instruction set that declares the rows and the operating system that declares no guarantee. - **riscv64, `m` and `zkt` selected.** The RISC-V Cryptography Extensions Volume I (scalar), chapter "Data Independent Execution Latency Subset: Zkt", table RVM, lists `mul`, `mulh`, `mulhsu`, `mulhu` and (rv64) `mulw`, and excludes `div` and `rem`. Zkt changes no instruction: it is a promise about the hart's timing, so like any selected extension it is assumed of every machine the binary runs on. riscv32 declares no rows yet. - **SPIR-V** has no timing model and declares nothing. **DOITM is not a multiply condition.** Intel's Data Operand Independent Timing Mode (`IA32_UARCH_MISC_CTL[0]`) hardens the data dependent prefetcher and the fast store forwarding predictor, memory-side predictors keyed on data values; it does not touch the multiplier, which is why the x86-64 rows hold without it (#3623). It is a model-specific register the kernel owns, user space cannot set it, and Linux does not enable it by default, so mach neither declares it nor offers a manifest key for it. A program that wants that hardening asks the operator for the kernel control, the same class as SMT and other side-channel mitigations. **Fail closed, never fall back.** A secret multiply the catalog does not admit is a compile error; the compiler never substitutes a bit-serial loop, a shift-add expansion or a call for it, and it has no multiply strength reduction, so a secret square and a secret multiply by a constant each reach the machine as the one multiply instruction. A per-ISA test decodes the emitted function and checks exactly that: one multiply, no call, branch or loop, on riscv64 through its Zkt rows, x86-64 through its unconditional rows and aarch64 through its DIT rows on linux and darwin. A program that compiles through a `PSTATE.DIT on` row and reaches a processor without the mode is refused at start by the std runtime, before any secret is multiplied, as the next section describes. #### PSTATE.DIT at run time An aarch64 multiply has data-independent timing only while the processor's DIT mode is on (Arm ARM DDI 0487, "About PSTATE.DIT"; the DIT register page of DDI 0601 lists the instructions: `madd`, `msub`, `smaddl`, `smsubl`, `smulh`, `umaddl`, `umsubl`, `umulh`). The mode is per-thread processor state that user code turns on with `msr dit, 1`, and a processor without FEAT_DIT has no such bit. So admitting the aarch64 rows takes two facts the compiler does not own, and the design splits them: **The operating system declares the guarantee.** A target's OS table declares, per instruction set, whether a process can set the mode at start, learn whether the processor has it, and keep it. The declaration is read through one accessor and never derived, and each carries its citation beside it: - **aarch64-linux declares it.** The kernel exposes the processor's FEAT_DIT as `HWCAP_DIT` (`1 << 24`) in `AT_HWCAP` since v4.17 (commit 7206dc93a58f, "arm64: Expose Arm v8.4 features"; `Documentation/arch/arm64/elf_hwcaps.rst`: "Functionality implied by ID_AA64PFR0_EL1.DIT == 0b0001"). PSTATE is saved to `SPSR_EL1` on every exception entry and restored on return, so the bit outlives a context switch. - **aarch64-darwin declares it.** Apple's "Writing ARM64 code for Apple platforms", section "Enable DIT for constant-time cryptographic operations", documents the mode as per-thread state user code turns on with `msr dit, #1` and reads back as bit 24 of `mrs dit`, and names the sysctl `hw.optional.arm.FEAT_DIT` as the check for a processor that has it. - **aarch64-windows and freestanding aarch64 declare nothing**, so a `DIT_MODE` row is refused there with a diagnostic that names the rows and the OS. **The compiler records the need and the runtime honours it.** The rows only say the instruction is safe *while the mode is on*; something must turn it on, and only for programs that need it. When a module's lowering admits a secret multiply through a `DIT_MODE` row, codegen marks the module: its object carries a local absolute symbol `__mach_needs_dit` (value 1) that reaches no linked image. When an executable is linked for a target whose OS declares the guarantee, the linker appends one synthetic input defining `__mach_dit_required`, a hidden one-byte read-only object that is `1` when any input carries the mark and `0` otherwise. The std start code reads that byte before `main`. When it is zero the program never touches DIT. When it is set, std checks that the processor has the mode (`HWCAP_DIT` on linux, `hw.optional.arm.FEAT_DIT` on darwin), turns it on with `msr dit, 1` followed by `dsb nsh; isb`, and does the same at the entry of every thread it creates, since the bit is per thread. A processor or kernel without the mode fails closed: std writes ``` std.runtime: this program contains a constant-time multiply that requires the processor's data-independent-timing mode (PSTATE.DIT), and this processor or kernel does not provide it (aarch64-linux: HWCAP_DIT absent; aarch64-darwin: hw.optional.arm.FEAT_DIT is 0); refusing to start ``` to stderr and terminates through its panic path (exit status 255) before any secret is multiplied. Two consequences of the design are worth knowing. The mark is per module, so a module that contains such a multiply sets the byte even when the linker dead-strips the function; the cost is one unneeded `msr`. And "a binary without the need carries nothing DIT-related" holds up to that one byte: the language has no weak import, so std reads a cell that is always defined on a declaring target, and a binary without the need carries it as `0`. A shared library gets no cell, because no start code of std's runs in one, and an object built by another compiler carries no mark. `$mach.build.ct_mul(op, width)` answers `1` for the aarch64 rows exactly where the OS declares the guarantee. #### Trusted base The only secret-to-public crossings are the explicit `:>T` cast and inline `asm` blocks (which a type system cannot check). Everything else is enforced. A proof is always relative to a leakage model. Its fidelity to real silicon is empirical. #### Assurance **The constant-time guarantee is incomplete. This support is an experimental preview and has not been audited. Do not build production cryptography on it at this version.** What holds today: the type system checks that the source respects the leakage model, `#[oblivious]` carries the obligation through codegen, and the translation validator independently re-checks the lowered MIR. All three are static. Two host-executed measurements sit under them, and each only means anything on the machine it runs on, which is why both are unit tests rather than build checks. **The timing harness is a tool, not a gate.** `mach.lang.ct.probe` is a dudect-style harness in the tree. `mach test` runs it at deliberately tiny sample counts and prints its table; it asserts **nothing about the numbers** - not a threshold, not positivity - because even "a mean is positive" proved to be a property of the host's clock resolution rather than of the harness (#3092). The run can only fail if the harness itself crashes. **No property of a measured time gates anything**, here or anywhere else in the suite. That is deliberate and it is not a gap. A dudect score is a statistic over wall-clock time on hardware nobody controls, so any threshold over it has a false-failure rate that belongs to the machine rather than to the code. A required check that can fail nondeterministically is worse than no check: it teaches a reader to re-run a red constant-time result until it turns green. An earlier form of this harness did assert on such thresholds, and failed a release gate on `x86_64-darwin` and then passed a re-run of the identical commit (#3070). **So the claims below are measurements, taken deliberately, not properties this suite enforces.** Raising the counts in `mach.lang.ct.probe` and reading the printed table is how they are re-taken, on a quiet machine, by a person who then reads the numbers. Performance and timing work belongs in [mach-bench](https://github.com/briar-systems/mach-bench), which is built for it, rather than in a correctness suite that must be deterministic. What such a measurement can establish is bounded. It works by refutation, because that is all a timing measurement can do: a clean score is consistent with a leak the instrument cannot see, so each mode carries a *planted* leak that must be detected and a constant-time reference that must not separate the way the planted one does. Read separations between the two class means, never `|t|` -- Welch's statistic divides by the sample variance, and concurrent load inflates the variance without moving the means, so `|t|` collapses under load while the leak is plainly still there. Measured at load average 27, six consecutive runs of one binary gave `|t|` between 7.51 and 11.24 against a threshold of 10 while the mean ratio never fell below 3.4. A third probe leaks nothing at all and says whether a run counts at all. Any separation *it* shows is the machine rather than the code, so a run in which it is not flat has no discriminating power and its verdict means nothing in either direction. Note that a flat null is necessary and not sufficient: it must also be sampled at a comparable cost to the probe it is bounding, or a quiet control at one magnitude certifies nothing about noise at another. **The x86-64 inline-asm flags table is measured, not inferred.** The twenty-two-row classification the `#[oblivious]` asm model rests on, naming which instructions *define* ZF and CF, which merely write them, and which read them, is re-derived on x86-64 hosts by `mach.lang.target.isa.x64.probe`, which runs each instruction twice with RFLAGS preset all-set and all-clear and compares what the CPU did against the transcription. `defines_flags` is the only fact that clears a taint, so a row that drifts from the silicon is a permission rather than a refusal. Three rows are exempt and classified by reasoning instead: `popfq`, `iretq` and `syscall` pass the writer probe cleanly and are still not definers, because the flags came from the stack, the interrupt frame, or an existing value masked through `IA32_FMASK`, and where a value came *from* is structural rather than measurable. A non-x86-64 host declines the probe by name, and an extension row runs only on a host whose cpuid reports the extension. **What a timing harness can and cannot assure.** The leakage model has three channels and no single sampling regime covers them (briar-systems/mach#2363): | channel | assured by | |---|---| | control-flow trace | the source-level branch gate; a latency-mode measurement can corroborate | | variable-latency operands | the sema/lowering gates; a latency-mode measurement can corroborate | | memory-address trace | the source-level **secret-index and secret-address gates**; only an address-mode measurement can corroborate | The two modes need opposite sampling and neither substitutes for the other. Latency mode times a large batch of calls per sample, which is what lifts a running-time difference above clock resolution. That same batching *hides* an address-trace leak: every call in a batch is handed the same input, so after the first call both input classes are reading a warm cache line and the single cache miss carrying the signal is averaged away. Measured on one function — a secret-indexed read over a table larger than the last-level cache — latency mode scored |t| ≈ 1–15 across runs (straddling its own threshold, so it neither confirmed nor denied) while address mode, at one call per sample, scored in the hundreds. Two consequences worth stating plainly: - A clean latency-mode number for a table lookup is **not** evidence of address-trace safety. It is the wrong instrument for that channel. - An address-trace leak is only *measurable* when the table exceeds the last-level cache. A cache-resident table leaks its index just as truly and no timing harness will see it. For small tables the property rests entirely on the secret-index gate and on reading the emitted code. **Where the check runs.** The constant-time contract is checked at the **IR level**: the validator runs over the lowered, target-independent MIR before width legalization, instruction selection, register allocation, spilling, frame insertion, and encoding, and trusts those stages to be timing-preserving. What it refuses, it refuses closed: an operation it does not know, an out-of-range register reference, and inline assembly without a complete effect declaration are rejected, never defaulted to public. A proof over the final allocated machine program — after selection, allocation, spills and frame insertion, over physical registers and flags — is planned additive work (#3591), not something this version claims. Where mach does not own the later stages at all — a whole-module emitter such as SPIR-V — the contract is refused rather than assumed. #### See also - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) - the compound type grammar ^ qualifies - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - tagged values and outer-secret ^Tag rules - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) - $is_secret, $holds_secret and the rest of the type-predicate family - [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) - the :: / :~ casts that preserve secrecy and :>T - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) - the #[oblivious] decorator reference - [grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md) - the formal grammar of ^ and :>T Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md ### Operators #### Arithmetic `+` `-` `*` `/` `%` — work on integer and floating-point scalars. On the seeded vector types they apply lane-wise, with the honest per-lane table in [SIMD vectors](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#simd-vectors) below (`+ - * /`, no vector `%`). ```mach val s: i64 = 10 + 20; val q: f32 = 1.5 * 2.0; ``` `%` is the remainder. On integers it is the native truncated remainder, taking the sign of the dividend (`-7 % 3 == -1`). On floats it is the truncated (C `fmod`) remainder `a - trunc(a / b) * b`, likewise taking the sign of the dividend (`5.5 % 3.0 == 2.5`, `-5.5 % 3.0 == -2.5`). For finite operands and a nonzero divisor, this applies across the finite operand range, including quotients beyond the `i64` range. A zero divisor or an infinite dividend gives NaN on every target, as IEEE 754 and C `fmod` do. So does a NaN operand. A finite dividend over an infinite divisor gives the dividend unchanged, its sign and a zero dividend's sign included, as C `fmod` does. Where the result is NaN, only the NaN is defined, not its sign or payload. In a constant expression a zero divisor is refused as division by zero, and the other cases fold to the same results as at run time. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val r: f64 = 5.5 % 3.0; # 2.5 val s: f64 = -5.5 % 3.0; # -2.5 val t: i64 = -7 % 3; # -1 print.printlnf("{} {} {}", r, s, t); ret 0; } ``` **Widening multiply.** A multiply whose operands are both conversions from one narrower integer type to a type exactly twice as wide is the full product of the narrow operands, and compiles to the target's widening instruction rather than a multiply at the wide width. This is how a 64 x 64 product is written at 128 bits, and the two halves a program selects from it are single instructions on every 64-bit target: ```mach fragment val full: u128 = (a::u128) * (b::u128); # a, b: u64; one widening multiply val hi: u64 = (full >> 64)::u64; # the high-half multiply val lo: u64 = full::u64; # the plain multiply ``` Both operands must be the same conversion (both zero-extensions or both sign-extensions); a mixed-sign product is an ordinary multiply at the wide width. See [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#128-bit-integers). **`*` on a secret operand.** `^T * T` and `^T * ^T` are the same wrapping, same-width product with a `^T` result: the operator means the same thing on a secret, and nothing declassifies. What the operand's secrecy changes is whether the target may execute it. A secret `/` or `%` is always refused, and a secret `*` compiles only where the instruction set declares the exact multiply it emits (the low half, a high half or the widening product, at that operand width) as data-independent-timing under a condition the build meets: on x86-64 every scalar cell unconditionally, on aarch64 under PSTATE.DIT on linux and darwin, on riscv64 with `m` and `zkt` selected, and nowhere else. An undeclared cell is a compile error at lowering, never a slower substitute. The widening form above carries through: `(a::^u128) * (b::^u128)` over 64-bit secrets is the 64-bit widening cell, and its halves are `^u64`. The per-instruction-set table and the conditions are in [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#constant-time-multiply-by-instruction-set), and `$mach.build.ct_mul(op, width)` answers the same question at comptime ([comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md)). ##### Half-precision arithmetic `+ - * /` on `f16` give the correctly rounded binary16 result, to nearest with ties to even, on every target. `%` is the truncated remainder of the other float widths, computed on the operands' exact widening to the format the operation runs in; the remainder of two `f16` values is itself an `f16`, so narrowing it back does not round. Unary `-` flips the sign bit, and a comparison relates the exact values, against an `f16` or any other float width ([Comparison](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#comparison)). No operator is added or removed for the width, and a secret `f16` operand is refused in each of them as it is at every float width ([secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md)). ```mach use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val third: f16 = 1.0f16 / 3.0; # 0.333251953125, the nearest binary16 val r: f16 = 5.5f16 % 3.0; # 2.5, exact val lt: u8 = third < 0.3333333333333333f64; # 1: compared exactly, not rounded if (third:~u16 != 0x3555 || r != 2.5 || lt != 1) { ret 1; } ret 0; } ``` The target's own half-precision instructions compute where it has them. Where it does not, the operands are widened exactly, the operation runs once in binary32 or binary64 and the result is narrowed once. binary32 carries 24 significand bits, more than twice binary16's 11 plus two, so that single rounding is already the correctly rounded binary16 result for `+ - * /`; binary64 carries more. A widening or narrowing the target has no instruction for is inline integer code, never a call: | target | `+ - * /` | comparisons | |---|---|---| | x86-64 | binary64, the conversions inline | binary64, the widening inline | | x86-64 with `f16c` (`x86-64-v3` and up) | binary32 through `vcvtph2ps` and `vcvtps2ph` | binary32 through `vcvtph2ps` | | aarch64 | binary64, the operands widened inline and the result narrowed by `fcvt` | binary32 through `fcvt` | | aarch64 with `fp16` | native on the `h` registers (`fadd`, `fsub`, `fmul`, `fdiv`) | native (`fcmp`) | | riscv64, riscv32 | binary64, the conversions inline | binary64, the widening inline | | riscv with `zfhmin` | binary32 through `fcvt.s.h` and `fcvt.h.s` | binary32 through `fcvt.s.h` | | riscv with `zfh` | native (`fadd.h`, `fsub.h`, `fmul.h`, `fdiv.h`) | native (`feq.h`, `flt.h`, `fle.h`) | | spirv with `float16` | native, the core float instructions on `OpTypeFloat 16` | native | | spirv without `float16` | binary32, the conversions inline | binary32, the widening inline | aarch64 without `fp16` computes in binary64 even though `fcvt` converts to binary32, because that conversion quiets a signaling operand and the half unit would give a signaling NaN priority. Where the NaN the operation makes is the target's (see [A NaN between float widths](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#a-nan-between-float-widths)), every path above makes the one the target's native half instruction would. x86-64's native rows under AVX-512 FP16 are #4159. That is proven, not assumed: the exhaustive proof `test/run.sh --f16proof` (#3804) checks every pair of `f16` operands for each operator, and every input of the `f16` conversions, against a reference computed in integers. On x86-64 with and without `x86-64-v3`, aarch64 with and without `fp16`, and riscv64 with and without `zfh` (under qemu-user), every result is bit-identical to it, NaNs by the target's rule, and so is the same operation widened to binary32 by hand and narrowed once. binary32 is enough for all four operators, and none needs binary64 for its rounding. #### Bitwise `&` `|` `^` `~` `<<` `>>` — work on integer scalars. On integer-lane vectors all six apply lane-wise, and a vector shift's count is a vector of the same shape (see [SIMD vectors](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#simd-vectors)). ```mach fragment val x: i64 = (a & b) | (c ^ d); val y: i64 = x << 2; ``` A shift's result has the left operand's type, and its count is any integer type. `<<` shifts zeros in from the right; `>>` on an unsigned operand shifts zeros in from the left and on a signed operand copies the sign bit in. A count at or above the left operand's width **saturates**: `<<` and an unsigned `>>` answer `0`, a signed `>>` answers the sign fill (`0` or `-1`). A count that is a compile-time constant at or above the width, or negative, is a compile error, since a program never means the saturated value by it: ```mach fragment val a: u32 = x << 31; # ok val b: u32 = x << 32; # error: shift count 32 is at least the width of `u32` (32 bits) val c: u32 = x >> n; # n: u8 at run time; 0 when n >= 32 val d: i32 = y >> n; # -1 or 0 when n >= 32, the sign of y ``` The saturation is branch-free, so a secret count admitted by the constant-time gates (see [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md)) stays admitted. A count the compiler proves below the width needs no saturation and compiles to the bare shift: one masked below it, such as `x << (n & 31)` on a `u32`, a loop counter its guard bounds, as `plane` in `for (plane < 8) { x >> plane }`, and a count a dominating compare bounds, as in `if (n < 32) { x << n }`, or sums, differences and products of such values, such as `(i & 3) * 16 + (i >> 2) * 4` under `i < 16`. A count a loop does not change has its range test computed once, before the loop. #### Comparison `==` `!=` `<` `>` `<=` `>=` — produce `u8` (`1` or `0`). Mach has no compiler `bool`; `bool` is a stdlib alias for `u8`. Comparisons relate **mathematical values**, so the result is identical in either operand order: - **integer vs integer** — any signedness and width mix is legal and compares the true values (e.g. a negative `i64` is never equal to, and always less than, any `u64`). Width aliases (`usize`, `isize`) follow their backing type. - **float vs float** — any width mix is legal; the narrower operand widens exactly (`f16` -> `f32` -> `f64`). - **integer vs float** — a compile error; cast one operand explicitly with `::`. An implicit widening would hide `f64` rounding above `2^53`. - **pointer vs pointer** — every one of the six operators accepts two pointer-like operands (a pointer, a `ptr`, a function, or `nil`), whatever their pointee types and whatever their pointees' secrecy. Addresses order as unsigned values of pointer width, and the result is a public `u8` even when both pointees are secret: ordering reveals no more than the `==` beside it, and no address comes back out of it. See [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#comparing-and-ordering-addresses). - **pointer vs integer** — a compile error. Ordering relates two addresses; it is not a route from an address to an integer. On the seeded vector types, a comparison produces a same-shape unsigned **mask** vector (lane-wise) — see [SIMD vectors](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#simd-vectors). `==` / `!=` on an **aggregate value** (a `rec`, `uni`, or whole `tag`) is a compile error. Comparing representations would silently relate padding bytes and unwritten union variants, so no whole-value structural equality is provided. Write an explicit field-wise comparison for records. Comparing pointers to aggregates is unaffected, and the rejection applies to a generic instantiated at an aggregate type as well as to a concrete one. For tagged values, `==` and `!=` are rejected entirely: there is no whole-tag equality, no payload equality and no ordering. Test which case is active with the `sel place.case` expression instead. See [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). #### Logical `&&` `||` `!` — short-circuiting. Operands are `u8` (`0` is false, nonzero is true); the result is `u8` (`1` or `0`). ```mach fragment val ok: u8 = (x > 0) && (y < 100); ``` #### Unary - `-` numeric negation - `~` bitwise NOT (integer) - `!` logical NOT (`u8`) On a float, `-` is the IEEE-754 sign-bit flip and is exact for every operand, so `-0.0` is negative zero — a constant distinct from `0.0`, whether it is folded at comptime or negated at run time. The two compare **equal** (`-0.0 == 0.0` is true), so code that must tell them apart compares bit patterns: `(-0.0):~u64` is `0x8000000000000000`. Note that `0.0 - x` is subtraction, not negation: it yields positive zero for either zero. #### Pointer - `?place` — address-of; produces a pointer to the operand. The operand must be a place: a binding, a field, an element, or a dereference (a field or element reached through a pointer counts). Taking the address of a call result, a literal, a cast, an operator result, or any other temporary is an error naming the operand kind. - `@ptr` — dereference; reads through the pointer. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var x: i64 = 9; var p: *i64 = ?x; @p = 11; # write through val v: i64 = @p; # read through print.printlnf("{}", v); ret 0; } ``` ```mach error cannot take the address of a call result fun g() i64 { ret 1; } fun addresses(x: i64) { val b: *i64 = ?g(); # error: cannot take the address of a call result val c: *i64 = ?42; # error: cannot take the address of a literal val d: *u64 = ?(x::u64); # error: cannot take the address of a cast result val e: *i64 = ?(x + 1); # error: cannot take the address of an operator result } ``` Each of those reads, in full, `cannot take the address of a call result: `?` applies to a place (a binding, a field, an element, or a dereference)`. Bind the temporary to a `var` and take that binding's address. #### Index - `x[i]`: one element of an array, an element through a pointer, or one lane of a vector. A lane index is a comptime constant, and a constant index into an array or a vector is bounds-checked at compile time ([types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#array)). - `x[start, count]`: a **range** of `count` consecutive elements or lanes from `start`. ##### Range `count` is a comptime constant of at least 1, a length and never an end index: there are no absolute ranges, no negative counts and no runtime counts, and each of those is an error at the count. `start` is any index expression, as for `x[i]`. | object | `x[start, count]` is | |---|---| | `TxN` | a `Txcount` vector of those lanes (`v[4, 4]` on an `i16x8` is an `i16x4`), with a constant `start` and a `count` of at least 2 | | `[N]T` | a `[count]T` value | | `*T` | a `[count]T` value, read through the pointer | A constant `start` over an array or a vector keeps the whole range inside it: `start + count` may equal `N` and may not pass it, reported as `range [3, 2] is out of bounds for `[4]i32` of length 4`. A pointer carries no length, so a range through one is not checked. A range is an assignment target: `x[start, count] = value` stores `count` elements or lanes starting at `start`, and `value` has exactly the range's type. A range read is a value, not a view onto the memory, so its address cannot be taken (`cannot take the address of a range`) and it cannot be written into. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var xs: [6]f32 = [6]f32{1.0, 2.0, 3.0, 4.0, 5.0, 6.0}; val p: *f32 = ?xs[0]; val a: f32x4 = p[1, 4]::f32x4; # the four floats at p[1] p[2, 4] = (a * a)::[4]f32; # stored at p[2] val tail: [2]f32 = xs[4, 2]; print.printlnf("{} {} {}", xs[2], tail[0], tail[1]); ret 0; } ``` The comma form cannot be confused with generic arguments: a comma list after a name reads both ways (`f[T, U]` is also a list of type arguments), and name resolution picks by what the name is, exactly as it does for `f[x]` ([grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#postfix)). #### Cast Two postfix cast operators, both written `expr OP Type`: - `expr::Type` — **value conversion**. Resizes integers (sign- or zero-extend, truncate), converts between integer and float (a numeric `CVT`), and is the identity on a same-type operand. Value-preserving where representable. Two vectors with the same lane count convert lane by lane with the scalar rule. An array and a vector (`[N]T` and `TxN`, either way) convert element by element and need the same element type and count, so `[8]i16::i32x4` is an error. Any other pair where either type is nonnumeric needs equal sizes, and the bits are reinterpreted. Constant expressions follow these rules at every nesting depth, including casts through type aliases. - `expr:~Type` — **bit reinterpret**. Reads the operand's exact bits as the target type with no conversion. Legal only when `Type` has the same byte size as the operand's type (a size mismatch is a compile error). The `~` recalls its bitwise heritage, so `:~` reads as "bit cast". Every cast (`::`, `:~` and the `:>` strip cast) is a [postfix](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#postfix), and a prefix operator (`@`, `?`, `-`, `~`, `!`) takes its operand together with the whole postfix chain that follows it ([grammar.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#prefix-atoms-and-unary)). So `@p::T` is `@(p::T)`: it casts the pointer `p` and then dereferences the result. To dereference first and convert the value read, parenthesize the dereference: `(@p)::T`. The same holds for `-x::T`, which is `-(x::T)`. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var x: i64 = -1; val p: *i64 = ?x; val bits: u64 = @p::*u64; # @(p::*u64): retype the pointer, read a u64 val wide: i128 = (@p)::i128; # read the i64, then sign-extend it print.printlnf("{:x} {}", bits, wide); ret 0; } ``` Reading `@p::u64` as dereference-then-convert is refused, because the cast applies to `p` and a `u64` cannot be dereferenced: ```mach error dereference of non-pointer type fun widen(p: *i64) u64 { ret @p::u64; # error: @(p::u64) dereferences a u64 } ``` On two vector types, `::` converts lane by lane: each lane goes through exactly the scalar `::` above, so `i32x4::f32x4` converts every lane numerically and `i8x4::i64x4` sign-extends every lane. Both sides need the same lane count (`i32x4::i64x2` is an error even though the two are the same size), and any pair of lane types is allowed, including equal-size integer and float lanes and a signedness change. A lane converts exactly as its scalar would on the same target, including NaN, the infinities and values outside the destination type. There is no cast between a vector and a scalar. The raw bits of a vector are `:~`, which, like every `:~`, needs only equal byte sizes (`i32x4:~f32x4`, `i32x4:~i64x2`). Between an array and a vector of the same shape, `[N]T` and `TxN`, `::` converts element by element in either direction. The element type and the count must be the same on both sides: `[4]i32::f32x4` and `[8]i16::i32x4` are errors, and a lane type change is a separate vector `::` after the array's lanes are in a vector. An array that lives in memory converted to a vector is one vector load, and a vector converted to an array and stored into a place is one vector store, which makes a [range](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#range) with `::` the load and store idiom: ```mach fragment val a: f32x4 = p[i, 4]::f32x4; # one vector load p[i, 4] = (a * a)::[4]f32; # one vector store ``` The two differ sharply on int<->float. `::` runs a numeric conversion, while `:~` reinterprets the raw bit pattern: ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val some_i64: i64 = -1; val a: u64 = some_i64::u64; # value conversion (resize) val p: *u8 = argv::*u8; # pointer value, retyped val n: u64 = 1.5::u64; # 1 (float -> int conversion) val b: u64 = 1.5:~u64; # 0x3FF8000000000000 (raw IEEE-754 bits) val f: f64 = b:~f64; # 1.5 (bits read back as a float) print.printlnf("{} {:x} {}", n, b, f); ret 0; } ``` ##### `f16` conversions `::` to and from `f16` follows the rule of the other float widths: - **`f16` to `f32` or `f64`** is exact: every binary16 value, subnormals included, is a binary32 and a binary64 value. - **`f32` or `f64` to `f16`** rounds once, to nearest with ties to even, straight from the source (never through binary32 on the way from binary64). A value past the largest finite `f16`, 65504, by half a unit or more rounds to infinity of its sign, and a nonzero value of at most half the smallest subnormal, 2^-25, rounds to zero of its sign. - **An integer to `f16`** rounds to nearest with ties to even, and one of magnitude 65520 or more overflows to infinity. Every integer width converts, `u8` to `u64` and `i8` to `i64`. - **`f16` to an integer** truncates toward zero. A value out of the integer's range, an infinity or a NaN gives the target's result for the same conversion from `f64`, the rule every float width follows. - **`:~`** reads an `f16`'s bits as a `u16` or an `i16` and back, exactly, NaN payloads included. ```mach use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val d: f64 = 0.1; val h: f16 = d::f16; # 0.0999755859375, rounded once val back: f64 = h::f64; # exact: the f16 value itself val n: i32 = (-2.75f16)::i32; # -2, truncated toward zero val big: f16 = 100000::f16; # past 65504: infinity val bits: u16 = h:~u16; # 0x2E66 if (back != 0.0999755859375 || n != -2 || big:~u16 != 0x7C00 || bits != 0x2E66) { ret 1; } ret 0; } ``` The target's own conversion instruction runs where it has one, and it gives the same result: - x86-64 with `f16c`: `vcvtph2ps` and `vcvtps2ph` to and from `f32`; - aarch64: `fcvt` to and from `f32` and `f64` on every aarch64 target, and with `fp16` the integer conversions on the `h` registers (`fcvtzs`, `fcvtzu`, `scvtf`, `ucvtf`); - riscv with `zfhmin`: `fcvt.s.h` and `fcvt.h.s`, and with `d` also `fcvt.d.h` and `fcvt.h.d`; with `zfh` the integer conversions (`fcvt.w.h`, `fcvt.h.w` and their unsigned and, on riscv64, 64-bit forms); - spirv with `float16`: the core float conversions. Every other conversion is inline integer code on the binary16 encoding, joined to the target's binary64 conversion for an integer, never a call. Vectors convert lane by lane with the same rule ([types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#simd-vectors)). ##### A NaN between float widths `::` between two float widths (`f16`, `f32`, `f64`) rounds a number once, to nearest with ties to even, the same on every target. What it makes of a NaN is the target's, because it is what the target's own conversion instruction does: | target | a NaN converted to another float width | |---|---| | x86-64, aarch64 | keeps its sign and the top of its payload, with the quiet bit set | | riscv32, riscv64 | the canonical NaN, positive and quiet with an empty payload | | spirv | the canonical NaN (SPIR-V leaves the payload unspecified, so a device may differ) | A conversion folded at compile time follows the build target's rule, so a folded `::` gives the bits the same `::` gives at run time on that target. An `f16` conversion the target has no instruction for follows the same rule. `:~` never converts, so it reads and writes a NaN's bits exactly on every target. Neither `::` nor `:~` may add or drop the `^` secret qualifier, and neither can erase a secret-welded pointer to `ptr`. Representation-changing `::` and `:~` casts are rejected when either by-value representation contains a tag, including through records, arrays, or union variants. Transparent aliases preserve the tag type. The only secrecy downgrade is the `:>T` strip cast, which removes outer secrecy from `^Tag` without altering the active case or inner payload qualifiers. See [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) and [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). #### SIMD vectors Vector types (see [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md)) carry lane-wise operators at every lane count, not only the 128-bit shapes. **Which operators are legal is target-independent.** The table below is the whole surface, and it is identical on every target and at every width — a `f32x8` add is as legal as a `f32x4` one, and the two differ only in how they are realized. | Lane family | `+` `-` | `*` | `/` | `%` | `& \| ^ ~` | `<< >>` | `== != < > <= >=` | |---|---|---|---|---|---|---|---| | float — `f16x8`, `f32x4`, `f64x2` | yes | yes | yes | no | — | no | → same-shape unsigned mask | | integer — `i8x16` `i16x8` `i32x4` `i64x2` (+ unsigned) | yes | yes | yes | no | yes | yes, by a same-shape count | → same-shape unsigned mask | Both operands of a binary operator must be the **same** vector shape: there is no implicit scalar↔vector mixing and no cross-shape widening. Anything the table marks `no` is a compile error, not a silent fallback: - no vector `%` on any lane type; - bitwise `& | ^ ~` and the shifts `<< >>` require integer lanes; - a shift's count is a vector of the shifted type, never a scalar: `vec << vec` and `vec >> vec` are the only two forms, and `v << 3` is an error that names the form to write instead. A vector shift shifts each lane by the count in the same lane, and each lane follows the scalar operator exactly: `>>` is arithmetic on signed lanes and logical on unsigned ones, a lane count at or above the lane width saturates (`0`, or the sign fill for an arithmetic `>>`), and a lane count that is a compile-time constant at or above the width is an error. A uniform count is the same count in every lane of a literal, which is the form every baseline instruction set shifts by in one instruction: ```mach use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val x: u32x4 = u32x4{1, 2, 0x80000000, 0xF0}; val n: u32 = argc::u32 + 31; # 32 at run time val h: u32x4 = x >> u32x4{4, 4, 4, 4}; # one psrld on x86_64 val k: u32x4 = x << u32x4{n, n, n, n}; # at the lane width: every lane is 0 val m: u32x4 = x << u32x4{0, 1, 2, 3}; # a count per lane if (h[3] != 0x0F || k[2] != 0 || m[1] != 4) { ret 1; } ret 0; } ``` ```mach error a vector shift count is a vector of the shifted type, never a scalar fun f(x: u32x4) u32x4 { ret x << 3; } ``` Integer division uses each lane's signedness and scalar division behavior, including truncation toward zero for signed quotients and the scalar behavior for division by zero or signed overflow. A secret dividend or divisor is rejected because integer division has variable latency. x86_64 and aarch64 realize integer vector division as scalar lane operations, as does RISC-V without a vector unit. ##### Legality is target-independent; realization is not A legal operator means the same thing on every target, but not every target has a packed instruction for every one. The backend picks, in order: the packed form where the target has one, else a **defined unrolled scalar expansion** with lane-identical results (see [policy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md)). Neither choice changes the answer, so nothing about the surface depends on it. The cost does change, so the build says where it is paid. Each operation that falls back to the scalar expansion is one warning at its own site, never one per lane. The warning names the operation, its lanes and the target: ``` warning[vector.scalarize]: vector divide on 4 lanes of 32-bit integers in 'app.main.kernel' scalarizes on x86_64: no packed form for it at any extension ``` Every packed row in a target's catalog declares the extension its instruction needs, and the warning reads those rows. When a row declared under an extension would pack the operation, the warning names that extension in place of "no packed form", for example "declaring `sse41` in the target's `extensions` packs it", and declaring it in [`[target.].extensions`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions) makes the warning go away. Code that is portable on purpose silences the key with `allow = ["vector.scalarize"]` in its profile, a declaration that scalarizes on purpose acknowledges it with [`#[expect("vector.scalarize")]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#expectkey--acknowledge-a-warning), and [`simd = "require"`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#profilename) turns the same sites into errors with the same text. Integer `*` is where this is most visible today: | shape | x86_64 (SSE2) | aarch64 (NEON) | riscv64 (no vector unit) | |---|---|---|---| | `i8x16 * i8x16` | scalar expansion | packed `mul .16b` | scalar expansion | | `i16x8 * i16x8` | packed `pmullw` | packed `mul .8h` | scalar expansion | | `i32x4 * i32x4` | packed `pmuludq` pair (`pmulld` under `sse41`) | packed `mul .4s` | scalar expansion | | `i64x2 * i64x2` | packed `pmuludq` triple (`vpmullq` under `avx512dq` and `avx512vl`) | scalar expansion (NEON has no `.2d` multiply) | scalar expansion | Shifts realize by the count's form: a count that is the same value in every lane shifts every lane by that one scalar, and any other count shifts each lane by its own. | shape | x86_64 (SSE2) | aarch64 (NEON) | riscv64 (no vector unit) | |---|---|---|---| | uniform `<<`, `>>` on 16-, 32- and 64-bit lanes | packed `psll*` / `psrl*` / `psra*` | packed `shl` / `ushr` / `sshr` by a constant, `dup` then `ushl` / `sshl` by a register | scalar expansion | | uniform arithmetic `>>` on 64-bit lanes | scalar expansion (`psraq` is AVX-512VL) | packed, as above | scalar expansion | | uniform `<<`, `>>` on 8-bit lanes | packed through the 16-bit shifts, each byte shifted with its neighbour cleared | packed, as above | scalar expansion | | per-lane count on 32- and 64-bit lanes | packed `vpsllv*` / `vpsrlv*` / `vpsravd` under `avx2`, else scalar expansion (the 64-bit arithmetic `vpsravq` is AVX-512VL) | packed `ushl` / `sshl` | scalar expansion | | per-lane count on 8- and 16-bit lanes | scalar expansion | packed `ushl` / `sshl` | scalar expansion | The x86_64 packed instructions saturate a count at or above the lane width on their own, so they need none of the scalar shift's range test. aarch64's `ushl` and `sshl` read only a count's low byte, as a signed amount that shifts right when negative, so a count is first clamped with `uqshl` and `ushr` to at most 127, which every lane width saturates at, and negated for a right shift. A constant count at or above the width is the fill itself. SPIR-V leaves an `OpShift*` by the component width or more undefined, so it packs the shift beside an `OpULessThan` against the width and an `OpSelect` of 0, or of the shift by the width less one for an arithmetic `>>`, where the count is out of range. A uniform count is splatted by `OpCompositeConstruct`, since `OpShift*` takes a count per component. Operators never widen implicitly, so a widening multiply is spelled as two lane casts and a multiply: `a::i32x4 * b::i32x4` for `a, b: i16x4`. When both operands are extensions of the same narrower vector type, with the same signedness, and the whole product fits one 128-bit register, the backend emits the target's widening multiply for that cell: | operands | x86_64 (SSE2) | aarch64 (NEON) | riscv64 (no vector unit) | |---|---|---|---| | `i8x8` / `u8x8` → 16-bit lanes | extend, then multiply | `smull` / `umull .8b` | extend, then multiply | | `i16x4` / `u16x4` → 32-bit lanes | `pmullw` + `pmulhw` / `pmulhuw` | `smull` / `umull .4h` | extend, then multiply | | `i32x2` → `i64x2` | extend, then multiply (`pmuldq` is SSE4.1) | `smull .2s` | extend, then multiply | | `u32x2` → `u64x2` | `pmuludq` | `umull .2s` | extend, then multiply | A wider product, such as `i16x8` → `i32x8`, is a 256-bit value and keeps the extend-then-multiply path where the vector register is 128 bits. x86-64 under `avx2` holds it in one `ymm` register, and the same widening multiply fills it: `pmullw` beside `pmulhw` over the operands, interleaved by `vpunpcklwd` into the 32-byte product. Either path gives the same lanes. `f16` lanes realize per target as [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#simd-vectors) lists: packed on aarch64 with `fp16` and on spirv with `float16`, through packed `f32` lanes on x86-64 with `f16c` (and for comparisons on aarch64 without `fp16`), and otherwise each lane's scalar `f16` operation ([Half-precision arithmetic](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#half-precision-arithmetic)). Every form gives each lane the bits of the scalar operation. A project that cannot afford a scalar expansion sets `simd = "require"` in its profile (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md)), which turns the shortfall into a build error naming the operation, its lane width, the function and the target. A comparison produces the same-shape **unsigned mask** vector — one lane per input lane, all-ones bits (`0xFF…`) for true and all-zeros for false, exactly what the hardware compare yields. There is no vector-bool type. The mask element is the unsigned integer of the input's lane width: `f32x4` / `i32x4` / `u32x4` → `u32x4`; `f64x2` / `i64x2` → `u64x2`; `f16x8` / `i16x8` → `u16x8`; `i8x16` → `u8x16`. Select/blend is not an operator; it is the library idiom `(mask & a) | (~mask & b)` over matching integer lanes (the tier-3 simd library, #2021). ```mach fragment val a: f32x4 = f32x4{1.0, 2.0, 3.0, 4.0}; val b: f32x4 = f32x4{4.0, 3.0, 2.0, 1.0}; val sum: f32x4 = a + b; # lane-wise -> {5.0, 5.0, 5.0, 5.0} val mask: u32x4 = a < b; # -> {0xFFFFFFFF, 0xFFFFFFFF, 0, 0} val m: i32x4 = i32x4{1, 2, 3, 4}; val n: i32x4 = i32x4{4, 3, 2, 1}; val z: i32x4 = (m & n) ^ n; # lane-wise bitwise on integer lanes ``` #### An operand typed by a type parameter Inside a generic, an operand whose type is a parameter has no operand class yet: `T` is not an integer, not a float and not a pointer, and asking would be asking about a placeholder. Every operator on such an operand is decided at the instantiation instead, against that instance's concrete type, and the table above is what it is checked against there. A type that does not support the operator is refused at the instantiation that asked for it. See [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md). #### See also - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) — how operators compose into expressions - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — an operator on a generic type parameter - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) — which types support which operators - [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md) — the `^` secret qualifier and the `:>T` strip cast Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md ### Expressions Expressions evaluate to values. They appear on the right side of bindings, as conditions, and as call arguments. Reading an aggregate captures its value at that evaluation point. A later argument, assignment destination expression, or `fin` body cannot change the captured value by modifying its original storage. Call arguments evaluate left to right. Assignment evaluates and captures the right side before evaluating the destination on the left side. #### Literals See [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md) — numeric, char, string, and `nil` forms. #### Names A bare identifier references a name in scope. Module-qualified names use the dot path: ```mach fragment counter # local or module-level binding core.add # symbol from module `core` ``` #### Record / array / union / tag literals A type name followed by a brace-delimited initializer: ```mach use std.types.error.err; use std.types.result.res; rec Point { x: i64; y: i64; } uni Number { i: i64; f: f64; } rec Pair[T, U] { left: T; right: U; } tag Reply: u8 { empty; value: i64; } tag MyErr: u8 { bad; } val p: Point = Point{x: 1, y: 2}; val a: [3]i64 = [3]i64{10, 20, 30}; val u: Number = Number{i: 99}; val pair: Pair[i64, u8] = Pair[i64, u8]{left: 5, right: 6u8}; val rep0: Reply = Reply.empty{}; val rep1: Reply = Reply.value{42}; val good: res[i64, MyErr] = res[i64, MyErr].ok{42}; val done: err[MyErr] = err[MyErr].ok{}; ``` For generics, the type arguments appear in brackets before the body. A tag value names its type and its case, and carries a positional payload only when the case declares one. Omitting a required payload, supplying a payload to a payloadless case, supplying more than one, naming the payload, or writing the tag type before a brace (`Type{...}`, which names no case) is a compile error. Vector literals (`f32x4{ 1.0, 2.0, 3.0, 4.0 }`) follow the same brace shape, but require one positional initializer per lane. See [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#simd-vectors). #### Field, index, and tag access ```mach use std.print; use std.runtime; rec Point { x: i64; y: i64; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val p: Point = Point{x: 1, y: 2}; val a: [3]i64 = [3]i64{10, 20, 30}; val x: i64 = p.x; # record field val first: i64 = a[0]; # array index print.printlnf("{} {}", x, first); ret 0; } ``` An index the compiler can fold is bounds-checked against a statically known length, such as a fixed array length `N` or a vector lane count. See [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md). For tagged values: - `sel tag_val.case` tests whether that case is currently selected. It reads only the discriminator and is an ordinary `bool`. - `tag_val.case` accesses the payload of that case. It is legal only inside a lexical guard for that place and case; an unguarded payload access is a compile error. - `TypeName.case` alone is a case selector, not a value. It cannot be stored or passed. See [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) for the guard rules. There is no other operator over a tag: failure handling is an ordinary `if`/`or` chain over `sel`, and the guard it opens is what makes the payload readable. #### Function calls ```mach fragment add(2, 3) identity[i64](42) # generic call: type args in [ ] sum(3, 10i64, 20i64, 30i64) # variadic pack call (see variadics.md) ``` A call to a pack-tailed function is monomorphized per distinct trailing type-list; `g(va...)` forwards a whole pack — see [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md). For comptime parameters, the value is passed positionally like a runtime argument — the function signature determines whether it must be comptime: ```mach fragment checked_add(MODE_FAST, 1, 2) # MODE_FAST is comptime-knowable ``` #### Operators See [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md). Operators combine expressions into larger expressions; precedence follows the usual C-family conventions. #### See also - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) — how expressions appear inside statements - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — function declarations and signatures Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md ### Statements Statement forms compose into function bodies and blocks. Statements end with `;` except where they end with a block `{...}`. #### `if` / `or` ```mach fragment if (cond) { ... } or (cond) { ... } or { ... # final else (no condition) } ``` - The `if` head opens the chain. - Each `or (cond) { ... }` adds another branch. - A trailing `or { ... }` is the catch-all. - Bodies are blocks; there is no one-statement-without-braces form. An `if`/`or` arm whose condition is exactly `sel P.c` guards the payload place `P.c` inside its block: ```mach tag Reply: u8 { empty; value: i64; } fun read(reply: Reply) i64 { if (sel reply.value) { ret reply.value; # guarded by the arm condition } or { ret 0; # reply holds empty on this path } } ``` A guard is a lexical region, not a flow fact. A chain whose every arm exits guards the remainder of the enclosing block for the case the chain left untested. See [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). #### `for` A single condition-loop form. There is no for-each. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var i: i64 = 0; for (i < 10) { i = i + 1; } print.printlnf("{}", i); ret 0; } ``` A `for` with no condition loops until a `brk` or a `ret` leaves it: ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var i: i64 = 0; for { i = i + 1; if (i == 3) { brk; } } print.printlnf("{}", i); ret 0; } ``` #### `ret` ```mach fragment ret expr; # return a value ret; # return from a void function ``` #### `brk` / `cnt` Loop control: `brk` exits the enclosing `for`; `cnt` continues to the next iteration. ```mach use std.print; use std.runtime; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { var i: i64 = 0; for (i < 10) { i = i + 1; if (i == 3) { cnt; } print.printf("{}", i); if (i == 8) { brk; } print.print(" "); } print.println(""); ret 0; } ``` Both are operand-less, so they are keywords only in their bare `brk;` / `cnt;` form. The same word followed by anything else is an ordinary identifier — a variable named `cnt` reads and assigns normally (`cnt = x;`), even inside a loop that also uses bare `cnt;` for control flow. #### `fin` — deferred block `fin` schedules a block to execute when its enclosing block exits, in reverse order of declaration. Useful for cleanup that should happen regardless of how the scope exits. ```mach use std.print; use std.runtime; var counter: i64 = 0; #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { { fin { counter = counter - 1; } fin { counter = counter * 2; } counter = 5; print.printf("{} ", counter); } # at block exit, in reverse order: counter * 2 runs first, then counter - 1 print.printlnf("{}", counter); ret 0; } ``` `fin` is block-scoped: it belongs to the block that declares it and covers only the statements after its declaration. A function body is a block, so the classic pattern — a `fin` at the top of a function running at every return — is the block rule applied to the outermost scope. - A **normal block exit** replays the block's fins, then execution continues after the block. - A **loop body** is a block: its fins replay at the end of every iteration, with that iteration's values. - **`brk` / `cnt`** replay the fins of every scope being exited — everything down through the targeted loop's body, innermost first — before branching. The loop body's fins replay under `cnt` too: the iteration is ending. - **`ret expr;`** fully evaluates the return expression first, then replays the fins of every open scope (innermost first), then returns the already-evaluated value — a fin's side effects are never observable in the returned value. A bare `ret` likewise, minus the expression. A `fin` declared inside another fin's body belongs to that body's block and replays when the body finishes. Control flow cannot cross a `fin` boundary outward: a `ret` inside a fin body is a compile error, as is a `brk` / `cnt` whose target loop encloses the fin. A loop fully inside the fin body uses `brk` / `cnt` normally. `fin` requires a block body (`fin { ... }`). The bare single-statement form (`fin stmt;`) is rejected, and so is a `ret` inside a `fin` body: ```mach error `ret` cannot appear inside a `fin` body fun leave() i64 { fin { ret 1; } ret 0; } ``` #### Block `{ ... }` introduces a new lexical scope. Statements inside are evaluated in order. Blocks can stand alone: ```mach fragment { val tmp: i64 = compute(); use_tmp(tmp); } ``` #### Expression statements An expression followed by a semicolon executes as a statement: ```mach fragment compute(); ``` Assignment is an expression (`x = y;` is an expression statement whose top operator is `=`), and so is a call whose result is discarded. #### Failure handling A function that can fail returns a tag. The caller tests the case with `sel` and exits the arm that handles the failure; the exiting chain guards the success payload for the rest of the block: ```mach use std.types.error.err; use std.types.result.res; tag WriteError: u8 { closed; full; } fun flush() err[WriteError] { ret err[WriteError].ok{}; } fun parse(input: u8) res[i64, WriteError] { ret res[i64, WriteError].ok{input::i64}; } fun increment(input: u8) res[i64, WriteError] { val flushed: err[WriteError] = flush(); if (sel flushed.err) { ret res[i64, WriteError].err{flushed.err}; } val r: res[i64, WriteError] = parse(input); if (sel r.err) { ret res[i64, WriteError].err{r.err}; } ret res[i64, WriteError].ok{r.ok + 1}; # r.ok is guarded: the chain above exits } ``` Every reachable path through the failure arm must leave the block, with `ret`, or with `brk` or `cnt` targeting a loop that encloses the chain; an arm that falls through opens no guard, and the payload read after it is rejected. See [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) for the guard rules. #### See also - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) - expressions and literals - [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) - tagged values, `sel` and guards - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) - the comptime counterpart Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime.md ### Comptime channel The `$` prefix opens the **comptime channel** — the compiler-owned namespace a program reads at compile time. It is read-only: `$` *selects and expands*, it never executes or mutates. Everything that touches comptime in Mach — conditional compilation, intrinsics, target queries — uses one of the shapes on this channel. #### The shapes | Shape | Meaning | Direction | |---|---|---| | `$mach.*` / `$project.*` / `$bin.*` | Rooted compiler-owned read | compiler → developer | | `$sym(args)` | Comptime function call (intrinsic) | call | | `$if`, `$or` | Comptime control flow | structural | > Per-declaration codegen attributes (symbol rename, library pin, inline, > align, section) are written as **`#[...]` decorators**, not `$`-comptime > shapes — see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md). A comptime directive takes no > `=`; a stray one is a parse error at the directive's terminator. The parser distinguishes these by structure: - `$.`, where `` is one of the reserved roots `mach`, `project`, `bin` — a read into a compiler-owned tree. The roots are reserved at the top of `$`; user symbols cannot collide with them. - `$ident(args)` — comptime call; the closed compiler-intrinsic set lives here. - `$if` / `$or` — comptime branches, structurally distinct from runtime `if` / `or`. #### Compiler-owned roots | Root | Reads | Source | |---|---|---| | `$mach.*` | resolved build os/arch/abi/mode tags, pointer width, compiler identity | active build + compiler | | `$mach.project.{id,version}` | the **owning** project's metadata | `[project]` in the `mach.toml` of the project that owns the module | | `$mach.project.version.{major,minor,patch}` | its structured version components | that `[project].version` | | `$mach.source.{file,line,module}` | the module's project-relative file, the line, the module name | the module | | `$project.{id,version}` | the **root** project's metadata | `[project]` in the root `mach.toml` | | `$project.version.{major,minor,patch}` | structured version components | the root's `[project].version` | | `$project.target.{os,arch,abi}` | the selected target's declared tuple, as strings | the selected `mach.toml` target | | `$bin.name` | the artifact being built | the selected build unit (`[artifact.*]`) | Two roots read a project's identity, and they differ in which project: - `$project.*` is the **root**: the project being built, the same in every module of the build, dependencies' modules included. A dependency built as another artifact's requirement is the root of that build. - `$mach.project.*` is the **owner**: the project whose source tree holds the module being compiled, read from its own manifest. It follows the module, as `{artifact..out}` does ([manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md)). In a root module the two read the same manifest. In a dependency's module `$project.version` is the consumer's version and `$mach.project.version` the dependency's own, so a library that reports its version reads `$mach.project.version`. `$project.target.*` carries the manifest's declared **string** spellings (`"linux"`, `"x86_64"`, `"sysv64"`), distinct from `$mach.build.*`'s numeric tags used for `$mach.{os,arch,abi}.*` comparison. Flat `$project.version` is the whole version **string** (`"2.0.0"`); the structured `$project.version.{major,minor, patch}` folds its integer components — both are available. `[project]` has exactly the keys `id`, `version`, `mach`, `src`, and `out` ([manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#project)), and `$project.*` carries exactly `id` and `version`. A path the root does not carry, `$project.name` and `$project.description` included, reports `` unknown `$project.*` path `` at the path. See [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md) for the `$mach.*` subtree. `$bin.name` is the table key of the artifact the module is compiled for. A build compiles one artifact, and a test build tests one, so every module reads that artifact's key. An editor session analyzes the project as the union of every artifact, and there a module reads the artifact whose walk reached it first: the primary artifact when its walk reaches the module, otherwise the first artifact in manifest order that does; a module first reached behind a gate decided later reads its importer's. ```mach fragment val ver: str = $project.version; # "2.0.0", from [project].version $if ($project.target.os == "windows") { ... } # the declared os string ``` #### The bare `$ident` is rejected A standalone bare `$ident` — `$mode`, `$foo` — is **none** of the shapes above and is rejected with one teaching diagnostic, owned by the comptime evaluator: > comptime parameters are referenced without `$`; comptime paths are rooted: > `` $mach ``, `` $project ``, `` $bin `` A comptime **parameter** is referenced by its bare name (no `$`); every comptime **path** is rooted. The rule applies identically in a `$if` gate and in value position — sema and lowering both defer to the evaluator's single verdict rather than each carrying their own. #### Which binding a name denotes An identifier in the comptime channel denotes the declaration it resolves to, under the ordinary scoping rules — never whichever binding happens to share its spelling. A block-scoped binding shadows an outer one of the same name here exactly as it does at runtime: ```mach error array length is not a comptime constant val N: i64 = 9; fun f(k: i64) i64 { val N: i64 = k; # shadows the module constant var xs: [N]i64; # error: array length is not a comptime constant ret xs[0]; } ``` The inner `N` is decided at runtime, so it has no comptime value, and the constant it shadows is not what the expression names. Reading such a name in a comptime position is refused: > identifier names a runtime binding, so it has no comptime value A binding marked `$` — a comptime value parameter, an `$each` loop variable — *is* a comptime binding, and shadows an outer name of its own spelling in the same way. Inside the `$each` below the name `N` is the element, not the 9: ```mach use std.print; use std.runtime; val N: i64 = 9; val ES: [2]i64 = [2]i64{1, 2}; fun g() i64 { var s: i64 = 0; $each N in ES { s = s + N; } # 3, not 18 ret s; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", g()); ret 0; } ``` #### What's not in the channel - No reflection-via-`$.*` subtree. Types are not first-class comptime values. - No decl-attached prefix sugar (`$inline pub fun ...` does not exist) — use `#[...]` decorators (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)). - No comptime function definitions, and no comptime *execution* — the channel selects and expands, it never runs a loop or mutates state. `$each` is not a counterexample: it splices its body once per element of a fixed comptime sequence (a variadic pack, `$fields(T)`, or a constant array `val`), a bounded structural expansion resolved at compile time, not an iterated computation. - No compile-time evaluation of user functions, ever: a call in a `val` initializer is a runtime call. A lookup table is committed as generator output next to its generator and held to it by a `test` (equality with the generator when it is deterministic, the property it searched for when it is randomised, and no startup fallback that repairs a bad constant before the test sees it), or built at startup into a module-private `var`; a module-scope `val` lands in read-only data. - No bare `$ident` — see above. #### See also - [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md) — the `$mach.*` namespace - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — `#[...]` codegen decorators - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) — `$size_of`, `$is_record`, `$type_name`, … - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` / `$or` Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md ### `$mach.*` — compiler-owned namespace The `$mach.*` subtree is the compiler's view of the world: the resolved build context, the compiler identity, the project that owns the module, and source position. All reads, all comptime constants. The tags `$mach.{os,arch,abi,mode}.*` exist for path-value comparison against the resolved-build facts. Every path is deterministic: a function of the source, the manifests and the selected build, never of the clock, the host machine or the checkout. The same source built at a different path, at a different time or on a different machine reads the same values, so a module's compiled output can be cached and reused. There is no build timestamp, host name or version-control state. A build that wants one generates a source file from a [build step](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps). #### Subtrees ##### `$mach.build.*` — what we're building for The resolved active build's facts. `os`/`arch`/`abi`/`mode` share the numeric tag space of `$mach.{os,arch,abi,mode}.*`, so a comparison is a plain integer compare (see [Comparison](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md#comparison)). ```mach fragment $mach.build.os # live; compared against $mach.os.* tags $mach.build.arch # live; compared against $mach.arch.* tags $mach.build.abi # live; compared against $mach.abi.* tags $mach.build.pointer_width # live; integer count of bytes $mach.build.mode # live; compared against $mach.mode.* tags $mach.build.pie # live; 1 when building position-independent, else 0 $mach.build.platform # live; the target's open platform tag as a string, "" when unset $mach.build.ct_mul(op, width) # live; 1 when a secret multiply of that cell is admitted, else 0 $mach.build.extensions. # live; 1 when the target selects that instruction-set extension, else 0 ``` The members above are the whole subtree, and no manifest key adds one. The `extensions` members are the selected isa's own vocabulary, not the manifest's. A `$mach.build.` that names none of them is a compile error at the use site (`` unknown `$mach.*` path ``). A project's own configuration constants are ordinary `val`s selected with `$if` over the facts above. ```mach error unknown `$mach.*` path val TRACING: u64 = $mach.build.TRACING; ``` ###### `$mach.build.ct_mul(op, width)` — the constant-time multiply catalog The one call-shaped fact. It folds to 1 exactly when the selected target admits a secret-operand multiply of that cell, and to 0 otherwise. The answer comes from the same decision the lowering gate and the `#[oblivious]` validators make, so a library can choose its hardware or bit-serial path without keeping a per-target list: ```mach $if ($mach.build.ct_mul(wide_u, 64) == 1) { # one widening multiply per limb product } $or { # the bit-serial product } ``` - `op` is a bare word, one of: - `low`: the low half of a same-width product; - `high_u`, `high_s`, `high_su`: the high half, with unsigned, signed or mixed operands; - `wide_u`, `wide_s`: the full double-width product. - `width` is the operand width in bits, a comptime integer: 8, 16, 32, 64 or 128. - A cell is admitted when the target declares it and the row's condition holds: - an always-safe instruction; - the extension it names, selected for the target; - a data-independent-timing mode the target guarantees. - riscv64 with Zkt selected (`rv64gc_zkt`) admits `low` at every width and the three high halves at 64. x86-64 admits `low`, `high_u`, `high_s` and `wide_u` / `wide_s` at every width on every OS. aarch64 declares `low` at every width, `high_u` and `high_s` at 64 and `wide_u` / `wide_s` at 32 under PSTATE.DIT, which linux and darwin declare they guarantee, so the query folds to 1 there and to 0 on aarch64-windows and freestanding aarch64 (#3508, see [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#pstatedit-at-run-time)). Every other target folds to 0. Lane multiplies are not part of the query. - The result is a `u8`, like `$mach.build.pie`. An unknown `op` or `width`, a missing argument, or arguments on any other path is a compile error that names what is accepted. ###### `$mach.build.extensions.` — instruction-set extensions One `u8` member per extension name of the selected isa, like `$mach.build.pie`. It is 1 when the target selects that extension and 0 otherwise. A build selects extensions with the target's `extensions` key and, on riscv, its isa string, closed over what each one implies (see [Instruction-set extensions](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions)), so `extensions = ["sse41"]` answers 1 for `ssse3` too. On spirv the target's `env` selects them. The names are the selected isa's vocabulary and nothing else: - `x86_64`: `ssse3`, `sse41`, `sha`, `fsgsbase`, `popcnt`, `lzcnt`, `bmi1`, `sse42`, `cx16`, `avx`, `avx2`, `bmi2`, `fma`, `movbe`, `f16c`, `avx512f`, `avx512bw`, `avx512cd`, `avx512dq`, `avx512vl`, `aes`, `pclmul`; - `aarch64`: `sha2`, `sb`, `aes`, `pmull`, `fp16`; - `riscv64` and `riscv32`: `i`, `m`, `a`, `f`, `d`, `c`, `zicond`, `zicsr`, `zifencei`, `zfhmin`, `zfh`, `zkt`; - `spirv`: `float16`, `zero_init_workgroup`, `storage_read_without_format`, `storage_write_without_format`, `subgroup_arithmetic`, `subgroup_clustered`, `subgroup_vote`, `subgroup_ballot`, `subgroup_shuffle`, `subgroup_shuffle_relative`, `subgroup_quad`, `subgroup_graphics_stages`, `buffer_int64_atomics`, `shared_int64_atomics`, `buffer_float32_atomics`, `buffer_float32_atomic_add`, `buffer_float32_atomic_min_max`, `buffer_float64_atomics`, `buffer_float64_atomic_add`, `buffer_float64_atomic_min_max`, `shared_float32_atomics`, `shared_float32_atomic_add`, `shared_float32_atomic_min_max`, `shared_float64_atomics`, `shared_float64_atomic_add`, `shared_float64_atomic_min_max`, `storage_image_multisample`, `resource_min_lod`, `image_gather_extended`, `maintenance8`, `image_int64_atomics`, `image_float32_atomics`, `image_float32_atomic_add`, `image_float32_atomic_min_max`, `vulkan_memory_model`, `vulkan_memory_model_device_scope`, `int8`, `int16`, `buffer_device_address`, `int64`, `float64`, `buffer_float16_atomics`, `buffer_float16_atomic_add`, `buffer_float16_atomic_min_max`, `shared_float16_atomics`, `shared_float16_atomic_add`, `shared_float16_atomic_min_max`, `storage_buffer_16bit_access`, `uniform_and_storage_buffer_16bit_access`, `storage_push_constant16`, `storage_input_output16`, `storage_buffer_8bit_access`, `uniform_and_storage_buffer_8bit_access`, `storage_push_constant8`. A name the selected isa does not declare is a compile error, never a silent 0, as `$mach.arch.*` refuses an unknown architecture: ``` `$mach.build.extensions.sha`: `sha` is not an extension of isa 'aarch64'; its extensions are: sha2, sb, aes, pmull, fp16 ``` So a source that serves several isas nests the extension question under an architecture guard. `&&` does not stand in for the nesting: every `$mach` path in a condition is checked on its own, so `$if ($mach.build.arch == $mach.arch.x86_64 && $mach.build.extensions.sha == 1)` is refused on aarch64 although the left side is false. Write the nested form: ```mach fragment $if ($mach.build.arch == $mach.arch.x86_64) { $if ($mach.build.extensions.sha == 1) { use backend: std.crypto.hash.sha256.x86_sha; } $or { use backend: std.crypto.hash.sha256.portable; } } $or ($mach.build.arch == $mach.arch.aarch64) { $if ($mach.build.extensions.sha2 == 1) { use backend: std.crypto.hash.sha256.arm_sha2; } $or { use backend: std.crypto.hash.sha256.portable; } } $or { use backend: std.crypto.hash.sha256.portable; } ``` A bare `$mach.build.extensions` is refused too. In an editor union build each target tuple answers for its own target, as `$mach.build.os` does. The member answers for the whole build. A function that uses an extension the target does not select, behind a run-time check, is marked [`#[extensions(...)]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#extensionsnames--an-outlier-function) instead, and the member stays 0 for it. ##### `$mach.version` — the compiler version ```mach fragment $mach.version # live; the version string, e.g. "2.0.0" $mach.version.major # live; integer component $mach.version.minor # live; integer component $mach.version.patch # live; integer component ``` ##### `$mach.compiler.*` — compiler identity ```mach fragment $mach.compiler.name # live $mach.compiler.version # live; same value as $mach.version ``` ##### `$mach.project.*` — the project that owns the module ```mach fragment $mach.project.id # the owning project's [project].id $mach.project.version # the owning project's [project].version string $mach.project.version.major # integer component $mach.project.version.minor # integer component $mach.project.version.patch # integer component ``` The project that owns the module being compiled, read from that project's own `mach.toml`. A module of the root project reads the root's manifest. A module of a dependency reads the dependency's, whichever project is being built. The top-level [`$project.*`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime.md#compiler-owned-roots) root is the other half: it names the project the build is for, the same in every module. In a root module the two agree. A library reports its own version with it, and the value stays right in every consumer's build: ```mach use std.types.string.str; pub val VERSION: str = $mach.project.version; ``` `$project.version` in the same place would read each consumer's version instead. A member the subtree does not carry is `` unknown `$mach.project.*` path ``. ##### `$mach.source.*` — current source position ```mach fragment $mach.source.file # the module's file, relative to its project's root: "src/lib/hedge.mach" $mach.source.line # the 1-based line the path is written on $mach.source.module # the module's fully qualified name: "hedge.lib.hedge" ``` `file` is relative to the root of the project that owns the module and never an absolute host path. It is `/`-separated on every host, Windows included, so a checkout at another path or on another host reads the same value. `line` is the line of the `$mach.source.line` read itself. `module` is the name a `use` imports the module by. ```mach val LINE: u64 = $mach.source.line; ``` A member the subtree does not carry is `` unknown `$mach.source.*` path ``. ##### `$mach.os.*`, `$mach.arch.*`, `$mach.abi.*`, `$mach.mode.*` — tag values ```mach fragment $mach.os.linux $mach.os.darwin $mach.os.windows $mach.os.freestanding # no OS / bare metal $mach.arch.x86_64 $mach.arch.aarch64 $mach.arch.riscv64 $mach.arch.riscv32 $mach.arch.spirv $mach.abi.sysv64 $mach.abi.win64 $mach.abi.aapcs64 $mach.abi.lp64 $mach.abi.lp64f $mach.abi.lp64d $mach.abi.ilp32 $mach.abi.ilp32f $mach.abi.ilp32d $mach.mode.debug $mach.mode.release ``` The tag names are the target registries' own spellings, read from them directly; there is no second list to keep in step. A tag name the registry does not carry is a compile error, never a silent fold. The x86-64 System V ABI is spelled `sysv64`, as the registry spells it; `$mach.abi.sysv` is an unknown tag, as is any other name the registry does not carry. #### Comparison Tag comparisons are path-value — no `.id` suffix or unwrapping. Both sides share one numeric space, so the comparison is an ordinary integer compare: ```mach fragment $if ($mach.build.os == $mach.os.linux) { ... } $if ($mach.build.arch == $mach.arch.x86_64) { ... } ``` #### Use in runtime values A `$mach.*` read can initialize a runtime binding. The compiler folds the RHS at compile time: ```mach use std.types.string.str; pub val IS_LINUX: u8 = $mach.build.os == $mach.os.linux; pub val COMPILER: *u8 = $mach.compiler.name; pub val VERSION: str = $mach.version; pub val MAJOR: u64 = $mach.version.major; pub val WIDTH: u64 = $mach.build.pointer_width; ``` #### See also - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` / `$or` using these reads - [comptime.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime.md) — the `$project.*` / `$bin.*` roots, and which project each root names - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) — binding compiler values into runtime constants Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md ### Intrinsics Intrinsics are compiler-shipped comptime functions. They have the same syntactic shape as user-defined function calls (`$name(args)`), but their names are reserved and their implementations are built into the compiler. The set is closed; adding a new intrinsic requires a compiler change. #### Value intrinsics Return comptime constant unsigned integers. The storage type is whatever the binding declares — Mach has no compiler-known `usize`. A layout measurement is an **untyped comptime integer**, exactly like an integer literal: it names a constant the compiler already knows, so putting it in a narrower binding is a choice of representation rather than a conversion. It adopts the binding's width and signedness, and is **refused** — never truncated — when the measured value does not fit: ```mach error value 300 is out of range for u8 use std.types.size.usize; rec Point { x: i64; y: i64; } val A: u8 = $size_of(Point); # 16, stored in one byte val B: i64 = $size_of(Point); # the same 16, stored in eight val C: usize = $size_of(usize); # correct at any pointer width rec Huge { a: [300]u8; } val D: u8 = $size_of(Huge); # error: value 300 is out of range for u8 (0..255) ``` Without a binding to read a width from — an array length, an `#[align(...)]` argument, a comparison against a typed value — the measurement behaves the way a literal does in the same position. ```mach fragment $size_of(T) # byte size of type T $length_of(T) # ELEMENT count of type T $align_of(T) # byte alignment of type T $offset_of(T, field) # byte offset of T's field ``` ```mach fragment pub val POINT_SIZE: i64 = $size_of(Point); pub val POINT_X: i64 = $offset_of(Point, x); ``` `T` is a **type**, written with the ordinary type grammar — not just a bare name. A generic instance, a pointer, an array, a `^` secret, and a qualified `module.Type` are all valid, including inside the generic that owns the parameter: ```mach fragment fun probe[T]() u64 { ret $size_of(Box[T]) + $align_of(Pair[T, u64]) + $size_of(*T) + $size_of([4]T); } ``` `$offset_of`'s **second** argument is the exception: a bare field name, or a tag's payload case name, resolved against the aggregate layout, never a type or a value. `$offset_of` adopts its binding's width like the other three, and answers from the same checked layout under the same complete-type rules as size and alignment (see [Where a layout intrinsic is constant](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md#where-a-layout-intrinsic-is-constant)). An unresolved or recursive layout is diagnosed, never guessed. ##### `$length_of` — elements, not bytes `$size_of` counts bytes and `$length_of` counts elements. For `[N]u8` the two answers are equal; for everything else they are not, and the difference is deliberately in the surface rather than in the caller's head — making a caller divide by an element size is exactly the silent-arithmetic error the `u8` case hides during development. ```mach fragment val PIXELS: [400]f32x4 = ...; $length_of(PIXELS) # 400 — elements $size_of(PIXELS) # 6400 — bytes $length_of(f32x4) # 4 — a vector's lane count ``` **Only a fixed array and a vector have an answer**, and everything else is refused rather than guessed: a pointer (`str` included) has a length the compiler does not know, and a record has a field count rather than an element count. The refusal names the type it was handed. `^` **is** stripped, which puts `$length_of` with `$size_of` rather than with the shape predicates: how many elements are stored is a storage question, and `^[4]u8` stores four of them. ##### The operand may be a binding Every intrinsic taking a type operand also accepts a **value binding** there, denoting that binding's type. A binding's type has no spelling — `val LOGO: [_]u8` really is a concrete `[7194]u8` that cannot be written — so without this a program can index an embed and pass it around and never learn its length: ```mach fragment #[embed("assets/logo.qoi")] val LOGO: [_]u8; $length_of(LOGO) # 7194 — elements $size_of(LOGO) # 7194 — bytes var i: u64 = 0; for (i < $length_of(LOGO)) { ... } ``` This is the operand slot's rule, not a per-intrinsic one, and it is **only** that slot. A name written where a type is expected and resolving to a value is still a mistake in every other position, and accepting it there would turn that mistake into a silently-typed binding. ##### Where a layout intrinsic is constant `$size_of`, `$length_of`, `$align_of`, `$offset_of` and `$type_id` fold in every **type** position, including ones resolved before layout would otherwise be known — the measured type's layout is established on demand when the measurement asks for it, so where the type is *declared* relative to where it is measured makes no difference: | position | `$size_of` / `$length_of` / `$align_of` / `$offset_of` / `$type_id` | |---|---| | `val` / `var` initializer | yes | | global `align` | yes | | record / union type `align` | yes | | array length `[N]T` | yes | | `$if` / `$or` condition, in a function body | yes | | `$if` / `$or` condition, in declaration scope | only when no arm of the chain declares anything | ```mach rec Pair { a: u64; b: u64; } #[align($align_of(Pair))] # a type's alignment rec Over { x: u8; } rec Holder { buf: [$size_of(Pair)]u8; } # an array length, inside a field type ``` `$offset_of` and a descriptor's `.offset` answer from the same checked layout that sizes the type, under the same complete-type rules as `$size_of`. A layout the walk cannot determine is reported, never guessed. A `$if` / `$or` condition is not a type position, so what it can measure depends on when the gate is decided. A gate in a function body, and a gate in declaration scope no arm of whose chain declares anything, are both decided during type checking and measure normally. A gate in declaration scope some arm of whose chain declares something is decided earlier, while names are being resolved and before any type is laid out, so a layout intrinsic there is rejected and reports why. See [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md). **Inside a generic, a predicate is answered per instantiation.** `$is_record(T)` in a `fun f[T]()` body is not decided against the template's placeholder — the gate is deferred and re-folded once for each instantiation with `T` substituted, so `f[SomeRecord]` and `f[u64]` take different arms from one template. A template that is never instantiated has no instance to answer for and reports nothing. **Cycles are refused, not resolved.** A measurement whose answer is one of its own inputs — `#[align($size_of(Self))]`, or two types each aligned to the other's size — is reported as a layout cycle naming the type that closes it. A pointer field does not create one: it stores an address of fixed width, so `rec Node { next: *Node; }` measures normally. #### Type predicates Ask about a type's **shape** rather than its storage. Each takes one type operand and folds in a `$if` / `$or` gate: ```mach fragment $is_record(T) # T is a record (or an instance of one) $is_union(T) # T is a union (or an instance of one) $is_tag(T) # T is a tagged value (or an instance of one) $is_pointer(T) # T is a reference: the raw `ptr` or a typed `*U` $is_integer(T) # T is an integer: i8..i64, u8..u64 $is_float(T) # T is a float: f32 or f64 $is_secret(T) # T is `^`-qualified at the outermost level $holds_secret(T) # any byte of T is secret ``` `$is_integer` and `$is_float` are the scalar half of the family, and they exist so that a walk can ask what it actually wants to know instead of spelling the scalar types out: ```mach fragment $each f in $fields(T) { $if ($is_integer(f.type) || $is_float(f.type)) { ... } # a scalar field $or ($is_record(f.type)) { ... } # descend } ``` A ten-way `f.type == i8 || f.type == i16 || ...` chain is the shape the language forced before these two, and every scalar type added later had to edit every one of those chains. Nothing has to edit a predicate. They are **comptime-only** — a gate condition selects an arm, and there is no runtime boolean for one to become, so using a predicate as a value is an error. **`^` is a constructor, and a predicate answers about the outermost one.** `^Pair` is a secret, not a record, so the three shape predicates answer false and a reflection walk refuses it instead of descending into secret storage. That is what keeps a predicate and `$fields` in agreement: `$fields(^Pair)` refuses, so a gate that called `^Pair` a record would send a walk into an operand the intrinsic then rejects. Outermost means outermost. `^*u8` is a secret pointer and `$is_pointer` answers false; `*^u8` is a **public** pointer to secret storage and is still a pointer, since the address is public. A generic instance answers as the declaration it instantiates, so `Box[i64]` is a record — which is the type a reflection loop actually meets — and `Box[^u64]` is a record too, because the instance is not itself secret; its field is, and the field is where a walk meets the question. `$is_secret` is the family's fourth member and its one **exception**, and the exception is coherent rather than special-cased: the other three ask about the shape *under* the wrapper, this one asks about the wrapper itself. It is what makes the other three's false readable — `$is_record(^Pair)` and `$is_record(u64)` are otherwise the same answer, so before this a library could only ever meet a secret as a fallthrough it had to refuse. ```mach fragment $each f in $fields(T) { $if ($is_secret(f.type)) { ... } # redact, refuse, or compare in constant time $or { ... } # an ordinary public field } ``` `$is_secret` draws the outermost line in exactly the same place the shape predicates do, so the four cannot disagree about what a type is: | operand | `$is_secret` | why | |---|---|---| | `^u64`, `^Pair`, `^[4]u8` | true | `^` is outermost | | `^*u8` | true | a secret **pointer**: the address is the secret | | `^^T` | true | `^^T` collapses to `^T` | | `*^u8` | false | a **public** pointer to secret storage | | `[4]^u8` | false | a public array of secret elements | | `rec S { k: ^u64; }` | false | the record is public, its **field** is secret | | `Box[^u64]` | false | the instance is a record; its field is secret | | `u64`, `Pair`, `ptr` | false | no `^` anywhere | `$is_integer(^u64)` and `$is_float(^f64)` answer **false** for the same reason `$is_pointer(^*u8)` does: the outermost constructor is `^`, and a secret is not the thing under it. `$is_secret` is where that question is asked. **It is not transitive, deliberately.** "Does this contain a secret anywhere" is a different question, and folding the two together would make the common case answer wrong: a `fmt` derive gating on a transitive answer would redact a whole record over one field, and could not tell which field to redact. The per-field question is the one a walk actually has, and `$is_secret(f.type)` is exactly it. Through a reference, the per-field question **composes**: `$is_secret($pointee_of(f.type))` says whether a pointer field points at secret storage. Note that `$pointee_of(^*U)` is refused, so there is no route through a secret pointer — which is correct, since `$is_secret` has already answered true for it and a walk should stop there. ##### `$holds_secret` `$holds_secret(T)` is the transitive question, asked by code that treats a type as bytes rather than walking it. It folds true when any part of `T` is secret, and it is the answer the compiler already computes when it refuses to erase a typed pointer to the raw `ptr`: one implementation, so a gate on it and that cast check cannot disagree. | operand | `$is_secret` | `$holds_secret` | |---|---|---| | `^u8`, `^Pair` | true | true | | `[32]^u8` | false | true | | `rec S { id: u64; key: ^[32]u8; }` | false | true | | `Box[^u64]` | false | true | | `*^u8` | false | true: the walk follows typed pointers, as the cast check does | | `u64`, `[32]u8`, `Pair`, `Box[u64]` | false | false | A memory primitive dispatches on byte class with the two together. A word kernel over `ptr` is legal only when no byte is secret, a wholly secret type takes the constant-time kernel, and a mixed one is handled element by element: ```mach fragment pub fun zero[T](p: *T, count: usize) { $if (!$holds_secret(T)) { raw_zero(p::ptr, $size_of(T) * count); } $or ($is_secret(T)) { ct.zeroize(p::*^u8, $size_of(T) * count); } $or { zero_typed[T](p, count); } } ``` A walk must not use it in place of `$is_secret`, for the reason above. The two questions and when each is asked are set out in [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#asking-about-secrecy-at-comptime). Because the shape predicates answer false for `^T`, "nothing classifies it" remains a usable signal on its own: a walk that gates on the shapes and refuses the fallthrough refuses secrets. `$is_secret` turns that refusal into a decision. ```mach fragment rec Inner { x: u64; y: u64; } rec Outer { i: Inner; n: u64; } $each f in $fields(Outer) { $if ($is_record(f.type)) { $each g in $fields(f.type) { ... } # descend } $or { ... } # a scalar field } ``` #### `$pointee_of(T)` — descend through a reference `$is_pointer` tells a walk that a field is a reference. `$pointee_of` says what it refers to, which is what makes the reference traversable rather than merely detectable: ```mach fragment $pointee_of(*U) # U $pointee_of(**U) # *U — one level, not all of them ``` It is a type **constructor**, in the same family as `*`, `[N]` and `^`, not a call that returns a value. So it is written wherever a type is written, including nested inside another intrinsic's operand and inside a generic argument list: ```mach fragment rec Node { value: i64; next: *Inner; } $each f in $fields(Node) { $if ($is_pointer(f.type)) { $each g in $fields($pointee_of(f.type)) { # gate, then descend total = total + (@(n.[f])).[g]; } } $or { total = total + n.[f]; } } ``` Because `str` is `def str: *char`, a `str` field is a reference field, and `$pointee_of(str)` is `u8` — which is what a formatter that renders a `str` field needs. **Everything that is not a typed reference is refused, and the refusal names what it was handed.** A plausible wrong type here flows into a `$fields` walk that then reports about the wrong record, so refusing is the only safe answer: | operand | result | |---|---| | `*U` | `U` | | `ptr` | refused — the raw pointer is untyped and carries no pointee | | `^*U` | refused — a `^` secret is not a reference | | anything else | refused, naming the type | `^` is **not** stripped, which puts `$pointee_of` with the predicates rather than with `$size_of`: `$is_pointer(^*U)` answers false, so descending through `^*U` would put the intrinsic and the gate that guards it back into disagreement, and would hand a walk secret storage the gate refused it. **Following references does not terminate structurally.** The by-value walk `$fields` supports does: a record cannot contain itself by value, so descending on `$is_record` reaches a finite set of types. A reference graph has no such property — `rec Grow[T] { p: *Grow[*T]; n: i64; }` is legal and has unboundedly many distinct instances, and a walk that follows `p` generates `Grow[i64]`, `Grow[*i64]`, `Grow[**i64]` without end. The compiler's generic-instantiation guard turns that into a diagnostic naming the derivation chain rather than a hang, but that is a backstop, not a termination story. A library that walks references owes its callers one of its own, which is why `std.derive` refuses reference fields by default and offers following as a separately named member. #### `$type_name(T)` — a type's spelling ```mach fragment $type_name(T) # the type's spelling, as a NUL-terminated string ``` Unlike the predicates this **is** a value (`*u8`), usable anywhere one is. The spelling is the same one diagnostics print, so a name a program reads and a name an error reports cannot drift. Composites spell compositely (`$type_name(*Pair)` is `"*Pair"`), and `^` spells too: `$type_name(^Pair)` is `"^Pair"`. Stripping it would be a drift on the one qualifier where a drift matters most, since a diagnostic about that type prints `^Pair`. #### `$type_id(T)` — a type's identity ```mach fragment $type_id(T) # a u64 unique to T ``` A value, like `$type_name`, and a compile-time constant: it folds wherever the layout intrinsics do (see [Where a layout intrinsic is constant](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md#where-a-layout-intrinsic-is-constant)), so it can gate a `$if`, initialize a `val`, or be compared at run time against a stored one. Its type is `u64`. It never adopts a narrower binding by value, because an identity cut to 32 bits is not one: `val id: u32 = $type_id(T)` is a type mismatch. The contract: - **Distinct for distinct types.** Types are nominal, so two records named `Point` in two modules have two identities. `^T` differs from `T`, and `*T`, `[N]T`, a function type and each instance of a generic differ by their parts: `$type_id(Box[u64])` is not `$type_id(Box[u32])`. - **Equal for the same type** wherever it is named: in any module, at any instantiation site of a generic, and in every compilation unit and library of one program. A `def` is a transparent alias, so `def Id: Point;` has `Point`'s identity. - **Unaffected by optimization and linking.** It is a constant the compiler folds, not an address, so no folding of identical code or data can merge two identities. Inside a generic, `$type_id(T)` is answered per instantiation, as the predicates are. ```mach use std.runtime; rec Point { x: u64; } rec Box[T] { v: T; } fun same[T, U]() u8 { ret $type_id(T) == $type_id(U); } #[symbol("main")] fun main() i32 { if ($type_id(^Point) == $type_id(Point)) { ret 1; } if ($type_id(Box[u64]) == $type_id(Box[u32])) { ret 2; } if (same[Box[Point], Box[Point]]() == 0) { ret 3; } ret 0; } ``` **Why a hash, and how often it collides.** A program's units are compiled separately, and a library may be compiled long before the program that links it, so no unit can see every type that will share the program. A numbering with no collisions needs that global view, and an address needs the linker and breaks at a shared library, whose internal symbols are hidden. So the identity is derived from the type alone: the first 64 bits of SHA-256 over the type's canonical recipe, which spells a nominal type by its owning module and name and every other type by its parts. Two distinct types of one program share an identity only by collision. For `n` distinct types that chance is at most `n(n-1)/2^65`: about `2^-26` for a million types, and far less for any real program. A check that pins a value to one type is defeated only when the two colliding types are the very pair it compares, and a source author aiming for that must find a SHA-256 match on 64 bits, about `2^64` work, while naming only types of their own program. The value is the same for every build of one program by one compiler. It is not promised across compiler versions, so a stored identity is a check inside a program, not a file format. #### Where `^` is stripped One rule covers the whole surface: **`^` is stripped only where the question is about storage.** | asks about | strips `^` | |---|---| | `$size_of` / `$length_of` / `$align_of` / `$offset_of` / `$discriminant_of` | yes: a secret occupies its base type storage and exposes storage width | | `$is_record` / `$is_union` / `$is_tag` / `$is_pointer` / `$is_integer` / `$is_float` | no: `^T` is a secret, not a `T` | | `$is_secret` / `$holds_secret` | no: they are the queries *about* the `^` | | `$pointee_of` | no: `^*U` is a secret, and is refused rather than followed | | `$type_name` | no: the spelling is `^T` | | `$type_id` | no: `^T` is a different type from `T` | | `$fields` / `$cases` | no: a secret aggregate is refused, not walked | | type comparison (`f.type == u64`) | no: `^u64` is not `u64` | #### The type operand `$size_of` / `$length_of` / `$align_of` / `$offset_of` / `$fields` and the queries above all take a **type** in argument 0, written with the ordinary type grammar — plus one extra form: a field descriptor's `f.type` inside a `$each` body. `$pointee_of` is part of that grammar rather than one of its consumers, so it composes with every one of them (`$size_of($pointee_of(f.type))`, `$fields($pointee_of(f.type))`). ```mach fragment $each f in $fields(T) { val n: u64 = $size_of(f.type); # the field's own size $each g in $fields(f.type) { ... } # its own fields } ``` The same form is valid in a **generic argument list**, which is what makes a walk recursive rather than merely descending (#2691): ```mach fragment fun eq[T](a: *T, b: *T) bool { $each f in $fields(T) { $if ($is_record(f.type)) { if (!eq[f.type](?a.[f], ?b.[f])) { ret false; } # re-enter at the field's type } $or { if (a.[f] != b.[f]) { ret false; } } } ret true; } ``` Without it, `$fields(f.type)` gives one level of descent per `$each` someone wrote, so a walk reaches only as deep as its author hand-unrolled. With it the walk is written once and reaches any depth. **Termination is structural and needs no depth limit.** Each descent instantiates at a field's own type, a record's fields are finite, and a record cannot contain itself by value — a self-reference must go through a pointer, which is a different type and which `$is_record` does not select. Note that a walk which followed references would not have this property: `rec Grow[T] { p: *Grow[*T]; }` is legal and has unboundedly many distinct instances reachable through its pointer. Inside the loop a field's type has no spelling, only the descriptor. A path that is genuinely a qualified type name (`mod.Type`) still reads as one, and a wrong one still reports against the type grammar — in the generic argument list exactly as in an intrinsic operand. A name that is not a `$each` loop variable is reported where it is written. #### Type intrinsic `$type_of(expr)` produces a comptime type value — the resolved type of its argument `expr`. Type values have no runtime representation; they are only meaningful as operands in comptime type comparisons. ```mach fragment $type_of(expr) # comptime type value of expr ``` Type values can be compared with `==` / `!=` inside `$if` conditions: ```mach fragment $if ($type_of(arg) == i64) { write_i64(w, arg); } $or ($type_of(arg) == str) { write_str(w, arg); } $or { $error("unsupported type"); } ``` A bare type name (e.g. `i64`, `str`, `Point`) is the other valid operand. The comparison selects one branch at compile time per monomorphization instance — useful for per-element type dispatch inside `$each` bodies. The operand is read at the **instance's** concrete type, in a plain generic as much as in a pack-tailed one, and whether it is a parameter, a local, a field, or a type derived from a generic parameter (`*T`). `$type_of(x) == T` against a generic parameter compares the instance's argument on both sides, so it holds. The provably-dead arms are **pruned** before type-checking, so each arm uses `arg` at its own concrete type with no per-arm cast: the `str` arm above is never checked against a `u64` element. Only the selected arm is type-checked and emitted. #### Field intrinsic and projection `$fields(T)` produces a comptime sequence of field descriptors for record type `T`, written with the full type grammar exactly as the layout intrinsics take it (`$fields(Box[T])`, `$fields(Pair[A, B])`, `$fields(mod.Rec)`). A union is refused: its variants overlap in storage, so a member walk over them would report distinct fields at distinct offsets that do not exist. Each descriptor carries three readable properties: | Property | Type | Value | |------------|----------|----------------------------------------------| | `f.name` | `*u8` | field name as a NUL-terminated string | | `f.type` | type val | comptime type value of the field's type | | `f.offset` | integer | byte offset of the field in `T`'s layout | `$fields(T)` is consumed by `$each f in $fields(T)`. Inside the loop body, `v.[f]` projects the concrete field off an instance `v` — it is an lvalue (readable and writable, including through a pointer receiver). ```mach fragment $fields(T) # comptime field sequence for record T v.[f] # comptime field projection: access the field f on v ``` ```mach use std.print; use std.runtime; rec Pair { x: i64; y: i64; } fun sum(p: Pair) i64 { var total: i64 = 0; $each f in $fields(Pair) { total = total + p.[f]; # p.x on iteration 1, p.y on iteration 2 } ret total; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", sum(Pair{x: 3, y: 4})); ret 0; } ``` `$each f in $fields(Empty)` expands to nothing when `T` has no fields. ##### Heterogeneous fields Because each `$each` iteration re-types `v.[f]` to the concrete field type, heterogeneous records work naturally: ```mach use std.print; use std.runtime; rec Mixed { a: i64; b: u8; } fun total(m: Mixed) i64 { var t: i64 = 0; $each f in $fields(Mixed) { t = t + m.[f]::i64; # m.a (i64) on iter 1, m.b (u8) cast to i64 on iter 2 } ret t; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", total(Mixed{a: 10, b: 2u8})); ret 0; } ``` ##### Descriptor reads Field descriptor properties can be read inside the loop body: ```mach use std.print; use std.runtime; rec Mixed { a: i64; b: u8; } fun offsum(m: Mixed) i64 { var s: i64 = 0; $each f in $fields(Mixed) { s = s + f.offset::i64; # 0 + 8 = 8 for Mixed { a: i64; b: u8; } } ret s; } fun count_i64(m: Mixed) i64 { var n: i64 = 0; $each f in $fields(Mixed) { $if (f.type == i64) { n = n + 1; } $or {} } ret n; } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { val m: Mixed = Mixed{a: 10, b: 2u8}; print.printlnf("{} {}", offsum(m), count_i64(m)); ret 0; } ``` Note: if a record has a field literally named `type`, that ordinary field access (`v.type`) is unaffected — `v.[f]` projection uses the `$each` loop variable, which is always a field descriptor, never a regular member. ##### Nested `$each` `$each` can be nested: ```mach fragment fun cross(p: Pair, q: Pair) i64 { var t: i64 = 0; $each f in $fields(Pair) { $each g in $fields(Pair) { t = t + p.[f] * q.[g]; } } ret t; } ``` #### Tag reflection: `$cases` and `$discriminant_of` The intrinsics described here are the reflection half of the tagged-value contract in [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md). `$cases(T)` produces a comptime sequence of owner-qualified case descriptors for a tag type `T`, in declaration order: ```mach fragment $cases(T) # comptime case descriptor sequence for tag T ``` `$cases(T)` is consumed by `$each case in $cases(T)`. Each case descriptor provides five readable properties: | Property | Type | Value | |--------------------|-----------|--------------------------------------------------------------| | `case.name` | `*u8` | case name as a NUL-terminated string | | `case.has_payload` | predicate | comptime predicate, valid only as a `$if` gate condition | | `case.type` | type val | comptime type value of the payload | | `case.offset` | `u64` | byte offset of the payload in `T`'s layout | | `case.code` | `u64` | declaration ordinal case code, the stored discriminator value | Accessing `case.type` or `case.offset` on a descriptor whose `has_payload` is false is a compile error; gate on `case.has_payload` first. `case.type` preserves all declared payload qualifiers, and over a generic instantiation it names the specialized payload type. `$is_tag(^T)` is false, and `$cases(^T)` is rejected. `$fields` refuses a tag and `$cases` refuses a record. Inside the loop body, `sel value.[case]` is the case test and `value.[case]` is the guarded payload place: ```mach use std.print; use std.runtime; tag Reply: u8 { empty; value: i64; } fun consume[T](v: T) { print.printlnf("value {}", v); } fun walk[T](value: T) { $each case in $cases(T) { if (sel value.[case]) { $if (case.has_payload) { consume[case.type](value.[case]); } } } } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { walk[Reply](Reply.value{42}); ret 0; } ``` `sel value.[case]` opens the same guards as `sel value.case`, so `value.[case]` is readable, writable and addressable exactly where the named place would be. Tags construct through the same single-case literal rule after specialization: `T.[case]{payload}` for a payload case and `T.[case]{}` for a payloadless one. A descriptor from another nominal type or another generic instantiation is rejected wherever it is used: in `sel`, in a projection and in a literal head. `$discriminant_of(T)` produces the actual unsigned integer type used to store the discriminator (`u8`, `u16`, `u32`, or `u64`). It may inspect outer-secret types (`$discriminant_of(^T)`) because it reports storage metadata rather than the active case. Conversely, `$cases(^T)` is refused, matching current shape query conventions. `$offset_of(T, payload_case)` reuses the layout intrinsic to report the common payload offset; naming a payloadless case is an error. Like `$size_of` and `$align_of`, layout answers for tags come from the checked target layout and are available during type checking. #### `$each` — compile-time unroll `$each` is a statement form that splices its body once per element of a comptime sequence. There are four sequence forms: ```mach fragment $each f in $fields(T) { ... } # one iteration per field of record T $each case in $cases(T) { ... } # one iteration per case of tag T $each a in va { ... } # one iteration per element of pack va $each x in ARR { ... } # one iteration per element of a constant array val ``` `$each` is valid only in statement scope (inside a function body). It is not a loop — the body is duplicated at compile time, not iterated at runtime. Enclosing runtime variables (e.g. an index or accumulator) are shared across all unrolled copies. See [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) for the pack form (`$each a in va`). ##### `$each` over a comptime-constant array `$each x in ARR` unrolls the body once per element of `ARR`, binding `x` to that element's compile-time constant per iteration. Unlike the pack and `$fields` forms, every element shares one type (the array's element type), so the loop variable is an ordinary constant value: it reads as a value, casts, dispatches a per-element `$if`, and — for a record element — projects fields with `x.field`. ```mach use std.print; use std.runtime; val PRIMES: [4]i64 = [4]i64{2, 3, 5, 7}; fun sum() i64 { var total: i64 = 0; $each x in PRIMES { total = total + x; # x is 2, then 3, then 5, then 7 } ret total; # 17 } #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{}", sum()); ret 0; } ``` A per-element `$if` selects its arm from the element's constant, so heterogeneous handling falls out of the unroll: ```mach fragment rec Rule { tag: i64; fn: fun(i64) i64; } val RULES: [3]Rule = [3]Rule{ Rule{tag: 1, fn: inc}, Rule{tag: 2, fn: dbl}, Rule{tag: 3, fn: neg}, }; fun run(n: i64) { $each r in RULES { $if (r.tag == 2) { use_double(r.fn(n)); } # r.fn folds to the element's function $or { use_other(r.tag, r.fn(n)); } } } ``` **Eligibility.** `ARR` must name an immutable `val` (never a `var`) declared in the current module, whose type is a fixed-size array `[N]E` fully initialized by an array literal of exactly `N` elements. `E` must be a scalar or record type; nested-array element types are not supported. An empty array (`[0]E`) unrolls to nothing. Each violation is reported with a teaching diagnostic. `x.field` on a record element projects the element's constant: a scalar field folds to a constant, a function-pointer field yields a function reference, and a record field materializes the nested literal. Projection is one level deep (`x.field`); `x` itself is a constant and has no address (`?x` is rejected). #### Diagnostic intrinsics `$error("msg")` fails compilation with `msg` when it is **reached** — on a live path: an unconditional position, or a `$if` / `$or` arm the compiler selects. A `$error` in a discarded (dead) arm never fires, so it is the natural total- coverage fallback for a `$type_of` dispatch — the unhandled-type `$or {}` arm fails the build at compile time instead of falling through to a runtime error. `$error` is valid in both declaration and statement scope and takes one string-literal message. ```mach fragment $error("msg") # fails compilation when reached $if (!supported) { $error("this target is not supported"); } $if ($type_of(arg) == i64) { write_i64(w, arg); } $or ($type_of(arg) == str) { write_str(w, arg); } $or { $error("no writer for this argument type"); } # compile error on an unhandled type ``` #### Not provided as intrinsics Code intrinsics — runtime-instruction emitters like `trap`, `fence`, `pause` — are not in the compiler-shipped set. They belong in stdlib as functions with per-arch `asm` bodies. See [policy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md). **`$assert` is not an intrinsic either**, and is not planned as one: `$if` and `$error` already compose to it exactly, so a dedicated directive would add spelling without adding capability. Write the composition directly. ```mach fragment # instead of $assert(cond, "msg") $if (!cond) { $error("msg"); } $if (!($mach.build.arch == $mach.arch.x86_64)) { $error("expected x86_64"); } ``` The composition inherits `$if`'s condition rules, which is the point: the same conditions fold there as in any other gate, and the ones that do not (a type query over an unbound generic parameter, anything asked of a chain that also declares something) refuse with their own cause rather than through a second surface that could describe them differently. A chain written this way declares nothing, so it is decided during type checking and can measure a type; see [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md). #### See also - [comptime.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime.md) — channel overview - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` / `$or` - [variadics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/variadics.md) — `$each a in va`, `va: ...`, `va.len`, `va...` Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md ### `$if` / `$or` — comptime control flow `$if` and `$or` branch on comptime-evaluable conditions. Only the taken branch compiles — the discarded branches are not resolved, type-checked, or emitted into the binary. This is fundamentally different from runtime `if` / `or`, which generates a branch at runtime. #### Grammar ```mach fragment $if (cond) { ... } $or (cond) { ... } $or { ... # comptime else } ``` The condition must be a comptime expression. Common shapes: - `$mach.*` reads for target / build conditions - Comparisons of comptime constants (`pub val` declarations) - Comparisons of a comptime function parameter (`$mode`) — see below A condition is a `u8`, and a constant it reads has the width and sign its declaration gives it. A constant declared through a `def`, as std's `bool` is `def bool: u8`, takes the integer primitive the `def` chain ends at, followed across `use` and `fwd`, so `$if (capability.HAS_FILES)` reads a `u8`. A chain that ends outside the integers (a record, a pointer, a float) is refused at the gate with a message naming the chain. A `def` declared inside a `$if` arm is not read this way: a declaring gate on a constant typed through one is decided once types are checked. A comptime comparison or arithmetic relates the **mathematical values** of its operands, exactly as the runtime operators do (see [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md)). A constant in `2^63 .. 2^64-1` is its true unsigned magnitude, so `$if (0xFFFFFFFFFFFFFFFF > 0)` is taken and a cross-sign comparison agrees with the runtime `if` — `$if (X < Y)` never selects a branch that `if (X < Y)` would not. Comptime arithmetic that overflows the value's range is a compile error rather than a silent wrap. #### Examples ##### Target-conditional code ```mach fragment $if ($mach.build.os == $mach.os.linux) { use full.os.linux; } $or ($mach.build.os == $mach.os.windows) { use full.os.windows; } $or { $error("unsupported OS"); } ``` ##### Dispatch on a comptime function parameter A comptime function parameter (`$mode: u8`) is a compile-time-known argument fixed per call site. `$if` / `$or` may branch on it: the compiler **monomorphizes** the function body once per distinct comptime-argument value, and each instance compiles only the arm its value selects. ```mach use std.print; use std.runtime; val MODE_DOUBLE: u8 = 0; val MODE_SQUARE: u8 = 1; fun apply($mode: u8, n: i64) i64 { $if (mode == MODE_DOUBLE) { ret n + n; } $or (mode == MODE_SQUARE) { ret n * n; } ret 0; } # apply(MODE_DOUBLE, ..) and apply(MODE_SQUARE, ..) emit two distinct bodies, # each carrying only its selected arm. #[symbol("main")] fun main(argc: i64, argv: **u8) i64 { print.printlnf("{} {}", apply(MODE_DOUBLE, 7), apply(MODE_SQUARE, 7)); ret 0; } ``` Rules: - The argument bound to a `$`-parameter must be a compile-time constant at the call site (a literal, a `pub val`, or another comptime parameter); a runtime value is rejected with `comptime argument is not a compile-time constant`. A cross-module constant works, whether imported by bare name or as a qualified member (`alias.CONST`). - Each arm gate of a comptime-parameter `$if` must itself be comptime-foldable: its identifiers must all be comptime (the comptime parameters or comptime constants). A gate referencing a runtime local/parameter is rejected. - A comptime parameter has no storage, so its address cannot be taken (`?$mode` is rejected with `cannot take the address of a comptime parameter`). - A comptime-parameter function is a template, not a value: it can only be called, not assigned, passed, or compared (`val fp = apply;` is rejected with `cannot reference a comptime-parameter function as a value`). - Comptime parameters carry no runtime cost: they are stripped from the lowered signature and ABI, so only the runtime parameters are passed. - A comptime parameter may be mixed freely with runtime parameters in any order. - A comptime parameter on a **generic** function (`fun f[T]($mode: u8, ...)`) is refused: a function has type instances or value instances, never both, and the declaration is reported as `comptime value parameters on a generic function are not yet supported`. - The function may live in any module: a value-parameter instance is emitted against its declaring module and folds its `$if` gates against that module's own comptime constants, so a library can export a comptime-parameter function gated on its own `pub val`s. - A comptime parameter may gate per-target asm safely, since each instance only compiles its taken arm: ```mach fragment pub fun load($order: Order, ptr: *i64) i64 { var result: i64 = 0; $if ($mach.build.arch == $mach.arch.aarch64) { $if (order == RELAXED) { asm aarch64 { ldr {result}, [{ptr}] } } $or (order == ACQUIRE) { asm aarch64 { ldar {result}, [{ptr}] } } } ret result; } ``` #### When a declaration-scope `$if` is decided A `$if` chain written in declaration scope runs at one of two times, and what its arms contain picks which. - **Some arm declares something.** The chain is decided while the modules load, because what it decides is which declarations exist and every later stage reads the resulting declaration set. Nothing has a type at that point, so the gate cannot ask a type question: `$size_of`, `$align_of`, `$length_of`, `$offset_of`, `$type_id`, `$type_of`, `$type_name` or an `$is_*` predicate there is rejected, with a message naming the question. So is one the gate reaches through a constant it reads, and the message names the constants it goes through. - **No arm declares anything.** The chain contributes no name and no type whichever arm is taken, so nothing depends on deciding it early. It is decided during type checking instead, where its gate may measure a type (`$size_of`, `$align_of`, `$length_of`), query one (`$is_record` and friends), or compare one. The question is answered from the **syntax**, over every arm (`$if`, every `$or`, and the final `$else`-style `$or {}`) before any gate is evaluated. One declaring arm anywhere keeps the whole chain at the earlier time. Per-arm answers are not possible: which stage runs the gate would then depend on which arm the gate selects, and the stage that would have to know that is the one being chosen. A `use` is a declaration, so a conditional import is always decided while the modules load. That is what makes the common target-gating form work. A declaring gate reads build facts (`$mach.*`, `$bin.*`, `$project.*`) and `val` constants from any module. The gates are decided in one pass that reads constants in the order they depend on each other, not the order they are written in: a gate may read a constant declared later in its own module, one declared under a later `$if` (whose gate is decided first), or one a later `use` imports. A `val` a gate reads may be typed through a `def`, followed across `use` and `fwd`, and a cast in the gate converts to the integer its destination names. Constants and gates that depend on each other have no order that decides them, so the gate is rejected with a message naming the cycle: ```mach error depend on each other, so no order decides them $if (A == 4) { pub val B: u32 = 4; } $if (B == 4) { pub val A: u32 = 4; } ``` ```mach rec MeshUniforms { model: [16]f32; } # no arm declares: decided during type checking, so the gate may measure $if ($size_of(MeshUniforms) != 64) { $error("MeshUniforms must be 64 bytes"); } ``` ```mach error `$size_of` asks a type question rec MeshUniforms { model: [16]f32; } # the second arm declares, so the whole chain is decided while the modules # load - and the gate is rejected there $if ($size_of(MeshUniforms) != 64) { $error("MeshUniforms must be 64 bytes"); } $or { val PADDING: u32 = 0; } ``` The one visible consequence is ordering. A `$error` reached under a chain that declares nothing is reported during type checking, so an unrelated name-resolution error elsewhere in the same module is reported before it rather than after. A `$if` inside a **function body** is always decided during type checking: it selects statements rather than declarations, so the question above does not arise and its gate may always ask about a type. #### Discarded branches A `$if` branch that isn't taken is entirely absent from the compiled output. For a target- or constant-gated `$if`, the compiler doesn't even resolve names inside the untaken branches — this is the mechanism that makes per-target asm blocks safe even when one block references registers the other backend doesn't know about. Inside a function body, a gate that names a constant from another module (for example `$if (capability.HOSTED) { ... }`) is decided during type checking, after names are resolved. Its arms are read the same way: a name that doesn't exist in an arm the gate discards is never reported, exactly as with C's `#if`. A name that doesn't exist in the arm the gate **selects** is reported as an ordinary `unresolved identifier` or `unresolved type name` error, once, however many times the enclosing function is instantiated. So such an arm may name what only exists on the targets that select it. ```mach fragment use capability: std.system.capability; fun field_token(c: *Cursor, error: io_error.Error) { $if (capability.HOSTED) { # names that only a hosted target provides put_quoted(c, io_error.message(error)); } } ``` A `$if` gated on a comptime **parameter** is the other exception: because arm selection happens per call site (at monomorphization), name resolution and type checking run over **all** arms structurally, and only the selected arm is emitted into each instance. Each arm must therefore be independently resolvable and type-checkable. A `$if` gated on a `$type_of` type comparison is *not* such an exception: at monomorphization the operand's concrete type is known, so the provably-dead arms are **pruned** and only the selected arm is type-checked (and emitted). Each arm may therefore use its value at its own concrete type with no per-arm cast — see [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md). Such a chain is decided only at an instantiation. While the operand's type still names a generic parameter no arm is selected and none is type-checked, so a `!=` gate and a bare `$or` fallback wait for the concrete type like every other gate, and a `$error` in an arm the instantiation does not select never fires. An instantiation that matches no arm selects the fallback, and its `$error` is reported at that instantiation. #### Two regimes inside a generic body A generic body is checked under two rules, and which one applies depends on what the question is about. Both are stated here because the boundary between them is the only thing a reader has to hold. | the question | when it is answered | what is checked | |---|---|---| | a `$if` gated on a comptime **value parameter** | per call site | **all** arms, structurally, before any is selected | | a `$if` gated on a `$type_of` or a type predicate | per instantiation | only the arm that instantiation selects | | an operator, cast, `:~`, literal or condition on a **type parameter** | per instantiation | the whole body, once per distinct instantiation | The first is the exception described under *Discarded branches*: arm selection happens per call site, so every arm must be independently resolvable and type-checkable and only the selected one is emitted. The second and third are the same rule applied to different constructs. Nothing about a type parameter is decided while it is still a parameter, because there is no concrete type to decide it against: a gate waits for the instantiation, and so does every operator. The template types the body so each instance and the lowering have an expression table to read, and reports nothing of its own. A refusal belongs to the instantiation that asked for the instance and names that instance's concrete type — see [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md). The consequence is worth stating plainly: a generic that nothing instantiates is not checked at all. #### See also - [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md) — `$mach.*` for target reads - [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md) — `asm` blocks gated by `$if` - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) — runtime `if` / `or` counterpart - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — generic type parameters and per-instance checking Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md ### Inline assembly Mach has one inline-assembly form: an ISA-tagged block of raw instructions with local-variable substitution. The compiler parses the instruction stream and infers operand direction and clobbers from the opcode semantics — no `in` / `out` declarations, no clobber list. #### Grammar ```mach fragment asm { # raw instructions, one per line, # for comments mov rcx, {ptr} mov rax, [rcx] mov {result}, rax } ``` - The ISA tag is mandatory. Bare `asm { ... }` does not exist. - The tag comes from a closed set: `x86_64`, `aarch64`, `riscv64`, `riscv32`. Each has a working assembler that emits native bytes; the first three run in CI (riscv64 under qemu, including a self-host smoke — riscv64 is a self-hosting target with a byte-identical fixpoint, #1852). - **The RISC-V tag names the machine, not the family.** An `asm riscv64` block is refused on a riscv32 target and the other way round, so a body reaching the assembler was written for the register width it is being assembled for. Guard the two with `$if ($mach.build.arch == $mach.arch.riscv32)` when a routine needs both. An RV64-only spelling (`ld`, `sd`, the `*W` group, the doubleword atomics) is refused under a `riscv32` tag, naming the instruction and the machine, because each instruction states the register widths it exists at and the assembler reads that rather than a second list of its own. - **A target with no assembler takes no `asm` block.** SPIR-V declares no assembly capability, so any `asm` block compiled for it, whatever its tag, is refused where it is written with `inline assembly is not available on instruction set 'spirv'`. Guard a routine shared with a shader behind `$if ($mach.build.arch != $mach.arch.spirv)`. - Each line is an instruction in the ISA's native syntax. - `#` introduces a line comment: everything from `#` to the end of the line is ignored, whatever it contains (`;`, `{}`, `%`, and so on are all inert inside a comment). #### Operand substitution `{name}` substitutes a local in scope. The compiler resolves the reference to a memory or register operand based on liveness and the instruction's expected operand class. In practice a `{name}` binds the local's storage — typically a stack slot — so a pointer local's pointee is reached by staging the pointer through a scratch register first (`mov rcx, {ptr}` then `mov rax, [rcx]`), never by a direct `[{ptr}]` indirection. Only an identifier inside braces names a local. Any other braced text belongs to the ISA's own syntax, such as aarch64's `{v0.16b}` register list, and reaches its grammar untouched. ```mach fragment pub fun add_via_asm(a: i64, b: i64) i64 { var result: i64 = 0; asm x86_64 { mov rax, {a} add rax, {b} mov {result}, rax } ret result; } ``` ##### A body that binds `{name}` may not move the stack pointer A `{name}` becomes a fixed displacement off a base register, measured once when the block is assembled. On aarch64 and riscv64 that base is the **stack pointer**, because it is the only one whose displacement stays inside those ISAs' immediate forms however deep the enclosing frame is. A statement that moves the stack pointer therefore moves every `{name}` in the block out from under its own address, and the compiler refuses the block rather than assembling a wrong one: ```mach fragment var x: i64 = 0; asm aarch64 { ldr x9, {x} stp x1, x2, [sp, -16]! # refused: this body binds {x} } ``` The refusal covers the whole block, not the statements after the write, because a backward branch reaches an earlier `{name}` again with the pointer already moved. Push and pop around the block instead, or drop the `{name}` and stage the address into a register yourself. A body that binds no `{name}` is unaffected and may do whatever it likes with the stack pointer, which is what a `#[naked]` function's hand-written prologue does. #### Calls and jumps (x86-64) `call` and `jmp` take the same three shapes, and which one a statement means is read off the operand: ```mach fragment asm x86_64 { call some_symbol # direct: E8 rel32, relocated against the symbol call rax # indirect through a register: FF /2, mod=11 call [0x100018] # indirect through an absolute address: ff 14 25 call [rax + 8] # indirect through a computed address jmp some_symbol # direct: E9 rel32 jmp rax # indirect through a register: FF /4 jmp [rax + 8] # ... and the same memory forms } ``` The absolute form exists for a fixed-address ABI — one whose entry points are addresses rather than symbols, like BareMetal's kernel call table at `0x100010..0x100040`. Its displacement is sign-extended to 64 bits, so an address outside signed 32-bit range is refused rather than silently truncated. `call [symbol]` and `jmp [symbol]` are refused too: the rip-relative form would mean "transfer to the pointer *stored* at the symbol", which is not what the direct form beside it means. Both indirect operands are fixed 64-bit in long mode, so a narrower register (`jmp eax`) is refused rather than widened. An indirect call clobbers exactly as a direct one does — the callee's caller-saved registers, which the surrounding block's barrier already covers. The register or memory holding the target is **read**, not written. #### Operand sizes (x86-64) A register operand states its own width, so `mov eax, [rcx]` is a four-byte load and needs nothing else. A memory operand states none, and where the instruction does not settle it either the width must be written out, in nasm's spelling: ```mach fragment asm x86_64 { movzx eax, word [rcx] # a two-byte load, zero-extended into eax movsx rax, dword [rcx] # a four-byte load, sign-extended (movsxd) mov dword [rcx], 1 # a four-byte store, not the machine word neg qword [rcx] # an eight-byte read-modify-write } ``` `byte`, `word`, `dword` and `qword` are accepted before a memory operand and nowhere else — on a register they would be redundant or contradictory, so `mov qword rax, rcx` is refused rather than ignored. GNU as's spelling with `ptr` (`mov dword ptr [rcx], 1`) means the same, and it is how the x86-64 `--emit-asm` listing writes every sized memory operand, so a listed instruction pastes back into a block unchanged. **A prefix that contradicts the instruction is a build error, not a dropped token.** What counts as a contradiction is per mnemonic: | shape | rule | |---|---| | most instructions | every operand shares one width, so a prefix must agree with any register operand, and two register operands of different widths (`add rax, ecx`) name no form; with no register operand the prefix *sets* the width | | `shl` / `shr` / `sar`, `shld` / `shrd` | the count is `cl` or an immediate whatever the width of the operand shifted | | `movzx` / `movsx` | the source is narrower by design, so a memory source **must** be sized, and the size must be strictly narrower than the destination | | `push` / `pop`, indirect `call` / `jmp` | fixed 64-bit in long mode, so any narrower prefix names no instruction | | `lidt` | its pseudo-descriptor is ten bytes, which no keyword names | So `mov eax, word [rcx]` is refused (two widths for one access), and `movzx eax, [rcx]` is refused too — an unsized source names no width at all, and reading it as a same-width move would silently assemble a plain `mov` where a zero-extending load was written. #### Bit scans, byte swap, multiply and double shifts (x86-64) ```mach fragment asm x86_64 { bsf rax, rcx # index of the lowest set bit; zf set and rax undefined when rcx is 0 bsr rdx, qword [rdi] # index of the highest set bit tzcnt r8, r9 # trailing zeros, 64 for a zero source (bmi1) lzcnt eax, dword [rsi] # leading zeros, 32 for a zero source (lzcnt) popcnt rax, rcx # set bits (popcnt) bswap rax # reverse the bytes of a 32- or 64-bit register imul rax, rcx # rax = rax * rcx, low half imul rax, rcx, 5 # rax = rcx * 5, low half, with a 32-bit signed immediate imul eax, dword [rdi], 7 # the source may be memory in either form shld rax, rcx, 5 # shift rax left, bits from rcx fill from the right shrd qword [rdi], rax, cl # the memory operand is shifted, rax feeds bits, cl counts } ``` The scans and counts take a 16-, 32- or 64-bit register destination and a register or memory source of the same width. `bsf` and `bsr` set ZF on a zero source and leave the destination and CF undefined, which the effect model reports as writing the flags. `popcnt`, `lzcnt` and `tzcnt` define ZF and CF, and need their extension, listed under [Extension instructions](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md#extension-instructions). `bswap` takes one 32- or 64-bit register, and has no 16-bit form. `imul` in two or three operands is the signed multiply with the low half kept; the one-operand widening form is not spelled. `shld` and `shrd` take the register or memory shifted, a register of the same width feeding bits, and a count of 0 to 255 or `cl`. A secret in `cl` is a variable-latency count for the constant-time check, as it is for `shl`; an immediate count is not. #### String instructions (x86-64) ```mach fragment asm x86_64 { mov rdi, {dst} mov rsi, {src} mov rcx, {n} rep movsb # copy rcx bytes from [rsi] to [rdi] xor eax, eax mov rcx, 8 rep stosq # store rax into 8 quadwords at [rdi] repne scasb # step rdi until [rdi] equals al, or rcx runs out } ``` `movs`, `stos`, `lods`, `cmps` and `scas` take a width suffix, `b`, `w`, `d` or `q`, and no operands: `movsd` here is the string move, since the scalar-double move is not an inline-asm row. Each reaches memory through its implicit registers: | mnemonic | reads | writes | memory | |---|---|---|---| | `movs` | `rsi`, `rdi` | `rsi`, `rdi` | loads `[rsi]`, stores `[rdi]` | | `stos` | `rax`, `rdi` | `rdi` | stores `[rdi]` | | `lods` | `rsi` | `rax`, `rsi` | loads `[rsi]` | | `cmps` | `rsi`, `rdi` | `rsi`, `rdi`, flags | loads both | | `scas` | `rax`, `rdi` | `rdi`, flags | loads `[rdi]` | A prefix repeats the instruction `rcx` times and counts `rcx` down, so it adds `rcx` to what the instruction reads and writes. `rep` repeats a move, store or load. A compare repeats under a condition instead, `repe` (or `repz`) while its elements are equal and `repne` (or `repnz`) while they differ, and stops early when the condition fails. `rep cmpsb` and `repe movsb` are refused, since the first says nothing about when to stop and the second names a condition nothing sets. The written registers are the block's clobbers like any other, so a value the allocator keeps live across the block never sits in one of them. Every step moves `rsi` and `rdi` forward when the direction flag is clear and backward when it is set. The compiler does not track DF. It assumes the ABI's guarantee that DF is clear at every call and return, so a block that sets it with a raw encoding must clear it with `cld` before the block ends. The constant-time check treats `rsi` and `rdi` as addresses, so a secret in either is refused as one. A single step is a load or store and passes data through as `mov` does. A repeated one runs once per element and a compare also stops on its data, so its timing is its count and its contents. In an `#[oblivious]` function it is refused whenever it reads a secret: a secret count, a secret value to store, or memory a pointer to a secret addresses. A public count storing public data is admitted even into secret memory, which is how a key buffer is cleared with `rep stosb`. #### Segment-relative memory (x86-64) A memory operand may lead with `fs:` or `gs:`, after any width keyword. The address is then relative to that segment's base, and the instruction carries the `0x64` or `0x65` prefix. The override belongs to the operand, so every instruction that takes a memory operand accepts it, the vector forms included: ```mach fragment asm x86_64 { mov rax, fs:[0x28] # an absolute offset from the fs base mov rax, qword gs:[rbx + 8] # a register-relative one movdqu xmm0, gs:[16] jmp qword fs:[rcx] } ``` The override is refused wherever it would change nothing or mislead: | operand | why it is refused | |---|---| | `es:`, `cs:`, `ss:`, `ds:` | long mode ignores these bases, so the prefix relocates nothing | | a register, an immediate or a `{name}` binding | only a bracketed memory operand has an address to relocate | | `lea rax, fs:[rbx]` | `lea` never accesses its address, so the result is not segment-relative | | `fs:[symbol]` | a symbol operand is RIP-relative, and the base would move it off the symbol | A segment override does not change which registers an instruction reads or writes, and the constant-time check treats `fs:[rbx]` exactly as `[rbx]`. The language itself has no thread-local storage. These forms only let a block reach a base that something else set up. In an `#[oblivious]` function the constant-time check reads two facts off each `{name}` binding: a pointer to a secret (`*^u32`) is a public address and a secret load, so `mov rax, {p}` then `mov ecx, [rax]` is admitted and `ecx` is a secret from then on, while a secret pointer (`^*u32`) is a secret address and is refused as one. #### Vector registers Both grammars take vector registers as operands. A vector register belongs to the floating-point and vector bank, so writing one adds it to the block's vector clobber set, and a live vector value crossing the block is kept the same way a general-purpose one is. A memory base or index is always a general-purpose register, and a vector register in a general-purpose form is refused by name. **x86-64** spells them `xmm0` to `xmm15`. A memory operand of a vector instruction is a whole 128-bit vector, written bare or as `xmmword [...]`, and a narrower width prefix is refused. `[symbol]` addresses RIP-relative data, and the relocation is correct after a trailing immediate. | form | mnemonics | |---|---| | move, in either direction between a register and memory | `movdqa`, `movdqu`, `movaps`, `movups` | | `xmm, xmm/m128` | `paddb` `paddw` `paddd` `paddq`, `psubb` `psubw` `psubd` `psubq` `psubusb` `psubusw`, `pmullw` `pmulhw` `pmulhuw` `pmuludq`, `psadbw`, `pand` `por` `pxor`, `pcmpeqb` `pcmpeqw` `pcmpeqd` `pcmpgtb` `pcmpgtw` `pcmpgtd`, `punpcklbw` `punpcklwd` `punpckldq` `punpckhbw` `punpckhwd` `punpckhdq`, `packsswb` `packssdw`, `addps` `subps` `mulps` `divps` `addpd` `subpd` `mulpd` `divpd`, `cvtdq2ps` `cvttps2dq` `cvtdq2pd` `cvttpd2dq` `cvtps2pd` `cvtpd2ps` | | `xmm, xmm/m128, imm8` | `pshufd`, `cmpps`, `cmppd` | | `xmm, r64` and `r64, xmm` | `movq` | | `xmm, imm8` | `pslldq` `psrldq`, `psllq` `psrlq` | | `xmm, xmm/m128` (the count is the source's low quadword) | `psllq` `psrlq` | `movq xmm, r64` writes the low quadword and zeroes the high one, and `movq r64, xmm` reads the low quadword. Only the 64-bit general register form is spelled. `pslldq` and `psrldq` shift the whole register by bytes, and `psllq` and `psrlq` shift each quadword by bits. A count past the width empties the register or the lane, as GNU as encodes it, so any immediate from 0 to 255 is accepted. All of them are baseline SSE2 and run in fixed time whatever the count, so the constant-time check lets a secret through them as data. **aarch64** spells them `vN.16b`, `vN.8h`, `vN.4s` or `vN.2d`. The suffix is the lane arrangement, and every operand of one instruction shares it. `add`, `sub`, `and`, `orr`, `eor` and `mov` select their vector form when their operands are vector registers. `and`, `orr`, `eor`, `mov` and `mvn` exist only at `.16b`, the float members only at `.4s` and `.2d`, and `mul` everywhere except `.2d`. | form | mnemonics | |---|---| | three registers | `add` `sub` `mul`, `and` `orr` `eor`, `cmeq` `cmgt` `cmge` `cmhi` `cmhs`, `fadd` `fsub` `fmul` `fdiv`, `fcmeq` `fcmgt` `fcmge` | | two registers | `mov`, `mvn`, `rev64` (not `.2d`) | | one element structure | `ld1 {vT.}, [Xn]`, `st1 {vT.}, [Xn]` | | three `.16b` registers and the first byte | `ext vD.16b, vN.16b, vM.16b, 8` | | a register from one lane, or from a general register | `dup vD.2d, vN.d[1]`, `dup vD.4s, wN` | | one lane from a lane of its width, or from a general register | `ins vD.d[1], xN`, `ins vD.s[0], vN.s[3]` | A lane is written `vN.b[i]`, `vN.h[i]`, `vN.s[i]` or `vN.d[i]`, with the index below the lane count. A lane taken from or put into a general register uses `x` for a `.d` lane and `w` for the narrower ones. `mov vD.d[1], xN` is the same instruction as `ins`, and the listing spells it that way, as objdump does. `ins` writes one lane and keeps the rest, so the register is read as well as written. `ext`'s byte index is bare like every other immediate, from 0 to 15. `rev x0, x1` and `rev w0, w1` reverse the bytes of a general register. All of them are baseline ASIMD and data-independent, so the constant-time check lets a secret through them. `ld1` and `st1` post-index their base by the structure size, `, 16`, or by an X register, `, x9`. Either form writes the base register: ```mach fragment asm aarch64 { ld1 {v0.16b}, [x1], 16 # load 16 bytes, then x1 += 16 ld1 {v1.16b}, [x1], 16 eor v0.16b, v0.16b, v1.16b st1 {v0.16b}, [x2], x9 # store, then x2 += x9 } ``` The multiply and float members are variable-latency operations, so the constant-time check treats them as it treats their scalar counterparts. It tracks secrets in vector registers separately from general-purpose ones. #### Extension instructions Some instructions exist only on processors that implement an extension beyond the instruction set's baseline. Each such row names its extension, and the encoder refuses it unless the target selects that extension (the manifest's [`extensions`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions) key) or the enclosing function admits it with [`#[extensions(...)]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#extensionsnames--an-outlier-function). The refusal names the instruction, the extension, and both ways to admit it. The one invariant behind every check: **an instruction that requires extension E is emitted only into a function whose admitted set holds E**, where a function's admitted set is the target's selection plus what its `#[extensions(...)]` names, each closed over what it implies. The encoder applies it to every row of an `asm` block, and the inliner applies the same predicate before moving one body into another, so an inlined body never carries an instruction its new home does not admit. An `asm` block has no spelling of its own: the tag is the isa, and the block inherits its function's set. | isa | extension | mnemonics | |---|---|---| | x86_64 | `ssse3` | `pshufb xmm, xmm/m128`, `palignr xmm, xmm/m128, imm8` | | x86_64 | `sse41` | `pblendw xmm, xmm/m128, imm8`, `ptest xmm, xmm/m128`, `pinsrd xmm, r32/m32, imm8`, `pextrd r32/m32, xmm, imm8`, `pmulld xmm, xmm/m128`, `pmovsxbw` `pmovsxwd` `pmovsxdq` `pmovzxbw` `pmovzxwd` `pmovzxdq` `xmm, xmm/m64`, `pmovsxbd` `pmovsxwq` `pmovzxbd` `pmovzxwq` `xmm, xmm/m32`, `pmovsxbq` `pmovzxbq` `xmm, xmm/m16` | | x86_64 | `sha` | `sha256rnds2 xmm, xmm/m128`, `sha256msg1 xmm, xmm/m128`, `sha256msg2 xmm, xmm/m128` | | x86_64 | `fsgsbase` | `rdfsbase r32/r64`, `rdgsbase r32/r64`, `wrfsbase r32/r64`, `wrgsbase r32/r64` | | x86_64 | `popcnt` | `popcnt r16/32/64, r/m` (CPUID leaf 1, ECX bit 23) | | x86_64 | `lzcnt` | `lzcnt r16/32/64, r/m` (CPUID leaf 0x80000001, ECX bit 5) | | x86_64 | `bmi1` | `tzcnt r16/32/64, r/m` (CPUID leaf 7, EBX bit 3) | | x86_64 | `aes` | `aesenc` `aesenclast` `aesdec` `aesdeclast` `xmm, xmm/m128`, `aesimc xmm, xmm/m128`, `aeskeygenassist xmm, xmm/m128, imm8` (CPUID leaf 1, ECX bit 25) | | x86_64 | `pclmul` | `pclmulqdq xmm, xmm/m128, imm8` (CPUID leaf 1, ECX bit 1) | | x86_64 | `avx2` | `vpsllvd` `vpsrlvd` `vpsravd` `vpsllvq` `vpsrlvq` `xmm, xmm, xmm/m128` (VEX.128, CPUID leaf 7, EBX bit 5) | | x86_64 | `avx512dq` and `avx512vl` | `vpmullq xmm, xmm, xmm/m128` (EVEX.128, CPUID leaf 7, EBX bits 17 and 31) | | aarch64 | `sha2` | `sha256h qN, qN, vN.4s`, `sha256h2 qN, qN, vN.4s`, `sha256su0 vN.4s, vN.4s`, `sha256su1 vN.4s, vN.4s, vN.4s` | | aarch64 | `sb` | `sb` (the FEAT_SB speculation barrier) | | aarch64 | `aes` | `aese vN.16b, vN.16b`, `aesd vN.16b, vN.16b`, `aesmc vN.16b, vN.16b`, `aesimc vN.16b, vN.16b` (FEAT_AES) | | aarch64 | `pmull` | `pmull vN.1q, vN.1d, vN.1d`, `pmull2 vN.1q, vN.2d, vN.2d` (FEAT_PMULL; brings `aes`) | A manifest [level](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#levels) (`extensions = ["x86-64-v2"]`) selects every member name, so it admits the rows of each: `pmulld` assembles under `x86-64-v2`, and `tzcnt`, `lzcnt` and `vpsllvd` under `x86-64-v3`, and `vpmullq` under `x86-64-v4`. The names a level brings that have no rows yet (`sse42`, `avx`, `avx512f`, `avx512bw`, `avx512cd`) admit nothing until an encoding lands for them. The `avx2` rows are VEX-encoded three-operand instructions: the destination, then the shifted vector, then the per-lane counts, as GNU as spells them. Each shifts every lane by the count in the same lane of the last operand, and a count at or above the lane width empties the lane (or, for `vpsravd`, fills it with the sign). The 128-bit form is the only one inline asm spells; `ymm` registers are not operands. `vpmullq` is EVEX-encoded and takes its operands the same way: the destination, then the two factors, and it keeps the low 64 bits of each lane's product. Its 128-bit form needs `avx512vl` beside `avx512dq`, so a target selecting only one of them refuses it, naming the other. A memory source's displacement is scaled by 16 when that fits a byte, as GNU as does. Opmask registers, broadcast, `xmm16` to `xmm31`, and the `ymm` and `zmm` forms are not operands inline asm spells. `sha256rnds2` also reads `xmm0`, the round keys, without naming it, and the constant-time check follows a secret through it. `ptest` defines ZF and CF, so a branch after a `ptest` of public data no longer counts a secret an earlier instruction left in the flags. The aes rounds and the carry-less multiplies are data-independent: Intel lists `aesenc`, `aesenclast`, `aesdec`, `aesdeclast`, `aesimc`, `aeskeygenassist` and `pclmulqdq` among the instructions whose latency does not depend on their data operands, and Arm lists `aese`, `aesd`, `aesmc`, `aesimc`, `pmull` and `pmull2` among those whose timing is independent of their data under DIT. The constant-time check admits them in an `#[oblivious]` function the way it admits the sha rows, and follows a secret through them, so what they compute from a secret is still a secret. On aarch64 `pmull` names its product `.1q` and its operands `.1d`, and `pmull2` reads the high halves at `.2d`. The `pmull` row brings `aes`, because the architecture reports both in one field (ID_AA64ISAR0_EL1.AES is 1 for the aes rounds and 2 for those and the 64-bit `pmull`). `cpuid` is baseline on x86_64 and is how a program finds out which extensions the processor has. It reads the leaf from `eax` and the subleaf from `ecx`, and writes all of `eax`, `ebx`, `ecx` and `edx`: ```mach fragment var b: u32 = 0; var c: u32 = 0; var d: u32 = 0; asm x86_64 { mov eax, 7 xor ecx, ecx cpuid mov {b}, ebx mov {c}, ecx mov {d}, edx } val has_sha: bool = ((b >> 29) & 1) == 1; ``` #### Privileged and systems instructions (x86-64) Beyond the ordinary surface, an OS-level block reaches: ```mach fragment asm x86_64 { cli / sti # the interrupt flag cld # clear DF before entering a program hlt # park the core in al, dx / out dx, al # port i/o, by immediate port or through dx rdtsc / rdmsr / wrmsr # the counter and the model-specific registers cpuid # the processor's identity and extensions lidt [rax] # install an interrupt descriptor table pushfq / popfq # save and restore RFLAGS swapgs # per-CPU state on a syscall entry rdfsbase rax / wrgsbase r9d # the fs or gs base, with the fsgsbase extension iretq # return from an interrupt handler mov rax, cr2 # the faulting address in a page-fault handler mov cr3, rax # switch page tables mov eax, cs # the live selector, for programming STAR } ``` Control registers are `cr0`, `cr2`, `cr3`, `cr4` and `cr8` — CR1 and CR5–CR7 are reserved and have no spelling. A control-register move takes a 64-bit general-purpose register on its other side, since there is no narrower form. Segment registers (`es`, `cs`, `ss`, `ds`, `fs`, `gs`) can be **read** into a general-purpose register; writing one through `mov` is not supported, because long mode does not admit it for CS and SS at all and the remaining loads carry descriptor-cache and interrupt-shadow effects. None of these writes a register the allocator tracks: the interrupt and direction flags, RFLAGS, the stack pointer and a segment base are all outside the allocated file, and `mov cr3, rax` writes a control register rather than any general-purpose one. `rdtsc` and `rdmsr` are the exceptions — both land their result in EDX:EAX, which the effect model reports — and so is `cpuid`, which writes EAX, EBX, ECX and EDX. `rdfsbase`, `rdgsbase`, `wrfsbase` and `wrgsbase` take one 32- or 64-bit general-purpose register. They need the `fsgsbase` extension (CPUID leaf 7, EBX bit 0), selected by the target or admitted by the function as in [Extension instructions](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md#extension-instructions). A read writes that register, and a write changes only the segment base. They also need the kernel to have enabled FSGSBASE, and they trap elsewhere. A base write is an address for the constant-time check, because every later `fs:` or `gs:` access goes through it. Writing a secret into a base is therefore refused, just as addressing memory with a secret is. `iretq` does not fall through, and the effect model has no way to say so: a `Mnemonic` names registers written and nothing else. That is sound — nothing can survive an instruction control never returns from — but it means statements after an `iretq` are unreachable without the compiler saying so. #### Raw encodings Four data directives emit their values verbatim, for an encoding the ISA's mnemonic table does not name. They work on every target: ```mach fragment asm x86_64 { .byte 0x0f, 0x01, 0xd0 # xgetbv } asm aarch64 { .word 0xd53be040 # mrs x0, cntvct_el0 } asm riscv64 { .word 0xc0102573 # csrr a0, time } ``` The widths are GNU as's, per target — `.word` is the one that differs: | directive | x86-64 | aarch64 / riscv64 | |---|---|---| | `.byte` | 1 byte | 1 byte | | `.word` | **2 bytes** | **4 bytes** | | `.long` | 4 bytes | 4 bytes | | `.quad` | 8 bytes | 8 bytes | Values are written in the target's byte order, so `.word 0xd503201f` is the aarch64 `nop` as its manual prints it. A non-negative value is read unsigned (`.quad 0xFFFFFFFFFFFFFFFF` is a legal address); a negative one is its two's complement at the directive's width. A value the width cannot hold is refused rather than truncated, and one directive carries a whole sequence — up to 256 payload bytes — not four. On aarch64 and riscv64 a statement must emit a whole number of instruction words: `.byte 0x1f, 0x20, 0x03` is refused, because three bytes would misalign every instruction after it. x86-64 has no such constraint. A raw encoding is an instruction stream the parser cannot read, so by default the block's clobber set becomes **every register in every bank**, and an `#[oblivious]` function may not contain one at all — which is why a real mnemonic is always preferable where one exists. #### Declaring a raw encoding's effects That default is correct and expensive: with nothing held live across the statement, the allocator spills every value in a callee-saved register and the prologue saves every callee-saved register the function could reach. A raw encoding may instead state what it writes, with a `::` clause on the directive: ```mach fragment asm x86_64 { mov dx, {port} .byte 0xee :: writes() # out dx, al - reads dx and al, writes nothing .byte 0x0f, 0x31 :: writes(rax, rdx) # rdtsc - lands its result in EDX:EAX } asm aarch64 { .word 0xd4000002 :: writes(x0, x1, x2, x3) # hvc #0, returning per SMCCC } asm riscv64 { .word 0xc0102573 :: writes(a0) # csrr a0, time } ``` This is the same move the system-register surface above already makes. The named set is deliberately not exhaustive because any register is *also* nameable by its encoding; declared clobbers do for instructions what `s3_3_c14_c0_2` does for registers. The curated mnemonic table stays the ergonomic path, and the escape hatch stops being a cliff — an unmodeled instruction can be used at full codegen quality on the day it is needed. `writes()` with an empty list is a **declaration**, not an omission: it says the encoding writes no register. Leaving the clause off entirely is what keeps today's conservative default, so nothing written before this existed changes behaviour. Registers are named in the target's own vocabulary, in either bank: | target | general-purpose | float / vector | |---|---|---| | x86-64 | `rax`–`r15` | `xmm0`–`xmm15` | | aarch64 | `x0`–`x30`, `sp`, `xzr` | `v0`–`v31` | | riscv64 | `x0`–`x31`, psABI aliases (`a0`, `t0`, `sp`, …) | `f0`–`f31`, psABI aliases (`fa0`, `ft0`, …) | A width is not a register: `eax`, `w0` and `v2.4s` are refused, because `eax` and `rax` are one register whose bits the allocator tracks as a unit and accepting both spellings would suggest the clobber set told them apart. ##### What the compiler still guarantees, and what it does not A declaration moves part of the correctness burden from the compiler to you, and it is worth being precise about which part. **Still checked.** The clause is accounted for like every other byte of the statement, so a malformed one fails the build rather than being silently dropped — a mistyped `writes(` is an error, not a declaration that quietly said nothing. Every register named must resolve in this target's own name table, so a typo fails rather than declaring a smaller set than you wrote. An attribute that does not exist is refused by name. **Structurally bounded.** A clause may only follow a **data directive**. A modeled instruction's effects come from its mnemonic and cannot be overridden, so a declaration can only ever narrow the maximally conservative default, at exactly the statements where the compiler had no information at all. It can never contradict something the compiler derived, and a wrong one affects only the single statement carrying it. **Not checked, and it cannot be.** Whether the bytes write what the clause says. The compiler does not decode the payload — decoding it is what the mnemonic table *is*. **If you under-declare, you get a miscompile**: the allocator will keep a value in a register your encoding overwrites. This is the one place in the inline-asm surface where correctness rests on the author. Declare what the instruction's manual says it writes, including any implicit destination, and prefer a real mnemonic whenever one exists. **A declaration buys allocation quality, not verification.** An `#[oblivious]` function still refuses a data directive whether or not it is declared. A write set says which registers change; it says nothing about whether the encoding branches on a secret or divides in variable time, which is what that check exists to catch. `writes(...)` is the only attribute today. The clause is a space-separated list so more can be added, but an attribute is only added once it changes what the compiler does — `noreturn` is not here yet because control flow reaching past an `asm` block is a CFG fact the effect model does not touch, and a `barrier` attribute would be inert, since every `asm` block is already assumed to modify arbitrary memory. #### System registers (aarch64) `mrs` and `msr` name a system register by its architectural name, in either case: ```mach fragment asm aarch64 { mrs x0, cntvct_el0 # the virtual counter mrs x1, CNTFRQ_EL0 # ... and its frequency, capitalized as ARM spells it msr vbar_el1, x2 # install an exception vector base msr daifset, 0xf # mask every interrupt } ``` The named set covers what freestanding code reaches for — the generic timer, the exception vectors and their syndrome registers, the MMU control registers, the thread pointers, and enough identification registers to detect a CPU. It is deliberately not exhaustive: **any** system register is also nameable by its encoding, exactly as ARM and GNU as spell it, which is what makes the surface complete rather than a list that always lags the architecture: ```mach fragment asm aarch64 { mrs x0, s3_3_c14_c0_2 # the same register as `mrs x0, cntvct_el0` } ``` `op0` must be 2 or 3 — the whole of the `mrs` / `msr` register space — and each remaining field is bounded by its own width. A field the architecture cannot hold is refused rather than truncated, because a truncated selector would name a *different* register than the text does. `msr , imm` writes a PSTATE field (`daifset`, `daifclr`, `spsel`, `pan`, `uao`, `ssbs`, `dit`, `tco`). The architecture spells these by name only, so there is no numeric escape for this form. The immediate is spelled bare, as every immediate in an `asm` block is: `#` opens a comment to the end of the line, so Arm's `#1` would leave the instruction without its operand. ```mach fragment asm aarch64 { msr dit, 1 # turn on data-independent timing for this thread dsb nsh # ... and make the write take effect before what follows isb } ``` Access permission is not checked: whether a register is writable depends on the exception level the code runs at, which the compiler does not know. Writing a register that is read-only at the current level traps at run time, as the architecture defines. #### Barriers (aarch64) The three architectural barriers take the Arm ARM's option vocabulary: ```mach fragment asm aarch64 { dmb ish # data memory barrier, inner shareable dsb nsh # data synchronization barrier, non-shareable isb # instruction synchronization barrier (`isb sy` spells the same word) sb # speculation barrier, under the `sb` extension } ``` `dmb` and `dsb` take one option from `sy`, `st`, `ld`, `ish`, `ishst`, `ishld`, `nsh`, `nshst`, `nshld`, `osh`, `oshst`, `oshld`. `isb` takes no option or `sy`, the one the architecture defines. `sb` exists only on a processor with FEAT_SB, so it is an [extension instruction](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md#extension-instructions) under `sb`. None of the four reads or writes a register, so an oblivious body may spell them beside a loaded secret, and a value stays live across them. #### Exception conduits and waits (aarch64) Three instructions generate an exception at a higher level, and they differ only in which level answers: ```mach fragment asm aarch64 { svc 0 # the kernel, at EL1 hvc 0 # the hypervisor, at EL2 smc 0 # the secure monitor, at EL3 } ``` `hvc` and `smc` are how PSCI is reached, which is the only way to power off or restart a `virt` board — `SYSTEM_OFF` and `SYSTEM_RESET` go through whichever conduit the firmware provides. Both follow the SMC Calling Convention, so the compiler declares them as destroying **x0–x17**: the result and scratch registers a service may use. x18–x30 and SP survive, which makes a conduit *cheaper* than an ordinary `bl` — a procedure call destroys x0–x18 and x30. Spelling the same instruction as `.word 0xd4000002` instead costs every register in every bank, because a raw payload is a stream the parser cannot read. The two waiting hints suspend the core until something wakes it: ```mach fragment asm aarch64 { wfi # ... until an interrupt: the correct idle loop wfe # ... until an event } ``` Neither writes anything, so an idle loop holds every live value across it. `yield` is the weaker hint of the three — it asks a hypervisor to schedule elsewhere and may do nothing at all, which is why `wfi` is what an idle loop should say. #### Jumps, calls and branches (riscv64) A jump or call target is a symbol or a numeric local label, and the row decides which it may be: ```mach fragment asm riscv64 { j some_symbol # jal x0: one J-type word under R_RISCV_JAL, +-1 MiB jal some_symbol # jal ra, the same word and relocation jal t0, some_symbol # ... with a named link register call some_symbol # auipc + jalr under R_RISCV_CALL_PLT, +-2 GiB, may reach a PLT stub tail some_symbol # the same pair through t1, no link j 1f # a numeric local label resolves in place, no relocation 1: beq a0, a1, 1b # a conditional branch takes a numeric local label only } ``` `j` and `jal` name a symbol under the 20-bit `R_RISCV_JAL` field, which the linker fills from the final placement and refuses when the target is more than 1 MiB away, so a target that may live in another module or a shared object is spelled `call` or `tail`, whose `auipc` + `jalr` pair reaches +-2 GiB and, for an imported function, the PLT stub. A conditional branch (`beq`, `bnez` and the rest) takes a numeric local label only: its 12-bit field has no relocation kind here, the same posture as aarch64's `b.cond`, so a symbol in that position is refused at the asm site rather than assembled to a word the linker cannot fill. On riscv32 the rows are the same. #### Control-and-status registers (riscv64) The Zicsr extension's six instructions — read-write, read-set and read-clear, each taking its source from a register or a five-bit immediate — reach a CSR by name: ```mach fragment asm riscv64 { csrrw a0, mstatus, a1 # read mstatus into a0, write a1 into it csrr a0, mtvec # csrrs a0, mtvec, x0 - the read-only pseudo csrw stvec, a1 # csrrw x0, stvec, a1 - install a trap vector rdtime a0 # csrrs a0, time, x0 - the unprivileged counters } ``` The named set covers what freestanding code reaches for — the machine and supervisor trap vector / exception-PC / cause registers, the interrupt enable / pending pairs, the address-translation root, the hart id an SMP boot path reads to tell cores apart, and the three unprivileged counters `rdtime` / `rdcycle` / `rdinstret` name. It is deliberately not exhaustive: the privileged spec defines several hundred addresses across three privilege levels, so **any** CSR is also reachable by its numeric address, exactly as a name resolves to one: ```mach fragment asm riscv64 { csrr a0, 0xc01 # the same register as `csrr a0, time` } ``` Unlike aarch64's system registers, RISC-V spells no separate escape syntax for this — a CSR operand simply parses as the ordinary integer literal it looks like, bounded to the twelve bits a CSR address occupies. `csrrwi` / `csrrsi` / `csrrci` (and their `csrwi` / `csrsi` / `csrci` pseudos) take a five-bit unsigned immediate in the same position a register would occupy in the non-`i` form. Access permission is not checked: whether a CSR is readable or writable depends on the privilege level the code runs at, which the compiler does not know. Accessing a CSR the current level cannot reach traps at run time, as the architecture defines. #### What the compiler infers - **Operand direction.** Position within an instruction determines whether an operand is read or written. - **Clobber set.** The compiler reads each instruction, knows what registers and flags it touches, and adds them to the surrounding function's clobber set. - **Memory clobber.** Every `asm` block is conservatively assumed to modify arbitrary memory. - **Raw encodings.** A data directive is a stream the parser cannot read, so it clobbers every register in every bank unless it [declares what it writes](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md#declaring-a-raw-encodings-effects). #### Multi-arch dispatch Different architectures use different mnemonics, registers, and calling conventions. There is no nested arch-block construct inside `asm`; instead, wrap each `asm` block in `$if` on `$mach.build.arch`: ```mach fragment $if ($mach.build.arch == $mach.arch.x86_64) { asm x86_64 { ... } } $or ($mach.build.arch == $mach.arch.aarch64) { asm aarch64 { ... } } ``` The discarded branches don't compile (see [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md)), so each `asm` block only needs to be valid for its tagged ISA. #### When to use - Truly target-specific operations: syscalls, register reads, stack-frame surgery. - Anything that doesn't have a 1:1 stdlib wrapper. For ops that exist as named stdlib functions (atomics, fences, traps, SIMD long-tail), use the stdlib API — those wrappers already contain the arch-dispatched `asm`. #### See also - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#naked--no-prologue-no-epilogue-body-as-written) — `#[naked]`, whose body may hold only `asm` - [policy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md) — compiler vs stdlib boundary - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` over `$mach.build.arch` Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/policy.md ### Backend abstraction policy Where things live — compiler vs stdlib. The boundary is drawn to keep the compiler small and the stdlib readable. #### Compiler handles Things that need to feel like the language: - **Type system.** Primitive types, pointers, arrays, function types, records, unions, generics. - **Control flow.** `if` / `or`, `for`, `ret`, `brk`, `cnt`, `fin`, blocks. - **SIMD operators on primitive vector types.** Lane-wise arithmetic, bitwise, comparison-to-mask, lane indexing, and full-arity vector literals over the seeded 128-bit vector types (the honest per-operator table is in [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md)). On a target with the hardware (SSE2 on x86_64, NEON on aarch64) the compiler emits one instruction per operator; on a target without it the compiler emits a **defined unrolled scalar expansion** of the same operator — scalarize operators, never algorithms — and reports the scalarization at build time. What to do on an incapable target is the `simd` profile lever (see [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md)), not a compiler default. - **`asm` parsing, encoding, and operand allocation** for each supported ISA. - **The comptime channel** — `$mach.*` reads, the closed intrinsic set, `$if` / `$or`. - **Comptime function parameter dispatch** — turning `$name: T` parameters into per-instantiation specializations. #### Stdlib handles Things that map 1:1 to specific instruction sequences: - **Atomics** — load, store, cas, RMW family. Per arch × per ordering, via `asm` bodies inside library functions. - **Memory fences and CPU hints** — pause, prefetch. - **Traps and unreachable markers** — `trap()` is a stdlib function with per-arch `asm { ud2 / brk 0 / ... }`. - **Syscalls** — per-platform syscall ABI wrappers. - **CPU feature detection** — CPUID-style reads at runtime. - **The long tail of SIMD ops** — shuffles, reductions, gather/scatter, saturating arithmetic, specialized math. Functions over the compiler-known SIMD types. - **Bit manipulation** — popcount, clz, ctz, bswap. Wrappers around the arch-specific instruction. - **String / number formatting, parsing, math, allocators** — pure Mach built on the primitives above. #### The dividing rule > Compiler handles things that need to feel like the language. Stdlib > handles things that map 1:1 to specific instruction sequences with > predictable lowerings. The compiler grows only when something genuinely cannot be expressed as a 1:1 instruction sequence per arch — autovectorization, 128-bit arithmetic that benefits from context-dependent lowering, and similar. #### The constant-time multiply fails closed A secret `*` is the one operator whose legality is a fact about the machine rather than the source, and it is split along the same line, with one rule on both sides: **a secret multiply either reaches the machine as the declared data-independent instruction or does not run at all.** Nothing on either side substitutes a slower or leakier path. - **The compiler owns the decision.** Each instruction set declares the multiply cells it can execute in data-independent time and the condition each holds under (always, PSTATE.DIT on, or extensions selected), and the lowering gate, the `#[oblivious]` validators and `$mach.build.ct_mul` read one admission function over those rows. An undeclared cell, an unmet condition or an operating system that declares no guarantee for the mode is a compile error naming what is missing. The compiler never emits a bit-serial loop, a shift-add expansion or a call in its place, and it has no multiply strength reduction, so a secret square or a secret multiply by a constant is still the one instruction. A library that must build on every target picks its own serial path under `$if ($mach.build.ct_mul(...))` rather than being handed one. - **The stdlib owns the mode.** Where a row holds only under PSTATE.DIT, the compiler marks the module and the linker writes one byte, `__mach_dit_required`, into every executable linked for an OS that declares the guarantee. std's start code reads it before `main` and at the entry of every thread it creates: a zero byte touches nothing, a nonzero byte turns the mode on when the OS reports the processor has it, and otherwise, or when the answer cannot be read, std writes its refusal line and ends the process with status 255 before any secret is multiplied. The decision is a pure function of the byte and the availability answer, unit tested on any processor. The per-instruction-set table, the conditions and the runtime rule are in [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#constant-time-multiply-by-instruction-set) and [PSTATE.DIT at run time](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#pstatedit-at-run-time). #### Why this works This boundary minimizes compiler intrinsics. Users can read stdlib source to see exactly what their code lowers to, and fork it for exotic needs. The compiler stays small because the stdlib does most of the platform work in plain Mach with `asm` bodies. #### See also - [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md) — the `asm` form stdlib functions are built on - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` for per-arch dispatch - [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md) — what IS a compiler intrinsic vs what's deferred to stdlib Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md ### Diagnostics Every diagnostic mach prints, from the compiler, the build, the manifest, the dependency tooling, the test runner and the command line, names its kind by a key, a stable code a tool can match instead of the wording: ``` error[name.unresolved]: unresolved identifier `helpr` --> src/main.mach:4:9 | 4 | ret helpr(x) * 2; | ^^^^^ | = help: did you mean `helper`? ``` The key sits between the severity and the message: `error[]` or `warning[]`. The message says what went wrong in words, and its wording may improve from release to release. The key does not change. A tool reads the key, the location and the rest as data with [`--diagnostics=json`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md) rather than from this text. #### Keys A key is dotted and named by the subject the diagnostic is about, not by the part of the compiler that raises it: `name.unresolved`, `call.arity`, `secret.branch`, `vector.scalarize`. Its leading components name a family, so `secret` covers every `secret.*` key. A family covers whole components only: `vec` is not a family of `vector.scalarize`. One key names one rule. Every place that reports the same rule reports it under the same key, whatever the message's wording and whichever phase finds it: a secret branch is `secret.branch` whether the type checker sees it in the source or the constant-time validator sees it in the emitted instructions. A failure the compiler does not blame on the program, a defect in mach itself, is `compiler.internal`. #### The registry The keys are rows of one table in the compiler, `src/lang/diagnostic/kind.mach`. Each row gives the key and its severity, and every site that raises a diagnostic names its row. A diagnostic that names no row is refused before it is recorded, so every diagnostic the compiler can print has a key. The same rows are what a profile's [`allow`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings) and a declaration's [`#[expect]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#expectkey--acknowledge-a-warning) select by, so a key there is exactly the kind the diagnostic carries. The table is append-only: - **A key is never reused.** Once a key has named a kind, it names that kind and no other, for good. - **A key is never removed.** A kind mach no longer raises is retired: its row stays, marked retired, so its key stays reserved and can never be given to a new kind. - **A key never changes severity.** A warning key stays a warning and an error key stays an error. - **A new kind gets a new key**, added at the end of the table. A test in the compiler holds the table to these rules: it fails when a key disappears, moves or changes severity. #### Families | Family | Covers | |---|---| | `syntax` | source that does not parse: an expected token, name, type, expression or declaration is missing, or nesting is too deep | | `source` | characters the source may not contain | | `literal` | a literal that is malformed, unterminated or out of range for its type | | `name`, `use`, `module`, `visibility`, `import`, `fwd` | names that do not resolve, collide, take a built-in type's name (`name.builtin_type`) or are not exported; `use` and `fwd` paths | | `decl`, `binding`, `global`, `const` | declaration forms: bindings, globals and constants | | `type`, `cast`, `operator`, `condition`, `assign`, `address`, `ptr` | type checking: mismatches, conversions, operators, conditions, assignment and addresses | | `call`, `variadic`, `pack`, `ret` | calls, C-variadic and pack parameters, and returned values | | `field`, `index`, `range`, `shift`, `array`, `uni`, `tag`, `sel`, `guard` | records, arrays, unions, tags, `sel` and guarded places | | `generic` | type arguments and instantiation | | `comptime`, `gate`, `each` | compile-time evaluation, intrinsics and descriptors, `$if` gates, `$each`, and the user's own `$error` (`comptime.user_error`) | | `decorator`, `expect`, `extension`, `embed`, `handle`, `abi_type`, `op`, `naked`, `fin`, `ext` | decorators and the declarations they shape | | `test`, `testing` | `test` blocks and `#[testing]` declarations | | `secret` | the secrecy rules, in the source and in the constant-time validation of emitted code (`secret.not_oblivious`) | | `vector` | vector types and operations, including the scalar fallback (`vector.scalarize`) | | `asm` | inline assembly: syntax, instructions, operands, extensions, labels and locals | | `target`, `layout`, `stack`, `alloca`, `spirv` | what a target cannot realize: widths, operations, frame sizes, SPIR-V rules | | `import.unused`, `decl.deprecated`, `doc.lint`, `float.inexact`, `fwd.instances`, `debug.dropped`, `target.skipped`, `target.default_deprecated`, `expect.unfulfilled` | the warnings, listed with what raises them under [Silencing warnings](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings) | | `manifest`, `toml`, `allow`, `selection`, `need`, `template`, `version` | `mach.toml`: its keys and values, profile `allow` lists, target, profile and artifact selection, `need` entries, path templates and version ranges | | `project`, `artifact`, `output`, `source`, `path`, `glob`, `step`, `clean` | the build: finding the project, artifacts and their outputs, build steps, and `mach clean` | | `dep`, `mach`, `git` | dependencies: declaration, resolution, realization and pins, the compiler range the closure accepts, and the Git operations behind them | | `test` | the test runner, alongside `test` blocks | | `cli`, `editor` | command-line flags, commands and operands, and the editor analysis entry points | | `fs`, `process`, `env` | the machine: a file, process or environment operation that failed | | `link`, `object`, `resource` | the link: its declared inputs, undefined and duplicate symbols, relocations that overflow or are unsupported, images over a format limit, and input objects, archives, libraries or resources that are malformed, unsupported or of another format | | `catalog` | a member of a closed catalog read from input that is malformed or that the target cannot honor | | `compiler` | `compiler.internal`, a defect in mach | The full list is the table itself. #### See also - [diagnostics-json.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md) — `--diagnostics=json`, the same diagnostics as versioned NDJSON records - [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings) — silencing warnings with `allow` - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#expectkey--acknowledge-a-warning) — acknowledging a warning with `#[expect]` Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md ### Machine-readable diagnostics `mach build`, `mach check` and `mach test` take `--diagnostics=`. `human`, the default, is the rendering shown in [diagnostics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md). `json` writes everything the command reports as JSON objects, one per line of stderr (NDJSON), for editors, CI and other tools: each diagnostic, each failure raised outside the compiler's diagnostics (a manifest, dependency, build-step, link or command-line refusal), each test result under `mach test`, and a closing summary. The records are the contract: the human text may change from release to release, and a record changes only under the [stability rule](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md#stability). The flag changes no exit code. ``` mach check . --diagnostics=json ``` A bare `--diagnostics` or any other value is refused with `error[cli.flag_value]`, in human text since no format has been chosen. `-v` and `-vv`, whose phase readout is text on stderr, are refused beside `json` with a `cli.flag_conflict` failure record. #### The stream - Records go to stderr, one per line, in the order the human rendering shows the same reports. Stdout keeps what the command writes there, such as the events of `mach test --format json` or a build step's banner. - Every line mach writes to stderr is a record, a JSON object on a line of its own, and the last one is the [summary](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md#the-summary-record). Output a build step or a test writes itself is not mach's and is not a record. - The text is ASCII. A non-ASCII character in a message, a label or a path is written as a `\u` escape, and a byte that is not valid UTF-8 as `�`, so every line is valid UTF-8 and valid JSON. - Under `json` the human tally (`N errors / M warnings`) and the `building ` banners are not written. #### A record ```json {"schema":1,"record":"diagnostic","severity":"error","code":"name.unresolved","message":"unresolved identifier `helpr`","origin":"resolve","primary":{"file":"src/main.mach","line":5,"column":18,"end_line":5,"end_column":23,"byte_start":107,"byte_end":112},"related":[],"notes":[],"help":["did you mean `helper`?"],"fixes":[{"label":"replace with `helper`","edits":[{"file":"src/main.mach","line":5,"column":18,"end_line":5,"end_column":23,"byte_start":107,"byte_end":112,"replacement":"helper"}]}]} ``` A type error in a generic instance, whose primary span covers two lines and whose related site is in another file: ```json {"schema":1,"record":"diagnostic","severity":"error","code":"operator.operand_mismatch","message":"type mismatch: incompatible operand types `R` and `i64`","origin":"sema","primary":{"file":"src/main.mach","line":6,"column":16,"end_line":7,"end_column":11,"byte_start":111,"byte_end":127},"related":[{"file":"src/lib.mach","line":3,"column":9,"end_line":3,"end_column":14,"byte_start":53,"byte_end":58,"label":"in this generic body, checked against this instance's type arguments"}],"notes":[],"help":[],"fixes":[]} ``` | Member | Type | Meaning | |---|---|---| | `schema` | integer | the schema version, `1` | | `record` | string | the record type, `"diagnostic"`, or `"failure"` for a [failure](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md#failure-records) | | `severity` | string | `"error"`, `"warning"` or `"note"` | | `code` | string | the diagnostic's key from the [registry](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md#the-registry), such as `name.unresolved`; the stable identity of the kind | | `message` | string | the primary text; its wording may change, the `code` does not | | `origin` | string | the phase that produced the diagnostic, see [Origins](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md#origins); absent when mach cannot attribute it | | `primary` | span or `null` | where the diagnostic points; `null` for a diagnostic with no location, such as `target.skipped` | | `related` | array of span | other sites the diagnostic names, each span with a `label`: a string, or `null` for a site shown without one | | `notes` | array of string | the notes, in order, including a related site's label when mach cannot locate that site | | `help` | array of string | the help lines, in order | | `fixes` | array of fix | the suggested fixes, in order | A **fix** is an object with a `label` (string, what the fix does) and `edits`, an array of spans each with a `replacement` string: applying a fix replaces the bytes of every one of its edits with that edit's replacement. The edits of one fix never overlap. ##### Spans A span is an object: | Member | Type | Meaning | |---|---|---| | `file` | string | the source file: relative to the project root for a file inside it, absolute for a file outside it, in the host's path spelling | | `line` | integer | the line of the span's first byte, from 1 | | `column` | integer | the column of the span's first byte, from 1, counted in UTF-8 bytes | | `end_line` | integer | the line of the position just past the span's last byte | | `end_column` | integer | the column of that position, from 1, in UTF-8 bytes | | `byte_start` | integer | the byte offset of the span's first byte in the file, from 0 | | `byte_end` | integer | the byte offset just past its last byte | The end is exclusive, as in the Language Server Protocol: a span of `helpr` at line 5 column 18 ends at line 5 column 23. An empty span has its end equal to its start. A column counts bytes, not characters, so a tool that needs UTF-16 columns converts from the line's text. `line` and `column` are the position the human rendering shows after `-->`. `primary` may carry a `label`, a string, when a diagnostic labels its primary span. ##### Origins | Origin | Produced by | |---|---| | `build` | build configuration: the manifest's target, profile and artifact selection, and build steps | | `load` | loading and parsing the module closure, and `$if` gates decided while loading | | `resolve` | name resolution, `use`, exports and visibility | | `sema` | type checking, including a declaration's unfulfilled `#[expect]` | | `lower` | lowering to IR and the optimizer, such as `vector.scalarize` | | `codegen` | code generation and the constant-time validation of emitted code | | `link` | linking an executable or library | | `test` | the test runner of `mach test`: running the collected tests | A key names a rule, not a phase, so one key can arrive from two origins: `secret.branch` is `sema` when the type checker finds it and `codegen` when the constant-time validator does. #### Failure records A failure raised outside the compiler's diagnostics, the ones the human rendering prints as an `error[]: ` line with no source excerpt, is a record with `"record": "failure"` and the members of a diagnostic record. Its `severity` is `"error"`, its `code` the failure's key, and `notes`, `help` and `fixes` are empty. ```json {"schema":1,"record":"failure","severity":"error","code":"link.entry_missing","message":"undefined entry symbol '_start'; ensure the target startup library is linked","origin":"link","primary":null,"related":[],"notes":[],"help":[],"fixes":[]} ``` `primary` is the span in the file that caused the failure, or `null` when no file did. A manifest refusal points into the `mach.toml` that made the claim it refuses: a value it rejects, the key token of a name or key it rejects, or the table a required key is missing from. A dependency's refusal points into the dependency's own manifest, not the root that imported it. Its `file` follows the rule for every span, so a dependency's manifest is `dep/libx/mach.toml`, and `dep\libx\mach.toml` on Windows. The human rendering shows the same place on the line after the headline, as `--> ::`, naming the file by the path it was read at, as it names a source file. ```json {"schema":1,"record":"failure","severity":"error","code":"version.invalid_range","message":"mach.toml: [project].mach = \"^x\": expected a number (clause 1)","origin":"build","primary":{"file":"mach.toml","line":2,"column":8,"end_line":2,"end_column":12,"byte_start":17,"byte_end":21},"related":[],"notes":[],"help":[],"fixes":[]} ``` A refusal raised later from the parsed manifest, while a build is laid out or its steps run, points at the entry that caused it the same way: - a path template that does not expand, at its value: `[project].out`, an artifact's `out`, a local link's `path`, and a step's `argv`, `env`, `in` and `out` entries - a step output under a directory the compiler owns, at the output - a local link found neither among the step outputs nor on disk, at its `path` - an artifact that does not build for the selected target, at its `targets` - no artifact building for the selected target, at every artifact's `targets` A refusal between several entries points at the first and lists the others in `related`, each with a `null` label: - two steps declaring one output, or two artifacts writing one path, at the first claim - more than one `default = true` profile or artifact, at the first `default` value - a selection several declarations could satisfy (`native` matching several targets, several profiles or artifacts with none marked default), at the first candidate's table key - `native` matching no declared target, at the first declared target's table key, the other declared targets related - a cycle among `need` entries, at the entry of its first edge, the entries of the other edges related The human rendering shows each related place on its own `-->` line after the primary's. A target or profile named on the command line that the manifest does not declare is caused by no entry in it, so its refusal has no `primary`. ```json {"schema":1,"record":"failure","severity":"error","code":"need.cycle","message":"mach.toml: build step 'a' is part of a 'need' cycle","origin":"build","primary":{"file":"mach.toml","line":12,"column":9,"end_line":12,"end_column":17,"byte_start":141,"byte_end":149},"related":[{"file":"mach.toml","line":18,"column":9,"end_line":18,"end_column":17,"byte_start":216,"byte_end":224,"label":null}],"notes":[],"help":[],"fixes":[]} ``` `origin` names the phase the failure came from: `build` for the manifest, dependency resolution and build steps, the phase that failed for a failure inside the compiler (`link` for the linker), and `test` for the test runner. A refusal of the command line itself, such as an unknown flag or a project path that names nothing, has no `origin`. #### Test records Under `mach test` each test's result is a record, in the order the tests finish: ```json {"schema":1,"record":"test","name":"app.main#parses","module":"app.main","file":"src/main.mach","line":10,"outcome":"exit","code":3,"origin":"test"} ``` | Member | Type | Meaning | |---|---|---| | `name` | string | the test's qualified name | | `module` | string | the module that declares it | | `file` | string | the source file, spelled as a span's `file` is | | `line` | integer | the line of its declaration, from 1 | | `outcome` | string | how it ended, the `kind` `mach test --format json` reports: `"pass"`, `"exit"`, `"signal"`, `"timeout"`, `"spawn"` or `"other"` | | `code` | integer | its exit code, or `0` when it has none | | `origin` | string | `"test"` | A test record carries no `severity` and counts in no summary total; the summary's `outcome` and `exit_code` say whether a test failed. #### The summary record Every run ends with one summary record, the last line on stderr: ```json {"schema":1,"record":"summary","errors":1,"warnings":0,"notes":0,"outcome":"failure","exit_code":1} ``` | Member | Type | Meaning | |---|---|---| | `errors` | integer | the diagnostic and failure records before it with severity `"error"` | | `warnings` | integer | those with severity `"warning"` | | `notes` | integer | those with severity `"note"` | | `outcome` | string | `"success"` when the command exits 0, otherwise `"failure"` | | `exit_code` | integer | the command's exit code | #### Stability - Every record carries `"schema": 1`. - **Compatible changes keep the version:** a new member in a record or a span, a new record type, and a new value of an enumerated member (`severity`, `origin`, `record`). A tool ignores members it does not know, skips records whose `record` it does not know, and accepts an `origin` it does not know. - **Any other change bumps the version:** removing or renaming a member, changing its type or meaning, or changing the span convention. - A diagnostic or failure `code` keeps its meaning for good, under the rules of the [registry](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md#the-registry). #### See also - [diagnostics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md) — keys, the registry and its never-reused rule - [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#json-output) — `mach test --format json`, the test runner's event stream on stdout Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md ### Mach grammar (EBNF) A formal grammar for the Mach dialect, derived from the parser (`src/lang/fe/lexer.mach`, `src/lang/fe/token.mach`, and `src/lang/fe/parser/`) and cross-checked against the per-element docs in this directory. Where the live parser diverges from a doc, the divergence is called out inline. Productions that could not be fully pinned to the parser are marked `(* approximate, verify *)`. The tag productions (`tag-decl`, `tag-literal`, `sel-expr`, the `$cases` sequence and the `.[desc]` projection) are the syntax of the tagged-value contract; see [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md) for the semantic restrictions the parser does not enforce (guards, places, payload arity). `test/doc-agreement.py` checks the keyword list below against the parser's `token.mach`. This is a reference grammar, not the parser's exact control flow. The parser is a hybrid recursive-descent / Pratt climber that is malformed-input tolerant (it recovers and continues rather than failing), so some ambiguities the grammar leaves open are resolved operationally — those are noted. #### Meta-notation | Notation | Meaning | |-----------------|------------------------------------------------------| | `::=` | production definition | | `\|` | alternation | | `{ x }` | zero or more repetitions of `x` | | `[ x ]` | optional `x` (zero or one) | | `( x )` | grouping | | `"abc"` | a literal terminal (exact source characters) | | `'a'` | a single literal character | | `UPPER` | a lexical token class (terminal produced by the lexer) | | `lower` | a grammar nonterminal | | `(* ... *)` | a note / annotation | Whitespace and comments may appear between any two tokens (see [Lexical grammar](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#lexical-grammar)); they are not shown in the productions below. #### Lexical grammar ##### Tokens The lexer (`lexer.mach`) emits these token kinds (`token.mach`): ```ebnf token ::= IDENT | LIT_INT | LIT_FLOAT | LIT_CHAR | LIT_STR | punctuation | operator | EOF | ERROR (* emitted for an unexpected character, see below *) punctuation ::= "(" | ")" | "{" | "}" | "[" | "]" | ";" | ":" | "," | "." | "?" | "@" | "$" | "#[" (* attribute-open: "#" immediately followed by "[" *) | "::" | ":~" | ":>" | "..." operator ::= "+" | "-" | "*" | "/" | "%" | "&" | "|" | "^" | "~" | "=" | "==" | "!" | "!=" | "<" | "<=" | "<<" | ">" | ">=" | ">>" | "&&" | "||" ``` Notes: - There is **no dedicated keyword token**. Keywords are ordinary `IDENT` tokens; the parser recognizes them contextually by their text (`at_kw` / `eat_kw` in `parser/state.mach`). The same applies to the primitive type names and `nil`. - The maximal-munch multi-character operators are `::`, `:~`, `:^`, `:>`, `...`, `==`, `!=`, `<=`, `<<`, `>=`, `>>`, `&&`, `||`. A leading `:` lexes as `::`, `:~`, `:^`, or `:>` before a bare `:`, so `expr:~Type` is one cast, never `:` then `~`. A space after the annotation colon (`k: ^[32]u8`) keeps `:` and `^` separate, the same adjacency rule `::`/`:~` rely on, so `k: >T` is still `:` then `>`. `:>` can never take a bare `:` away from an existing program: `>` neither begins an expression nor begins a type, so a `:` immediately followed by `>` was a syntax error in every position. - The lexer recognizes the standalone characters `=`, `!`, `<`, `>`, `&`, `|` and the two-character forms above. There is no `+=`, `-=`, etc. — compound assignment does not exist. - `ERROR` (`KIND_ERROR`) is a real token kind the lexer emits for an unexpected character: the input is recorded as a `LEX_ERR_UNEXPECTED_CHAR` in the sibling error buffer, a one-byte `ERROR` token is pushed, and lexing continues. `EOF` (`KIND_EOF`) terminates every stream. Neither appears in a well-formed production; they are listed here for completeness. ##### Identifiers and keywords ```ebnf IDENT ::= ident-start { ident-char } ident-start ::= 'a'..'z' | 'A'..'Z' | '_' ident-char ::= ident-start | '0'..'9' ``` The reserved keywords (matched as `IDENT` text by the parser) are: ``` asm brk cnt def each error ext fin for fun fwd if in nil or pub rec ret sel tag test uni use val var ``` `nil` is an expression literal and `sel` is a prefix expression operator. `each`, `error` and `in` are recognized only in the comptime forms that spell them (`$each x in ...`, `$error(...)`; see [Statements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#statements)). The rest are statement, declaration, or type introducers. Note these are *contextual*: nothing in the lexer prevents a binding or field from being named after one, but the parser will treat the keyword in its keyword position. The operand-less statement keywords `brk` and `cnt` are keywords only in their bare form (`brk;` / `cnt;`); the same word followed by anything else is an ordinary identifier, so `cnt = x;` is an assignment to a variable named `cnt`. The primitive type names (`u8`, `u16`, `u32`, `u64`, `i8`, `i16`, `i32`, `i64`, `f32`, `f64`, `ptr`) are *not* keywords — they are ordinary identifiers resolved by name later. ##### Integer literals ```ebnf LIT_INT ::= dec-int | hex-int | bin-int | oct-int dec-int ::= digit { digit | "_" } [ int-suffix ] hex-int ::= "0x" hex-digit { hex-digit | "_" } [ int-suffix ] bin-int ::= "0b" ( "0" | "1" ) { "0" | "1" | "_" } [ int-suffix ] oct-int ::= "0o" oct-digit { oct-digit | "_" } [ int-suffix ] int-suffix ::= "u8" | "u16" | "u32" | "u64" | "i8" | "i16" | "i32" | "i64" digit ::= '0'..'9' hex-digit ::= digit | 'a'..'f' | 'A'..'F' oct-digit ::= '0'..'7' ``` A leading `0x` / `0b` / `0o` selects the base; the suffix is scanned on any non-float integer literal regardless of base. The lexer scans the digit run, then — when the token is not a float — takes a trailing `u`, `i` or `f` with the digits after it into the same token, so a suffix that names no integer type (`7u7`, `1f32`) is diagnosed as a bad literal rather than lexed as a number followed by an identifier. Underscores are permitted as digit separators in every base. A suffix is the spelling of the type the literal has; see [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md#typed-suffixes). ##### Float literals ```ebnf LIT_FLOAT ::= digit { digit | "_" } ( frac [ exponent ] | exponent ) [ float-suffix ] frac ::= "." digit { digit | "_" } exponent ::= ( "e" | "E" ) [ "+" | "-" ] digit { digit | "_" } float-suffix ::= "f32" | "f64" ``` A token is a float when it has a fractional part (`.` followed by a digit) and/or a scientific exponent. The `.` must be immediately followed by a digit, otherwise it lexes as a separate `.` token (so `1.field` is `1` `.` `field`, not a float). ##### Char and string literals ```ebnf LIT_CHAR ::= "'" char-body "'" LIT_STR ::= '"' { str-char } '"' ``` The lexer captures the raw span between the quotes and treats `\` as an escape that consumes the next character (so an escaped quote does not terminate the literal). Escape decoding happens later — in `comptime.eval_lit_char` / `comptime.eval_lit_str` for a comptime-evaluated literal, and in `me/lower/expr.lit_decode_escape` (an intentional mirror of the same table, #2472) for one lowered as ordinary runtime code. The recognized escapes are: ``` char escapes: \n \t \r \\ \' \0 \xHH string escapes: (char escapes) + \" ``` A string literal is single-line (see [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md)): the lexer scans from the opening `"` to the next unescaped `"`, stopping at an unescaped newline or EOF. A char literal scans likewise to the next `'`. An unterminated `'` or `"` (the scan hits a newline or EOF first) is a lexer error but still produces a token span covering the rest of that line. ##### Comments and whitespace ```ebnf comment ::= "#" { any-char-except-newline } (* but "#[" is ATTR_OPEN, not a comment *) whitespace ::= " " | "\t" | "\n" | "\r" ``` `#` begins a line comment that runs to (but not including) the next newline — **unless** it is immediately followed by `[`, which opens an attribute decorator (`#[`, `KIND_ATTR_OPEN`; see [Decorators](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#decorators)). A literal comment that begins `#[` therefore needs a separating space (`# [...]`). There is no block-comment form. Whitespace and comments separate tokens and are otherwise discarded. ##### Unexpected characters The lexer has no fall-through quote or reserved class: any byte that begins none of the token forms above is an **unexpected character**. The lexer records a `LEX_ERR_UNEXPECTED_CHAR`, emits a one-byte `ERROR` token (`KIND_ERROR`) for it, and continues. The backtick `` ` `` is not a token: it is an unexpected character wherever it appears. `#[` is the two-byte attribute-open token (`KIND_ATTR_OPEN`) that opens a decorator (see [Decorators](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#decorators) below). #### Module A source file is a single module: a sequence of declarations. ```ebnf module ::= { decl } ``` #### Decorators Zero or more leading `#[...]` decorator clauses may appear before any declaration. Each clause is a comptime directive name, optionally followed by a parenthesized argument list of comptime expressions. ```ebnf decorator ::= "#[" IDENT [ "(" [ expr { "," expr } ] ")" ] "]" decorated-decl ::= { decorator } decl ``` - Decorators attach to the **immediately following** declaration and do not bleed across declarations. - Clauses may appear on the same line (space-separated) or one per line. - The argument list is optional; a bare `#[inline]` carries no arguments. - Arguments are comptime expressions (not types): `$size_of(T)` is a valid argument; `T` as a raw type name is not. A layout intrinsic is accepted on both a global's `align` and a record/union type's, see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md). - The directive set is closed and enforced by sema, not the parser; the full list (`symbol`, `library`, `inline`, `noinline`, `align`, `packed`, `section`, `oblivious`, `scalar`, `naked`, `embed`, and the shader and target-type directives) is in [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md). - `#[...]` is the only decorator surface; a backtick is an unexpected character. #### Declarations ```ebnf decl ::= comptime-decl | flags ( use-decl | fwd-decl | fun-decl | rec-decl | uni-decl | tag-decl | bind-decl | def-decl | test-decl ) flags ::= { "pub" | "ext" } (* any order, any count; deduped to a bitfield *) ``` A leading `$` at declaration scope routes to `comptime-decl` *before* flags are parsed (so `$if` / `$`-directives cannot carry `pub`/`ext`). ##### `use` and `fwd` ```ebnf use-decl ::= "use" [ IDENT ":" ] dotted-path ";" fwd-decl ::= "fwd" [ IDENT ":" ] dotted-path ";" dotted-path ::= IDENT { "." IDENT } ``` `use` imports; `fwd` re-exports. `fwd` always publishes, so `pub fwd` is rejected by the parser (the `pub` flag is dropped and an error is emitted). ##### `def` — type alias ```ebnf def-decl ::= "def" IDENT ":" type ";" ``` ##### `rec` and `uni` ```ebnf rec-decl ::= "rec" IDENT [ generic-params ] field-block uni-decl ::= "uni" IDENT [ generic-params ] field-block field-block ::= "{" { typed-name ";" } "}" ``` `rec` is a struct (sequential layout); `uni` is a raw union (overlapping layout). Both may be generic and both share the same field-block grammar. ##### `tag` - tagged value ```ebnf tag-decl ::= "tag" IDENT [ generic-params ] ":" discriminator "{" { tag-case ";" } "}" discriminator ::= "u8" | "u16" | "u32" | "u64" tag-case ::= IDENT [ ":" type ] ``` `tag` defines a discriminated aggregate value with one active case at any time. Each case specifies a name and either one payload type or no payload. The discriminator type is mandatory and must be able to number every case. ##### `fun` — function ```ebnf fun-decl ::= "fun" IDENT [ generic-params ] param-list [ type ] ( block | ";" ) ``` - A `{ ... }` block is a defined function; a bare `;` is a forward/external signature (used with `ext`, see [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md)). - The optional `type` between the parameter list and the body is the return type. Its absence (the next token is `{` or `;`) means no return value. ```ebnf param-list ::= "(" [ params ] ")" params ::= typed-name { "," typed-name } [ "," pack-param ] | typed-name { "," typed-name } [ "," c-variadic ] | pack-param pack-param ::= IDENT ":" "..." c-variadic ::= "..." typed-name ::= [ "$" ] IDENT ":" type ``` - A trailing `name: ...` declares a **comptime variadic pack parameter** (`TYPE_KIND_PACK`); it must be the last parameter. - A trailing **bare** `...` declares a **C-variadic** signature: the preceding parameters are the fixed arity and a call may pass further arguments, placed by the target's C variadic rules (see [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#c-variadic-imports)). It is accepted only on an `ext` declaration and only after at least one fixed parameter. On any other declaration a bare `...` is not a parameter and the parser reports the missing parameter name; a comptime variadic pack is a named parameter (`va: ...`). (A function *pointer type* also carries `...`; see `fun-type-params` below.) - A leading `$` on a `typed-name` marks it a **comptime value parameter**. > **Divergence (parser vs. doc).** `typed-name` is shared between function > parameters and `rec`/`uni` fields, so the parser will *accept* a leading > `$` on a named `rec`/`uni` field too. [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) states comptime > value parameters apply to function parameters only; anonymous inline > `rec {...}` / `uni {...}` field types hardcode the comptime flag off. A > `$` on a *named-type* field parses but is not the documented surface. ##### `val` / `var` — bindings ```ebnf val-decl ::= "val" IDENT ":" type "=" expr ";" var-decl ::= "var" IDENT ":" type [ "=" expr ] ";" bind-decl ::= val-decl | var-decl ``` `val` is immutable, `var` mutable. The type annotation is mandatory for both (Mach has no type inference — see [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md)); the parser rejects a binding with no `: type`. A `val` requires an initializer; a `var` may omit it and is default-initialized. A `val`/`var` may also appear as a local statement (see [Statements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#statements)). ##### `test` ```ebnf test-decl ::= "test" IDENT block ``` A `test` carries a name and a block body. The name lives in the module's test namespace, apart from every other declaration (see [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#grammar)). ##### Generic parameters ```ebnf generic-params ::= "[" [ IDENT { "," IDENT } [ "," ] ] "]" ``` A bracketed list of bare type-parameter names with no constraints. An empty `[]` is accepted. #### Comptime declarations and directives A `$` at declaration scope is either a comptime `$if` chain or a comptime directive (`$intrinsic(args);`). ```ebnf comptime-decl ::= comptime-if-decl | comptime-directive comptime-if-decl ::= "$" "if" "(" expr ")" decl-branch-body { "$" "or" "(" expr ")" decl-branch-body } [ "$" "or" decl-branch-body ] decl-branch-body ::= "{" { decl } "}" comptime-directive ::= expr-no-assign ";" ``` - `comptime-directive` is a bare **comptime intrinsic / directive call** (`$error("msg");`). Per-declaration codegen attributes are written as `#[...]` decorators (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)); a directive takes no `=`, so a stray one after the target is a parse error at the directive's `;`. The target is parsed at a binding power above assignment so that `=` is never swallowed into the expression. - `expr-no-assign` is `expr` parsed with the assignment operator excluded at the top level (binding power >= 2; see [Expressions](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#expressions)). It still begins with the leading `$` because the first prefix atom is a `comptime-ident`. The `$if` chain also exists as a **statement** form with the same shape but a statement-list body (see [Statements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#statements)). #### Types ```ebnf type ::= secret-type | ptr-type | array-type | fun-type | rec-type | uni-type | pointee-of-type | field-type | named-type secret-type ::= "^" type ptr-type ::= "*" type array-type ::= "[" ( expr | "_" ) "]" type pointee-of-type ::= "$" "pointee_of" "(" type ")" field-type ::= IDENT "." "type" named-type ::= dotted-path [ type-args ] type-args ::= "[" [ type { "," type } [ "," ] ] "]" fun-type ::= "fun" "(" [ fun-type-params ] ")" [ type ] fun-type-params ::= "..." | type { "," type } [ "," "..." ] rec-type ::= "rec" anon-field-block uni-type ::= "uni" anon-field-block anon-field-block ::= "{" { IDENT ":" type ";" } "}" ``` Notes: - `^T` marks `T` as carrying secret data for the constant-time guarantee (#1643). It is a prefix qualifier binding to the type immediately to its right, so it nests with `*` and `[N]` in any order (`*^u8`, `^*u8`, `[N]^u8`). The grammar accepts it in every type position. Its flow-typing semantics - the secrecy lattice and join, the leakage-model gates, and the welded-storage pointer rules - are enforced by sema (#1645) and documented in [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md). - `*T` is a pointer; the untyped pointer type is the primitive name `ptr` (an ordinary `named-type`, not its own syntax). - `$pointee_of(T)` is the type a typed reference `*U` refers to (#2693). It is a type **constructor** rather than an intrinsic call, which is what lets it nest inside another intrinsic's operand and inside a generic argument list. `ptr`, `^*U`, and any non-reference operand are refused by sema, each with its own cause — see [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md). - `f.type` is a field descriptor's own type inside a `$each` body (#2691), and is a type spelling in its own right so a constructor can take one (`$pointee_of(f.type)`). `type` is contextual, not a keyword: a module or record member genuinely named `type` still reads as a path (`mod.type.T`, `mod.type[A]`). - `[N]T` is a fixed-length array; `N` is a full expression (a comptime constant), including `$size_of(T)` / `$align_of(T)` — see [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md). Nesting (`[N][M]T`) falls out of the recursion. - `[_]T` is an inferred array length (#2507): the length is not written but taken from elsewhere. It is legal **only** on a `val` carrying `#[embed(...)]`, where the length comes from the embedded file's byte count; written anywhere else it is rejected — "an inferred array length `[_]` is only valid on an `#[embed(...)]` declaration". See [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#embedstr--compile-time-file-embedding). - `named-type` covers both plain names (`i64`, `Point`) and generic instantiations (`Pair[i64, u8]`, `Map[str, u32]`). The dotted path allows module-qualified names (`core.Thing`). - A `fun(...)` type's return type is optional; it is omitted when the next token cannot begin a type (`{`, `;`, `,`, `)`, `]`, `=`, EOF). - `fun`-type parameter lists carry a `variadic` flag via a trailing `...`, mirroring `fun` declarations. - Anonymous inline `rec {...}` / `uni {...}` types use a field block whose entries do **not** accept the leading `$` comptime marker. There is **no `?T` option-type sugar and no special Result keyword sugar in the grammar.** `?` is exclusively the prefix address-of operator (below). The std failure tags `res[T, E]`, `opt[T]`, and `err[E]` are ordinary `named-type` generic instantiations of declarations imported from `std.types.result`, `std.types.option` and `std.types.error`; the parser and the compiler know nothing of the three names. #### Expressions Expressions are a Pratt climber over a primary (prefix + postfix) grammar. ```ebnf expr ::= prefix { postfix } { binary-op prefix { postfix } } ``` (The braces above are illustrative; the real binding is precedence-driven — see the precedence table.) ##### Prefix (atoms and unary) ```ebnf prefix ::= LIT_INT | LIT_FLOAT | LIT_CHAR | LIT_STR | "nil" | IDENT | comptime-ident | typed-literal | tag-literal | array-literal | sel-expr | unary-op prefix { postfix } | "(" expr ")" sel-expr ::= "sel" prefix { postfix } comptime-ident ::= "$" IDENT unary-op ::= "-" (* numeric negation *) | "!" (* logical not *) | "~" (* bitwise not *) | "?" (* address-of *) | "@" (* dereference *) ``` - A unary operator binds its operand as `prefix` followed by any postfix chain, so `@p.field` and `?arr[i]` apply member/index *inside* the unary. - `sel` binds the same way and requires the result to be a member access or a descriptor projection (`sel r.ok`, `sel v.[c]`), so `sel r.ok && r.ok > 3` parses as `(sel r.ok) && (r.ok > 3)`; any other operand is a parse error (`` `sel` tests one case of a place: write `sel place.case` or `sel place.[case]` ``). - `(expr)` is a plain grouping; there is no tuple form. ##### Postfix ```ebnf postfix ::= call-args | generic-call | index | member | project | cast call-args ::= "(" [ call-arg { "," call-arg } [ "," ] ] ")" call-arg ::= expr [ "..." ] (* trailing "..." makes it a va... spread *) generic-call ::= type-args call-args (* callee[T, U](args) *) index ::= "[" expr [ "," expr ] "]" (* x[i], or the range x[start, count] *) member ::= "." IDENT project ::= "." "[" expr "]" (* v.[f]: comptime field projection *) cast ::= ( "::" | ":~" ) type | ":>" type (* secret-qualifier strip cast *) ``` `::` is a value conversion and `:~` a same-size bit reinterpret; see [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#cast). Neither `::` nor `:~` may add or drop the `^` secret qualifier. `:>` is the only downgrade: it strips `^` from the operand's type, producing a new public value (#1643, [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md)). Its target type is required and names the operand's stripped public type; a bare `:>` is a parse error, and `:>` never reinterprets storage. `:^` is not a token: `x:^u32` lexes as `x`, `:`, `^`, `u32`, the postfix chain stops at the colon, and the enclosing statement reports its own terminator error there. All casts bind as postfix. The second `expr` of an `index` makes it a range: `count` elements or lanes from `start` ([operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md#range)). `count` is a length, never an end index. Disambiguating a postfix `[`: the bracket may open a generic argument list (`callee[T, U](args)`, or `f[T]` naming an instance as a value) or be an index (`obj[idx]`, a range `obj[start, count]`, or the index of an index-then-call `callee[idx](args)`). The two grammars overlap — a bare name is both a type and a value — so the parser cannot decide locally. It probes the payload against both grammars, and the same four outcomes apply whether or not a `(` follows the `]`: - reads cleanly **only** as a type list (`*T`, `[N]T`, a list of three or more, or a nested generic like `res[bool, ParseError]`) → generic arguments; - reads cleanly **only** as an expression (`i + 1`, `f(x)`, a literal, or a range with such a part like `p[i, 4]`) → an index, so `table[0]()` and `table[i + 1]()` are index-then-call; - reads cleanly as **both** (a bare identifier, a dotted path, a generic application like `v[k]`, or two of those separated by a comma like `v[a, b]`) → the decision is deferred to name resolution, which picks by the object's resolved kind. In **callee** position a value (or any non-name callee, e.g. an index/call result) makes `[x]` an index, while a function, type, or imported name takes `[x]` as a type argument. In **value** position the predicate is narrower — only a name that resolves to a *function declaration* takes type arguments — because with no `(` to key on every subscript in the language reaches this path, and `mod.ARR[i]` must stay an index. So `table[i]()` calls the indexed function pointer, `make[T]()` instantiates the generic, and `make[T]` names that instance as a value. A two-part payload follows the same rule: `v[a, b]` on a value is a range, and `f[T, U]` on a generic is its type arguments, in either position; - reads cleanly as **neither** → the committing index parse reports the error (never a silent wrong parse). In value position, a payload that is a type and nothing else has no index reading to fall back on, and is reported against the object instead. - The struct-literal `Name[T]{...}` form is recognized at the **prefix** stage (`typed-literal`) and never reaches postfix. ##### Literals with a type prefix ```ebnf typed-literal ::= named-type "{" [ member-init { "," member-init } [ "," ] ] "}" tag-literal ::= named-type ( "." IDENT | "." "[" expr "]" ) "{" [ expr ] "}" array-literal ::= array-type "{" [ expr { "," expr } [ "," ] ] "}" member-init ::= IDENT ":" expr | expr ``` - `typed-literal` is a record or union literal: a named type (optionally generic) followed by a brace-delimited initializer list, where each member is a `field: value` pair (`Point{ x: 1, y: 2 }`, `Pair[i64, u8]{ left: 5, right: 6u8 }`). Numeric vector types use positional expressions, as in `f32x4{1.0, 2.0, 3.0, 4.0}`. - `tag-literal` names the type, the case and the payload (`Reply.value{42}`, `Reply.empty{}`, `res[i64, E].ok{42}`). The payload is positional and exactly one; a payloadless case takes empty braces. The descriptor form `T.[case]{...}` names the case through a comptime case descriptor bound by `$each` and obeys the same single-case rule after specialization. The parser commits to a literal via a lookahead (`Name (.Name)* ([...])? (.Name | .[...])? {`); where the head reads as `A.b` the resolver decides whether `b` is a case name or the final segment of a type path. The parser accepts named, positional and bare initializers in every literal; sema enforces that a tag literal carries at most one positional payload and a record literal names its fields. - `array-literal` is `[N]T{ e0, e1, ... }`, an array type followed by a brace-delimited positional element list. ##### Operator precedence From `token.infix_precedence` / `token.is_right_assoc`. Higher binds tighter; all binary operators are left-associative **except** `=` (assignment), which is right-associative. Unary prefix operators and the postfix chain (call/index/member/cast) bind tighter than every binary operator. | Prec | Operators | Assoc | Meaning | |------|------------------------|--------|---------------------------------| | 11 | `*` `/` `%` | left | multiply, divide, remainder | | 10 | `+` `-` | left | add, subtract | | 9 | `<<` `>>` | left | shift left / right | | 8 | `<` `<=` `>` `>=` | left | relational | | 7 | `==` `!=` | left | equality | | 6 | `&` | left | bitwise and | | 5 | `^` | left | bitwise xor | | 4 | `\|` | left | bitwise or | | 3 | `&&` | left | logical and (short-circuit) | | 2 | `\|\|` | left | logical or (short-circuit) | | 1 | `=` | right | assignment | ```ebnf binary-op ::= "*" | "/" | "%" | "+" | "-" | "<<" | ">>" | "<" | "<=" | ">" | ">=" | "==" | "!=" | "&" | "^" | "|" | "&&" | "||" | "=" ``` Assignment is an expression (precedence 1), not a statement form; an assignment statement is just an expression statement whose top operator is `=`. #### Statements ```ebnf stmt ::= block | if-stmt | for-stmt | ret-stmt | brk-stmt | cnt-stmt | fin-stmt | asm-stmt | local-decl-stmt | comptime-if-stmt | comptime-each-stmt | expr-stmt block ::= "{" { stmt } "}" if-stmt ::= "if" "(" expr ")" block { or-arm } [ or-else ] or-arm ::= "or" "(" expr ")" block or-else ::= "or" block for-stmt ::= "for" [ "(" expr ")" ] block (* no condition => infinite loop *) ret-stmt ::= "ret" [ expr ] ";" brk-stmt ::= "brk" ";" cnt-stmt ::= "cnt" ";" fin-stmt ::= "fin" block (* defer: runs block at enclosing-block exit *) local-decl-stmt ::= bind-decl (* a "val"/"var" used as a statement *) expr-stmt ::= expr ";" comptime-if-stmt ::= "$" "if" "(" expr ")" stmt-branch-body { "$" "or" "(" expr ")" stmt-branch-body } [ "$" "or" stmt-branch-body ] stmt-branch-body ::= "{" { stmt } "}" comptime-each-stmt ::= "$" "each" IDENT "in" expr stmt-branch-body ``` `$each` is a compile-time unroll: the body is duplicated once per element of the sequence, which must be `$fields(T)`, `$cases(T)`, a variadic pack identifier, or a comptime-constant array `val` (see [comptime-intrinsics.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-intrinsics.md)). `in` is a contextual keyword. Notes: - `if` / `or`: the `or` chain models both `else if` (`or (cond) { ... }`) and `else` (`or { ... }`). Arms are parsed greedily; an `or` with a condition continues the chain, an `or` without one terminates it. Each arm body is a block. - `fin { ... }` registers a block to run when the enclosing block exits — Mach's defer, block-scoped (see [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) for the exit rules). The body must be a block; the bare `fin stmt;` form is rejected. - `comptime-if-stmt` is the statement-scope `$if`/`$or` chain (the declaration-scope variant is under [Comptime](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/grammar.md#comptime-declarations-and-directives)). A `$` only begins this form when the next token is the keyword `if`; otherwise a leading `$` at statement position is parsed as an `expr-stmt` whose first atom is a `comptime-ident`. ##### Inline assembly ```ebnf asm-stmt ::= "asm" IDENT "{" asm-body "}" ``` - The `IDENT` after `asm` is the mandatory **ISA tag** (`x86_64`, `aarch64`, … — a closed set). Bare `asm { ... }` is rejected. - `asm-body` is **raw text**, not a token grammar: the parser captures the source span between the opening `{` and its brace-matched `}` and hands it to the backend verbatim. Nested `{ }` are balanced by depth. Local substitution uses `{name}` references inside the body; operand direction and clobbers are inferred from the instruction stream, so there is no operand or clobber list. `#` introduces a line comment inside the body. ```ebnf asm-body ::= (* raw source text, brace-balanced; not tokenized *) ``` #### Comptime surface (syntactic forms) These are not separate grammar productions — they reuse `comptime-ident`, `call-args`, and `member` — but are listed here as the recognized comptime shapes for reference. They are accepted syntactically; which ones the compiler actually resolves is a semantic concern (several are documented stubs). ```ebnf comptime-ident ::= "$" IDENT (* $size_of, $mach, $type_of, ... *) intrinsic-call ::= comptime-ident call-args (* $size_of(T), $fields(T), $type_of(e) *) mach-read ::= comptime-ident { member } (* $mach.build.os, $mach.arch.x86_64 *) ``` - Intrinsic calls (`$size_of(T)`, `$length_of(T)`, `$align_of(T)`, `$offset_of(T, field)`, `$type_of(e)`, `$fields(T)`, `$cases(T)`, `$discriminant_of(T)`, `$is_tag(T)`, `$is_record(T)`, `$is_union(T)`, `$is_pointer(T)`, `$is_secret(T)`, `$holds_secret(T)`, `$type_name(T)`, `$type_id(T)`, `$error("msg")`) are syntactically a `comptime-ident` callee with `call-args`. - The **type-taking** intrinsics — `$size_of`, `$length_of`, `$align_of`, `$offset_of`, `$fields`, `$cases`, `$discriminant_of`, and the type predicates with `$type_name` and `$type_id` — parse their **first argument with the `type` production**, not the expression grammar, so the whole type language is spellable there: `$fields(Box[T])`, `$size_of(Pair[A, B])`, `$size_of(*T)`, `$size_of([4]u16)`, `$size_of(^u32)`, `$fields(mod.Rec)`. Every other argument is an ordinary expression; `$offset_of`'s second is a bare field name resolved against the record or tag payload case. `$type_of(e)` takes a value expression and produces a comptime type value. - A **type comparison** operand (`$type_of(x) == Name`) is the one type spelling read with the expression grammar, because the comparison is only recognizable once both sides are parsed. Only a bare name or `module.Name` is available there; a generic instance is not. - `$mach.*` reads are a `comptime-ident` followed by a `.`-member chain. They appear in `$if` conditions and as comptime initializers. - A bare **directive** `$intrinsic(args);` (e.g. `$error("msg");`) is the `comptime-directive` declaration form above. The `$mach.*` tag/path set is closed and documented in [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md); this grammar treats every `$ident` and `.`-chain off it uniformly. ##### Field projection (`v.[f]`) `v.[f]` is a postfix form that projects the field described by a comptime field descriptor `f` (bound by a `$each f in $fields(T)` loop) off an instance `v`. Syntactically: `.` followed immediately by `[expr]`. It is disambiguated from a regular member access `v.name` by the `[` lookahead: `.` then `[` = projection; `.` then `IDENT` = member. #### Verification notes Productions verified directly against the parser source: - **Lexical grammar** — `lexer.mach` / `token.mach`: token set (incl. `KIND_ATTR_OPEN`), operator maximal-munch, number/char/string scanning and escapes, comment and whitespace handling (incl. the `#[` attribute-open exception), the "keywords are `IDENT`s" model. - **Precedence ladder** — `token.infix_precedence` / `token.is_right_assoc` (the table is a direct transcription; only `=` is right-associative). - **Decorators** — `parser/grammar.mach` `parse_decorators` / `parse_one_decorator`: leading `#[name(args)]` clauses (one Decorator node), closed directive set. - **Declarations** — `parser/grammar.mach`: `use`, `fwd` (incl. `pub fwd` rejection), `fun` (generics, params, variadic `...`, named pack `name: ...`, comptime `$` params, optional return type, block-or-`;` body), `rec`, `uni`, `tag` (mandatory discriminator, cases with an optional payload type), `val`/`var` (type annotation required; `val x = 42;` is rejected), `def`, `test`, `flags` (`pub`/`ext` any order/count), the decl-scope `$if`/`$or` chain, and the `comptime-directive` (attribute-write vs. bare directive) form. - **Statements** — `parser/grammar.mach`: `block`, `if`/`or` chain, `for` (optional condition), `ret`/`brk`/`cnt`/`fin`, local `val`/`var`, the stmt-scope `$if`/`$or` chain, `$each … in … { }`, and `expr-stmt`. - **Expressions** — `parser/grammar.mach`: prefix atoms, `sel`, all five unary prefix operators (`-`, `!`, `~`, `?`, `@`), the postfix chain (call with optional `...` spread on arguments, generic-call, index, member, field projection `.[f]`, cast), struct/array/tag literals and the typed-literal lookahead, the generic-call-vs-index `[` disambiguation, and `comptime-ident`. - **Types** — `parser/grammar.mach`: `*T`, `[N]T`, `fun(...) R` (with variadic, pack `name: ...`, and optional return), anonymous `rec {...}` / `uni {...}`, and named types with generic args / dotted paths. - **Inline asm** — `parser/iasm.mach`: mandatory ISA tag, raw brace-balanced body, no operand/clobber list. Doc-only (intended surface, not a distinct parser production): - The concrete escape sets (`\n \t \r \\ \' \0 \xHH`, plus `\"` for strings) — the lexer only treats `\` as "consume next char"; the actual escape set is decoded in `comptime.eval_lit_char` / string lowering and documented in [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md). - The closed ISA-tag set (`x86_64`, `aarch64`, `riscv64`, `riscv32`) and the closed `$mach.*` tag/path set — the parser accepts any `IDENT` / `$`-chain; the closed sets are enforced later (see [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md), [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md)). - The closed intrinsic set (`$size_of`, `$length_of`, `$align_of`, `$offset_of`, `$type_of`, `$fields`, `$cases`, `$discriminant_of`, `$is_tag`, `$is_record`, `$is_union`, `$is_pointer`, `$is_secret`, `$holds_secret`, `$type_name`, `$error`) — syntactically indistinguishable from any other `comptime-ident` call. - The closed decorator directive set ([decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)) — the parser accepts any `IDENT` after `#[`; sema enforces the closed set. Divergences flagged inline: - `$` comptime marker is grammatically accepted on named `rec`/`uni` fields (shared `typed-name`), though [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) scopes comptime value parameters to functions only. No production above is left unverified against the parser; nothing here is invented. The only "approximate" surface is the asm body, which is deliberately *not* a token grammar (it is raw text by design). #### See also - [literals.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/literals.md) — literal forms and escapes - [types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md) — type semantics - [operators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/operators.md) — operator semantics and precedence prose - [expressions.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/expressions.md) — expression composition - [statements.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/statements.md) — control-flow semantics - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) — function declarations, generics, variadics, comptime params - [comptime-control.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-control.md) — `$if` / `$or` - [comptime-mach.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md) — `$mach.*` namespace - [asm.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/asm.md) — inline assembly Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/documentation.md ### Documentation Mach source-level documentation uses `#` comments immediately above declarations. Each docstring is one summary followed by an optional component block. `mach doc` renders them, and the compiler's docstring lint checks the component block of every `pub fun`, `pub rec`, `pub uni`, and `pub tag` against the declaration it documents. A docstring states **what** a declaration is and **how** it is used, and a module's docstring states the contract the module holds its callers to. The language, command and manifest references under `doc/` say what the user sees; when something changed belongs in the changelog. Documentation never changes generated code. #### Grammar ``` # # --- # : # : ``` - **Summary.** `` is a single sentence by convention. Additional paragraphs may follow, separated by blank `#` lines, and the summary runs to the separator or to the end of the docstring. - **Separator.** `# ---` is present when one or more component lines follow; absent when the docstring is summary-only. - **Component lines.** Each is `# : `. One line per element the declaration exposes, in declaration order. A component head has zero or one space between `#` and the name; a line with two or more spaces after `#` is a continuation of the previous description, which is how a long description wraps. - Prose is lowercase except for proper nouns and type names. #### Component identifiers | Declaration element | Component identifier | |---|---| | Function parameter | parameter name | | Comptime parameter | the `$name` form | | Generic type parameter | the `[T]` form | | Return value | `ret` | | Record field | field name | | Union variant | variant name | | Tag case | case name | #### What the lint checks For a `pub fun`, `pub rec`, `pub uni`, or `pub tag` whose docstring has a component block, every component line must name an element of that declaration, carry a description, and appear in declaration order (generics, then parameters, fields, or cases, then `ret` for functions). Each violation is a warning naming the line: ``` documented component matches no parameter, field, case, generic, or `ret` of this declaration documented component has no description documented components are out of declaration order ``` Tags have no return value, so a `ret:` component line on a tag is rejected. A bare generic name without brackets (such as `T:` instead of `[T]:`) is also rejected. A summary-only docstring, a docstring on a `val`, `var`, `def`, `use`, or `fwd` (which are summary-only forms), and a non-`pub` declaration are not checked. A component block need not be complete: an element with no line is not a warning, a line with no element is. `mach doc` renders every `pub` declaration whether or not it is documented. #### Placement The lexer folds line-adjacent `#` lines that each start their own line into one run, and a run becomes a declaration's docstring when it ends on the line directly above the declaration or directly above the declaration's `#[...]` decorators. A blank line between the run and the declaration breaks the attachment, and a decorator line is not a comment, so it never joins a run. A comment that shares its line with code (`val x: i32 = 1; # note`) is not documentation: it neither starts a run, nor joins the comment on the next line, nor attaches to the declaration below it. The docstring is therefore the first thing above the declaration, with decorators between it and the declaration: ```mach fragment # terminate the program with a message # --- # msg: text to emit before terminating #[symbol("panic")] pub fun panic(msg: *u8) { ... } ``` #### Function ```mach fragment # read the wall-clock time # --- # out: pointer to Timespec to populate # ret: 0 on success, negative errno on failure pub fun realtime(out: *Timespec) i64 { ... } ``` Generic and comptime parameters appear in the component block under their syntactic form: ```mach fragment # atomic load through a typed pointer # --- # [T]: element type # $order: memory ordering constraint # ptr: pointer to load from # ret: loaded value pub fun load[T]($order: Order, ptr: *T) T { ... } ``` Summary-only, with no separator and no component block: ```mach fragment # yield the CPU to other threads pub fun spin_hint() { ... } ``` #### Record / union / def ```mach # a 2D Cartesian point with i64 coordinates # --- # x: horizontal coordinate # y: vertical coordinate pub rec Point { x: i64; y: i64; } ``` ```mach # holds either an integer or a float # --- # i: integer interpretation # f: float interpretation pub uni Number { i: i64; f: f64; } ``` ```mach # an i64 representing years since birth pub def Age: i64; ``` #### Tag ```mach # a value that may be absent # --- # [T]: value type # none: empty case # some: payload case pub tag Maybe[T]: u8 { none; some: T; } ``` ```mach # a unit outcome with a typed failure # --- # [E]: error type # err: failure case # ok: success case pub tag Outcome[E]: u8 { err: E; ok; } ``` Cases appear in declaration order after generic parameters. Tags have no return value, so `ret:` is refused. Undocumented cases are permitted, but any documented case must exist on the tag; `doclint` warns at the component otherwise: ```mach # a unit outcome with a typed failure # --- # [E]: error type # err: failure case # done: no such case pub tag Outcome[E]: u8 { err: E; ok; } ``` #### Module A `.mach` file begins with a module docstring as the first content in the file, before any `use`, `fwd`, or decorator. ```mach # cross-platform OS interface # # forwards the portable intersection of all supported targets. # for platform-specific functionality, import the target module # directly (e.g. myproj.system.os.linux). ``` Modules may extend beyond the summary with additional paragraphs separated by blank `#` lines. Other declaration kinds do not extend beyond the summary and the component block. #### Value ```mach # maximum counter value before saturation pub val MAX: i64 = 100; # module-local request counter var calls: i64 = 0; ``` #### Comments that are not docstrings A comment that is not directly above a declaration is an ordinary comment. Comments are brief, lowercase, and single-line, and appear only where the code is not self-evident. There are no sectional or separator comments. A line comment that begins `#[` with no space opens a decorator; write such a comment as `# [...]`. #### See also - [fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/fun.md) - function declaration grammar - [rec.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/rec.md), [uni.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/uni.md), [tag.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/tag.md), [def.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/def.md) - type forms - [val-var.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/val-var.md) - binding declarations - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) - module structure and file layout - `mach help doc` - the command that renders docstrings Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md ### `mach.toml` — the project manifest A Mach project is described by a `mach.toml` at its root: its identity, the platforms it targets, the artifacts it produces, the build variants it offers, its external link requirements, the build steps that produce them, and its dependencies. Every `mach` subcommand takes the project explicitly — a directory (whose `mach.toml` is read) or a manifest file directly — so the manifest a build uses is never guessed from the working directory. `mach help ` describes the path argument. The manifest is built from `[category.name]` tables in seven sections — `[project]`, `[target.X]`, `[profile.X]`, `[artifact.X]`, `[link.X]`, `[step.X]`, and `[dep.X]`. TOML itself enforces name uniqueness within a section. #### Totality A manifest is required. A project directory without one does not build: ``` error[project.no_manifest]: no mach.toml in the project directory ``` Nothing is inferred from the directory layout. `mach init` writes a complete manifest so a new project never starts from that error (`mach help init`). A table you *declare*, you declare completely. Every field of a declared table is required; a missing field is a strict-parse error, not a silent default. This is the manifest twin of Mach's explicitness: there are no field defaults to memorize, because "any" and "none" are said out loud — - `"*"` is the explicit any-token for a filter axis or `targets` entry; - `[]` is the explicit empty list ("none"). The sole exception is **shape-dependence**: a field whose presence follows another value in the same table. A dependency is `git` *or* `path`; a `[link.X]` names a `name` *or* a `path` according to its `source`. Nothing else defaults. A manifest refusal points at what it refuses, on the line after the message: the value it rejects, the key token of a name or key it rejects, or the table a required key is missing from, in the manifest that states it, a dependency's own included. [`--diagnostics=json`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics-json.md#failure-records) carries the same place as the failure's `primary` span. ``` error[manifest.invalid_value]: mach.toml: profile 'debug': opt must be 0, 1, or 2 (got 7) --> mach.toml:15:7 ``` Unknown sections and unknown keys are always errors. A path value is always `/`-separated; a literal `\` is rejected (`manifest paths use '/'`), so the same manifest is portable and is normalized to the host separator at the filesystem boundary. ##### Root vs. dependency strictness A dependency's `mach.toml` is read by the same closed schema, once, and an unknown or removed key in it fails the consumer's build naming the dependency: ``` error[manifest.unknown_key]: dep 'std': mach.toml: unknown key 'bogus' in [project] ``` What a consumer *uses* from a dependency's manifest is its export surface: the project id, the module a bare `use ;` binds (see [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md#bare-project-id-imports)), its `export = true` link entries, the steps those entries demand, and what its `default = true` library artifact requires — the artifacts and steps named in that artifact's `need` (see [Dependency requirements travel](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#dependency-requirements-travel)). Nothing else travels: a `bin` artifact's `need`, a non-default library's, and every other requirement of the dependency stay its own. A dependency's `[profile.*]` tables are never read to build the consumer, which resolves its own profile and builds everything with it. A dependency's `[target.*]` tables are read for exactly one purpose: the targets its travelling requirements name, `env` included, since those artifacts are built for the targets the dependency declares for them. No other `[target.*]` entry is read. #### The schema at a glance ```toml [project] id = "demo" # required: identifier; root of every module path version = "0.1.0" # required mach = "^5.3" # required in a root manifest: the compiler range src = "src" # required: source dir, project-root-relative out = "out/{target.name}/{profile.name}" # required: output-path template root [target.linux] # a platform: a fully-spelled tuple isa = "x86_64" os = "linux" abi = "sysv64" [profile.debug] # a build variant; at least one is required opt = 0 # 0 (debug pipeline) | 1 | 2 (release pipeline) debug = true # emit debug info for this profile simd = "scalarize" # SIMD lever: "scalarize" | "require" vectorize = false # auto-vectorization lever float_reassoc = false # float reassociation permission # allow = ["import.unused"] # optional: warning keys this profile silences [artifact.demo] # a produced artifact kind = "bin" # "bin" | "static" | "shared" entry = "main.mach" # entry source, relative to src out = "bin/demo{artifact.suffix}" # output path, relative to the project out targets = ["*"] # which declared targets build it ("*" = all) link = [] # [link.X] names this artifact links need = [] # step.X / artifact.X requirements # subsystem = "gui" # optional: windows console/GUI selector # icon = "assets/demo.ico" # optional: PE executable icon # manifest = "assets/demo.manifest" # optional: PE application manifest [dep.std] # a dependency git = "https://github.com/briar-systems/mach-std" ref = "branch/main" ``` `[link.X]` and `[step.X]` each have their own section below. #### `[project]` | Key | Type | Meaning | |-----------|--------|---------| | `id` | string | Root segment of every module path the project exposes: a file at `/foo/bar.mach` is reachable as `.foo.bar`. Must be a plain identifier — letters, digits, `_`, `-` — since it names the dependency store and keys step stamp files. Read by `$project.id`. | | `version` | string | Project version. Read by `$project.version` and `$project.version.{major,minor,patch}`, and stamped into a Windows executable's version resource. | | `src` | string | Source root, project-root-relative. Module paths resolve under it. | | `out` | string | The output-path template root, referenced as `{project.out}` by artifact `out`, step paths, and `cmd`s. Expanded over `{target.name}`/`{target.isa}`/`{target.os}`/`{target.abi}`/`{profile.name}` (see [Path templates](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#path-templates)). | | `mach` | string | The compiler versions this project builds with, as a [version range](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#version-ranges) (`"^5.3"`). Required in a root manifest. See [Compiler range](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#compiler-range). | `[project]` is exactly these five keys. Any other key, `name` and `description` included, is an unknown-key error (`mach.toml: unknown key 'name' in [project]`), in a root manifest and a dependency's alike. `[profile.]` likewise carries no `emit_ir` or `emit_asm`: emission is `--emit-ir`/`--emit-asm` on the command line. ##### The output directory Everything a build and its cache write lands under the expanded `out`, in one layout: | Path | Holds | | --- | --- | | `obj/` | one object per module, each carrying its cache key; it is the [object cache](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps) and is read by developers and tooling too | | `ir/`, `asm/` | the human-readable views `--emit-ir` and `--emit-asm` write | | `.cache/` | compiler-only state, read and written by nothing but the compiler | | `.cache/steps/` | one fingerprint stamp per [build step](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps) | | `.stage/` | build step scratch space, one directory per step, reset before the step runs | | `test//dispatch.o` | the test dispatcher object of a tested artifact | | `test//` | the test dispatcher executable | | `test//log/` | a failing test's captured output | | `dep//` | a dependency's artifact outputs (see [Dependency requirements travel](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#dependency-requirements-travel)) | Test objects sit in `obj/` beside the module objects, as `obj//.test.o`. Artifact outputs go wherever their own `out` names under the directory. `mach clean` removes `obj/`, `ir/`, `asm/`, `.cache/`, `.stage/`, `test/` and `dep/` along with every artifact output, for every declared target and profile, so the build after it is a cold one that reuses nothing. A step output may not name a path inside `obj//`, `.cache/` or `.stage/` (see [build steps](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps)). ##### Compiler range `mach` states which compilers a project builds with, and every command that reads the dependency closure (build, test, check, `mach dep verify`, the language server) checks it for the root and for every realized dependency. A compiler outside any of those ranges is refused once, with every unmet requirement and the chain that states it, pointing at the first unmet range in the manifest that states it: ``` error[mach.version_unaccepted]: this is mach 5.2.1, and the dependency closure does not accept it: app (mach.toml) requires mach ^5.3 app -> gfx -> glfw requires mach >=5.4, <6 --> mach.toml:2:11 ``` A root manifest must state `mach`. One without it is refused with the line to add (`mach.toml: [project] states no compiler range; add mach = "^5.3", the oldest release that reads the key, and raise it when the project uses a later feature`), pointing at its `[project]` header. A dependency without it states no constraint. `mach init` writes the same range. It is the oldest release of the running compiler's major that reads the key: `^5.3` for every 5.x compiler, since 5.3.0 is the first release that accepts `mach`, and `^N.0` for a later major N, since a caret cannot span majors. The range depends only on the running major, so two authors on one project write the same line. The compiler's version is the last release it was built from. A build from an unreleased tree reports that release, so a project cannot require an unreleased feature by version: a feature is a compatibility promise only once it is released. ##### Version ranges `[project].mach` and `[dep.].version` share one range grammar, defined here and pinned by the compiler's tests: ``` range = clause *( "," clause ) ; the intersection of every clause clause = op partial op = "^" / "~" / ">=" / ">" / "<=" / "<" / "=" partial = major [ "." minor [ "." patch [ "-" pre ] ] ] ``` A version satisfies a range when it satisfies every clause. Each clause names its operator, so a bare `1.2` is refused (`a clause needs an operator, such as ^1.2 or >=1.2`). Whitespace is allowed around `,` and between an operator and its version, and nowhere else. A missing component is 0. | clause | means | |---|---| | `^1.2.3` | `>=1.2.3, <2.0.0` | | `^1.2` | `>=1.2.0, <2.0.0` | | `^1` | `>=1.0.0, <2.0.0` | | `~1.2.3` | `>=1.2.3, <1.3.0` | | `~1.2` | `>=1.2.0, <1.3.0` | | `~1` | `>=1.0.0, <2.0.0` | | `>=1.2`, `>1.2`, `<=1.2`, `<2` | the bound with missing components as 0: `>1.2` is `>1.2.0` | | `=1.2.3` | exactly `1.2.3`; `=` needs all three components | Below 1.0 a minor release is breaking, so caret fixes everything up to the first nonzero component: | clause | means | |---|---| | `^0.4.2` | `>=0.4.2, <0.5.0` | | `^0.4` | `>=0.4.0, <0.5.0` | | `^0.0.3` | `=0.0.3` | | `^0.0` | `>=0.0.0, <0.1.0` | | `^0` | `>=0.0.0, <1.0.0` | A pre-release version (`1.3.0-rc.1`) satisfies a range only when one of its clauses names a pre-release of that same release. So `^1.2` never selects `1.3.0-rc.1`, and `>=1.3.0-rc.1, <2` does. Build metadata (`+...`) is refused in a range and ignored in a release's version. There is no `*` and no `||`. Either can be added later without changing what an existing range means. #### `[target.]` Each `` is a selector you pass to `-t `. A target is a fully-spelled platform tuple; nothing is inferred from another key. `native` is a reserved name — declaring `[target.native]` is an error, because `native` resolves to whichever *declared* target matches the host. | Key | Required | Meaning | |-------|----------|---------| | `isa` | yes | Instruction-set architecture. Read by `$project.target.arch`. | | `os` | yes | Operating system. Read by `$project.target.os`. | | `abi` | yes | Application binary interface. Read by `$project.target.abi`. | | `of` | no | Object-format override; defers to the os's format when omitted. See [Object-format override](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#object-format-override). | | `base` | no | Load-address override (integer). Overrides the os's default base virtual address; defers to it (`0` for `freestanding`) when omitted. | | `platform` | no | Open platform tag (string), surfaced to comptime as `$mach.build.platform` (empty when unset). A support library keys its backend on it; the compiler treats it as opaque. See [Platform targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#platform-targets-bare-metal). | | `stack_reserve` | no | Thread stack reserve in bytes. See [Image stack size](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#image-stack-size). | | `stack_commit` | no | Thread stack commit in bytes. See [Image stack size](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#image-stack-size). | | `default` | no | Deprecated and ignored. It once marked the target `native` fell back to when none matched the host. The key is still accepted, so a published dependency keeps building, and warns as `target.default_deprecated`; it will be removed in a later major release. See [`native` target resolution](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#native-target-resolution). | | `extensions` | no | Array of instruction-set extension names the target may assume, such as `["sha", "ssse3"]`. Each name must be in the isa's vocabulary. See [Instruction-set extensions](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#instruction-set-extensions). | | `env` | no | Consumer environment (string). The values are owned by the target's isa: an `env` the isa does not define is a manifest error naming the target and the known values, and an isa that defines none refuses the key outright. Today only `spirv` defines any; see [Finished-module targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets). | ##### Instruction-set extensions `extensions` lists the extensions a target may assume beyond its isa's baseline: ```toml [target.linux-x86_64-sha] isa = "x86_64" extensions = ["sha", "ssse3", "sse41"] os = "linux" abi = "sysv64" ``` Each isa owns its vocabulary. The names are identifiers, so each one is also a comptime member, `$mach.build.extensions.` (see [`$mach`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-mach.md)): | `isa` | Baseline | Extensions | |-------|----------|------------| | `x86_64` | SSE2 | `ssse3`, `sse41`, `sse42`, `sha`, `fsgsbase`, `popcnt`, `lzcnt`, `bmi1`, `bmi2`, `cx16`, `avx`, `avx2`, `fma`, `movbe`, `f16c`, `avx512f`, `avx512bw`, `avx512cd`, `avx512dq`, `avx512vl`, `aes`, `pclmul` | | `aarch64` | AdvSIMD | `sha2`, `sb`, `aes`, `pmull`, `fp16` | | `riscv64`, `riscv32` | the isa string's selection | `i`, `m`, `a`, `f`, `d`, `c`, `zicond`, `zicsr`, `zifencei`, `zfhmin`, `zfh`, `zkt` | | `spirv` | | the type device features `float16`, `int8`, `int16`, `int64` and `float64`; `zero_init_workgroup`, `storage_read_without_format`, `storage_write_without_format`, `vulkan_memory_model`, `vulkan_memory_model_device_scope` and `buffer_device_address`, which `vulkan1.3` selects; the subgroup device features `subgroup_arithmetic`, `subgroup_clustered`, `subgroup_vote`, `subgroup_ballot`, `subgroup_shuffle`, `subgroup_shuffle_relative`, `subgroup_quad` and `subgroup_graphics_stages`; the atomic device features `buffer_int64_atomics`, `shared_int64_atomics`, `buffer_float32_atomics`, `buffer_float32_atomic_add`, `buffer_float32_atomic_min_max`, `buffer_float64_atomics`, `buffer_float64_atomic_add`, `buffer_float64_atomic_min_max`, `shared_float32_atomics`, `shared_float32_atomic_add`, `shared_float32_atomic_min_max`, `shared_float64_atomics`, `shared_float64_atomic_add` `shared_float64_atomic_min_max`, `buffer_float16_atomics`, `buffer_float16_atomic_add`, `buffer_float16_atomic_min_max`, `shared_float16_atomics`, `shared_float16_atomic_add`, `shared_float16_atomic_min_max`, `image_int64_atomics`, `image_float32_atomics`, `image_float32_atomic_add` and `image_float32_atomic_min_max` ([decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#optarget-set-name--a-function-that-is-a-target-instruction)); the image device features `storage_image_multisample` ([types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#handles)) and the sampling device features `resource_min_lod`, `image_gather_extended` and `maintenance8` ([decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#optarget-set-name--a-function-that-is-a-target-instruction)); the storage device features `storage_buffer_16bit_access`, `uniform_and_storage_buffer_16bit_access`, `storage_push_constant16`, `storage_input_output16`, `storage_buffer_8bit_access`, `uniform_and_storage_buffer_8bit_access` and `storage_push_constant8` | A name the selected isa does not hold is refused when the target resolves, with the names it does hold: ``` error[target.invalid]: target: `sha2` is not an extension or level of isa 'x86_64'; its extensions are: ssse3, sse41, sha, fsgsbase, popcnt, lzcnt, bmi1, sse42, cx16, avx, avx2, bmi2, fma, movbe, f16c, avx512f, avx512bw, avx512cd, avx512dq, avx512vl, aes, pclmul; its levels are: x86-64-v2, x86-64-v3, x86-64-v4 ``` The array must hold strings, each an identifier (`sse41`, not `sse4.1`) or a level spelling (`x86-64-v2`), listed once. A level is a bundle, never an axis of its own: each name may imply others, and the selection is closed over that once, when the target resolves. `sse41` brings `ssse3` (the chain stops there; SSE3 is not modelled), `sse42` brings `sse41`, `avx` brings `sse42`, `avx2`, `fma` and `f16c` bring `avx`, and every `avx512*` set brings `avx512f`, which brings `avx2`. On aarch64 `pmull` brings `aes`. On riscv `d` brings `f`, `zfh` brings `zfhmin`, `zfhmin` brings `f`, and `f` brings `zicsr`, as the isa string's own grammar has it, so `extensions = ["d"]` on `rv64i` selects `rv64ifd` with Zicsr. The isa string and the list feed one set: `isa = "rv64i"` with `extensions = ["m"]` selects the same machine as `isa = "rv64im"`. A name nothing in the compiler encodes against yet (`avx`, `avx512f`) is still a declared requirement: the inline assembler has no rows to admit under it, so today it records only the promise the binary makes about its hosts, and the promise is the program's to check. ###### Levels x86-64 also spells the published microarchitecture levels. A level is a manifest spelling that expands to its member names, so it may stand alone or beside names (`["x86-64-v2", "sha"]`), and it includes every lower level. It has no bit of its own: `$mach.build.extensions.x86-64-v2` and `#[extensions("x86-64-v2")]` do not exist, programs ask about the members (`.avx2`). `sse2`, `cmpxchg8b` and `lahf-sahf` are the `x86_64` baseline and are not names. The table here is the compiler's, held together by a test: | Level | Members beyond the level before | |-------|----------------------------------| | `x86-64-v2` | `ssse3`, `sse41`, `sse42`, `popcnt`, `cx16` | | `x86-64-v3` | `avx`, `avx2`, `bmi1`, `bmi2`, `fma`, `lzcnt`, `movbe`, `f16c` | | `x86-64-v4` | `avx512f`, `avx512bw`, `avx512cd`, `avx512dq`, `avx512vl` | `native` is not a spelling and never a default: the hardware requirement is visible in the manifest, never inferred from the build machine. The list is never part of `{target.isa}`. That placeholder is the `isa` value as written (`rv64i`, `x86_64`), on every isa; the list belongs to the target's identity and to `{target.name}`. "Selects" means the extension is assumed of every machine the binary runs on: the inline assembler admits its rows, `$mach.build.extensions.` answers 1, and a property the extension declares (Zkt's data-independent timing, which the constant-time multiply rows read) is taken as given. It never means a mode is on. A row such as a `dit` would admit `msr dit`, not set it. The `extensions` list is the manifest's only lever over the constant-time multiply, and only on riscv64, where `zkt` is what admits a secret `*`. There is no key that declares or overrides a timing mode. On aarch64 the condition is PSTATE.DIT, which the operating system declares it guarantees (linux and darwin) and the linked program's start code turns on for a binary that needs it; a manifest cannot assert it for an OS that declares nothing. On x86-64 the multiply rows hold unconditionally on Intel and AMD, so nothing is there to declare, and Intel's DOITM is a kernel-owned model-specific register that hardens memory-side predictors rather than the multiplier, so mach offers no key for it either. The rows, their conditions and the DOITM note are in [secrecy.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/secrecy.md#constant-time-multiply-by-instruction-set). Some rows are the target's alone. On riscv `i` is the baseline, `c` is a code-size selection mach never emits, and `f` and `d` select the float register file and the calling convention's float registers, and `zkt` is a promise about the machine's execution timing that the constant-time rows read, so none of them may be named in [`#[extensions(...)]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#extensionsnames--an-outlier-function); the refusal says why. So are spirv `float16`, `int8`, `int16`, `int64` and `float64`, the device features that let the whole module use a type of that width, and spirv `zero_init_workgroup`, the device feature that zero-initializes the workgroup memory of every [`#[shared]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#inputn--outputn--builtinstr--uniformset-binding--storageset-binding--samplerset-binding--push--specid--shared--shader-interface) variable in the module, and spirv `storage_read_without_format` and `storage_write_without_format`, the device features that let a storage image of `Unknown` format be read and written, spirv `storage_image_multisample`, the device feature that lets a storage image be multisampled, spirv `resource_min_lod`, `image_gather_extended` and `maintenance8`, the device features that let a sample clamp its level of detail, a gather take a run-time offset, and a fetch or a sample take one too, and spirv `vulkan_memory_model` and `vulkan_memory_model_device_scope`, the device features that select the Vulkan memory model for the whole module and let it use the `Device` scope, and spirv `buffer_device_address`, the device feature that lets a module hold physical pointers. Every x86_64 and aarch64 row, and riscv `m`, `a`, `zicond`, `zicsr`, `zifencei`, `zfhmin` and `zfh`, may be. Selecting an extension is a promise about **every** machine the binary runs on. The inline assembler admits the extension's mnemonics anywhere in the build, and a host without the extension faults on the first one it executes. A portable binary keeps the target at its baseline instead. It confines the extension instructions to [`#[extensions(...)]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#extensionsnames--an-outlier-function) functions and picks one of those at run time, after detecting the host's features. A mnemonic that needs an extension the target does not select, outside such a function, is refused. The refusal names the line to add: ``` error[asm.extension]: inline-asm instruction 'sha256rnds2' needs the `sha` extension, which this target does not select; add `extensions = ["sha"]` to the target, or mark the function `#[extensions(sha)]` and call it only after detecting the extension at run time ``` The selected set is part of the target's identity: two targets that differ only in `extensions` never share cached products. The selection also reaches code generation. Every vector operation is legal on every target and its shape never depends on the extension list; what moves is the lowering. Each packed row of a target's catalog declares the extension its instruction needs, and the lowering reads those rows: a cell with a baseline row and a row gated on an extension lowers to the gated instruction when the selected set holds it (`i32x4 * i32x4` is `pmulld` under `sse41` on x86_64 and the `pmuludq` pair without). A cell whose only packed row is gated scalarizes without the extension, and the warning at each such site names it (see `simd` below), and [`simd = "require"`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#profilename) judges against the selected set, so a kernel refused on the baseline is accepted once the target declares the extension it needs. SSE2 is the x86_64 baseline. A build never infers the build machine's features: what the binary assumes is what the manifest declares. ##### Image stack size `stack_reserve` and `stack_commit` set the thread stack an image asks its loader for, as byte counts: ```toml [target.windows] isa = "x86_64" os = "windows" abi = "win64" stack_reserve = 0x800000 # 8 MiB stack_commit = 0x1000 # 4 KiB ``` Both are optional. Omitting them keeps the format's conventional default, so an image built without them is byte-identical to one built before the keys existed. On PE that default is a 1 MiB reserve and a 4 KiB commit. A **reserve** is address space, not committed memory, so raising it costs nothing until the stack is actually used. That is also why these live on the target rather than the artifact: two artifacts built for one target share the value, and a small tool inheriting a large program's reserve pays nothing for it. **Only some object formats carry a stack size.** PE keeps both in its optional header and Mach-O keeps a reserve in `LC_MAIN.stacksize`. ELF has nowhere to put one — a linux main thread's stack is the kernel's and `ulimit`'s business — and a raw flat image has no header at all. Either key on a target whose resolved object format carries no stack size is refused when the manifest is read, naming the key, the target and the format, before anything builds: ``` error[target.stack_size]: mach.toml: [target.lin].stack_reserve is not expressible on the `elf` object format, which carries no stack size in its image headers ``` Every declared target is checked, not only the one being built, so the mistake is found on the first build rather than whenever someone happens to build that cell. A function whose own stack frame exceeds the reserve is refused at build time, naming the function, its frame size and the reserve: ``` error[stack.reserve_exceeded]: `main` needs a 1107824-byte stack frame, which its target's 1048576-byte stack reserve cannot hold; raise `stack_reserve` on the target, or move the large locals off the stack ``` That is a proof rather than an estimate - one frame against one reserve, with no call graph and no input dependence - so it is an error and has no false positives. Targets whose stack is not bounded by the image, such as every ELF target, are not checked. On Mach-O the value reaches only a **position-independent** image. A non-PIE one enters through `LC_UNIXTHREAD`, which has no stacksize member, and a stack size requested for such an image is refused at link rather than silently dropped. ##### Accepted tuple values | Axis | Values | |-------|--------| | `isa` | `x86_64`, `aarch64`, `riscv64`, `riscv32`, a canonical RISC-V extension string such as `rv32imc`, `spirv` | | `os` | `linux`, `windows`, `darwin`, `freestanding` | | `abi` | `sysv64`, `win64`, `aapcs64`, `lp64`, `lp64f`, `lp64d`, `ilp32`, `ilp32f`, `ilp32d`, `spirv` | A `[target.*]` naming an `isa`, `os` or `abi` outside these lists is refused through the registry's own lookup (`no isa implementation registered for 'mos6502' (registered: ...)`). `x86_64`/`linux`/`sysv64` is the primary host and target. `aarch64`-linux builds and runs natively in CI on every PR; `riscv64`-linux runs under qemu and self-hosts (#1852). `windows` is a supported cross-compilation target (PE/COFF, Win64 ABI). `darwin` is validated end-to-end on both architectures: each self-hosts to a three-generation fixpoint on a native macOS runner and ships a release archive. `freestanding` targets a raw flat image with no OS runtime; a bare-metal platform such as BareMetal (`bmos`) is a `freestanding` target plus a `platform` tag and `base` override (see [Platform targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#platform-targets-bare-metal)), not an os of its own. `spirv` is not a machine at all — it emits a finished GPU module rather than machine code (see [Finished-module targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). `riscv64` and `riscv32` are width-only spellings, and each names a **default profile**: `riscv64` is `rv64gc` and `riscv32` is `rv32imac`. A canonical extension string such as `rv32imc` or `rv64imafd` selects a smaller machine. The retained vocabulary is I, M, A, F, D, C, Zicond, Zicsr, Zifencei and Zkt, written in lowercase canonical order with multi-letter names after an underscore, so `rv64gc_zicond`; `g` expands to IMAFD plus Zicsr and Zifencei. F carries its required Zicsr, and D requires F. Zicond adds `czero.eqz` and `czero.nez`, which a branch-free select compiles to where a selection holds it and the xor-and-mask sequence it replaces does otherwise. Zkt changes no instruction. It states that the listed operations run in data-independent time, which is what lets a secret multiply compile (see `secrecy.md`). An optional version must be the one mach models: I 2.1, M 2.0, A 2.1, F and D 2.2, C 2.0, Zicond 1.0, Zicsr and Zifencei 2.0, Zkt 1.0. Unknown extensions, other versions, duplicates, noncanonical order and the E base are refused rather than rounded up to the default machine. The selected ISA bounds what the compiler generates and what named inline assembly may use: an instruction needing an extension the selection lacks is refused with a diagnostic naming that extension. A foreign object's `Tag_RISCV_arch` must declare only selected extensions at the modeled versions and the same register width, and its header flags may not claim compressed code without C; linking never widens the selection. A raw `.word` directive is the documented unchecked encoding boundary. C is accepted as a capability of the selected machine, but the emitter writes full-width instructions only. The ABI still selects the calling convention on its own, and it must fit the selected machine: `rv32imc` has no floating-point registers, so it takes `ilp32`, and `riscv32` (rv32imac) is refused with `ilp32f` or `ilp32d`; spell `rv32imafdc` when RV32 hardware floating point is wanted. `mach init` scaffolds `riscv32`/`freestanding` with `ilp32` for that reason. `lp64`, `lp64f`, `lp64d`, `ilp32`, `ilp32f`, and `ilp32d` are the RISC-V psABI calling-convention family, one `abi` per member. The lp64 three target `riscv64` and the ilp32 three target `riscv32`; an ilp32 member on a `riscv64` target or an lp64 member on `riscv32` is refused when that target is selected (`calling convention 'ilp32' does not target instruction set 'riscv64'`), since XLEN is part of what the id means. Within each width, the members differ only in how floating-point arguments travel: `lp64` and `ilp32` are **soft float** — every float argument rides an integer register — while the `f` and `d` suffixes are **hard float**, passing `f32` (`f`) or both `f32` and `f64` (`d`) in the `fa0`-`fa7` register bank. `lp64d` is what a `riscv64-linux-gnu` toolchain means by "riscv64", and is the convention `mach init` scaffolds for a `riscv64`/`linux` target; a manifest that wants soft float, or the `f`-only convention, must name it explicitly, since `abi` has no default of its own (see [Totality](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#totality)). `linux` accepts all three on `riscv64` because its kernel ABI is integer-only, whereas every operating system accepts only the conventions it declares per instruction set (`linux` and `darwin` refuse `win64` on `x86_64`, `windows` refuses `sysv64`) and `freestanding` accepts every convention the instruction set covers. `riscv32` currently only reaches a `freestanding` target — `mach info targets` lists `freestanding-riscv32` rows and no `linux`/`darwin` riscv32 row. `mach info targets` prints every tuple this binary can actually build, one `(os, isa, abi, object)` cell per line with every dimension spelled; it is derived from the same declarations composition reads, so it never advertises a tuple that would fail to resolve. `mach info` alone prints the tuple the host resolves to. A value outside its axis's set is a strict-parse error, so a typo is caught rather than silently never matching. ##### Object-format override `of` overrides the object format a target implies. Each os has a default format — `linux` → `elf`, `windows` → `coff`, `darwin` → `macho`, `freestanding` → `raw` — and `of` names a different one from the same closed set: `elf`, `coff`, `macho`, `raw`, `spv`. `of` is optional; omit it to take the default. An os accepts only the formats it can load, so an override the os cannot enter is refused. ```toml [target.metal] isa = "x86_64" os = "freestanding" # os default object format is "raw" abi = "sysv64" of = "elf" # override: emit an ELF object instead ``` The default is a function of the whole tuple, not the os alone. An os default carries relocatable machine text, which a whole-module emitter does not produce, so a `spirv` target resolves to `spv` — the format that carries a finished module — regardless of the os it names. An `of` naming a format whose emission shape does not match the instruction set's is refused at composition (`instruction set 'spirv' emits finished modules, but object format 'raw' carries linkable objects`), so the override cannot compose a tuple that would emit nothing. The page an image is laid out at is a function of the same tuple. A format a loader maps by page (`elf`, `coff`, `macho`) places every load segment on a page of its own, so no two segments with different permissions share one: the os's page where the os declares one (`linux` on `aarch64` lays out at 64 KiB, the largest page a kernel may use, `darwin` on `aarch64` at 16 KiB, 4 KiB elsewhere), and the instruction set's hardware page (4 KiB on `x86_64`, `aarch64`, `riscv64` and `riscv32`) where the os has no loader of its own, which is what `freestanding` with `of = "elf"` gives a bootloader such as Limine or GRUB. A flat image (`raw`) and a finished module (`spv`) have no page and are laid out byte-tight. ##### Finished-module targets A `spirv` target's object output is a complete, self-contained module rather than a link input. Each module is written to `/obj/.spv` like any other target's objects, and the entry module already carries every function its stages reach, so it is the whole deliverable. A `bin` artifact therefore needs no linker: the build publishes the entry module at the artifact's resolved `out` (or `-o`), and `--emit obj` stops at the module tree: ```toml [target.gpu] isa = "spirv" os = "freestanding" abi = "spirv" # no `of`: the finished-module format resolves on its own ``` ``` mach build . --target gpu # writes the entry module at the artifact's out, and out/gpu//obj/.spv ``` With `debug` on, each module carries its debug information inside it, written as core instructions that need no capability or extension and so fit every `env`: `OpString` and `OpSource` name the source files, `OpName` names functions, interface variables and locals, `OpName` and `OpMemberName` name a uniform or storage block's record and its fields, and `OpLine` attributes each instruction to its source line and column. A required shader artifact built for a consumer's debug profile therefore builds, and validation layers and capture tools report names and source lines. `env` is a general target key whose values are owned by the target's isa; a `spirv` target uses it to declare the environment its modules are consumed in. The environment fixes the SPIR-V version word and the capability ceiling: the compiler derives the minimal capability set a module needs and refuses a module that needs more than the ceiling, naming the capability and the environment. The ceiling only admits a capability. One that Vulkan leaves to an optional device feature also needs the target to select that feature, as below. Without `env` a module is written as SPIR-V 1.6 with no ceiling, and holds every feature. The environment also selects the `zero_init_workgroup` extension from `vulkan1.3`, where `shaderZeroInitializeWorkgroupMemory` is core, and a module without `env` has it too. A target for an earlier version selects it with `extensions` when its consumer enables `VK_KHR_zero_initialize_workgroup_memory`. The environment also selects `storage_read_without_format` and `storage_write_without_format` from `vulkan1.3`, which accepts the `StorageImageReadWithoutFormat` and `StorageImageWriteWithoutFormat` capabilities with no feature enabled, and a module without `env` has them too. A module reading or writing a storage image of `Unknown` format for an earlier version is refused unless the target selects the matching extension, which it does when its consumer enables `shaderStorageImageReadWithoutFormat` or `shaderStorageImageWriteWithoutFormat`. The memory model follows the same selection. A module is written under the GLSL450 memory model unless its target selects `vulkan_memory_model`, Vulkan's `vulkanMemoryModel` feature, and then under the Vulkan memory model, declaring the `VulkanMemoryModel` capability and, below SPIR-V 1.5, the `SPV_KHR_vulkan_memory_model` extension. `vulkan1.3` requires the feature of every device, so it selects the extension, and a module without `env` has it too. A target for `vulkan1.1` or `vulkan1.2` selects it with `extensions` when its consumer enables the feature, and one for `vulkan1.0` is refused, since the model needs SPIR-V 1.3. Under the Vulkan model a memory scope of `Device` needs `vulkanMemoryModelDeviceScope` as well, which `vulkan1.3` also requires: an instruction taking that scope is refused unless the target selects `vulkan_memory_model_device_scope`, which brings `vulkan_memory_model` with it. Under GLSL450 the `QueueFamily` scope and the `MakeAvailable`, `MakeVisible` and `Volatile` memory semantics, which only the Vulkan model defines, are refused. What `"coherent"` and `#[shared]` mean under each model is in [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#inputn--outputn--builtinstr--uniformset-binding--storageset-binding--samplerset-binding--push--specid--shared--shader-interface). Physical pointers follow it too. A module holding one, a pointer stored in memory or made from an address ([types.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/types.md#pointers-on-spir-v)), is refused unless the target selects `buffer_device_address`, Vulkan's `bufferDeviceAddress` feature, and then declares the `PhysicalStorageBufferAddresses` capability, the `PhysicalStorageBuffer64` addressing model and, below SPIR-V 1.5, the `SPV_KHR_physical_storage_buffer` extension. `vulkan1.3` requires the feature of every device, so it selects the extension, and a module without `env` has it too. A target for an earlier version selects it with `extensions` when its consumer enables the feature. No environment selects `storage_image_multisample`, Vulkan's `shaderStorageImageMultisample`, which every version leaves optional. A module declaring a multisampled storage image is refused unless the target selects it, and with it the module declares `StorageImageMultisample`, and `ImageMSArray` as well for an arrayed one. A multisampled sampled image needs no feature. No environment selects `resource_min_lod` either, Vulkan's `shaderResourceMinLod`. A sample passing the `MinLod` image operand is refused unless the target selects it, and with it the module declares `MinLod`. Nor does any select `image_gather_extended`, `shaderImageGatherExtended`, which a run-time `Offset` image operand needs and with which the module declares `ImageGatherExtended`, or `maintenance8`, under which Vulkan admits that operand on a fetch or a sample as well as a gather. No environment selects a type's feature, since every Vulkan version leaves them optional: `int8`, `int16`, `int64`, `float16` and `float64` are Vulkan's `shaderInt8`, `shaderInt16`, `shaderInt64`, `shaderFloat16` and `shaderFloat64`, and the target selects one with `extensions` when its consumer enables it. A module without `env` has all five. The ceiling below still bounds them, so `int8` or `float16` under `vulkan1.0` is refused for the environment, not the feature. Under `float16` an `f16` is the native `OpTypeFloat 16` wherever it lives, computed, negated and converted by the core float instructions, so an `f16` shader needs `float16` alone. A local, a parameter or a result is that type, and so is an `f16` in memory the host or the workgroup shares, a stage input or output, a storage buffer, a uniform or push block, a record a physical pointer reaches or a `#[shared]` variable, so an atomic can operate on it and a whole record moves between that memory and a local as it is. Reading an `f16`'s bits with `:~` into a `u16` or `i16` local, or back, needs no `int16` either. Without it an `f16` is the software expansion on its 16 bits, which computes in binary32 on 32-bit integers, so it needs neither `int64` nor `float64` of its own. An `f64` it converts to or from needs `float64` as any `f64` does. `%` needs no `int64` at any float width. A stage input or output is `OpTypeFloat 16` with or without `float16`, so a pipeline interpolates an `f16` varying as it does an `f32` one, and only an integer or 64-bit fragment input is `Flat`. Under `int8` and `int16` an integer of that width computes at its own width. Without the feature, an integer of that width is carried, wherever it lives in a function, in a 32-bit integer, which needs no capability: a `u8`, `i8`, `u16` or `i16` local, and a `bool`, is wrapped and extended at its own width where the program can tell. A vector is carried lane by lane the same way, so a `u8x4`, `i16x4` or `u16x8` local, and without `float16` an `f16x4`, needs no feature either, and its lanes are wrapped and extended at their own width where the program can tell, a reinterpret with `:~` included. A member of an aggregate keeps its declared width, so an 8-bit or 16-bit one, or a vector of them, in a local or in `#[shared]` memory needs its feature, and the module is refused, naming it, without. Memory the host shares keeps its width too, under the storage feature below rather than `int8` or `int16`, and a load from it or a store to it converts to and from the wider integer the function computes in. Nothing carries a 64-bit type, so a module holding a `u64`, `i64` or `f64` anywhere needs `int64` or `float64`. An 8- or 16-bit scalar, an `f16` included, in memory the host shares needs Vulkan's storage feature for its width and memory, which no environment selects: a target names each one its consumer enables, and a target naming no `env` holds them all. Each enables the capability of its name, declared with its SPIR-V extension below the version that took it into the core. | Memory | 16-bit | 8-bit | |---|---|---| | `#[storage(...)]`, or a record a physical pointer reaches | `storage_buffer_16bit_access` (`StorageBuffer16BitAccess`) | `storage_buffer_8bit_access` (`StorageBuffer8BitAccess`) | | `#[uniform(...)]` | `uniform_and_storage_buffer_16bit_access` (`UniformAndStorageBuffer16BitAccess`) | `uniform_and_storage_buffer_8bit_access` (`UniformAndStorageBuffer8BitAccess`) | | `#[push]` | `storage_push_constant16` (`StoragePushConstant16`) | `storage_push_constant8` (`StoragePushConstant8`) | | `#[input(n)]`, `#[output(n)]` | `storage_input_output16` (`StorageInputOutput16`) | refused | A stage output starts at its initializer, and at zero without one. A constant of a 16-bit type needs the type's own feature, `int16` or `float16`, so an output holding a 16-bit scalar without it carries no zero constant: each stage that reaches it stores the zero first, converted from a 32-bit one, and the output needs `storage_input_output16` alone. One initialized to anything but zero starts at that constant, which needs the type's feature as well. The 16-bit capabilities need `SPV_KHR_16bit_storage` below SPIR-V 1.3 and the 8-bit ones `SPV_KHR_8bit_storage` below SPIR-V 1.5. Vulkan defines no 8-bit stage input or output, so one is refused, and an 8-bit member of a `#[storage(...)]` buffer is refused under `vulkan1.0`, whose storage buffer is a `BufferBlock` in the `Uniform` class, which the 8-bit feature does not reach. ```toml [target.gpu] isa = "spirv" os = "freestanding" abi = "spirv" env = "vulkan1.0" extensions = ["int16", "int64", "float64"] ``` | `env` | SPIR-V | capabilities the ceiling admits | |---|---|---| | `vulkan1.0` | 1.0 | Int16, Int64, Float64, Sampled1D, SampledCubeArray, Image1D, ImageCubeArray, SampledBuffer, ImageBuffer, StorageImageExtendedFormats | | `vulkan1.1` | 1.3 | same as `vulkan1.0` | | `vulkan1.2` | 1.5 | the above plus Int8, Float16 | | `vulkan1.3` | 1.6 | same as `vulkan1.2` | The ceiling is what a conforming implementation of that Vulkan version can enable through core device features alone, with no device extension: from the Vulkan specification's "Vulkan Environment for SPIR-V" appendix, the capabilities table maps `Int64`, `Int16`, `Float64` and `SampledCubeArray` to the `shaderInt64`, `shaderInt16`, `shaderFloat64` and `imageCubeArray` features of Vulkan 1.0, `ImageCubeArray` to `imageCubeArray` as well, `Sampled1D`, `Image1D`, `SampledBuffer`, `ImageBuffer` and `StorageImageExtendedFormats` to core, and `Int8` and `Float16` to `shaderInt8` and `shaderFloat16`, which became core features in Vulkan 1.2 (promoted from `VK_KHR_shader_float16_int8`). The SPIR-V version per Vulkan version is the appendix's required version: 1.0, 1.3, 1.5 and 1.6. ```toml [target.gpu] isa = "spirv" os = "freestanding" abi = "spirv" env = "vulkan1.2" ``` The artifact's `out` template and `-o` name that delivered module, so `{artifact.suffix}` gives it `.spv`, and `{artifact..out}` names it for a consumer that embeds it (see [Artifact requirements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-requirements)). A `static` or `shared` artifact kind, and `mach test`, are refused by name, since there is no archive, shared object, or executable form for a module. ##### Platform targets (bare metal) A bare-metal platform — such as [BareMetal](https://github.com/ReturnInfinity/BareMetal) (`bmos`), Return Infinity's x86-64 exokernel — is not its own `os`. It is `os = "freestanding"` plus two optional keys: a `base` load-address override and an open `platform` tag a support library keys its backend on (surfaced to comptime as `$mach.build.platform`). A BareMetal target: ```toml [target.bmos] isa = "x86_64" os = "freestanding" abi = "sysv64" base = 0xFFFF800000000000 # BareMetal copies the flat image here and calls it platform = "bmos" # selects the mach-bmos backend # no `of`: freestanding's default object format is "raw" ``` - The artifact is a **flat binary** — no header, no sections, no entry record. The loader copies the file's bytes verbatim to `base`. - The load address is set by **`base`** and the loader relocates nothing, so the image is never position-independent. - Execution begins at the **first byte of the image**, which the loader reaches with a `call`. The entry function is marked `#[symbol("_start")]`, and it must be the only function or the first one emitted, since a flat image cannot say where else to enter. An entry anywhere but the base is refused at link. - A program **exits by returning**: the entry function's `ret` goes back to its caller. There is no exit syscall. The compiler encodes no BareMetal knowledge — `base` places the image and `platform` is an opaque string. The kernel-call machinery lives in the `mach-bmos` platform package, which gates on `$mach.build.platform == "bmos"`; because that is a library, a non-x86-64 bmos build fails at the package's own `$mach.build.arch` gate rather than in the compiler. One fact the kernel leaves to the program: **the stack is not guaranteed 16-byte aligned at entry**, so a startup shim must align it before calling anything that may use SSE. BareMetal's own `crt0.c` does exactly this. Zero-initialized data needs no such step. A flat image spans its whole memory extent, so `.bss` is stored as the zero bytes it is and arrives zeroed with the rest of the image (#2402) — an image costs its bss size in file bytes, and nothing has to zero anything at startup. #### `[profile.]` A profile is one explicit compilation policy: a build variant. The optimization level, the debug-emission toggle and the three SIMD levers live here because they are variant concerns, and every one of them is stated. A root manifest declares at least one profile; nothing is synthesized. Values that are *derived* rather than declared live elsewhere: a target's object format and naming come from `[target.*]` facts, and an absent optional feature such as a `[link.*]` filter axis is spelled `"*"` where it applies, not defaulted here. | Key | Type | Meaning | |---------|---------|---------| | `opt` | integer | Optimization level: `0` selects the debug pipeline (the always-on passes only), `1` and `2` select the release pipeline. `1` and `2` currently share a pass set, which includes loop auto-vectorization (see `vectorize` below). Any other integer — or a non-integer — is a manifest error. | | `debug` | bool | Emit debug info for this profile: DWARF in ELF, Mach-O and COFF objects alike, and the core SPIR-V debug instructions on a `spirv` target (see [Finished-module targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). A PE image carries its DWARF in `.debug_*` sections, which gdb, lldb and the LLVM tools read and Visual Studio and WinDbg do not. Gates emission only, never the optimizer, so a `release` profile can keep symbols with `debug = true`. A non-boolean is a manifest error. | | `simd` | string | SIMD scalarization lever. `"scalarize"` emits a defined unrolled scalar expansion wherever the target has no packed instruction for a vector operator, with one `vector.scalarize` warning at each such operation naming the operation, its lanes, the function, the target and the extension that would pack it (or that none would). `"require"` makes each of those sites a hard error with the same text. It applies **per operation on every target**, not only to targets with no vector unit: x86-64's SSE2 baseline has no 32-bit lane integer multiply and NEON has no 64-bit one, so a capable target scalarizes too. Any other string is a manifest error. | | `vectorize` | bool | Auto-vectorization lever. When `true`, the release pipeline rewrites provably-safe counted loops to SIMD at the target's vector width (128 bits, and 256 on x86-64 under `avx2`, or `avx` for float lanes) on a target with hardware vectors; `false` skips the pass, so release output stays scalar. A non-boolean is a manifest error. | | `float_reassoc` | bool | Permission to treat floating-point addition and multiplication as **associative**. It lets the vectorizer reduce an `f32`/`f64` accumulator through lane-count partial sums, which changes the result — see [Float reassociation](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#float-reassociation) for what that costs and what it buys. A non-boolean is a manifest error. | | `default` | bool | **Optional.** `true` marks the profile a build uses when several are declared and `--profile` is absent. Exactly one profile may carry it. See [Profile requirement and selection](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#profile-requirement-and-selection). | | `allow` | array of strings | **Optional.** The warnings this profile silences, each entry a key or a family of keys from the [warning key table](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#silencing-warnings), named once. An unknown key, an entry that covers only errors, a repeated entry or a non-string entry is a manifest error. Absent, nothing is silenced. | Five keys (`opt`, `debug`, `simd`, `vectorize`, `float_reassoc`) are required in a declared profile, in a root and in a dependency manifest alike; only `default` and `allow` are optional. A missing key is a manifest error naming the table and the key: ``` error[manifest.missing_required]: mach.toml: [profile.debug] is missing required key 'vectorize'; a profile declares opt, debug, simd, vectorize and float_reassoc ``` ##### Profile requirement and selection A root manifest declares at least one `[profile.*]` table. A root that declares none does not build: ``` error[manifest.missing_required]: mach.toml: no [profile.] table is declared; a build needs an explicit profile declaring opt, debug, simd, vectorize and float_reassoc ``` `mach init` writes `debug` (`opt = 0`, `debug = true`, `default = true`) and `release` (`opt = 2`, `vectorize = true`) in full, so a scaffold never starts from that error. A dependency manifest that declares no profile is still read (its profiles are never used to build the consumer, see [Root vs. dependency strictness](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#root-vs-dependency-strictness)); it gets the two built-in profiles `debug` and `release` for its own `{profile.name}` templates. That synthesis is a dependency-only convenience and applies to no root. Which profile a build uses follows one rule, the same one that selects a target and an artifact: 1. an explicit `-p ` wins; 2. otherwise a sole declared profile is chosen; 3. otherwise the one marked `default = true` is chosen. Table order carries no meaning. A manifest that declares several profiles and marks none is refused wherever a command must pick one: ``` error[selection.ambiguous]: mach.toml: several profiles are declared and none is marked `default = true`; no profile is selected by table order: mark exactly one [profile.] with `default = true` or select one with -p ``` Emission of the human-readable IR and assembly side-artifacts is **not** a profile concern — it is controlled only by the `--emit-ir` / `--emit-asm` flags of `mach build`. The `vectorize` lever only ever *subtracts*. The pass it gates runs in the release pipeline on targets that report 128-bit vector support (SSE2 on x86-64, NEON on aarch64, and `OpTypeVector` on spirv) and rewrites counted, unit-stride loops whose dependence analysis proves independence — element-wise maps behind a runtime alias guard, and associative-exact integer reductions. A loop it cannot prove safe stays scalar, and a target without hardware vectors (riscv64) never enters the pass, so `vectorize = false` changes performance and never semantics. For a single function, the `#[scalar]` decorator is the finer-grained opt-out (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md)). `float_reassoc` is the one lever here that *adds*, and the only profile key that can change a program's computed answer. It widens that same pass to float reductions and does nothing else; `vectorize = false` switches the pass off wholesale and so overrides it. The `simd`, `vectorize` and `float_reassoc` levers are always the **consumer's**. Consistent with the root-vs-dependency strictness above, a dependency's `[profile.*]` is parsed by the same schema and never read to build the consumer, so a library's values are inert — the effective levers come from the consumer's resolved profile. Libraries set nothing SIMD-specific and inherit the consumer's choice; there is no ecosystem fork and no dual API. ##### Silencing warnings Every diagnostic kind has a dotted key, named by the subject it concerns, and every diagnostic that has one prints it: ``` warning[vector.scalarize]: vector divide on 4 lanes of 32-bit integers in 'app.main.kernel' scalarizes on x86_64: no packed form for it at any extension ``` `allow` names warning keys the build does not report. A silenced warning is dropped before it is printed or counted, so the summary's warning count leaves it out as well. Nothing else changes: the same code is built, and an error is never silenced. ```toml [profile.release] opt = 2 debug = false simd = "scalarize" vectorize = true float_reassoc = false allow = ["import.unused", "target"] ``` A key's leading components name a family: `"target"` covers `target.skipped` and `target.default_deprecated`, and `"vector"` covers `vector.scalarize`. A family covers whole components only, so `"vec"` is not a key. To acknowledge one warning where it is raised instead of across the whole build, put [`#[expect]`](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#expectkey--acknowledge-a-warning) on the declaration that raises it. Every diagnostic kind has one row in one table in the compiler (`src/lang/diagnostic/kind.mach`). Each warning names its row where it is raised, and `allow`, `#[expect]` and the printed key read the same rows, so a key here is exactly the kind the warning carries. The list is closed: a key no row declares is refused, naming the warning keys there are. A key is never reused for a different kind: see [the registry](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/diagnostics.md#the-registry). | Key | Warns when | Decided by source | |---|---|---| | `import.unused` | a symbol import names something the module never uses | yes | | `decl.deprecated` | code outside a `#[deprecated]` declaration's module uses it | yes | | `doc.lint` | a doc comment's component list names no parameter, field, generic or `ret` of its declaration, leaves a component undescribed, or lists them out of declaration order | yes | | `float.inexact` | a float literal is not exact at its type and its digits are not the shortest spelling of the value stored | yes | | `fwd.instances` | a shared library `fwd`s a generic, comptime-parameter or pack declaration, which exports no symbol | no | | `debug.dropped` | the linker leaves out an object's debug info that it cannot merge | no | | `target.skipped` | multi-target analysis skips a declared target this build does not support | no | | `target.default_deprecated` | a `[target.*]` table carries the deprecated `default = true`, which selects nothing | no | | `vector.scalarize` | a vector operation falls back to scalar code on the target (see `simd`); portable code silences it | no | | `expect.unfulfilled` | an `#[expect]` names a key decided by source and no such warning is raised inside its declaration | no | A kind decided by source warns or not from the source text alone, whatever the target, goal or profile. Only such a key is reported when an `#[expect]` naming it goes unfulfilled; the others depend on what the build compiles and for which target, so an expectation of one can be quiet in a given build. Only warnings can be silenced. The table also keys error kinds, which print their key as well, and naming one in `allow` is refused rather than read as unknown: | Key | Error | |---|---| | `secret.not_oblivious` | a function performs a constant-time operation on a secret value without `#[oblivious]` | ``` error[allow.error_key]: mach.toml: [profile.release].allow entry "secret.not_oblivious" names an error; only a warning can be silenced ``` Like the SIMD levers, `allow` is the consumer's: a dependency's profiles are never read to build it, so a library cannot silence the consumer's warnings. ##### Float reassociation `float_reassoc = true` grants the optimizer exactly one liberty: it may treat floating-point `+` and `*` as **associative**, and regroup a reduction accordingly. Nothing else changes. It does not license reciprocal substitution for division, assumptions that operands are finite or non-NaN, contraction into a fused multiply-add, or flushing subnormals to zero. Each of those would be its own key, with its own argument. What it unlocks is the reduction vectorizer. `s = s + a[i]` is a serial dependence chain: every iteration must wait for the previous one's rounded result, so it cannot be done four lanes at a time without regrouping the additions. With the key set, the loop becomes lane-count independent partial accumulators plus a horizontal combine at the end — which computes a *differently grouped*, and therefore differently rounded, sum. Element-wise float loops (`a[i] = b[i] * c[i]`) never needed the key and are unaffected: each lane performs exactly the operation the scalar loop performed. The reductions it admits are sum (`s = s + x`), product (`s = s * x`), and the dot / matmul inner loop (`s = s + a[i]*b[i]`). Subtraction and division reductions are not associative in real arithmetic either, so they are refused with the key set exactly as without it. **The accuracy cost, measured.** Summing one million `f64` values of `0.1`, scored against a compensated (Kahan) reference on x86-64: | ordering | result | relative error | |---|---|---| | strict, sequential | `100000.00000133288` | 1.33e-11 | | reassociated, 2 lanes | `99999.9999991058` | 8.94e-12 | The two answers differ from **each other** by 2.2e-11 relative — about 153,000 `f64` ULPs at that magnitude. The same experiment in `f32` (4 lanes) differs by 1.2e-2 relative: strict gives `100958.34`, reassociated `99759.85`, against a true value of `100000.0`. Note the direction. In both cases the *reassociated* answer is the more accurate one, because splitting into per-lane partials keeps each running total smaller and so grinds off fewer low bits — the same reason pairwise summation beats sequential summation. But that is a property of this input, not a guarantee. The honest statement is that the result **changes**, by roughly the accumulated rounding error of the sum, in a direction that depends on the data. Code whose correctness depends on the exact bit pattern of a float reduction — a checksum, a reproducibility requirement, a comparison against a reference implementation — must leave the key off. Two exactness properties are preserved rather than traded away. The idle lanes are seeded with the op's **exact** IEEE identity — `-0.0` for addition (`x + (-0.0)` is `x` for every `x`, where `+0.0` would turn a negative-zero sum positive) and `1.0` for multiplication — so no signed-zero or NaN behaviour changes. And a trip count below the lane count never enters the vector loop at all, so short reductions are bit-identical regardless of the key. The key is profile-wide. For a single function, `#[scalar]` opts out of vectorization entirely and takes precedence over it, so a routine that must stay IEEE-strict inside an otherwise-reassociating build has a spelling. There is no per-function opt-*in*: whether a reduction may be reassociated is the caller's tolerance to decide, not the callee author's. Integer reductions are untouched by this key. They vectorize unconditionally and are bit-identical to the scalar reference, because integer add / xor / or / and reassociate exactly. The CLI selects and overrides at invocation time: `-p ` picks the profile; `-g` forces `debug` on for one build regardless of the profile's key (precedence `-g` > profile > off — there is no flag to force it off over a `debug = true` profile; edit the manifest or pick another profile). `-O0` and `-O2` override the profile's `opt` the same way; `-O1` is rejected (`-O1 was removed; use -O0 or -O2`). #### `[artifact.]` Every artifact is declared explicitly and named by its table key. `$bin.name` reads the selected artifact's name. | Key | Required | Meaning | |-----------|----------|---------| | `kind` | yes | `"bin"`, `"static"`, or `"shared"` (see below). | | `entry` | yes | Entry source, relative to the project `src` dir (e.g. `main.mach` for `src/main.mach`). The entry module's FQN is `.`, `/` turned into `.`. | | `out` | yes | This artifact's output path, **relative to the expanded project `out`** and rooted there automatically — write `bin/demo`, not `{project.out}/bin/demo`. Use `{artifact.suffix}` for the target extension, or write a literal filename. See [Artifact filenames and identity](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-filenames-and-identity). | | `targets` | yes | Array of declared target names this artifact builds for; `["*"]` means every declared target. | | `link` | yes | Array of `[link.X]` names this artifact links (see below). `[]` for none. A name with no table is a manifest error naming the artifact and the declared tables (`[artifact.p1].link names no [link.*] table: 'nosuch' (declared: [link.kernel32])`). | | `need` | yes | Array of category-qualified requirements such as `step.generate`, `artifact.support`, and `artifact.shader-*`. Each glob matches only its named category. `[]` for none. See [Artifact requirements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-requirements). | | `subsystem` | no | `"console"` (default) or `"gui"` — the environment a windows executable declares it runs under; refused on a target whose image format has no subsystem (see below). | | `icon` | no | Project-root-relative `.ico` path embedded in a Windows executable's PE resources. Non-empty path string; `bin` artifacts only. | | `manifest` | no | Project-root-relative application-manifest path embedded byte-for-byte in a Windows executable's PE resources. Non-empty path string; `bin` artifacts only. | | `default` | no | `true` puts the artifact in the [default selection](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#selection-and-the-build-matrix): with no `-a`, `mach build` and `mach check` take the marked artifacts among those supporting the selected target (every one of them when none is marked), and a command that needs one artifact (`mach test`, `mach run`, the editor's union build) takes the marked one. A command that needs one artifact refuses two marked candidates; an explicit `-a` always wins, and a sole candidate needs no marker. | `entry` is the build cell's source root. The build follows its active `use` and `fwd` edges transitively and compiles that reachable module set; another file under `src` is not part of the cell merely because it shares the project directory. This is what lets one project declare host and accelerator artifacts with disjoint target sets. `mach build`, `mach check` and `mach test` all operate on the selected artifact's closure: a test build compiles and tests exactly the modules the artifact under test reaches, so a module no selected artifact reaches is not loaded under any of them (see [test.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/test.md#which-tests-run)). - **`bin`** links an executable at the resolved `out` path. On a finished-module target such as `spirv` it is the entry module, written there unlinked. - **`static`** materialises a real `ar` archive at the resolved `out` path — the per-module objects with an archive symbol index, the deliverable a consumer links as a `.a` (#1997). - **`shared`** links a dynamic library at the resolved `out`. Only ELF targets write one today: `linux` on `x86_64`, `aarch64` and `riscv64` produce a `.so` whose `SONAME` is its file name. The Mach-O `.dylib` and PE `.dll` writers are not built yet, so a `darwin` or `windows` target refuses with `link: object format cannot write shared libraries` (#3588). A `freestanding` target never writes one: its default `raw` format refuses with `a flat-image object format produces only executables`, and setting `of = "elf"` moves the refusal to the link, `link: a shared library needs a loader to map it, and os = "freestanding" has none`, because a shared library only exists to be mapped by a loader the os provides. - **Exports.** The library exports the root project's `pub` functions and variables and every name its modules re-export with `fwd`, including a dependency's. A dependency's own `pub` surface is not exported unless it is re-exported. `#[symbol("name")]` sets the name an export carries and does not make anything visible: a `pub` function exports under its `#[symbol]` name, and a non-`pub` one stays hidden whatever its name. - **Internals.** Every other definition still links inside the library but is absent from `.dynsym`. In the `.so` it is a `LOCAL` symbol in `.symtab`, and in the per-module object it is a `GLOBAL` symbol whose visibility the format spells its own way: ELF `STV_HIDDEN`, Mach-O `N_PEXT`, and COFF, which has no visibility bit, a `.drectve` section listing every exported definition as ` /EXPORT:`, so a global definition the directives do not name is hidden. A COFF object with no `.drectve` says nothing about visibility and is read as it was. A `fwd` re-export of a dependency's symbol is a request the object carries separately: on COFF it is one more `/EXPORT:` token, and on ELF and Mach-O it rides in a non-loaded mach section (`.mach.exports`, or `__MACH,__mach_exports`) the parser consumes. A relocatable object emitted and parsed back therefore keeps the same visibility and the same requests. - **Refusals.** A shared artifact that exports nothing is refused: ``` link: shared library '' exports nothing: a shared library needs at least one `pub` declaration in the project, or a `fwd` re-export of one ``` A `freestanding` target is refused as well. With its default `raw` format the artifact fails naming (`artifact naming: this object format has no shared-library form`), and with `of = "elf"` the link refuses with `link: a shared library needs a loader to map it, and os = "freestanding" has none`. Per-target extension or per-target entry is not a per-cell exception table — it is a second artifact stanza, so the condition stays visible like everything else. ##### Artifact filenames and identity Use `{artifact.suffix}` in an artifact's `out` to select its target filename extension. Literal paths stay literal. No prefix is inserted, so a library may spell its desired `lib` prefix directly. ```toml [artifact.app] kind = "bin" entry = "main.mach" out = "bin/app{artifact.suffix}" targets = ["*"] link = [] need = [] ``` This produces `app.exe` on Windows and `app` on Linux and Darwin. The artifact name and `$bin.name` remain `app` on every target. `mach init` generates one artifact using this form. Build, run, clean, required-artifact paths and plan inspection use the same expansion. | Target output format | `bin` suffix | `static` suffix | `shared` suffix | | --- | --- | --- | --- | | ELF on Linux or freestanding | empty | `.a` | `.so` (refused on freestanding) | | Mach-O on Darwin | empty | `.a` | `.dylib` (not written yet, #3588) | | COFF/PE on Windows | `.exe` | `.lib` | `.dll` (not written yet, #3588) | | Raw image | empty | unsupported | unsupported | | SPIR-V module | `.spv` | unsupported | unsupported | The selected object format supplies the naming rules, including explicit target format overrides. Unsupported library forms are errors. A SPIR-V `bin` artifact is its entry module, delivered at `out` with the per-module objects still written under `obj/` (see [Finished-module targets](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#finished-module-targets)). `{artifact.suffix}` is available only in an artifact output template. It does not expand in project output roots, link paths, step arguments or source embeds. Output collisions are checked after expansion among artifacts selected for the target. An explicit literal such as `bin/app.exe` can therefore collide with `bin/app{artifact.suffix}` on Windows. ##### `subsystem` — the windows console/GUI selector ```toml [artifact.game] kind = "bin" entry = "main.mach" out = "bin/game.exe" targets = ["*"] link = [] need = [] subsystem = "gui" ``` A PE executable records in its optional header which environment it wants, and the Windows loader honours it: `"console"` gets a console window attached to the process, `"gui"` does not. mach defaults to `"console"`, which is what every PE it has ever emitted declares, so an artifact that omits the key is byte-identical to one built before the key existed. A graphical application sets `"gui"` to stop an empty console from opening behind it on launch. Only a PE image carries the field. A key written on an artifact that builds for a target whose format has none (ELF, Mach-O, a flat image) is refused as unsupported, naming the key, the target and the format: ``` mach.toml: artifact.game.subsystem: Subsystem gui is unsupported by elf (target 'host' produces a elf image, which declares no subsystem) ``` An omitted key is the console default and is never refused, so an artifact that declares no subsystem still builds everywhere. An artifact that needs the key and also targets a non-windows cell declares two artifacts, one per format, the way `[link.X]` entries carry `os`/`isa`/`abi` axes: the manifest never carries a declaration a build silently ignores. `--subsystem console|gui` overrides the key for one invocation and is refused the same way on a target whose format has no subsystem (`mach help build`). ##### `icon` / `manifest` — Windows executable resources ```toml [artifact.game] kind = "bin" entry = "main.mach" out = "bin/game.exe" targets = ["*"] link = [] need = [] icon = "assets/game.ico" manifest = "assets/game.manifest" ``` On a Windows target, either key adds a `.rsrc` section. `icon` must name a valid ICO container; mach emits each contained image as `RT_ICON` and an `RT_GROUP_ICON` that indexes them. `manifest` is emitted unchanged as `RT_MANIFEST`. A `VS_VERSIONINFO` (`RT_VERSION`) accompanies the declared resources with these schema-derived values: | Version field | Value | |---------------|-------| | `FileVersion`, `ProductVersion` | `[project].version` | | `InternalName`, `ProductName` | the `[artifact.]` table key | | `OriginalFilename` | basename of the resolved executable output, retaining an extension such as `.exe` | There is no `FileDescription`: the live manifest schema has no accepted description field for an artifact. Strings are converted from strict UTF-8 to UTF-16, including surrogate pairs; malformed text, malformed/empty resources, and values that exceed PE's 16/32-bit fields fail the build instead of being truncated. Paths use the same portable `/` spelling as other manifest paths and are resolved against the project root. A generating `[step.X]` must appear in `need` and write the named path before linking. Resource paths and contents participate in the link fingerprint, so changing an asset at the same path relinks a warm build. The keys remain valid in a multi-target artifact, but are completely inert off Windows: mach does not resolve or read either path and ELF, Mach-O, and raw output remain unchanged. `static` and `shared` artifacts reject these executable-only keys. #### `[link.]` — link requirements A `[link.X]` is a named external link requirement. Artifacts reference entries by name in their `link = [...]`; an entry whose filters do not match the build cell is skipped. An entry with `export = true` also applies to any project that links this project's modules, so a platform link requirement lives once — in the manifest that needs it — and cascades to consumers. A standalone build and a consumed build use the same entries, so nothing behaves differently as a dependency. | Key | Required | Meaning | |-----------|----------|---------| | `source` | yes | `"system"` (a system library resolved by name), `"framework"` (a macOS framework), or `"local"` (a file on disk). | | `name` | shape | Library/framework name — required for `source = "system"`/`"framework"`, forbidden for `"local"`. | | `path` | shape | File path — required for `source = "local"`, forbidden otherwise. A template (see below). | | `library` | no | Stable logical name used by `#[library("...")]`; defaults to the `[link.]` table name. | | `symbols` | no | Array of symbol names this dependency provides, attributing imports that have no `ext` declaration to decorate (see below). Written as **source-level** names; the target's C symbol prefix is applied by Mach. Omit for none. | | `os` | yes | Filter axis: a canonical `os` value, `"*"` (any), an array of values, or `[]` (none). | | `isa` | yes | Filter axis over `isa`, same forms. | | `abi` | yes | Filter axis over `abi`, same forms. | | `export` | yes | `true` cascades this entry to consumers; `false` keeps it to this project's own builds. | | `include` | no | `"always"` (the default) names the dynamic library in the linked image whether or not anything imports from it; `"referenced"` names it only when a live import references it, so an unused provider leaves no load command behind. Any other value is a manifest error (`[link.k].include must be "always" or "referenced"`). | The `os`/`isa`/`abi` axes select the build cells an entry applies to. Each takes a single canonical value, `"*"` for any, or an array — `os = "linux"` and `os = ["linux"]` filter identically. `[]` matches nothing (an entry deliberately switched off). A non-canonical spelling is a strict-parse error. An entry applies to a cell when all three axes match. A `local` entry's `path` must, at build time, either match a `[step.X]`'s `out` (which demands that step) or already exist on disk — anything else is an up-front error, so a typo never silently drops an input. A `local` path naming a shared library is validated for the selected target before it is recorded: an ELF `.so` that is not a loadable shared object for the target's architecture (a linker script, a foreign-architecture file) is refused (`'' is not a loadable ELF shared object for the selected architecture`). `library` decouples source attribution from platform loader spelling. Give mutually exclusive platform entries the same logical value when they provide the same API; one unconditional `#[library("glfw")]` can then bind against `libglfw.so.3` on Linux, an `LC_ID_DYLIB` install name on Darwin, and `glfw3.dll` on Windows. Exact canonical loader names remain accepted for compatibility. Selecting two dependencies that map the same logical name to different loader names in one build is an error. A logical name that equals a different dependency's canonical loader name is likewise rejected, so attribution never depends on requirement order. A `#[library]` resolves against the **effective** link set: the artifact's own referenced entries, plus every entry a dependency exports. A binding project therefore names its libraries once and a consumer writing its own `ext fun` against them adds nothing but a `dep` entry. A logical name may belong to an entry that resolves to a **static** input, and that is not something an import can bind to: a static input defines symbols rather than importing them, so a pin naming one means the symbol must come out of that object or archive. When it does, the pin is inert and the link is normal. When it does not, the symbol is undefined, and mach says exactly that — naming the entry, its kind, and the undefined symbol, rather than claiming the library is missing from the link. `symbols` names the symbols the dependency provides. On a two-level-namespace format (PE, Mach-O) every import must identify its provider, and `#[library]` can only attribute a symbol your Mach source declares. A **vendored static archive** leaves its own undefined references — the Win32 calls inside a `glfw3.a`, say — with no declaration to decorate, so the entry that provides them claims them: ```toml [link.kernel32] source = "system" name = "kernel32.dll" library = "kernel32" symbols = ["Sleep", "CreateFileW", "CloseHandle"] os = "windows" isa = "*" abi = "*" export = true ``` Each name is written the way you would write it in **source**, without the target's C symbol prefix. Mach applies that prefix itself, exactly as it does for an `ext fun` declaration, so `symbols = ["Sleep"]` attributes the `Sleep` a Linux or Windows object names and the `_Sleep` a Mach-O object names, and one manifest is correct on every target. The prefix is only ever added, never stripped: `_exit` is a real C symbol whose Mach-O object spelling is `__exit`, so "already prefixed" is not something a spelling can be checked for. Writing the mangled form yourself therefore does not work — on darwin `symbols = ["_Sleep"]` claims `__Sleep`, which nothing imports. Nothing reads a library's export table to derive this, so the claim is what makes cross-linking a PE from a Linux host work with no target DLL present. Claims travel with the entry, so `export = true` cascades them to consumers and a C-binding project declares them once. A symbol may be claimed only once per link: two selected entries claiming it, or a claim contradicting a `#[library]` decorator, is an error naming both claimants rather than an order-dependent win — repeating the *same* claim is fine. Listing a symbol twice within one entry is rejected, and so is a claim on an entry that resolves to a **static** input, which defines symbols rather than importing them. On ELF the key is accepted and validated but changes no emitted bytes, since that loader resolves imports by global search. Whether an input links **statically** or **dynamically** follows the resolved file — a loose `.o`/`.obj` or static `.a`/`.lib` links statically; ELF `.so`, Mach-O `.dylib`, and PE `.dll` inputs are recorded using their format's canonical loader name. An `@rpath/` Mach-O install name also retains the directory where resolution found the dylib, which the executable records as `LC_RPATH`. Darwin frameworks use a version-independent system framework path. See [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md#linking-external-objects) for the `ext fun` workflow that consumes these inputs. #### `[step.]` — build steps A step is a command, make-recipe style, that produces files a build consumes (typically a `local` link input, e.g. a vendored-C object). `` must be a plain identifier — it keys the step's stamp file. | Key | Required | Meaning | |--------|----------|---------| | `argv` | yes | Nonempty array of strings, spawned directly with no shell. `argv[0]` names the executable, by path or resolved on the planner `PATH`. Templates expand in every element. To use a shell, spell it: `["sh", "-c", "…"]`. | | `env` | no | Table of string values added to the step process's environment. | | `in` | yes | Declared input file list. Accepts globs (`*`, `**`), expanded sorted for a stable fingerprint; a glob that matches nothing is a hard error. | | `out` | yes | Declared output file list. Concrete paths only — a glob here is an error, since the demand match and cache key expand `out` verbatim. | | `need` | yes | Array of `step.` requirements or `step.` globs this step must run after. Steps may require only steps. Cycles are manifest errors. `[]` for none. | | `timeout` | no | Duration string (`"30ms"`, `"30s"`, `"5m"`, `"1h"`) after which the step's process group is terminated and the build fails. Omit for an unbounded step. | Steps carry **no filters** and **never run automatically**. A step runs only when **demanded**: - by a selected `[link.X]` whose `local` `path` matches the step's `out`; - by another step's `need`; - by an artifact's `need` (for outputs that are not link inputs), by name or through a glob; - in a dependency, by the `need` of its `default = true` library artifact, which travels to every consumer (see [Dependency requirements travel](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#dependency-requirements-travel)). Because a step has no filter of its own, the condition for running it lives in the link entry that demands it: on a build cell where that entry filters out, the step is never demanded and never runs. A step is cached by content: its declared inputs, resolved executable, expanded arguments, and effective environment contribute to its fingerprint. An unchanged step whose outputs still exist is skipped. Changing an inherited environment value received by the child also invalidates the step. The fingerprint of the last successful run is kept as a stamp in `{project.out}/.cache/steps/`, and a step's outputs under `{project.out}` are written into scratch space in `{project.out}/.stage//` and published only once the step succeeds. Both belong to the compiler: a declared `out` inside `{project.out}/.cache/` or `{project.out}/.stage/` fails at manifest load, naming the step and the path. `mach clean` removes both, so the next build runs every demanded step again. **Bounding a step.** `timeout` gives the step a deadline measured from the moment it is spawned. When the deadline passes, the step's whole process group is signalled — a compiler or archiver the step's shell invoked dies with it rather than outliving the build — the child is reaped, and the build fails naming the step and the bound. Omitting the key leaves the step unbounded, which is the default and the behaviour of every step that does not set it. The value is a string holding a positive integer and a unit, `ms`, `s`, `m` or `h`, the same grammar as `mach test --timeout`. A bare number, zero, a fraction and any other unit are rejected at manifest load. The bound is not part of the step's cache key: changing it does not invalidate a cached step, because it cannot change what the step produces. A source file's `#[embed(...)]` decorator (see [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md#embedstr--compile-time-file-embedding)) is a build input under the same content-based principle, by a different mechanism: it has no `[step]` stanza of its own. The embedded file's content digest is published into a `Q_EMBED_FILE` query input that the embedding module's sema depends on, so an edited asset invalidates that module's cached sema and an untouched asset is a cache hit — the same guarantee `in` gives a step, keyed to one file instead of a step's whole input list, and by digest rather than timestamp either way. **Output homing.** `{project.out}` resolves to the **root** project's expanded `out` in every manifest of the closure. A dependency's step outputs land in the consumer's output tree — exactly as a dependency's compiled modules do — and the dependency's own checkout is never written to. For an exported dependency step, the command receives `{project.out}` as an absolute path rooted at the consumer; it does not depend on the directory from which the consumer was invoked. An ordinary relative consumer root (including `./` segments) and its normalized absolute spelling produce the same expanded command and cache key. **The module object tree is reserved.** `{project.out}/obj//` belongs to the compiler: every module of the project compiles to one object in it, named after the module's path (`src/window.mach` in project `glfw` becomes `{project.out}/obj/glfw/window.o`). A step that writes there collides with those objects by name, and because the link takes whichever file survived, the result is a binary that is subtly wrong rather than a build that fails. **The object tree is the object cache.** `mach build` and `mach test` read it by default: a module whose object in `obj/` was built under the same key is reused as it is: the module is neither lowered nor generated again, and when nothing the build still compiles imports it, it is not resolved or type-checked either. Each module has its own key: the compiler identity, the build configuration, the module's own source and embedded files, and the surface of every module it imports, directly or not. A module's surface is its source without the bodies of its tests and of its functions that are neither generic nor take a comptime parameter, since no importer compiles those; a release build inlines function bodies across modules, so there the whole source is the surface. Editing such a body rebuilds that module alone, and editing a declaration rebuilds the module and the modules that import it. With debug information the surface also covers where each retained declaration sits, so an edit that moves one to another line rebuilds its importers. A reused module reports again the warnings it reported when it was compiled, so a warm build prints what a cold one prints. Each object carries its key in a section no link loads: `.mach.cache` on ELF (not allocated) and COFF (`IMAGE_SCN_LNK_INFO | IMAGE_SCN_LNK_REMOVE`), and `__MACH,__mach_cache` on Mach-O (debug-attributed). The parser consumes it, so an object links exactly as it would without it. The section also carries the image the compiler generated for the module, and a reused module links that image rather than what the object format spells, since COFF and Mach-O cannot spell everything a link reads, such as symbol sizes. A warm build therefore links the binary a cold one links, byte for byte. An object that is missing, has no key, a damaged one or another key is rebuilt, never linked stale. `obj/` holds one object per module, the latest, and each object is written to a sibling temporary and renamed into place, so an interrupted build leaves the previous object or the new one, never a torn file. The digests of the sources the keys read are remembered in `{project.out}/.cache/digests` under each file's path, size, modification time and identity, so an unchanged file is not hashed again for its key; a missing or damaged memo is rebuilt. `--no-cache` forces an uncached build: it reuses no object and writes each one without a key, and `mach clean` removes the tree with the rest of the output. A step output is therefore rejected in that subtree. A declared `out` inside it fails at manifest load, naming the step and the path, before any step runs. A step that writes an object there without declaring it is caught after it runs, with the same message — this covers the common case of a vendored `make` dropping every object it built into the output directory. Pick any other subtree of `{project.out}`. The conventional choice for a vendored library is a directory named after the library rather than after the project, e.g. `{project.out}/obj/miniaudio/` for a project whose own id is `audio`; note that this only stays clear of the reserved tree while the two names differ, so prefer a distinct sibling such as `{project.out}/vendor//`. **Target environment.** Every step process additionally receives the active build cell's target tuple as `MACH_TARGET_ISA`, `MACH_TARGET_OS`, and `MACH_TARGET_ABI`, so the script `argv` invokes can branch on the target without threading it through the template — e.g. `cc --target=$MACH_TARGET_ISA-…`. The step inherits the planner's environment, then applies its declared `env` values, then assigns those three target variables. Names use host identity: case-sensitive on Unix and ordinal case-insensitive on Windows, including Unicode names. A declaration cannot contain names differing only in ASCII case on any host, or names that alias under Windows Unicode comparison on Windows. Each name has one value. A `MACH_TARGET_*` value inherited from an enclosing build or declared by the step is overwritten by the cell's own. The same three values are available in the `argv` templates as the `{target.*}` keys (see [Path templates](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#path-templates)). `cmd` was the pre-4.x spelling of a step's command and is a removed key: `[step.s] uses removed key 'cmd'; use a nonempty 'argv' array`. #### `[dep.]` A dependency is named by its **project id**, and that one name is used in three places: the manifest key `[dep.]`, the directory `dep//`, and the head segment of every module path the dependency exposes (`use .x;`). The compiler checks all three agree: `dep//mach.toml` must declare `id = ""`. A dependency whose id is the declaring project's own is refused where it is declared, since one head segment cannot name two projects. ```toml [dep.std] git = "https://github.com/briar-systems/mach-std" version = "^7.4" ``` A stanza declares exactly one source: | Key | Meaning | |--------|---------| | `git` | Git URL. The dependency is a git **submodule** at `dep//`, pinned by the gitlink the root repository commits. Requires `ref` or `version`. | | `ref` | Selector for `git`: `branch/`, `tag/`, or `commit/`. Any other spelling is rejected (`[dep.std].ref must be branch/, tag/, or commit/`). | | `version` | A [version range](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#version-ranges) over the dependency's releases (see [Releases and resolution](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#releases-and-resolution)). Valid only with `git`; a path dependency has no releases. | | `path` | Local project tree, never fetched. A relative `path` is resolved relative to this manifest's directory. `mach dep add --path` copies its files into `dep//` without the source's own `dep/` or Git metadata. No repository or index is required for a path dependency, and copied files are not automatically staged. Forbids `ref` and `version`. | `git` and `path` are mutually exclusive and exactly one is required. A `git` dependency also names exactly one selector, `ref` or `version`. `ref` and `version` are mutually exclusive (`[dep.std] names both 'ref' and 'version'; keep one`). `ref` selects one exact commit or tag, or follows a branch. `version` selects among releases. ##### Releases and resolution A **release** of a git dependency is a tag `vX.Y.Z` (optionally `vX.Y.Z-pre`) together with the `mach.toml` at that tag. A tag whose manifest does not load, as an old tag's written for an earlier manifest schema, or whose `[project].version` differs from the tag name is not a candidate. Resolution sets it aside and keeps looking. When nothing fits, the error lists it among the requirements (`gl 0.1.0 is not a candidate: its mach.toml does not load: unknown key 'name' in [project]`, or `... its [project].version does not match the tag`), so the error never reads as one in the project's own `mach.toml`. Resolution runs in exactly three places: `mach dep add`, `mach dep update` and `mach dep outdated`. **Builds never resolve.** They verify, offline (see [What a build verifies](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#what-a-build-verifies)). For every identity in the closure that some manifest selects by `version`, resolution picks one release such that: 1. every requirer's range contains it; 2. its own `[project].mach` contains the running compiler; 3. the closure its own manifest implies also resolves. A requirer is the root, a release resolution chose, or a dependency the root reaches by `ref` or `path`. The last declares its range in the closure directly, so two such dependencies naming one identity by range are two requirements of the same problem, and the error names each by its chain (`root -> c requires b <1.2`). Among the choices that satisfy all three, it takes the highest release of each identity. The result is written as gitlinks, like any other pin; there is still no lock file. `mach dep update ` keeps every other identity at its pinned release (the release its recorded gitlink carries) while that release still fits, so an update moves as little as it can. `--all` resolves from scratch. A pin the range no longer admits, as after the range is raised past it, is re-pinned to the release resolution picks; `update` never keeps it. The recorded gitlink is the pin, not whatever the checkout holds: a checkout that drifted from its gitlink is moved to the chosen release and the gitlink staged in the same run, even when the drifted checkout already sits at that release. When nothing fits, the error lists every requirement that took part and names the identity the root can settle: ``` error[dep.unsatisfiable]: no set of releases satisfies every requirement: root requires b ^1.2 a 1.0.0 requires b ^2.0 the root decides by declaring the identity itself, for example: [dep.b] version = "" ``` A release that needs a newer compiler appears as one of those lines (`b 2.0.0 requires mach ^6, and this is mach 5.2.1`). Resolution never silently settles for a lower release than the ranges allow. What `mach dep add` writes: - with `--git ` alone, `version = "^X.Y.Z"`, where `X.Y.Z` is the release resolution picked. The lower bound is the release actually tested when the dependency was added, and the caret follows the pre-1.0 rule; - with `--version `, that range; - with `--ref `, that selector, as before. `mach init` adds std the same way, so a new project names the std release that works with the compiler that created it. std is an ordinary dependency, with no std-specific command. `mach init --no-deps` still resolves and writes the range and skips only the checkout, so it needs the network too. Offline it fails and writes no `[dep.std]` table. A tool that needs the std for a given compiler runs `mach init` and `mach dep pull` in a scratch project and takes what resolution chose. **`--offline`.** `add`, `update` and `outdated` read candidates from each dependency's repository: one `git ls-remote --tags` per URL, and the manifest at a release through a shallow fetch of its tag. With `--offline` they use only the tags already present in the realized checkouts, and they say so (`resolving from releases already fetched (--offline)`). A resolution that needs a candidate it doesn't have fails, naming the identity. Realizing a checkout fetches nothing either. A missing checkout is initialized from the submodule store Git kept for it, and a pin that no local checkout or store holds fails, naming the dependency and the commit. A dependency added over a retained checkout or store is checked out at its selector as held there, and one with neither is not cloned. A selector nothing local holds fails, naming the dependency and the selector. **`--lowest`.** `mach dep update --all --lowest` picks the lowest release every range accepts. A library's CI runs it in a scratch checkout and then builds and tests, which proves the lower bounds it declares are honest. Without that check, `^3.2.0` can quietly depend on something only 3.4 has. It belongs in a release or manually dispatched job, never a scheduled one. **`mach dep outdated `** prints, for each version-selected identity, the pinned release, the highest release resolution would pick now, and the highest release published. A newer release held back by a range or by the compiler is marked as such. **No yanking.** A bad release that is otherwise compatible is fixed forward with a new release. Nothing marks a published version as withdrawn. A consumer that must avoid one raises its range's lower bound (`^3.2.1`). **Forks.** Identity is the project id, not the URL. A root that declares a fork's URL makes that fork the candidate source, so its `vX.Y.Z` tags compete under the same ranges. A fork that wants to stay distinguishable tags pre-releases (`v1.4.3-fork.1`), and a consumer opts in by naming the pre-release in its range. ##### Root declarations: narrowing and overriding A root `version` for an identity **narrows**: it is intersected with every requirer's range, and resolution and verification hold the pin to all of them. A root `ref` or `path` **overrides**: the requirers' ranges and selectors for that identity no longer apply. A range is therefore never widened silently, and the escape hatch is one visible line in the root manifest. An override is always reported. `mach dep pull`, `mach dep update` and `mach dep add` print one note on stderr for each requirement a root declaration replaced, whatever `--quiet` says, naming the identity, the root's winning selection, the requirer chain and what that chain asked for: ``` note: dependency 'std': the root declares ref = "tag/v2.0.0", overriding hedgeacme -> hedge -> std which requires ref = "tag/v2.1.0"; nothing checks that 'std' supports the root's selection ``` A requirement that asks for exactly the root's selection is no override and is not noted. `mach dep list` shows each root declaration's winning selection (`ref=`, `version=` or `path=`, and the recorded `pin=`), its state (`realized`, `missing`, or, for a `version` selection whose range excludes the pinned release, `out of range (the pinned release 7.0.2 is outside ^8.0)`) and, under it, every requirement it overrides (`overrides hedgeacme -> hedge -> std, which requires ref = "tag/v2.1.0"`). `mach dep outdated` names the requirements of a chosen release that a root or a fixed dependency overrides (`root declares b by ref "branch/main", overriding root -> a 1.0.0 requires b ^1.2`). An override is not checked against the requirers. Because the closure is flat, a requirer's `use b.*` binds to whatever the root selected, even a major that requirer was never built or tested against. Nothing proves the requirer supports it: a passing build only shows that the code the build reached compiled, so it is evidence and not a guarantee. `mach dep verify` prints a note for every edge an override replaced, without failing (`note: dependency 'b': the root declares ref = "tag/v2.0.0", overriding root -> a -> b which requires version = "^1.2"; nothing checks that 'b' supports the root's selection`). Treat each note as a claim to confirm, by testing the requirer at that selection or by checking its own range. ##### A release selects only releases The rule follows how a manifest was reached, not where it sits: - A dependency reached through a **release** (a `version` range or an exact `tag/`) may itself select dependencies only by `version` or `tag/`. A release is then reproducible from its tag, all the way down. - A dependency reached through a `branch/` or `commit/` selection is in development, and its manifest may use any selector. A release that breaks the rule is refused wherever it is reached. Resolution stops when it reaches one, and verification (and so every build) refuses it, naming the chain and the offending line: ``` error[dep.release_selector]: root -> a is a release (ref = "tag/v1.1.0"), and its manifest selects [dep.b] by ref = "branch/main"; a release may select its dependencies only by `version` or an exact `tag/`, so it cannot be reproduced from its tag ``` `mach dep verify --release` holds the project itself to the same rule, so a library's release workflow catches the mistake before it tags the release, not when its first consumer resolves it. ##### Pins are gitlinks; there is no lock file The record of which commit a dependency is at is the **gitlink** committed in the root repository, generated into `.gitmodules` by `mach dep`. Nothing else records a pin: there is no `mach.lock`, and a file of that name in the project root is an unrelated file no command reads. A project does not need its own Git repository. In a repository root, Git dependencies use the staged gitlinks as their pins. A subproject, a project in a subdirectory of a repository, uses the gitlink the enclosing repository commits under its prefix (`test/consumer/dep/std` for a subproject at `test/consumer`) the same way: `pull` realizes that gitlink's commit and `update` moves it and stages it. Without such a gitlink, and in a filesystem project, Git dependencies are plain clones whose own checkout commits are verified. Local path dependencies are verified from their filesystem realizations, independently of any Git index. A version range is resolved for the whole closure, not for the root's own declarations alone: a range a dependency declares, whether that dependency was reached by a range, a `ref` or a `path`, is pinned under the root's `dep/` by `mach dep add` and `mach dep update`. When the root has no checkout of the identity yet, resolution starts from the declaring dependency's own committed gitlink for it, so a dependency brings the pin it was tested with; `update --all` moves every range to the highest release all of them admit; and a root declaration of the same identity by `ref` or `path` overrides the range, noted as above. `pull` refuses a range with neither a gitlink nor a checkout under the root and names `mach dep update `, which pins it wherever in the closure it is declared. ##### The root owns the flat closure The root's `dep/` holds every identity in its **transitive** closure, one directory each, one level deep. The root manifest declares only what the root uses directly (plus any override, below); a dependency's own dependencies reach the root's `dep/` through closure computation and never need a declaration in the consumer. A consumed dependency's own `dep/` is never initialized. A package cloned on its own is a root and realizes its own flat `dep/`. So with a root that declares `a`, and `a` that declares `b`, the layout is `dep/a/` and `dep/b/`, and `a`'s `use b.*` resolves against the root's `dep/b/`. Git materializes `a`'s own gitlink as an empty `dep/a/dep/b/` directory; that entry is neither realized, verified, nor descended into. ##### One identity, one commit Identity is the project id, not the key and not the URL. One identity resolves to exactly one commit per build, with no exception for majors: two majors of one identity in one closure is a **clash**, not a case the build accommodates. The diagnostic prints both requiring chains and the exact root declaration that would resolve it: ``` error[dep.conflict]: dependency conflict: project id 'b' is reached with two different selections: root -> a -> b requires git @ tag/v1.0.0 root -> c -> b requires git @ tag/v2.0.0 the root decides by declaring the identity itself, for example: [dep.b] git = "" ref = "tag/v2.0.0" ``` (`` stands for the repository URL as declared.) The root resolves by declaring the identity with a `ref`, which may point at upstream or at a fork carrying the same id. A fork slots in without any consumer source or manifest change, because identity is not the URL. Two unrelated packages claiming one id is a collision and is rejected. URL disagreement is a mirror, not a conflict: the root's declared URL wins, else the first declaring path's; verification compares commits, never URLs. A realized checkout whose remote points somewhere else (a `.git` suffix, another host, a local mirror) verifies by its commit alone. ##### What a build verifies Builds never fetch and never write under `dep/`. Every build (and `mach dep verify` as a command) checks, offline, that: 1. every Git dependency is a clean checkout at its applicable pin (`dependency 'std': checkout is dirty: M mach.toml`), and every path dependency is a contained filesystem tree without repository metadata; 2. its project id equals the directory name; 3. the closure computed from the realized manifests equals the set of directories under `dep/`: nothing missing, nothing extra. A dependency with nothing checked out, whether `dep/` is absent or is the empty directory Git leaves for an uninitialized gitlink on a fresh clone, is refused naming the command that realizes it (`dependency 'std' is not realized (nothing checked out at 'dep/std'); run `mach dep pull ` for project ''`, and for a `version` selection also `or `mach dep update std` when it has no pin yet`); 4. there are no cycles (reported as the chain); 5. every realized manifest's `[project].mach` accepts the running compiler (see [Compiler range](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#compiler-range)); 6. for every identity selected by `version`, the pinned commit carries a release tag, read from the checkout's own refs, and that release is inside every requirer's range, the root's included; and every release in the closure selects only releases (see [A release selects only releases](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#a-release-selects-only-releases)). A pin outside a range names the requirer chain, the range, the pinned release and a runnable remedy (`dependency 'vb': root -> vb requires version '^1.2' but the pinned release is 1.1.0; run `mach dep update vb` for project '' to re-pin it, or declare the identity at the root to override`). A pin no tag in the checkout names reads `... but no release tag in its checkout names the pinned commit ''`, and names `mach dep pull` beside `mach dep update`. A checkout that holds no tags at all, as a shallow clone (`git submodule update --depth 1`) does, cannot say which release its pin is, and the refusal says so rather than claim the pin is outside the range (`... but the release of the pinned commit '' cannot be read: its checkout at holds no tags (a shallow clone fetches none); run `mach dep pull ` for project '' to fetch them`). The recorded gitlink is the pin, and two kinds of drift from it are refused, each naming the identity and the command that fixes it: - a `dep/` checkout at another commit than its gitlink (`dependency 'b': the checkout is at '' but the recorded gitlink is ''; run `mach dep pull ` for project '' to restore the recorded pin, or `mach dep update b` to re-pin it to the manifest's selection`); - a gitlink outside the manifest's selection. The root's own `tag/` or `commit/` must be satisfied by the pin (`dependency 'b': exact ref 'tag/v1.0.0' required by root -> b resolves to '' but the realized commit is ''; run `mach dep update b` for project '' to re-pin it to the root's selection`; a root `commit/` that does not match reads `exact commit ref 'commit/' is not satisfied by the realized commit ''`), and so must every range the root declares (item 6). `pull` realizes the gitlink as it is, so it never cures this; `update` moves the gitlink to the selection. A root `ref` or `path` is the override for its identity, so the requirers' selectors are not checked against its pin (the override notes above name them). For an identity the root does **not** declare, every requirer's exact selector (`tag/`, resolved through the checkout's own refs, or `commit/`) must be satisfied by the realized commit; a mismatch names both commits and the two remedies (`dependency 'b': exact ref 'tag/v1.0.0' required by root -> a -> b resolves to '' but the realized commit is ''; run `mach dep update b` for project '' to re-pin it, or declare the identity at the root to override`). A `branch/` selector is an input to `update`, never a verify fact. The verifier reads the git **index**, so a freshly realized dependency is verifiable before it is committed. `mach dep pull` reads what a Git dependency's `dep/` holds together with its record (the staged gitlink, its `.gitmodules` entry, and any module directory Git retained) and takes the one step that brings it to what a build verifies: - a staged gitlink with nothing checked out, as on a fresh clone or after the directory was deleted, is initialized in place (`realized std @ … (initialized the committed gitlink)`), first restoring its `.gitmodules` entry from the manifest if that entry is gone; - a checkout at another commit than its gitlink is checked out at the gitlink; - a clean checkout of its own with no gitlink is registered, moved to the declared selector from the declared source, and staged (`(registered the existing checkout)`); - with neither, the submodule is added at the selector, reusing a module directory Git retained from an earlier removal. Then, when the checkout lacks the tag its selection is read by (the tag a `tag/` names, or a release tag naming a `version` selection's pin), pull fetches its tags (`fetched the tags of std`). `update` and `outdated` fetch them the same way to read a pin's release, unless `--offline`. A symlink, a file, a directory that is not a checkout of its own, and a dirty checkout that would be registered are refused and left as they are. `mach dep add` takes the same step for its Git source, so re-adding a dependency whose checkout `remove` retained registers that checkout, and it refuses a dirty one. No gitlink command ever runs against a path that is not a checkout of its own. A path dependency has no pin, so `mach dep pull` syncs its `dep/` with the declared `path` every time, and `mach dep update` does the same, once per command. Unless `--quiet`, each says whether the copy was refreshed or reused (`realized hedge from ../.. (copy refreshed from its source)`, or `(copy reused: it already matched its source)`), so a copy left from another checkout cannot pass unnoticed. A build reads the copy as it is and never syncs it. A changed `path` realizes the new source. A file the source no longer has is removed and named, and a file whose content differs from the source is overwritten and named (`replaced 'src/lib.mach' with its source's content`), so local edits to the copy do not survive a pull. A `dep/` that is a symlink is refused and left as it is. A project root is identified by its own `mach.toml`, not by an enclosing git repository; `dep/` is resolved relative to the project root. A project nested inside an unrelated repository or without any repository builds. Git dependencies in these projects are verified from their own plain checkouts. ##### Selection on `update` `mach dep update` is the only command that moves a pin. It advances every `branch/` selector to its current remote tip and re-stages the gitlink, and it moves an exact selector to the commit it names, so an identity realized at a dependency's selection lands on the root's declaration once the root declares one (`b: 0564… -> e508… (pinned to the exact selector)`, or `(exact selector, already pinned)` when nothing moves). `` is looked up in the whole dependency closure, so `mach dep update b` for an identity the root does not declare moves its checkout to the selector its requirers declare, and `--all` moves every selector in the closure. Resolution reads a dependency selected by `ref` at the commit its selector names, never from its checkout, so a requirer that has just moved its selector is resolved against the manifest it now selects. A name outside the closure is refused (`dependency 'x' is not in the dependency closure`). For an identity reached by more than one path, one rule decides: the root's selector wins if the root declares the identity; otherwise agreement among the requirers is taken; otherwise the command stops, prints both chains, and names the root declaration that would decide (the diagnostic above). A consumed dependency's own gitlink records a tested commit, readable without initializing that dependency's `dep/`. It is not a compatibility floor. Compatibility is stated by ranges, and `update` resolves every version-selected identity as described in [Releases and resolution](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#releases-and-resolution). ##### Removed forms Two shapes of a realized closure are refused by `pull`, `verify` and every build: - a **key that is not the project id**, a `[dep.]` whose realized project declares a different id: `[dep.foo] realizes project 'std': the manifest key, the directory under dep/, and the project id are one name, so rename the table to [dep.std] and the directory to dep/std`; - a **nested realization**, a `dep//dep//mach.toml`: `dependency 'a': dep/a/dep/b is a nested realization: the root's dep/ owns the flat closure and a dependency's own dep/ is never realized, so delete dep/a/dep`. The empty directory git materializes for a consumed dependency's own gitlink is not a realization and passes. Command-line usage (`pull`, `verify`, `add`, `update`, `remove`, `list`) is documented by `mach help dep`. A dependency's export surface — all a consumer sees — is its source module tree (addressed by the dep's id), the module a bare `use ;` binds, its `export = true` link entries, and the steps those entries demand. Nothing else in a dependency's manifest applies to consumers. #### Path templates Paths and `cmd`s expand over a closed, final set of eight variables: - `{project.out}` — the **root** project's expanded `[project].out`, in every manifest of the closure. - `{target.name}` — the resolved target name (never the literal `native`). - `{target.isa}` — the resolved target's `isa` (e.g. `x86_64`). - `{target.os}` — the resolved target's `os` (e.g. `linux`). - `{target.abi}` — the resolved target's `abi` (e.g. `sysv64`). - `{profile.name}` — the selected profile name. - `{artifact.suffix}` — the conventional filename suffix of the artifact being named, for its kind on the selected target (`.exe`, `.a`, `.so`, ...). Available only in an artifact's own `out`. See [Artifact filenames and identity](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-filenames-and-identity). - `{artifact..out}` — the output path of a required artifact, relative to the root project's directory exactly as `{project.out}` is. In a dependency's module it names a requirement of that dependency's default library artifact, homed under `dep/`. See [Artifact requirements](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#artifact-requirements). The three `{target.*}` tuple keys are also exported to every step process as `MACH_TARGET_ISA`/`MACH_TARGET_OS`/`MACH_TARGET_ABI` (see [build steps](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#stepname--build-steps)). An artifact's `out` is relative to the expanded project `out` and is rooted there automatically — write `bin/demo`, not `{project.out}/bin/demo`. Step `out` lists and local link `path`s are **not** auto-rooted: they name `{project.out}` explicitly, which is what homes a dependency's build products into the *consumer's* output tree rather than the dependency's checkout. There are no `{name}`/`{ext}` or bare `{target}`/`{profile}` aliases. An unresolvable `{...}` reference, or an unterminated `{`, is a strict-parse error. `{project.out}` is not available inside `[project].out` itself (it would be self-referential), and `{artifact..out}` is not available inside an artifact's own `out` for the same reason. `{artifact.suffix}` is available nowhere but an artifact's own `out`. Two artifacts selected for one target that resolve to the same `out` path collide and fail at build start. #### Artifact requirements An artifact's `need` names what must exist before it is built. Every entry identifies its category: `step.generate` selects `[step.generate]`, while `artifact.support` selects `[artifact.support]`. A `*` glob applies only within that category, so `artifact.shader-*` never selects a similarly named step. A step and an artifact may share a name. List both qualified names to require both. A required artifact is built before its consumer, for the consumer's profile and for the required artifact's own targets: the consumer's target when the requirement declares it, and otherwise every target the requirement names. A requirement reached from several consumers is built once. Inside the consumer, `{artifact..out}` expands to that artifact's output path. It is an error to name an artifact the consumer does not require, or one that builds for several targets here and therefore has no single output. ```toml [artifact.shader-blur] kind = "bin" entry = "shaders/blur.mach" out = "shaders/blur.spv" targets = ["vulkan"] link = [] need = [] [artifact.app] kind = "bin" entry = "main.mach" out = "bin/app" targets = ["linux-x86_64"] link = [] need = ["artifact.shader-*"] ``` ```mach fragment #[embed("{artifact.shader-blur.out}")] val BLUR: [_]u8; ``` An `#[embed]` argument holding a template resolves against the **project root**, as the template's own value does; a literal `#[embed]` path keeps resolving against the declaring file's directory. Only `{artifact..out}` may appear in an `#[embed]` path. Requirements are written within one manifest: a `need` entry never names another project's artifact or step, and a consumer cannot add to a dependency's `need`. A dependency's own requirements reach the consumer by travelling, below. These are errors: - a missing or malformed category prefix, including bare names; - a name or glob matching no declaration in its named category; - an explicit self-requirement, or a glob matching only the declaring item; - an artifact requirement in a step's `need` list; - artifact cycles and step cycles, including cycles formed by globs. Globs exclude the declaring item. Matching declarations retain manifest order, and transitive prerequisites run before their consumers. Both root and dependency manifests receive these checks during parsing, before planning can execute a step. A required artifact that fails to build fails its consumer, naming the requirement, and the consumer is not attempted. `mach check` builds nothing, and that includes required artifacts. An `#[embed]` is read in the frontend, so a consumer that embeds a requirement's output cannot be checked on a tree that has never been built: the check reports the output it cannot read, names the artifact whose output it is, and says to build first. One `mach build` produces the outputs and every later check of unchanged requirements is clean, so a pipeline that checks before it builds should build first. ##### Dependency requirements travel A library that embeds what it builds cannot be consumed if its requirements stop at its own manifest, and a consumer has no way to declare them. So a dependency's **`default = true` library artifact** carries its requirements as part of its export surface: the artifacts and steps its `need` names are built for any project whose dependency closure holds that dependency, before the cells that compile against the closure. Only that artifact's `need` travels. Several library artifacts may share the `default` marker and the public entry; their requirements travel together. ```toml # the dependency's mach.toml [artifact.shlib] default = true kind = "static" entry = "lib.mach" out = "lib/shlib{artifact.suffix}" targets = ["linux-x86_64"] link = [] need = ["artifact.shader-frag"] [artifact.shader-frag] kind = "bin" entry = "shaders/frag.mach" out = "spv/frag{artifact.suffix}" targets = ["spirv"] link = [] need = [] ``` ```mach fragment # the dependency's src/lib.mach, compiled by every consumer #[embed("{artifact.shader-frag.out}")] val FRAG: [_]u8; ``` The rules: - **The consumer's profile, the dependency's targets.** A travelling requirement is built with the profile the consumer resolved, for every target it names in the dependency's own manifest. The consumer's target never selects among them, because target names of two manifests are unrelated; a requirement naming several targets has no single `{artifact..out}`, exactly as within one manifest. - **Homed in the consumer, namespaced by id.** The output is `/dep//`, so two dependencies that both declare `shader-quad` produce two files and neither writes into its own checkout. `{project.out}` keeps meaning the root's out in every manifest of the closure, as it does for a dependency's steps, and the cell's objects sit in the root's `obj/` beside every other module's. `mach clean` removes that home and those objects with the rest of the output, reading the realized dependency manifests for the target names the root never declares. - **Scope follows the module.** `{artifact..out}` in a module the dependency owns is read in that dependency's manifest and checked against its default library artifact; the root's modules keep reading the root's manifest and the requirements of the artifact being built. Naming anything else is the ordinary refusal, located at the scope it was read in. - **Its own closure, recursively.** A travelling requirement compiles against the dependency's own transitive closure, realized in the root's flat `dep/`, and the requirements of *those* dependencies' default library artifacts travel to it in turn. A requirement reached through several consumers is built once, and a failure names the dependency chain (`dependency app -> boom -> shader: ...`). - **Steps too.** A step the default library artifact's `need` names runs for the consumer, alongside the steps its `export = true` link entries demand. Nothing here makes an `#[embed]` an edge in the build graph: a missing embedded file is still a compile error and still triggers nothing (#2887). What runs is the requirement the dependency declared. #### Selection and the build matrix A build cell is one artifact × one target × one profile. Every command that builds, checks, tests, runs or documents cells selects them with three options, one per axis: - `-a, --artifact ` selects `[artifact.]` entries; - `-t, --target ` selects `[target.]` entries; - `-p, --profile ` selects `[profile.]` entries. A pattern is an exact name, which must be declared, or a glob in which `*` matches any run of characters and `?` any one character, which must match at least one entry. Each option repeats, and the axis takes every entry any of its patterns names, in declaration order. Artifact names are unique table keys, so an artifact's kind never needs naming. Quote a glob so the shell leaves it alone: `-a '*'`. A value that names no entry but does name a file or directory in the working directory is refused as the shell's expansion of an unquoted wildcard, with a hint to quote it. `--all` fills every axis no option names with `*`: `mach build . --all` builds every artifact on every target it supports in every profile, and `mach test . --all -p debug` does the same in `debug` only. An axis no option names, without `--all`, takes the manifest's default: - the target is the [`native` target](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#native-target-resolution), or the one a named artifact settles (see [below](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#a-named-artifact-can-settle-the-target)); - the profile is the sole declared one, or the one marked `default = true`; - the artifacts are the **default selection** for each selected target: of the artifacts whose `targets` includes it, those marked `default = true` when any is marked, and every one of them when none is. `mach build` and `mach check` take the whole default selection. `mach test` and `mach doc` need one artifact and take the default selection when it holds one; several with none marked are refused. `mach run` takes the sole `bin` the target builds. No default is chosen by table order: several candidates with none marked are refused, naming them. - `mach build ` and `mach check ` build and check every selected cell. A selection that spans several profiles plans and runs one profile after another. - `mach test ` builds a test dispatcher for every selected cell as `mach build` would build the cell, its closure, its `link` entries, its `need` and exported dependency entries, and links the dispatcher in place of its entry. Tests then run once per (target, profile): the tests every selected artifact reaches there are combined, each qualified name running once. Only a target whose `os` and `isa` are the host's runs; every other (target, profile) is built, reported on a `skip` line, and not run, whatever emulation the host has. `--runner ` runs a foreign target's tests through a command and needs the selection to resolve to one cell. A run in which nothing was runnable exits `1`, so a green run always ran something. - `mach run ` and `mach doc ` consume exactly one cell and refuse a selection that resolves to several, naming them. `mach run` takes no `--all`, and `mach doc` selects with `-a` and `-t` only. ##### Enumerated cells are filtered; named ones are not A cell whose artifact does not list the cell's target is a cell the manifest never declared, so a selection that reaches it through a glob skips it. Naming both halves of that pair exactly is a different act: `-a kernel -t linux-x86_64` is refused by name, because you asked for a cell that does not exist. `-a kernel -t '*'` globs the target axis and so filters back to the targets `kernel` declares. If a selection is well-formed but holds no cell — a `-t` no artifact lists, or globs that only pair unsupported cells — it fails naming what it selected, rather than succeeding with an empty plan. ##### A named artifact can settle the target `-a ` with no `-t` lets the artifact decide, since its `targets` list may already leave only one answer: - exactly one declared target: that target is used, and `-t` would only repeat what the manifest already said. A hosted target that does not match the host is refused instead (see [`native` target resolution](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#native-target-resolution)) - several, one of which matches the host: the host target, as before - several, none matching the host: refused, naming the targets the artifact does declare so the choice is visible without opening `mach.toml` An explicit `-t` always wins, including when it names a target the artifact does not list — that pair is still refused by name. This only applies to a named artifact: an artifact axis left to the default keeps the target fixed for the whole matrix, so a bare `mach build ` never widens into a target it was not asked for. ##### `-o` names one output `-o` is accepted exactly when the selection resolves to a single (artifact, target, profile), and refused otherwise, naming the cells it resolved to. Two artifacts collide on one output path the same way two targets or two profiles do: each would link over the previous, leaving only the last with no warning. Narrow with `-a`, `-t` and `-p`. `-o` names a canonical path inside the project root, as an artifact's `out` does: relative, `/`-separated, with no `.` or `..` component and no empty one. `-o ../mach`, `-o ./mach` and `-o /tmp/mach` are refused with `-o must name a canonical path inside the project root`, so a build never writes outside the tree it was asked to build. ##### When one cell fails Every cell is attempted; a failure does not abandon the ones after it. Each cell's diagnostics are reported under its own heading as it happens, and every cell that succeeded leaves its artifact on disk at its own path — nothing is rolled back. The exit code is `0` when all cells succeeded, and otherwise the code of the worst failure among them: `2` when any cell failed internally, else `3` when any failed for the environment, and `1` otherwise. Artifacts cannot share an output path: a manifest whose expanded `out` templates collide is rejected before the build starts, and so is a selection spanning profiles whose `[project].out` has no `{profile.name}` to keep them apart. ##### `native` target resolution `native` resolves the host's `(isa, os)` against the **declared** targets only — never a synthesized tuple, and never a target the host cannot run. Exactly one host match is chosen; several matching tuples is an ambiguity error naming the candidates. With no match `native` is an error, however many targets are declared and however they are marked, and it is raised before any step runs. A declared target that does not match the host is built only when `-t` names it: ``` error[selection.no_host_target]: mach.toml: 'native' matches no declared target: the host is aarch64-linux and the declared targets are linux-x86_64 (x86_64-linux), windows-x86_64 (x86_64-windows); declare a [target.] for the host or select one with -t ``` A cross-only project whose targets are hosted (`linux`, `darwin`, `windows`) selects its target with `-t`. With no `-t`, [an artifact can settle the target](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md#a-named-artifact-can-settle-the-target) when its `targets` list leaves one answer. The same rule holds there. An artifact whose only target is hosted and does not match the host is refused with `selection.no_host_target`, naming the artifact, because building it would be the same fallback. An artifact whose only target no host runs as `native` (a `freestanding` target, including a finished-module target such as `spirv`) is still pinned to it: such a target is never `native`, so naming the artifact selects it explicitly, and a host artifact that `need`s it builds it on any host. This path is taken whenever an artifact is settled before its target: `mach build` and `mach check` with `-a`, `mach run` and `mach test` (which also settle on a sole artifact), and editor analysis. A plain `mach build` or `mach check` resolves `native` first and never reaches it. A manifest with no `[target.*]` table has the synthesized host target. `[target.*] default = true` no longer means anything: it is accepted and ignored, with a `target.default_deprecated` warning, and will be removed in a later major release. Migrate by deleting the key and declaring a target for each host the project builds on, or passing `-t`. The same rule applies to `[profile.*]` and to `[artifact.*]` when a command needs one artifact. #### Worked example: a consumer of C bindings and vendored C A project that uses a system-lib binding (`glfw`) and a vendored-C library (`miniz`), building a native binary and cross-compiling to windows. The platform shim is built by steps and linked through `local` entries; the OS-specific shims are gated by their link entries' `os` axis, so the x11 step runs on a linux build and never on a windows one. ```toml [project] id = "demo" version = "0.1.0" mach = "^5.3" src = "src" out = "out/{target.name}/{profile.name}" [dep.std] git = "https://github.com/briar-systems/mach-std" ref = "tag/v0.17.0" [dep.glfw] git = "https://github.com/briar-systems/mach-glfw" ref = "tag/v0.2.1" [dep.mz] git = "https://github.com/briar-systems/mach-miniz" ref = "tag/v1.0.3" [link.shim] source = "local" path = "{project.out}/obj/platform/shim.o" os = "*" isa = "*" abi = "*" export = false [link.shim-x11] source = "local" path = "{project.out}/obj/platform/x11.o" os = "linux" isa = "*" abi = "*" export = false [link.shim-win32] source = "local" path = "{project.out}/obj/platform/win32.o" os = "windows" isa = "*" abi = "*" export = false [link.gl] source = "system" name = "GL" os = "linux" isa = "*" abi = "*" export = false [artifact.demo] kind = "bin" entry = "main.mach" out = "bin/demo" targets = ["linux", "windows"] link = ["shim", "shim-x11", "shim-win32", "gl"] need = [] [step.shim] argv = ["cc", "-c", "-O2", "-fPIC", "-Ivendor/platform", "-o", "{project.out}/obj/platform/shim.o", "vendor/platform/shim.c"] in = ["vendor/platform/shim.c"] out = ["{project.out}/obj/platform/shim.o"] need = [] [step.shim-x11] argv = ["cc", "-c", "-O2", "-fPIC", "-Ivendor/platform", "-o", "{project.out}/obj/platform/x11.o", "vendor/platform/x11.c"] in = ["vendor/platform/x11.c"] out = ["{project.out}/obj/platform/x11.o"] need = [] [step.shim-win32] argv = ["cc", "-c", "-O2", "-fPIC", "-Ivendor/platform", "-o", "{project.out}/obj/platform/win32.o", "vendor/platform/win32.c"] in = ["vendor/platform/win32.c"] out = ["{project.out}/obj/platform/win32.o"] need = [] [target.linux] isa = "x86_64" os = "linux" abi = "sysv64" [target.windows] isa = "x86_64" os = "windows" abi = "win64" [profile.debug] opt = 0 debug = true simd = "scalarize" [profile.release] opt = 2 debug = false simd = "scalarize" ``` The `gl` and `shim-x11` entries carry `os = "linux"`, so on a windows build cell they filter out (and `shim-x11`'s step is never demanded); `shim-win32` carries `os = "windows"` and applies only there. The unconditional `shim` entry (`os = "*"`) applies to both. #### Worked example: a C-binding dependency's export `mach-glfw` exports its `system`/`framework` link entries — a consumer that imports its modules inherits every `export = true` entry that matches the build: linux and darwin pull the `glfw` system library, a windows build pulls `glfw3.dll`, and the darwin frameworks apply only on darwin. ```toml [project] id = "glfw" version = "0.3.0" src = "src" out = "out/{target.name}/{profile.name}" [link.glfw] source = "system" name = "glfw" library = "glfw" os = ["linux", "darwin"] isa = "*" abi = "*" export = true [link.glfw-win] source = "system" name = "glfw3.dll" library = "glfw" os = ["windows"] isa = "*" abi = "*" export = true [link.Cocoa] source = "framework" name = "Cocoa" os = ["darwin"] isa = "*" abi = "*" export = true ``` Both GLFW entries expose the logical name `glfw`, so the binding can use the same attribution on every target: ```mach #[library("glfw")] pub ext fun glfwInit() i32; ``` #### Worked example: a vendored-C dependency `mach-miniz` exports one `local` entry whose path is produced by a step; importing its surface pulls the entry, and the entry's path demands the step in the consumer's output tree: ```toml [project] id = "mz" version = "1.0.3" src = "src" out = "out/{target.name}/{profile.name}" [link.miniz] source = "local" path = "{project.out}/obj/miniz/miniz.o" os = "*" isa = "*" abi = "*" export = true [step.miniz] argv = ["cc", "-c", "-O2", "-fPIC", "-Ivendor/miniz", "-o", "{project.out}/obj/miniz/miniz.o", "vendor/miniz/miniz.c"] in = ["vendor/miniz/*.c", "vendor/miniz/*.h"] out = ["{project.out}/obj/miniz/miniz.o"] need = [] ``` #### The compiler's own manifest Mach builds itself from a manifest that declares six targets, two explicit profiles, two binary artifacts with literal output paths, and one dependency, `std`: ```toml [project] id = "mach" version = "5.0.0" mach = "^5.3" src = "src" out = "out/{target.name}/{profile.name}" [target.linux-x86_64] isa = "x86_64" os = "linux" abi = "sysv64" [target.windows-x86_64] isa = "x86_64" os = "windows" abi = "win64" stack_reserve = 0x800000 [profile.debug] default = true opt = 0 debug = false simd = "scalarize" vectorize = true float_reassoc = false [profile.release] opt = 2 debug = false simd = "scalarize" vectorize = true float_reassoc = false [artifact.mach] kind = "bin" entry = "bin/main.mach" out = "bin/mach" targets = ["linux-x86_64", "linux-arm64", "linux-riscv64", "darwin-x86_64", "darwin-aarch64"] link = [] need = [] [artifact.mach-windows] kind = "bin" entry = "bin/main.mach" out = "bin/mach.exe" targets = ["windows-x86_64"] link = [] need = [] [dep.std] git = "https://github.com/briar-systems/mach-std" ref = "commit/e6fc41251e442eb736d4d15de677903a4ba52461" ``` (The full manifest declares all six targets; `version` is whatever the tree's current release is, and the `std` selector is the exact std 2.0.0 commit the tree builds against until the tag is cut.) `mach build .` selects the host-matching target via `native`, compiles `src/bin/main.mach` and its transitive imports — including modules from `std` at `dep/std/` — and links `out///bin/mach`. The two artifacts keep literal outputs rather than one `bin/mach{artifact.suffix}` because the published seed compiler that bootstraps this tree predates the template; a manifest the seed must parse stays within what the seed accepts. `dep/std` is a git submodule whose gitlink is the pin; the build resolves it purely by that directory and verifies the gitlink from the repository's index, fetching nothing. #### See also - [files.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/files.md) — file layout and `lib.mach` / `main.mach` - [modules.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/modules.md) — how files map to module paths - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) — linking against external symbols Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ir-output.md ### IR output `mach build --emit-ir` writes the middle-end IR of each module beside its object, as `/ir//.ir`. The flag takes an optional form: ``` --emit-ir the ir-debug dump (the default) --emit-ir=debug the same dump, named --emit-ir=listing a readable listing ``` `--emit-asm` is the same idea one stage later, and takes no form. On x86-64 the asm listing is GNU as source in intel syntax: a module's listing assembles with `as --64` to the same `.text` bytes as the module's object. Every sized memory operand is written `qword ptr [...]`, a direct branch carries `{disp32}` because mach always encodes a 32-bit displacement, and a block's label is `.L_`, numbered within the module. A symbol a module calls or addresses stays undefined in the listing, as it is a relocation in the object. An `asm` block's branch names its target as the block does: the symbol, or the numbered label (`1f`, `1b`) that the listing defines where the block does, which GNU as resolves to the same definition. #### `debug` — the ir-debug dump The dump is the compiler's own debugging view. Every instruction carries its whole record, so no two distinct instructions can print the same text: ``` %5 = mul !0 %p0 {kind=1 ty=!0 secret=false}, %p0 {kind=1 ty=!0 secret=false} ; state{ty=!0, aux=!nil, kind=2, operands=2, flags=0, loc=1:82, secret=false, pure=false, ...} ``` Types are type-table rows (`!0`), positions are `file:byte-offset`, and every flag prints whether it is set or not. That completeness is the point: a compiler change that alters an instruction alters its text. **The dump's text is not a contract.** It tracks the IR's internal shape and changes with it. Do not parse it. #### `listing` — the readable form The listing prints the instruction and its source position and nothing else: ``` ir-listing stage="post-codegen" target="linux-x86_64" isa="x86_64" os="linux" abi="sysv64" of="elf" module demo.main { fun @demo.main.square(%p0: i64) i64 [inline] { ; src/main.mach bb0: %4 = mul.pure i64 %p0, %p0 ; 2:30 ret %4 ; 2:26 } } ``` - **Types are spelled the way the language spells them**: `i64`, `*i64`, `[4]i64`, `i32x4`, `rec{i64, f32}`, `fun(i64) i64`. A type that nests deeper than the spelling allows, which a self-referential record does, falls back to its type-table row (`!7`). - **Positions are `line:column`**, one-based, resolved against the source at print time. The IR itself still holds a byte offset; nothing about what it stores changed. - **A function's header names the file its body lives in**, and a bare `line:column` is measured against that file. An instruction from another file — what an inlined callee leaves behind — prints `path:line:column` instead, so every line stands on its own. - **An instruction with no position prints no comment at all.** - **Flags that change what an instruction means still print**, as suffixes on the opcode: `.nsw`, `.nuw`, `.exact`, `.volatile`, `.secret`, `.pure`. A secret operand keeps its `~`. Everything else — the type-table rows, the operand records, the raw flag words, the metadata ids — is elided. - If the source an offset indexes is no longer loaded, the position prints as `file#@` rather than a line and column it cannot compute. **The listing is the form tooling reads.** It is what a consumer such as Compiler Explorer maps back to editor lines. #### See also - [manifest.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/manifest.md) — the `mach.toml` manifest reference - `mach help build` — the full command-line reference Source: https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/comptime-attrs.md ### Symbol attributes Symbol attributes are expressed as **decorators** — leading, per-declaration codegen directives on the declared symbol. A decorator is written as an attribute: `#[name]` for a bare flag or `#[name(args)]` for a directive with arguments. See [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) for the full reference. ```mach fragment #[symbol("main")] fun entry(argc: i64, argv: **u8) i64 { ... } #[library("ws2_32.dll")] #[symbol("WSAStartup")] ext fun wsa_startup(ver: u16, data: *u8) i32; #[align(64)] pub var cache_line: u8 = 0; ``` A backtick is not a token: one anywhere in source is a lexer error. A comptime directive takes no `=`; a stray one is a parse error at the directive's terminator. #### See also - [decorators.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/decorators.md) — full decorator reference - [ext-fun.md](https://github.com/briar-systems/mach/blob/v6.10.1/doc/language/ext-fun.md) — `ext` imports and `library` / `symbol` use cases