zasm performance
zasm is a typed SSA language whose checker proves memory safety and data-race freedom, so compiled code carries no runtime checks. Here it is measured against Rust and C on 13 programs, each written with the same algorithm and data layout in all three languages.
Release builds, program by program
zasm's run time divided by each other build's. Left of the parity line zasm is
faster. Log scale. Rust is rustc -C opt-level=3 -C target-cpu=native; C is
-O3 -march=native -ffp-contract=off.
Table: run times in seconds
Development builds
The fast-to-compile builds against each other: zasm's Cranelift backend,
rustc -C opt-level=0 and clang -O0. How many times longer the other
binary runs than zasm's; right of the parity line zasm is faster. Log scale.
Table: run times in seconds
Compile time
Whole build including linking, geometric mean over the 13 programs, in milliseconds.
Development builds
Release builds
Table: compile times in milliseconds
How to read this
- Where zasm is faster, its compiler does something the others' do not, on the same source. zasm-opt, an optimizer on zasm itself, runs 16 rows of a loop nest together so that LLVM vectorizes across rows: for an inner loop summing floats in order (spectralnorm) this is gcc's outer-loop vectorization, which no LLVM-based compiler does; for an inner loop whose length depends on the data (mandelbrot) finished rows are masked off, as hand-written SIMD does, and no compiler here does that. It also inlines small recursive functions into themselves (binarytrees), which gcc does and LLVM never does. Results stay bit-for-bit identical.
- It is not tuned to these 13 programs. 48 more programs were written by
agents that did not know the optimizer, in three rounds; each round was measured before the
optimizer changed, and the next one checked that the changes carry over (inlining a function
that holds the inner loop, flattening branches in it, counters of any kind). With the final
compiler, zasm-opt changes 14 of the 48 and makes each faster (0.28x the cycles over those
14, up to 12x on a Newton iteration), and leaves the other 34 with identical code. See
bench/heldout/README.md. - The optimizer is not trusted. Every function it changes carries the facts its new code needs and goes through the checker again; it is kept only if the checker accepts it. The borrow rules make the rows' independence a local check: loads only through shared borrows, stores only after the inner loop, in row order. Fuzzers compare optimized and original programs in the interpreter and in native builds.
- Two LLVM settings differ. LLVM hoists a vectorized loop's runtime alias checks into the enclosing loop; in Floyd-Warshall one row always overlaps, so the hoisted check always failed and the scalar fallback always ran. zasm turns that off (floyd 2.7x faster). And on x86-64 it keeps conditional moves that LLVM would turn into branches, which miss on data-dependent choices: chosen on held-out programs, it is neutral on these 13 and 0.96x on 45 programs, from 0.30x (kmeans) to 1.13x (smoothing). Both are compiler options, not language properties: clang and rustc users can pass them too.
- Where every compiler does the same thing, zasm is at parity with Rust and C compiled by the same LLVM (hashtable, matmul, sieve, quicksort, psort, wordcount). What remains: gcc is faster on binarytrees (it inlines deeper, and the C builder has no bounds checks, which zasm keeps because its facts cannot express a tree's size) and on fannkuch; nbody is 1.10x behind Rust (LLVM's SLP vectorizer packs zasm's loop into narrower vectors, with the same instruction count).
- binarytrees_region is an idiom comparison, shown separately: trees built in
zasm regions against safe std-only Rust with a
Boxper node and C with amallocper node. - Four surprising results were flawed comparisons and are gone. Two zasm wins over Rust (wordcount, binarytrees) came from weaker Rust code. Two C wins came from the setup: binarytrees rebuilt an identical tree every iteration, so clang 23 built one and multiplied, and clang fused multiply-adds by default, a different rounding from Rust and zasm.