zasm performance
zasm is a typed SSA language whose checker proves memory safety and data-race freedom, so compiled code carries no runtime checks. Here it is measured against Rust and C on 13 programs, each written with the same algorithm and data layout in all three languages.
Release builds, program by program
zasm's run time divided by each other build's. Left of the parity line zasm is
faster. Log scale. Rust is rustc -C opt-level=3 -C target-cpu=native; C is
-O3 -march=native -ffp-contract=off.
Table: run times in seconds
Development builds
The fast-to-compile builds against each other: zasm's Cranelift backend,
rustc -C opt-level=0 and clang -O0. How many times longer the other
binary runs than zasm's; right of the parity line zasm is faster. Log scale.
Table: run times in seconds
Compile time
Whole build including linking, geometric mean over the 13 programs, in milliseconds.
Development builds
Release builds
Table: compile times in milliseconds
How to read this
- Where zasm is faster, its compiler does something the others' do not, on the same source. zasm-opt, an optimizer on zasm itself, runs 16 rows of a loop nest together so that LLVM vectorizes across rows: for an inner loop summing floats in order (spectralnorm) this is gcc's outer-loop vectorization, which no LLVM-based compiler does; for an inner loop whose length depends on the data (mandelbrot) finished rows are masked off, as hand-written SIMD does, and no compiler here does that. It also inlines small recursive functions into themselves (binarytrees), which gcc does and LLVM never does. Results stay bit-for-bit identical.
- The optimizer is not trusted. Every function it changes carries the facts its new code needs and goes through the checker again; it is kept only if the checker accepts it. The borrow rules make the rows' independence a local check: loads only through shared borrows, stores only to each row's own element. Fuzzers compare optimized and original programs in the interpreter and in native builds.
- One LLVM setting differs. LLVM hoists a vectorized loop's runtime alias checks into the enclosing loop; in Floyd-Warshall one row always overlaps, so the hoisted check always failed and the scalar fallback always ran. zasm turns that off (floyd 2.7x faster). It is a compiler option, not a language property: clang and rustc users can pass it too.
- Where every compiler does the same thing, zasm is at parity with Rust and C compiled by the same LLVM (hashtable, matmul, sieve, quicksort, psort, wordcount). What remains: gcc is faster on binarytrees (it inlines deeper, and the C builder has no bounds checks, which zasm keeps because its facts cannot express a tree's size) and on fannkuch; nbody is 1.10x behind Rust (LLVM's SLP vectorizer packs zasm's loop into narrower vectors, with the same instruction count).
- binarytrees_region is an idiom comparison, shown separately: trees built in
zasm regions against safe std-only Rust with a
Boxper node and C with amallocper node. - Four surprising results were flawed comparisons and are gone. Two zasm wins over Rust (wordcount, binarytrees) came from weaker Rust code. Two C wins came from the setup: binarytrees rebuilt an identical tree every iteration, so clang 23 built one and multiplied, and clang fused multiply-adds by default, a different rounding from Rust and zasm.