zasm performance
zasm is a typed SSA language whose checker proves memory safety and data-race freedom, so compiled code carries no runtime checks. Here it is measured against Rust and C on 13 programs, each written with the same algorithm and data layout in all three languages.
Release builds, program by program
zasm's run time divided by each other build's. Left of the parity line zasm is
faster. Log scale. Rust is rustc -C opt-level=3 -C target-cpu=native; C is
-O3 -march=native -ffp-contract=off.
Table: run times in seconds
Development builds
The fast-to-compile builds against each other: zasm's Cranelift backend,
rustc -C opt-level=0 and clang -O0. How many times longer the other
binary runs than zasm's; right of the parity line zasm is faster. Log scale.
Table: run times in seconds
Compile time
Whole build including linking, geometric mean over the 13 programs, in milliseconds.
Development builds
Release builds
Table: compile times in milliseconds
How to read this
- Like for like, zasm is at parity with Rust and clang. All three go through the same LLVM. zasm emits no bounds checks and passes LLVM what its checker proved (exclusive and read-only borrows, no-wrap arithmetic, non-null aligned pointers, slice-length ranges), but in these loops LLVM removes most of Rust's checks too, and C has none.
- gcc's lead comes from three optimizations LLVM does not make, so Rust and clang trail it just as much: on spectralnorm it vectorizes the outer loop, computing 8 rows at once with 512-bit divides while keeping each element's operation order; on binarytrees it clones the recursive builder for each constant depth; on floyd it uses AVX-512 masked stores where LLVM writes every element back. gcc is slower on wordcount and nbody.
- binarytrees_region is an idiom comparison, shown separately: trees built in
zasm regions against safe std-only Rust with a
Boxper node and C with amallocper node. With an index arena in every language (binarytrees) zasm is within 3% of Rust and 12% behind clang, whose builder has no bounds checks; zasm, like Rust, checks each store because its facts cannot express the size of a tree. - What remains. nbody (1.08 against Rust): the same instruction count, but
LLVM's SLP vectorizer packs zasm's loop into narrower vectors; nothing principled moved it.
fannkuch (1.03): LLVM turns zasm's proved-in-bounds rotation loop into a libc
memmovefor 1 to 10 elements, 13% of the run, which Rust's bounds checks happen to prevent; with that transformation off zasm runs at 0.99 of Rust. - Four surprising results were flawed comparisons and are gone. Two zasm wins over Rust (wordcount, binarytrees) came from weaker Rust code. Two C wins came from the setup: binarytrees rebuilt an identical tree every iteration, so clang 23 built one and multiplied (0.008 s against 0.12 s; each tree now starts at an offset that depends on the running total, in every language), and clang fused multiply-adds by default, a different rounding from Rust and zasm worth 7-8% on mandelbrot and nbody.