Overview
Ziren is an open-source zero-knowledge virtual machine (zkVM) for the MIPS32r2 instruction set architecture (ISA), developed by ZKM. A program compiled for MIPS32r2, for example from Rust or Go, runs inside Ziren, and Ziren produces a proof that the execution was correct. A verifier checks the proof far faster than it could replay the execution, either natively or on-chain after the proof is wrapped into a SNARK.
Ziren proves the complete user-mode integer instruction set of MIPS32r2, branch-delay slot included, and proves Ethereum mainnet blocks end to end in production. It is used by the Entangled Rollup protocol for native cross-chain asset circulation, with deployments including the GOAT Network Bitcoin L2 and the Metis Hybrid Rollup.
This documentation describes Ziren V2.0.
Architectural Workflow
-
Compilation. Guest source code (Rust, or Go) is compiled by the Ziren toolchain for the
mipsel-zkm-zkvm-elftarget into a MIPS32r2 ELF binary. The ELF image, loaded into memory, fixes the program the proof is about. -
Execution. The executor runs the ELF, either with an interpreter or with a just-in-time compiler, and cuts the run into shards. For each shard it records the events (instructions, memory accesses, syscalls and precompile calls) the prover needs.
-
Arithmetization. Every executed instruction becomes one row of the chip for its opcode family; there is no central CPU table. The branch-delay slot is carried in the machine state as a pair
(pc, next_pc). Chips exchange values over buses (program fetch, register and memory accesses, byte lookups, syscalls), which are checked by a lookup argument. -
Shard proving. Each shard is proved with one LogUp-GKR lookup argument for all of its buses and one zerocheck for all of its constraints, over the 31-bit KoalaBear field. A jagged polynomial commitment reduces the claims on the shard's columns, which have differing heights, to a single evaluation that is opened by one batched WHIR proof. Hashing uses Poseidon2.
-
Recursion and compression. A recursion tree composes the shard proofs: a leaf program verifies one shard proof, and compose programs merge contiguous ranges of proofs up to a single root. Every recursion program's verifying key must belong to an enumerated allowlist, committed to as a Merkle root (the vk root). The result is a compressed proof whose size does not grow with the execution.
-
SNARK wrapping. For on-chain verification the compressed proof is shrunk and wrapped into a Groth16 or PLONK proof over BN254.
-
Verification. Groth16 and PLONK proofs are verified by Solidity contracts or by the
zkm-verifiercrate; core and compressed proofs are verified natively by the SDK.
Design Choices
-
MIPS32r2 execution. The guest executes the MIPS32r2 integer instructions listed in MIPS ISA, including the branch-delay slot, with a minimal Linux ABI for runtimes such as Go (see Linux ABI).
-
Per-opcode chips. Instructions are proved by narrow per-family chips rather than one wide table, so a common instruction pays only for its own columns. An addition or bitwise operation costs 33 to 37 committed cells per row, of which 29 to 32 are the shared instruction frame.
-
Cross-shard memory consistency by multiset hashing. Accesses to memory that crosses shard boundaries are accumulated as points on an elliptic curve over a degree-7 extension of KoalaBear (see Memory Consistency Checking), rather than by Merkle hashing.
-
KoalaBear field. Arithmetic is over the prime \(2^{31} - 2^{24} + 1\), with a degree-4 extension for challenges.
-
GPU proving. A CUDA implementation of the same protocol generates traces and proves shards on the GPU; the host only executes the guest and coordinates.
-
Formal determinism. A Lean 4 statement that each core chip's outputs are determined by its inputs is extracted mechanically from the chip's constraints (
crates/fv). -
Foundations. Ziren builds on Plonky3 and adapts SP1's circuit builder, recursion compiler and precompiles for MIPS32.
Target Use Cases
Ziren enables verifiable computation for general programs, including:
-
Bitcoin L2. GOAT Network is a Bitcoin L2 built on Ziren and BitVM2 to improve the scalability and interoperability of Bitcoin. See Use Cases.
-
ZK-OP (hybrid rollups). Combines an optimistic rollup's cost with validity-proof verifiability, letting users choose between a fast, higher-cost withdrawal and a slow, lower-cost one.
-
Entangled Rollup. Entangles rollups for trustless cross-chain communication, with a universal L2 extension that addresses fragmented liquidity through proof-of-burn (for example, cross-chain asset transfers).
-
zkML verification. Verifies the result of a machine-learning computation without exposing the model or its inputs (for example, validating a diagnosis without revealing patient data).
Installation
Ziren is available for Linux and macOS.
Requirements
- Git
- Rust via
rustup. The Ziren repository pins its nightly toolchain inrust-toolchain.toml, whichrustupinstalls on first use. - Go 1.23 or later, only for generating Groth16 or PLONK proofs (the SNARK backend is built through a Go FFI).
Get Started
Option 1: Quick Install
The zkmup installer installs the Ziren guest toolchain, which compiles programs for the mipsel-zkm-zkvm-elf target. Run:
curl --proto '=https' --tlsv1.2 -sSf https://raw.githubusercontent.com/ProjectZKM/toolchain/refs/heads/main/setup.sh | sh
The script downloads zkmup to ~/.zkm-toolchain/bin, installs the latest toolchain release (zkmup install) and writes ~/.zkm-toolchain/env (zkmup setup). Load the environment in each shell that builds guest programs:
. ~/.zkm-toolchain/env
List the available toolchain releases, and install or select a specific one:
~/.zkm-toolchain/bin/zkmup list-available
~/.zkm-toolchain/bin/zkmup install -v 20260917
~/.zkm-toolchain/bin/zkmup setup -v 20260917
Use release 20260917 or later. It fixes an LLVM MIPS backend bug, present in earlier releases, in which the delay-slot filler could move a store across a call and miscompile some guest programs.
Guest programs are built by running cargo from PATH, which after loading ~/.zkm-toolchain/env is the Ziren toolchain's cargo; the environment file also sets ZIREN_ZKM_CC for C dependencies. If cargo is instead a rustup proxy, the build selects the rustup toolchain named by ZKM_GUEST_TOOLCHAIN (default zkm).
You can now run the Ziren examples or unit tests:
git clone https://github.com/ProjectZKM/Ziren
cd Ziren && cargo test -r
Troubleshooting
The following error may occur:
cargo build --release
cargo: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.32' not found (required by cargo)
cargo: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.33' not found (required by cargo)
cargo: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.34' not found (required by cargo)
The prebuilt toolchain binaries are built for Ubuntu 22.04 and macOS. On systems with an older GLIBC, build the toolchain from source.
Option 2: Building from Source
See the toolchain repository.
Choosing a Prover
The SDK's ProverClient::new() selects its prover from the ZKM_PROVER environment variable:
ZKM_PROVER | Prover |
|---|---|
local or cpu (default) | CPU prover in the current process |
cuda | GPU prover, reached over RPC at CUDA_ENDPOINT (default http://localhost:3000/twirp/) |
network | ZKM Prover Network (requires the SDK's network feature) |
mock | Mock prover for testing, which produces no real proof |
Groth16 and PLONK circuit artifacts are downloaded on first use into ~/.zkm/circuits/{groth16,plonk}/<version>, where <version> is the SDK's circuit version (v2.0.0 for Ziren 2.0.0).
Quickstart
Get started with Ziren by executing, generating and verifying a proof for your custom program.
Overview of all the steps to create your Ziren proof:
- Create a new project using the Ziren project template or CLI
- Compile and execute your guest program
- Generate a ZK proof of your program locally or via the proving network
- Verify the proof of your program, including on-chain verification
Creating a new project
After installing the Ziren toolchain, you can create a new project either directly via the CLI or by cloning the project template.
Using the CLI
Install the CLI locally from source:
#![allow(unused)] fn main() { cd Ziren/crates/cli cargo install --locked --force --path . }
You can now create a bare new project using the new command:
#![allow(unused)] fn main() { cargo ziren new --bare <NEW_PROJECT> }
To view additional CLI commands, including build for compiling a program and vkey for displaying the guest’s verification key hash:
#![allow(unused)] fn main() { cargo ziren --help }
Using the Project Template:
You can also create a new project by cloning the Ziren Project Template, which includes:
- Guest and host Rust programs for proving a Fibonacci sequence
- Solidity contracts for on-chain verification
- Sample inputs/outputs and test data
The template's host/Cargo.toml and guest/Cargo.toml take zkm-sdk, zkm-build and zkm-zkvm as git dependencies. To build against Ziren 2.0, point them at https://github.com/ProjectZKM/Ziren (the repository's former name, zkMIPS, still appears in older copies of the template).
The project directory has the following structure:
.
├── contracts
│ ├── lib
│ ├── script
│ │ ├── ZKMVerifierGroth16.s.sol
│ │ └── ZKMVerifierPlonk.s.sol
│ ├── src
│ │ ├── Fibonacci.sol
│ │ ├── IZKMVerifier.sol
│ │ ├── fixtures
│ │ │ ├── groth16-fixture.json
│ │ │ └── plonk-fixture.json
│ │ └── v1.0.0
│ │ ├── Groth16Verifier.sol
│ │ ├── PlonkVerifier.sol
│ │ ├── ZKMVerifierGroth16.sol
│ │ └── ZKMVerifierPlonk.sol
│ └── test
│ ├── Fibonacci.t.sol
│ ├── ZKMVerifierGroth16.t.sol
│ └── ZKMVerifierPlonk.t.sol
├── guest
│ ├── Cargo.toml
│ └── src
│ └── main.rs
├── host
│ ├── Cargo.toml
│ ├── bin
│ │ ├── evm.rs
│ │ └── vkey.rs
│ ├── build.rs
│ ├── src
│ │ └── main.rs
│ └── tool
│ ├── ca.key
│ ├── ca.pem
│ └── certgen.sh
There are three main directories in the project:
guest: contains the guest program executed inside the zkVM.
./guest/src/main.rs: guest program that contains your core program logic to be executed inside the zkVM.
host: contains the host program that controls the end-to-end process of compiling and executing the guest, generating a proof and outputting verifier artifacts.
./host/src/main.rs: host program that contains the host logic for your program../host/bin/evm.rsand./host/bin/vkey.rs: used to generate EVM-compatible verifier artifacts and print the verifying key../host/tool/: contains the certificates and scripts for network proving../host/build.rs: contains custom build logic including compiling the guest to an ELF artifact whenever the host is built.
contracts: contains the Solidity verifier smart contracts and test scripts for on-chain verification.
contracts/src/Fibonacci.sol: a sample Solidity contract demonstrating input/output structure for the Fibonacci program.IZKMVerifier.sol: implemented Ziren interface for verifiers.fixtures/: contains the public outputs, proof and verification keys in JSON format.v1.0.0/: contains the Groth16 and PLONK verifier implementations and wrapper contracts for circuit version v1.0.0. The verifier contracts must match the circuit version of the SDK that produced the proof. For Ziren 2.0 (circuit versionv2.0.0), the SDK downloads them with the circuit artifacts into~/.zkm/circuits/groth16/v2.0.0/(Groth16Verifier.sol,ZKMVerifierGroth16.sol) and~/.zkm/circuits/plonk/v2.0.0/(PlonkVerifier.sol,ZKMVerifierPlonk.sol); copy them intocontracts/src/v2.0.0/.contracts/script/: contains forge scripts to deploy the verifier contracts.contracts/test/: contains Foundry tests to validate verifier functionality.
./guest/src/main.rs and ./host/src/main.rs contain the guest and host implementation of the Fibonacci example. You can change the content of these programs to fit your intended use case.
Building the program
The guest program is compiled into a MIPS executable ELF through ./host/build.rs. The generated ELF file is stored in ./target/elf-compilation
Executing the program
You can execute the program and display the output without generating a proof to check its output:
cd host
cargo run --release -- --execute
Sample output for the Fibonacci program example in the template:
n: 20
Program executed successfully.
n: 20
a: 6765
b: 10946
Values are correct!
Number of cycles: 6755
Generating a proof for the program
The generated ELF binary will be used for proof generation. Generate a proof for your guest program with the following command:
#![allow(unused)] fn main() { cargo run --release -- --<PROOF_TYPE> // for core and compressed proofs cargo run --release --bin evm -- --system <PROOF_TYPE> // for EVM-compatible proofs }
Within the project template, there are three types of proofs that can be generated. The host requires exactly one of --execute, --core or --compressed:
- Core proof: one proof per shard.
#![allow(unused)] fn main() { cargo run --release -- --core }
- Compressed proof: the shard proofs composed into one proof whose size does not grow with the execution:
#![allow(unused)] fn main() { cargo run --release -- --compressed }
- EVM-compatible proof: the compressed proof wrapped into a PLONK or Groth16 proof over BN254.
For Groth16 proofs, recommended for on-chain verification:
#![allow(unused)] fn main() { cd host cargo run --release --bin evm -- --system groth16 }
The output includes public values, the verifying key hash, and proof bytes. An example output, abbreviated (the Groth16 constraint count, timings, key hash and proof depend on the circuit version and the program; the Ziren 2.0 Groth16 circuit has about 31.9 million constraints):
n: 20
Proof System: Groth16
Reading R1CS took ...
Reading proving key took ...
... DBG prover done acceleration=none backend=groth16 curve=bn254 nbConstraints=... took=...
Generating proof took ...
... DBG verifier done backend=groth16 curve=bn254 took=...
Verification Key: 0x009f93857b6ce5bea9e982a82efa735aa57c0af27165b85ad17f21f9d3aae01a
Public Values: 0x00000000000000000000000000000000000000000000000000000000000000140000000000000000000000000000000000000000000000000000000000001a6d0000000000000000000000000000000000000000000000000000000000002ac2
Proof Bytes: 0x00cd9fa*f08b496bc4ff57a76d69719938a872d4b95f2f638b5f21f2b0cef825606bc14032748e2a86a8dadb00de79f88c650fa2f83813e4f69661b15600d09e9f328b7332dddaeb8dc27519c906bea438929e5474fe76dc940ea6ec3b50ce898a54c339d1c1ece67e746335652d5afb0d950d19960ef34dc7885d99296cad3b1385f9c161ab54408afffb8143708737c89c54988b3512478b79affac241bd08ba4dc75da105764b63b6101fdf06cdbd136a371f932d769c2cd6ba4b699fa61a2bfb5ad5215b4b64f1388484291240faf87c09a04d8ee2a7c47d3f68c6509ab172fecfffd02a2edee4e7deaf93557e22bc99fd3d7f2bf1af15b40b7d09685e6b8d528a122
For PLONK proofs, which are larger than Groth16 proofs but need no circuit-specific trusted setup:
#![allow(unused)] fn main() { cargo run --release --bin evm -- --system plonk }
Proof fixtures will be saved in ./contracts/src/fixtures/ to be used for on-chain verification. These contain the public outputs, proof and verification keys in JSON format. An example of groth16-fixture.json :
{
"a": 6765,
"b": 10946,
"n": 20,
"vkey": "0x009f93857b6ce5bea9e982a82efa735aa57c0af27165b85ad17f21f9d3aae01a",
"publicValues": "0x00000000000000000000000000000000000000000000000000000000000000140000000000000000000000000000000000000000000000000000000000001a6d0000000000000000000000000000000000000000000000000000000000002ac2",
"proof": "0x00cd9faf08b496bc4ff57a76d69719938a872d4b95f2f638b5f21f2b0cef825606bc14032748e2a86a8dadb00de79f88c650fa2f83813e4f69661b15600d09e9f328b7332dddaeb8dc27519c906bea438929e5474fe76dc940ea6ec3b50ce898a54c339d1c1ece67e746335652d5afb0d950d19960ef34dc7885d99296cad3b1385f9c161ab54408afffb8143708737c89c54988b3512478b79affac241bd08ba4dc75da105764b63b6101fdf06cdbd136a371f932d769c2cd6ba4b699fa61a2bfb5ad5215b4b64f1388484291240faf87c09a04d8ee2a7c47d3f68c6509ab172fecfffd02a2edee4e7deaf93557e22bc99fd3d7f2bf1af15b40b7d09685e6b8d528a122"
}
Note: EVM-compatible proofs e.g., Groth16, are more computationally intensive to generate but are required for on-chain verification. See here for more detailed explanations on the types of proofs Ziren offers.
Local proving is the default (ZKM_PROVER=local, or unset). For heavier workloads, use a GPU prover (ZKM_PROVER=cuda) or ZKM's Prover Network. To enable network proving, follow the instructions here and update your .env file:
#![allow(unused)] fn main() { ZKM_PROVER=network ZKM_PRIVATE_KEY=<your_key> SSL_CERT_PATH=<path_to_cert> SSL_KEY_PATH=<path_to_key> }
Verifying a proof on-chain
Once you’ve generated a proof, you can compile the verifier contract. To compile and execute all Foundry Solidity test files in the contracts directory:
#![allow(unused)] fn main() { cd contracts forge test }
Foundry will detect and run all test contracts to verify that
- The proof fixtures in
./contracts/src/fixtures(either PLONK or Groth16) are correctly accepted or rejected. - The Solidity verifier contracts
ZKMVerifierGroth16.solandZKMVerifierPlonk.solbehave correctly. - The application contract
Fibonacci.solintegrates correctly with the verifiers.
An example output with all passing tests:
[⠊] Compiling...
[⠊] Compiling 32 files with Solc 0.8.28
[⠒] Solc 0.8.28 finished in 902.25ms
Compiler run successful!
Ran 2 tests for test/Fibonacci.t.sol:FibonacciGroth16Test
[PASS] testRevert_InvalidFibonacciProof() (gas: 28279)
[PASS] test_ValidFibonacciProof() (gas: 28825)
Suite result: ok. 2 passed; 0 failed; 0 skipped; finished in 1.37ms (617.07µs CPU time)
Ran 2 tests for test/Fibonacci.t.sol:FibonacciPlonkTest
[PASS] testRevert_InvalidFibonacciProof() (gas: 29569)
[PASS] test_ValidFibonacciProof() (gas: 29995)
Suite result: ok. 2 passed; 0 failed; 0 skipped; finished in 1.42ms (608.57µs CPU time)
Ran 2 tests for test/ZKMVerifierGroth16.t.sol:ZKMVerifierGroth16Test
[PASS] test_RevertVerifyProof_WhenGroth16() (gas: 209970)
[PASS] test_VerifyProof_WhenGroth16() (gas: 209948)
Suite result: ok. 2 passed; 0 failed; 0 skipped; finished in 9.24ms (15.99ms CPU time)
Ran 2 tests for test/ZKMVerifierPlonk.t.sol:ZKMVerifierPlonkTest
[PASS] test_RevertVerifyProof_WhenPlonk() (gas: 282131)
[PASS] test_VerifyProof_WhenPlonk() (gas: 282132)
Suite result: ok. 2 passed; 0 failed; 0 skipped; finished in 10.94ms (18.99ms CPU time)
Ran 4 test suites in 11.53ms (22.97ms CPU time): 8 tests passed, 0 failed, 0 skipped (8 total tests)
Once you’ve compiled the verifier contracts, you can deploy it to Sepolia or another EVM-compatible test network:
#![allow(unused)] fn main() { forge script script/ZKMVerifierGroth16.s.sol:ZKMVerifierGroth16Script \ --rpc-url <RPC_URL> \ --private-key <YOUR_PRIVATE_KEY> --broadcast }
<RPC_URL>: Replace with your RPC endpoint (e.g., from Alchemy or Infura)<YOUR_PRIVATE_KEY>: Replace with the private key of your wallet
The command executes the ZKMVerifierGroth16Script script, which deploys the ZKMVerifierGroth16 contract to the network.
To deploy to a different verifier, e.g., PLONK, replace the script name accordingly:
#![allow(unused)] fn main() { forge script script/ZKMVerifierPlonk.s.sol:ZKMVerifierPlonkScript \ --rpc-url <RPC_URL> \ --private-key <YOUR_PRIVATE_KEY> --broadcast }
The successful deployment output for a Groth16 proof on Sepolia:
Script ran successfully.
## Setting up 1 EVM.
==========================
Chain 11155111
Estimated gas price: 0.002525776 gwei
Estimated total gas used for script: 3185830
Estimated amount required: 0.00000804669295408 ETH
==========================
##### sepolia
✅ [Success] Hash: 0x2143e3239579833092460969bffe71b8f2b8cc8cc360a34f01203c22bc465abb
Contract Address: 0x750Ad1b02000F6cC9Bc4E1F2dE2a85534D681841
Block: 8664583
Paid: 0.000004480767952712 ETH (2450639 gas * 0.001828408 gwei)
✅ Sequence #1 on sepolia | Total Paid: 0.000004480767952712 ETH (2450639 gas * avg 0.001828408 gwei)
==========================
ONCHAIN EXECUTION COMPLETE & SUCCESSFUL.
Transactions saved to: /zkm-project-template/contracts/broadcast/ZKMVerifierGroth16.s.sol/11155111/run-latest.json
Sensitive values saved to: /zkm-project-template/contracts/cache/ZKMVerifierGroth16.s.sol/11155111/run-latest.json
Performance
Metrics
Three quantities describe a zkVM's performance:
- Instruction efficiency: the committed trace area (cells) and bus interactions one executed instruction costs. It is a property of the arithmetization, independent of the hardware.
- Proving throughput: guest cycles proved per second on named hardware, from reading the input to writing the compressed proof. It is a rate, not a clock frequency. One cycle is one executed MIPS instruction; a precompile call is one cycle and many rows of its chip.
- Proof size: the bytes of the compressed proof a verifier reads.
Proving cost follows from throughput and the price of the hardware: the cost of a proof is its proving time multiplied by the cost per second of the machine. ethproofs.org reports proving time, proof size and cost per Mgas for Ethereum mainnet blocks proved by each zkVM.
To reproduce measurements on your own hardware, use the zkvm-benchmarks suite.
Measurements
The figures below are from the Ziren V2.0 paper (docs/paper). The workload is Ethereum mainnet blocks executed by a MIPS32 build of the reth execution client. The hardware is NVIDIA RTX 5090 GPUs (32 GB) in a host with an AMD EPYC 9355 processor and 925 GB of memory.
The proving times and proof size were measured on revisions of the v2.0.0 branch that precede the release in two respects: the Poseidon2 permutation used 13 partial rounds instead of the released 20, and the lookup challenge was not ground. The released configuration therefore differs from these numbers by an unmeasured amount.
Instruction efficiency
Committed cells and bus interactions for one row of the most frequent chips. The frame (fetch, operand reads, register write and the (pc, next_pc) pair) is shared by every instruction chip:
| Chip | cells: frame | cells: own | cells: row | interactions: row |
|---|---|---|---|---|
AddSubImm | 29 | 4 | 33 | 16 |
AddSub | 32 | 4 | 36 | 20 |
Bitwise | 32 | 5 | 37 | 24 |
Lt | 32 | 18 | 50 | 23 |
ShiftLeftImm | 26 | 19 | 45 | 20 |
ShiftRight | 32 | 57 | 89 | 45 |
Branch | 29 | 27 | 56 | ≤ 25 |
LoadWord | 29 | 19 | 48 | 25 |
StoreWord | 29 | 23 | 52 | 25 |
LoadNarrow | 29 | 30 | 59 | 26 |
Mul | 32 | 42 | 74 | 42 |
DivRem | 32 | 131 | 163 | 64 |
Over a whole block the cost per instruction is higher, mainly because of the per-shard memory-argument rows. Block 25,955,640 (495.6 million cycles) commits 29.2 G cells and 12.8 G (row, interaction) pairs: 59.5 cells and 26.2 pairs per executed instruction. The cross-shard Global chip holds 15% of the area (8.9 cells per instruction), because every word a shard touches costs two of its rows. The figure depends on the workload: a block that uses more precompiles has more cells per cycle.
Proving throughput
One RTX 5090 proves a 288-million-cycle block at 5.9 MHz. Multi-GPU results, on warm wall-clock time of three consecutive proofs:
| Block (guest cycles) | 1 GPU (s) | 2 GPUs (s) | 4 GPUs (s) | 1 GPU (MHz) | 2 GPUs (MHz) | 4 GPUs (MHz) |
|---|---|---|---|---|---|---|
| 420 M | 65.9 | 34.3 | 21.0 | 6.4 | 12.2 | 20.0 |
| 530 M | 84.3 | 44.7 | 26.5 | 6.3 | 11.9 | 20.0 |
| 912 M | 124.3 | 65.0 | 38.0 | 7.3 | 14.0 | 24.0 |
Four GPUs reach 3.1 to 3.3 times the throughput of one. The gap to linear scaling is the serial recursion tail over the last shards and the start-up interval before every GPU has a shard.
On a GPU, the lookup argument (LogUp-GKR) takes 42% of kernel time and WHIR commitment and opening 19%, in a ten-shard profile. The lookup cost scales with (row, interaction) pairs, so memory instructions, which carry 25 to 26 interactions per row, account for 47% of all pairs.
Execution
On one core of an AMD EPYC 9355:
| Executor | Rate |
|---|---|
| JIT compiler, once compiled (496 M-cycle block) | 198 MHz |
| JIT compiler, including its compilation pass | 166 MHz |
| Interpreter, without the event record | 40.6 MHz |
| Interpreter, with the event record | 14.6 MHz |
With one GPU, execution does not limit proving. With several GPUs fed by one host, each GPU worker re-executes the shard it proves.
Proof size
The compressed proof is 603 KiB (617,618 bytes in an instrumented run) and does not grow with the length of the execution. Openings of the first WHIR oracle account for 79% of it. Groth16 and PLONK proofs wrapped from it are constant-size SNARKs for on-chain verification.
Use Cases: GOAT Network
GOAT Network is a Bitcoin L2, the first L2 to be implemented on ZKM’s Entangled Network. As a full-stack Bitcoin-based zkRollup, GOAT leverages Ziren (ZKM’s in-house zkVM) to enable real-time proving, trustless bridging, and Bitcoin-native settlement.
Proving Block Execution
Ziren is used to generate ZK proofs for individual blocks processed by GOAT’s sequencers. Each proof demonstrates that a block was executed correctly and that its resulting state transitions are valid. To amortize costs, proofs are generated over a specific period rather than for each transaction individually.
The proof contains the hash of the associated block, root hash of the sequencer set and new state root. These values are bundled and submitted to the Bitcoin L1. The block hash and the root hash serve as public inputs to the ZK circuit and are required to match the corresponding values that have already been posted on Bitcoin L1.
Real-Time Proving & Peg-Out Proofs
Ziren enables the generation of proofs in real time for every block on GOAT, which powers the BitVM2 bridge used for withdrawals during the peg-out process. This capability eliminates the delays that can otherwise occur during withdrawals, allowing funds to be released in sync with block production rather than requiring days (or even weeks) of waiting.
During the peg-out process, operators act as provers to generate ZK proofs for peg-out transactions. Challengers can contest invalid proofs within an optimistic challenge window.
The proof generation pipeline for peg-outs consists of several stages:
- Core prover: proves each shard of the block's execution.
- Compress prover: composes the shard proofs through the recursion tree into one compressed proof.
- SNARK prover: wraps the compressed proof into a Groth16 proof over BN254, a constant-size, EVM-compatible format used for verification on Bitcoin through BitVM2.
This architecture enables low-latency withdrawals and the rapid inclusion of new blocks. View proofs being generated in real-time for GOAT, including details for each proof in the pipeline here: https://bitvm2.testnet3.goat.network/proof.
Proof Verification via BitVM Paradigm
Once generated, proof commitments are submitted to a BitVM2-anchored covenant on Bitcoin L1 and verified through BitVM2’s fraud-proof challenge mechanism. Verification involves confirming that the public inputs of the ZK proof correspond to a valid transaction on Bitcoin L1 and that this transaction is finalized within a sufficiently long proof-of-work chain.
All ZK proofs are tied to the committed state; any attempt to manipulate the state would produce a different public input, which would cause verification to fail. Both the L2 state and its execution are fully auditable by any participant.
Transaction Lifecycle
The following is an overview of the transaction lifecycle on GOAT and Ziren’s role in it:
- A user first submits a transaction to the L2.
- The sequencer batches and executes transactions into an L2 block, which is then processed by Ziren.
- Ziren generates a ZK proof that attests to the correctness of the block’s execution and its state transitions.
- The ZK proof generated by Ziren, together with the state commitment, are submitted to a BitVM2-anchored covenant on the Bitcoin L1.
Entangled Rollup Use Case
ZKM’s primary use case is Entangled Network, using validity proofs to verify cross-chain messaging between L2s without implementing the architecture of a bridge. In this way, security is natively inherited by the underlying L1 e.g., Bitcoin in the case of GOAT, as opposed to a third-party bridge which has its own security tradeoffs such as custodial risk, use of wrapped assets, additional trust assumptions, and multisig control.
By treating the L2s as bridges, all L2s part of the Entangled Rollup Network can share state, liquidity and execution. Ziren is used to generate valid proofs of state transitions, proving the execution of blocks in the L2s to be verified on their underlying L2s.
Ziren underpins GOAT’s cross-chain interoperability through the Entangled Rollup Network:
- Rollup proofs double as bridge receipts: a single validity proof can verify execution on one L2 and unlock assets on another.
- Enables native asset transfers across incompatible chains (e.g., Bitcoin ↔ Ethereum L2s).
- Validity proofs ensure L1-final settlement across chains, preserving each chain’s native security.
Read more about how Entangled Network enables native assets and unified liquidity between (even incompatible) chains here.
Independent Evaluations
- Ziren (v1.1.4) underwent an independent evaluation by Prooflab.
Recommended use cases for Ziren as noted in the report are Bitcoin L2 implementations, hybrid (combining ZK and optimistic) rollups, zkML verification and cross-chain applications. View the full evaluation report for Ziren here.
- SoK: Understanding zkVM: From Research to Practice (Yang, Cheng, Tang, Yang, Zhang and Ren; Zhejiang University, University of Sussex and Singapore Management University; IACR ePrint 2026/525).
The survey decomposes zkVMs into an ISA layer, a VM layer and a proving layer and evaluates representative systems. It describes Ziren as taking a deliberate architectural bet on MIPS32, prioritizing instruction regularity and constraint uniformity over RISC-V ecosystem convenience (zkVM-level instruction efficiency).
- Efficient Branch-and-Bound Testing and Verification of zkVMs (Takahashi, Jana and Yang; Columbia University; arXiv 2609.15020).
The paper presents ZEBRA, a framework that checks zkVM constraint tables against their instruction semantics by interval-based branch-and-bound search, and evaluates it on five Plonky3-based zkVMs: Pico, SP1, Sphinx, Valida and Ziren, the only one targeting MIPS.
MIPS VM
Ziren is a verifiable computation infrastructure based on MIPS32, designed to generate zero-knowledge proofs for programs written in Rust (and Go). Ziren adopts the MIPS32r2 instruction set. The MIPS VM, one of the core components of Ziren, is the execution framework for MIPS32r2 instructions. Below we briefly introduce the advantages of MIPS32r2 over RV32IM and the execution flow of the MIPS VM.
Advantages of MIPS32r2 over RV32IM
1. MIPS32r2 is more consistent and offers more complex opcodes
- The J/JAL instructions support jump ranges of up to 256MiB, offering greater flexibility for large-scale data processing and complex control flow scenarios.
- MIPS32r2 has rich set of bit manipulation instructions and additional conditional move instructions (such as MOVZ and MOVN) that ensure precise data handling.
- MIPS32r2 has integer multiply-add/sub instructions, which can improve arithmetic computation efficiency.
- MIPS32r2 has SEH and SEB sign extension instructions, which make it very convenient to perform sign extension operations on char and short type data.
2. MIPS32r2 has a long-established ecosystem
- MIPS32r2 is a fixed, complete specification that has been in wide use for more than 20 years, so there are no optional extensions whose combinations the zkVM must track.
- MIPS is used by Optimism's fault-proof VM (Cannon).
3. The branch-delay slot is part of the proved semantics
- A MIPS32 branch or jump takes effect after the instruction in its delay slot. Ziren carries the machine state as the pair
(pc, next_pc), so a transfer sets the second component and the delay-slot instruction still executes. The instruction set is proved as specified, delay slot included.
Execution Flow of MIPS VM
The execution flow of MIPS VM is as follows:
Before execution, the developer's Rust or Go program is compiled by the Ziren toolchain for the mipsel-zkm-zkvm-elf target into a MIPS32r2 ELF binary.
The MIPS VM executes the ELF as follows:
- The ELF is loaded into a Program: all data is loaded into the memory image, and all code is decoded into the Instruction list. The program image is fixed by the verifying key.
- The executor runs the instructions from the ELF entry point until the program halts (the
HALTsyscall, orexit_groupfor Linux-ABI programs), updating the registers,HI/LOand memory at each step. It runs either as an interpreter or with a just-in-time compiler. The run is cut into shards: a shard closes when the committed trace area it would produce reaches a fixed limit, so the number of shards depends on the instruction mix, not only on the cycle count. - For each shard, the executor records the events the prover needs: executed instructions, memory accesses, syscalls and precompile calls. In a multi-GPU deployment the coordinator runs the program once and records only each shard's starting state and inputs, and each GPU worker re-executes its own shard.
After execution, the prover uses the execution record:
- Each opcode family has its own chip, and every executed instruction becomes one row of its chip; there is no central CPU table. Memory accesses, byte lookups, syscalls and precompiles have their own chips.
- The chips' traces are proved per shard, and the shard proofs are composed into one proof (see Prover Architecture).
Memory Layout for guest program
The memory layout for guest program is controlled by VM, runtime and toolchain.
Rust guest program
Two kinds of allocators are provided to rust guest program
- bump allocator: both normal memory and program I/O is allocated from the heap. And the heap address is always increased and cannot be reused.
| Section | Start | Size | Access | Controlled-by |
|---|---|---|---|---|
| registers | 0x00 | 36 | rw | VM |
| Stack | 0x7f000000 | (stack grows down) | rw | runtime |
| Code | ||||
| .text | .text size | ro | toolchain | |
| .rodata | .rodata size | ro | toolchain | |
| .eh_frame | .eh_frame size | ro | toolchain | |
| .bss | .bss size | ro | toolchain | |
| Heap (contains program I/O) | _end | 0x7f000000 - _end | rw | runtime |
- embedded allocator: Program I/O address space is reserved and split from heap address space. A TLS heap is used for heap management.
| Section | Start | Size | Access | Controlled-by |
|---|---|---|---|---|
| registers | 0x00 | 36 | rw | VM |
| Stack | 0x7f000000 | (stack grows down) | rw | runtime |
| Code | ||||
| .text | .text size | ro | toolchain | |
| .rodata | .rodata size | ro | toolchain | |
| .eh_frame | .eh_frame size | ro | toolchain | |
| .bss | .bss size | ro | toolchain | |
| Program I/O | 0x3f000000 | 0x40000000 | rw | runtime |
| Heap | _end | 0x3f000000 - _end | rw | runtime |
Go guest program
Go guest program is similar to embedded-mode rust guest program, except that the initial args is set by VM at the top of the stack. The memory layout is as follows:
| Section | Start | Size | Access | Controlled-by |
|---|---|---|---|---|
| registers | 0x00 | 36 | rw | VM |
| Stack | 0x7f000000 | (stack grows down) | rw | runtime |
| Initial args | 0x7effc000 | 0x4000 | ro | VM |
| Code | ||||
| .text | .text size | ro | toolchain | |
| .rodata | .rodata size | ro | toolchain | |
| .eh_frame | .eh_frame size | ro | toolchain | |
| .bss | .bss size | ro | toolchain | |
| Program I/O | 0x3f000000 | 0x40000000 | rw | runtime |
| Heap | _end | 0x3f000000 - _end | rw | runtime |
MIPS ISA
The Opcode enum organizes MIPS instructions into several functional categories, each serving a specific role in the instruction set:
#![allow(unused)] fn main() { pub enum Opcode { // ALU ADD = 0, // ADDSUB SUB = 1, // ADDSUB MUL = 2, // MUL MULT = 3, // MUL MULTU = 4, // MUL DIV = 5, // DIVREM DIVU = 6, // DIVREM MOD = 7, // DIVREM MODU = 8, // DIVREM SLL = 9, // SLL SRL = 10, // SR SRA = 11, // SR ROR = 12, // SR SLT = 13, // LT SLTU = 14, // LT AND = 15, // BITWISE OR = 16, // BITWISE XOR = 17, // BITWISE NOR = 18, // BITWISE CLZ = 19, // CLO_CLZ CLO = 20, // CLO_CLZ // Control FLow BEQ = 21, // BRANCH BGEZ = 22, // BRANCH BGTZ = 23, // BRANCH BLEZ = 24, // BRANCH BLTZ = 25, // BRANCH BNE = 26, // BRANCH Jump = 27, // JUMP Jumpi = 28, // JUMP JumpDirect = 29, // JUMP SYSCALL = 30, // SYSCALL // Memory Op LB = 31, // LOAD LBU = 32, // LOAD LH = 33, // LOAD LHU = 34, // LOAD LW = 35, // LOAD LWL = 36, // LOAD LWR = 37, // LOAD LL = 38, // LOAD SB = 39, // STORE SH = 40, // STORE SW = 41, // STORE SWL = 42, // STORE SWR = 43, // STORE SC = 44, // STORE // Misc INS = 45, // INS MADDU = 46, // MADDSUB MSUBU = 47, // MADDSUB MADD = 48, // MADDSUB MSUB = 49, // MADDSUB MEQ = 50, // MOVCOND MNE = 51, // MOVCOND WSBH = 52, // WSBH EXT = 53, // EXT TEQ = 54, // TEQ SEXT = 55, // SEXT // Syscall UNIMPL = 0xff, } }
All MIPS instructions can be divided into the following taxonomies:
ALU Operators
This category includes the fundamental arithmetic logical operations and count operations. It covers addition (ADD) and subtraction (SUB), several multiplication, division and remainder variants (MULT, MULTU, MUL, DIV, DIVU, MOD, MODU), as well as bit shifting and rotation operations (SLL, SRL, SRA, ROR), comparison operations like set less than (SLT, SLTU) a range of bitwise logical operations (AND, OR, XOR, NOR) and count operations like CLZ counts the number of leading zeros, while CLO counts the number of leading ones. These operations are useful in bit-level data analysis.
Memory Operations
This category is dedicated to moving data between memory and registers. It contains a comprehensive set of load instructions, such as LH (load halfword), LWL (load word left), LW (load word), LB (load byte), LBU (load byte unsigned), LHU (load halfword unsigned), LWR (load word right), and LL (load linked), as well as corresponding store instructions like SB (store byte), SH (store halfword), SWL (store word left), SW (store word), SWR (store word right), and SC (store conditional). These operations ensure that data is correctly and efficiently read from or written to memory.
Branching Instructions
Instructions BEQ (branch if equal), BGEZ (branch if greater than or equal to zero), BGTZ (branch if greater than zero), BLEZ (branch if less than or equal to zero), BLTZ (branch if less than zero), and BNE (branch if not equal) are used to change the flow of execution based on comparisons. These instructions are vital for implementing loops, conditionals, and other control structures.
Jump Instructions
Jump-related instructions, including Jump, Jumpi, and JumpDirect, are responsible for altering the execution flow by redirecting it to different parts of the program. They are used for implementing function calls, loops, and other control structures that require non-sequential execution, ensuring that the program can navigate its code dynamically.
Syscall Instructions
SYSCALL triggers a system call, allowing the program to request services from the zkvm operating system. The service can be a precompile computation, such as do sha extend operation by SHA_EXTEND precompile. It can also be an input/output operation such as SYSHINTREAD and WRITE, or a Linux system call (see Linux ABI).
Misc Instructions
This category includes other instructions. TEQ is typically used to test equality conditions between registers. MADDU/MSUBU is used for multiply accumulation. SEB/SEH (executor opcode SEXT) sign-extend a byte or halfword. EXT/INS extract and insert bit fields. WSBH swaps the bytes within each halfword. MOVZ/MOVN (executor opcodes MEQ/MNE) are conditional moves.
Supported instructions
The support instructions are as follows:
| instruction | Op [31:26] | rs [25:21] | rt [20:16] | rd [15:11] | shamt [10:6] | func [5:0] | function | chip |
|---|---|---|---|---|---|---|---|---|
| ADD | 000000 | rs | rt | rd | 00000 | 100000 | rd = rs + rt | AddSub |
| ADDI | 001000 | rs | rt | imm | imm | imm | rt = rs + sext(imm) | AddSubImm |
| ADDIU | 001001 | rs | rt | imm | imm | imm | rt = rs + sext(imm) | AddSubImm |
| ADDU | 000000 | rs | rt | rd | 00000 | 100001 | rd = rs + rt | AddSub |
| AND | 000000 | rs | rt | rd | 00000 | 100100 | rd = rs & rt | Bitwise |
| ANDI | 001100 | rs | rt | imm | imm | imm | rt = rs & zext(imm) | BitwiseImm |
| BEQ | 000100 | rs | rt | offset | offset | offset | PC = PC + sext(offset<<2), if rs == rt | Branch |
| BGEZ | 000001 | rs | 00001 | offset | offset | offset | PC = PC + sext(offset<<2), if rs >= 0 | Branch |
| BGTZ | 000111 | rs | 00000 | offset | offset | offset | PC = PC + sext(offset<<2), if rs > 0 | Branch |
| BLEZ | 000110 | rs | 00000 | offset | offset | offset | PC = PC + sext(offset<<2), if rs <= 0 | Branch |
| BLTZ | 000001 | rs | 00000 | offset | offset | offset | PC = PC + sext(offset<<2), if rs < 0 | Branch |
| BNE | 000101 | rs | rt | offset | offset | offset | PC = PC + sext(offset<<2), if rs != rt | Branch |
| CLO | 011100 | rs | rt | rd | 00000 | 100001 | rd = count_leading_ones(rs) | CloClz |
| CLZ | 011100 | rs | rt | rd | 00000 | 100000 | rd = count_leading_zeros(rs) | CloClz |
| DIV | 000000 | rs | rt | 00000 | 00000 | 011010 | (hi, lo) = (rs % rt, rs / rt), signed; division by zero is rejected | DivRem |
| DIVU | 000000 | rs | rt | 00000 | 00000 | 011011 | (hi, lo) = (rs % rt, rs / rt), unsigned; division by zero is rejected | DivRem |
| MOD | 000000 | rs | rt | rd | 00011 | 011010 | rd = rs % rt, signed (MIPS32 Release 6 encoding) | DivRem |
| MODU | 000000 | rs | rt | rd | 00011 | 011011 | rd = rs % rt, unsigned (MIPS32 Release 6 encoding) | DivRem |
| J | 000010 | instr_index | instr_index | instr_index | instr_index | instr_index | PC = instr_index || 00 (the PC[31..28] region bits are taken as 0) | Jump |
| JAL | 000011 | instr_index | instr_index | instr_index | instr_index | instr_index | r31 = PC + 8, PC = instr_index || 00 (the PC[31..28] region bits are taken as 0) | Jump |
| JALR | 000000 | rs | 00000 | rd | hint | 001001 | rd = PC + 8, PC = rs | Jump |
| JR | 000000 | rs | 00000 | 00000 | hint | 001000 | PC = rs | Jump |
| LB | 100000 | base | rt | offset | offset | offset | rt = sext(mem_byte(base + offset)) | LoadNarrow |
| LBU | 100100 | base | rt | offset | offset | offset | rt = zext(mem_byte(base + offset)) | LoadNarrow |
| LH | 100001 | base | rt | offset | offset | offset | rt = sext(mem_halfword(base + offset)) | LoadNarrow |
| LHU | 100101 | base | rt | offset | offset | offset | rt = zext(mem_halfword(base + offset)) | LoadNarrow |
| LL | 110000 | base | rt | offset | offset | offset | rt = mem_word(base + offset) | LoadWord |
| LUI | 001111 | 00000 | rt | imm | imm | imm | rt = imm<<16 | AddSubImm |
| LW | 100011 | base | rt | offset | offset | offset | rt = mem_word(base + offset) | LoadWord |
| LWL | 100010 | base | rt | offset | offset | offset | rt = rt merge most significant part of mem(base+offset) | MemoryUnaligned |
| LWR | 100110 | base | rt | offset | offset | offset | rt = rt merge least significant part of mem(base+offset) | MemoryUnaligned |
| MFHI | 000000 | 00000 | 00000 | rd | 00000 | 010000 | rd = hi | AddSubImm |
| MFLO | 000000 | 00000 | 00000 | rd | 00000 | 010010 | rd = lo | AddSubImm |
| MOVN | 000000 | rs | rt | rd | 00000 | 001011 | rd = rs, if rt != 0 (executor opcode MNE) | MovCond |
| MOVZ | 000000 | rs | rt | rd | 00000 | 001010 | rd = rs, if rt == 0 (executor opcode MEQ) | MovCond |
| MTHI | 000000 | rs | 00000 | 00000 | 00000 | 010001 | hi = rs | AddSubImm |
| MTLO | 000000 | rs | 00000 | 00000 | 00000 | 010011 | lo = rs | AddSubImm |
| MUL | 011100 | rs | rt | rd | 00000 | 000010 | rd = rs * rt | Mul |
| MULT | 000000 | rs | rt | 00000 | 00000 | 011000 | (hi, lo) = rs * rt | Mul |
| MULTU | 000000 | rs | rt | 00000 | 00000 | 011001 | (hi, lo) = rs * rt | Mul |
| NOR | 000000 | rs | rt | rd | 00000 | 100111 | rd = !(rs | rt) | Bitwise |
| OR | 000000 | rs | rt | rd | 00000 | 100101 | rd = rs | rt | Bitwise |
| ORI | 001101 | rs | rt | imm | imm | imm | rt = rs | zext(imm) | BitwiseImm |
| SB | 101000 | base | rt | offset | offset | offset | mem_byte(base + offset) = rt | StoreNarrow |
| SC | 111000 | base | rt | offset | offset | offset | mem_word(base + offset) = rt, rt = 1, if atomic update, else rt = 0 | StoreWord |
| SH | 101001 | base | rt | offset | offset | offset | mem_halfword(base + offset) = rt | StoreNarrow |
| SLL | 000000 | 00000 | rt | rd | sa | 000000 | rd = rt<<sa | ShiftLeftImm |
| SLLV | 000000 | rs | rt | rd | 00000 | 000100 | rd = rt << rs[4:0] | ShiftLeft |
| SLT | 000000 | rs | rt | rd | 00000 | 101010 | rd = rs < rt | Lt |
| SLTI | 001010 | rs | rt | imm | imm | imm | rt = rs < sext(imm) | LtImm |
| SLTIU | 001011 | rs | rt | imm | imm | imm | rt = rs < sext(imm) | LtImm |
| SLTU | 000000 | rs | rt | rd | 00000 | 101011 | rd = rs < rt | Lt |
| SRA | 000000 | 00000 | rt | rd | sa | 000011 | rd = rt >> sa | ShiftRightImm |
| SRAV | 000000 | rs | rt | rd | 00000 | 000111 | rd = rt >> rs[4:0] | ShiftRight |
| SYNC | 000000 | 00000 | 00000 | 00000 | stype | 001111 | sync (nop) | AddSubImm |
| SRL | 000000 | 00000 | rt | rd | sa | 000010 | rd = rt >> sa | ShiftRightImm |
| SRLV | 000000 | rs | rt | rd | 00000 | 000110 | rd = rt >> rs[4:0] | ShiftRight |
| SUB | 000000 | rs | rt | rd | 00000 | 100010 | rd = rs - rt | AddSub |
| SUBU | 000000 | rs | rt | rd | 00000 | 100011 | rd = rs - rt | AddSub |
| SW | 101011 | base | rt | offset | offset | offset | mem_word(base + offset) = rt | StoreWord |
| SWL | 101010 | base | rt | offset | offset | offset | store most significant part of rt | MemoryUnaligned |
| SWR | 101110 | base | rt | offset | offset | offset | store least significant part of rt | MemoryUnaligned |
| SYSCALL | 000000 | code | code | code | code | 001100 | syscall | SyscallInstrs |
| XOR | 000000 | rs | rt | rd | 00000 | 100110 | rd = rs ^ rt | Bitwise |
| XORI | 001110 | rs | rt | imm | imm | imm | rt = rs ^ zext(imm) | BitwiseImm |
| BAL | 000001 | 00000 | 10001 | offset | offset | offset | RA = PC + 8, PC = PC + sign_extend(offset || 00) | Jump |
| SYNCI | 000001 | base | 11111 | offset | offset | offset | sync (nop) | AddSubImm |
| PREF | 110011 | base | hint | offset | offset | offset | prefetch(nop) | AddSubImm |
| TEQ | 000000 | rs | rt | code | code | 110100 | trap if rs == rt (the execution is rejected) | MiscInstrs |
| ROTR | 000000 | 00001 | rt | rd | sa | 000010 | rd = rotate_right(rt, sa) | ShiftRightImm |
| ROTRV | 000000 | rs | rt | rd | 00001 | 000110 | rd = rotate_right(rt, rs[4:0]) | ShiftRight |
| WSBH | 011111 | 00000 | rt | rd | 00010 | 100000 | rd = swaphalf(rt) | MovCond |
| EXT | 011111 | rs | rt | msbd | lsb | 000000 | rt = rs[msbd+lsb..lsb] | MiscInstrs |
| SEH | 011111 | 00000 | rt | rd | 11000 | 100000 | rd = signExtend(rt[15..0]) | MiscInstrs |
| SEB | 011111 | 00000 | rt | rd | 10000 | 100000 | rd = signExtend(rt[7..0]) | MiscInstrs |
| INS | 011111 | rs | rt | msb | lsb | 000100 | rt = rt[32:msb+1] || rs[msb+1-lsb : 0] || rt[lsb-1:0] | MiscInstrs |
| MADDU | 011100 | rs | rt | 00000 | 00000 | 000001 | (hi, lo) = rs * rt + (hi,lo) | MiscInstrs |
| MADD | 011100 | rs | rt | 00000 | 00000 | 000000 | (hi, lo) = (hi,lo) + rs * rt (signed) | MiscInstrs |
| MSUBU | 011100 | rs | rt | 00000 | 00000 | 000101 | (hi, lo) = (hi,lo) - rs * rt | MiscInstrs |
| MSUB | 011100 | rs | rt | 00000 | 00000 | 000100 | (hi, lo) = (hi,lo) - rs * rt (signed) | MiscInstrs |
Supported syscalls
| syscall number | function |
|---|---|
| SYSHINTLEN = 0x00_00_00F0, | Return length of current input data. |
| SYSHINTREAD = 0x00_00_00F1, | Read current input data. |
| SYSVERIFY = 0x00_00_00F2, | Verify pre-compile program. |
| HALT = 0x00_00_0000, | Halts the program. |
| WRITE = 0x00_00_0002, | Write to the output buffer. |
| ENTER_UNCONSTRAINED = 0x00_00_0003, | Enter unconstrained block. |
| EXIT_UNCONSTRAINED = 0x00_00_0004, | Exit unconstrained block. |
| SHA_EXTEND = 0x30_01_0005, | Executes the SHA_EXTEND precompile. |
| SHA_COMPRESS = 0x01_01_0006, | Executes the SHA_COMPRESS precompile. |
| ED_ADD = 0x01_01_0007, | Executes the ED_ADD precompile. |
| ED_DECOMPRESS = 0x00_01_0008, | Executes the ED_DECOMPRESS precompile. |
| KECCAK_SPONGE = 0x01_01_0009, | Executes the KECCAK_SPONGE precompile. |
| SECP256K1_ADD = 0x01_01_000A, | Executes the SECP256K1_ADD precompile. |
| SECP256K1_DOUBLE = 0x00_01_000B, | Executes the SECP256K1_DOUBLE precompile. |
| SECP256K1_DECOMPRESS = 0x00_01_000C, | Executes the SECP256K1_DECOMPRESS precompile. |
| BN254_ADD = 0x01_01_000E, | Executes the BN254_ADD precompile. |
| BN254_DOUBLE = 0x00_01_000F, | Executes the BN254_DOUBLE precompile. |
| COMMIT = 0x00_00_0010, | Executes the COMMIT precompile. |
| COMMIT_DEFERRED_PROOFS = 0x00_00_001A, | Executes the COMMIT_DEFERRED_PROOFS precompile. |
| VERIFY_ZKM_PROOF = 0x00_00_001B, | Executes the VERIFY_ZKM_PROOF precompile. |
| BLS12381_DECOMPRESS = 0x00_01_001C, | Executes the BLS12381_DECOMPRESS precompile. |
| UINT256_MUL = 0x01_01_001D, | Executes the UINT256_MUL precompile. |
| U256XU2048_MUL = 0x01_01_002F, | Executes the U256XU2048_MUL precompile. |
| BLS12381_ADD = 0x01_01_001E, | Executes the BLS12381_ADD precompile. |
| BLS12381_DOUBLE = 0x00_01_001F, | Executes the BLS12381_DOUBLE precompile. |
| BLS12381_FP_ADD = 0x01_01_0020, | Executes the BLS12381_FP_ADD precompile. |
| BLS12381_FP_SUB = 0x01_01_0021, | Executes the BLS12381_FP_SUB precompile. |
| BLS12381_FP_MUL = 0x01_01_0022, | Executes the BLS12381_FP_MUL precompile. |
| BLS12381_FP2_ADD = 0x01_01_0023, | Executes the BLS12381_FP2_ADD precompile. |
| BLS12381_FP2_SUB = 0x01_01_0024, | Executes the BLS12381_FP2_SUB precompile. |
| BLS12381_FP2_MUL = 0x01_01_0025, | Executes the BLS12381_FP2_MUL precompile. |
| BN254_FP_ADD = 0x01_01_0026, | Executes the BN254_FP_ADD precompile. |
| BN254_FP_SUB = 0x01_01_0027, | Executes the BN254_FP_SUB precompile. |
| BN254_FP_MUL = 0x01_01_0028, | Executes the BN254_FP_MUL precompile. |
| BN254_FP2_ADD = 0x01_01_0029, | Executes the BN254_FP2_ADD precompile. |
| BN254_FP2_SUB = 0x01_01_002A, | Executes the BN254_FP2_SUB precompile. |
| BN254_FP2_MUL = 0x01_01_002B, | Executes the BN254_FP2_MUL precompile. |
| SECP256R1_ADD = 0x01_01_002C, | Executes the SECP256R1_ADD precompile. |
| SECP256R1_DOUBLE = 0x00_01_002D, | Executes the SECP256R1_DOUBLE precompile. |
| SECP256R1_DECOMPRESS = 0x00_01_002E, | Executes the SECP256R1_DECOMPRESS precompile. |
| POSEIDON2_PERMUTE = 0x00_01_0030, | Executes the POSEIDON2_PERMUTE precompile. |
| SYS_MMAP = 4210, | Linux mmap: allocate memory from the heap. |
| SYS_MMAP2 = 4090, | Linux mmap2: same as mmap. |
| SYS_BRK = 4045, | Linux brk: return the program break. |
| SYS_CLONE = 4120, | Linux clone: simulated, returns 1. |
| SYS_EXT_GROUP = 4246, | Linux exit_group: halt with an exit code. |
| SYS_READ = 4003, | Linux read: only stdin, returns 0 bytes. |
| SYS_WRITE = 4004, | Linux write: stdout, stderr, or the public-values (3) and hint (4) descriptors. |
| SYS_FCNTL = 4055, | Linux fcntl: F_GETFD and F_GETFL on fds 0-2. |
Linux syscalls not listed above but handled as no-ops (open, close, munmap, rt_sigaction, uname, futex_time64, prctl and others) are listed in Linux ABI. All Linux syscalls are proved by one chip, SysLinux; in the proof they are grouped under the code SYS_LINUX = 4000, which is not itself a syscall. Any other syscall number is rejected by the executor (UnsupportedSyscall).
Linux ABI Support
This document describes how Ziren supports the Linux ABI inside its MIPS zkVM, covering execution, proving, and cross-shard verification.
Overview
Ziren runs MIPS guest programs compiled against a Linux userspace ABI. The guest issues SYSCALL instructions just like a real MIPS/Linux process. The zkVM intercepts these and either:
- Executes the syscall in the host executor (producing a concrete result), then
- Proves that the result is correct via AIR constraints in the
SysLinuxChip.
The guest never touches real kernel code. The zkVM emulates a minimal Linux kernel that supports memory management, basic I/O, and process lifecycle, enough to run programs compiled with standard C/Go/Rust toolchains targeting MIPS.
Architecture
Guest Program (MIPS binary)
|
SYSCALL (V0 = syscall number)
|
v
+-------------------------------+
| Executor |
| |
| execute_operation() |
| -> SyscallCode::from_u32() |
| -> get_syscall() |
| -> handler.execute() |
| -> emit_syscall_event() |
+-------------------------------+
| |
LinuxEvent LinuxEvent
| |
v v
+--------------------+ +------------------------+
| Core Shard | | Precompile Shard |
| | | |
| SyscallInstrsChip | | SyscallChip(Precompile)|
| send_syscall() | | receive_syscall() |
| | | | | |
| SyscallChip(Core) | | SysLinuxChip |
| receive_syscall() | | |
+--------|-----------+ +---------|------------- +
| |
| global lookup message: |
| [shard, clk, |
| syscall_id, |
| arg1, arg2, |
| result_lo, result_hi] |
| |
v v
+--------------------------------------+
| GlobalChip |
| Verify send/receive multiplicities |
| Ensure result consistency |
+--------------------------------------+
Register Convention (MIPS ABI)
All registers are 32-bit (u32). In the AIR, each is represented as Word<T> = 4 x u8 (little-endian, each byte range-checked to [0, 255]).
| Register | Width | Role |
|---|---|---|
V0 | 32-bit | Syscall number (input) / return value (output) |
A0 | 32-bit | First argument |
A1 | 32-bit | Second argument |
A2 | 32-bit | Third argument (read via memory when needed) |
A3 | 32-bit | Error code output (0 = success, 9 = EBADF) |
Supported Linux Syscalls
SYS_MMAP (4210) / SYS_MMAP2 (4090): Memory Mapping
Used by the guest allocator to request memory pages.
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | Requested address. 0 = allocate from heap. |
a1 | 32-bit | Size in bytes. Rounded up to page boundary. |
return v0 | 32-bit | Allocated address (heap pointer when a0 == 0, or a0 itself). |
output A3 | 32-bit | Always 0x00000000. |
Execution logic:
a0 == 0: returns current heap pointer, increments heap by page-aligned size.a0 != 0: returnsa0(reuse existing mapping).
Page alignment in the AIR:
The 32-bit size a1 is decomposed at the byte level to separate the 12-bit page offset from the 20-bit page-aligned upper address:
a1 = [byte0 : 8-bit] [byte1 : 8-bit] [byte2 : 8-bit] [byte3 : 8-bit]
├── lo nibble: 4-bit (boolean bit decomposition)
└── hi nibble: 4-bit (boolean bit decomposition)
page_offset = byte0 + lo_nibble * 256 (12-bit, range [0, 4095])
upper_address = hi_nibble * 4096 + byte2 * 65536 + byte3 * 16777216
The mmap_size is constrained byte-by-byte (not via field reduce()) to avoid KoalaBear prime collisions:
mmap_size[0] = 0
mmap_size[1] = hi_nibble * 16 + 16 * not_aligned - carry[0] * 256
mmap_size[2] = a1[2] + carry[0] - carry[1] * 256
mmap_size[3] = a1[3] + carry[1]
Where not_aligned = 1 when page_offset != 0 (round up to next page). The carry bits handle byte overflow from the +0x1000 addition. Page alignment (low 12 bits = 0) is structural: byte0 = 0 and every term in byte1 is a multiple of 16.
Heap update uses bytewise AddOperation (4 x u8 + 3 carry bits):
new_heap = old_heap + mmap_size.
SYS_BRK (4045): Program Break
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | New break address. |
a1 | 32-bit | Unused. |
return v0 | 32-bit | max(a0, current_brk). |
output A3 | 32-bit | Always 0x00000000. |
AIR uses GtColsBytes (bytewise greater-than with complementary LTU lookups) to compare a0 against the BRK register.
SYS_CLONE (4120): Process Clone (Simulated)
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | Clone flags (ignored). |
a1 | 32-bit | Child stack (ignored). |
return v0 | 32-bit | Always 0x00000001. |
output A3 | 32-bit | Always 0x00000000. |
Threading is not implemented. The syscall always returns 1 (simulated parent PID).
SYS_EXT_GROUP (4246): exit_group, Terminate Execution
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | Exit code. |
a1 | 32-bit | Unused. |
return v0 | 32-bit | Always 0x00000000. |
output A3 | 32-bit | Always 0x00000000. |
Sets next_pc = 0 and records the exit code. Equivalent to HALT: the executor reports a nonzero exit code as an execution error (HaltWithNonZeroExitCode).
SYS_READ (4003): Read from File Descriptor
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | File descriptor. Only 0 (stdin) is valid. |
a1 | 32-bit | Buffer address. |
return v0 | 32-bit | 0 (end of file) for stdin, or 0xFFFFFFFF on error. |
output A3 | 32-bit | 0x00000000 on success, 0x00000009 (EBADF) on invalid fd. |
Only stdin (fd 0) is accepted, and it reads no bytes: guest input is provided through the hint syscalls (SYSHINTLEN, SYSHINTREAD), not through read. All other fds return -1 with EBADF.
SYS_WRITE (4004): Write to File Descriptor
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | File descriptor. |
a1 | 32-bit | Buffer address. |
A2 (implicit) | 32-bit | Byte count (read from A2 register via memory). |
return v0 | 32-bit | Bytes written (= A2 value). |
output A3 | 32-bit | Always 0x00000000. |
AIR constrains inorout.value == inorout.prev_value (read-only guard on A2 memory access).
SYS_FCNTL (4055): File Control
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | File descriptor (0/1/2 valid). |
a1 | 32-bit | Command: 1 = F_GETFD, 3 = F_GETFL. |
return v0 | 32-bit | Flags/fd value, or 0xFFFFFFFF on error. |
output A3 | 32-bit | 0x00000000 on success, 0x00000009 on error. |
Full case matrix:
cmd (a1) | fd (a0) | result (v0) | error (A3) |
|---|---|---|---|
| 1 (F_GETFD) | 0 | 0x00000000 | 0x00000000 |
| 1 (F_GETFD) | 1 | 0x00000001 | 0x00000000 |
| 1 (F_GETFD) | 2 | 0x00000002 | 0x00000000 |
| 1 (F_GETFD) | other | 0xFFFFFFFF | 0x00000009 |
| 3 (F_GETFL) | 0 | 0x00000000 (O_RDONLY) | 0x00000000 |
| 3 (F_GETFL) | 1 | 0x00000001 (O_WRONLY) | 0x00000000 |
| 3 (F_GETFL) | 2 | 0x00000001 (O_WRONLY) | 0x00000000 |
| 3 (F_GETFL) | other | 0xFFFFFFFF | 0x00000009 |
| other | any | 0xFFFFFFFF | 0x00000009 |
AIR uses bidirectional IsZeroOperation decoders on a0 (3 decoders) and a1 (2 decoders) with exhaustive branch constraints.
NOP Syscalls: No Operation
| Arg | Width | Semantics |
|---|---|---|
a0 | 32-bit | Ignored. |
a1 | 32-bit | Ignored. |
return v0 | 32-bit | Always 0x00000000. |
output A3 | 32-bit | Always 0x00000000. |
| Syscall | Number |
|---|---|
| SYS_OPEN | 4005 |
| SYS_CLOSE | 4006 |
| SYS_MUNMAP | 4091 |
| SYS_NANOSLEEP | 4166 |
| SYS_RT_SIGACTION | 4194 |
| SYS_RT_SIGPROCMASK | 4195 |
| SYS_SIGALTSTACK | 4206 |
| SYS_FSTAT64 | 4215 |
| SYS_MADVISE | 4218 |
| SYS_GETTID | 4222 |
| SYS_SCHED_GETAFFINITY | 4240 |
| SYS_CLOCK_GETTIME | 4263 |
| SYS_OPENAT | 4288 |
| SYS_PRLIMIT64 | 4338 |
| SYS_UNAME | 4122 |
| SYS_PRCTL | 4192 |
| SYS_FUTEX_TIME64 | 4422 |
The last three are called by newer Go runtimes. From Go 1.25 the runtime calls prctl to name memory regions and threads, and ignores the result. From Go 1.27 it calls uname at startup to decide whether futex_time64 exists; the no-op returns success with a zeroed utsname, which the runtime cannot parse, so the runtime probes futex_time64, receives 0, and uses the 64-bit time path.
The executor rejects any syscall number not listed on this page or in MIPS ISA with UnsupportedSyscall. In the SysLinuxChip constraints, a Linux syscall ID that is not one of the handled calls (MMAP, MMAP2, BRK, CLONE, EXIT_GROUP, READ, WRITE, FCNTL) is treated as a no-op, so every ID the executor accepts as a no-op is proved by the NOP branch. The circuit is wider than the executor here: a Linux ID the executor rejects would also satisfy the NOP branch (result 0, A3 0, no memory access), and SyscallInstrsChip accepts a non-Linux ID it does not recognise with v0 unchanged. Neither leaves the prover a choice, so a proof over such a row is a deterministic execution; it is not one the executor performs, and the honest prover never produces it, since the executor refuses the program. This has been checked by construction: a guest issuing syscall 4999 (or 0x55) is refused by the executor, while a prover that handles the number the way the chips accept it obtains a proof that the unmodified verifier accepts, with the public values the no-op semantics give (0, or v0 unchanged). Closing it means either decoding every accepted number in the two chips (the verifying keys move) or making the executor accept unknown numbers exactly as the circuit does.
Cross-Shard Verification
Linux syscalls execute across two shards:
- Core shard: CPU decodes
SYSCALL, writes(shard, clk, syscall_id, arg1, arg2)toSyscallChip(Core). - Precompile shard:
SysLinuxChipcomputes the result, receives the same tuple fromSyscallChip(Precompile).
Both SyscallChip instances send two global lookup messages:
Global #1 (Syscall kind):
[shard, clk, syscall_id, arg1_lo, arg1_hi, arg2_lo, arg2_hi]
Global #2 (SyscallResult kind):
[shard, clk, syscall_id, result_lo, result_hi, 0, 0]
The GlobalChip verifies that send/receive multiplicities match for both messages. This ensures:
- Argument integrity: The Core and Precompile shards process the same 32-bit arguments (half-word packed to prevent
reduce()collisions modulo the KoalaBear prime). - Result consistency: Both shards agree on the syscall return value.
Arguments and results are packed as half-words (lo = byte0 + byte1 * 256, hi = byte2 + byte3 * 256), each U16Range-checked to [0, 65535]. This decomposition is injective for 32-bit values, unlike reduce() which can collide.
Linux vs Non-Linux Syscalls
Ziren supports two categories of syscalls, distinguished by the encoding of the syscall code in register V0:
| Property | Linux Syscalls | Non-Linux (Precompile) Syscalls |
|---|---|---|
| Syscall ID | MIPS ABI numbers (4003, 4045, 4210, ...) | Ziren-defined codes (0x00000005, 0x01010006, ...) |
| ID encoding | Byte 1 of V0 is non-zero | Byte 1 of V0 is zero |
| Proving chip | SysLinuxChip | Dedicated chip per precompile (SHA256, Poseidon2, etc.) |
| Argument handling | Full byte-level: a0/a1 as Word<T> (4 x u8) | Reduced field element: arg1 = reduce(op_b) |
| Result handling | result as Word<T>, A3 error code | Return value in V0 (precompile-specific) |
| Global linkage | Half-word packed args + result via two global lookups | Half-word packed args via one global lookup |
| Local linkage | send_syscall_result_packed with is_linux multiplicity | send_syscall / receive_syscall (reduced args only) |
How detection works
The SyscallInstrsChip examines byte 1 of the syscall code (prev_a_value[1]):
byte[1] != 0→ Linux syscall, routed toSysLinuxChipvia lookupbyte[1] == 0→ Precompile syscall, routed to the precompile's dedicated chip
This is enforced bidirectionally via an IsZeroOperation, preventing a malicious prover from misrouting a precompile call into the Linux path or vice versa.
Argument representation across the pipeline
SyscallInstrsChip (Core shard):
send_syscall(reduce(op_b), reduce(op_c)) -- reduced, for both types
send_syscall_result(op_a_word, op_b_word, op_c_word) -- byte-level, linux only
SyscallChip (bridge):
arg1 = arg1_lo + arg1_hi * 65536 -- derived inline (not stored)
arg1_lo, arg1_hi: U16Range-checked -- for ALL syscalls
Global #1: [shard, clk, id, a1_lo, a1_hi, a2_lo, a2_hi] -- collision-resistant
Global #2: [shard, clk, id, result_lo, result_hi, 0, 0] -- result linkage
SysLinuxChip (Precompile shard, linux only):
receive_syscall_result(result_word, a0_word, a1_word) -- byte-level constraints
Constrains result based on a0/a1 byte values
Other precompile chips (Precompile shard, non-linux):
receive_syscall(reduce(arg1), reduce(arg2)) -- reduced args only
Constrains output based on memory at arg addresses
The key distinction: linux syscalls need byte-level argument constraints (e.g., MMAP checks page alignment of a1 bytes, FCNTL checks a0 == 0/1/2 as integers). Non-linux precompiles typically use arguments as memory pointers and operate on the data at those addresses, so reduced field elements suffice.
AIR Soundness Properties
The SysLinuxChip enforces these key properties (all proven bidirectionally):
-
Syscall routing: Each Linux syscall ID maps to exactly one handler branch, and IDs outside the handled set take the NOP branch. Bidirectional
IsZeroOperationdecoders prevent misrouting (e.g., CLONE cannot be routed to NOP). -
Argument decoding:
a0 == 0/1/2anda1 == 1/3flags are bidirectional: when the argument matches a known value, the flag MUST be set. -
Result correctness: Every branch constrains both
result(V0) andoutput(A3) to specific values matching the executor semantics. -
Memory consistency: Read-only memory accesses (BRK read, A2 read) enforce
value == prev_value. Write accesses (HEAP update) use bytewiseAddOperation. -
Page alignment: MMAP size is constrained byte-by-byte. Low 12 bits of
mmap_sizeare structurally zero (byte0 = 0, byte1 is always a multiple of 16). No fieldreduce()is used. -
Cross-shard linkage: Two global lookups per syscall: one for collision-resistant argument matching (half-word packed), one for result consistency. U16Range checks on all half-words ensure canonical decomposition for both linux and non-linux syscalls.
Developer Tutorial
In essence, the "computation problem" in Ziren is the given program, and its "solution" is the execution trace produced when running that program. This trace details every step of the program execution, with each row corresponding to a single step (or a cycle) and each column representing a fixed CPU variable or register state.
Proving a program means checking that every step in the trace follows the corresponding MIPS instruction, encoding the trace columns as polynomials, and committing to those polynomials with a polynomial commitment scheme.
Below is the workflow of Ziren.

High-Level Workflow of Ziren
Referring to the above diagram, Ziren follows a structured pipeline composed of the following stages:
-
Guest Program A program written in a high-level language such as Rust, Go or C/C++, containing the application logic that needs to be proved.
-
MIPS Compiler The high-level program is compiled into a MIPS32r2 little-endian ELF binary by the Ziren toolchain.
-
ELF Loader The ELF loader reads the ELF file and prepares it for execution within the MIPS VM: it places the code and data segments at their virtual addresses, initializes memory, and sets the program's entry point.
-
MIPS VM The MIPS virtual machine runs the loaded ELF file. It records every step of execution, including register states, memory accesses and instruction addresses, as events from which the execution trace is generated.
-
Execution Trace The trace is the data the proof is about. Each row of a chip's table records one operation, and the constraints of every chip check that the rows follow the semantics of the MIPS instructions.
-
Prover The prover takes the execution trace and generates a zero-knowledge proof that the program ran from its entry point to a halt and produced the committed public values, without revealing private inputs.
-
Verifier Proofs can be verified natively with the
zkm-verifiercrate (also inno_stdand WASM builds), or on EVM-compatible chains with the Solidity verifier contracts shipped with each release.
Prover Internal Proof Generation Steps
Within the prover, Ziren processes the execution trace in several stages, ultimately producing a proof suitable for on-chain verification:
-
Shard A long execution is split into shards so that each shard's trace fits in memory. Each shard is proved independently, and the shard proofs are later joined by recursion.
-
Chip Each instruction in a shard generates one or more events, and each event is recorded in the table of a specific chip (for example
AddSub,LoadWord,Branch, or a precompile chip such asKeccakSponge), each with its own set of constraints. -
Lookup Lookups serve two purposes:
- Cross-chip communication: a chip sends the facts it cannot check itself (for example a byte range or a memory access) to the chip that checks them.
- Memory consistency: the value a memory read returns is the value last written to that address.
Within a shard both are proved with the LogUp argument, evaluated with the GKR protocol (LogUp-GKR). Across shards, memory is reconciled with a multiset hash on an elliptic curve over a degree-7 extension of the KoalaBear field.
-
Core Proof The core proof is the list of shard proofs. Each shard proof commits to all of the shard's chip tables as one jagged multilinear polynomial and opens it with the WHIR polynomial commitment scheme.
-
Compressed Proof The shard proofs are aggregated into a single constant-size proof by recursion. A normalize (leaf) program verifies one shard proof, and compose programs of arity up to three merge adjacent ranges of shards until one root remains. Every recursion program's verifying key must belong to a published allowlist of keys (a Merkle tree, the "vk map"), whose root is a public value of the compressed proof.
-
SNARK Proof The compressed proof is shrunk and wrapped into a proof over the BN254 field, which is then proved with either Groth16 or PLONK, resulting in a final Groth16 or PLONK proof that can be verified on-chain.
In summary, Ziren compiles a high-level program into MIPS instructions, runs those instructions to produce an execution trace, proves the trace shard by shard with a STARK built from LogUp-GKR, a zerocheck and the jagged-over-WHIR polynomial commitment, aggregates the shard proofs by recursion, and wraps the result in a Groth16 or PLONK proof.
Program
In Ziren, a prover runs a public program on private inputs and wants to convince a verifier that the program executed correctly and produced the asserted output, without revealing anything about the inputs or the intermediate state of the computation.

All inputs are private; the program and its committed output are public.
From a developer's perspective a Ziren application has two parts: the program to be proved and the program that proves it.
The former is called the guest, and the latter the host.
Host Program
In a Ziren application, the host is the machine that runs the zkVM. The host is an untrusted agent that sets up the zkVM environment, supplies inputs to the guest, and collects its outputs and proofs.
Example: Fibonacci
This host program sends the input n = 1000 to the guest program, executes it, proves it, and verifies the proof.
use zkm_sdk::{include_elf, utils, ProverClient, ZKMProofWithPublicValues, ZKMStdin}; /// The ELF we want to execute inside the zkVM. const ELF: &[u8] = include_elf!("fibonacci"); fn main() { utils::setup_logger(); // The input stream that the guest reads with `zkm_zkvm::io::read`. The types written here // must match the types the guest reads, in the same order. let n = 1000u32; let mut stdin = ZKMStdin::new(); stdin.write(&n); // The prover is selected by the `ZKM_PROVER` environment variable (CPU by default). let client = ProverClient::new(); // Execute the guest without generating a proof. let (_, report) = client.execute(ELF, &stdin).run().unwrap(); println!("executed program with {} cycles", report.total_instruction_count()); // Generate the proving and verifying keys, then a proof (core mode by default). let (pk, vk) = client.setup(ELF); let mut proof = client.prove(&pk, stdin).run().unwrap(); println!("generated proof"); // Read the values the guest committed with `zkm_zkvm::io::commit`, in the same order. let _ = proof.public_values.read::<u32>(); let a = proof.public_values.read::<u32>(); let b = proof.public_values.read::<u32>(); println!("a: {}", a); println!("b: {}", b); // Verify the proof and its public values. client.verify(&proof, &vk).expect("verification failed"); // Proofs can be saved and loaded. proof.save("proof-with-pis.bin").expect("saving proof failed"); let deserialized_proof = ZKMProofWithPublicValues::load("proof-with-pis.bin").expect("loading proof failed"); client.verify(&deserialized_proof, &vk).expect("verification failed"); println!("successfully generated and verified proof for the program!") }
Note that execute takes the input stream by reference (&stdin), while prove takes it by value.
For more details, see the Prover page.
Guest Program
In Ziren, the guest program is the code that is executed and proven by the zkVM.
Programs written in Rust, Go, C/C++ and other languages can be compiled into a MIPS32r2 little-endian ELF executable with the Ziren toolchain, as long as the result satisfies the zkVM's specification.
Ziren provides Rust runtime libraries for guest programs to handle input and output:
zkm_zkvm::io::read::<T>()reads a value of typeTfrom the input stream.zkm_zkvm::io::commit::<T>(&value)commits a value of typeTto the public values.
T must implement serde::Deserialize (for read) or serde::Serialize (for commit). For raw bytes, the following functions bypass serialization and use fewer cycles:
zkm_zkvm::io::read_vec()reads the next input as aVec<u8>.zkm_zkvm::io::commit_slice(&[u8])commits raw bytes to the public values.
Ziren also provides a Go runtime library, github.com/ProjectZKM/Ziren/crates/go-runtime/zkvm_runtime, with:
zkvm_runtime.Read[T any]()for reading structured data;zkvm_runtime.Commit[T any](value)for committing structured data;zkvm_runtime.RuntimeExit(code)for exiting the program.
Guest Program Example
Below are guest programs written in Rust, Go and C/C++.
Rust Example: Fibonacci
//! A simple program that takes a number `n` as input, and writes the `n-1`th and `n`th fibonacci //! number as an output. // These two lines are necessary for the program to properly compile. // // Under the hood, we wrap your main function with some extra code so that it behaves properly // inside the zkVM. #![no_std] #![no_main] zkm_zkvm::entrypoint!(main); pub fn main() { // Read an input to the program. Behind the scenes, this is a system call that reads from // the input stream the host provided. let n = zkm_zkvm::io::read::<u32>(); // Commit n to the public values. zkm_zkvm::io::commit(&n); // Compute the n'th fibonacci number, using normal Rust code. let mut a = 0; let mut b = 1; for _ in 0..n { let mut c = a + b; c %= 7919; // Modulus to prevent overflow. a = b; b = c; } // Commit the outputs of the program. zkm_zkvm::io::commit(&a); zkm_zkvm::io::commit(&b); }
Go Example: Simple-Go
package main
import (
"log"
"github.com/ProjectZKM/Ziren/crates/go-runtime/zkvm_runtime"
)
func main() {
a := zkvm_runtime.Read[uint32]()
if a != 10 {
log.Fatal("%x != 10", a)
}
zkvm_runtime.Commit[uint32](a)
}
A Go guest is compiled with the standard Go toolchain for GOOS=linux GOARCH=mipsle GOMIPS=softfloat, with the ziren build tag and the overlay that zkm_build::generate_go_overlay produces. The example's host/build.rs shows the full command; the host then embeds the resulting binary with include_bytes!.
C/C++ Example: Fibonacci_C
For other languages, compile the code to a static library and link it into a Rust guest through FFI. The example compiles add.cpp and modulus.c with the cc crate in the guest's build.rs:
fn main() { cc::Build::new().file("src/c_lib/add.cpp").compile("libadd.a"); cc::Build::new().file("src/c_lib/modulus.c").compile("libmodulus.a"); }
The cc crate compiles for the guest target, so it needs a C compiler that emits MIPS32r2 objects; the host's default cc emits objects the guest linker rejects as incompatible. Without a MIPS cross compiler, clang works when it is given the guest target's ABI (soft float, no abicalls, static relocation):
FLAGS="--target=mipsel-unknown-none-elf -march=mips32r2 -msoft-float -mno-abicalls -fno-pic -ffreestanding"
export CC_mipsel_zkm_zkvm_elf=clang CXX_mipsel_zkm_zkvm_elf=clang++ AR_mipsel_zkm_zkvm_elf=llvm-ar
export CFLAGS_mipsel_zkm_zkvm_elf="$FLAGS" CXXFLAGS_mipsel_zkm_zkvm_elf="$FLAGS -fno-exceptions -fno-rtti"
The same settings build the bitcoin example, whose secp256k1-sys dependency compiles C.
add.cpp:
extern "C" {
unsigned int add(unsigned int a, unsigned int b) {
return a + b;
}
}
The Rust guest declares and calls the C functions:
#![no_std] #![no_main] zkm_zkvm::entrypoint!(main); // Use the add and modulus functions from the static libraries. extern "C" { fn add(a: u32, b: u32) -> u32; fn modulus(a: u32, b: u32) -> u32; } pub fn main() { let n = zkm_zkvm::io::read::<u32>(); zkm_zkvm::io::commit(&n); let mut a = 0; let mut b = 1; unsafe { for _ in 0..n { let mut c = add(a, b); c = modulus(c, 7919); a = b; b = c; } } zkm_zkvm::io::commit(&a); zkm_zkvm::io::commit(&b); }
Compiling Guest Program
The guest program has to be compiled to an ELF file that the zkVM executes. The Ziren toolchain must be installed (see Installation).
To build the guest automatically when compiling or running the host crate, add a build.rs file to your host/ directory (next to the host crate's Cargo.toml) that uses the zkm-build crate:
.
├── guest
└── host
├── build.rs # Add this file
├── Cargo.toml
└── src
build.rs:
fn main() { zkm_build::build_program("../guest"); }
The host then embeds the compiled ELF with zkm_sdk::include_elf!("<guest package name>").
The Ziren crates are not published on crates.io; depend on them from the Ziren repository. In host/Cargo.toml:
[dependencies]
zkm-sdk = { git = "https://github.com/ProjectZKM/Ziren" }
[build-dependencies]
zkm-build = { git = "https://github.com/ProjectZKM/Ziren" }
and in guest/Cargo.toml:
[dependencies]
zkm-zkvm = { git = "https://github.com/ProjectZKM/Ziren" }
Pin a tag or a revision (tag = "..." or rev = "...") so that the guest ELF, and therefore the program's verifying key, does not change when the repository moves.
Advanced Build Options
The build can be configured by passing a BuildArgs struct to build_program_with_args(). BuildArgs selects features, extra rustc flags, packages, binaries, the output ELF name and directory, and static C/C++ libraries to link.
For example, the following build.rs builds every guest program in a guests workspace with the default arguments:
use std::{ io::{Error, Result}, path::PathBuf, }; use zkm_build::build_program_with_args; fn main() -> Result<()> { let workspace_path = [env!("CARGO_MANIFEST_DIR"), "guests"].iter().collect::<PathBuf>().canonicalize()?; build_program_with_args( workspace_path.to_str().ok_or_else(|| { Error::other(format!("expected {workspace_path:?} to be valid UTF-8")) })?, Default::default(), ); Ok(()) }
Example Walkthrough - Best Practices
This page walks through the Fibonacci example in the Ziren repository. It has the standard layout of a Ziren application:
examples/fibonacci
├── guest
│ ├── Cargo.toml
│ └── src/main.rs # the program that is proved
└── host
├── Cargo.toml
├── build.rs # compiles the guest to a MIPS ELF
├── bin/ # one binary per proof mode
└── src/main.rs # executes, proves and verifies the guest
Guest
guest/src/main.rs:
//! A simple program that takes a number `n` as input, and writes the `n-1`th and `n`th fibonacci //! number as an output. // These two lines are necessary for the program to properly compile. // // Under the hood, we wrap your main function with some extra code so that it behaves properly // inside the zkVM. #![no_std] #![no_main] zkm_zkvm::entrypoint!(main); pub fn main() { // Read an input to the program. Behind the scenes, this is a system call that reads from // the input stream the host provided. let n = zkm_zkvm::io::read::<u32>(); // Commit n to the public values. zkm_zkvm::io::commit(&n); // Compute the n'th fibonacci number, using normal Rust code. let mut a = 0; let mut b = 1; for _ in 0..n { let mut c = a + b; c %= 7919; // Modulus to prevent overflow. a = b; b = c; } // Commit the outputs of the program. zkm_zkvm::io::commit(&a); zkm_zkvm::io::commit(&b); }
guest/Cargo.toml:
[package]
name = "fibonacci"
version = "1.1.0"
edition = "2021"
publish = false
[dependencies]
zkm-zkvm = { path = "../../../crates/zkvm/entrypoint", features= ["embedded"] }
The package name (fibonacci) is the name the host passes to include_elf!. The embedded feature selects an allocator that can free memory, in place of the default bump allocator. Outside the Ziren repository, replace the path dependency by a git dependency, as shown in Guest Program.
Host
host/build.rs compiles the guest whenever the host is built:
fn main() { zkm_build::build_program("../guest"); }
host/Cargo.toml (abbreviated; the example also declares the other binaries in bin/):
[package]
name = "fibonacci-host"
version = { workspace = true }
edition = { workspace = true }
default-run = "fibonacci-host"
publish = false
[dependencies]
hex = "0.4.3"
zkm-sdk = { workspace = true }
[build-dependencies]
zkm-build = { workspace = true }
[[bin]]
name = "groth16_bn254"
path = "bin/groth16_bn254.rs"
[[bin]]
name = "fibonacci-host"
path = "src/main.rs"
host/src/main.rs executes the guest, generates a core proof, reads the public values, and verifies the proof; it is listed on the Host Program page. The bin/ directory holds one host per mode:
| Binary | What it does |
|---|---|
execute | executes the guest and prints the execution report |
compressed | generates and verifies a compressed proof |
groth16_bn254 | generates and verifies a Groth16 proof for on-chain use |
plonk_bn254 | generates and verifies a PLONK proof for on-chain use |
host/bin/groth16_bn254.rs:
use zkm_sdk::{include_elf, utils, HashableKey, ProverClient, ZKMStdin}; /// The ELF we want to execute inside the zkVM. const ELF: &[u8] = include_elf!("fibonacci"); fn main() { utils::setup_logger(); let n = 500u32; let mut stdin = ZKMStdin::new(); stdin.write(&n); let client = ProverClient::new(); let (pk, vk) = client.setup(ELF); println!("vk: {:?}", vk.bytes32()); let proof = client.prove(&pk, stdin).groth16().run().unwrap(); println!("generated proof"); let public_values = proof.public_values.as_slice(); println!("public values: 0x{}", hex::encode(public_values)); let solidity_proof = proof.bytes().expect("the proof has a byte encoding"); println!("proof: 0x{}", hex::encode(solidity_proof)); client.verify(&proof, &vk).expect("verification failed"); proof.save("fibonacci-groth16.bin").expect("saving proof failed"); println!("successfully generated and verified proof for the program!") }
vk.bytes32() is the program verifying key hash that an on-chain verifier checks the proof against, proof.public_values.as_slice() are the committed public values, and proof.bytes() is the proof in the encoding the Solidity verifier expects (see Verifier).
Running the Example
From examples/fibonacci/host:
# execute, prove (core mode) and verify
RUST_LOG=info cargo run --release
# execute only
RUST_LOG=info cargo run --release --bin execute
# Groth16 proof for on-chain verification
RUST_LOG=info cargo run --release --bin groth16_bn254
Best Practices
Guest
-
Start every Rust guest with
#![no_std],#![no_main]andzkm_zkvm::entrypoint!(main);. -
Read inputs in the order the host wrote them, with the same types.
read::<T>()deserializes with bincode; for raw bytes,read_vec()andcommit_slice()skip serialization and cost fewer cycles. -
Commit only what the verifier needs to see. Everything committed with
commitorcommit_slicebecomes public. -
If a Solidity contract consumes the public values, commit them in an ABI encoding (for example with
alloy_sol_types::SolType::abi_encode) and decode them on-chain withabi.decode. -
Use the precompiles for cryptographic operations; they are much cheaper than the same code compiled to MIPS instructions. For example, the keccak-precompile example hashes with:
#![allow(unused)] fn main() { use zkm_zkvm::lib::keccak256::keccak256; let output = keccak256(&input.as_slice()); }Many common crates (
sha2,k256,p256,substrate-bn, and others) have patched versions that call the precompiles. -
Keep the guest's dependencies pinned. Any change to the guest ELF changes the program's verifying key.
Host
- Run
executebeforeprove. It is much faster, catches guest panics, and its report gives the cycle count (see Optimizations). - Read the public values in the order the guest committed them. In the Fibonacci example the first value is
n. - Choose the proof mode by the consumer: core or compressed proofs for off-chain verification, Groth16 or PLONK for on-chain verification.
- Publish
vk.bytes32()with the application. A verifier contract must be configured with this value, and it only changes when the guest ELF changes. - Save proofs with
proof.save(...)and load them withZKMProofWithPublicValues::load(...)to verify them elsewhere.
Prover
The zkm_sdk crate provides the tools for proof generation. Its ProverClient lets you:
- generate the proving and verifying keys with
setup(); - execute a program without proving it with
execute(); - generate proofs with
prove(); - verify proofs with
verify().
When generating Groth16 or PLONK proofs, the ProverClient downloads the circuit artifacts of the current release (the proving key from the trusted setup, the verifying key and the Solidity verifier) on first use with try_install_circuit_artifacts(), and caches them in ~/.zkm/circuits/{groth16,plonk}/<version> (for example v2.0.0).
Example: Fibonacci
The following code uses zkm_sdk in a host program.
use zkm_sdk::{include_elf, utils, ProverClient, ZKMProofWithPublicValues, ZKMStdin}; /// The ELF we want to execute inside the zkVM. const ELF: &[u8] = include_elf!("fibonacci"); fn main() { utils::setup_logger(); let n = 1000u32; let mut stdin = ZKMStdin::new(); stdin.write(&n); let client = ProverClient::new(); let (_, report) = client.execute(ELF, &stdin).run().unwrap(); println!("executed program with {} cycles", report.total_instruction_count()); let (pk, vk) = client.setup(ELF); let mut proof = client.prove(&pk, stdin).run().unwrap(); println!("generated proof"); let _ = proof.public_values.read::<u32>(); let a = proof.public_values.read::<u32>(); let b = proof.public_values.read::<u32>(); println!("a: {}", a); println!("b: {}", b); client.verify(&proof, &vk).expect("verification failed"); proof.save("proof-with-pis.bin").expect("saving proof failed"); let deserialized_proof = ZKMProofWithPublicValues::load("proof-with-pis.bin").expect("loading proof failed"); client.verify(&deserialized_proof, &vk).expect("verification failed"); println!("successfully generated and verified proof for the program!") }
Proof Types
The proof mode is chosen on the prove() builder. The proof itself is a ZKMProof:
#![allow(unused)] fn main() { pub enum ZKMProof { /// A proof generated by the core proof mode. /// /// The proof size scales linearly with the number of cycles. Core(Vec<ShardProof<CoreSC>>), /// A proof generated by the compress proof mode. /// /// The proof size is constant, regardless of the number of cycles. Compressed(Box<ZKMReduceProof<InnerSC>>), /// A proof generated by the Plonk proof mode. Plonk(PlonkBn254Proof), /// A proof generated by the Groth16 proof mode. Groth16(Groth16Bn254Proof), /// A proof generated by the DV-SNARK proof mode. DvSnark(DvSnarkBn254Proof), /// Compressed-proof-to-Groth16 conversion. CompressToGroth16, } }
A proof is returned as a ZKMProofWithPublicValues, which bundles the proof, the committed public values and the Ziren version, and can be written and read back with save() and load().
Core Proof (Default)
The default mode produces one STARK proof per shard; the total size grows linearly with the length of the execution.
#![allow(unused)] fn main() { let client = ProverClient::new(); client.prove(&pk, stdin).run().unwrap(); }
Compressed Proof
The compressed mode aggregates the shard proofs by recursion into a single STARK proof of constant size. It is verified natively (for example with zkm_verifier::StarkVerifier), but is too large for on-chain verification.
#![allow(unused)] fn main() { let client = ProverClient::new(); client.prove(&pk, stdin).compressed().run().unwrap(); }
Groth16 Proof (Recommended)
The Groth16 mode wraps the compressed proof into a Groth16 proof over BN254 of 260 bytes (a 4-byte verifier selector and 8 field elements), verifiable on-chain.
#![allow(unused)] fn main() { let client = ProverClient::new(); client.prove(&pk, stdin).groth16().run().unwrap(); }
PLONK Proof
The PLONK mode wraps the compressed proof into a PLONK proof over BN254 of about 868 bytes, also verifiable on-chain. PLONK uses a universal setup (the Aztec Ignition SRS) instead of a circuit-specific trusted setup ceremony.
#![allow(unused)] fn main() { let client = ProverClient::new(); client.prove(&pk, stdin).plonk().run().unwrap(); }
Other Modes and Options
compress_to_groth16()converts an existing compressed proof into a Groth16 proof. The input stream must hold exactly one compressed proof (written withstdin.write_proof) and the bincode-encoded public values as its only input.- The builder also accepts
shard_size(),shard_batch_size(),cycle_limit(),with_hook(), andtimeout()(network prover only).
Immutable Wrap Verifying Key
By default the Groth16 circuit is specific to one Ziren release. In the immutable-wrap-vk mode (ZKM_IMM_WRAP_VK=1 when building the guest and proving), a single Groth16 verifying key serves all releases: the guest hashes its public values with BLAKE3 instead of SHA-256, and the commitment and start pc of the release's partial STARK verifying key are hashed into the program key hash. zkm-build then builds the guest with its imm-wrap-vk feature, so the guest's Cargo.toml must forward it:
[features]
imm-wrap-vk = ["zkm-zkvm/imm-wrap-vk"]
Such proofs are verified with Groth16Verifier::verify_by_imm_groth16_vk, and their artifacts live in ~/.zkm/circuits/groth16/imm-wrap-vk/<version>. See the imm-wrap-vk-add example.
Hardware Acceleration
GPU Acceleration
Ziren provides a CUDA-based GPU prover, which proves with much lower latency and better cost-performance than the CPU prover.
Software Requirements
- CUDA 12
- Docker with the NVIDIA Container Toolkit; the user must be allowed to run
docker.
Hardware Requirements
- Processor: 4-core CPU or higher
- System memory: 16 GB RAM or higher
- Graphics card: 24 GB VRAM or higher
The NVIDIA GPU must have a compute capability of at least 8.6. You can check yours in the official NVIDIA documentation.
Usage
Build a CUDA client in one of two ways:
- Option A (environment variable): use
ProverClient::new()withZKM_PROVER=cuda. - Option B (direct method): use
ProverClient::cuda().
By default the client starts the GPU prover in a Docker container. Because the container receives the private input stream, the image must be pinned by digest:
export ZKM_GPU_IMAGE=projectzkm/ziren-gpu@sha256:<digest> # a reviewed image digest
A mutable tag such as projectzkm/ziren-gpu:latest is refused unless ZKM_ALLOW_MUTABLE_GPU_IMAGE=1 is set, which is meant for local development only. Further options:
export CUDA_VISIBLE_DEVICE_INDEX=<index> # GPU the container uses
export CUDA_PORT=<port> # port of the container's prover server
export CUDA_RUN_DOCKER=false # connect to an already running GPU server instead
export CUDA_ENDPOINT=<url> # its endpoint (default: http://localhost:3000/twirp/)
With the client built, generate proofs with the standard methods.
CPU Acceleration
Ziren supports AVX2/AVX512 acceleration on x86 CPUs through Plonky3.
Check your CPU's AVX support with:
grep avx /proc/cpuinfo
and look for avx2 or avx512 in the output.
To enable AVX2, add these flags to your RUSTFLAGS environment variable:
RUSTFLAGS="-C target-cpu=native" cargo run --release
To enable AVX512, add these flags to your RUSTFLAGS environment variable:
RUSTFLAGS="-C target-cpu=native -C target-feature=+avx512f" cargo run --release
Network Prover
The ZKM Proof Network proves programs remotely. The SDK's network prover (the network feature of zkm-sdk) produces compressed and Groth16 proofs, and converts compressed proofs to Groth16 (compress_to_groth16()); it does not produce core or PLONK proofs.
A network proof passes through several stages (queuing, splitting, proving and finalizing), and each stage may take a different amount of time.
Requirements
- Register your address to gain access.
- A client certificate and key issued for the network, and the certificate of the CA that signs the network's certificate.
- SDK dependency: add
zkm-sdkwith thenetworkfeature to yourCargo.toml:
zkm-sdk = { git = "https://github.com/ProjectZKM/Ziren", features = ["network"] }
Environment Variable Setup
Before running your application, export the following environment variables:
export ZKM_PRIVATE_KEY=<your_private_key> # Private key corresponding to your registered public key
export SSL_CERT_PATH=<path_to_ssl_certificate> # Path to the SSL client certificate (e.g., ssl.pem)
export SSL_KEY_PATH=<path_to_ssl_key> # Path to the SSL client private key (e.g., ssl.key)
export CA_CERT_PATH=<path_to_ca_certificate> # Required when SSL_CERT_PATH/SSL_KEY_PATH are set
The repository's crates/sdk/tool directory contains a certgen.sh script and a test CA (ca.pem, ca.key). The test CA's private key is public, so it is only for local testing: the SDK uses it only when ZKM_ALLOW_INSECURE_TEST_CA=1 is set, and never for real witness data (see INSECURE-TEST-PKI.md there).
Optional: customize the network prover's behavior:
export SHARD_SIZE=<shard_size> # Shard (segment) size requested from the network
export MAX_PROVER_NUM=<max_prover_num> # Maximum number of provers to use in parallel
export SINGLE_NODE=<true|false> # Whether to use a single node for proving (default: false)
export ZKM_PROOF_POLL_INTERVAL=<seconds> # How often to poll for the proof status
To use your own proof network endpoint:
export ENDPOINT=<proof_network_endpoint> # Proof network endpoint (default: https://152.32.186.45:20002)
export DOMAIN_NAME=<domain_name> # TLS domain name (default: "stage")
Example
The following host uses the network prover:
use zkm_sdk::{include_elf, utils, ProverClient, ZKMStdin}; const FIBONACCI_ELF: &[u8] = include_elf!("fibonacci"); fn main() { utils::setup_logger(); let mut stdin = ZKMStdin::new(); stdin.write(&10u32); // Create a network client (or set ZKM_PROVER=network and use ProverClient::new()). let client = ProverClient::network(); let (pk, vk) = client.setup(FIBONACCI_ELF); let proof = client.prove(&pk, stdin).groth16().run().unwrap(); client.verify(&proof, &vk).unwrap(); }
Verifier
Proof Verification Overview
A verifier checks a proof against:
- the program's verifying key (or its 32-byte hash);
- the public values the program committed;
- the proof itself.
Both the proving key and the verifying key are derived from the compiled guest ELF by setup. The 32-byte verifying key hash identifies the program: a proof verifies only under the key of the program that produced it, and any change to the guest ELF changes the key. Retrieve the hash with the SDK:
#![allow(unused)] fn main() { use zkm_sdk::{HashableKey, ProverClient}; let client = ProverClient::new(); let (_pk, vk) = client.setup(ELF); let vkey_hash = vk.bytes32(); // 0x-prefixed hex string of the 32-byte program vk hash }
Within a host, client.verify(&proof, &vk) verifies any proof kind. The rest of this page covers verification outside the SDK: on-chain, with the zkm-verifier crate, in WASM, inside the zkVM, and in BitVM.
On-chain verification
A verifier smart contract lets anyone check a proof on an EVM chain. The proof is generated off-chain; the chain only pays for verification, whose cost does not depend on the size of the proved computation.
STARK proofs are too large to verify economically on Ethereum, so Ziren wraps them into Groth16 or PLONK proofs over BN254. A Groth16 proof is 260 bytes and a PLONK proof about 868 bytes, and both are checked with a constant number of BN254 precompile calls. Every on-chain proof has three public inputs:
- the program verifying key hash (
vk.bytes32()); - the digest of the public values: SHA-256 of the committed bytes, masked to 253 bits so it fits in a BN254 field element;
- the root of the recursion verifying key allowlist (the "vk map") of the Ziren release that produced the proof.
Verifier Contracts
The verifier contracts of each release are generated when the circuit is built, from the templates in crates/recursion/gnark-ffi/assets, and are shipped in the release's circuit artifacts. After the SDK has installed the artifacts (see Prover), they are in:
~/.zkm/circuits/groth16/<version>/:ZKMVerifierGroth16.solandGroth16Verifier.sol;~/.zkm/circuits/plonk/<version>/:ZKMVerifierPlonk.solandPlonkVerifier.sol.
The contracts are:
- IZKMVerifier is the verifier interface.
- ZKMVerifierGroth16.sol and ZKMVerifierPlonk.sol (both define a contract named
ZKMVerifier) check the verifier selector, compute the public inputs, and call the proof system verifier. - Groth16Verifier.sol and PlonkVerifier.sol implement the pairing checks of the proof systems; they are generated by gnark from the circuit's verifying key.
IZKMVerifier.sol:
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.20;
/// @title Ziren Verifier Interface
/// @author ZKM Labs
/// @notice This contract is the interface for the Ziren Verifier.
interface IZKMVerifier {
/// @notice Verifies a proof with given public values and vkey.
/// @dev It is expected that the first 4 bytes of proofBytes must match the first 4 bytes of
/// target verifier's VERIFIER_HASH.
/// @param programVKey The verification key for the MIPS program.
/// @param publicValues The public values encoded as bytes.
/// @param proofBytes The proof of the program execution the Ziren zkVM encoded as bytes.
function verifyProof(
bytes32 programVKey,
bytes calldata publicValues,
bytes calldata proofBytes
) external view;
}
interface IZKMVerifierWithHash is IZKMVerifier {
/// @notice Returns the hash of the verifier.
function VERIFIER_HASH() external pure returns (bytes32);
}
verifyProof takes the program verifying key hash, the public values and the proof. VERIFIER_HASH is the SHA-256 hash of the Groth16 or PLONK verifying key; the first 4 bytes of every proof (as returned by proof.bytes()) must equal its first 4 bytes. The check rejects a proof sent to the wrong verifier, for example a PLONK proof sent to the Groth16 verifier or a proof from another release.
The Groth16 template, ZKMVerifierGroth16.txt, from which each release's ZKMVerifierGroth16.sol is generated:
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.20;
import {IZKMVerifier, IZKMVerifierWithHash} from "../IZKMVerifier.sol";
import {Groth16Verifier} from "./Groth16Verifier.sol";
/// @title Ziren Verifier
/// @author ZKM Labs
/// @notice This contracts implements a solidity verifier for Ziren.
contract ZKMVerifier is Groth16Verifier, IZKMVerifierWithHash {
/// @notice Thrown when the verifier selector from this proof does not match the one in this
/// verifier. This indicates that this proof was sent to the wrong verifier.
/// @param received The verifier selector from the first 4 bytes of the proof.
/// @param expected The verifier selector from the first 4 bytes of the VERIFIER_HASH().
error WrongVerifierSelector(bytes4 received, bytes4 expected);
/// @notice Thrown when the proof is invalid.
error InvalidProof();
function VERSION() external pure returns (string memory) {
return "{ZKM_CIRCUIT_VERSION}";
}
/// @inheritdoc IZKMVerifierWithHash
function VERIFIER_HASH() public pure returns (bytes32) {
return {VERIFIER_HASH};
}
/// @notice The root of the Merkle tree of recursion verifying keys this verifier accepts.
/// @dev Inside the proof tree this root is a witness the prover supplies, so the in-circuit
/// checks only establish that every child key lies in a tree with *that* root. Supplying this
/// value as a public input here, rather than taking it from the caller, is what binds a proof
/// to the published recursion programs and rules out one built around a substituted compose,
/// leaf or shrink program.
function VK_ROOT() public pure returns (bytes32) {
return {VK_ROOT};
}
/// @notice Hashes the public values to a field elements inside Bn254.
/// @param publicValues The public values.
function hashPublicValues(
bytes calldata publicValues
) public pure returns (bytes32) {
return sha256(publicValues) & bytes32(uint256((1 << 253) - 1));
}
/// @notice Verifies a proof with given public values and vkey.
/// @param programVKey The verification key for the MIPS program.
/// @param publicValues The public values encoded as bytes.
/// @param proofBytes The proof of the program execution the Ziren zkVM encoded as bytes.
function verifyProof(
bytes32 programVKey,
bytes calldata publicValues,
bytes calldata proofBytes
) external view {
bytes4 receivedSelector = bytes4(proofBytes[:4]);
bytes4 expectedSelector = bytes4(VERIFIER_HASH());
if (receivedSelector != expectedSelector) {
revert WrongVerifierSelector(receivedSelector, expectedSelector);
}
bytes32 publicValuesDigest = hashPublicValues(publicValues);
uint256[3] memory inputs;
inputs[0] = uint256(programVKey);
inputs[1] = uint256(publicValuesDigest);
inputs[2] = uint256(VK_ROOT());
uint256[8] memory proof = abi.decode(proofBytes[4:], (uint256[8]));
this.Verify(proof, inputs);
}
}
The release build fills in {ZKM_CIRCUIT_VERSION} (for example v2.0.0), {VERIFIER_HASH} and {VK_ROOT}. The contract:
- checks that the first 4 bytes of
proofBytesmatchVERIFIER_HASH; - binds the proof to the program through
programVKey; - binds it to the public values through their SHA-256 digest;
- binds it to the release's recursion programs through
VK_ROOT, which is a constant of the contract and not an argument, so a caller cannot substitute another allowlist; - calls
Groth16Verifier.Verify, which reverts if the proof is invalid.
The PLONK contract is the same except for the last step: it passes the three inputs as a uint256[] together with the raw proof bytes to PlonkVerifier.Verify, and reverts with InvalidProof() when that returns false.
A proof verifies only against the contracts of the release that produced it: a different release has a different VERIFIER_HASH or VK_ROOT.
Application Contracts
An application contract stores the program verifying key hash and calls a deployed ZKMVerifier through the interface. The following contract accepts proofs of a Fibonacci guest that commits its outputs ABI-encoded as (uint32 n, uint32 a, uint32 b) (for example with alloy_sol_types::SolType::abi_encode):
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.20;
import {IZKMVerifier} from "./IZKMVerifier.sol";
struct PublicValuesStruct {
uint32 n;
uint32 a;
uint32 b;
}
contract Fibonacci {
/// @notice The address of the Ziren verifier contract.
IZKMVerifier public verifier;
/// @notice The verification key hash for the fibonacci program.
bytes32 public fibonacciProgramVKey;
constructor(IZKMVerifier _verifier, bytes32 _fibonacciProgramVKey) {
verifier = _verifier;
fibonacciProgramVKey = _fibonacciProgramVKey;
}
function verifyFibonacciProof(bytes calldata _publicValues, bytes calldata _proofBytes)
public
view
returns (uint32, uint32, uint32)
{
verifier.verifyProof(fibonacciProgramVKey, _publicValues, _proofBytes);
PublicValuesStruct memory publicValues = abi.decode(_publicValues, (PublicValuesStruct));
return (publicValues.n, publicValues.a, publicValues.b);
}
}
The arguments come from the host (see examples/fibonacci/host/bin/groth16_bn254.rs):
_fibonacciProgramVKeyisvk.bytes32();_publicValuesisproof.public_values.as_slice();_proofBytesisproof.bytes().
Deployment
The contracts are plain Solidity and can be deployed with any tool. With Foundry, place IZKMVerifier.sol in src/ and the release's two Groth16 contracts in src/<version>/ (the layout their imports expect), and deploy the verifier with a script:
// SPDX-License-Identifier: UNLICENSED
pragma solidity ^0.8.20;
import {Script} from "forge-std/Script.sol";
import {ZKMVerifier} from "../src/v2.0.0/ZKMVerifierGroth16.sol";
contract ZKMVerifierGroth16Script is Script {
function run() public {
vm.startBroadcast();
new ZKMVerifier();
vm.stopBroadcast();
}
}
forge script script/ZKMVerifierGroth16.s.sol:ZKMVerifierGroth16Script \
--rpc-url $RPC_URL --private-key $PK --broadcast
Then deploy the application contract with the verifier's address and the program verifying key hash. Anyone can then call verifyProof (directly or through the application contract) with a proof; the call reverts if the proof is invalid.
Off-chain verification
Off-chain verification checks compressed STARK, Groth16 and PLONK proofs without a chain, and so without gas. Compressed proofs can be verified directly, without the STARK-to-SNARK wrapping. The result is only known to whoever runs the check.
The zkm-verifier Crate
The zkm-verifier crate verifies proofs without the prover. It supports no_std (disable the default std feature), so it also builds for WASM and for the zkVM itself. It embeds the verifying keys of its Ziren release:
GROTH16_VK_BYTESandPLONK_VK_BYTES: the Groth16 and PLONK verifying keys;VK_ROOT_BYTES: the recursion verifying key allowlist root, the third public input;IMM_GROTH16_VK_BYTESandPART_STARK_VK_BYTES: the keys for the immutable-wrap-vk mode.
Its verifiers take the byte encodings the SDK produces (proof.bytes(), proof.public_values.to_vec(), vk.bytes32()):
| Function | Verifies |
|---|---|
Groth16Verifier::verify(proof, public_values, vkey_hash, groth16_vk) | a Groth16 proof |
Groth16Verifier::verify_by_imm_groth16_vk(proof, public_values, vkey_hash, imm_groth16_vk, part_stark_vk) | a Groth16 proof made in the immutable-wrap-vk mode; Groth16Verifier::get_part_stark_vk(version) returns a bundled release's partial STARK key |
PlonkVerifier::verify(proof, public_values, vkey_hash, plonk_vk) | a PLONK proof |
StarkVerifier::verify(proof, public_values, vk) | a compressed proof (proof.bytes()) against the bincode-serialized ZKMVerifyingKey, and checks that the public values match the digest committed in the proof |
StarkVerifier::verify_proof(proof, vk) | a compressed proof, without the public values check |
For example:
#![allow(unused)] fn main() { use zkm_verifier::{Groth16Verifier, GROTH16_VK_BYTES}; Groth16Verifier::verify(&proof_bytes, &public_values, &vkey_hash, *GROTH16_VK_BYTES) .expect("invalid proof"); }
With the ark feature, the crate also converts proofs to arkworks types and verifies them with ark-groth16.
WASM Verification
Ziren provides WASM bindings for verifying Groth16, PLONK, and STARK proofs in-browser. These bindings are generated from Rust functions in the zkm_verifier crate and exposed to JavaScript via wasm_bindgen. See the ziren-wasm-verifier repository for setup instructions.
The repository includes the following:
- The guest program (with an example on computing the Fibonacci sequence) takes an input
n, commits it to the public values, computes the sequence, and writes outputs (a,b) to public values via zkVM syscalls. - The host program compiles the guest to an ELF using
zkm_build, runs setup withProverClient, generates the proof, and saves the proof, public values, and verifying key hash. - The verifier directory wraps the Rust verifier functions with
wasm_bindgenso they can be invoked from JavaScript after runningwasm-pack build.
The following is an example verifier wrapper (verifier/lib.rs):
#![allow(unused)] fn main() { //! A simple wrapper around the `zkm_verifier` crate. use wasm_bindgen::prelude::wasm_bindgen; use zkm_verifier::{ Groth16Verifier, PlonkVerifier, StarkVerifier, GROTH16_VK_BYTES, PLONK_VK_BYTES, }; /// Wrapper around [`zkm_verifier::StarkVerifier::verify`]. #[wasm_bindgen] pub fn verify_stark(proof: &[u8], public_inputs: &[u8], zkm_vk: &[u8]) -> bool { StarkVerifier::verify(proof, public_inputs, zkm_vk).is_ok() } /// Wrapper around [`zkm_verifier::Groth16Verifier::verify`]. /// /// We hardcode the Groth16 VK bytes to only verify Ziren proofs. #[wasm_bindgen] pub fn verify_groth16(proof: &[u8], public_inputs: &[u8], zkm_vk_hash: &str) -> bool { Groth16Verifier::verify(proof, public_inputs, zkm_vk_hash, *GROTH16_VK_BYTES).is_ok() } /// Wrapper around [`zkm_verifier::PlonkVerifier::verify`]. /// /// We hardcode the Plonk VK bytes to only verify Ziren proofs. #[wasm_bindgen] pub fn verify_plonk(proof: &[u8], public_inputs: &[u8], zkm_vk_hash: &str) -> bool { PlonkVerifier::verify(proof, public_inputs, zkm_vk_hash, *PLONK_VK_BYTES).is_ok() } }
All Rust functions in the zkm_verifier crate encoding the verification logic for the proof systems are wrapped to generate WASM bindings.
The repository also contains a few examples. The wasm_example script demonstrates verifying proofs in Node.js:
/**
* This script verifies the proofs generated by the script in `example/host`.
*
* It loads json files in `example/json` and verifies them using the wasm bindings
* in `example/verifier/pkg/zkm_wasm_verifier.js`.
*/
import * as wasm from "../../verifier/pkg/zkm_wasm_verifier.js"
import fs from 'node:fs'
import path from 'node:path'
// Convert a hexadecimal string to a Uint8Array
export const fromHexString = (hexString) =>
Uint8Array.from(hexString.match(/.{1,2}/g).map((byte) => parseInt(byte, 16)));
const files = fs.readdirSync("../json");
// Iterate through each file in the data directory
for (const file of files) {
try {
// Read and parse the JSON content of the file
const fileContent = fs.readFileSync(path.join("../json", file), 'utf8');
const proof_json = JSON.parse(fileContent);
// Determine the ZKP type (Groth16 or Plonk) based on the filename
const file_name = file.toLowerCase();
const zkpType = file_name.includes('groth16') ? 'groth16' : file_name.includes('plonk')? 'plonk' : 'stark';
const proof = fromHexString(proof_json.proof);
const public_inputs = fromHexString(proof_json.public_inputs);
const vkey_hash = proof_json.vkey_hash;
// Get the values using DataView.
const view = new DataView(public_inputs.buffer);
// Read each 32-bit (4 byte) integer as little-endian
const n = view.getUint32(0, true);
const a = view.getUint32(4, true);
const b = view.getUint32(8, true);
console.log(`n: ${n}`);
console.log(`a: ${a}`);
console.log(`b: ${b}`);
if (zkpType == 'stark') {
const vkey = fromHexString(proof_json.vkey);
const startTime = performance.now();
const result = wasm.verify_stark(proof, public_inputs, vkey);
const endTime = performance.now();
console.log(`${zkpType} verification took ${endTime - startTime}ms`);
console.assert(result, "result:", result, "proof should be valid");
console.log(`Proof in ${file} is valid.`);
} else {
// Select the appropriate verification function and verification key based on ZKP type
const verifyFunction = zkpType === 'groth16' ? wasm.verify_groth16 : wasm.verify_plonk;
const startTime = performance.now();
const result = verifyFunction(proof, public_inputs, vkey_hash);
const endTime = performance.now();
console.log(`${zkpType} verification took ${endTime - startTime}ms`);
console.assert(result, "result:", result, "proof should be valid");
console.log(`Proof in ${file} is valid.`);
}
} catch (error) {
console.error(`Error processing ${file}: ${error.message}`);
}
}
The following logic is included in the script:
- Loads proof JSON files from
example/json/. - Decodes hex-encoded proof and public inputs.
- Dispatches verification to the appropriate WASM binding (
verify_stark,verify_groth16, orverify_plonk).
STARK verification requires converting the verifying key to bytes and passing it explicitly, while Groth16 and PLONK require the vkey_hash.
This example logs:
- Input values (n, a, b)
- Verification time (in ms)
- Whether the proof is valid
The eth_wasm example demonstrates in-browser STARK verification for Ethereum block proofs for EthProofs.
main.js:
/**
* This script verifies the proofs generated by the script in `example/host`.
*
* It loads json files in `example/json` and verifies them using the wasm bindings
* in `example/verifier/pkg/zkm_wasm_verifier.js`.
*/
import * as wasm from "../../verifier/pkg/zkm_wasm_verifier.js"
import fs from 'node:fs'
const vkey = fs.readFileSync('../binaries/eth_vk.bin');
// Download the proof from https://ethproofs.org/blocks/23174100 > ZKM
const proof = fs.readFileSync('../binaries/23174100_ZKM_167157.txt');
const startTime = performance.now();
const result = wasm.verify_stark_proof(proof, vkey);
const endTime = performance.now();
console.log(`stark verification took ${endTime - startTime}ms`);
console.assert(result, "result:", result, "proof should be valid");
console.log(`ETH proof is valid.`);
The script reads a STARK verifying key and an Ethereum block proof downloaded from EthProofs, and calls verify_stark_proof, which wraps StarkVerifier::verify_proof. Unlike verify_stark, it takes no separate public values: it checks the proof against the verifying key only, and the block's public values are read from the proof. The proof and the verifying key must come from the same Ziren release as the verifier.
The EthProofs project uses a modified version of the WASM verifier, published as an npm package: @ethproofs/ziren-wasm-stark-verifier.
For the Fibonacci example with n = 1000, the script prints n: 1000, a: 5965 and b: 3651, the verification time, and whether each proof is valid.
The WASM verifier embeds the verifying keys of one Ziren release, so it verifies only proofs of that release.
no_std Verification
Because zkm-verifier is no_std, it can run where the Rust standard library is unavailable:
- in resource-constrained or bare-metal environments;
- inside the zkVM, as a guest program.
A verifier guest reads a proof, its public values and the program verifying key hash from the input stream and verifies the proof. Proving that guest yields a proof of verification. The groth16 example executes such a guest on a Groth16 proof of the Fibonacci program.
The host generates the Fibonacci proof and executes the verifier guest on it:
//! A script that generates a Groth16 proof for the Fibonacci program, and verifies the //! Groth16 proof in ZKM. use zkm_sdk::{include_elf, utils, HashableKey, ProverClient, ZKMStdin}; /// The ELF for the Groth16 verifier program. const GROTH16_ELF: &[u8] = include_elf!("groth16-verifier"); /// The ELF for the Fibonacci program. const FIBONACCI_ELF: &[u8] = include_elf!("fibonacci"); /// Generates the proof, public values, and vkey hash for the Fibonacci program in a format that /// can be read by `zkm-verifier`. /// /// Returns the proof bytes, public values, and vkey hash. fn generate_fibonacci_proof() -> (Vec<u8>, Vec<u8>, String) { let n = 20u32; let mut stdin = ZKMStdin::new(); stdin.write(&n); let client = ProverClient::new(); let (pk, vk) = client.setup(FIBONACCI_ELF); println!("vk: {:?}", vk.bytes32()); let proof = client.prove(&pk, stdin).groth16().run().unwrap(); (proof.bytes().expect("the proof has a byte encoding"), proof.public_values.to_vec(), vk.bytes32()) } fn main() { utils::setup_logger(); let (fibonacci_proof, fibonacci_public_values, vk) = generate_fibonacci_proof(); let mut stdin = ZKMStdin::new(); stdin.write_vec(fibonacci_proof); stdin.write_vec(fibonacci_public_values); stdin.write(&vk); let client = ProverClient::new(); let (_, report) = client.execute(GROTH16_ELF, &stdin).run().unwrap(); println!("executed groth16 program with {} cycles", report.total_instruction_count()); println!("{}", report); }
The verifier guest:
//! A program that verifies a Groth16 proof in ZKM. #![no_main] zkm_zkvm::entrypoint!(main); use zkm_verifier::Groth16Verifier; pub fn main() { let proof = zkm_zkvm::io::read_vec(); let zkm_public_values = zkm_zkvm::io::read_vec(); let zkm_vkey_hash: String = zkm_zkvm::io::read(); let groth16_vk = *zkm_verifier::GROTH16_VK_BYTES; println!("cycle-tracker-start: verify"); let result = Groth16Verifier::verify(&proof, &zkm_public_values, &zkm_vkey_hash, groth16_vk); println!("cycle-tracker-end: verify"); match result { Ok(()) => { println!("Proof is valid"); } Err(e) => { println!("Error verifying proof: {:?}", e); } } }
Groth16Verifier::verify checks that the proof's selector matches the Groth16 verifying key, then verifies the proof against the three public inputs: the program verifying key hash, the digest of the public values, and VK_ROOT_BYTES.
BitVM Verifier
BitVM-based verification combines both off-chain and on-chain steps depending on the phase of the process. Ziren integrates with GOAT Network’s BitVM2 node to support verification in a Bitcoin-native environment, which allows ZK proofs to ultimately inherit Bitcoin’s security guarantees.
Ziren generates a proof for each L2 block, which is then ingested and stored by the BitVM2 node. These individual block proofs are recursively aggregated, and can be wrapped (for example into Groth16 format) to support other verification environments. BitVM2 implements an optimistic fraud-proof mechanism. In the default case, proofs are accepted based on off-chain checks and periodic sequencer commitments. If, however, an invalid proof is suspected, honest validators can trigger a challenge protocol. This protocol reduces the entire computation trace to a single disputed step, which is then resolved directly on Bitcoin L1. By doing so, the protocol ensures that verified proofs settled through BitVM2 achieve security equivalent to Bitcoin consensus while enabling efficient ZK-based verification.
Proofs are verified off-chain for peg-ins, peg-outs and sequencer commitments. During a peg-in, an SPV proof is generated proving that the user’s transaction of their deposit was included in a valid block. The GOAT contract and committee verify this proof off-chain. During a peg-out, Ziren generates a ZK proof that attests to the correctness of the PegBTC burn and the L2 state. This proof is initially checked off-chain by watchers, the committee, and potential challengers.
If no dispute is raised, the operator is reimbursed without requiring any further on-chain action. If a dispute is raised, the on-chain challenge process begins. The Watchtower generates the longest chain proof and verifies that the operator’s Kickoff commitment matches the canonical longest chain. If this chain-level check passes, challengers can proceed to the circuit-level dispute. At that stage, the entire execution trace is revealed and challenge protocol narrows the disagreement down to the disputed computation. The Bitcoin covenant then executes the check on-chain to verify if the step belongs to the committed state and if the state transition is either valid or invalid. As a result, direct costs are only incurred during disputes.
In addition to bridging operations, the BitVM2 protocol also requires sequencer set commitments. Periodically, the committee commits the sequencer set (sequencer public keys) to Bitcoin L1. Merkle proofs of individual sequencers can then be verified off-chain against this root, ensuring that the sequencers producing L2 blocks are consistent with the commitments.
Overall, this design makes it so that proofs are not posted to Bitcoin for every block. Although proofs are generated per block, they are stored and aggregated off-chain, and verification by default also takes place off-chain. The Bitcoin L1 is only involved in periodic sequencer commitments and in disputes, where a single step is checked on-chain.
Note that in the EVM verification based setting, every proof must be submitted and verified on-chain, with each verification incurring a gas cost. Verification is also immediate and canonical in Ethereum state. By contrast, BitVM2 verification avoids per-proof L1 costs, only escalating to Bitcoin L1 in the presence of fraud dispute, while still anchoring security in Bitcoin consensus.

BitVM transaction flow. See the full paper here.
Proof Aggregation
With aggregation, multiple proofs can be combined together into a single aggregated proof by verifying each proof generated by Ziren inside the Ziren zkVM. A toy example of using aggregation for Fibonacci proofs can be found in Ziren's examples directory here.
In this example, multiple proofs proving the execution of a Fibonacci sequence for different values of n are combined into a single higher-level “aggregated” proof. This higher-level proof proves that the collection of all the other Fibonacci individual proofs are valid.
Instead of verifying each proof one by one, a verifier only needs to check a single aggregated proof. The batching of many small computations into a single proof reduces verification costs and enables applications such as block aggregation, where many transactions in a block can be proven with one single succinct proof.
The host generates individual proofs, the guest recursively verifies them, and the final output aggregated proof can be cheaply verified.
The following is the guest program implementation in the example:
guest > main.rs:
//! A simple program that aggregates the proofs of multiple programs proven with the zkVM. #![no_main] zkm_zkvm::entrypoint!(main); use sha2::{Digest, Sha256}; pub fn main() { // Read the verification keys. let vkeys = zkm_zkvm::io::read::<Vec<[u32; 8]>>(); // Read the public values. let public_values = zkm_zkvm::io::read::<Vec<Vec<u8>>>(); // Verify the proofs. assert_eq!(vkeys.len(), public_values.len()); for i in 0..vkeys.len() { let vkey = &vkeys[i]; let public_values = &public_values[i]; let public_values_digest = Sha256::digest(public_values); zkm_zkvm::lib::verify::verify_zkm_proof(vkey, &public_values_digest.into()); } // TODO: Do something interesting with the proofs here. // // For example, commit to the verified proofs in a merkle tree. For now, we'll just commit to // all the (vkey, input) pairs. let commitment = commit_proof_pairs(&vkeys, &public_values); zkm_zkvm::io::commit_slice(&commitment); } pub fn words_to_bytes_le(words: &[u32; 8]) -> [u8; 32] { let mut bytes = [0u8; 32]; for i in 0..8 { let word_bytes = words[i].to_le_bytes(); bytes[i * 4..(i + 1) * 4].copy_from_slice(&word_bytes); } bytes } /// Encode a list of vkeys and committed values into a single byte array. In the future this could /// be a merkle tree or some other commitment scheme. /// /// ( vkeys.len() || vkeys || committed_values[0].len as u32 || committed_values[0] || ... ) pub fn commit_proof_pairs(vkeys: &[[u32; 8]], committed_values: &[Vec<u8>]) -> Vec<u8> { assert_eq!(vkeys.len(), committed_values.len()); let mut res = Vec::with_capacity( 4 + vkeys.len() * 32 + committed_values.len() * 4 + committed_values.iter().map(|vals| vals.len()).sum::<usize>(), ); // Note we use big endian because abi.encodePacked in solidity does also res.extend_from_slice(&(vkeys.len() as u32).to_be_bytes()); for vkey in vkeys.iter() { res.extend_from_slice(&words_to_bytes_le(vkey)); } for vals in committed_values.iter() { res.extend_from_slice(&(vals.len() as u32).to_be_bytes()); res.extend_from_slice(vals); } res }
The guest first reads a list of verification keys and public values for each individual proof, supplied by the host. For each verification key and public value pair, the program will:
- Hash the public values with SHA-256 to obtain a fixed-size digest.
- Call
zkm_zkvm::lib::verify::verify_zkm_proof(vkey, &public_values_digest.into())to recursively check the individual proof against its associated verification key and digest.
verify_zkm_proof is available when the guest enables the verify feature of zkm-zkvm:
[dependencies]
zkm-zkvm = { git = "https://github.com/ProjectZKM/Ziren", features = ["verify"] }
After verifying all individual proofs, the guest commits to the entire batch of verification keys and public values pairs to build a combined byte string. This commitment becomes the public output of the aggregated proof.
The following is the host program implementation in the example:
host > main.rs:
//! A simple example showing how to aggregate proofs of multiple programs with ZKM. use zkm_sdk::{ include_elf, HashableKey, ProverClient, ZKMProof, ZKMProofWithPublicValues, ZKMStdin, ZKMVerifyingKey, }; /// A program that aggregates the proofs of the simple program. const AGGREGATION_ELF: &[u8] = include_elf!("aggregation"); /// A program that just runs a simple computation. const FIBONACCI_ELF: &[u8] = include_elf!("fibonacci"); /// An input to the aggregation program. /// /// Consists of a proof and a verification key. struct AggregationInput { pub proof: ZKMProofWithPublicValues, pub vk: ZKMVerifyingKey, } fn main() { // Setup the logger. zkm_sdk::utils::setup_logger(); // Initialize the proving client. let client = ProverClient::new(); // Setup the proving and verifying keys. let (aggregation_pk, _) = client.setup(AGGREGATION_ELF); let (fibonacci_pk, fibonacci_vk) = client.setup(FIBONACCI_ELF); // Generate the fibonacci proofs. let proof_1 = tracing::info_span!("generate fibonacci proof n=10").in_scope(|| { let mut stdin = ZKMStdin::new(); stdin.write(&10); client.prove(&fibonacci_pk, stdin).compressed().run().expect("proving failed") }); let proof_2 = tracing::info_span!("generate fibonacci proof n=20").in_scope(|| { let mut stdin = ZKMStdin::new(); stdin.write(&20); client.prove(&fibonacci_pk, stdin).compressed().run().expect("proving failed") }); let proof_3 = tracing::info_span!("generate fibonacci proof n=30").in_scope(|| { let mut stdin = ZKMStdin::new(); stdin.write(&30); client.prove(&fibonacci_pk, stdin).compressed().run().expect("proving failed") }); // Setup the inputs to the aggregation program. let input_1 = AggregationInput { proof: proof_1, vk: fibonacci_vk.clone() }; let input_2 = AggregationInput { proof: proof_2, vk: fibonacci_vk.clone() }; let input_3 = AggregationInput { proof: proof_3, vk: fibonacci_vk.clone() }; let inputs = vec![input_1, input_2, input_3]; // Aggregate the proofs. tracing::info_span!("aggregate the proofs").in_scope(|| { let mut stdin = ZKMStdin::new(); // Write the verification keys. let vkeys = inputs.iter().map(|input| input.vk.hash_u32()).collect::<Vec<_>>(); stdin.write::<Vec<[u32; 8]>>(&vkeys); // Write the public values. let public_values = inputs.iter().map(|input| input.proof.public_values.to_vec()).collect::<Vec<_>>(); stdin.write::<Vec<Vec<u8>>>(&public_values); // Write the proofs. // // Note: this data will not actually be read by the aggregation program, instead it will be // witnessed by the prover during the recursive aggregation process inside Ziren itself. for input in inputs { let ZKMProof::Compressed(proof) = input.proof.proof else { panic!() }; stdin.write_proof(*proof, input.vk.vk); } // Generate the plonk bn254 proof. client.prove(&aggregation_pk, stdin).plonk().run().expect("proving failed"); }); }
In the host program, during the setup, the host will compile and load two guest programs as compiled ELF binaries: the AGGREGATION_ELF (representing the aggregation guest program) and FIBONACCI_ELF, whose corresponding guest program implementation can be found here. These programs are passed to client.setup() to generate the proving and verifying keys and to client.prove() to execute inside the zkVM and generate proofs. In this example, the host generates three compressed proofs proving the correct computation of Fibonacci for inputs n=10, 20, 30.
The host then prepares inputs corresponding to each Fibonacci proof for the aggregation guest. Specifically, the host feeds the verification key hashes and raw public values as inputs and supplies the compressed proofs and full verification keys as witness data. The guest hashes the public values to a digest and calls verify_zkm_proof(vk_hash, digest) to check each proof. After all pass, the guest commits to the batch. That commitment becomes the public output of the aggregated proof.
The individual proofs must be compressed proofs; the host passes them to the prover with stdin.write_proof(proof, vk). The executor checks each proof natively when the guest calls verify_zkm_proof, but the core proof of the aggregation program only accumulates the (verifying key, public values digest) pairs into a deferred-proofs digest. The inner proofs are verified in-circuit in the recursion stage, when the prover compresses the aggregation proof. A core proof of the aggregation program therefore does not cover the inner proofs; request a compressed, Groth16 or PLONK proof.
The aggregation program is run inside the zkVM and generates a PLONK proof (representing the aggregated proof) that certifies the validity of the three individual Fibonacci proofs.
As an overview what aggregation entails in Ziren:
- Generate individual proofs.
- Collect the verification keys and public outputs of the individual proofs.
- Inside another zkVM program (the aggregation guest), recursively verify all proofs.
- Commit to the batch as a single public commitment and generate a succinct new proof proving the correct execution of all individual proofs (the aggregated proof).
For computationally heavy applications, proving logic can be divided into multiple proofs and later aggregated into a single proof. In block-level aggregation, instead of re-executing transactions individually on-chain (which can incur high gas costs), a succinct proof attesting to the validity of all transactions in a block can be generated off-chain and verified on-chain. The aggregated proof can also be in other proof formats, such as STARK or Groth16. In addition to verification via smart contract deployment, the aggregated proof can be verified off-chain using Ziren's WASM verifier.
Precompiles
Precompiles are operations, mostly cryptographic, that Ziren proves with a dedicated chip instead of executing them as MIPS instructions. A precompile call costs one syscall instruction plus rows in the precompile's own table, far fewer cycles than the same operation compiled to MIPS code. Hashing, signature verification and pairing-based cryptography in a guest should therefore go through the precompiles, usually via a patched crate.
Within the zkVM, precompiles are invoked with the MIPS syscall instruction. Register $v0 holds the system call code, and $a0 and $a1 hold the arguments, usually pointers to the operands in memory. The result is written back in place, to the memory the first argument points to.
Specification
A system call code is a 32-bit integer with the following little-endian layout:
| Byte 0 | Byte 1 | Byte 2 | Byte 3 |
|---|---|---|---|
| ID0 | ID1 | Table | Cycles |
- Bytes 0 and 1 are the system call identifier.
- Byte 2 is 1 if the system call is proved in its own chip table, and 0 otherwise.
- Byte 3 is the number of additional cycles the system call takes, which bounds its memory accesses.
The system calls and precompiles, from crates/core/executor/src/syscalls/code.rs:
| Syscall | Code | Operation |
|---|---|---|
HALT | 0x00_00_00_00 | Halts the program with an exit code. |
WRITE | 0x00_00_00_02 | Writes to a file descriptor (stdout, stderr, hooks). |
ENTER_UNCONSTRAINED | 0x00_00_00_03 | Enters unconstrained execution. |
EXIT_UNCONSTRAINED | 0x00_00_00_04 | Exits unconstrained execution. |
COMMIT | 0x00_00_00_10 | Commits a word of the public values digest. |
COMMIT_DEFERRED_PROOFS | 0x00_00_00_1A | Commits a word of the deferred proofs digest. |
VERIFY_ZKM_PROOF | 0x00_00_00_1B | Verifies a Ziren proof (deferred to recursion). |
SYSHINTLEN | 0x00_00_00_F0 | Returns the length of the next input. |
SYSHINTREAD | 0x00_00_00_F1 | Reads the next input into memory. |
SYSVERIFY | 0x00_00_00_F2 | Verifies a Ziren proof (Go runtime). |
SHA_EXTEND | 0x30_01_00_05 | SHA-256 message schedule extension. |
SHA_COMPRESS | 0x01_01_00_06 | SHA-256 compression. |
KECCAK_SPONGE | 0x01_01_00_09 | Keccak-256 sponge over a padded input. |
POSEIDON2_PERMUTE | 0x00_01_00_30 | Poseidon2 permutation over KoalaBear. |
ED_ADD | 0x01_01_00_07 | Ed25519 point addition. |
ED_DECOMPRESS | 0x00_01_00_08 | Ed25519 point decompression. |
SECP256K1_ADD | 0x01_01_00_0A | secp256k1 point addition. |
SECP256K1_DOUBLE | 0x00_01_00_0B | secp256k1 point doubling. |
SECP256K1_DECOMPRESS | 0x00_01_00_0C | secp256k1 point decompression. |
SECP256R1_ADD | 0x01_01_00_2C | secp256r1 point addition. |
SECP256R1_DOUBLE | 0x00_01_00_2D | secp256r1 point doubling. |
SECP256R1_DECOMPRESS | 0x00_01_00_2E | secp256r1 point decompression. |
BN254_ADD | 0x01_01_00_0E | BN254 G1 point addition. |
BN254_DOUBLE | 0x00_01_00_0F | BN254 G1 point doubling. |
BN254_FP_ADD / SUB / MUL | 0x01_01_00_26 to 0x01_01_00_28 | BN254 base field operations. |
BN254_FP2_ADD / SUB / MUL | 0x01_01_00_29 to 0x01_01_00_2B | BN254 quadratic extension field operations. |
BLS12381_ADD | 0x01_01_00_1E | BLS12-381 G1 point addition. |
BLS12381_DOUBLE | 0x00_01_00_1F | BLS12-381 G1 point doubling. |
BLS12381_DECOMPRESS | 0x00_01_00_1C | BLS12-381 G1 point decompression. |
BLS12381_FP_ADD / SUB / MUL | 0x01_01_00_20 to 0x01_01_00_22 | BLS12-381 base field operations. |
BLS12381_FP2_ADD / SUB / MUL | 0x01_01_00_23 to 0x01_01_00_25 | BLS12-381 quadratic extension field operations. |
UINT256_MUL | 0x01_01_00_1D | 256-bit modular multiplication. |
U256XU2048_MUL | 0x01_01_00_2F | 256-bit by 2048-bit multiplication. |
The executor also emulates the subset of Linux MIPS system calls (codes 4000 and above, for example SYS_BRK, SYS_MMAP and SYS_READ) that the Go runtime needs.
Guest Interface
The zkm_zkvm::syscalls module implements each system call as a #[no_mangle] extern "C" function that issues the syscall instruction; a guest calls them directly, for example zkm_zkvm::syscalls::syscall_sha256_extend. The zkm-lib crate, re-exported as zkm_zkvm::lib, declares the same functions for crates that do not depend on zkm-zkvm (the patched crates use it), and provides higher-level wrappers such as zkm_zkvm::lib::keccak256::keccak256. Its declarations, from crates/zkvm/lib/src/lib.rs:
#![allow(unused)] fn main() { //! Syscalls for the Ziren zkVM. //! //! Documentation for these syscalls can be found in the zkVM entrypoint //! `zkm_zkvm::syscalls` module. pub mod bls12381; pub mod bn254; #[cfg(feature = "ecdsa")] pub mod ecdsa; pub mod ed25519; pub mod io; pub mod keccak256; pub mod poseidon2; pub mod secp256k1; pub mod secp256r1; pub mod sha3; pub mod unconstrained; pub mod utils; #[cfg(feature = "verify")] pub mod verify; extern "C" { /// Halts the program with the given exit code. pub fn syscall_halt(exit_code: u8) -> !; /// Writes the bytes in the given buffer to the given file descriptor. pub fn syscall_write(fd: u32, write_buf: *const u8, nbytes: usize); /// Reads the bytes from the given file descriptor into the given buffer. pub fn syscall_read(fd: u32, read_buf: *mut u8, nbytes: usize); /// Executes the SHA-256 extend operation on the given word array. pub fn syscall_sha256_extend(w: *mut [u32; 64]); /// Executes the SHA-256 compress operation on the given word array and a given state. pub fn syscall_sha256_compress(w: *mut [u32; 64], state: *mut [u32; 8]); /// Executes an Ed25519 curve addition on the given points. pub fn syscall_ed_add(p: *mut [u32; 16], q: *const [u32; 16]); /// Executes an Ed25519 curve decompression on the given point. pub fn syscall_ed_decompress(point: &mut [u8; 64]); /// Executes an Sepc256k1 curve addition on the given points. pub fn syscall_secp256k1_add(p: *mut [u32; 16], q: *const [u32; 16]); /// Executes an Secp256k1 curve doubling on the given point. pub fn syscall_secp256k1_double(p: *mut [u32; 16]); /// Executes an Secp256k1 curve decompression on the given point. pub fn syscall_secp256k1_decompress(point: &mut [u8; 64], is_odd: bool); /// Executes an Secp256r1 curve addition on the given points. pub fn syscall_secp256r1_add(p: *mut [u32; 16], q: *const [u32; 16]); /// Executes an Secp256r1 curve doubling on the given point. pub fn syscall_secp256r1_double(p: *mut [u32; 16]); /// Executes an Secp256r1 curve decompression on the given point. pub fn syscall_secp256r1_decompress(point: &mut [u8; 64], is_odd: bool); /// Executes a Bn254 curve addition on the given points. pub fn syscall_bn254_add(p: *mut [u32; 16], q: *const [u32; 16]); /// Executes a Bn254 curve doubling on the given point. pub fn syscall_bn254_double(p: *mut [u32; 16]); /// Executes a BLS12-381 curve addition on the given points. pub fn syscall_bls12381_add(p: *mut [u32; 24], q: *const [u32; 24]); /// Executes a BLS12-381 curve doubling on the given point. pub fn syscall_bls12381_double(p: *mut [u32; 24]); /// Executes the Keccak Sponge pub fn syscall_keccak_sponge(input: *const u32, result: *mut [u32; 17]); /// Executes the Poseidon2 permutation pub fn syscall_poseidon2_permute(state: *mut [u32; 16]); /// Executes an uint256 multiplication on the given inputs. pub fn syscall_uint256_mulmod(x: *mut [u32; 8], y: *const [u32; 8]); /// Executes a 256-bit by 2048-bit multiplication on the given inputs. pub fn syscall_u256x2048_mul( x: *const [u32; 8], y: *const [u32; 64], lo: *mut [u32; 64], hi: *mut [u32; 8], ); /// Enters unconstrained mode. pub fn syscall_enter_unconstrained() -> bool; /// Exits unconstrained mode. pub fn syscall_exit_unconstrained(); /// Defers the verification of a valid Ziren zkVM proof. pub fn syscall_verify_zkm_proof(vk_digest: &[u32; 8], pv_digest: &[u8; 32]); /// Returns the length of the next element in the hint stream. pub fn syscall_hint_len() -> usize; /// Reads the next element in the hint stream into the given buffer. pub fn syscall_hint_read(ptr: *mut u8, len: usize); /// Allocates a buffer aligned to the given alignment. pub fn sys_alloc_aligned(bytes: usize, align: usize) -> *mut u8; /// Decompresses a BLS12-381 point. pub fn syscall_bls12381_decompress(point: &mut [u8; 96], is_odd: bool); /// Computes a big integer operation with a modulus. pub fn sys_bigint( result: *mut [u32; 8], op: u32, x: *const [u32; 8], y: *const [u32; 8], modulus: *const [u32; 8], ); /// Executes a BLS12-381 field addition on the given inputs. pub fn syscall_bls12381_fp_addmod(p: *mut u32, q: *const u32); /// Executes a BLS12-381 field subtraction on the given inputs. pub fn syscall_bls12381_fp_submod(p: *mut u32, q: *const u32); /// Executes a BLS12-381 field multiplication on the given inputs. pub fn syscall_bls12381_fp_mulmod(p: *mut u32, q: *const u32); /// Executes a BLS12-381 Fp2 addition on the given inputs. pub fn syscall_bls12381_fp2_addmod(p: *mut u32, q: *const u32); /// Executes a BLS12-381 Fp2 subtraction on the given inputs. pub fn syscall_bls12381_fp2_submod(p: *mut u32, q: *const u32); /// Executes a BLS12-381 Fp2 multiplication on the given inputs. pub fn syscall_bls12381_fp2_mulmod(p: *mut u32, q: *const u32); /// Executes a BN254 field addition on the given inputs. pub fn syscall_bn254_fp_addmod(p: *mut u32, q: *const u32); /// Executes a BN254 field subtraction on the given inputs. pub fn syscall_bn254_fp_submod(p: *mut u32, q: *const u32); /// Executes a BN254 field multiplication on the given inputs. pub fn syscall_bn254_fp_mulmod(p: *mut u32, q: *const u32); /// Executes a BN254 Fp2 addition on the given inputs. pub fn syscall_bn254_fp2_addmod(p: *mut u32, q: *const u32); /// Executes a BN254 Fp2 subtraction on the given inputs. pub fn syscall_bn254_fp2_submod(p: *mut u32, q: *const u32); /// Executes a BN254 Fp2 multiplication on the given inputs. pub fn syscall_bn254_fp2_mulmod(p: *mut u32, q: *const u32); /// Reads a buffer from the input stream. pub fn read_vec_raw() -> ReadVecResult; } #[repr(C)] pub struct ReadVecResult { pub ptr: *mut u8, pub len: usize, pub capacity: usize, } }
Guest Example: syscall_sha256_extend
This guest calls the SHA-256 message schedule precompile three times:
#![no_std] #![no_main] zkm_zkvm::entrypoint!(main); use zkm_zkvm::syscalls::syscall_sha256_extend; pub fn main() { let mut w = [1u32; 64]; syscall_sha256_extend(&mut w); syscall_sha256_extend(&mut w); syscall_sha256_extend(&mut w); }
Patched Crates
Patching a crate means replacing the implementation of a specific interface within the crate with a call to the corresponding zkVM precompile, which reduces the number of cycles, and therefore the proving cost, substantially.
Supported Crates
The patches are maintained in the ziren-patches organization. The following are used by the examples in the Ziren repository (examples/Cargo.toml):
| Crate Name | Repository | Versions |
|---|---|---|
| sha2 | sha2-v0-10-8 = { git = "https://github.com/ziren-patches/RustCrypto-hashes", package = "sha2", branch = "patch-sha2-0.10.8" } | 0.10.8 |
| curve25519-dalek | curve25519-dalek = { git = "https://github.com/ziren-patches/curve25519-dalek", branch = "patch-4.1.3" } | 4.1.3 |
| secp256k1 | secp256k1 = { git = "https://github.com/ziren-patches/rust-secp256k1", branch = "patch-0.29.1" } | 0.29.1 |
| substrate-bn | substrate-bn = { git = "https://github.com/ziren-patches/bn", branch = "patch-0.6.0" } | 0.6.0 |
| rsa | rsa = { git = "https://github.com/ziren-patches/RustCrypto-RSA.git", branch = "patch-rsa-0.9.6" } | 0.9.6 |
| ecdsa | ecdsa-core = { git = "https://github.com/ziren-patches/signatures", package = "ecdsa", branch = "patch-ecdsa-0.16.9" } | 0.16.9 |
| k256 | k256 = { git = "https://github.com/ziren-patches/elliptic-curves", branch = "patch-k256-0.13.4" } | 0.13.4 |
| p256 | p256 = { git = "https://github.com/ziren-patches/elliptic-curves", branch = "patch-p256-0.13.2" } | 0.13.2 |
The following are used by the Ethereum block prover reth-processor (patched in bin/guest/Cargo.toml; kzg-rs is a workspace dependency in the root Cargo.toml), in addition to substrate-bn, k256 and p256 above:
| Crate Name | Repository | Versions |
|---|---|---|
| sha2 | sha2 = { git = "https://github.com/ziren-patches/RustCrypto-hashes", branch = "patch-sha2-0.10.9", package = "sha2" } | 0.10.9 |
| alloy-primitives | alloy-primitives-v1-4-1 = { git = "https://github.com/ziren-patches/core.git", package = "alloy-primitives", branch = "patch-alloy-primitives-1.4.1" } | 1.4.1 |
| kzg-rs | kzg-rs = { git = "https://github.com/ziren-patches/kzg-rs", branch = "patch-0.2.7", default-features = false } | 0.2.7 |
Using Patched Crates
There are two approaches to using patched crates.
Option 1: add the patched crate directly as a dependency in the guest program's Cargo.toml. For example:
[dependencies]
sha2 = { git = "https://github.com/ziren-patches/RustCrypto-hashes.git", package = "sha2", branch = "patch-sha2-0.10.8" }
Option 2: keep the crates.io dependency and add a patch entry to the guest's Cargo.toml. This also replaces the crate when it is a transitive dependency. For example:
[dependencies]
sha2 = "0.10.8"
[patch.crates-io]
sha2 = { git = "https://github.com/ziren-patches/RustCrypto-hashes.git", package = "sha2", branch = "patch-sha2-0.10.8" }
When the crate comes from a git repository rather than crates.io, the patch section must name that source repository. For example:
[dependencies]
ed25519-dalek = { git = "https://github.com/dalek-cryptography/curve25519-dalek" }
[patch."https://github.com/dalek-cryptography/curve25519-dalek"]
ed25519-dalek = { git = "https://github.com/ziren-patches/curve25519-dalek", branch = "patch-4.1.3" }
A patch only takes effect if its version matches the version Cargo resolves. Check that the patched crate appears in Cargo.lock with the ziren-patches source; Cargo prints a warning for patches it did not use.
How to Patch a Crate
First, the operation must exist as a zkVM precompile (for example syscall_keccak_sponge), with its chip and constraints in Ziren. The available precompiles are listed on the Precompiles page; since a new precompile needs circuit work, open an issue to request one.
Then replace the crate's implementation with a call to the precompile, under #[cfg(target_os = "zkvm")]. For example, Ziren implements keccak256 with syscall_keccak_sponge, and the patched alloy-primitives uses it for keccak256:
#![allow(unused)] fn main() { if #[cfg(target_os = "zkvm")] { let output = zkm_zkvm::lib::keccak256::keccak256(bytes); B256::from(output) } }
Finally, patch the crate in the guest as shown above, as reth-processor does.
Optimizations
There are several ways to reduce the cost of proving a program:
- measure where the cycles go with cycle tracking, and optimize those parts;
- route cryptographic operations through precompiles, directly or via patched crates;
- prove on a GPU, or enable AVX on the CPU prover;
- avoid unnecessary work in the guest, such as copying data or serializing and deserializing it more often than needed.
Testing Your Program
Test your program and check its outputs before generating proofs; execution is much faster than proving.
To execute your program without generating a proof, call ProverClient::execute from the host:
#![allow(unused)] fn main() { let client = ProverClient::new(); let (_, report) = client.execute(ELF, &stdin).run().unwrap(); println!("executed program with {} cycles", report.total_instruction_count()); }
execute returns the public values the program committed with zkm_zkvm::io::commit and an ExecutionReport, which holds the instruction count per opcode (opcode_counts), the system call count per system call (syscall_counts) and the cycle tracker results (cycle_tracker). The report implements Display, so println!("{}", report) prints all of them.
Acceleration Options
Acceleration via Precompiles
Precompiles are dedicated chips for common cryptographic operations, such as SHA-256, Keccak-256, elliptic curve arithmetic over secp256k1, secp256r1, Ed25519, BN254 and BLS12-381, and 256-bit modular multiplication. A precompile call costs far fewer cycles than the same operation compiled to MIPS instructions.
A guest can call the precompiles directly with system calls. The Precompiles page lists them and has an example guest program.
Alternatively, use the patched crates, which replace the implementation of common crates (sha2, k256, p256, substrate-bn, and others) with precompile calls, so that existing code uses the precompiles without changes.
The Ethereum block prover reth-processor is an example; note the patch entries for sha2, bn, k256, p256 and alloy-primitives in its guest's Cargo.toml.
Acceleration via Hardware
Ziren supports hardware acceleration for proof generation on both GPU and CPU:
- a CUDA-based GPU prover, selected with the
ZKM_PROVER=cudaenvironment variable or theProverClient::cuda()constructor; - AVX2/AVX512 optimizations on x86 CPUs via Plonky3, enabled through
RUSTFLAGS.
For setup and examples, see the Prover page.
Cycle Tracking
Cycle counts show where a program spends its execution and which parts to optimize. More cycles mean longer proving. Proving cost more precisely follows the number of rows the execution fills across all chip tables, and precompile calls and memory accesses add rows of their own; the cycle count is a good first proxy.
The guest marks a region with cycle-tracker-start and cycle-tracker-end lines printed to stdout, or a whole function with the #[zkm_derive::cycle_tracker] attribute (from the zkm-derive crate). The executor then logs the cycles each region takes. With cycle-tracker-report-start and cycle-tracker-report-end, it also stores the count in the execution report under the region's name.
The cycle-tracking example has two guest programs. normal.rs logs the cycles of its regions:
#![no_main] zkm_zkvm::entrypoint!(main); #[zkm_derive::cycle_tracker] pub fn expensive_function(x: usize) -> usize { let mut y = 1; for _ in 0..100 { y *= x; y %= 7919; } y } pub fn main() { let mut nums = vec![1, 1]; println!("cycle-tracker-start: setup"); for _ in 0..100 { let mut c = nums[nums.len() - 1] + nums[nums.len() - 2]; c %= 7919; nums.push(c); } println!("cycle-tracker-end: setup"); println!("cycle-tracker-start: main-body"); for i in 0..2 { let result = expensive_function(nums[nums.len() - i - 1]); println!("result: {}", result); } println!("cycle-tracker-end: main-body"); }
report.rs uses cycle-tracker-report-start: setup and cycle-tracker-report-end: setup instead, and the host reads the result from the report:
#![allow(unused)] fn main() { let (_, report) = client.execute(REPORT_ELF, &ZKMStdin::new()).run().expect("proving failed"); let setup_cycles = report.cycle_tracker.get("setup").unwrap(); }
Run the example with RUST_LOG=info from examples/cycle-tracking/host to see the logged cycle counts:
RUST_LOG=info cargo run --release
Design
Ziren is a zkVM for the MIPS32r2 instruction set. It proves that a guest program, compiled to a MIPS ELF, ran to completion on given inputs and produced given public outputs. This section describes Ziren V2.0: the machine that is proved, the proof system that proves it, and the recursion that turns many shard proofs into one small proof.
-
The machine
The execution trace is split into chips: fixed-width tables over the KoalaBear field \(p = 2^{31} - 2^{24} + 1\). There is no central CPU table. Every executed instruction occupies one row in the chip of its opcode family (AddSub, Branch, LoadWord, ...), and each such row carries an instruction frame: the program fetch, the three register accesses and the
(clk, pc)hand-off to the next instruction. Constraints are local to one row and of degree at most three; every relation between rows or between chips is an interaction on a named bus. Memory, byte and range lookups, syscalls and precompiles are chips on the same buses. See State Machine and Arithmetization. -
The shard argument
A long execution is cut into shards, each proved on its own. A shard proof commits to all of the shard's traces with one jagged polynomial commitment, proves every bus balance with one LogUp-GKR argument and every chip's constraints with one zerocheck, and opens the committed columns with WHIR. See Lookup Arguments and STARK Protocol.
-
Memory across shards
Within a shard, memory is checked with timestamped read/write tuples on the memory bus. Across shards, every value a shard leaves for a later one is sent as a point on an elliptic curve over the degree-7 extension of KoalaBear, and the sum of all shards' points must be the identity. See Memory Consistency Checking and Continuation.
-
Recursion and the SNARK
A recursion machine verifies shard proofs inside its own programs: a normalize (leaf) program verifies one shard, compose programs merge up to three adjacent results, and a shrink and a wrap step prepare the final proof for a BN254 SNARK. The wrap proof is verified in a Groth16 or PLONK circuit for on-chain use. See Prover Architecture.
-
Proof composition
A guest can verify other Ziren proofs through the
verify_zkm_proofsyscall; those proofs are checked by the recursion tree as deferred proofs. See Proof Composition.
State Machine
The Ziren state machine is the MIPS32r2 machine as a set of chips (tables) connected by buses. A chip is a fixed-width table over KoalaBear whose constraints are local to one row. Everything that relates two rows, or two chips, is an interaction: a row sends or receives a tuple on a named bus, and the lookup argument proves that on every bus the sent and received multisets are equal.
The chip set is the MipsAir enum in crates/core/machine/src/mips/mod.rs. The families are:
-
Program
A preprocessed table of the program image: one row per instruction, pinning
pc = pc_base + 4·indexand the decoded instruction. Instruction rows fetch their instruction from it on the program bus. -
Instruction chips
One chip per opcode family, each row one executed instruction:
- ALU:
AddSub,AddSubImm,Bitwise,BitwiseImm,Mul,DivRem,Lt,LtImm,CloClz,ShiftLeft,ShiftLeftImm,ShiftRight,ShiftRightImm(see ALU); - control flow:
Branch,Jump(see Flow Control); - memory instructions:
LoadWord,LoadNarrow,StoreWord,StoreNarrow,MemoryUnaligned(see Memory); - others:
MovCond(MOVZ/MOVN and WSBH),MiscInstrs(INS, EXT, SEB/SEH, MADD/MADDU/MSUB/MSUBU, TEQ) andSyscallInstrs(thesyscallinstruction, including halt).
Each instruction chip embeds the instruction frame described in CPU: there is no separate CPU table.
- ALU:
-
Memory
MemoryLocal,MemoryGlobalInit,MemoryGlobalFinalandMemoryBump(see Memory and Memory Consistency Checking). -
Global
The
Globalchip turns every message that crosses a shard boundary (memory hand-offs, syscalls sent to a precompile shard) into a point on a septic elliptic curve and accumulates the points into the shard's digest. -
Lookup tables
ByteLookup, a preprocessed table of byte operations over all byte pairs (AND, OR, XOR, NOR, SLL, LTU, MSB, shift-carry, u8 and u16 range checks), andRangeLookup, a preprocessed table of(a, bits)fora < 2^bits,bits ≤ 10. -
Syscalls and precompiles
SyscallCorerecords each syscall made in an execution shard,SyscallPrecompilereceives it in the precompile shard that proves it, andSysLinuxcomputes the results of the supported Linux syscalls. Precompile chips cover SHA-256 extend and compress, the Keccak sponge, the Poseidon2 permutation, Weierstrass and Edwards curve operations, BN254 and BLS12-381 field arithmetic and 256-bit multiplication. SHA-256 and Keccak, whose operations span several rows, split into a worker chip and a control chip chained on thePrecompileChainbus. See Precompiles.
The state bus
Instead of a transition constraint between consecutive rows, the machine's control state travels on the State bus. Every instruction row receives (shard, clk, pc, next_pc) and sends (shard, clk + 5 + extra, next_pc, next_next_pc), where extra is the number of additional cycles a syscall takes. The shard's public values send the initial state and receive the final one. Because each real row receives exactly one tuple and sends exactly one with a strictly larger clk, bus balance forces the rows to form a single chain from the shard's initial state to its final state, in any order in the tables.
Two program counters travel on the bus because MIPS has a branch delay slot: next_pc is the address of the instruction after the current one, and next_next_pc the one after that, which a taken branch or a jump sets to its target.
Buses
| Bus | Carries |
|---|---|
Program | (pc, instruction) fetches against the program table |
State | (shard, clk, pc, next_pc) from one instruction to the next |
Memory | timestamped (shard, clk, addr, value) accesses |
Byte, Range | byte-operation and range-check lookups |
Syscall, SyscallResult | syscall arguments and results between the syscall chips |
Global | messages that cross shards, consumed by the Global chip |
GlobalAccumulation | the running sum of the Global chip's points |
MemoryGlobalInitControl, MemoryGlobalFinalizeControl | the ordered chain of initialized and finalized addresses |
PrecompileChain | the per-row state of multi-row precompiles |
The State, GlobalAccumulation and the two global-memory control buses are closed by endpoints the shard's public values supply; every other bus balances within the shard's traces.
CPU
Ziren V2.0 has no CPU chip. The work a CPU table used to do (fetch the instruction, read and write registers, advance the clock and the program counter) is done by an instruction frame: a fixed group of columns and constraints that every instruction chip embeds in its rows. An executed instruction therefore costs one row in the chip of its opcode family and nothing else. The frame is defined in crates/core/machine/src/frame/mod.rs.
Frame columns
InstructionFrameCols has:
shard: the shard number, range-checked to 16 bits.clk_16bit_limb,clk_high_limb: the per-shard clock \( clk = clk_{16} + 2^{16} \cdot clk_{high} \). The low limb is range-checked to 16 bits and the high limb to 10 bits, so timestamps are 26-bit values.instruction: the decoded instruction (opcode,op_a,op_b,op_c, the immediate flagsimm_b,imm_c, andop_a_0, which marks a write to$zero).op_a_access,op_b_access,op_c_access: the three register accesses.op_ais a read-write access (previous and new value),op_bandop_care reads.
Jump, MovCond and MiscInstrs use this full frame. Narrower variants save columns where the instruction format allows it:
ITypeFrameCols:op_aandop_bregisters and an immediateop_cword. Used byAddSubImm,BitwiseImm,LtImm,CloClz,Branchand the memory-instruction chips.RTypeFrameCols: three registers, with the opcode and operand indices stored as single columns. Used byAddSub,Bitwise,Lt,Mul,DivRem,ShiftLeft,ShiftRightandSyscallInstrs.ShamtFrameCols:op_aandop_bregisters and a 5-bit shift amount held in one column. Used byShiftLeftImmandShiftRightImm.
Each register access holds the value, the clock of the previous access to that register, and one 16-bit limb of the clock difference: six columns. The previous access is always in the same shard because the MemoryBump chip inserts a read of every touched register at (shard, 0) (see Memory).
Frame constraints
eval_instruction_frame constrains, for a real row:
- Program fetch. The row sends
(pc, instruction)on theProgrambus, so the instruction must be the one the preprocessed program table stores atpc. The instruction's opcode must equal the chip's opcode, which the chip computes from its own selector columns. - Clock. The shard and the two clock limbs are range-checked.
- Immediates. If
imm_b(orimm_c) is set, the operand value equals the immediate in the instruction and no register is read. - Registers. The accesses happen at fixed sub-cycle offsets:
op_cat \( clk + 1 \),op_bat \( clk + 2 \),op_aat \( clk + 3 \). A memory access, where there is one, uses \( clk + 0 \), and theHIregister written byMul,DivRemand the multiply-accumulate instructions ofMiscInstrsuses \( clk + 4 \). Ifop_a_0is set, the written value is zero. The written value's four bytes are range-checked. - State hand-off. The row receives
(shard, clk, pc, next_pc)on theStatebus and sends(shard, clk + 5 + extra, next_pc, next_next_pc), whereextrais the number of extra cycles a syscall takes (zero for every other chip).
The chip supplies pc, next_pc and next_next_pc as expressions. pc and next_pc are whatever the row receives from its predecessor; in particular a delay-slot instruction receives a branch target as its next_pc. A sequential chip only fixes \( next\_next\_pc = next\_pc + 4 \). The branch and jump chips carry next_pc and next_next_pc as range-checked words and compute the target. The halt row of the syscall chip receives its predecessor's pc + 4 as next_pc and sends next_pc = 0, the exit marker.
Because every instruction advances clk by at least 5 and consumes one State tuple while producing one, balancing the State bus forces the rows of all instruction chips to form one chain from the shard's initial state to its final state (see State Machine).
Program counters and the delay slot
MIPS executes the instruction after a branch or jump (the delay slot) before the target. The frame carries two program counters:
next_pcis the address of the next instruction to execute; for a branch or jump this is the delay slot.next_next_pcis the address after that. For sequential instructions it isnext_pc + 4; a branch or jump sets it to the taken target or the fall-through address.
The delay-slot instruction receives (next_pc, next_next_pc) and passes the target on as its own next_pc.
The shard's public values carry both pairs, (start_pc, start_next_pc) and (next_pc, next_next_pc), but a pending branch target never crosses a shard boundary. The executor never closes a shard in front of a delay slot, and the recursion leaf enforces it: for a shard that executes instructions it asserts start_next_pc = start_pc + 4 and next_next_pc = next_pc + 4, so both boundaries are sequential. The recursion then chains only pc: each shard's start_pc must equal the previous shard's next_pc.
Memory
Registers and memory share one address space and one memory argument. The 36 registers (the 32 general-purpose registers, LO, HI and two internal registers for the program break and heap) occupy addresses 0 to 35, and memory instructions may not address them. Every access is a pair of tuples on the Memory bus: the accessing row sends the previous access (prev_shard, prev_clk, addr, prev_value) and receives the current one (shard, clk, addr, value), and asserts that the current timestamp is strictly larger. The argument is described in Memory Consistency Checking. This page describes the chips.
The memory chips are the five memory-instruction chips, MemoryBump, MemoryLocal, MemoryGlobalInit and MemoryGlobalFinal. The sources are in crates/core/machine/src/memory/.
Memory instructions
The load and store opcodes are split by width and direction so that each chip carries only the columns its opcodes need:
| Chip | Opcodes |
|---|---|
LoadWord | LW, LL |
LoadNarrow | LB, LBU, LH, LHU |
StoreWord | SW, SC |
StoreNarrow | SB, SH |
MemoryUnaligned | LWL, LWR, SWL, SWR |
Each chip embeds the I-type instruction frame and a common block that constrains:
- The effective address \( addr = op_b + op_c \), computed inline with an addition gadget (value and carries).
- That the address word is a canonical KoalaBear value with byte-checked limbs, and that \( addr \ge 36 \), so a memory instruction cannot touch a register.
- The two low address bits, and the memory access at the aligned address \( addr - (addr \bmod 4) \), at timestamp \( clk + 0 \).
- \( next\_next\_pc = next\_pc + 4 \): memory instructions are sequential.
Per chip:
LoadWordandStoreWordpin the low address bits to zero. A load's value isop_a; a store's memory value isop_a.SCstores the previous value ofop_aand setsop_ato 1. The machine is single-threaded, so a store-conditional always succeeds andLLis an ordinary load.LoadNarrowselects the byte or half-word with three offset flags and extends it. For a signed load whose top bit is set, the bytes above the loaded value are0xFF. This is a byte constraint, not an ALU lookup.StoreNarrowuses the offset flags to replace one byte or half-word of the previous memory word.MemoryUnalignedcombines the memory word with the previous value ofop_aaccording to the offset, asLWL/LWR/SWL/SWRspecify.
The memory-instruction chips do not range-check the bytes of the loaded or stored word. A stored word comes from a register, and a loaded word goes to one. Values written into memory from any other source (global initialization, precompiles) are byte-checked where they enter.
MemoryBump
A register access compares only clocks, not shards. To make that sound, MemoryBump has one row per (register, shard) for every register the shard touches: a shadow read of the register at timestamp (shard, 0). Its previous access may be in any earlier shard, so this row uses the full access columns with a shard comparison. Because clk restarts at 0 in every shard and real register accesses sit at \( clk + 1 \) to \( clk + 4 \), the shadow read is always the first access of the shard to that register. Every other register access then has prev_shard = shard. This reduces each register access to six columns (value, previous clock and one 16-bit limb of the clock difference) instead of nine.
MemoryLocal
MemoryLocal has one row per address the shard accesses. The row opens and closes the shard's chain for that address:
- It receives the initial tuple
(initial_shard, initial_clk, addr, initial_value)on theMemorybus, where the shard's first access sends it as its "previous" access. It also sends the same tuple to theGlobalbus as a receive message. - It sends the final tuple
(final_shard, final_clk, addr, final_value), matching the shard's last access, and sends it to theGlobalbus as a send message.
The initial side is otherwise a free witness, so the row range-checks both shards to 16 bits, both clocks to 26 bits (a 16-bit and a 10-bit limb) and every value limb to a byte.
On the Global bus, a shard's final value for an address cancels against the next accessing shard's initial value (see Memory Consistency Checking).
MemoryGlobalInit and MemoryGlobalFinal
These chips cover the lifetime of an address across the whole execution:
MemoryGlobalInithas a row for each address the program touches that is not in the program image. It sends(0, 0, addr, value)on theGlobalbus with its initial value.MemoryGlobalFinalhas a row for each touched address. It receives the last(shard, clk, addr, value)from theGlobalbus.
Initial values of the program image do not have rows. They are folded into the verifying key as initial_global_cumulative_sum, which the leaf of the first shard adds to the global sum.
Each address may be initialized and finalized once. Both chips enforce it the same way:
- The value is witnessed as 32 bits and the address is bit-decomposed.
- Consecutive rows are chained on a control bus (
MemoryGlobalInitControlorMemoryGlobalFinalizeControl) carrying(index, addr, valid). - Each row asserts
prev_addr < addrwith a 32-bit comparison.
The chain continues across shards: each shard's first row receives the previous shard's last address from the public values (previous_init_addr_bits, last_init_addr_bits and the finalize equivalents), and the leaf checks that the ranges of consecutive shards join. The only row exempt from the comparison is the genesis row, index 0 with previous address 0, which initializes and finalizes address 0 to zero.
ALU
Each arithmetic and logic opcode family has its own chip, and most families have a second chip for the immediate form. A row is one executed instruction: it embeds an instruction frame (see CPU), which reads the operands and writes the result register, and adds only the columns needed to prove that result. The ALU chips are sequential: each sends \( next\_next\_pc = next\_pc + 4 \). There is no ALU bus: an ALU row proves its own result and receives nothing from other instruction chips. The sources are in crates/core/machine/src/alu/.
| Chip | Frame | Opcodes (MIPS instructions) |
|---|---|---|
AddSub | R-type | ADD (ADD, ADDU), SUB (SUB, SUBU) |
AddSubImm | I-type | ADD, SUB with an immediate (ADDI, ADDIU, LUI) |
Bitwise | R-type | AND, OR, XOR, NOR |
BitwiseImm | I-type | AND, OR, XOR with an immediate (ANDI, ORI, XORI) |
Lt | R-type | SLT, SLTU |
LtImm | I-type | SLT, SLTU with an immediate (SLTI, SLTIU) |
Mul | R-type | MUL, MULT, MULTU |
DivRem | R-type | DIV, DIVU, MOD, MODU |
CloClz | I-type | CLO, CLZ |
ShiftLeft | R-type | SLL by a register (SLLV) |
ShiftLeftImm | shift-amount | SLL by an immediate |
ShiftRight | R-type | SRL, SRA, ROR by a register (SRLV, SRAV, ROTRV) |
ShiftRightImm | shift-amount | SRL, SRA, ROR by an immediate |
The decoder also maps some instructions onto these opcodes: MFHI, MTHI, MFLO and MTLO become ADD with HI (register 33) or LO (register 32) as an operand and zero as the other, and LUI becomes ADD of $zero and the shifted immediate.
Techniques
- Addition and subtraction.
AddSubdoes not witness carries. With byte limbs \( a_i, b_i, c_i \), the carry out of limb \( i \) is the expression \( carry_i = (b_i + c_i - a_i + carry_{i-1}) \cdot 256^{-1} \), and the chip asserts that each carry is boolean. The bytes of \( a \) are range-checked by the frame. Subtraction is checked as the addition \( a + c = b \). See Arithmetization for the full column layout. - Bitwise. One byte lookup per byte,
(op, a_i, b_i, c_i), against the byte table. The lookups are skipped when the destination is$zero. - Comparison.
Ltlocates the most significant byte where the operands differ, using one-hot byte flags and an inverse witness, and compares that byte pair with theLTUbyte lookup. ForSLTthe top byte is first masked to 7 bits with anAND 0x7flookup, and the result is combined with the two sign bits. - Multiplication. Both operands are extended to 64 bits (sign-extended for
MULT) and the product is checked limb by limb with witnessed carries.MULwrites the low word to the destination;MULTandMULTUwrite the low word toLOand the high word toHI, the second write at \( clk + 4 \). - Division.
DivRemwitnesses the quotient and remainder and checks \( b = c \cdot q + r \) in 64-bit arithmetic with an inline multiplication. It asserts that \( r \) has the sign of \( b \) and that \( |r| < |c| \), and handles the overflow case \( b = -2^{31}, c = -1 \). Division by zero traps in the executor, and the chip rejects \( c = 0 \).DIVandDIVUwrite the quotient toLOand the remainder toHI;MODandMODUwrite the remainder to the destination register. - Shifts. The 5-bit shift amount is split into a byte shift \( \lfloor c/8 \rfloor \) and a bit shift \( c \bmod 8 \). Left shifts multiply by \( 2^{c \bmod 8} \) and move bytes. Right shifts use the
ShrCarrybyte lookup for the bit part and extend the operand to 64 bits so thatSRAfills with the sign bit;RORalso feeds the shifted-out bits back in at the top. - Leading bits.
CLZwitnesses the result \( n \) and checks it with a right shift: \( b = 0 \) gives \( n = 32 \), and otherwise \( b \gg (31 - n) = 1 \).CLOis computed asCLZof \( \lnot b \).
Flow Control
Branches and jumps are proved by the Branch and Jump chips (crates/core/machine/src/control_flow/). Both set the post-delay-slot address next_next_pc that the instruction frame sends on the State bus. The instruction in the delay slot at next_pc runs first and passes that address on as its own next_pc (see CPU). Both chips carry next_pc and next_next_pc as 32-bit words and range-check them to canonical KoalaBear values. Neither chip can end a shard: the executor never closes a shard before a delay slot, and the recursion leaf requires sequential shard boundaries.
Branch chip
Branch proves BEQ, BNE, BLTZ, BLEZ, BGTZ and BGEZ. It uses the I-type frame. op_a and op_b are the compared registers; for the compare-with-zero branches op_b is $zero. op_c is the sign-extended offset shifted left by two.
Columns, besides the frame: pc, next_pc and next_next_pc with their range checkers, an addition gadget target_add, the six opcode selectors, is_branching, the equality witnesses (eq_lo, eq_hi and their inverses, a_eq_b), msb_a and a_gt_0.
Constraints:
- Selectors. The six selectors are boolean and their sum is
is_real, so a real row has exactly one. The opcode passed to the frame is their weighted sum. The registerop_ais only read: its written value equals its previous value. - Next address. If
is_branching, \( next\_next\_pc = next\_pc + op_c \), proved bytarget_add. Otherwise \( next\_next\_pc = next\_pc + 4 \).is_branchingis boolean and zero on padding rows. - Equality. With \( d_{lo} \) and \( d_{hi} \) the differences of the low and high 16-bit halves of
op_aandop_b, the chip checks \( eq_{lo} \cdot d_{lo} = 0 \) and \( eq_{lo} = 1 - d_{lo} \cdot inv_{lo} \) (and the same for the high half), so \( a\eq\b = eq{lo} \cdot eq{hi} \) is 1 exactly when the registers are equal. - Sign. For the compare-with-zero branches,
msb_ais the top bit ofop_a, taken from theMSBbyte lookup, and \( a\_gt\_0 = (1 - msb\_a)(1 - a\_eq\_b) \). - Condition.
is_branchingmust equal the branch condition:a_eq_bforBEQ, its negation forBNE,msb_aforBLTZ, its negation forBGEZ,a_gt_0forBGTZand its negation forBLEZ.
Jump chip
Jump proves the three jump opcodes the decoder produces:
| Opcode | MIPS instructions | Target |
|---|---|---|
Jump | JR, JALR | the register op_b |
Jumpi | J, JAL | the immediate op_b, the 26-bit index shifted left by two |
JumpDirect | BAL | \( next\_pc + op_b \), with op_b the shifted offset |
The chip uses the full instruction frame. Its columns are pc, next_pc and next_next_pc with range checkers, an addition gadget target_add, a range checker for the link value, and the selectors is_jump, is_jumpi and is_jumpdirect.
Constraints:
- Selectors. The selectors are boolean and their sum is
is_real. The opcode passed to the frame is their weighted sum. - Link. Unless the destination is
$zero(op_a_0), the value written toop_ais \( next\_pc + 4 \), the address after the delay slot. The value is range-checked to a canonical word.JRandJdecode with$zeroas the destination, so they link nothing. - Target. For
JumpandJumpi, \( next\_next\_pc = op_b \). ForJumpDirect,target_addproves \( next\_next\_pc = next\_pc + op_b \).
Other Components
Besides the ALU, flow-control and memory chips, the core machine has the program table, two lookup tables, the chips for the remaining instructions, the syscall and precompile chips, and the Global chip. The recursion machine that verifies shard proofs is a separate machine with its own chips; it is described in Recursive STARK.
Program chip
A preprocessed table with one row per instruction of the program image. The preprocessed columns are pc and the decoded instruction (opcode, op_a, op_b, op_c, imm_b, imm_c, op_a_0), and the only main column is a multiplicity. Every instruction row sends (pc, instruction) on the Program bus and the program chip receives it with the multiplicity. The preprocessed commitment is part of the verifying key, so a prover can only execute instructions of the committed program at the addresses where they were loaded.
Lookup tables
ByteLookupis preprocessed over all pairs of bytes, \( 2^{16} \) rows. For each pair it lists the results ofAND,OR,XOR,NOR,SLL,ShrCarryandLTU, theMSBof a byte, and theU8RangeandU16Rangechecks, with one multiplicity column per operation. Lookups have the form(op, a, b, c).RangeLookupis preprocessed with the pairs(a, bits)for \( a < 2^{bits} \) and \( bits \le 10 \). Its main use is the 10-bit high limb of 26-bit timestamps.
Other instruction chips
MovCondproves the conditional movesMOVZandMOVN(opcodesMEQandMNE) and the byte swapWSBH.MiscInstrsprovesINS,EXT, the sign extensionsSEBandSEH(opcodeSEXT), the multiply-accumulate instructionsMADD,MADDU,MSUBandMSUBU, which read and writeHIandLO, and the trapTEQ.SyscallInstrsproves thesyscallinstruction. The syscall number and two arguments are the registers$v0,$a0and$a1. The row:- sends the call on the
Syscallbus when a precompile or a Linux syscall must prove it; - advances
clkby the call's extra cycles; - writes the result to
$v0; - handles
COMMITandCOMMIT_DEFERRED_PROOFSby checking the committed word against the shard's publiccommitted_value_digestordeferred_proofs_digest; - on
HALTorexit_group, setsnext_pcto 0 and requires the exit code to equal the publicexit_code.
- sends the call on the
Syscalls and precompiles
-
SyscallCoreandSyscallPrecompileconnect execution shards to precompile shards. Precompile events are proved in separate shards, so a syscall made in an execution shard is sent as aGlobalmessage bySyscallCoreand received bySyscallPrecompilein the shard that proves it. -
SysLinuxproves the results of the supported Linux syscalls:mmap/mmap2,brk,clone,exit_group,fcntl,readandwrite. -
The precompile chips each prove one operation over memory:
- hashing:
Sha256ExtendandSha256Compress(each with a control chip),KeccakSponge(with a control chip) andPoseidon2Permute; - elliptic curves:
Ed25519Add,Ed25519Decompress, the add, double and decompress chips for secp256k1, secp256r1 and BLS12-381, and add and double for BN254; - field arithmetic: base-field
Fp,Fp2MulandFp2AddSubfor BN254 and BLS12-381; - big integers:
Uint256Mul(multiplication modulo a 256-bit modulus) andU256x2048Mul.
A precompile row receives the call from the
Syscallbus with its(shard, clk), reads and writes memory at that timestamp through the memory argument, and range-checks every word it writes. The list of syscalls is in MIPS ISA, and the guest interface in Precompiles. - hashing:
Global chip
The Global chip receives every message on the Global bus, that is, every value that must travel between shards: the initial and final value of each memory address a shard touches (from MemoryLocal, MemoryGlobalInit and MemoryGlobalFinal) and each syscall sent to a precompile shard. A message is seven field elements plus a kind and a send or receive flag. The chip maps it to a point on an elliptic curve over the degree-7 extension of KoalaBear and adds the points up with the GlobalAccumulation bus. The shard's resulting sum is a public value, and the recursion checks that the sums of all shards add to zero. The construction is described in Memory Consistency Checking.
Arithmetization
Ziren expresses the execution of a MIPS program as an Algebraic Intermediate Representation (AIR): a set of tables whose cells are KoalaBear field elements, together with polynomial constraints and bus interactions that the tables satisfy exactly when the execution is valid.
Key concepts
- Chips. The trace is split into chips (see State Machine). A chip is a table with a fixed number of columns and a height that depends on the execution, typically one row per event (an executed instruction, a memory address, a precompile call). Some chips also have preprocessed columns, which are fixed by the program and committed once at setup. The program table and the byte and range tables are preprocessed.
- Row constraints. Each chip has polynomial constraints over the columns of a single row, the public values and, for preprocessed chips, the row's preprocessed columns. Every constraint has degree at most 3. Constraints do not refer to the next row: there are no transition constraints and no first-row or last-row selectors.
- Interactions. A row may send or receive tuples of column expressions on named buses, each with a multiplicity expression. A shard is valid when every chip's constraints hold on every row and every bus balances. The relations that a classic AIR states between consecutive rows, such as the program counter moving from one instruction to the next, are bus interactions in Ziren (see the
Statebus in CPU). The balance of all buses is proved by the lookup argument described in Lookup Arguments. - Padding. Traces are padded with rows whose selectors are zero. Every constraint and multiplicity is gated by those selectors, so padding rows satisfy the constraints and send nothing.
The constraints are proved with a zerocheck over the multilinear extensions of the columns, not by dividing by a vanishing polynomial (see STARK Protocol).
The AddSub chip as an example
Instructions
The AddSub chip proves the register forms of addition and subtraction. The decoder maps both the trapping and the non-trapping MIPS forms to one opcode each, so the chip needs no overflow logic. The immediate forms (ADDI, ADDIU) are proved by the AddSubImm chip in the same way.
| instruction | op [31:26] | rs [25:21] | rt [20:16] | rd [15:11] | shamt [10:6] | func [5:0] | function | opcode |
|---|---|---|---|---|---|---|---|---|
| ADD | 000000 | rs | rt | rd | 00000 | 100000 | rd = rs + rt | ADD |
| ADDU | 000000 | rs | rt | rd | 00000 | 100001 | rd = rs + rt | ADD |
| SUB | 000000 | rs | rt | rd | 00000 | 100010 | rd = rs - rt | SUB |
| SUBU | 000000 | rs | rt | rd | 00000 | 100011 | rd = rs - rt | SUB |
Columns
#![allow(unused)] fn main() { pub struct AddSubCols<T> { pub pc: T, pub next_pc: T, pub add_gate: T, // is_add * (1 - op_a_0) pub sub_gate: T, // is_sub * (1 - op_a_0) pub is_add: T, pub is_sub: T, pub frame: RTypeFrameCols<T>, } }
The R-type frame holds the shard, the two clock limbs, the opcode, the three register indices, op_a_0, and the three register accesses. op_a is a read-write access (previous value, value, previous clock, clock-difference limb); op_b and op_c are reads (value, previous clock, clock-difference limb). With 4-byte words the chip has 36 columns: 6 of its own and 30 in the frame. The chip needs no columns for the result or the carries. The result is the value of the op_a access, whose bytes the frame range-checks, and the carries are expressions.
Constraints
Write \( a_i, b_i, c_i \) for byte \( i \) of the values of op_a, op_b and op_c. For addition, the carry out of byte \( i \) is
\[ k_i = (b_i + c_i - a_i + k_{i-1}) \cdot 256^{-1}, \qquad k_{-1} = 0, \]
a degree-1 expression in the columns. \( a = b + c \bmod 2^{32} \) holds exactly when every \( k_i \) is 0 or 1, so the chip asserts
\[ add\_gate \cdot k_i \cdot (k_i - 1) = 0, \qquad i = 0, \dots, 3. \]
Subtraction \( a = b - c \) is checked as the addition \( a + c = b \), with carries \( k'i = (a_i + c_i - b_i + k'{i-1}) \cdot 256^{-1} \) and \( sub\_gate \cdot k'_i \cdot (k'_i - 1) = 0 \).
The gate is a column rather than the product \( is\_add \cdot (1 - op\_a\_0) \): with the product inline the constraint would have degree 4. The chip asserts \( add\_gate = is\_add \cdot (1 - op\_a\_0) \) separately. When the destination is $zero the frame forces the written value to 0, which need not equal \( b + c \), so the check is switched off.
The remaining constraints are that is_add, is_sub and their sum is_real are boolean, and the frame's interactions with is_real as multiplicity:
- the
Programfetch of(pc, ADD or SUB, op_a, op_b, op_c, ...); - the three register accesses on the
Memorybus; - the
U16Range,RangeandU8Rangebyte lookups for the shard, the clock and the result bytes; - the
Statereceive of(shard, clk, pc, next_pc)and send of(shard, clk + 5, next_pc, next_pc + 4).
Example trace
Consider this fragment:
| pc | instruction | effect |
|---|---|---|
| 0x100 | addu $7, $5, $6 | \( 13 + 13685 = 13698 \) |
| 0x104 | slt $2, $6, $7 | (proved by the Lt chip) |
| 0x108 | subu $5, $7, $4 | \( 13698 - 10 = 13688 \) |
The AddSub chip gets two rows. The slt goes to the Lt chip, and the State bus connects the rows of the two chips. The value columns in little-endian bytes:
| pc | next_pc | is_add | is_sub | add_gate | sub_gate | op_a_0 | a (value of op_a) | b | c |
|---|---|---|---|---|---|---|---|---|---|
| 0x100 | 0x104 | 1 | 0 | 1 | 0 | 0 | [130, 53, 0, 0] | [13, 0, 0, 0] | [117, 53, 0, 0] |
| 0x108 | 0x10c | 0 | 1 | 0 | 1 | 0 | [120, 53, 0, 0] | [130, 53, 0, 0] | [10, 0, 0, 0] |
In the first row \( k_0 = (13 + 117 - 130) / 256 = 0 \) and \( k_1 = (0 + 53 - 53 + 0)/256 = 0 \). In the second row \( k'_0 = (120 + 10 - 130)/256 = 0 \). All carries are boolean and the rows are valid. Had the prover put 131 in \( a_0 \) of the first row, \( k_0 = -1/256 \) would be neither 0 nor 1 and the constraint would fail.
The rows the chip has beyond the executed instructions are padding: all selectors are 0, so every gated constraint holds and every interaction has multiplicity 0.
Preprocessed traces
Tables that do not depend on the execution (the program image, the byte table, the range table) are preprocessed. Their columns are committed during setup, the commitment is part of the verifying key, and the prover only supplies the multiplicity columns in the main trace. A proof for a different program therefore fails against the verifying key, because the program table's commitment differs.
Lookup Arguments
A lookup argument proves that every value a table looks up appears in another table. Ziren uses one general form of it for every relation between rows and between chips: each chip row sends or receives tuples on named buses, and a single argument per shard proves that on every bus the multiset of sent tuples equals the multiset of received tuples. A byte range check, an instruction fetch, a memory access and the hand-off from one instruction to the next are all instances. See State Machine for the list of buses.
LogUp
Ziren uses the logarithmic-derivative form of the multiset check, LogUp. A multiset \( \{ f_j \} \) with multiplicities \( m_j \) equals a multiset \( \{ t_i \} \) with multiplicities \( m'_i \) exactly when, as rational functions of \( X \),
\[ \sum_j \frac{m_j}{X + f_j} = \sum_i \frac{m'_i}{X + t_i}. \]
The verifier checks the identity at a random \( X = \alpha \). A false identity survives only if \( \alpha \) is a root of a nonzero polynomial whose degree is bounded by the total number of terms, so the error is at most that number divided by the size of the extension field.
A tuple \( (v_1, \dots, v_k) \) on bus \( b \) is compressed to a single field element, its fingerprint, with further random challenges:
\[ f = \alpha + \beta_0 \cdot b + \sum_{j=1}^{k} \beta_j \cdot v_j. \]
Each interaction contributes the fraction \( m / f \), with \( m \) the row's multiplicity expression for a send and \( -m \) for a receive. The multiplicity is a column expression, typically is_real or a selector, and is zero on padding rows. The argument holds when the sum of all fractions over all rows of all chips in the shard is zero. The exception is buses that the shard's public values close (State, GlobalAccumulation and the two global-memory control buses), where the sum must equal the fractions of the boundary tuples the verifier computes from the public values.
The challenges are drawn from the degree-4 extension of KoalaBear after the prover has committed to all traces.
Proving the sum with GKR
The sum has one term per (row, interaction) pair, many millions per shard. Ziren does not commit to running-sum columns. It proves the sum with the GKR protocol for fractional sums (LogUp-GKR):
- Each chip's
(numerator, denominator)pairs form a table indexed by row and interaction. Chips are padded to a common number of rows and interactions with the neutral fraction \( 0/1 \). - A layered circuit adds the fractions pairwise: \( \frac{n_0}{d_0} + \frac{n_1}{d_1} = \frac{n_0 d_1 + n_1 d_0}{d_0 d_1} \). Each layer halves the row dimension, and the last layer combines interactions and chips into one fraction.
- The prover sends the output fraction, and the verifier checks that it matches the expected total. Then, layer by layer, a sumcheck reduces a claim about one layer's multilinear extensions to a claim about the layer below at a new random point.
- At the bottom, the claim is about the numerator and denominator at a random point. These are low-degree expressions in the chips' columns, so it reduces to claims about the multilinear extensions of the trace columns at that point.
These column claims are not checked by the lookup argument itself. They are passed to the zerocheck, which batches them with the constraint check, and the resulting openings are proved by the polynomial commitment (see STARK Protocol). Before the first GKR challenge is sampled, the prover grinds a proof of work of ZIREN_LOGUP_GRINDING_BITS bits (default 22). The grind raises the soundness of this step to the per-component target.
Byte and range lookups
The two lookup tables are preprocessed chips that receive on the Byte and Range buses.
ByteLookuphas one row for each pair of bytes \( (b, c) \), \( 2^{16} \) rows. The preprocessed columns hold the result of each byte operation on that pair (AND,OR,XOR,NOR,SLL, shift-right with carry,LTU,MSB), and the pair itself doubles as a 16-bit value forU16Range. A lookup(op, a, b, c)claims thatais the result ofoponbandc(or, for range checks, that the operand is in range). The main trace has one multiplicity column per operation.RangeLookupholds the pairs(a, bits)with \( a < 2^{bits} \) for \( bits \le 10 \). The machine uses it for limbs narrower than a byte or a half-word, chiefly the 10-bit high limb of a 26-bit timestamp.
For example, the check that a word consists of four bytes is two U8Range lookups, (U8Range, 0, b_0, b_1) and (U8Range, 0, b_2, b_3). Each tuple exists in the table only if both bytes are below 256. The table row for that pair counts the requests in its U8Range multiplicity column, and bus balance forces the counts to be right.
Memory Consistency Checking
Offline memory checking lets a prover show that a read/write memory was used correctly: every read returns the value most recently written to that address. Unlike online checking with Merkle paths, it checks nothing per access. Each access adds tuples to two multisets, and one equality check between the multisets at the end covers all accesses. Ziren uses it for registers and memory alike, within a shard through the lookup argument and across shards through a multiset hash on an elliptic curve.
Read set and write set
The read set \( RS \) and write set \( WS \) are multisets of tuples \( (a, v, c) \): an address, a value and a timestamp.
- Initialization. \( RS = WS = \emptyset \). For every address \( a_i \) with initial value \( v_i \), add \( (a_i, v_i, 0) \) to \( WS \).
- Access. To access address \( a \) at time \( c_{now} \), take the last tuple \( (a, v, c) \) written for \( a \), add it to \( RS \), and add \( (a, v', c_{now}) \) to \( WS \). For a read \( v' = v \); for a write \( v' \) is the new value.
- Post-processing. For every address, add its last tuple in \( WS \) to \( RS \).
The memory was used correctly if:
- the sets were initialized correctly;
- at every access \( c < c_{now} \), so the timestamps of an address strictly increase;
- a read adds the same value to \( RS \) and \( WS \);
- after post-processing, \( RS = WS \).
Suppose the first incorrect read of address \( a \) returns \( (a, v', c') \) instead of the last written \( (a, v, c) \). All tuples in \( WS \) are distinct because timestamps strictly increase, and \( (a, v', c') \) was never written. So \( RS \) contains a tuple that \( WS \) does not contain, and no later step can remove it. Then \( RS \neq WS \).
Within a shard
A shard's accesses are tuples (shard, clk, addr, value) on the Memory bus. The timestamp is the pair (shard, clk).
- An access sends its previous tuple
(prev_shard, prev_clk, addr, prev_value)(the read) and receives its new tuple (the write). It asserts that(shard, clk)is strictly larger than(prev_shard, prev_clk). When the shards are equal, the differenceclk - prev_clk - 1is split into a 16-bit and a 10-bit limb and both are range-checked, which provesclk > prev_clkbecause both clocks are below \( 2^{26} \). Otherwise the same check is applied to the shards. MemoryLocalhas one row per address the shard touches. The row supplies the address's first tuple in the shard and consumes its last one.- Registers are addresses 0 to 35.
MemoryBumpinserts a read of every touched register at(shard, 0), so the check for every other register access compares clocks only (see Memory).
The Memory bus balances when the sends equal the receives, which is condition 4 for the shard. The LogUp-GKR argument proves it together with all other buses (see Lookup Arguments).
Across shards
The first and last tuple of each address in a shard must also match the neighbouring shards. These tuples go on the Global bus:
MemoryLocalreceives the address's initial tuple and sends its final tuple.MemoryGlobalInitsends the initial value of each address outside the program image, at timestamp(0, 0).MemoryGlobalFinalreceives the final value of each touched address.- The initial values of the program image are not rows. Their contribution is precomputed at setup as the verifying key's
initial_global_cumulative_sum.
Because shards are proved independently, the lookup argument, which only balances within one shard, cannot match these messages. Instead each message is hashed to a point on an elliptic curve, the Global chip adds up the shard's points, and the shard's sum is a public value. The recursion adds the sums of all shards and the verifying key's initial sum, and the root checks that the total is the neutral digest. A matching send and receive map to opposite points and cancel. The same mechanism carries syscalls from an execution shard to the precompile shard that proves them.
Multiset hashing
A multiset hash maps a multiset to a short value such that it is infeasible to find two different multisets with the same hash, and such that the hash can be updated one element at a time in any order. Ziren maps each element to a point on an elliptic curve and hashes a multiset to the sum of its points.
To map a message \( m = (m_0, \dots, m_6) \) to a point, Ziren follows Constraint-Friendly Map-to-Elliptic-Curve-Group Relations and Their Applications and uses the message directly as the \( x \)-coordinate, without hashing it first:
- \( x_0 = m_0 + 2^{16} \cdot kind \), where \( m_0 \) is range-checked to 16 bits and \( kind \) names the bus the message came from;
- \( x_i = m_i \) for \( 1 \le i \le 5 \);
- \( x_6 = 256 \cdot m_6 + t \), where \( m_6 \) is a byte and \( t \) is an 8-bit tweak.
The prover tries tweaks \( t = 0, 1, \dots \) until \( x \) is on the curve. For the square root \( y \), the top coefficient \( y_6 \) fixes the sign. A received message must have \( 1 \le y_6 \le 63 \cdot 2^{24} \), and a sent message \( 2^{30} + 1 \le y_6 \le p - 1 \). The two ranges are disjoint and mirror each other under \( y \mapsto -y \), so a send and the matching receive give opposite points. Each row sets exactly one of is_send and is_receive.
The Global chip adds the points in a chain on the GlobalAccumulation bus using the chord formula. It witnesses \( (x_2 - x_1)^{-1} \) at each step so that doubling and adding an inverse are not provable. Sums start at a fixed point derived from \( \sqrt{2} \), and that point is the neutral digest.
The parameters are:
- base field KoalaBear, \( p = 2^{31} - 2^{24} + 1 \);
- extension \( \mathbb{F}_{p^7} = \mathbb{F}_p[z]/(z^7 + 2z - 8) \);
- curve \( y^2 = x^3 + 3z \cdot x - 3 \).
Elliptic curve selection over the KoalaBear extension field
Objective
Find an elliptic curve over the degree-7 extension of KoalaBear, \( p = 2^{31} - 2^{24} + 1 \), with more than 100 bits of security against known attacks and cheap arithmetic.
Code location
The search is in septic-curve-over-koalabear, a fork of Cheetah, which finds a curve over a sextic extension of the Goldilocks prime \( 2^{64} - 2^{32} + 1 \).
Construction
-
Step 1: sparse irreducible polynomial.
- Requirements: few nonzero coefficients, small coefficients, irreducible over the base field.
- Implementation (
septic_search.sage):poly = find_sparse_irreducible_poly(Fpx, extension_degree, use_root=True). - Result: \( z^7 + 2z - 8 \).
-
Step 2: candidate curves.
- Form \( y^2 = x^3 + ax + b \) with small coefficients.
- Search in
septic_search.sage:for i in range(wid, 1000000000, processes): coeff_a = 3 * a # Fixed coefficient scaling coeff_b = i - 3 E = EllipticCurve(extension, [coeff_a, coeff_b]) - Result: \( a = 3z \), \( b = -3 \), with \( z \) the generator of the extension.
-
Step 3: security checks.
- Pollard rho: the largest prime factor of the group order has more than 210 bits.
prime_order = list(ecm.factor(n))[-1] assert prime_order.nbits() > 210 - Embedding degree:
embedding_degree = calculate_embedding_degree(E) assert embedding_degree.nbits() > EMBEDDING_DEGREE_SECURITY - The same two checks for the quadratic twist.
- Pollard rho: the largest prime factor of the group order has more than 210 bits.
-
Step 4: complex discriminant. With \( n \) the order of the curve, \( D = (p^7 + 1 - n)^2 - 4p^7 \) must be a large negative integer (absolute value above 100 bits) whose square-free part exceeds 100 bits. Run
sage verify.sageto check.
STARK Protocol
This page describes how one shard is proved: the shard argument. Its inputs are the shard's chip traces (see Arithmetization) and its public values. The argument has four parts, run in this order under one Fiat-Shamir transcript:
- a commitment to all main traces of the shard with one jagged polynomial commitment;
- a LogUp-GKR argument that every bus balances;
- a zerocheck that every chip's constraints hold on every row;
- an opening of the committed columns at one random point, proved with WHIR.
The code is in crates/pcs/src/shard_level/ (prover and verifier), crates/pcs/src/jagged*.rs and crates/pcs/src/whir/.
Fields, hash and transcript
- Base field: KoalaBear, \( p = 2^{31} - 2^{24} + 1 \).
- Challenge field: the degree-4 extension,
BinomialExtensionField<KoalaBear, 4>. - Hash: Poseidon2 over KoalaBear with width 16. Merkle trees use a padding-free sponge for leaves and a truncated permutation for internal nodes, with 8-element digests.
- Transcript: a duplex challenger on the same permutation.
The wrap proof, which is verified inside a BN254 SNARK, uses a different hash and challenger (see STARK to SNARK).
Traces as multilinear polynomials
Each column of a chip trace with \( 2^n \) rows is read as the multilinear extension of its values over the Boolean hypercube \( \{0,1\}^n \). There are no cyclic domains, no "next row" and no quotient polynomial. A chip's constraints are polynomials \( C_j \) in the values of one row (main and preprocessed columns, plus public values), of degree at most 3.
Transcript prologue
The verifier observes the verifying key (the preprocessed commitment, the start pc, the initial global digest and the chip layout). For each shard it then observes the public values, the main commitment, and the number, heights and names of the chips present. Binding the heights before any challenge prevents the prover from choosing trace dimensions after seeing randomness.
Lookup argument
The first sub-argument proves that every bus balances. The prover grinds a proof of work, then runs LogUp-GKR over all interactions of all chips (see Lookup Arguments). It ends with claimed evaluations of the columns that appear in interactions, at a random point.
Zerocheck
The zerocheck proves that every constraint vanishes on every real row. Constraints are batched within a chip with powers of a challenge \( \alpha \), giving \( C(x) = \sum_j \alpha^j C_j(x) \), and across chips with powers of a challenge \( \lambda \). The GKR's column claims are folded into the same sum with a third challenge. The prover then runs one sumcheck for
\[ \sum_{b \in \{0,1\}^n} \mathrm{eq}(r, b) \cdot C(b) = \text{(the combination of the GKR claims)}, \]
where \( r \) is a random point and \( \mathrm{eq} \) is the multilinear equality polynomial. Each round polynomial has degree 4 (degree 3 from the constraints, 1 from \( \mathrm{eq} \)), and there is one round per row variable of a fixed cube, \( 2^{22} \) rows for core shards (CORE_MAX_LOG_ROW_COUNT). A chip shorter than the cube is extended virtually, and the contribution of the virtual rows is removed analytically, so they need not satisfy the constraints. The sumcheck ends with claimed evaluations of every column of every chip at one point \( z^* \).
Jagged commitment
A shard has dozens of chips with different heights and widths. Committing each separately would cost one Merkle tree and one opening per chip. Ziren uses a jagged polynomial commitment (Jagged Polynomial Commitments). All columns of all chips are concatenated into one dense vector
\[ q = [\, T_0[:,0] \mid T_0[:,1] \mid \cdots \mid T_N[:,w_N - 1] \,] \]
without padding the columns to a common height. Prefix sums \( t_k \) of the column heights record where each column starts. The sparse "table of all columns" is related to \( q \) by
\[ p(z_r, z_c) = \sum_j q(j) \cdot \mathrm{eq}(\mathrm{row}(j), z_r) \cdot \mathrm{eq}(\mathrm{col}(j), z_c), \]
where \( \mathrm{col}(j) \) and \( \mathrm{row}(j) \) are the column and row that position \( j \) belongs to according to \( t \).
The dense vector is cut into stripes of \( 2^{21} \) values (the stacking height, DEFAULT_LOG_STACKING_HEIGHT). Each stripe is Reed-Solomon encoded, and all stripes of a round are committed in one Merkle tree. The preprocessed traces form a round committed at setup, whose root is in the verifying key; the main traces form a round committed per shard.
To open the columns at \( z^* \), a sumcheck reduces the per-column claims to one claim about \( q \) at a random point. The verifier evaluates the jagged indicator from the heights it already observed. The heights are part of the transcript, so the verifier checks the layout it was committed to. The claim about \( q \) is proved with WHIR.
WHIR
WHIR is a multilinear polynomial commitment built on Reed-Solomon proximity testing. The stripes of both committed rounds are batched into one virtual polynomial \( F = \sum_i \mu^i \cdot \mathrm{stripe}_i \), and \( \mu \) is drawn after the claims are fixed. The prover grinds before the batching challenge. WHIR then alternates folding sumchecks, commitments to the folded codeword at a lower rate, out-of-domain samples and queries into the previous codeword. Each query opens a Merkle path, and a round-0 query opens the same row of every stripe.
The production schedule for a core shard (core_whir_config in crates/pcs/src/whir/jagged.rs) is:
| Parameter | Value |
|---|---|
| Stacking height | \( 2^{21} \) |
| Starting rate | \( \rho = 2^{-2} \); each committed round divides it by 8 |
| Folding factors | 3, then 6, 6 |
| Queries per round | 124, 88, 85 (final: 85) |
| Out-of-domain samples | 2 per committed round |
| Query grinding | 22 bits (ZIREN_WHIR_QUERY_GRINDING_BITS) |
| Batching grinding | 14 bits (ZIREN_WHIR_BATCH_GRINDING_BITS) |
| LogUp-GKR grinding | 22 bits (ZIREN_LOGUP_GRINDING_BITS) |
The query counts are not listed in the code but solved. In the unique-decoding regime, a query into a code of rate \( \rho \) is worth \( -\log_2((1 + \rho)/2) \) bits, so a round with \( q \) queries and \( g \) grinding bits gives \( q \cdot (-\log_2((1+\rho)/2)) + g \) bits. The solver picks the least \( q \) that reaches the per-component target, ZIREN_SOUNDNESS_TARGET_BITS (default 106). At the defaults:
\[ 124 \cdot 0.678 + 22 = 106.08, \quad 88 \cdot 0.956 + 22 = 106.09, \quad 85 \cdot 0.994 + 22 = 106.52. \]
The target is 106 rather than 100 because a shard transcript has about two dozen components, and a union bound over them should still exceed 100 bits: \( 100 + \log_2 24 \approx 104.6 \). Lowering a grinding parameter through the environment raises the query counts instead of weakening a round. These are bounds on the interactive protocol; the Fiat-Shamir transformation costs a further factor in the number of hash queries an adversary makes.
Verification
The verifier, verify_shard in crates/pcs/src/shard_level/verifier.rs:
- replays the prologue;
- checks the LogUp-GKR proof, using the public values for the endpoints of the
State,GlobalAccumulationand global-memory control buses; - checks the zerocheck sumcheck and recomputes every chip's batched constraint at \( z^* \) from the claimed column values;
- checks the jagged reduction and the WHIR proof against the preprocessed root in the verifying key and the shard's main commitment.
Checks between shards (the pc chain, shard numbers, memory address ranges, the global digest) are not part of the shard argument. The recursion enforces them (see Recursive STARK).
Prover Architecture
The Ziren prover turns one execution of a guest ELF into a proof in four stages. Each stage is a function of ZKMProver in crates/prover/src/lib.rs, and the SDK calls them in this order:
- Execute and prove shards (
prove_core). The executor runs the program and cuts the execution into shards. Each shard's events become chip traces, and each shard gets its own shard proof (see STARK Protocol). The result is aZKMCoreProof, a list of shard proofs. - Compress (
compress). A recursion tree verifies all shard proofs, and any deferred proofs, and reduces them to one recursion proof over the whole execution (see STARK Aggregation and Recursive STARK). - Shrink and wrap (
shrink,wrap_bn254). The compressed proof is verified by a fixed-shape shrink program and then by a wrap program. The wrap program is proved with a BN254-friendly hash so that a SNARK can verify it. - SNARK (
wrap_groth16_bn254,wrap_plonk_bn254,wrap_dvsnark_bn254). A gnark circuit verifies the wrap proof and produces a Groth16, PLONK or designated-verifier SNARK over BN254 (see STARK to SNARK).
The SDK proof kinds stop at different stages: core after stage 1, compressed after stage 2, and groth16, plonk or dvsnark after stage 4.
Execution and sharding
The executor (crates/core/executor) interprets the MIPS32r2 program and records events: one per executed instruction, memory access, syscall and precompile call. It closes the current shard and starts a new one when any of these limits is reached:
- the cycle budget (
shard_size); - the per-shard clock, so that every timestamp stays below the \( 2^{26} \) range the memory argument checks;
- the estimated trace area of the shard, in cells;
- the height of the tallest chip, which must stay below \( 2^{22} \) rows, the size the recursion verifier is built for.
A shard is never closed in front of a delay slot. Precompile events are moved into separate precompile shards, and the memory initialization and finalization events of the whole execution are placed in the last shards. Each shard records its public values: start and end pc, shard and execution-shard numbers, the memory address ranges it initialized and finalized, the committed-value and deferred-proof digests, the exit code and its global digest.
Shard proofs
Shard proofs are independent of each other, so they are generated in parallel. Every constraint between shards is a public value or a message on the global bus, and the recursion checks them. Trace generation and proving can run on the CPU prover or on GPUs.
Recursion and SNARK
The recursion machine is a separate STARK machine (RecursionAir) over KoalaBear. It runs recursion programs, which are verifiers of shard proofs compiled to the recursion machine's instruction set, and proves their execution with the same shard argument. Each proof it produces is again a shard proof, so recursion proofs can be verified recursively. The final wrap proof uses the same machine with a Poseidon2 hash over BN254, which the gnark circuit can check.
STARK Aggregation
ZKMProver::compress reduces the shard proofs of an execution, together with any deferred proofs, to a single recursion proof. It does this with a tree of recursion programs whose leaves are shard proofs.
Leaves: normalize
Each core shard proof is verified by a normalize program, one shard per program. The program:
- runs the shard verifier (LogUp-GKR, zerocheck, jagged WHIR opening) against the program's verifying key;
- checks the conditions that depend on the shard alone;
- outputs the shard's public values in the recursion format (
RecursionPublicValues), which covers a shard range[start_shard, next_shard).
The first shard must start at the verifying key's pc_start and contributes the key's initial global digest. An execution shard must have sequential boundaries (start_next_pc = start_pc + 4, next_next_pc = next_pc + 4) and exit code 0.
Deferred proofs are verified by a separate deferred program in batches. A deferred batch covers a range placed after the last execution shard (see Proof Composition).
Inner nodes: compose
A compose program verifies up to three (REDUCE_BATCH_SIZE) recursion proofs of adjacent ranges and outputs the public values of their union. It checks that each input starts where the previous one ends:
start_pcequals the previousnext_pc;- shard and execution-shard numbers continue;
- the initialized and finalized address ranges join;
- the deferred-proof digest chain continues;
- the committed-value and deferred-proof digests agree.
It also adds the inputs' global digests. Each input's verifying key must be in the enumerated set of recursion keys: the program checks a Merkle proof of the key's hash against vk_root, a tree of height 14. The check is compiled into the program when VERIFY_VK is true, the default.
Scheduling
The tree is keyed by shard range, not by depth (crates/prover/src/compress_tree.rs). When a proof completes, the scheduler looks for an adjacent range. Once a contiguous run reaches three proofs, it is dispatched to a compose program; a shorter run waits. Levels therefore overlap: a range is reduced as soon as its neighbours are ready. The compose program asserts continuity between its inputs, so every batch must be a contiguous run in order.
A leaf is never the root. An execution of one shard is closed by a compose of arity one.
Root
The compose that covers the whole execution runs with is_complete set and checks that the proof describes a complete, successful run:
- the execution starts at shard 1 and contains an execution shard;
- it ends with
next_pc = 0; - the reconstructed deferred digest starts at zero and ends at the digest the guest committed;
- the sum of all global digests is the neutral digest, so every cross-shard memory and syscall message was consumed.
The output is a ZKMReduceProof, the compressed proof, which the SDK can verify directly or pass on to the SNARK stages.
STARK to SNARK
A compressed proof is a KoalaBear STARK. Verifying it on a blockchain would cost too much, so Ziren verifies it inside a SNARK over BN254. There are two recursion steps before the SNARK and then the SNARK itself.
1. Shrink
ZKMProver::shrink(reduced_proof, opts) runs the shrink program. The program verifies the compressed proof, including the Merkle proof that its verifying key is in the recursion key set, and is proved with the recursion machine (ShrinkAir) at a fixed shape. The compressed proof's shape depends on the execution; the shrink proof's does not. This gives the next step a single input shape. The shrink proof still uses the inner configuration: KoalaBear, Poseidon2 over KoalaBear and jagged WHIR.
2. Wrap
ZKMProver::wrap_bn254(shrink_proof, opts) runs the wrap program, which verifies the shrink proof. The wrap program is proved with a different configuration (OuterSC in crates/recursion/core/src/stark/config.rs), chosen to be cheap to verify in a BN254 circuit:
- the field is still KoalaBear with its degree-4 extension, but Merkle trees and the transcript use Poseidon2 over the BN254 scalar field (width 3). The transcript is a
MultiField32Challenger, which packs KoalaBear elements into BN254 elements; - the dense polynomial commitment under the jagged layer is BaseFold at rate \( 2^{-3} \), with 94 queries and 26 bits of query grinding (
ZIREN_WRAP_QUERY_GRINDING_BITS); - the wrap machine (
WrapAir) allows constraints of degree 9, so its Poseidon2 chip uses the degree-9 column layout with fewer intermediate columns than the degree-3 layout of the compress and shrink machines.
3. SNARK
The wrap proof is verified by a gnark circuit. build_outer_circuit in crates/prover/src/build.rs compiles the wrap verifier into constraints over BN254, and the circuit is proved with one of:
wrap_groth16_bn254: Groth16, the smallest proof and cheapest on-chain verification;wrap_plonk_bn254: PLONK, with a universal setup;wrap_dvsnark_bn254: a designated-verifier SNARK.
The circuit has three public inputs:
| Input | Meaning |
|---|---|
vkey_hash | the digest of the guest program's verifying key (zkm_vk_digest) |
committed_values_digest | the digest of the guest's public outputs (SHA-256; BLAKE3 when the guest is built with the imm-wrap-vk feature), as 32 bytes packed into one BN254 element |
vk_root | the root of the recursion verifying-key set |
The circuit also pins the wrap verifying key: in the standard build it asserts the wrap key's preprocessed commitment and pc_start against the values the circuit was built from, so a change to the recursion programs requires a new circuit and setup. The build with ZKM_IMM_WRAP_VK instead folds the wrap key into vkey_hash with Poseidon2, so the circuit does not change when the wrap key does.
A verifier of the SNARK proof takes the program's vkey_hash and the public values, recomputes committed_values_digest from the public values, and checks the proof against those inputs and the expected vk_root.
Proof Composition
A guest program can verify other Ziren proofs. The guest states which proofs it relies on, and the recursion tree checks them as deferred proofs, so the guest does not run a STARK verifier inside the MIPS machine.
Use cases
- Aggregation. Combine many independent proofs, for example of blocks or transactions, into one proof.
- Modular programs. Split an application into programs that are proved separately and composed, so that one part can change without re-proving the rest.
- Pipelining. Prove the parts of a long computation in parallel and join them in a final program.
Interface
In the guest, with the verify feature of zkm-zkvm:
#![allow(unused)] fn main() { zkm_zkvm::lib::verify::verify_zkm_proof(&vk_digest, &public_values_digest); }
vk_digest: &[u32; 8] is the digest of the inner program's verifying key, and public_values_digest: &[u8; 32] is the digest of its public values. The call does not verify anything by itself. It issues the VERIFY_ZKM_PROOF syscall (0x1B) and folds the pair into the guest's running deferred-proof digest:
\[ D \leftarrow \mathrm{Poseidon2}(D, \mathit{vk\_digest}, \mathit{pv\_digest}), \qquad D_0 = 0. \]
When the guest halts, it commits \( D \) word by word with the COMMIT_DEFERRED_PROOFS syscall. The SyscallInstrs chip checks each word against the shard's public deferred_proofs_digest.
On the host, the inner proofs must be compressed proofs (ZKMReduceProof). They are passed in with the input, in the order the guest verifies them:
#![allow(unused)] fn main() { stdin.write_proof(inner_proof, inner_vk); }
During execution, each VERIFY_ZKM_PROOF call takes the next proof from this list, and unless deferred-proof verification is disabled in the executor options, the executor checks it against the call's (vk_digest, pv_digest). This check only reports errors early; soundness comes from the recursion.
How the recursion checks it
compress passes the inner proofs to the deferred program in batches (get_recursion_deferred_inputs_basefold). For each batch, the deferred program:
- verifies each inner compressed proof, including the Merkle proof that its verifying key is in the recursion key set;
- requires each proof to be complete;
- folds each proof's verifying-key digest and committed-value digest into the reconstructed digest with the same Poseidon2 hash, starting from the previous batch's value.
Deferred batches are placed in the shard range after the last execution shard, so the compose programs join them to the execution like any other range and chain start_reconstruct_deferred_digest to end_reconstruct_deferred_digest. At the root, assert_complete requires the reconstruction to start at zero and to end at the deferred_proofs_digest the guest committed. A guest that claims a proof the prover did not supply, or a different one, therefore produces a digest that the reconstruction cannot match.
Verification
The outer proof is verified like any other Ziren proof, against the outer program's verifying key. Its public values are only the outer guest's. The inner proofs and their public values are bound through the deferred digest and are not needed by the verifier.
Recursive STARK
This page follows one execution through the proving pipeline and names the code for each step. STARK Aggregation describes what the recursion tree checks, and STARK Protocol describes the shard argument that every proof in the pipeline uses.
Shard proof generation
1. Setup
-
ZKMProver::get_program(elf)loads the ELF into aProgram: the instruction image,pc_startand the initial memory image. -
ZKMProver::setup(elf)builds the proving and verifying keys. The verifying key holds:- the commitment to the preprocessed traces (the program table, the byte and range tables);
pc_start;initial_global_cumulative_sum, the global digest of the initial memory image;- the chip layout.
Its digest (
zkm_vk_digest) identifies the program in every later proof.
2. Execution and trace generation
ZKMProver::prove_core(pk, program, stdin, opts, context) runs the executor, which cuts the execution into shards (see Continuation), and generates the traces of each shard. The machine is MipsAir (crates/core/machine/src/mips/mod.rs), an enum with one variant per chip: the instruction chips (each with its instruction frame, see CPU), the memory chips, the lookup tables, the syscall and precompile chips and the Global chip. A chip with no events in a shard is left out of that shard's proof.
3. Shard proofs
Each shard is proved with the shard argument: jagged commitment, LogUp-GKR, zerocheck and WHIR. The result, ZKMCoreProof, is a list of ShardProofs. Each holds the main commitment, the public values, the chip heights, the LogUp-GKR and zerocheck proofs, the opened values and the jagged evaluation proof.
Recursive aggregation
RecursionAir
Recursion programs run on a small virtual machine over KoalaBear whose execution is proved by RecursionAir (crates/recursion/core/src/machine.rs). Its chips are:
| Chip | Role |
|---|---|
MemoryConst, MemoryVar | the program's memory: constants and variables |
BaseAlu, ExtAlu | arithmetic in KoalaBear and in its degree-4 extension |
Poseidon2Wide | the Poseidon2 permutation, for Merkle paths and the transcript |
Select | conditional selection |
Ext2Felt | decomposition of an extension element into base-field coefficients |
PublicValues | the program's public values |
The chips have no special gates for FRI or sumcheck. The recursion verifier's sumchecks, Merkle path checks and WHIR folding are compiled to these operations. The machine is instantiated at three constraint degrees: CompressAir and ShrinkAir at degree 3 and WrapAir at degree 9. The degree changes only the Poseidon2 column layout.
Recursion programs
The recursion compiler (crates/recursion/compiler) builds each program from the verifier code in crates/recursion/circuit:
| Program | Built by | Verifies |
|---|---|---|
| normalize | recursion_program_basefold | one core shard proof |
| compose | compose_program_basefold | one to three adjacent recursion proofs |
| deferred | deferred_program_basefold | a batch of deferred compressed proofs |
| shrink | shrink_program_basefold | the compressed proof |
| wrap | wrap_bn254_program_basefold | the shrink proof, for the BN254 SNARK |
The program is fixed by the shape of the proofs it verifies. A normalize program depends on the chip set of the shard and its height class. A compose program depends only on its arity, since every recursion proof has the same shape. The prover caches programs and their keys.
Verifying-key set
Programs are generated, not written by hand, so the set of valid recursion keys is enumerated in advance: the normalize programs for every shard shape, the compose programs of every arity, and the deferred programs. Their hashes form a Merkle tree of height 14 (VK_MERKLE_TREE_HEIGHT) whose root, vk_root, is a public value of every recursion proof and a public input of the SNARK. The enumeration ships with the prover as vk_map.bin. The compose and deferred programs check a Merkle proof that the key of each proof they verify is in the set, and the shrink program checks it for the compressed proof. Without this check, a prover could verify a proof of an arbitrary program in place of a normalize proof.
Compression
ZKMProver::compress(vk, core_proof, deferred_proofs, opts):
- builds the leaf inputs with
get_first_layer_inputs: one normalize input per shard (get_recursion_core_inputs_basefold) and the deferred batches (get_recursion_deferred_inputs_basefold); - runs the range-keyed reduction tree of compose programs (
compress_tree.rs) until one proof covers the whole execution; - returns the root proof as a
ZKMReduceProof.
shrink and wrap_bn254 continue from there (see STARK to SNARK).
Source mapping
| Stage | Functions | Machine |
|---|---|---|
| Setup | get_program, setup | MipsAir |
| Execution, traces, shard proofs | prove_core | MipsAir |
| Leaf inputs | get_first_layer_inputs, get_recursion_core_inputs_basefold, get_recursion_deferred_inputs_basefold | |
| Compression | compress, compose_program_basefold, deferred_program_basefold | CompressAir |
| Shrink | shrink, shrink_program_basefold | ShrinkAir |
| Wrap | wrap_bn254, wrap_bn254_program_basefold | WrapAir over OuterSC |
| SNARK | wrap_groth16_bn254, wrap_plonk_bn254, wrap_dvsnark_bn254 | gnark over BN254 |
Continuation
Ziren proves an execution of any length by splitting it into shards, proving each shard on its own, and joining the shard proofs in the recursion. This has three advantages:
- Bounded proofs. Each shard proof has a bounded size and cost, however long the execution.
- Parallelism. Shards are proved independently, on many cores, GPUs or machines.
- State continuity. The recursion checks that each shard starts where the previous one ended, and memory consistency checking across shards ensures that every shard sees the memory the previous shards left.
Sharding
The executor runs the whole program and closes a shard when any of these limits is reached (see Prover Architecture): the cycle budget, the per-shard clock limit, the estimated trace area, or the height of the tallest chip. It never closes a shard in front of a branch delay slot. Precompile calls are proved in separate precompile shards, which the execution shards reach through syscall messages on the global bus.
The per-shard clock clk restarts at 0 in every shard, and a timestamp is the pair (shard, clk).
Shard state
A shard's proof exposes its interface as public values:
| Public values | Meaning |
|---|---|
start_pc, start_next_pc, next_pc, next_next_pc | the program counters on entry and exit |
shard, execution_shard | the shard's number and its number among execution shards |
previous_init_addr_bits, last_init_addr_bits, previous_finalize_addr_bits, last_finalize_addr_bits | the range of addresses whose global initialization and finalization the shard covers |
committed_value_digest, deferred_proofs_digest | the guest's public output digest and deferred-proof digest, once committed |
exit_code | the exit code on halt |
global_cumulative_sum | the sum of the shard's cross-shard messages on the septic curve |
No register file or memory image is passed between shards. Registers are memory addresses 0 to 35. The last value of each address a shard touches leaves the shard as a message on the global bus, and the next shard that touches the address takes it from there (see Memory Consistency Checking).
Key constraints
- Shard validity. Each shard proof verifies against the program's verifying key.
- Initial state. The first shard starts at
pc_startfrom the verifying key, and its global digest is added to the key's digest of the initial memory image. - Transitions. Each shard's
start_pcequals the previous shard'snext_pc. Shard boundaries are sequential (start_next_pc = start_pc + 4,next_next_pc = next_pc + 4), so no pending branch target crosses them. Shard numbers increase by one, and the initialized and finalized address ranges of consecutive shards join. - Completion. The last shard ends with
next_pc = 0and exit code 0. The sum of all shards' global digests and the initial digest is the neutral digest, so every memory value handed from one shard to another was consumed exactly once.
The normalize programs check the conditions on single shards, the compose programs check the transitions, and the root checks completion (see STARK Aggregation).