Skip to content

Latest commit

 

History

History
2500 lines (1987 loc) · 99.8 KB

File metadata and controls

2500 lines (1987 loc) · 99.8 KB

Quake - End-to-end testing management tool

Quake is a tool for deploying Arc testnets and running end-to-end tests.

Quake testnets vs. the Arc Testnet: Arc has a public, persistent Testnet open to external developers and validators. Quake testnets are different: they are private, ephemeral networks spun up on demand for development and CI, then torn down when testing is complete. All mentions of "testnet" in this repository refer to Quake testnets unless explicitly stated otherwise.

Features

  • Describe the testnet topology, node configurations, and test scenarios in TOML manifest files.
  • CLI interface for deploying, managing, monitor testnets, and run tests on them.
  • Deploy testnets locally in your machine or remotely in AWS infrastructure.
  • Test the resilience of a running testnet:
    • perturb the nodes (disconnect, kill, pause, and restart them),
    • perform chaos testing (apply random perturbations, at random intervals, to random subsets of nodes or containers),
    • change the voting power of validators in the validator set of a node,
    • upgrade the running version of individual nodes.
  • Emulate network latency between nodes by assigning data-center regions to nodes and injecting artificial latency between regions.
  • Web-based topology viewer for real-time visualization of nodes, connections, peer status, and network health.
  • MCP (Model Context Protocol) server for AI-assisted testnet management via Claude Code, Cursor, and other MCP-compatible clients.

Table of contents

Usage

Build Quake from the repository root directory:

cargo build -p quake
ln -s target/debug/quake quake # create a symbolic link to simplify calling the tool

Quick start (local deployment)

Deploying and starting a local testnet is as simple as running start:

./quake -f crates/quake/scenarios/examples/10nodes.toml start

This will generate all necessary files to run the testnet and start running Arc in the nodes.

Now you can display comprehensive information about the testnet and check its status:

./quake info

The start command implicitly performs other commands (mainly build and setup), which we describe below. Check out the command workflow for a complete picture of all available commands.

Generate the files needed to start the testnet. If needed, you can manually modify the generated configuration files before starting the nodes.

./quake -f crates/quake/scenarios/examples/10nodes.toml setup

Tip

A command with the -f or --file option creates or updates a .quake/.last_manifest file containing the path to the most recently used manifest. This allows subsequent Quake commands to automatically reference the last manifest without requiring the -f flag.

# Build Docker images for the Consensus Layer (CL) and Execution Layer (EL)
./quake build

# Start the nodes. Implicitly calls setup.
./quake start

# Shows comprehensive information about the testnet's nodes.
./quake info

# Check the state of the testnet by displaying the latest heights of each node
./quake info heights -n 5

# Wait for the given nodes (or all nodes if unspecified) to reach height 100
./quake wait height 100 validator1 validator2

# Wait for execution clients to finish syncing
./quake wait sync

# Apply a pause for a random amount of time to the consensus layer containers of all validators
./quake perturb pause val*_cl

# Send 1000 transactions per second during 30 seconds to one validator node
./quake load -t 30 -r 1000 --targets validator1

# Mixed EIP-1559 and legacy transfer load (70/30 split)
./quake load -t 30 -r 1000 --mix transfer=70,legacy=30 --targets validator1

# Mixed ERC-20 and native transfer load (70/30 split)
./quake load -t 30 -r 1000 --mix transfer=70,erc20=30 --targets validator1

# ERC-20 with mixed functions: 60% transfer, 30% approve, 10% transferFrom
./quake load -t 30 -r 1000 --mix erc20=100 --erc20-fn-weights transfer=60,approve=30,transfer-from=10 --targets validator1

# Stop one node (both CL and EL containers)
./quake stop val*3

# Stop all the nodes and remove generated files
./quake clean

# Shorthand: clean and start the testnet in one step
./quake restart

Command workflow

A bold arrow from command A to B means that calling B will automatically execute A. For example, ./quake start implicitly calls setup (if it finds that no config files were generated) and build (if it finds no Docker images exist).

graph LR
    init_local(("local<br>deployment")):::hidden
    init_remote(("remote<br>deployment")):::hidden

    subgraph INIT["initialize infra"]
        create["remote create"]
        build
    end

    subgraph START["<div style="white-space: nowrap;">generate config files and start testnet</div\>"]
        setup
        provision["remote provision"]
        ssm["remote ssm start"]
        start
    end

    subgraph ACTIONS["actions"]
        info
        wait
        perturb
        logs
        valset
        load
        web
        ssh["remote ssh"]
        export["remote export"]
        import["remote import"]
        monitor["remote monitor"]
    end

    subgraph STOP["stop testnet and clean data"]
        stop
        destroy["remote destroy"]
        clean
        restart
    end

    init_local ~~~ init_remote
    init_remote --> preinit
    init_local --> build
    preinit --> create ==> setup
    build ==> start --> STOP
    provision ==> setup ==> start --> ACTIONS --> STOP
    ssm ==> setup
    stop ==> clean
    destroy ==> clean
    ACTIONS ~~~ restart --> ACTIONS

    %% use invisible edges to give some order
    info ~~~ wait
    valset ~~~ perturb
    load ~~~ logs
    export ~~~ import
    ssh ~~~ monitor

    classDef hidden width:0px;

    %% remote commands
    classDef grey fill:#777
    class preinit grey
    class create grey
    class provision grey
    class ssm grey
    class destroy grey
    class ssh grey
    class export grey
    class import grey
    class monitor grey
Loading

Basic commands

Create the testnet files, including the node config files and genesis:

./quake -f crates/quake/scenarios/examples/10nodes.toml setup

This will create a directory .quake/10nodes/ with:

  • compose.yaml: Docker compose file with all the testnet containers.
  • Node directories: Each node has its configuration files and a script for configuring latency emulation (if set in the manifest; see below). These directories also will store the Malachite app and Reth databases.
  • assets/: Directory with files common to all nodes, such as genesis.json and prometheus.yml.
  • monitoring/: Directory for config files and data of Prometheus and Grafana.

Build Docker images (as defined in the generated .quake/10nodes/compose.yaml):

./quake build

Start the nodes:

./quake start

It will run setup, if not done before.

Check out the logs of a node's CL or EL:

./quake logs validator1_cl
./quake tail validator1_el

Alternatively, the logs of all CL and EL containers are accessible in .quake/<testnet_name>/logs/.

Wait for the given nodes (or all nodes if unspecified) to reach height 100:

./quake wait height 100 validator1 validator2

Wait for execution clients to finish syncing (useful after node restarts or upgrades):

./quake wait sync validator1 validator2

The sync subcommand waits until eth_syncing returns false for all specified nodes. You can customize the timeout and number of retries for transient RPC failures:

./quake wait sync --timeout 120 --max-retries 5

Tip

When a command takes node or container names directly, we can use a wildcard *. For example, val*_cl expands to the names of all consensus-layer containers of validator nodes (validator1_cl, validator2_cl, etc.). This does not apply to load or spam --targets, which accept exact node names and manifest node groups.

Apply a pause of 300ms to the consensus layer containers of all validators:

./quake perturb pause val*_cl --time-off 300ms

If no time is specified, Quake will apply the perturbation for a random amount of time. For more on perturbations, see below.

Send 1000 transactions per second during 30 seconds to one validator node:

./quake load -t 30 -r 1000 --targets validator1

Send a mixed workload (ERC-20 and native transfers) at 500 TPS:

./quake load -t 60 -r 500 --mix transfer=50,erc20=50 --targets validator1

Send ERC-20 traffic with diverse function calls (approve, transferFrom alongside transfer):

./quake load -t 60 -r 500 --mix erc20=100 --erc20-fn-weights transfer=60,approve=30,transfer-from=10 --targets validator1

Stop the nodes

./quake stop

Remove generated files

./quake clean

It will stop the nodes first if needed. Monitoring services are managed separately with quake monitoring.

Clean and restart the testnet in one step:

./quake restart

This is equivalent to running clean followed by start. It accepts clean scope flags such as --all and --data, plus the regular start flags:

# Clean everything (including monitoring data) and restart without monitoring services
./quake restart --all --monitoring=false

# Clean all nodes and restart specific nodes
./quake restart validator1 validator2

If you want to restart just some nodes leaving the others running, use quake perturb restart <nodes>.

The load and spam commands

Both commands send transaction load to a running testnet using the spammer. Transactions are dispatched over WebSocket JSON-RPC. They accept the same flags as the spammer and differ only in the adopted send mode:

Command Send mode Nonce handling Error recovery Best for
load Backpressure Advances only on acceptance Re-queries nonce on rejection, skips after 3 consecutive failures Correctness-sensitive workloads, reproducible tests
spam Fire-and-forget Incremented optimistically None (transactions may be lost) Peak-throughput stress tests

Backpressure mode (load) waits for each JSON-RPC response for a transaction before submitting the next one. If a transaction is rejected the sender re-queries the node for the correct nonce and retries. This is slower but guarantees that every submitted transaction has a valid nonce and is accepted by RPC endpoint.

Fire-and-forget mode (spam) pushes transactions into a 10,000-item buffered channel and dispatches them without waiting for responses. Nonces are incremented at generation time, so nonce gaps can occur when transactions are rejected. Use --wait-response (-w) to optionally wait for each response while still using optimistic nonces; expect multiple nonce too low errors to be produced.

Both commands support blending transaction types with --mix:

# 1000 TPS of native transfers for 30 seconds (backpressure)
./quake load -t 30 -r 1000 --targets validator1

# Same workload in fire-and-forget mode
./quake spam -t 30 -r 1000 --targets validator1

# Mixed workload: 70% native transfers, 30% ERC-20
./quake load -t 60 -r 500 --mix transfer=70,erc20=30 --targets validator1

# Gas-intensive workload with diverse guzzler functions
./quake load -t 60 -r 200 --mix guzzler=100 \
  --guzzler-fn-weights hash-loop=70@2000,storage-write=30@600 \
  --targets validator1

# Fire-and-forget at high throughput, targeting all nodes
./quake spam -t 120 -r 5000

--targets accepts a comma-separated list of explicit node names or manifest node groups such as ALL_VALIDATORS, ALL_NON_VALIDATORS, ALL_NODES, or custom groups defined under [node_groups]. If --targets is omitted, transactions are sent to all manifest nodes.

Common flags (see ./quake load --help for the full list):

Flag Short Default Description
--rate -r 1000 Target TPS across all generators
--time -t 0 Max duration in seconds (0 = unlimited)
--num-txs -n 0 Max total transactions (0 = unlimited)
--num-generators -g 1 Parallel generators, each with its own account slice
--mix transfer=100 Transaction type blend: transfer, erc20, guzzler
--tx-latency false Record submit-to-finalized latency to CSV

Latency tracking (--tx-latency)

The --tx-latency flag measures end-to-end transaction latency: the wall-clock time between eth_sendRawTransaction submission and finalized block inclusion. See the spammer README for full details on the tracking architecture, CSV output format, and analysis tools.

./quake load -t 30 -r 1000 --tx-latency --targets validator1
./quake load -t 30 -r 1000 --tx-latency --csv-dir .quake/results --targets validator1

Recording behavior by mode:

  • Backpressure (load): only transactions accepted by the node are tracked. Rejected and transient errors are skipped.
  • Fire-and-forget with --wait-response (spam -w): same, only accepted transactions.
  • Fire-and-forget without --wait-response (spam): all dispatched transactions are tracked, including those later rejected by the node. These unmatched entries remain in memory and are evicted after 5 minutes without appearing in the CSV.

Transactions that are never included in a block (dropped from mempool, rejected after submission) do not appear in the CSV.

The run saturation command

quake run saturation orchestrates a multi-phase load experiment that ramps offered TPS across a configured rate list and decides when the cluster hits saturation. Each phase: warms the spammer state, runs the spammer at the target rate for a fixed window, drains the mempool, snapshots a wide set of Prometheus + RPC metrics, and appends a PhaseRecord to experiment.json. Six saturation signals (gas plateau, TPS plateau, TPS-ratio drop, latency spike, mempool growth, EL-CPU saturation) are evaluated between adjacent phases; the first phase where any signal fires is reported as the saturation point.

The testnet must already be set up and running — the runner never starts or stops the cluster, only drives load against it. It works for both local (Docker Compose) and remote (AWS EC2) testnets transparently.

# Canonical saturation sweep against the non-validator submission tier.
# Ramps 1000 → 2000 TPS in 100-TPS increments, 3 min per phase, with 30 s
# warmup at the first rate and 30 s mempool drain between phases.
./quake run saturation \
    --targets ALL_NON_VALIDATORS \
    --rampup 30s --cooldown 30s --phase-duration 3m \
    --rates 1000-2000:100

# Same against a remote testnet.
./quake -f crates/quake/scenarios/examples/27nodes-saturation.toml \
    run saturation --targets ALL_NON_VALIDATORS --rates 1000-2000:100

# Discrete rate list instead of a range.
./quake run saturation --targets ALL_NON_VALIDATORS --rates 500,1000,2000,4000

--rates accepts either a comma-separated list of single rates (500,1000,4000) or one or more inclusive ranges with a step (1000-2000:100 → 1000, 1100, …, 2000). The parser rejects degenerate inputs (step > range, start == end, non-evenly-dividing step, expansion above 256 phases) so an operator typo fails fast instead of producing a year-long run.

Common flags (see ./quake run saturation --help for the full list):

Flag Default Description
--rates 500,1000,2000,4000 Offered-TPS sweep (singles, ranges, or a mix)
--rampup 90s Warmup phase duration at the first rate
--phase-duration / -d 5m Measured window per rate
--cooldown 90s Mempool drain between phases
--max-duration 3h Wall-clock hard limit
--generators 10 Parallel spammer generators
--targets every node Spammer submission targets (names or manifest groups)
--mix transfer=35,legacy=25,erc20=25,guzzler=15 Tx type blend
--guzzler-fn-weights hash-loop=77@200,storage-write=3@1,storage-read=20@35 Per-function weights + args for guzzler calls
--erc20-fn-weights transfer=100 Per-function weights for ERC-20 calls
--tx-input-size 0 Extra random bytes appended to each transaction's input field
--output-dir .quake/experiments Where <experiment-id>/ is written

Each run produces an artifact directory like .quake/experiments/saturation-20260626T140000Z/ containing:

  • experiment.json — full metadata, parameters, the manifest used, and every phase's metric snapshot
  • phase_<rate>/tx_latency_*.csv — per-tx submit-to-finalized latencies
  • metrics.tar.gz — Prometheus snapshot for the experiment window (remote testnets only)

The HTML report scripts in scripts/ (saturation_report.py and compare_saturation.py) consume experiment.json to render self-contained charts and per-rate delta tables; see the "Saturation reports" section below for their invocation.

Saturation reports

Two Python scripts in scripts/ render self-contained HTML reports from a saturation experiment's experiment.json. Both require pip install matplotlib jinja2.

# Single-experiment report: per-phase charts (TPS, gas/s, latency, CPU,
# memory, mempool subpools), saturation-point callout, topology summary,
# and a full manifest dump for reproducibility.
python3 scripts/saturation_report.py .quake/experiments/saturation-20260626T140000Z

# Side-by-side comparison of two experiments (e.g. baseline vs. a config
# change). Writes a single HTML file with paired charts and a per-rate
# row-by-row delta table.
python3 scripts/compare_saturation.py \
    .quake/experiments/saturation-20260626T140000Z-baseline \
    .quake/experiments/saturation-20260626T140000Z-candidate \
    /tmp/compare.html \
    --label-a baseline --label-b candidate

The reports are entirely self-contained — charts are inlined as base64 PNGs and the manifest is embedded — so they can be attached to a Jira ticket or emailed without external assets.

The perturb command

Quake offers a number of perturbations that can be applied to nodes in the testnet. These are available as subcommands of the perturb command:

  • disconnect: Disconnect the node from the network
  • kill: Kill the node using SIGKILL, wait for a given amount of time, and restart it
  • pause: Pause the node, wait for a given amount of time, and unpause it
  • restart: Restart the node
  • upgrade: Upgrade the consensus and/or execution layer of a node to newer Docker image (will stop the node to do so).
  • chaos: Apply random perturbations, at random intervals, to a list of nodes or containers.

All perturbations accept a single node, a list of nodes, a single container, or a list of containers as arguments:

  • With "node" we refer to the logical node as defined in the manifest. For example, with these manifest files:

    [[nodes]]
    [validator1]
    [validator2]

    or

    [nodes.validator1]
    [nodes.validator2]

    you should pass validator1 or validator2 as arguments to the perturb command. Passing the node name will apply the perturbation to both the consensus layer (CL) and execution layer (EL) containers of the node.

  • With "container" we refer to the actual docker container running in the testnet. For example, if you have a node named validator1, it will have two containers named validator1_cl and validator1_el. You can pass either of these container names as arguments to the perturb command to apply the perturbation to a specific container.

  • You can also use wildcards to match multiple nodes or containers. For example, val* will match all nodes starting with val, and val*_cl will match all consensus layer containers.

We'll go a bit deeper into each perturbation in the following sections.

Disconnect

The disconnect perturbation disconnects one or more nodes/containers from the network defined in the docker compose file, waits for a given, configurable amount of time, and reconnects them to the network. By default, the disconnect time is a random value between 250ms and 10s. Note that the command will return an error if the disconnect time is less than 250ms, or greater than 10s.

Usage examples:

# Disconnect a single node (both CL and EL containers) for a random amount of time
quake perturb disconnect validator1

# Disconnect multiple nodes (both CL and EL containers) for 500ms
quake perturb disconnect validator1 validator2 --time-off 500ms

# Disconnect a single consensus container for 3 seconds
quake perturb disconnect validator1_cl -t 3s

# Disconnect multiple execution containers for a random amount of time
quake perturb disconnect val*_el

Kill

The kill perturbation kills one or more nodes/containers using SIGKILL (i.e., ungraceful shutdown), waits for a given, configurable amount of time, and restarts them. By default, the time it waits before restarting them is a random value between 250ms and 10s. Note that the command will return an error if the wait time is less than 250ms, or greater than 10s.

Usage examples:

# Kill a single node (both CL and EL containers) for a random amount of time
quake perturb kill validator1

# Kill multiple nodes (both CL and EL containers) for 500ms
quake perturb kill validator1 validator2 --time-off 500ms

# Kill a single consensus container for 3 seconds
quake perturb kill validator1_cl -t 3s

# Kill multiple execution containers for a random amount of time
quake perturb kill val*_el

Pause

The pause perturbation pauses one or more nodes/containers, waits for a given, configurable amount of time, and resumes them. By default, the time it waits before resuming them is a random value between 250ms and 10s. Note that the command will return an error if the pause time is less than 250ms, or greater than 10s.

Usage examples:

# Pause a single node (both CL and EL containers) for a random amount of time
quake perturb pause validator1

# Pause multiple nodes (both CL and EL containers) for 500ms
quake perturb pause validator1 validator2 -t 500ms

# Pause a single consensus container for 3 seconds
quake perturb pause validator1_cl --time-off 3s

# Pause multiple execution containers for a random amount of time
quake perturb pause val*_el

Restart

The restart perturbation restarts one or more nodes/containers. The node/containers are gracefully stopped before restarting them, unlike the kill perturbation, which uses SIGKILL. Note that this command does not accept a time argument, since the restart is immediate.

Usage examples:

# Restart a single node (both CL and EL containers)
quake perturb restart validator1

# Restart multiple nodes (both CL and EL containers)
quake perturb restart validator1 validator2

# Restart a single consensus container
quake perturb restart validator1_cl

# Restart multiple execution containers
quake perturb restart val*_el

Upgrade

The upgrade perturbation upgrades one or more nodes/containers to a new Docker image. The node/containers are gracefully stopped before upgrading them, and then restarted with the new image. Note that this command does not accept a time argument, since the upgrade is immediate.

You must declare the new image name and tag in the manifest file before running this command, otherwise it will fail. For example:

image_cl="arc_consensus:current"     # Starting image for CL containers
image_el="arc_execution:current"     # Starting image for EL containers
image_cl_upgrade="arc_consensus:new" # Upgrade image for CL containers
image_el_upgrade="arc_execution:new" # Upgrade image for EL containers

[[nodes]]
... node definitions ...

The image_cl and image_el fields are optional and specify which Docker image to use when starting the testnet. If not specified, the default images from the deployment YAML files will be used (namely, arc_consensus:latest and arc_execution:latest).

The image_cl_upgrade and image_el_upgrade fields specify which Docker image to use when upgrading nodes with the upgrade command.

image_cl_upgrade/image_el_upgrade are global: upgrade switches every targeted node to the same upgrade image. The base image_cl/image_el, by contrast, can be overridden per node or per node group to boot a mixed-version network from genesis; see Per-node and per-group images.

Usage examples:

# Upgrade a single node (both CL and EL containers)
quake perturb upgrade validator1

# Upgrade multiple nodes (both CL and EL containers)
quake perturb upgrade validator1 validator2

Important Note: The upgrade command is a one-time operation. You should run it only once per node.

Chaos testing

The perturb chaos sub-command applies random perturbations on a running testnet.

For example, the following command will run chaos testing for 30 minutes, randomly killing, pausing, or restarting up to a third of the targeted containers, waiting between 5 s and 20 s between actions, and keeping affected containers offline for at most 1 minute.

./quake perturb --max-time-off 1m chaos --time 30m --min-wait 5s --max-wait 20s --perturbations kill,pause,restart

The valset command

This command updates the voting power of one or more validators. Under the hood, each validator sends a transaction to its local reth instance's validator manager smart contract via RPC (the updateValidatorVotingPower call).

Because each validator performs this update independently, the changes are not atomic and it may take several heights for the changes to take effect.

Example run:

quake valset validator1:30 validator2:0

Setting a validator's voting power to 0 will remove it from the validator set, while keeping its controller account.

Notes:

  • the command accepts only node names (validator1, validator2, etc.), that is, do not use container names (validator1_cl, validator2_cl).
  • it does not accept * wildcards.

The rpc command

The rpc command fans an RPC request out to one or more nodes in parallel and prints each node's result. It targets two different protocols:

Subcommand Layer Protocol Server
quake rpc el Execution Layer JSON-RPC over HTTP Reth
quake rpc cl Consensus Layer REST over HTTP Malachite
quake rpc list both shows CL endpoint catalog + Reth docs link —

Requests run concurrently against every selected node. The process exits 0 only when every node succeeded; per-node errors are reported in the output without aborting the fan-out.

quake rpc el — Execution Layer JSON-RPC

quake rpc el <METHOD> [TARGET] [PARAMS...] [--raw '<JSON>']
              [--timeout SECS] [--retries N] [--format json|table|raw]
  • <METHOD> is the JSON-RPC method, e.g. admin_clearTxpool, eth_blockNumber.
  • [TARGET] is a comma-separated list of node names or manifest groups (ALL_NODES, ALL_VALIDATORS, ALL_NON_VALIDATORS, or any custom group). When omitted, defaults to all manifest nodes.
  • [PARAMS...] are positional JSON-RPC params. Each one is auto-promoted via serde_json::from_str: 42 becomes a number, true a boolean, [1,2,3] an array, {"a":1} an object; anything that doesn't parse stays a JSON string (so 0xabc and latest pass through unchanged).
  • --raw '<JSON>' lets you supply the params as a literal JSON array (mutually exclusive with positional params, for cases where auto-promotion is awkward).

Important

When you want to pass params alongside the default ALL_NODES target, write the target slot explicitly: quake rpc el eth_getBalance ALL_NODES 0xabc latest. The first positional after the method is always the target.

Examples:

# Wipe every node's mempool (the original use case)
quake rpc el admin_clearTxpool

# Read the latest block number from every validator
quake rpc el eth_blockNumber ALL_VALIDATORS

# Single-node read with raw output suitable for piping
quake rpc el eth_blockNumber validator1 --format raw

# Account balance lookup against all nodes
quake rpc el eth_getBalance ALL_NODES 0xaaaa...bbbb latest

# Complex params via --raw
quake rpc el eth_call validator1 --raw '[{"to":"0x...","data":"0x..."},"latest"]'

Reth's full JSON-RPC reference is at https://reth.rs/jsonrpc/intro.

quake rpc cl — Consensus Layer REST

quake rpc cl <PATH> [TARGET] [--method GET|POST|DELETE|PUT|PATCH]
              [--body '<JSON>'] [--timeout SECS] [--retries N]
              [--format json|table|raw]
  • <PATH> is the REST path. The leading / is optional and prepended automatically (so consensus-state and /consensus-state are equivalent). May include a query string (e.g. '/commit?height=42').
  • [TARGET] mirrors the EL form; defaults to all consensus-enabled nodes.
  • --method defaults to GET. --body supplies a JSON body for mutating verbs.

Examples:

# Catalog of every CL endpoint (live, fetched from a running node)
quake rpc list

# Application status on every validator (leading slash optional)
quake rpc cl /status ALL_VALIDATORS
quake rpc cl status ALL_VALIDATORS

# Latest commit certificate from one validator
quake rpc cl /commit validator1

# Commit at a specific height
quake rpc cl '/commit?height=42' validator1

# Add a persistent peer
quake rpc cl /persistent-peers validator1 \
  --method POST --body '{"addr":"/ip4/.../tcp/26656/p2p/12D3KooW..."}'

Output formats

--format When to use
json (default) Newline-delimited JSON: {"node":"validator1","result":...} per line. Pipe to jq.
table Two-column `NODE
raw Prints the result value only (no node/result envelope). Requires a single target node.

The web command

The web command starts a browser-based topology viewer that visualizes the testnet in real time and allows you to control the testnet. Open http://localhost:7777 in a browser to see the web application. Currently, it only works in local mode.

There is one tab per topology:

  • Manifest: expected topology from manifest peers and subnets (always available)
  • CL Consensus / Liveness / Proposal Parts: gossipsub mesh per topic (live)
  • EL Peers: execution layer devp2p peer connections (live)

Two views: Graph (force-directed layout with subnet clustering) and Map (world map with nodes at their AWS region coordinates).

Data architecture

  • CL data (mesh topology, proposer, rounds): Fetched via HTTP from each node's /network-state and /status endpoints during each topology poll.
  • EL data (block heights, peers, mempool): Collected via a single WebSocket connection per node. Block heights arrive in real-time via eth_subscribe(newHeads). Peer data (admin_peers) and mempool status (txpool_status) are polled periodically on the same connection.
  • Container statuses: Tracked by two background tasks: a docker events subscriber for real-time state changes (start, stop, pause, die) and a periodic docker inspect poller for network disconnect detection.

For a deeper dive (server state, background tasks, topology assembly, frontend rendering pipeline), see docs/web-architecture.md.

Options

Flag Default Description
--host 127.0.0.1 Bind address for the web server
--port 7777 Web server port
--refresh-ms 1000 Frontend topology poll interval (ms)
--el-refresh-ms 1000 EL peer refresh poller interval (ms)
--container-refresh-ms 1000 Docker container status poller interval (ms)

The mcp command

The mcp command starts a Model Context Protocol (MCP) server that exposes Quake's testnet tools to AI assistants like Claude Code, Cursor, and other MCP-compatible clients. This lets you observe, manage, and test a running testnet through natural language.

The server automatically discovers the most recently used testnet via .quake/.last_manifest, so there's no need to specify a manifest path.

Quick setup

  1. Start a testnet as usual:

    ./quake -f crates/quake/scenarios/examples/10nodes.toml start
  2. Start Claude Code or your preferred MCP client (or restart it if it was already running). It will read .mcp.json and automatically spawn quake mcp as a subprocess using stdio transport.

  3. Interact with the testnet:

    > What's the current status of the testnet?
    > Pause validator3 for 5 seconds
    > Run the probe tests
    

Transport

The MCP server uses stdio transport. This is the standard mode used by Claude Code, Cursor, and similar clients that spawn the server as a subprocess.

quake mcp

Available tools

The MCP server exposes 20 tools organized into five categories:

  • Observability (read-only): testnet_status, list_nodes, get_block_heights, get_mempool, get_peers
  • Lifecycle: start_nodes, stop_nodes, restart_testnet, clean_testnet
  • Perturbations: perturb_disconnect, perturb_kill, perturb_pause, perturb_restart, perturb_upgrade
  • Testing: run_tests, wait_height, valset_update
  • Remote (remote testnets only): remote_ssh, remote_ssm, remote_provision

Resources

The server also exposes two MCP resources:

URI Description
quake://manifest The current testnet manifest (TOML configuration file)
quake://nodes All node metadata as JSON

The generate command

The generate command (alias gen) creates random manifest files for testing. It is useful for nightly or ad‑hoc runs that exercise many topologies and configurations without writing manifests by hand.

Behavior

  • Writes one or more TOML manifests into the given output directory.
  • For certain combinations of topology, height strategy, and region strategy, it generates count manifests (each with a different seed).
  • Seeding: If you pass --seed S (before the generate command), that value is used as the base seed. The first manifest gets seed S, the next gets S+1, then S+2, and so on. This makes runs reproducible: the same --seed and options produce the same manifests. Without --seed, the base seed is chosen at random (different each run), so output varies between invocations.
    • Example: quake --seed 42 gen -o out -c 2 produces 2 x 9 combinations = 18 manifests with seeds 42, 43, 44, ... , 59; running again with the same arguments yields identical manifests.

Randomization strategies

Dimension Options
Network topology 1 node | 5 nodes | complex
Height start All nodes at 0 | some nodes start at 100
Region assignment Single region | uniform random | clustered

The complex topology creates a sentry architecture with the following structure:

  • Sentry group 1: 1–3 validators (randomly chosen per manifest), fully meshed with each other and connected to sentry-1
  • Sentry group 2: 1–3 validators (randomly chosen per manifest), fully meshed with each other and connected to sentry-2
  • Sentries: sentry-1 and sentry-2 are connected to each other, to their respective validator groups, and to the relayer
  • Relayer: Connected to both sentries and to 1 full node
  • Full node: full-1 is connected to the relayer

All connections use persistent peers, creating a structured network topology that isolates validators behind sentry nodes.

For each combination, it generates count manifests (default 1). The combinations are:

  • 1 combination with single node (no region or height strategy variation)
  • 6 combinations with 5 nodes (2 height strategies × 3 region strategies)
  • 2 combinations with complex topology (2 height strategies, all nodes within a single region)

Total manifests generated = 9 × count.

Randomized per manifest

  • Consensus Layer: logging, p2p transport (tcp/quic), value_sync parameters, runtime flavor, pruning, and related options.
  • Execution Layer: txpool, builder, and engine options (within safe ranges).
  • Manifest-level: engine API connection (IPC vs RPC), initial hardfork (e.g. zero6/zero5).

Consensus, value_sync, and RPC are always enabled so that setup → start → wait height → test works on every generated manifest.

Options

Option Short Default Description
--output-dir -o .quake/generated Directory to write manifest files into.
--count -c 1 Number of manifests to generate per combination.
--seed — (random) Base seed for the RNG (use for reproducible runs).

Examples

# Generate 1 manifest per combination (9 total) with a fixed seed
quake --seed 42 generate --output-dir target/manifests

# Generate 10 manifests per combination (90 total), reproducible
quake --seed 123 generate -o target/manifests -c 10

# Generate 1 per combination with a random seed (different each run)
quake generate -o target/manifests

Nightly CI

A nightly workflow runs daily at 3 AM UTC and can be triggered:

  • Scheduled/PR runs: Use seed 42 for reproducibility.
  • Manual dispatch (workflow_dispatch): Optionally specify a base seed (must be a non-negative integer; leave blank to use 42) and a per-job count (positive integer ≤ 50; leave blank to use 3).

The workflow runs 10 parallel jobs (matrix indices 0–9). Each job computes its effective seed as BASE_SEED + 10000 × index, then calls quake --seed <SEED> generate --count <COUNT> to produce COUNT × 9 manifests (default: 3 × 9 = 27 per job, 270 total across all jobs). Each job then runs the full test pipeline on every manifest: setup → start → wait height 140 → test → clean. Logs and reports are uploaded to the artifacts bucket with per-job artifact names (-idx<N>). The driving script is scripts/scenarios/nightly-random-manifests.sh.

PRs labeled test-random will also trigger this workflow.

The clean command

By default, clean stops the testnet and removes all node data and configuration, but leaves monitoring data alone. The following flags control what is removed:

Flag Short Description
--all -a Remove everything, including monitoring services and their data. Cannot be combined with data flags.
--data -d Remove only execution and consensus layer data, preserving configuration. Cannot be combined with --execution-data or --consensus-data.
--execution-data -x Remove only execution layer (Reth) data. Cannot be combined with --data or --consensus-data.
--consensus-data -c Remove only consensus layer (Malachite) data. Cannot be combined with --data or --execution-data.
# Remove node data only (keep config, monitoring intact)
./quake clean --data

# Remove only execution layer data
./quake clean --execution-data

# Remove only consensus layer data
./quake clean --consensus-data

# Remove everything including monitoring
./quake clean --all

The monitoring command

Monitoring services (Prometheus, Grafana, cAdvisor, Blockscout) can be controlled with the monitoring command:

# Start monitoring services
./quake monitoring start

# Stop monitoring services
./quake monitoring stop

# Stop monitoring services and remove monitoring data
./quake monitoring clean

# Download a Prometheus metrics snapshot + a single-node database snapshot
./quake download

# Or download just one:
./quake download metrics
./quake download db

The download subcommand group queries the running Prometheus directly over HTTP (local Docker port for local testnets, SSM-tunnelled port for remote) and bundles each metric's query_range response into one archive. The db subcommand archives node database files (remote) or logs their on-disk paths (local). Common options:

# Limit to a time range
./quake download metrics --from 2024-01-15T10:30:00Z --to 2024-01-15T12:00:00Z

# Download only specific metrics (names go after `--`)
./quake download metrics -- reth_db_size_bytes go_goroutines

# Save to a custom path
./quake download metrics -o /tmp/my-metrics.tar.gz

# Download db from a specific node (default: first node in manifest)
./quake download db -- validator1

Without --from, the start defaults to Prometheus' headStats.minTime (the current head block start, typically the last ~2 h). Without --to, defaults to now. Without --step, the step is auto-sized to keep the response below Prometheus' 11 000-point limit. Archives land in .quake/metrics/<testnet>/ and .quake/db/<testnet>/ by default; pass -o to override. quake remote download {metrics,db} is deprecated but kept as a backward-compatible alias.

Manifest File Format

The manifest is a TOML file.

Before parsing, all ${VAR_NAME} patterns are replaced with values from the process environment and .env files. This allows any field to reference environment variables. For example:

image_cl="${IMAGE_REGISTRY_URL}/arc-consensus:abc123"

Basic Structure

Optional top-level settings:

  • name: Name of the test scenario
  • description: Description of the test scenario
  • engine_api_connection: Connection method between Consensus Layer (CL) and Execution Layer (EL). Valid values: "ipc" (default), "rpc".
  • image_cl and image_el: Docker images for CL and EL containers. If omitted, defaults to
    • for local mode: arc_consensus:latest and arc_execution:latest, or
    • for remote mode: ${IMAGE_REGISTRY_URL}/arc-consensus:<version> and ${IMAGE_REGISTRY_URL}/arc-execution:<version>, where IMAGE_REGISTRY_URL is taken from the .env file (see Custom Docker images). These are the network-wide defaults; individual nodes or node groups can override them (see Per-node and per-group images).
  • image_cl_upgrade, image_el_upgrade: Docker images to use when upgrading containers with quake perturb upgrade. Required for upgrade scenarios; not supported in remote mode.
  • group_images: Per-node-group image overrides, declared as [group_images.<group>] with image_cl/image_el keys. See Per-node and per-group images.
  • node_size: EC2 instance type for validator/full nodes (e.g. "m6a.4xlarge"). Equivalent to the --node-size CLI flag. See Instance sizing for available options. Remote mode only — ignored in local mode (a warning is printed).
  • cc_size: EC2 instance type for the Control Center. Equivalent to --cc-size. Remote mode only — ignored in local mode.
  • node_disk_gb: Root EBS volume size in GiB for each node. Must be ≥ 8. Equivalent to --node-disk-gb. Omit to keep the AMI default. Remote mode only — ignored in local mode.
  • cc_disk_gb: Root EBS volume size in GiB for the Control Center. Must be ≥ 8. Equivalent to --cc-disk-gb. Remote mode only — ignored in local mode.
  • node_volume_type: AWS EBS volume type for each node's root disk; tunes disk cost/performance (e.g. match a production disk profile). General Purpose SSD (gp2, gp3), Provisioned IOPS SSD (io1, io2), Throughput Optimized HDD (st1), Cold HDD (sc1). Default: gp3. See AWS EBS volume types. Equivalent to --node-volume-type. Remote mode only.
  • node_volume_iops: Provisioned IOPS for the node root EBS volume; raises the I/O ceiling above the volume type's baseline. Only valid with gp3, io1, io2; range 100–256000. Default: AMI's baseline IOPS for the chosen type. Equivalent to --node-volume-iops. Remote mode only.
  • node_data_on_instance_store: When true, mounts the local instance-store NVMe at the node data directory so the EL/CL databases live on local disk instead of the root EBS volume. Requires an instance type with local NVMe (e.g. i4i.*, i3.*, m6id.*); multiple instance-store volumes are striped RAID0 into one device. A no-op on instance types without instance store, leaving the data directory on EBS. Independent of node_volume_type/node_volume_iops, which keep configuring the root EBS volume. Equivalent to --node-data-on-instance-store. Remote mode only.
  • el_cpu_limit: Hard CPU cap for each EL container; reproduces production CPU quotas on the testnet. Whole or fractional CPUs (e.g. 0.5). Maps to Docker Compose cpus. Default: no limit (container uses all host CPUs).
  • el_memory_limit_gb: Hard memory cap for each EL container in GiB; fractional values (e.g. 2.5) are allowed. Maps to Docker Compose mem_limit. Default: no limit locally; 2.5 GiB on remote.
  • cl_cpu_limit: Hard CPU cap for each CL container; same semantics as el_cpu_limit.
  • cl_memory_limit_gb: Hard memory cap for each CL container in GiB; fractional values allowed. Default: no limit locally; 1 GiB on remote.

Nodes

Nodes are defined as individual TOML sections with names starting with validator or node.

[[nodes]]
[validator1]
[validator2]
[node1]
[node2]

Node Configuration

Consensus Layer (CL) configuration is set under cl.config.* keys. The schema matches the StartCmd struct in crates/malachite-cli/src/cmd/start.rs. Keys are flat and map 1:1 to the arc-node-consensus start CLI flags (e.g. cl.config.log_level = "debug" → --log-level=debug). Quake translates the merged config into CLI flags at setup time; the CL does not read a config.toml.

Matching Flags to the Target Image Version

Upgrade scenarios can pin an older arc_consensus image tag (e.g. v0.6.0). Quake derives CLI flags from the StartCmd definition compiled into its own binary, which may have gained, renamed, or removed flags since that image shipped. Before handing the flags to the container, Quake rewrites them to match the target version, i.e., older images receive a compatible subset, and "latest", missing, or unparsable tags pass through unchanged.

If pinning an older image fails with unexpected argument, that version likely needs a new compatible entry. See apply_version_compat in src/cli_version.rs for the rustdoc describing how to add one.

The default configuration of Reth (Execution Layer) is defined in crates/quake/src/manifest.rs. It can be set globally or for each node by prefixing the config field with el.config..

For example:

# Global settings that apply to all nodes
engine_api_connection = "rpc"  # or "ipc" (default)
cl.config.log_level = "debug"
el.config.disable-discovery = true

[[nodes]]
[validator1]
# Node-specific settings
cl.config.discovery_num_outbound_peers = 30
[validator2]
# Node-specific settings
el.config.builder.deadline = 5
[node1]
[node2]

In general, node configuration options are applied with the following precedence, from lowest to highest priority:

  1. Global manifest configs: defined at the top level of your manifest file
  2. Per-node manifest configs: defined within each node's section

Higher-priority configs override lower-priority ones when their keys match.

Execution Layer Configuration

In addition to the general node configuration described above, the EL configuration adds another layer of defaults defined in crates/quake/src/manifest.rs. This means that Quake manages Reth (Execution Layer) flags through a three-tier configuration system. Configs are applied with the following precedence, from lowest to highest priority:

  1. Default configs: defined in crates/quake/src/manifest.rs
  2. Global manifest configs: defined at the top level of your manifest file
  3. Per-node manifest configs: defined within each node's section

As before, higher-priority configs override lower-priority ones when their keys match. For a full list of reth CLI flags that you can set in your manifest, see the Reth documentation.

Configuration Syntax

EL configs use TOML table syntax under the el.config key. Boolean flags (like --http or --disable-discovery) use true/false, while flags with values use their appropriate types (strings, integers, arrays).

# Global EL config that applies to all nodes
[el.config]
http.enable = true
http.api = ["admin", "net", "eth"]
engine.persistence-threshold = 5
disable-discovery = true

[nodes.validator1]
# Per-node config that overrides global and defaults for this node only
el.config.engine.persistence-threshold = 10
el.config.builder.deadline = 5

[nodes.validator2]
# No per-node config, so it inherits global + defaults

[nodes.full1]
# Per-node config can also use table syntax
[nodes.full1.el.config]
txpool.nolocals = false

Special syntax notes:

  • http.enable = true produces --http (the .enable suffix is stripped)
  • ws.enable = true produces --ws
  • disable-discovery = true produces --disable-discovery
  • disable-discovery = false omits the flag entirely
  • Array values like http.api = ["admin", "net"] produce --http.api=admin,net
  • Array values are replaced, not merged. If a node defines http.api = ["admin"], it completely overrides the default array, not appends to it.
  • Quake does not validate flag names. Misspelled flags (e.g., htpp.port = 8545) will be passed to Reth, which will fail at startup with an unrecognized flag error.

Default Flags

The following flags are applied to all nodes by default. They are defined in crates/quake/src/manifest.rs and can be overridden in your manifest:

Flag Default Value Description
http.enable true Enable the HTTP-RPC server
http.api ["admin", "net", "eth", "web3", "debug", "txpool", "trace", "reth"] APIs exposed over HTTP
ws.enable true Enable the WebSocket-RPC server
ws.api ["admin", "net", "eth", "web3", "debug", "txpool", "trace", "reth"] APIs exposed over WebSocket
engine.persistence-threshold 0 Persistence threshold for engine payloads
engine.memory-block-buffer-target 0 Memory block buffer target
enable-arc-rpc true Enable Arc-specific RPC methods
rpc.txfeecap 1000 Maximum transaction fee cap
txpool.nolocals true Treat all transactions equally (no local priority)

To override a default, simply define the flag in your manifest's global or per-node el.config section.

Reserved Flags (Do Not Override)

The following flags are managed by Docker Compose templates and must not be set in manifests. If present, they are silently ignored:

Flag Value (Local) Value (Remote) Notes
datadir /data/reth/execution-data /data/reth/execution-data Data directory path
chain /app/assets/genesis.json /app/assets/genesis.json Genesis file path
http.port 8545 8545 HTTP-RPC port
http.addr 0.0.0.0 0.0.0.0 HTTP-RPC bind address
http.corsdomain * * CORS allowed origins
ws.port 8546 8546 WebSocket-RPC port
ws.addr 0.0.0.0 0.0.0.0 WebSocket-RPC bind address
ws.origins * * WebSocket allowed origins
metrics 0.0.0.0:9001 0.0.0.0:9001 Metrics endpoint
authrpc.addr 0.0.0.0 0.0.0.0 Auth server address to listen on (RPC mode)
authrpc.port 8551 8551 Auth server port to listen on (RPC mode)
authrpc.jwtsecret /app/assets/jwtsecret /assets/jwtsecret JWT secret path (RPC mode)
ipcdisable (set) (set) Disable IPC (RPC mode only)
ipcpath /sockets/reth.ipc /sockets/reth.ipc IPC socket path (IPC mode)
auth-ipc (set) (set) Enable authenticated IPC (IPC mode)
auth-ipc.path /sockets/auth.ipc /sockets/auth.ipc Auth IPC socket path (IPC mode)
p2p-secret-key /data/reth/execution-data/nodekey /data/reth/execution-data/nodekey Pre-generated secp256k1 key for P2P identity
trusted-peers (auto-generated) (auto-generated) Comma-separated enode URLs of all other nodes

The IPC vs RPC flags are automatically selected based on the connection mode. Use quake setup --rpc to switch from IPC (default) to RPC connections between the Consensus Layer and Execution Layer. You can also set this in your manifest using the engine_api_connection top-level key.

Examples

Example 1: Override a default flag globally

# Disable the txpool.nolocals default for all nodes
[el.config]
txpool.nolocals = false

[nodes.validator1]
[nodes.validator2]

Example 2: Per-node override

[nodes.validator1]
# This node uses a custom persistence threshold
el.config.engine.persistence-threshold = 10
el.config.builder.deadline = 5

[nodes.validator2]
# Uses defaults only

Example 3: Mixed global and per-node configuration

# Global: disable discovery for all nodes
[el.config]
disable-discovery = true
engine.persistence-threshold = 5

[nodes.validator1]
# Override: re-enable discovery for this node
el.config.disable-discovery = false

[nodes.validator2]
# Uses global config (discovery disabled, threshold=5)

[nodes.full1]
# Override: different persistence threshold
el.config.engine.persistence-threshold = 20

Note

When you set enable-arc-rpc = true (the default), --arc-rpc-upstream-url=<URL> is automatically added to Reth's configuration. You don't need to include it manually.

Environment Variables

In addition to CLI flags (el.config/cl.config), you can set environment variables on a node's containers via the el.env (Execution Layer) and cl.env (Consensus Layer) tables. They follow the same precedence as config: global values are inherited by every node and per-node values override matching keys.

# Global: applies to every node's containers
[el.env]
RUST_LOG = "info"
[cl.env]
RUST_LOG = "info"

[nodes.validator1.el.env]
# Override the global value for this node's EL container only
RUST_LOG = "debug,net::discovery=trace"

[nodes.validator2.cl.env]
# Halt this node's CL at a given height (testing graceful shutdown)
ARC_HALT_AT_BLOCK_HEIGHT = 100
  • el.env is applied to the EL (Reth) container; cl.env to the CL (Malachite) container. Use the right table for the layer you want to affect.
  • Keys must be valid environment variable names (^[A-Za-z_][A-Za-z0-9_]*$).
  • Values may be strings, integers, floats, or booleans; non-scalars (arrays/tables) are rejected. All values are emitted as strings.
  • Quake sets some environment variables by default (e.g. RUST_LOG and ARC_LOG_FILE on the EL container, ARC_HALT_AT_BLOCK_HEIGHT on the CL container). Setting the same key in el.env/cl.env replaces the default rather than duplicating it.
  • Works in both local and remote deployments.

Voting Power

By default every validator in genesis receives a voting power of 20. To override this, set cl_voting_power on each validator node:

[nodes.validator-1]
cl_voting_power = 2000

[nodes.validator-2]
cl_voting_power = 2000

[nodes.validator-3]
cl_voting_power = 1000

[nodes.full1]

If cl_voting_power is specified for any validator, it must be specified for all validators (all-or-nothing). This prevents accidental power imbalances where one validator silently defaults to 20 while others are set to much higher values. Non-validator nodes ignore this field.

Node Groups and Persistent Peers

Note: The cl_persistent_peers setting described below applies to the Consensus Layer (Malachite) P2P connections. Execution Layer (Reth) P2P: during setup, a secp256k1 nodekey is pre-generated for each node. Reth's --trusted-peers is built from each node's el.config.trusted_peers when set (same format as cl_persistent_peers: node names or group names, resolved to enodes); when el.config.trusted_peers is not set for a node, that node gets a full mesh of all other nodes.

You can define custom groups of nodes and use them to configure peer connections. This is useful for setting up network topologies where certain nodes should only connect to specific subsets of other nodes.

Pre-defined node groups:

  • ALL_NODES - All nodes in the manifest
  • ALL_VALIDATORS - All validator nodes, that is, nodes with names starting with val (e.g., validator1, val2)
  • ALL_NON_VALIDATORS - All nodes that are not validators

These names are reserved built-ins and cannot be redefined under [node_groups].

Custom node groups are defined in the [node_groups] section. Groups can reference individual node names, pre-defined groups, or other groups previously declared:

[node_groups]
FULL_NODES = ["full1", "full2"]
TRUSTED = ["ALL_VALIDATORS", "FULL_NODES", "other_node"]

[nodes.validator1]
cl_persistent_peers = ["TRUSTED"]
[nodes.validator2]
[nodes.validator3]
[nodes.validator4]
[nodes.full1]
cl_persistent_peers = ["ALL_NON_VALIDATORS"]
[nodes.full2]
cl_persistent_peers = ["ALL_VALIDATORS"]
[nodes.sentry]
cl_persistent_peers = ["ALL_NODES"]
[nodes.other_node]

In this example:

  • FULL_NODES is a custom group containing full1 and full2
  • TRUSTED combines the ALL_VALIDATORS group, the FULL_NODES group, and the individual node other_node
  • validator1 will have persistent peers: validator2, validator3, validator4, full1, full2, other_node (the TRUSTED group, excluding itself)
  • full1 will connect to all non-validators: full2, sentry, other_node
  • sentry will connect to all nodes except itself

The same group names can also be used as quake load and quake spam targets. For example:

./quake load -t 60 -r 500 --targets ALL_VALIDATORS
./quake spam -t 30 -r 1000 --targets TRUSTED

To distinguish group references from individual nodes in peer lists, by convention we use lowercase for node names and uppercase for node group names.

Note: A node is automatically excluded from its own persistent peers list.

Default behavior for cl_persistent_peers:

  • If cl_persistent_peers is not specified for a node, it will connect to all other nodes in the network (default behavior for simple testnets).
  • If cl_persistent_peers is specified as an empty array (cl_persistent_peers = []), the node will have no persistent peers.
  • If cl_persistent_peers is specified with values, the node will connect only to those specific peers.

el.config.trusted_peers (Execution Layer): identical behavior to cl_persistent_peers.

Per-node and per-group images

By default every node runs the global image_cl/image_el (see Basic Structure). You can override the base image for individual nodes or whole node groups, which boots a mixed-version network from genesis without a rolling perturb upgrade. Useful for cross-version consensus testing, pre-rollout validation, and reproducing version skew.

  • Per node: set image_cl/image_el under a [nodes.<name>] section.
  • Per group: set them under [group_images.<group>], keyed by any node group (custom or a built-in such as ALL_VALIDATORS); applied to every member.

Precedence, lowest to highest: global image < node-group override < per-node override. A node covered by two image-declaring groups for the same layer is rejected as ambiguous.

image_cl = "arc_consensus:latest"   # global base images
image_el = "arc_execution:latest"

[node_groups]
OLDIES = ["validator4", "validator5"]

# validator4 and validator5 run the pinned release instead of the global image
[group_images.OLDIES]
image_cl = "${IMAGE_REGISTRY_URL}/arc-consensus:0.6.0"
image_el = "${IMAGE_REGISTRY_URL}/arc-execution:0.6.0"

[nodes.validator1]
[nodes.validator4]
[nodes.full1]
image_el = "${IMAGE_REGISTRY_URL}/arc-execution:0.6.0"  # inline override wins

Compatibility (operator's responsibility). Every image in a mixed network must agree on genesis state, the hardfork schedule, and on-disk db format. Quake enforces only arc_consensus >= v0.5.0 and, in remote mode, a ghcr.io/ registry for each image; genesis or db mismatches are not caught statically and fail loudly at startup. The base image is per-node, while the image_cl_upgrade/image_el_upgrade used by perturb upgrade stay global.

A worked example ships at scenarios/mixed-version.toml.

Starting height

By default, all nodes (Consensus Layer and Execution Layer containers) will start when the start CLI command is invoked, unless a node has a start_at height set in the manifest.

[nodes.validator1]
[nodes.validator2]
[nodes.full1]
start_at = 30

Subnets

By default, all nodes are connected to a single Docker network named default. You can isolate nodes into separate sub-networks and create bridge nodes that connect multiple sub-networks by using the subnets field.

Each subnet is assigned a dedicated private IP address range. In local mode, subnets use 172.<N>.0.0/16 CIDR blocks where N starts at 21 and increments for each subnet.

Subnets work in both local and remote deployments:

  • Local: Isolation is enforced via separate Docker networks. Each subnet is assigned a dedicated private IP address range using 172.<N>.0.0/16 CIDR blocks where N starts at 21. Containers in different networks cannot communicate directly. Network perturbations (disconnect/connect) use docker network disconnect and docker network connect to detach and reattach containers from their subnet networks.
  • Remote: Isolation is enforced at the AWS infrastructure level. Each logical network maps to a separate VPC subnet with its own security group. Nodes belonging to multiple networks (bridge nodes) have multiple network interfaces (ENIs) attached, one per network. Network perturbations (disconnect/connect) use host-level iptables rules to block/unblock traffic between nodes' VPC IPs.

The following example has 5 nodes across multiple isolated networks: trusted, untrusted, and default:

[nodes.validator1]
subnets = ["trusted"]

[nodes.validator2]
subnets = ["trusted"]

[nodes.validator3]
subnets = ["trusted", "untrusted"]

[nodes.validator4]
subnets = ["untrusted", "default"]

[nodes.full1]
# No subnets specified, defaults to ["default"]

In this example:

  • validator1 and validator2 are isolated in the trusted subnet
  • validator3 bridges the trusted and untrusted subnets
  • validator4 bridges the untrusted and default subnets
  • full1 is in the default subnet

Quake validates that the network topology forms a connected graph. If networks are completely isolated from each other (no bridge nodes), the manifest validation will fail.

Implementation approach

Subnet isolation is enforced at two levels:

  1. Infrastructure level:

    • Local: Each subnet maps to a separate Docker network configured as internal: true, which prevents routing through Docker's gateway. This enforces isolation even on Docker Desktop where bridge networks would otherwise be able to communicate. Containers are also connected to a shared host-access network (non-internal) for port publishing. Network perturbations (disconnect/connect) use docker network disconnect and docker network connect to detach and reattach containers from their subnet networks.
    • Remote: Each subnet maps to a separate VPC subnet with its own security group that only allows traffic within that subnet. Bridge nodes get multiple ENIs (one per subnet). Network perturbations use iptables DROP rules installed only on the target node's EC2 host, blocking peer IPs in the INPUT, OUTPUT, and FORWARD chains. This unidirectional approach avoids altering peer hosts. On reconnect, the rules are removed and Malachite's persistent peer reconnection handles re-establishing connections on both sides.
  2. Application level: Consensus layer nodes are configured with cl_persistent_peers that only include nodes sharing at least one subnet. This ensures nodes only attempt to connect to peers they can actually reach.

Both layers are necessary for robust isolation. Infrastructure-level isolation works reliably on native Linux Docker and AWS, but Docker Desktop (Mac/Windows) does not fully isolate bridge networks. The application-level peer configuration provides defense-in-depth and ensures correct behavior across all platforms.

Latency emulation

All nodes in the testnet are typically deployed to the same private network configuration either in a local machine or remotely in one cloud region. We can emulate latency between nodes by artificially increasing the latency of outbound traffic with the Linux tc (traffic control) command.

Unless we set latency_emulation = false, latency emulation will be enabled by default. We can assign in the manifest an AWS-region to each node. Nodes that don't have an explicit region in the manifest will be assigned a random one. Then, Quake will simulate network latency between containers in different regions, using real-world average latency values between each region.

latency_emulation = true
[[nodes]]
[validator1]
region = "eu-central-1"
[validator2]
[validator3]

This latency emulation mechanism is a re-implementation of the one in CometBFT's e2e framework. In turn, the latter was adapted from https://github.com/paulo-coelho/latency-setter.

Remote deployment

Quake can deploy a testnet to remote infrastructure (AWS EC2 instances).

The remote setup consists of:

  • one EC2 instance per node, where
    • each node consists of a Consensus Layer and a Execution Layer, each in its own Docker container
  • one extra instance for a Control Center (CC) server
    • for monitoring services, and
    • for generating and sending transaction load to the nodes.

Currently all instances are deployed in one AWS region (us-east-1 by default) and we rely on latency emulation to make the node communication behavior more realistic.

graph TB
    subgraph nodes[" "]
        direction LR
        subgraph node_3["Node 3 instance"]
            direction TB
            CL3["CL"]
            EL3["EL"]
        end
        subgraph node_2["Node 2 instance"]
            direction TB
            CL2["CL"]
            EL2["EL"]
        end
        subgraph node_1["Node 1 instance"]
            direction TB
            CL1["CL"]
            EL1["EL"]
        end
    end
    CC

    CC --> node_1
    CC --> node_2
    CC --> node_3
    CL1 --- EL1
    CL2 --- EL2
    CL3 --- EL3
Loading

Communication via Control Center

All communication with remote nodes is routed through the Control Center (CC):

  • RPC requests are routed via an nginx reverse proxy on CC
  • Pprof requests are routed via a separate nginx reverse proxy on CC
  • SSH/SCP commands are routed via CC using nodes' private IPs

This architecture requires only a single SSM session to CC, avoiding AWS API throttling that would occur with per-node SSM sessions.

flowchart LR
    subgraph local [Local Machine]
        Quake
    end
    subgraph aws [AWS VPC]
        CC[Control Center]
        RPC[RPC proxy :8080]
        PProf[pprof proxy :6060]
        N0[Node 0]
        N1[Node 1]
        NN[Node N]
    end

    Quake -->|SSM Session| CC
    CC --- RPC
    CC --- PProf
    RPC -->|Private IP :8545| N0
    RPC -->|Private IP :8545| N1
    RPC -->|Private IP :8545| NN
    PProf -->|Private IP :6060/:6161| N0
    PProf -->|Private IP :6060/:6161| N1
    PProf -->|Private IP :6060/:6161| NN
    CC -->|SSH via Private IP| N0
    CC -->|SSH via Private IP| N1
    CC -->|SSH via Private IP| NN
Loading

RPC Proxy:

  • A single SSM tunnel forwards localhost:18080 to CC's port 8080
  • The nginx proxy routes requests by URL path:
    • /<node-name>/el forwards to the node's Execution Layer (Reth JSON-RPC) on port 8545
    • /<node-name>/el/ws forwards to the node's Execution Layer (Reth WebSocket) on port 8546
    • /<node-name>/cl forwards to the node's Consensus Layer (Malachite RPC) on port 31000
    • /<node-name>/cl/metrics forwards to the node's Consensus Layer metrics on port 29000
  • Additional endpoints: /health (health check), /nodes (list available nodes)

Pprof Proxy:

  • A single SSM tunnel forwards localhost:16060 to CC's port 6060
  • The nginx proxy routes by URL path: /pprof/cl/<node-name>/... forwards to the node's CL pprof (port 6060), /pprof/el/<node-name>/... to the EL pprof (port 6161)
  • The proxy has a 600-second read timeout because CPU profiling requests block for the sampling duration (?seconds=N, default 30s). For CPU profiles longer than 10 minutes, SSH into the node and query pprof locally.
  • Additional endpoints: /health (health check), /nodes (list available nodes)

SSH/SCP Routing:

  • SSH to CC uses a direct SSM session
  • SSH to nodes is routed through CC: Quake SSHs to CC, then CC SSHs to the node using its private IP
  • Parallel commands to multiple nodes run from a single SSH session to CC

Benefits:

  • Avoids AWS API throttling: only one SSM session needed instead of N+1
  • Simpler session management with only one tunnel to maintain
  • Low latency between CC and nodes (same VPC/region)

Quick start

All basic commands for local deployment work in remote mode. It just requires a few extra steps.

To quickly deploy and start a remote testnet with one command, run:

quake -f crates/quake/scenarios/examples/5nodes.toml start --remote

which is equivalent to

quake -f crates/quake/scenarios/examples/5nodes.toml remote create --yes
quake start

Check out quake remote --help for details on every sub-command.

For testnets running longer than ~20 hours, add --node-size t3.large (see Instance sizing).

Setting up the required tools to make this work requires a few extra steps, described below.

Instance sizing

The --node-size and --cc-size flags let you override the default EC2 instance types when creating remote infrastructure. The --node-disk-gb and --cc-disk-gb flags set the root EBS volume size in GiB for nodes and the Control Center; omit them to keep the AMI default volume size.

# Use larger nodes for a multi-day testnet
quake remote create --node-size t3.large --cc-size t3.2xlarge

# Larger root volume for long runs (disk fills before RAM on default volume)
quake remote create --node-size t3.large --node-disk-gb 100 --cc-disk-gb 100

# Or with the shorthand
quake start --remote --node-size t3.large --node-disk-gb 100

Local NVMe vs EBS storage

By default the node data directory lives on the root EBS volume. The --node-data-on-instance-store flag instead mounts the instance's local NVMe instance store at the data directory, so the EL/CL databases run on local disk. It requires an instance type that ships local NVMe (i4i.*, i3.*, m6id.*, c6id.*); on any other type it is a no-op and the data directory stays on EBS. Instance types with multiple instance-store volumes are striped RAID0 into a single device. The flag is independent of --node-volume-type/--node-volume-iops, which keep tuning the root EBS volume, so the two storage backends can be compared directly.

# Datadir on io2 EBS
quake remote create --node-size i4i.xlarge \
  --node-volume-type io2 --node-volume-iops 64000 --node-disk-gb 1000

# Datadir on local NVMe (same instance type, only the storage backing differs)
quake remote create --node-size i4i.xlarge --node-data-on-instance-store

The instance store is ephemeral: its contents are lost when the instance stops or terminates. That is fine for benchmarking and short-lived testnets, but never use it for state you need to keep.

Node instances

Each node runs an Execution Layer (EL) and a Consensus Layer (CL) container, plus the OS, Docker daemon, NFS client, and SSM agent. The EL is the dominant consumer of both memory (~2.5 GiB) and disk (debug logs grow at ~200 MiB/hr).

Instance vCPU RAM Max testnet duration Best for
t3.medium (default) 2 4 GiB ~20 hours Short tests, CI smoke runs
t3.large 2 8 GiB 1–3 days Day-long testnets, moderate load
t3.xlarge 4 16 GiB Multi-day Heavy load, large state, long-running

The duration estimates assume debug-level logging with no log rotation on a 20 GiB root volume. The primary constraint is disk space: the 4 GiB swap file, ~9 GiB of Docker images, and growing log files fill the default 20 GiB volume in roughly 20 hours. Larger instances do not increase disk size; use --node-disk-gb for that. Larger instances do provide more RAM headroom, reducing swap pressure and making the node more resilient to memory spikes.

Tip

For testnets that need to run longer than 20 hours, consider both upgrading the instance size (for RAM) and passing --node-disk-gb (and --cc-disk-gb if the CC needs more space) for disk.

Control Center (CC) instance

The CC runs Prometheus, Grafana, Blockscout (backend + frontend + DB), an RPC reverse proxy, a pprof reverse proxy, node-exporter, and optionally spammer containers.

Instance vCPU RAM Notes
t3.large 2 8 GiB Insufficient — Blockscout + Prometheus exceed 8 GiB
t3.xlarge (default) 4 16 GiB Standard monitoring stack
t3.2xlarge 8 32 GiB Many nodes (>15) or heavy Blockscout indexing

Requirements

Install:

  • aws CLI tool
  • aws's Session Manager plugin
  • Terraform:
    brew tap hashicorp/tap
    brew install hashicorp/tap/terraform

Custom Docker images

By default, remote nodes pull the latest arc-consensus and arc-execution Docker images from Circle's private container registry. To access them, you must provide a GitHub personal access token (PAT) that has permission to read packages.

You can also override the defaults to use images from your own registry, to test changes from a branch that hasn't been merged or published yet. For example, if stored in GHCR, set image_cl and image_el in your manifest to:

image_cl = "ghcr.io/<org|user>/<repo>/arc-consensus:<tag>"
image_el = "ghcr.io/<org|user>/<repo>/arc-execution:<tag>"

Both Terraform (for pre-pulling images during provisioning) and the per-node Docker Compose files will use these values.

The subsections below explain how to set up credentials to access the private registry and, optionally, build and push your own images.

Setting up a GitHub token

Example setup for GitHub Container Registry (ghcr.io):

  1. Visit https://github.com/settings/tokens/new
  2. Create a Classic personal access token (PAT) with the read:packages scope.
    • Make sure to select Classic PAT, not a Fine-Grained token.
    • Add a descriptive note and set an expiration date (recommended).
  3. If the image is hosted under an organization, you might need to authorize the PAT for that organization (e.g. Under your tokens list in GitHub, Configure SSO on the token, and select the org).
  4. Store your GitHub username and the generated PAT (it starts with ghp_) in a file named .env in the repository root directory.
    GITHUB_USER=<github-username>
    GITHUB_TOKEN=<github-classic-token>
    

Step-by-step: build locally and deploy to remote

  1. Prerequisites

    • Docker installed and running.
    • A .env file with GITHUB_USER and GITHUB_TOKEN as described above. Your PAT needs the additional write:packages scope (to push images) on top of the read:packages scope required for pulling.
  2. Log in to ghcr.io

    source .env
    echo $GITHUB_TOKEN | docker login ghcr.io -u $GITHUB_USER --password-stdin
  3. Build images for linux/amd64

    EC2 instances run on amd64. Do not use ./quake build for this — on Apple Silicon Macs it produces arm64 images that will fail on EC2 with a misleading unauthorized error (see troubleshooting below).

    Build each image explicitly with --platform linux/amd64. Replace <org|user> with your GitHub username (e.g. your-github-username) and <tag> with your chosen tag (e.g. dev).

    From the repository root:

    # Execution layer
    docker build --platform linux/amd64 \
      --build-context certs=deployments/certs \
      --build-arg GIT_COMMIT_HASH=$(git rev-parse HEAD) \
      --build-arg GIT_VERSION=$(git describe --tags --always) \
      --build-arg GIT_SHORT_HASH=$(git rev-parse --short HEAD) \
      --target dev-runtime \
      -t ghcr.io/<org|user>/<repo>/arc-execution:<tag> \
      -f deployments/Dockerfile.execution .
    
    # Consensus layer
    docker build --platform linux/amd64 \
      --build-context certs=deployments/certs \
      --build-arg GIT_COMMIT_HASH=$(git rev-parse HEAD) \
      --build-arg GIT_VERSION=$(git describe --tags --always) \
      --build-arg GIT_SHORT_HASH=$(git rev-parse --short HEAD) \
      --target dev-runtime \
      -t ghcr.io/<org|user>/<repo>/arc-consensus:<tag> \
      -f deployments/Dockerfile.consensus .
  4. Push images to your registry

    docker push ghcr.io/<org|user>/<repo>/arc-execution:<tag>
    docker push ghcr.io/<org|user>/<repo>/arc-consensus:<tag>
  5. Set manifest fields

    Make sure image_cl and image_el in your manifest match what you pushed:

    image_cl = "ghcr.io/<org|user>/<repo>/arc-consensus:<tag>"
    image_el = "ghcr.io/<org|user>/<repo>/arc-execution:<tag>"
  6. Create and start the remote testnet

    ./quake clean                   # remove any previous local state
    ./quake -f <path_to_the_manifest_file> start --remote     # start the network

Important

Your ghcr.io packages must remain private. EC2 nodes authenticate using the GITHUB_USER and GITHUB_TOKEN from your .env file, which are passed to the instances during provisioning.

Troubleshooting: misleading "unauthorized" errors from ghcr.io

When pulling images from ghcr.io, Docker returns an unauthorized error for several different problems, not just missing credentials. The error always looks like this:

Error response from daemon: Head "https://ghcr.io/v2/<owner>/<repo>/<image>/manifests/<tag>": unauthorized

This same error appears when:

  1. The image tag does not exist — e.g. you pushed an image with tag dev but your manifest specifies a different tag in image_cl or image_el. Docker reports unauthorized instead of "not found".
  2. The image was built for the wrong architecture — e.g. you built on an Apple Silicon Mac (arm64) and pushed the image, but EC2 nodes are amd64. The manifest exists but has no matching platform, and Docker reports unauthorized instead of a platform mismatch.
  3. Actual authorization failure — e.g. the .env token is expired, missing read:packages scope, or not authorized for SSO.

Remote commands

Initialize Terraform plugins and state. This step is required only once.

./quake remote preinit

Create EC2 instances for each node in the testnet, plus one extra for the Control Center (CC) server.

./quake [-f <manifest>] remote create [--dry-run] [--yes] [--node-size <type>] [--cc-size <type>] [--node-disk-gb <GIB>] [--cc-disk-gb <GIB>] [--node-data-on-instance-store]

See Instance sizing for recommended instance types.

This will create a <testnet-dir>/infra.json file with the names, IP addresses, and instance IDs of the created nodes.

For accessing the instances via SSH, we need to start SSM sessions that create tunnels from local ports to remote ports.

./quake remote ssm start

Note that tunnels are closed automatically after 20 minutes of inactivity. For long-running experiments, keep them alive in a separate terminal:

./quake remote ssm keep-alive 2h

Once the SSM session are established, we can log in via SSH to a node or CC, or run commands in the instances directly from the terminal:

./quake remote ssh validator1
./quake remote ssh cc docker ps

Show information on the EC2 instances just created.

./quake info

Create configuration files locally and upload them to the remote nodes.

./quake setup

After creating the configuration files, this command implicitly will run:

  • remote provision, to upload the generated files to the remote nodes, and
  • remote ssm start, to create SSM tunnels from local ports to ports in remote nodes.

Start Arc on the testnet nodes

./quake start

Check that the validators are creating blocks.

./quake info heights

Monitor health of all nodes and the Control Center:

./quake remote monitor                  # all hosts, one-shot
./quake remote monitor --follow         # all hosts, 30s continuous refresh
./quake remote monitor validator1       # single node, one-shot
./quake remote monitor validator1 -f    # single node, 5s time-series
./quake remote monitor cc -f            # Control Center only, 5s time-series
./quake remote monitor -f -i 10         # all hosts, custom 10s refresh

The dashboard combines block heights, peer counts (via RPC), and memory/CPU/disk usage with per-container memory in MiB (via SSH). Without --follow (-f), data is collected and printed once. With --follow, the multi-host dashboard refreshes continuously and the single-host view appends a new row each tick. Press Ctrl+C to stop.

Send transaction load:

./quake load --targets validator1,validator2 -r 1000 -t 60

quake load and quake spam auto-dispatch based on testnet type: local testnets run the spammer directly, remote testnets forward to the Control Center via SSH. All Spammer options are supported. --targets accepts comma-separated selectors including manifest node groups such as ALL_VALIDATORS or custom [node_groups].

Download diagnostic artifacts from the testnet via the mode-agnostic quake download command group (quake remote download {metrics,db} is deprecated but kept as a backward-compatible alias):

# Download both metrics and a single-node db snapshot
./quake download

# Metrics only — covers the current head block (~2 h) by default
./quake download metrics

# Metrics for a specific time range
./quake download metrics --from 2024-01-15T10:30:00Z --to 2024-01-15T12:00:00Z

# Specific metrics only (metric names go after --)
./quake download metrics -- reth_db_size_bytes go_goroutines

# Save to a custom output path
./quake download metrics -o /tmp/my-metrics.tar.gz

# Download node database (defaults to the first node in the manifest)
./quake download db

# Execution layer only, from specific nodes
./quake download db --execution-only -- validator1 validator2

# Save to a custom output path
./quake download db -o /tmp/my-db.tar.gz

Archives land in .quake/metrics/<testnet>/quake-metrics-<timestamp>.tar.gz and .quake/db/<testnet>/quake-db-<timestamp>.tar.gz by default; pass -o to override.

Once finished with your tests, remember to destroy the remote infrastructure!

./quake remote destroy

or

./quake clean

Sharing a remote testnet

You can give a colleague or another user access to a running remote testnet without them having to create the infrastructure. The export / import commands bundle and restore all the files Quake needs into a single JSON file.

The export bundle contains:

  • Manifest content
  • Infrastructure metadata (instance IDs, IPs)
  • SSH private key
  • Controllers config (validator keys and addresses)
  • Terraform state (so the importer can also destroy the infrastructure)

On your machine (the one that created the testnet):

./quake remote export

This writes .quake/<testnet-dir>/<testnet-name>-export.json. Send this file to your colleague.

To create a lighter bundle without Terraform state (e.g. for read-only access where the recipient won't need to destroy the infrastructure):

./quake remote export --exclude-terraform

On the colleague's machine:

./quake remote import path/to/export.json

remote import is supported on Unix-like platforms only, because Quake must set SSH private key permissions to 0600.

After importing, all Quake commands work — including quake info, quake remote ssh, and quake remote destroy. If the bundle includes Terraform state, the importer can also tear down the infrastructure when done. Bundles exported with --exclude-terraform support all commands except quake remote destroy and quake clean (which skip the infrastructure destroy step gracefully).

Important

The export file contains the SSH private key and Terraform state for the remote infrastructure. Treat it as sensitive material: transfer it securely and do not commit it to version control.

Cleaning up orphaned AWS resources

quake clean relies on Terraform state to delete the AWS infrastructure for a remote testnet. If that state is lost (e.g. the .quake/ directory was deleted, or quake remote create crashed before state was written), the nodes, VPC, security groups, and related resources stay in AWS with no supported way for quake to remove them.

crates/quake/scripts/aws-resources.sh is the recovery path. It discovers resources by the project name they were created with (arc-<testnet>-testnet-<user>) and removes them using the AWS CLI directly, no Terraform state required.

# Summarize every orphaned project for the current user.
./crates/quake/scripts/aws-resources.sh list

# Show the full resource plan for a single testnet (no deletion).
./crates/quake/scripts/aws-resources.sh list <testnet>

# Delete orphaned resources for a testnet. Without --yes, the script prints the
# plan and prompts before deleting; with --yes, it proceeds non-interactively.
./crates/quake/scripts/aws-resources.sh remove <testnet>
./crates/quake/scripts/aws-resources.sh remove <testnet> --yes

list is read-only. remove defaults to an interactive confirmation, so a bare invocation never deletes anything without an explicit y. The script scopes to the current user ($GITHUB_USER, or --user NAME); remove and list TESTNET refuse to touch a project belonging to another user unless --user is passed explicitly, and the bare list summary surfaces a notice when --user overrides the current $GITHUB_USER. Run --help for the full option set.

Important

Prefer quake clean whenever the Terraform state is still available. This script is a recovery tool for orphaned infrastructure only. It matches resources on AWS-side names and tags alone, so a mistyped testnet or the wrong --region can tear down infrastructure belonging to an active testnet.

Profiling

Both the consensus (CL) and execution (EL) binaries support heap and CPU profiling via a pprof-compatible HTTP server. This is gated behind the pprof Cargo feature flag and requires building with the profiling build profile (release optimizations + debug symbols for readable flamegraphs).

Prerequisites

Analyzing profiles requires either Go or the standalone pprof tool. Install one of them:

# Option 1: standalone pprof via Homebrew (no Go required)
brew install pprof

# Option 2: use go tool pprof (bundled with any Go installation)
# Install Go from https://go.dev/dl/

Flamegraph rendering requires Graphviz:

brew install graphviz

Feature environment variables

Each layer has its own feature environment variable so they can be configured independently:

Variable Layer Default
CL_FEATURES Consensus (empty)
EL_FEATURES Execution default js-tracer

Building with profiling

Note

quake start skips building when the image tag already exists. If you previously built without profiling, run quake build explicitly before quake start to rebuild the images with the pprof feature.

Use quake build with -p profiling to select the profiling Cargo build profile. Set the feature variable for the layer(s) you want to profile:

# CL only
CL_FEATURES=pprof ./quake -f <manifest.toml> build -p profiling

# EL only (include base features)
EL_FEATURES="default js-tracer pprof" \
  ./quake -f <manifest.toml> build -p profiling

# Both layers
CL_FEATURES=pprof EL_FEATURES="default js-tracer pprof" \
  ./quake -f <manifest.toml> build -p profiling

The same variables work with make build-docker:

CL_FEATURES=pprof BUILD_PROFILE=profiling make build-docker

After building, start the testnet normally:

./quake -f <manifest.toml> start

quake build builds for the local architecture only. For remote testnets, follow the Custom docker images instructions to build linux/amd64 images and push them to your registry. Add the profiling build args to the docker build commands:

# Execution layer — add these flags:
--build-arg BUILD_PROFILE=profiling \
--build-arg FEATURES="default js-tracer pprof"

# Consensus layer — add these flags:
--build-arg BUILD_PROFILE=profiling \
--build-arg FEATURES=pprof

Available profiles

Endpoint Type Description
/debug/pprof/allocs Heap Snapshot of in-use memory allocations
/debug/pprof/heap Heap Alias for /debug/pprof/allocs
/debug/pprof/profile?seconds=N CPU CPU sampling over N seconds

Pulling profiles

Layer Internal port Local host port
CL 6060 6060 + node_index
EL 6061 6161 + node_index * 100
# CL heap profile from the first validator
pprof -text http://localhost:6060/debug/pprof/heap
go tool pprof -text http://localhost:6060/debug/pprof/heap

# EL heap profile from the first validator
pprof -text http://localhost:6161/debug/pprof/heap
go tool pprof -text http://localhost:6161/debug/pprof/heap

# 30-second CPU profile from the first validator (CL)
pprof -text 'http://localhost:6060/debug/pprof/profile?seconds=30'
go tool pprof -text 'http://localhost:6060/debug/pprof/profile?seconds=30'

# Interactive web UI with flamegraph
pprof -http :8080 http://localhost:6060/debug/pprof/heap
go tool pprof -http :8080 http://localhost:6060/debug/pprof/heap

Note

When the pprof feature is not compiled in, the server is a no-op. Pprof ports are always mapped in compose but nothing will listen unless the binary was built with the pprof feature.

Remote testnets

Pprof requests are proxied through CC (see Communication via Control Center). Start the SSM tunnel with quake remote ssm start, then access:

# CL heap profile for first validator
pprof -http :8080 http://localhost:16060/pprof/cl/validator1/debug/pprof/allocs
go tool pprof -http :8080 http://localhost:16060/pprof/cl/validator1/debug/pprof/allocs

# EL 30-second CPU profile for first validator
pprof -text 'http://localhost:16060/pprof/el/validator1/debug/pprof/profile?seconds=30'
go tool pprof -text 'http://localhost:16060/pprof/el/validator1/debug/pprof/profile?seconds=30'

Testing with Testnet

The test command provides runtime validation of a running testnet. Tests verify connectivity, sync status, peer connections, and other operational aspects of the testnet. The idea is to have a test suite to run against arbitrary quake scenarios.

TODO: more testing scenarios to come

Running tests

Run all tests (except excluded groups: validation, health, validator_set, perf):

./quake test

Run an excluded group explicitly:

./quake test perf:block_time
./quake test validation:basic

Run all tests in a specific group:

./quake test probe

Run a single test:

./quake test probe:connectivity

Run multiple specific tests:

./quake test probe:connectivity,sync

Glob pattern matching - Use * (any characters) and ? (single character) for flexible test selection. Quote patterns to prevent shell expansion:

# Run tests in groups starting with 'n'
./quake test 'n*'

# Run all tests named 'sync' in any group
./quake test '*:sync'

# Run probe tests starting with 'conn'
./quake test 'probe:conn*'

# Run tests containing 'peer' in groups starting with 'n'
./quake test 'n*:*peer*'

# Run tests starting with 's' in groups starting with 'p'
./quake test 'p*:s*'

List tests without running - Use --dry-run to see which tests would be executed:

# List all available test groups and tests
./quake test --dry-run

# List tests in a specific group
./quake test probe --dry-run

# List tests matching a pattern
./quake test 'n*:*peer*' --dry-run

Configure RPC timeout for tests (default is 1 second):

./quake test --rpc-timeout 5s

Available test groups

probe - Basic connectivity and sync validation:

  • connectivity - Verifies all nodes are reachable via RPC and returns their current block height
  • sync - Checks that all nodes have completed syncing (not currently syncing)

net - Network peer validation:

  • peer_count - Ensures all nodes have at least one peer connection
  • cl_persistent_peers - Verifies that persistent peers defined in the manifest are actually connected

infra - Substrate-level readiness checks (excluded from the default quake test run):

  • latency_emulation - Verifies the tc netem rules inside each node's CL and EL containers match the manifest. When latency_emulation = true, cross-checks each peer's expected delay against AWS_LATENCY_MATRIX within ±10% tolerance. When latency_emulation = false, asserts no netem qdiscs exist (catches stale rules left over from a prior latency run).

    Known limitation: the check probes only the container's primary interface (eth0). Bridge nodes attached to multiple subnets (sentries, relayers) and local compose containers connected to host-access carry additional eth* interfaces that also receive tc netem rules; the check does not verify those. A node passing the check today guarantees the matrix on its primary interface only. Broken or stale rules on secondary interfaces are not caught.

The infra group is intended as a pre-experiment readiness gate on a running testnet, not a CI signal. It shell-execs into every container, so cost scales with node count; especially relevant on remote testnets where each probe involves an SSH hop.

Adding new tests

Tests automatically register themselves using the #[quake_test] macro. To add new tests:

  1. Create a new test file in crates/quake/src/tests/ (e.g., foobar.rs) or add to an existing file:

    use tracing::debug;
    
    use super::{quake_test, in_parallel, CheckResult, RpcClientFactory, TestOutcome, TestResult};
    use crate::testnet::Testnet;
    
    /// Test description
    #[quake_test(group = "foobar", name = "my_test")]
    fn my_test<'a>(
        testnet: &'a Testnet,
        factory: &'a RpcClientFactory,
    ) -> TestResult<'a> {
        Box::pin(async move {
            debug!("Running my test...");
    
            // Example: Check all nodes are reachable
            let node_urls = testnet.nodes_metadata.all_execution_urls();
            let results = in_parallel(&node_urls, factory, |client| async move {
                client.get_block_number().await
            })
            .await;
    
            // Use structured test results
            let mut outcome = TestOutcome::new();
    
            for (name, url, result) in results {
                match result {
                    Ok(block_number) => {
                        outcome.add_check(CheckResult::success(
                            name,
                            format!("{} (block #{})", url, block_number),
                        ));
                    }
                    Err(e) => {
                        outcome.add_check(CheckResult::failure(
                            name,
                            format!("{} - Error: {}", url, e),
                        ));
                    }
                }
            }
    
            outcome.with_summary("All nodes checked").into_result()
        })
    }
  2. Add module declaration in crates/quake/src/tests/mod.rs:

    mod foobar;
  3. Run your new test:

    ./quake test foobar:my_test

Key points:

  • Tests register automatically via #[quake_test(group = "...", name = "...")] macro
  • Each group:name combination must be unique (enforced at compile time)
  • Tests use async functions that return Pin<Box<dyn Future<Output = Result<()>>>> wrapped in Box::pin(async move { ... })
  • The RpcClientFactory provides RPC clients with consistent timeout configuration
  • Use the in_parallel() helper for concurrent RPC operations across nodes
  • Use TestOutcome and CheckResult for structured test reporting with consistent formatting
  • All test output goes to stdout with ✓ for success and ✗ for failure

nightly-chaos-testing

Tests network resilience under chaos conditions with transaction load and validator set changes.

./scripts/scenarios/nightly-chaos-testing.sh [scenario] [spam_duration_in_seconds] [tx_rate]

Our CI workflow runs this script every night with:

./scripts/scenarios/nightly-chaos-testing.sh crates/quake/scenarios/nightly-chaos-testing.toml 3600 1000

Which loads the network for 3600s (1 hour) at a 1000tx/s rate, while continuously running chaos testing.