Quake is a tool for deploying Arc testnets and running end-to-end tests.
Quake testnets vs. the Arc Testnet: Arc has a public, persistent Testnet open to external developers and validators. Quake testnets are different: they are private, ephemeral networks spun up on demand for development and CI, then torn down when testing is complete. All mentions of "testnet" in this repository refer to Quake testnets unless explicitly stated otherwise.
- Describe the testnet topology, node configurations, and test scenarios in TOML manifest files.
- CLI interface for deploying, managing, monitor testnets, and run tests on them.
- Deploy testnets locally in your machine or remotely in AWS infrastructure.
- Test the resilience of a running testnet:
- perturb the nodes (disconnect, kill, pause, and restart them),
- perform chaos testing (apply random perturbations, at random intervals, to random subsets of nodes or containers),
- change the voting power of validators in the validator set of a node,
- upgrade the running version of individual nodes.
- Emulate network latency between nodes by assigning data-center regions to nodes and injecting artificial latency between regions.
- Web-based topology viewer for real-time visualization of nodes, connections, peer status, and network health.
- MCP (Model Context Protocol) server for AI-assisted testnet management via Claude Code, Cursor, and other MCP-compatible clients.
Table of contents
- Quake - End-to-end testing management tool - Features
Build Quake from the repository root directory:
cargo build -p quake
ln -s target/debug/quake quake # create a symbolic link to simplify calling the toolDeploying and starting a local testnet is as simple as running start:
./quake -f crates/quake/scenarios/examples/10nodes.toml startThis will generate all necessary files to run the testnet and start running Arc in the nodes.
Now you can display comprehensive information about the testnet and check its status:
./quake infoThe start command implicitly performs other commands (mainly build and setup), which we describe below.
Check out the command workflow for a complete picture of all available commands.
Generate the files needed to start the testnet. If needed, you can manually modify the generated configuration files before starting the nodes.
./quake -f crates/quake/scenarios/examples/10nodes.toml setupTip
A command with the -f or --file option creates or updates a
.quake/.last_manifest file containing the path to the most recently used
manifest. This allows subsequent Quake commands to automatically reference the
last manifest without requiring the -f flag.
# Build Docker images for the Consensus Layer (CL) and Execution Layer (EL)
./quake build
# Start the nodes. Implicitly calls setup.
./quake start
# Shows comprehensive information about the testnet's nodes.
./quake info
# Check the state of the testnet by displaying the latest heights of each node
./quake info heights -n 5
# Wait for the given nodes (or all nodes if unspecified) to reach height 100
./quake wait height 100 validator1 validator2
# Wait for execution clients to finish syncing
./quake wait sync
# Apply a pause for a random amount of time to the consensus layer containers of all validators
./quake perturb pause val*_cl
# Send 1000 transactions per second during 30 seconds to one validator node
./quake load -t 30 -r 1000 --targets validator1
# Mixed EIP-1559 and legacy transfer load (70/30 split)
./quake load -t 30 -r 1000 --mix transfer=70,legacy=30 --targets validator1
# Mixed ERC-20 and native transfer load (70/30 split)
./quake load -t 30 -r 1000 --mix transfer=70,erc20=30 --targets validator1
# ERC-20 with mixed functions: 60% transfer, 30% approve, 10% transferFrom
./quake load -t 30 -r 1000 --mix erc20=100 --erc20-fn-weights transfer=60,approve=30,transfer-from=10 --targets validator1
# Stop one node (both CL and EL containers)
./quake stop val*3
# Stop all the nodes and remove generated files
./quake clean
# Shorthand: clean and start the testnet in one step
./quake restartA bold arrow from command A to B means that calling B will automatically execute
A. For example, ./quake start implicitly calls setup (if it finds that no
config files were generated) and build (if it finds no Docker images exist).
graph LR
init_local(("local<br>deployment")):::hidden
init_remote(("remote<br>deployment")):::hidden
subgraph INIT["initialize infra"]
create["remote create"]
build
end
subgraph START["<div style="white-space: nowrap;">generate config files and start testnet</div\>"]
setup
provision["remote provision"]
ssm["remote ssm start"]
start
end
subgraph ACTIONS["actions"]
info
wait
perturb
logs
valset
load
web
ssh["remote ssh"]
export["remote export"]
import["remote import"]
monitor["remote monitor"]
end
subgraph STOP["stop testnet and clean data"]
stop
destroy["remote destroy"]
clean
restart
end
init_local ~~~ init_remote
init_remote --> preinit
init_local --> build
preinit --> create ==> setup
build ==> start --> STOP
provision ==> setup ==> start --> ACTIONS --> STOP
ssm ==> setup
stop ==> clean
destroy ==> clean
ACTIONS ~~~ restart --> ACTIONS
%% use invisible edges to give some order
info ~~~ wait
valset ~~~ perturb
load ~~~ logs
export ~~~ import
ssh ~~~ monitor
classDef hidden width:0px;
%% remote commands
classDef grey fill:#777
class preinit grey
class create grey
class provision grey
class ssm grey
class destroy grey
class ssh grey
class export grey
class import grey
class monitor grey
Create the testnet files, including the node config files and genesis:
./quake -f crates/quake/scenarios/examples/10nodes.toml setupThis will create a directory .quake/10nodes/ with:
compose.yaml: Docker compose file with all the testnet containers.- Node directories: Each node has its configuration files and a script for configuring latency emulation (if set in the manifest; see below). These directories also will store the Malachite app and Reth databases.
assets/: Directory with files common to all nodes, such asgenesis.jsonandprometheus.yml.monitoring/: Directory for config files and data of Prometheus and Grafana.
Build Docker images (as defined in the generated .quake/10nodes/compose.yaml):
./quake buildStart the nodes:
./quake startIt will run setup, if not done before.
Check out the logs of a node's CL or EL:
./quake logs validator1_cl
./quake tail validator1_el
Alternatively, the logs of all CL and EL containers are accessible in .quake/<testnet_name>/logs/.
Wait for the given nodes (or all nodes if unspecified) to reach height 100:
./quake wait height 100 validator1 validator2Wait for execution clients to finish syncing (useful after node restarts or upgrades):
./quake wait sync validator1 validator2The sync subcommand waits until eth_syncing returns false for all specified nodes.
You can customize the timeout and number of retries for transient RPC failures:
./quake wait sync --timeout 120 --max-retries 5Tip
When a command takes node or container names directly, we can use a
wildcard *. For example, val*_cl expands to the names of all
consensus-layer containers of validator nodes (validator1_cl,
validator2_cl, etc.). This does not apply to load or spam
--targets, which accept exact node names and manifest node groups.
Apply a pause of 300ms to the consensus layer containers of all validators:
./quake perturb pause val*_cl --time-off 300msIf no time is specified, Quake will apply the perturbation for a random amount of time. For more on perturbations, see below.
Send 1000 transactions per second during 30 seconds to one validator node:
./quake load -t 30 -r 1000 --targets validator1Send a mixed workload (ERC-20 and native transfers) at 500 TPS:
./quake load -t 60 -r 500 --mix transfer=50,erc20=50 --targets validator1Send ERC-20 traffic with diverse function calls (approve, transferFrom alongside transfer):
./quake load -t 60 -r 500 --mix erc20=100 --erc20-fn-weights transfer=60,approve=30,transfer-from=10 --targets validator1Stop the nodes
./quake stopRemove generated files
./quake cleanIt will stop the nodes first if needed. Monitoring services are managed
separately with quake monitoring.
Clean and restart the testnet in one step:
./quake restartThis is equivalent to running clean followed by start. It accepts clean
scope flags such as --all and --data, plus the regular start flags:
# Clean everything (including monitoring data) and restart without monitoring services
./quake restart --all --monitoring=false
# Clean all nodes and restart specific nodes
./quake restart validator1 validator2If you want to restart just some nodes leaving the others running, use quake perturb restart <nodes>.
Both commands send transaction load to a running testnet using the spammer. Transactions are dispatched over WebSocket JSON-RPC. They accept the same flags as the spammer and differ only in the adopted send mode:
| Command | Send mode | Nonce handling | Error recovery | Best for |
|---|---|---|---|---|
load |
Backpressure | Advances only on acceptance | Re-queries nonce on rejection, skips after 3 consecutive failures | Correctness-sensitive workloads, reproducible tests |
spam |
Fire-and-forget | Incremented optimistically | None (transactions may be lost) | Peak-throughput stress tests |
Backpressure mode (load) waits for each JSON-RPC response for a transaction before submitting
the next one. If a transaction is rejected the sender re-queries the
node for the correct nonce and retries. This is slower but guarantees that every
submitted transaction has a valid nonce and is accepted by RPC endpoint.
Fire-and-forget mode (spam) pushes transactions into a 10,000-item
buffered channel and dispatches them without waiting for responses. Nonces are
incremented at generation time, so nonce gaps can occur when transactions are
rejected. Use --wait-response (-w) to optionally wait for each response
while still using optimistic nonces; expect multiple nonce too low errors to be produced.
Both commands support blending transaction types with --mix:
# 1000 TPS of native transfers for 30 seconds (backpressure)
./quake load -t 30 -r 1000 --targets validator1
# Same workload in fire-and-forget mode
./quake spam -t 30 -r 1000 --targets validator1
# Mixed workload: 70% native transfers, 30% ERC-20
./quake load -t 60 -r 500 --mix transfer=70,erc20=30 --targets validator1
# Gas-intensive workload with diverse guzzler functions
./quake load -t 60 -r 200 --mix guzzler=100 \
--guzzler-fn-weights hash-loop=70@2000,storage-write=30@600 \
--targets validator1
# Fire-and-forget at high throughput, targeting all nodes
./quake spam -t 120 -r 5000--targets accepts a comma-separated list of explicit node names or manifest
node groups such as ALL_VALIDATORS, ALL_NON_VALIDATORS, ALL_NODES, or
custom groups defined under [node_groups]. If --targets is omitted,
transactions are sent to all manifest nodes.
Common flags (see ./quake load --help for the full list):
| Flag | Short | Default | Description |
|---|---|---|---|
--rate |
-r |
1000 |
Target TPS across all generators |
--time |
-t |
0 |
Max duration in seconds (0 = unlimited) |
--num-txs |
-n |
0 |
Max total transactions (0 = unlimited) |
--num-generators |
-g |
1 |
Parallel generators, each with its own account slice |
--mix |
transfer=100 |
Transaction type blend: transfer, erc20, guzzler |
|
--tx-latency |
false |
Record submit-to-finalized latency to CSV |
The --tx-latency flag measures end-to-end transaction latency: the wall-clock
time between eth_sendRawTransaction submission and finalized block inclusion.
See the spammer README for
full details on the tracking architecture, CSV output format, and analysis tools.
./quake load -t 30 -r 1000 --tx-latency --targets validator1
./quake load -t 30 -r 1000 --tx-latency --csv-dir .quake/results --targets validator1Recording behavior by mode:
- Backpressure (
load): only transactions accepted by the node are tracked. Rejected and transient errors are skipped. - Fire-and-forget with
--wait-response(spam -w): same, only accepted transactions. - Fire-and-forget without
--wait-response(spam): all dispatched transactions are tracked, including those later rejected by the node. These unmatched entries remain in memory and are evicted after 5 minutes without appearing in the CSV.
Transactions that are never included in a block (dropped from mempool, rejected after submission) do not appear in the CSV.
quake run saturation orchestrates a multi-phase load experiment that ramps
offered TPS across a configured rate list and decides when the cluster hits
saturation. Each phase: warms the spammer state, runs the spammer at the
target rate for a fixed window, drains the mempool, snapshots a wide set of
Prometheus + RPC metrics, and appends a PhaseRecord to experiment.json.
Six saturation signals (gas plateau, TPS plateau, TPS-ratio drop, latency
spike, mempool growth, EL-CPU saturation) are evaluated between adjacent
phases; the first phase where any signal fires is reported as the saturation
point.
The testnet must already be set up and running — the runner never starts or stops the cluster, only drives load against it. It works for both local (Docker Compose) and remote (AWS EC2) testnets transparently.
# Canonical saturation sweep against the non-validator submission tier.
# Ramps 1000 → 2000 TPS in 100-TPS increments, 3 min per phase, with 30 s
# warmup at the first rate and 30 s mempool drain between phases.
./quake run saturation \
--targets ALL_NON_VALIDATORS \
--rampup 30s --cooldown 30s --phase-duration 3m \
--rates 1000-2000:100
# Same against a remote testnet.
./quake -f crates/quake/scenarios/examples/27nodes-saturation.toml \
run saturation --targets ALL_NON_VALIDATORS --rates 1000-2000:100
# Discrete rate list instead of a range.
./quake run saturation --targets ALL_NON_VALIDATORS --rates 500,1000,2000,4000--rates accepts either a comma-separated list of single rates (500,1000,4000)
or one or more inclusive ranges with a step (1000-2000:100 → 1000, 1100, …,
2000). The parser rejects degenerate inputs (step > range, start == end,
non-evenly-dividing step, expansion above 256 phases) so an operator typo
fails fast instead of producing a year-long run.
Common flags (see ./quake run saturation --help for the full list):
| Flag | Default | Description |
|---|---|---|
--rates |
500,1000,2000,4000 |
Offered-TPS sweep (singles, ranges, or a mix) |
--rampup |
90s |
Warmup phase duration at the first rate |
--phase-duration / -d |
5m |
Measured window per rate |
--cooldown |
90s |
Mempool drain between phases |
--max-duration |
3h |
Wall-clock hard limit |
--generators |
10 |
Parallel spammer generators |
--targets |
every node | Spammer submission targets (names or manifest groups) |
--mix |
transfer=35,legacy=25,erc20=25,guzzler=15 |
Tx type blend |
--guzzler-fn-weights |
hash-loop=77@200,storage-write=3@1,storage-read=20@35 |
Per-function weights + args for guzzler calls |
--erc20-fn-weights |
transfer=100 |
Per-function weights for ERC-20 calls |
--tx-input-size |
0 |
Extra random bytes appended to each transaction's input field |
--output-dir |
.quake/experiments |
Where <experiment-id>/ is written |
Each run produces an artifact directory like
.quake/experiments/saturation-20260626T140000Z/ containing:
experiment.json— full metadata, parameters, the manifest used, and every phase's metric snapshotphase_<rate>/tx_latency_*.csv— per-tx submit-to-finalized latenciesmetrics.tar.gz— Prometheus snapshot for the experiment window (remote testnets only)
The HTML report scripts in scripts/ (saturation_report.py
and compare_saturation.py) consume experiment.json to render
self-contained charts and per-rate delta tables; see the "Saturation reports"
section below for their invocation.
Two Python scripts in scripts/ render self-contained HTML
reports from a saturation experiment's experiment.json. Both require
pip install matplotlib jinja2.
# Single-experiment report: per-phase charts (TPS, gas/s, latency, CPU,
# memory, mempool subpools), saturation-point callout, topology summary,
# and a full manifest dump for reproducibility.
python3 scripts/saturation_report.py .quake/experiments/saturation-20260626T140000Z
# Side-by-side comparison of two experiments (e.g. baseline vs. a config
# change). Writes a single HTML file with paired charts and a per-rate
# row-by-row delta table.
python3 scripts/compare_saturation.py \
.quake/experiments/saturation-20260626T140000Z-baseline \
.quake/experiments/saturation-20260626T140000Z-candidate \
/tmp/compare.html \
--label-a baseline --label-b candidateThe reports are entirely self-contained — charts are inlined as base64 PNGs and the manifest is embedded — so they can be attached to a Jira ticket or emailed without external assets.
Quake offers a number of perturbations that can be applied to nodes in the testnet.
These are available as subcommands of the perturb command:
disconnect: Disconnect the node from the networkkill: Kill the node usingSIGKILL, wait for a given amount of time, and restart itpause: Pause the node, wait for a given amount of time, and unpause itrestart: Restart the nodeupgrade: Upgrade the consensus and/or execution layer of a node to newer Docker image (will stop the node to do so).chaos: Apply random perturbations, at random intervals, to a list of nodes or containers.
All perturbations accept a single node, a list of nodes, a single container, or a list of containers as arguments:
-
With "node" we refer to the logical node as defined in the manifest. For example, with these manifest files:
[[nodes]] [validator1] [validator2]
or
[nodes.validator1] [nodes.validator2]
you should pass
validator1orvalidator2as arguments to theperturbcommand. Passing the node name will apply the perturbation to both the consensus layer (CL) and execution layer (EL) containers of the node. -
With "container" we refer to the actual docker container running in the testnet. For example, if you have a node named
validator1, it will have two containers namedvalidator1_clandvalidator1_el. You can pass either of these container names as arguments to theperturbcommand to apply the perturbation to a specific container. -
You can also use wildcards to match multiple nodes or containers. For example,
val*will match all nodes starting withval, andval*_clwill match all consensus layer containers.
We'll go a bit deeper into each perturbation in the following sections.
The disconnect perturbation disconnects one or more nodes/containers from the
network defined in the docker compose file, waits for a given, configurable amount of
time, and reconnects them to the network.
By default, the disconnect time is a random value between 250ms and 10s.
Note that the command will return an error if the disconnect time is less than 250ms,
or greater than 10s.
Usage examples:
# Disconnect a single node (both CL and EL containers) for a random amount of time
quake perturb disconnect validator1
# Disconnect multiple nodes (both CL and EL containers) for 500ms
quake perturb disconnect validator1 validator2 --time-off 500ms
# Disconnect a single consensus container for 3 seconds
quake perturb disconnect validator1_cl -t 3s
# Disconnect multiple execution containers for a random amount of time
quake perturb disconnect val*_el
The kill perturbation kills one or more nodes/containers using SIGKILL (i.e.,
ungraceful shutdown), waits for a given, configurable amount of time, and restarts them.
By default, the time it waits before restarting them is a random value between 250ms and 10s.
Note that the command will return an error if the wait time is less than 250ms,
or greater than 10s.
Usage examples:
# Kill a single node (both CL and EL containers) for a random amount of time
quake perturb kill validator1
# Kill multiple nodes (both CL and EL containers) for 500ms
quake perturb kill validator1 validator2 --time-off 500ms
# Kill a single consensus container for 3 seconds
quake perturb kill validator1_cl -t 3s
# Kill multiple execution containers for a random amount of time
quake perturb kill val*_el
The pause perturbation pauses one or more nodes/containers, waits for a given,
configurable amount of time, and resumes them.
By default, the time it waits before resuming them is a random value between 250ms
and 10s.
Note that the command will return an error if the pause time is less than 250ms,
or greater than 10s.
Usage examples:
# Pause a single node (both CL and EL containers) for a random amount of time
quake perturb pause validator1
# Pause multiple nodes (both CL and EL containers) for 500ms
quake perturb pause validator1 validator2 -t 500ms
# Pause a single consensus container for 3 seconds
quake perturb pause validator1_cl --time-off 3s
# Pause multiple execution containers for a random amount of time
quake perturb pause val*_el
The restart perturbation restarts one or more nodes/containers.
The node/containers are gracefully stopped before restarting them, unlike the kill
perturbation, which uses SIGKILL.
Note that this command does not accept a time argument, since the restart is immediate.
Usage examples:
# Restart a single node (both CL and EL containers)
quake perturb restart validator1
# Restart multiple nodes (both CL and EL containers)
quake perturb restart validator1 validator2
# Restart a single consensus container
quake perturb restart validator1_cl
# Restart multiple execution containers
quake perturb restart val*_el
The upgrade perturbation upgrades one or more nodes/containers to a new Docker image.
The node/containers are gracefully stopped before upgrading them, and then restarted
with the new image.
Note that this command does not accept a time argument, since the upgrade is immediate.
You must declare the new image name and tag in the manifest file before running this command, otherwise it will fail. For example:
image_cl="arc_consensus:current" # Starting image for CL containers
image_el="arc_execution:current" # Starting image for EL containers
image_cl_upgrade="arc_consensus:new" # Upgrade image for CL containers
image_el_upgrade="arc_execution:new" # Upgrade image for EL containers
[[nodes]]
... node definitions ...The image_cl and image_el fields are optional and specify which Docker image
to use when starting the testnet. If not specified, the default images
from the deployment YAML files will be used (namely, arc_consensus:latest and
arc_execution:latest).
The image_cl_upgrade and image_el_upgrade fields specify which Docker image
to use when upgrading nodes with the upgrade command.
image_cl_upgrade/image_el_upgrade are global: upgrade switches every
targeted node to the same upgrade image. The base image_cl/image_el, by
contrast, can be overridden per node or per node group to boot a mixed-version
network from genesis; see
Per-node and per-group images.
Usage examples:
# Upgrade a single node (both CL and EL containers)
quake perturb upgrade validator1
# Upgrade multiple nodes (both CL and EL containers)
quake perturb upgrade validator1 validator2
Important Note: The upgrade command is a one-time operation.
You should run it only once per node.
The perturb chaos sub-command applies random perturbations on a running testnet.
For example, the following command will run chaos testing for 30 minutes, randomly killing, pausing, or restarting up to a third of the targeted containers, waiting between 5 s and 20 s between actions, and keeping affected containers offline for at most 1 minute.
./quake perturb --max-time-off 1m chaos --time 30m --min-wait 5s --max-wait 20s --perturbations kill,pause,restart
This command updates the voting power of one or more validators.
Under the hood, each validator sends a transaction to its local reth instance's
validator manager smart contract via RPC (the updateValidatorVotingPower call).
Because each validator performs this update independently, the changes are not atomic and it may take several heights for the changes to take effect.
Example run:
quake valset validator1:30 validator2:0Setting a validator's voting power to 0 will remove it from the validator set, while keeping its controller account.
Notes:
- the command accepts only node names (
validator1,validator2, etc.), that is, do not use container names (validator1_cl,validator2_cl). - it does not accept
*wildcards.
The rpc command fans an RPC request out to one or more nodes in parallel and
prints each node's result. It targets two different protocols:
| Subcommand | Layer | Protocol | Server |
|---|---|---|---|
quake rpc el |
Execution Layer | JSON-RPC over HTTP | Reth |
quake rpc cl |
Consensus Layer | REST over HTTP | Malachite |
quake rpc list |
both | shows CL endpoint catalog + Reth docs link | — |
Requests run concurrently against every selected node. The process exits 0 only when every node succeeded; per-node errors are reported in the output without aborting the fan-out.
quake rpc el <METHOD> [TARGET] [PARAMS...] [--raw '<JSON>']
[--timeout SECS] [--retries N] [--format json|table|raw]
<METHOD>is the JSON-RPC method, e.g.admin_clearTxpool,eth_blockNumber.[TARGET]is a comma-separated list of node names or manifest groups (ALL_NODES,ALL_VALIDATORS,ALL_NON_VALIDATORS, or any custom group). When omitted, defaults to all manifest nodes.[PARAMS...]are positional JSON-RPC params. Each one is auto-promoted viaserde_json::from_str:42becomes a number,truea boolean,[1,2,3]an array,{"a":1}an object; anything that doesn't parse stays a JSON string (so0xabcandlatestpass through unchanged).--raw '<JSON>'lets you supply the params as a literal JSON array (mutually exclusive with positional params, for cases where auto-promotion is awkward).
Important
When you want to pass params alongside the default ALL_NODES target, write
the target slot explicitly: quake rpc el eth_getBalance ALL_NODES 0xabc latest.
The first positional after the method is always the target.
Examples:
# Wipe every node's mempool (the original use case)
quake rpc el admin_clearTxpool
# Read the latest block number from every validator
quake rpc el eth_blockNumber ALL_VALIDATORS
# Single-node read with raw output suitable for piping
quake rpc el eth_blockNumber validator1 --format raw
# Account balance lookup against all nodes
quake rpc el eth_getBalance ALL_NODES 0xaaaa...bbbb latest
# Complex params via --raw
quake rpc el eth_call validator1 --raw '[{"to":"0x...","data":"0x..."},"latest"]'Reth's full JSON-RPC reference is at https://reth.rs/jsonrpc/intro.
quake rpc cl <PATH> [TARGET] [--method GET|POST|DELETE|PUT|PATCH]
[--body '<JSON>'] [--timeout SECS] [--retries N]
[--format json|table|raw]
<PATH>is the REST path. The leading/is optional and prepended automatically (soconsensus-stateand/consensus-stateare equivalent). May include a query string (e.g.'/commit?height=42').[TARGET]mirrors the EL form; defaults to all consensus-enabled nodes.--methoddefaults toGET.--bodysupplies a JSON body for mutating verbs.
Examples:
# Catalog of every CL endpoint (live, fetched from a running node)
quake rpc list
# Application status on every validator (leading slash optional)
quake rpc cl /status ALL_VALIDATORS
quake rpc cl status ALL_VALIDATORS
# Latest commit certificate from one validator
quake rpc cl /commit validator1
# Commit at a specific height
quake rpc cl '/commit?height=42' validator1
# Add a persistent peer
quake rpc cl /persistent-peers validator1 \
--method POST --body '{"addr":"/ip4/.../tcp/26656/p2p/12D3KooW..."}'--format |
When to use |
|---|---|
json (default) |
Newline-delimited JSON: {"node":"validator1","result":...} per line. Pipe to jq. |
table |
Two-column `NODE |
raw |
Prints the result value only (no node/result envelope). Requires a single target node. |
The web command starts a browser-based topology viewer that visualizes the testnet in real time and allows you to control the testnet.
Open http://localhost:7777 in a browser to see the web application.
Currently, it only works in local mode.
There is one tab per topology:
- Manifest: expected topology from manifest peers and subnets (always available)
- CL Consensus / Liveness / Proposal Parts: gossipsub mesh per topic (live)
- EL Peers: execution layer devp2p peer connections (live)
Two views: Graph (force-directed layout with subnet clustering) and Map (world map with nodes at their AWS region coordinates).
- CL data (mesh topology, proposer, rounds): Fetched via HTTP from each node's
/network-stateand/statusendpoints during each topology poll. - EL data (block heights, peers, mempool): Collected via a single WebSocket connection per node. Block heights arrive in real-time via
eth_subscribe(newHeads). Peer data (admin_peers) and mempool status (txpool_status) are polled periodically on the same connection. - Container statuses: Tracked by two background tasks: a
docker eventssubscriber for real-time state changes (start, stop, pause, die) and a periodicdocker inspectpoller for network disconnect detection.
For a deeper dive (server state, background tasks, topology assembly, frontend rendering pipeline), see docs/web-architecture.md.
| Flag | Default | Description |
|---|---|---|
--host |
127.0.0.1 |
Bind address for the web server |
--port |
7777 |
Web server port |
--refresh-ms |
1000 |
Frontend topology poll interval (ms) |
--el-refresh-ms |
1000 |
EL peer refresh poller interval (ms) |
--container-refresh-ms |
1000 |
Docker container status poller interval (ms) |
The mcp command starts a Model Context Protocol (MCP) server that exposes Quake's testnet tools to AI assistants like Claude Code, Cursor, and other MCP-compatible clients. This lets you observe, manage, and test a running testnet through natural language.
The server automatically discovers the most recently used testnet via .quake/.last_manifest, so there's no need to specify a manifest path.
-
Start a testnet as usual:
./quake -f crates/quake/scenarios/examples/10nodes.toml start
-
Start Claude Code or your preferred MCP client (or restart it if it was already running). It will read
.mcp.jsonand automatically spawnquake mcpas a subprocess using stdio transport. -
Interact with the testnet:
> What's the current status of the testnet? > Pause validator3 for 5 seconds > Run the probe tests
The MCP server uses stdio transport. This is the standard mode used by Claude Code, Cursor, and similar clients that spawn the server as a subprocess.
quake mcpThe MCP server exposes 20 tools organized into five categories:
- Observability (read-only):
testnet_status,list_nodes,get_block_heights,get_mempool,get_peers - Lifecycle:
start_nodes,stop_nodes,restart_testnet,clean_testnet - Perturbations:
perturb_disconnect,perturb_kill,perturb_pause,perturb_restart,perturb_upgrade - Testing:
run_tests,wait_height,valset_update - Remote (remote testnets only):
remote_ssh,remote_ssm,remote_provision
The server also exposes two MCP resources:
| URI | Description |
|---|---|
quake://manifest |
The current testnet manifest (TOML configuration file) |
quake://nodes |
All node metadata as JSON |
The generate command (alias gen) creates random manifest files for testing. It is useful for nightly or ad‑hoc runs that exercise many topologies and configurations without writing manifests by hand.
Behavior
- Writes one or more TOML manifests into the given output directory.
- For certain combinations of topology, height strategy, and region strategy, it generates
countmanifests (each with a different seed). - Seeding: If you pass
--seed S(before thegeneratecommand), that value is used as the base seed. The first manifest gets seedS, the next getsS+1, thenS+2, and so on. This makes runs reproducible: the same--seedand options produce the same manifests. Without--seed, the base seed is chosen at random (different each run), so output varies between invocations.- Example:
quake --seed 42 gen -o out -c 2produces2 x 9 combinations = 18 manifestswith seeds42,43,44, ... ,59; running again with the same arguments yields identical manifests.
- Example:
Randomization strategies
| Dimension | Options |
|---|---|
| Network topology | 1 node | 5 nodes | complex |
| Height start | All nodes at 0 | some nodes start at 100 |
| Region assignment | Single region | uniform random | clustered |
The complex topology creates a sentry architecture with the following structure:
- Sentry group 1: 1–3 validators (randomly chosen per manifest), fully meshed with each other and connected to
sentry-1 - Sentry group 2: 1–3 validators (randomly chosen per manifest), fully meshed with each other and connected to
sentry-2 - Sentries:
sentry-1andsentry-2are connected to each other, to their respective validator groups, and to therelayer - Relayer: Connected to both sentries and to 1 full node
- Full node:
full-1is connected to therelayer
All connections use persistent peers, creating a structured network topology that isolates validators behind sentry nodes.
For each combination, it generates count manifests (default 1). The combinations are:
- 1 combination with single node (no region or height strategy variation)
- 6 combinations with 5 nodes (2 height strategies × 3 region strategies)
- 2 combinations with complex topology (2 height strategies, all nodes within a single region)
Total manifests generated = 9 × count.
Randomized per manifest
- Consensus Layer: logging, p2p transport (tcp/quic), value_sync parameters, runtime flavor, pruning, and related options.
- Execution Layer: txpool, builder, and engine options (within safe ranges).
- Manifest-level: engine API connection (IPC vs RPC), initial hardfork (e.g. zero6/zero5).
Consensus, value_sync, and RPC are always enabled so that setup → start → wait height → test works on every generated manifest.
Options
| Option | Short | Default | Description |
|---|---|---|---|
--output-dir |
-o |
.quake/generated |
Directory to write manifest files into. |
--count |
-c |
1 |
Number of manifests to generate per combination. |
--seed |
— | (random) | Base seed for the RNG (use for reproducible runs). |
Examples
# Generate 1 manifest per combination (9 total) with a fixed seed
quake --seed 42 generate --output-dir target/manifests
# Generate 10 manifests per combination (90 total), reproducible
quake --seed 123 generate -o target/manifests -c 10
# Generate 1 per combination with a random seed (different each run)
quake generate -o target/manifestsNightly CI
A nightly workflow runs daily at 3 AM UTC and can be triggered:
- Scheduled/PR runs: Use seed
42for reproducibility. - Manual dispatch (
workflow_dispatch): Optionally specify a base seed (must be a non-negative integer; leave blank to use42) and a per-job count (positive integer ≤ 50; leave blank to use3).
The workflow runs 10 parallel jobs (matrix indices 0–9). Each job computes its effective seed as BASE_SEED + 10000 × index, then calls quake --seed <SEED> generate --count <COUNT> to produce COUNT × 9 manifests (default: 3 × 9 = 27 per job, 270 total across all jobs). Each job then runs the full test pipeline on every manifest: setup → start → wait height 140 → test → clean. Logs and reports are uploaded to the artifacts bucket with per-job artifact names (-idx<N>). The driving script is scripts/scenarios/nightly-random-manifests.sh.
PRs labeled test-random will also trigger this workflow.
By default, clean stops the testnet and removes all node data and configuration,
but leaves monitoring data alone. The following flags control what is removed:
| Flag | Short | Description |
|---|---|---|
--all |
-a |
Remove everything, including monitoring services and their data. Cannot be combined with data flags. |
--data |
-d |
Remove only execution and consensus layer data, preserving configuration. Cannot be combined with --execution-data or --consensus-data. |
--execution-data |
-x |
Remove only execution layer (Reth) data. Cannot be combined with --data or --consensus-data. |
--consensus-data |
-c |
Remove only consensus layer (Malachite) data. Cannot be combined with --data or --execution-data. |
# Remove node data only (keep config, monitoring intact)
./quake clean --data
# Remove only execution layer data
./quake clean --execution-data
# Remove only consensus layer data
./quake clean --consensus-data
# Remove everything including monitoring
./quake clean --allMonitoring services (Prometheus, Grafana, cAdvisor, Blockscout) can be controlled with the monitoring command:
# Start monitoring services
./quake monitoring start
# Stop monitoring services
./quake monitoring stop
# Stop monitoring services and remove monitoring data
./quake monitoring clean
# Download a Prometheus metrics snapshot + a single-node database snapshot
./quake download
# Or download just one:
./quake download metrics
./quake download dbThe download subcommand group queries the running Prometheus directly over
HTTP (local Docker port for local testnets, SSM-tunnelled port for remote) and
bundles each metric's query_range response into one archive. The db
subcommand archives node database files (remote) or logs their on-disk paths
(local). Common options:
# Limit to a time range
./quake download metrics --from 2024-01-15T10:30:00Z --to 2024-01-15T12:00:00Z
# Download only specific metrics (names go after `--`)
./quake download metrics -- reth_db_size_bytes go_goroutines
# Save to a custom path
./quake download metrics -o /tmp/my-metrics.tar.gz
# Download db from a specific node (default: first node in manifest)
./quake download db -- validator1Without --from, the start defaults to Prometheus' headStats.minTime (the
current head block start, typically the last ~2 h). Without --to, defaults
to now. Without --step, the step is auto-sized to keep the response below
Prometheus' 11 000-point limit. Archives land in
.quake/metrics/<testnet>/ and .quake/db/<testnet>/ by default; pass -o
to override. quake remote download {metrics,db} is deprecated but kept as a
backward-compatible alias.
The manifest is a TOML file.
Before parsing, all ${VAR_NAME} patterns are
replaced with values from the process environment and .env files. This allows
any field to reference environment variables. For example:
image_cl="${IMAGE_REGISTRY_URL}/arc-consensus:abc123"Optional top-level settings:
- name: Name of the test scenario
- description: Description of the test scenario
- engine_api_connection: Connection method between Consensus Layer (CL) and Execution Layer (EL). Valid values: "ipc" (default), "rpc".
- image_cl and image_el: Docker images for CL and EL containers. If omitted, defaults to
- for local mode:
arc_consensus:latestandarc_execution:latest, or - for remote mode:
${IMAGE_REGISTRY_URL}/arc-consensus:<version>and${IMAGE_REGISTRY_URL}/arc-execution:<version>, whereIMAGE_REGISTRY_URLis taken from the.envfile (see Custom Docker images). These are the network-wide defaults; individual nodes or node groups can override them (see Per-node and per-group images).
- for local mode:
- image_cl_upgrade, image_el_upgrade: Docker images to use when upgrading containers with
quake perturb upgrade. Required for upgrade scenarios; not supported in remote mode. - group_images: Per-node-group image overrides, declared as
[group_images.<group>]withimage_cl/image_elkeys. See Per-node and per-group images. - node_size: EC2 instance type for validator/full nodes (e.g.
"m6a.4xlarge"). Equivalent to the--node-sizeCLI flag. See Instance sizing for available options. Remote mode only — ignored in local mode (a warning is printed). - cc_size: EC2 instance type for the Control Center. Equivalent to
--cc-size. Remote mode only — ignored in local mode. - node_disk_gb: Root EBS volume size in GiB for each node. Must be ≥ 8. Equivalent to
--node-disk-gb. Omit to keep the AMI default. Remote mode only — ignored in local mode. - cc_disk_gb: Root EBS volume size in GiB for the Control Center. Must be ≥ 8. Equivalent to
--cc-disk-gb. Remote mode only — ignored in local mode. - node_volume_type: AWS EBS volume type for each node's root disk; tunes disk
cost/performance (e.g. match a production disk profile). General Purpose SSD
(
gp2,gp3), Provisioned IOPS SSD (io1,io2), Throughput Optimized HDD (st1), Cold HDD (sc1). Default:gp3. See AWS EBS volume types. Equivalent to--node-volume-type. Remote mode only. - node_volume_iops: Provisioned IOPS for the node root EBS volume; raises the I/O
ceiling above the volume type's baseline. Only valid with
gp3,io1,io2; range 100–256000. Default: AMI's baseline IOPS for the chosen type. Equivalent to--node-volume-iops. Remote mode only. - node_data_on_instance_store: When
true, mounts the local instance-store NVMe at the node data directory so the EL/CL databases live on local disk instead of the root EBS volume. Requires an instance type with local NVMe (e.g.i4i.*,i3.*,m6id.*); multiple instance-store volumes are striped RAID0 into one device. A no-op on instance types without instance store, leaving the data directory on EBS. Independent ofnode_volume_type/node_volume_iops, which keep configuring the root EBS volume. Equivalent to--node-data-on-instance-store. Remote mode only. - el_cpu_limit: Hard CPU cap for each EL container; reproduces production CPU quotas
on the testnet. Whole or fractional CPUs (e.g.
0.5). Maps to Docker Composecpus. Default: no limit (container uses all host CPUs). - el_memory_limit_gb: Hard memory cap for each EL container in GiB; fractional values
(e.g.
2.5) are allowed. Maps to Docker Composemem_limit. Default: no limit locally; 2.5 GiB on remote. - cl_cpu_limit: Hard CPU cap for each CL container; same semantics as
el_cpu_limit. - cl_memory_limit_gb: Hard memory cap for each CL container in GiB; fractional values allowed. Default: no limit locally; 1 GiB on remote.
Nodes are defined as individual TOML sections with names starting with validator or node.
[[nodes]]
[validator1]
[validator2]
[node1]
[node2]Consensus Layer (CL) configuration is set under cl.config.* keys. The
schema matches the StartCmd struct in
crates/malachite-cli/src/cmd/start.rs.
Keys are flat and map 1:1 to the arc-node-consensus start CLI flags
(e.g. cl.config.log_level = "debug" → --log-level=debug). Quake
translates the merged config into CLI flags at setup time; the CL does not
read a config.toml.
Upgrade scenarios can pin an older arc_consensus image tag (e.g.
v0.6.0). Quake derives CLI flags from the StartCmd definition
compiled into its own binary, which may have gained, renamed, or removed
flags since that image shipped. Before handing the flags to the
container, Quake rewrites them to match the target version, i.e., older
images receive a compatible subset, and "latest", missing, or
unparsable tags pass through unchanged.
If pinning an older image fails with unexpected argument, that version
likely needs a new compatible entry. See apply_version_compat in
src/cli_version.rs for the rustdoc describing
how to add one.
The default configuration of Reth (Execution Layer) is defined in
crates/quake/src/manifest.rs. It can be set globally
or for each node by prefixing the config field with el.config..
For example:
# Global settings that apply to all nodes
engine_api_connection = "rpc" # or "ipc" (default)
cl.config.log_level = "debug"
el.config.disable-discovery = true
[[nodes]]
[validator1]
# Node-specific settings
cl.config.discovery_num_outbound_peers = 30
[validator2]
# Node-specific settings
el.config.builder.deadline = 5
[node1]
[node2]In general, node configuration options are applied with the following precedence, from lowest to highest priority:
- Global manifest configs: defined at the top level of your manifest file
- Per-node manifest configs: defined within each node's section
Higher-priority configs override lower-priority ones when their keys match.
In addition to the general node configuration described above, the EL configuration
adds another layer of defaults defined in crates/quake/src/manifest.rs.
This means that Quake manages Reth (Execution Layer) flags through a three-tier
configuration system. Configs are applied with the following precedence, from
lowest to highest priority:
- Default configs: defined in
crates/quake/src/manifest.rs - Global manifest configs: defined at the top level of your manifest file
- Per-node manifest configs: defined within each node's section
As before, higher-priority configs override lower-priority ones when their keys match. For a full list of reth CLI flags that you can set in your manifest, see the Reth documentation.
EL configs use TOML table syntax under the el.config key.
Boolean flags (like --http or --disable-discovery) use true/false,
while flags with values use their appropriate types (strings, integers, arrays).
# Global EL config that applies to all nodes
[el.config]
http.enable = true
http.api = ["admin", "net", "eth"]
engine.persistence-threshold = 5
disable-discovery = true
[nodes.validator1]
# Per-node config that overrides global and defaults for this node only
el.config.engine.persistence-threshold = 10
el.config.builder.deadline = 5
[nodes.validator2]
# No per-node config, so it inherits global + defaults
[nodes.full1]
# Per-node config can also use table syntax
[nodes.full1.el.config]
txpool.nolocals = falseSpecial syntax notes:
http.enable = trueproduces--http(the.enablesuffix is stripped)ws.enable = trueproduces--wsdisable-discovery = trueproduces--disable-discoverydisable-discovery = falseomits the flag entirely- Array values like
http.api = ["admin", "net"]produce--http.api=admin,net - Array values are replaced, not merged. If a node defines
http.api = ["admin"], it completely overrides the default array, not appends to it. - Quake does not validate flag names. Misspelled flags (e.g.,
htpp.port = 8545) will be passed to Reth, which will fail at startup with an unrecognized flag error.
The following flags are applied to all nodes by default. They are defined in
crates/quake/src/manifest.rs and can be overridden in your manifest:
| Flag | Default Value | Description |
|---|---|---|
http.enable |
true |
Enable the HTTP-RPC server |
http.api |
["admin", "net", "eth", "web3", "debug", "txpool", "trace", "reth"] |
APIs exposed over HTTP |
ws.enable |
true |
Enable the WebSocket-RPC server |
ws.api |
["admin", "net", "eth", "web3", "debug", "txpool", "trace", "reth"] |
APIs exposed over WebSocket |
engine.persistence-threshold |
0 |
Persistence threshold for engine payloads |
engine.memory-block-buffer-target |
0 |
Memory block buffer target |
enable-arc-rpc |
true |
Enable Arc-specific RPC methods |
rpc.txfeecap |
1000 |
Maximum transaction fee cap |
txpool.nolocals |
true |
Treat all transactions equally (no local priority) |
To override a default, simply define the flag in your manifest's global or
per-node el.config section.
The following flags are managed by Docker Compose templates and must not be set in manifests. If present, they are silently ignored:
| Flag | Value (Local) | Value (Remote) | Notes |
|---|---|---|---|
datadir |
/data/reth/execution-data |
/data/reth/execution-data |
Data directory path |
chain |
/app/assets/genesis.json |
/app/assets/genesis.json |
Genesis file path |
http.port |
8545 |
8545 |
HTTP-RPC port |
http.addr |
0.0.0.0 |
0.0.0.0 |
HTTP-RPC bind address |
http.corsdomain |
* |
* |
CORS allowed origins |
ws.port |
8546 |
8546 |
WebSocket-RPC port |
ws.addr |
0.0.0.0 |
0.0.0.0 |
WebSocket-RPC bind address |
ws.origins |
* |
* |
WebSocket allowed origins |
metrics |
0.0.0.0:9001 |
0.0.0.0:9001 |
Metrics endpoint |
authrpc.addr |
0.0.0.0 |
0.0.0.0 |
Auth server address to listen on (RPC mode) |
authrpc.port |
8551 |
8551 |
Auth server port to listen on (RPC mode) |
authrpc.jwtsecret |
/app/assets/jwtsecret |
/assets/jwtsecret |
JWT secret path (RPC mode) |
ipcdisable |
(set) | (set) | Disable IPC (RPC mode only) |
ipcpath |
/sockets/reth.ipc |
/sockets/reth.ipc |
IPC socket path (IPC mode) |
auth-ipc |
(set) | (set) | Enable authenticated IPC (IPC mode) |
auth-ipc.path |
/sockets/auth.ipc |
/sockets/auth.ipc |
Auth IPC socket path (IPC mode) |
p2p-secret-key |
/data/reth/execution-data/nodekey |
/data/reth/execution-data/nodekey |
Pre-generated secp256k1 key for P2P identity |
trusted-peers |
(auto-generated) | (auto-generated) | Comma-separated enode URLs of all other nodes |
The IPC vs RPC flags are automatically selected based on the connection mode.
Use quake setup --rpc to switch from IPC (default) to RPC connections between
the Consensus Layer and Execution Layer. You can also set this in your manifest using
the engine_api_connection top-level key.
Example 1: Override a default flag globally
# Disable the txpool.nolocals default for all nodes
[el.config]
txpool.nolocals = false
[nodes.validator1]
[nodes.validator2]Example 2: Per-node override
[nodes.validator1]
# This node uses a custom persistence threshold
el.config.engine.persistence-threshold = 10
el.config.builder.deadline = 5
[nodes.validator2]
# Uses defaults onlyExample 3: Mixed global and per-node configuration
# Global: disable discovery for all nodes
[el.config]
disable-discovery = true
engine.persistence-threshold = 5
[nodes.validator1]
# Override: re-enable discovery for this node
el.config.disable-discovery = false
[nodes.validator2]
# Uses global config (discovery disabled, threshold=5)
[nodes.full1]
# Override: different persistence threshold
el.config.engine.persistence-threshold = 20Note
When you set enable-arc-rpc = true (the default), --arc-rpc-upstream-url=<URL>
is automatically added to Reth's configuration. You don't need to include it
manually.
In addition to CLI flags (el.config/cl.config), you can set environment
variables on a node's containers via the el.env (Execution Layer) and
cl.env (Consensus Layer) tables. They follow the same precedence as config:
global values are inherited by every node and per-node values override matching
keys.
# Global: applies to every node's containers
[el.env]
RUST_LOG = "info"
[cl.env]
RUST_LOG = "info"
[nodes.validator1.el.env]
# Override the global value for this node's EL container only
RUST_LOG = "debug,net::discovery=trace"
[nodes.validator2.cl.env]
# Halt this node's CL at a given height (testing graceful shutdown)
ARC_HALT_AT_BLOCK_HEIGHT = 100el.envis applied to the EL (Reth) container;cl.envto the CL (Malachite) container. Use the right table for the layer you want to affect.- Keys must be valid environment variable names (
^[A-Za-z_][A-Za-z0-9_]*$). - Values may be strings, integers, floats, or booleans; non-scalars (arrays/tables) are rejected. All values are emitted as strings.
- Quake sets some environment variables by default (e.g.
RUST_LOGandARC_LOG_FILEon the EL container,ARC_HALT_AT_BLOCK_HEIGHTon the CL container). Setting the same key inel.env/cl.envreplaces the default rather than duplicating it. - Works in both local and remote deployments.
By default every validator in genesis receives a voting power of 20. To override this, set cl_voting_power on each validator node:
[nodes.validator-1]
cl_voting_power = 2000
[nodes.validator-2]
cl_voting_power = 2000
[nodes.validator-3]
cl_voting_power = 1000
[nodes.full1]If cl_voting_power is specified for any validator, it must be specified for all validators (all-or-nothing). This prevents accidental power imbalances where one validator silently defaults to 20 while others are set to much higher values. Non-validator nodes ignore this field.
Note: The
cl_persistent_peerssetting described below applies to the Consensus Layer (Malachite) P2P connections. Execution Layer (Reth) P2P: during setup, a secp256k1 nodekey is pre-generated for each node. Reth's--trusted-peersis built from each node'sel.config.trusted_peerswhen set (same format ascl_persistent_peers: node names or group names, resolved to enodes); whenel.config.trusted_peersis not set for a node, that node gets a full mesh of all other nodes.
You can define custom groups of nodes and use them to configure peer connections. This is useful for setting up network topologies where certain nodes should only connect to specific subsets of other nodes.
Pre-defined node groups:
ALL_NODES- All nodes in the manifestALL_VALIDATORS- All validator nodes, that is, nodes with names starting withval(e.g.,validator1,val2)ALL_NON_VALIDATORS- All nodes that are not validators
These names are reserved built-ins and cannot be redefined under
[node_groups].
Custom node groups are defined in the [node_groups] section. Groups can reference individual node names, pre-defined groups, or other groups previously declared:
[node_groups]
FULL_NODES = ["full1", "full2"]
TRUSTED = ["ALL_VALIDATORS", "FULL_NODES", "other_node"]
[nodes.validator1]
cl_persistent_peers = ["TRUSTED"]
[nodes.validator2]
[nodes.validator3]
[nodes.validator4]
[nodes.full1]
cl_persistent_peers = ["ALL_NON_VALIDATORS"]
[nodes.full2]
cl_persistent_peers = ["ALL_VALIDATORS"]
[nodes.sentry]
cl_persistent_peers = ["ALL_NODES"]
[nodes.other_node]In this example:
FULL_NODESis a custom group containingfull1andfull2TRUSTEDcombines theALL_VALIDATORSgroup, theFULL_NODESgroup, and the individual nodeother_nodevalidator1will have persistent peers:validator2,validator3,validator4,full1,full2,other_node(theTRUSTEDgroup, excluding itself)full1will connect to all non-validators:full2,sentry,other_nodesentrywill connect to all nodes except itself
The same group names can also be used as quake load and quake spam
targets. For example:
./quake load -t 60 -r 500 --targets ALL_VALIDATORS
./quake spam -t 30 -r 1000 --targets TRUSTEDTo distinguish group references from individual nodes in peer lists, by convention we use lowercase for node names and uppercase for node group names.
Note: A node is automatically excluded from its own persistent peers list.
Default behavior for cl_persistent_peers:
- If
cl_persistent_peersis not specified for a node, it will connect to all other nodes in the network (default behavior for simple testnets). - If
cl_persistent_peersis specified as an empty array (cl_persistent_peers = []), the node will have no persistent peers. - If
cl_persistent_peersis specified with values, the node will connect only to those specific peers.
el.config.trusted_peers (Execution Layer): identical behavior to cl_persistent_peers.
By default every node runs the global image_cl/image_el (see
Basic Structure). You can override the base image for
individual nodes or whole node groups, which boots a mixed-version network
from genesis without a rolling perturb upgrade. Useful for cross-version
consensus testing, pre-rollout validation, and reproducing version skew.
- Per node: set
image_cl/image_elunder a[nodes.<name>]section. - Per group: set them under
[group_images.<group>], keyed by any node group (custom or a built-in such asALL_VALIDATORS); applied to every member.
Precedence, lowest to highest: global image < node-group override < per-node override. A node covered by two image-declaring groups for the same layer is rejected as ambiguous.
image_cl = "arc_consensus:latest" # global base images
image_el = "arc_execution:latest"
[node_groups]
OLDIES = ["validator4", "validator5"]
# validator4 and validator5 run the pinned release instead of the global image
[group_images.OLDIES]
image_cl = "${IMAGE_REGISTRY_URL}/arc-consensus:0.6.0"
image_el = "${IMAGE_REGISTRY_URL}/arc-execution:0.6.0"
[nodes.validator1]
[nodes.validator4]
[nodes.full1]
image_el = "${IMAGE_REGISTRY_URL}/arc-execution:0.6.0" # inline override winsCompatibility (operator's responsibility). Every image in a mixed network
must agree on genesis state, the hardfork schedule, and on-disk db format. Quake
enforces only arc_consensus >= v0.5.0 and, in remote mode, a ghcr.io/
registry for each image; genesis or db mismatches are not caught statically and
fail loudly at startup. The base image is per-node, while the
image_cl_upgrade/image_el_upgrade used by perturb upgrade stay global.
A worked example ships at
scenarios/mixed-version.toml.
By default, all nodes (Consensus Layer and Execution Layer containers) will
start when the start CLI command is invoked, unless a node has a
start_at height set in the manifest.
[nodes.validator1]
[nodes.validator2]
[nodes.full1]
start_at = 30By default, all nodes are connected to a single Docker network named default.
You can isolate nodes into separate sub-networks and create bridge nodes that
connect multiple sub-networks by using the subnets field.
Each subnet is assigned a dedicated private IP address range. In local mode,
subnets use 172.<N>.0.0/16 CIDR blocks where N starts at 21 and increments
for each subnet.
Subnets work in both local and remote deployments:
- Local: Isolation is enforced via separate Docker networks. Each subnet is
assigned a dedicated private IP address range using
172.<N>.0.0/16CIDR blocks whereNstarts at 21. Containers in different networks cannot communicate directly. Network perturbations (disconnect/connect) usedocker network disconnectanddocker network connectto detach and reattach containers from their subnet networks. - Remote: Isolation is enforced at the AWS infrastructure level. Each
logical network maps to a separate VPC subnet with its own security group.
Nodes belonging to multiple networks (bridge nodes) have multiple network
interfaces (ENIs) attached, one per network. Network perturbations
(disconnect/connect) use host-level
iptablesrules to block/unblock traffic between nodes' VPC IPs.
The following example has 5 nodes across multiple isolated networks: trusted, untrusted, and default:
[nodes.validator1]
subnets = ["trusted"]
[nodes.validator2]
subnets = ["trusted"]
[nodes.validator3]
subnets = ["trusted", "untrusted"]
[nodes.validator4]
subnets = ["untrusted", "default"]
[nodes.full1]
# No subnets specified, defaults to ["default"]In this example:
validator1andvalidator2are isolated in thetrustedsubnetvalidator3bridges thetrustedanduntrustedsubnetsvalidator4bridges theuntrustedanddefaultsubnetsfull1is in thedefaultsubnet
Quake validates that the network topology forms a connected graph. If networks are completely isolated from each other (no bridge nodes), the manifest validation will fail.
Subnet isolation is enforced at two levels:
-
Infrastructure level:
- Local: Each subnet maps to a separate Docker network configured as
internal: true, which prevents routing through Docker's gateway. This enforces isolation even on Docker Desktop where bridge networks would otherwise be able to communicate. Containers are also connected to a sharedhost-accessnetwork (non-internal) for port publishing. Network perturbations (disconnect/connect) usedocker network disconnectanddocker network connectto detach and reattach containers from their subnet networks. - Remote: Each subnet maps to a separate VPC subnet with its own security
group that only allows traffic within that subnet. Bridge nodes get multiple
ENIs (one per subnet). Network perturbations use
iptablesDROP rules installed only on the target node's EC2 host, blocking peer IPs in the INPUT, OUTPUT, and FORWARD chains. This unidirectional approach avoids altering peer hosts. On reconnect, the rules are removed and Malachite's persistent peer reconnection handles re-establishing connections on both sides.
- Local: Each subnet maps to a separate Docker network configured as
-
Application level: Consensus layer nodes are configured with
cl_persistent_peersthat only include nodes sharing at least one subnet. This ensures nodes only attempt to connect to peers they can actually reach.
Both layers are necessary for robust isolation. Infrastructure-level isolation works reliably on native Linux Docker and AWS, but Docker Desktop (Mac/Windows) does not fully isolate bridge networks. The application-level peer configuration provides defense-in-depth and ensures correct behavior across all platforms.
All nodes in the testnet are typically deployed to the same private network configuration either in a local machine or remotely in one cloud region.
We can emulate latency between nodes by artificially increasing the latency of outbound traffic with the Linux tc (traffic control) command.
Unless we set latency_emulation = false, latency emulation will be enabled by default.
We can assign in the manifest an AWS-region to each node.
Nodes that don't have an explicit region in the manifest will be assigned a random one.
Then, Quake will simulate network latency between containers in different regions, using real-world average latency values between each region.
latency_emulation = true
[[nodes]]
[validator1]
region = "eu-central-1"
[validator2]
[validator3]
This latency emulation mechanism is a re-implementation of the one in CometBFT's e2e framework. In turn, the latter was adapted from https://github.com/paulo-coelho/latency-setter.
Quake can deploy a testnet to remote infrastructure (AWS EC2 instances).
The remote setup consists of:
- one EC2 instance per node, where
- each node consists of a Consensus Layer and a Execution Layer, each in its own Docker container
- one extra instance for a Control Center (CC) server
- for monitoring services, and
- for generating and sending transaction load to the nodes.
Currently all instances are deployed in one AWS region (us-east-1 by default)
and we rely on latency emulation to make the node
communication behavior more realistic.
graph TB
subgraph nodes[" "]
direction LR
subgraph node_3["Node 3 instance"]
direction TB
CL3["CL"]
EL3["EL"]
end
subgraph node_2["Node 2 instance"]
direction TB
CL2["CL"]
EL2["EL"]
end
subgraph node_1["Node 1 instance"]
direction TB
CL1["CL"]
EL1["EL"]
end
end
CC
CC --> node_1
CC --> node_2
CC --> node_3
CL1 --- EL1
CL2 --- EL2
CL3 --- EL3
All communication with remote nodes is routed through the Control Center (CC):
- RPC requests are routed via an nginx reverse proxy on CC
- Pprof requests are routed via a separate nginx reverse proxy on CC
- SSH/SCP commands are routed via CC using nodes' private IPs
This architecture requires only a single SSM session to CC, avoiding AWS API throttling that would occur with per-node SSM sessions.
flowchart LR
subgraph local [Local Machine]
Quake
end
subgraph aws [AWS VPC]
CC[Control Center]
RPC[RPC proxy :8080]
PProf[pprof proxy :6060]
N0[Node 0]
N1[Node 1]
NN[Node N]
end
Quake -->|SSM Session| CC
CC --- RPC
CC --- PProf
RPC -->|Private IP :8545| N0
RPC -->|Private IP :8545| N1
RPC -->|Private IP :8545| NN
PProf -->|Private IP :6060/:6161| N0
PProf -->|Private IP :6060/:6161| N1
PProf -->|Private IP :6060/:6161| NN
CC -->|SSH via Private IP| N0
CC -->|SSH via Private IP| N1
CC -->|SSH via Private IP| NN
RPC Proxy:
- A single SSM tunnel forwards
localhost:18080to CC's port8080 - The nginx proxy routes requests by URL path:
/<node-name>/elforwards to the node's Execution Layer (Reth JSON-RPC) on port 8545/<node-name>/el/wsforwards to the node's Execution Layer (Reth WebSocket) on port 8546/<node-name>/clforwards to the node's Consensus Layer (Malachite RPC) on port 31000/<node-name>/cl/metricsforwards to the node's Consensus Layer metrics on port 29000
- Additional endpoints:
/health(health check),/nodes(list available nodes)
Pprof Proxy:
- A single SSM tunnel forwards
localhost:16060to CC's port6060 - The nginx proxy routes by URL path:
/pprof/cl/<node-name>/...forwards to the node's CL pprof (port 6060),/pprof/el/<node-name>/...to the EL pprof (port 6161) - The proxy has a 600-second read timeout because CPU profiling requests block for the sampling duration (
?seconds=N, default 30s). For CPU profiles longer than 10 minutes, SSH into the node and query pprof locally. - Additional endpoints:
/health(health check),/nodes(list available nodes)
SSH/SCP Routing:
- SSH to CC uses a direct SSM session
- SSH to nodes is routed through CC: Quake SSHs to CC, then CC SSHs to the node using its private IP
- Parallel commands to multiple nodes run from a single SSH session to CC
Benefits:
- Avoids AWS API throttling: only one SSM session needed instead of N+1
- Simpler session management with only one tunnel to maintain
- Low latency between CC and nodes (same VPC/region)
All basic commands for local deployment work in remote mode. It just requires a few extra steps.
To quickly deploy and start a remote testnet with one command, run:
quake -f crates/quake/scenarios/examples/5nodes.toml start --remotewhich is equivalent to
quake -f crates/quake/scenarios/examples/5nodes.toml remote create --yes
quake startCheck out quake remote --help for details on every sub-command.
For testnets running longer than ~20 hours, add --node-size t3.large (see Instance sizing).
Setting up the required tools to make this work requires a few extra steps, described below.
The --node-size and --cc-size flags let you override the default EC2 instance
types when creating remote infrastructure. The --node-disk-gb and --cc-disk-gb
flags set the root EBS volume size in GiB for nodes and the Control Center;
omit them to keep the AMI default volume size.
# Use larger nodes for a multi-day testnet
quake remote create --node-size t3.large --cc-size t3.2xlarge
# Larger root volume for long runs (disk fills before RAM on default volume)
quake remote create --node-size t3.large --node-disk-gb 100 --cc-disk-gb 100
# Or with the shorthand
quake start --remote --node-size t3.large --node-disk-gb 100By default the node data directory lives on the root EBS volume. The
--node-data-on-instance-store flag instead mounts the instance's local NVMe
instance store at the data directory, so the EL/CL databases run on local disk.
It requires an instance type that ships local NVMe (i4i.*, i3.*, m6id.*,
c6id.*); on any other type it is a no-op and the data directory stays on EBS.
Instance types with multiple instance-store volumes are striped RAID0 into a single
device. The flag is independent of --node-volume-type/--node-volume-iops, which
keep tuning the root EBS volume, so the two storage backends can be compared directly.
# Datadir on io2 EBS
quake remote create --node-size i4i.xlarge \
--node-volume-type io2 --node-volume-iops 64000 --node-disk-gb 1000
# Datadir on local NVMe (same instance type, only the storage backing differs)
quake remote create --node-size i4i.xlarge --node-data-on-instance-storeThe instance store is ephemeral: its contents are lost when the instance stops or terminates. That is fine for benchmarking and short-lived testnets, but never use it for state you need to keep.
Each node runs an Execution Layer (EL) and a Consensus Layer (CL) container, plus the OS, Docker daemon, NFS client, and SSM agent. The EL is the dominant consumer of both memory (~2.5 GiB) and disk (debug logs grow at ~200 MiB/hr).
| Instance | vCPU | RAM | Max testnet duration | Best for |
|---|---|---|---|---|
t3.medium (default) |
2 | 4 GiB | ~20 hours | Short tests, CI smoke runs |
t3.large |
2 | 8 GiB | 1–3 days | Day-long testnets, moderate load |
t3.xlarge |
4 | 16 GiB | Multi-day | Heavy load, large state, long-running |
The duration estimates assume debug-level logging with no log rotation on a 20
GiB root volume. The primary constraint is disk space: the 4 GiB swap file,
~9 GiB of Docker images, and growing log files fill the default 20 GiB volume in
roughly 20 hours. Larger instances do not increase disk size; use --node-disk-gb
for that. Larger instances do provide more RAM headroom, reducing swap pressure
and making the node more resilient to memory spikes.
Tip
For testnets that need to run longer than 20 hours, consider both upgrading
the instance size (for RAM) and passing --node-disk-gb (and --cc-disk-gb
if the CC needs more space) for disk.
The CC runs Prometheus, Grafana, Blockscout (backend + frontend + DB), an RPC reverse proxy, a pprof reverse proxy, node-exporter, and optionally spammer containers.
| Instance | vCPU | RAM | Notes |
|---|---|---|---|
t3.large |
2 | 8 GiB | Insufficient — Blockscout + Prometheus exceed 8 GiB |
t3.xlarge (default) |
4 | 16 GiB | Standard monitoring stack |
t3.2xlarge |
8 | 32 GiB | Many nodes (>15) or heavy Blockscout indexing |
Install:
awsCLI toolaws's Session Manager plugin- Terraform:
brew tap hashicorp/tap brew install hashicorp/tap/terraform
By default, remote nodes pull the latest arc-consensus and arc-execution
Docker images from Circle's private container registry. To access them, you must
provide a GitHub personal access token (PAT) that has permission to read
packages.
You can also override the defaults to use images from your own registry, to test
changes from a branch that hasn't been merged or published yet. For example, if
stored in GHCR, set image_cl and image_el in your manifest to:
image_cl = "ghcr.io/<org|user>/<repo>/arc-consensus:<tag>"
image_el = "ghcr.io/<org|user>/<repo>/arc-execution:<tag>"Both Terraform (for pre-pulling images during provisioning) and the per-node Docker Compose files will use these values.
The subsections below explain how to set up credentials to access the private registry and, optionally, build and push your own images.
Example setup for GitHub Container Registry (ghcr.io):
- Visit https://github.com/settings/tokens/new
- Create a Classic personal access token (PAT) with the
read:packagesscope.- Make sure to select Classic PAT, not a Fine-Grained token.
- Add a descriptive note and set an expiration date (recommended).
- If the image is hosted under an organization, you might need to authorize the PAT for that organization (e.g. Under your tokens list in GitHub, Configure SSO on the token, and select the org).
- Store your GitHub username and the generated PAT (it starts with
ghp_) in a file named.envin the repository root directory.GITHUB_USER=<github-username> GITHUB_TOKEN=<github-classic-token>
-
Prerequisites
- Docker installed and running.
- A
.envfile withGITHUB_USERandGITHUB_TOKENas described above. Your PAT needs the additionalwrite:packagesscope (to push images) on top of theread:packagesscope required for pulling.
-
Log in to ghcr.io
source .env echo $GITHUB_TOKEN | docker login ghcr.io -u $GITHUB_USER --password-stdin
-
Build images for
linux/amd64EC2 instances run on
amd64. Do not use./quake buildfor this — on Apple Silicon Macs it producesarm64images that will fail on EC2 with a misleadingunauthorizederror (see troubleshooting below).Build each image explicitly with
--platform linux/amd64. Replace<org|user>with your GitHub username (e.g.your-github-username) and<tag>with your chosen tag (e.g.dev).From the repository root:
# Execution layer docker build --platform linux/amd64 \ --build-context certs=deployments/certs \ --build-arg GIT_COMMIT_HASH=$(git rev-parse HEAD) \ --build-arg GIT_VERSION=$(git describe --tags --always) \ --build-arg GIT_SHORT_HASH=$(git rev-parse --short HEAD) \ --target dev-runtime \ -t ghcr.io/<org|user>/<repo>/arc-execution:<tag> \ -f deployments/Dockerfile.execution . # Consensus layer docker build --platform linux/amd64 \ --build-context certs=deployments/certs \ --build-arg GIT_COMMIT_HASH=$(git rev-parse HEAD) \ --build-arg GIT_VERSION=$(git describe --tags --always) \ --build-arg GIT_SHORT_HASH=$(git rev-parse --short HEAD) \ --target dev-runtime \ -t ghcr.io/<org|user>/<repo>/arc-consensus:<tag> \ -f deployments/Dockerfile.consensus .
-
Push images to your registry
docker push ghcr.io/<org|user>/<repo>/arc-execution:<tag> docker push ghcr.io/<org|user>/<repo>/arc-consensus:<tag>
-
Set manifest fields
Make sure
image_clandimage_elin your manifest match what you pushed:image_cl = "ghcr.io/<org|user>/<repo>/arc-consensus:<tag>" image_el = "ghcr.io/<org|user>/<repo>/arc-execution:<tag>"
-
Create and start the remote testnet
./quake clean # remove any previous local state ./quake -f <path_to_the_manifest_file> start --remote # start the network
Important
Your ghcr.io packages must remain private. EC2 nodes authenticate
using the GITHUB_USER and GITHUB_TOKEN from your .env file,
which are passed to the instances during provisioning.
When pulling images from ghcr.io, Docker returns an unauthorized error for
several different problems, not just missing credentials. The error always
looks like this:
Error response from daemon: Head "https://ghcr.io/v2/<owner>/<repo>/<image>/manifests/<tag>": unauthorized
This same error appears when:
- The image tag does not exist — e.g. you pushed an image with tag
devbut your manifest specifies a different tag inimage_clorimage_el. Docker reportsunauthorizedinstead of "not found". - The image was built for the wrong architecture — e.g. you built on an
Apple Silicon Mac (arm64) and pushed the image, but EC2 nodes are amd64.
The manifest exists but has no matching platform, and Docker reports
unauthorizedinstead of a platform mismatch. - Actual authorization failure — e.g. the
.envtoken is expired, missingread:packagesscope, or not authorized for SSO.
Initialize Terraform plugins and state. This step is required only once.
./quake remote preinitCreate EC2 instances for each node in the testnet, plus one extra for the Control Center (CC) server.
./quake [-f <manifest>] remote create [--dry-run] [--yes] [--node-size <type>] [--cc-size <type>] [--node-disk-gb <GIB>] [--cc-disk-gb <GIB>] [--node-data-on-instance-store]See Instance sizing for recommended instance types.
This will create a <testnet-dir>/infra.json file with the names, IP addresses,
and instance IDs of the created nodes.
For accessing the instances via SSH, we need to start SSM sessions that create tunnels from local ports to remote ports.
./quake remote ssm startNote that tunnels are closed automatically after 20 minutes of inactivity. For long-running experiments, keep them alive in a separate terminal:
./quake remote ssm keep-alive 2hOnce the SSM session are established, we can log in via SSH to a node or CC, or run commands in the instances directly from the terminal:
./quake remote ssh validator1
./quake remote ssh cc docker psShow information on the EC2 instances just created.
./quake infoCreate configuration files locally and upload them to the remote nodes.
./quake setupAfter creating the configuration files, this command implicitly will run:
remote provision, to upload the generated files to the remote nodes, andremote ssm start, to create SSM tunnels from local ports to ports in remote nodes.
Start Arc on the testnet nodes
./quake startCheck that the validators are creating blocks.
./quake info heightsMonitor health of all nodes and the Control Center:
./quake remote monitor # all hosts, one-shot
./quake remote monitor --follow # all hosts, 30s continuous refresh
./quake remote monitor validator1 # single node, one-shot
./quake remote monitor validator1 -f # single node, 5s time-series
./quake remote monitor cc -f # Control Center only, 5s time-series
./quake remote monitor -f -i 10 # all hosts, custom 10s refreshThe dashboard combines block heights, peer counts (via RPC), and memory/CPU/disk
usage with per-container memory in MiB (via SSH). Without --follow (-f), data
is collected and printed once. With --follow, the multi-host dashboard refreshes
continuously and the single-host view appends a new row each tick. Press Ctrl+C to
stop.
Send transaction load:
./quake load --targets validator1,validator2 -r 1000 -t 60quake load and quake spam auto-dispatch based on testnet type: local testnets
run the spammer directly, remote testnets forward to the Control Center via SSH.
All Spammer options are supported. --targets accepts comma-separated selectors
including manifest node groups such as ALL_VALIDATORS or custom [node_groups].
Download diagnostic artifacts from the testnet via the mode-agnostic
quake download command group (quake remote download {metrics,db} is deprecated but kept as a backward-compatible alias):
# Download both metrics and a single-node db snapshot
./quake download
# Metrics only — covers the current head block (~2 h) by default
./quake download metrics
# Metrics for a specific time range
./quake download metrics --from 2024-01-15T10:30:00Z --to 2024-01-15T12:00:00Z
# Specific metrics only (metric names go after --)
./quake download metrics -- reth_db_size_bytes go_goroutines
# Save to a custom output path
./quake download metrics -o /tmp/my-metrics.tar.gz
# Download node database (defaults to the first node in the manifest)
./quake download db
# Execution layer only, from specific nodes
./quake download db --execution-only -- validator1 validator2
# Save to a custom output path
./quake download db -o /tmp/my-db.tar.gzArchives land in .quake/metrics/<testnet>/quake-metrics-<timestamp>.tar.gz
and .quake/db/<testnet>/quake-db-<timestamp>.tar.gz by default; pass -o
to override.
Once finished with your tests, remember to destroy the remote infrastructure!
./quake remote destroyor
./quake cleanYou can give a colleague or another user access to a running remote testnet
without them having to create the infrastructure. The export / import
commands bundle and restore all the files Quake needs into a single JSON file.
The export bundle contains:
- Manifest content
- Infrastructure metadata (instance IDs, IPs)
- SSH private key
- Controllers config (validator keys and addresses)
- Terraform state (so the importer can also destroy the infrastructure)
On your machine (the one that created the testnet):
./quake remote exportThis writes .quake/<testnet-dir>/<testnet-name>-export.json. Send this file to
your colleague.
To create a lighter bundle without Terraform state (e.g. for read-only access where the recipient won't need to destroy the infrastructure):
./quake remote export --exclude-terraformOn the colleague's machine:
./quake remote import path/to/export.jsonremote import is supported on Unix-like platforms only, because Quake must set
SSH private key permissions to 0600.
After importing, all Quake commands work — including quake info, quake remote ssh,
and quake remote destroy. If the bundle includes Terraform state, the importer can
also tear down the infrastructure when done. Bundles exported with
--exclude-terraform support all commands except quake remote destroy and
quake clean (which skip the infrastructure destroy step gracefully).
Important
The export file contains the SSH private key and Terraform state for the remote infrastructure. Treat it as sensitive material: transfer it securely and do not commit it to version control.
quake clean relies on Terraform state to delete the AWS infrastructure for a
remote testnet. If that state is lost (e.g. the .quake/ directory was deleted,
or quake remote create crashed before state was written), the nodes, VPC,
security groups, and related resources stay in AWS with no supported way for
quake to remove them.
crates/quake/scripts/aws-resources.sh is the recovery path. It discovers
resources by the project name they were created with
(arc-<testnet>-testnet-<user>) and removes them using the AWS CLI directly,
no Terraform state required.
# Summarize every orphaned project for the current user.
./crates/quake/scripts/aws-resources.sh list
# Show the full resource plan for a single testnet (no deletion).
./crates/quake/scripts/aws-resources.sh list <testnet>
# Delete orphaned resources for a testnet. Without --yes, the script prints the
# plan and prompts before deleting; with --yes, it proceeds non-interactively.
./crates/quake/scripts/aws-resources.sh remove <testnet>
./crates/quake/scripts/aws-resources.sh remove <testnet> --yeslist is read-only. remove defaults to an interactive confirmation, so a
bare invocation never deletes anything without an explicit y. The script
scopes to the current user ($GITHUB_USER, or --user NAME); remove and
list TESTNET refuse to touch a project belonging to another user unless
--user is passed explicitly, and the bare list summary surfaces a notice
when --user overrides the current $GITHUB_USER. Run --help for the full
option set.
Important
Prefer quake clean whenever the Terraform state is still available. This
script is a recovery tool for orphaned infrastructure only. It matches
resources on AWS-side names and tags alone, so a mistyped testnet or the
wrong --region can tear down infrastructure belonging to an active
testnet.
Both the consensus (CL) and execution (EL) binaries support heap and CPU
profiling via a pprof-compatible HTTP server. This is gated behind the
pprof Cargo feature flag and requires building with the profiling
build profile (release optimizations + debug symbols for readable
flamegraphs).
Analyzing profiles requires either Go or the standalone pprof
tool. Install one of them:
# Option 1: standalone pprof via Homebrew (no Go required)
brew install pprof
# Option 2: use go tool pprof (bundled with any Go installation)
# Install Go from https://go.dev/dl/Flamegraph rendering requires Graphviz:
brew install graphvizEach layer has its own feature environment variable so they can be configured independently:
| Variable | Layer | Default |
|---|---|---|
CL_FEATURES |
Consensus | (empty) |
EL_FEATURES |
Execution | default js-tracer |
Note
quake start skips building when the image tag already exists.
If you previously built without profiling, run quake build
explicitly before quake start to rebuild the images with the
pprof feature.
Use quake build with -p profiling to select the profiling Cargo
build profile. Set the feature variable for the layer(s) you want to
profile:
# CL only
CL_FEATURES=pprof ./quake -f <manifest.toml> build -p profiling
# EL only (include base features)
EL_FEATURES="default js-tracer pprof" \
./quake -f <manifest.toml> build -p profiling
# Both layers
CL_FEATURES=pprof EL_FEATURES="default js-tracer pprof" \
./quake -f <manifest.toml> build -p profilingThe same variables work with make build-docker:
CL_FEATURES=pprof BUILD_PROFILE=profiling make build-dockerAfter building, start the testnet normally:
./quake -f <manifest.toml> startquake build builds for the local architecture only. For remote testnets,
follow the Custom docker images instructions to build
linux/amd64 images and push them to your registry. Add the profiling build
args to the docker build commands:
# Execution layer — add these flags:
--build-arg BUILD_PROFILE=profiling \
--build-arg FEATURES="default js-tracer pprof"
# Consensus layer — add these flags:
--build-arg BUILD_PROFILE=profiling \
--build-arg FEATURES=pprof| Endpoint | Type | Description |
|---|---|---|
/debug/pprof/allocs |
Heap | Snapshot of in-use memory allocations |
/debug/pprof/heap |
Heap | Alias for /debug/pprof/allocs |
/debug/pprof/profile?seconds=N |
CPU | CPU sampling over N seconds |
| Layer | Internal port | Local host port |
|---|---|---|
| CL | 6060 | 6060 + node_index |
| EL | 6061 | 6161 + node_index * 100 |
# CL heap profile from the first validator
pprof -text http://localhost:6060/debug/pprof/heap
go tool pprof -text http://localhost:6060/debug/pprof/heap
# EL heap profile from the first validator
pprof -text http://localhost:6161/debug/pprof/heap
go tool pprof -text http://localhost:6161/debug/pprof/heap
# 30-second CPU profile from the first validator (CL)
pprof -text 'http://localhost:6060/debug/pprof/profile?seconds=30'
go tool pprof -text 'http://localhost:6060/debug/pprof/profile?seconds=30'
# Interactive web UI with flamegraph
pprof -http :8080 http://localhost:6060/debug/pprof/heap
go tool pprof -http :8080 http://localhost:6060/debug/pprof/heapNote
When the pprof feature is not compiled in, the server is a no-op.
Pprof ports are always mapped in compose but nothing will listen
unless the binary was built with the pprof feature.
Pprof requests are proxied through CC (see
Communication via Control Center).
Start the SSM tunnel with quake remote ssm start, then access:
# CL heap profile for first validator
pprof -http :8080 http://localhost:16060/pprof/cl/validator1/debug/pprof/allocs
go tool pprof -http :8080 http://localhost:16060/pprof/cl/validator1/debug/pprof/allocs
# EL 30-second CPU profile for first validator
pprof -text 'http://localhost:16060/pprof/el/validator1/debug/pprof/profile?seconds=30'
go tool pprof -text 'http://localhost:16060/pprof/el/validator1/debug/pprof/profile?seconds=30'The test command provides runtime validation of a running testnet. Tests verify connectivity, sync status, peer connections, and other operational aspects of the testnet. The idea is to have a test suite to run against arbitrary quake scenarios.
TODO: more testing scenarios to come
Run all tests (except excluded groups: validation, health, validator_set, perf):
./quake testRun an excluded group explicitly:
./quake test perf:block_time
./quake test validation:basicRun all tests in a specific group:
./quake test probeRun a single test:
./quake test probe:connectivityRun multiple specific tests:
./quake test probe:connectivity,syncGlob pattern matching - Use * (any characters) and ? (single character) for flexible test selection. Quote patterns to prevent shell expansion:
# Run tests in groups starting with 'n'
./quake test 'n*'
# Run all tests named 'sync' in any group
./quake test '*:sync'
# Run probe tests starting with 'conn'
./quake test 'probe:conn*'
# Run tests containing 'peer' in groups starting with 'n'
./quake test 'n*:*peer*'
# Run tests starting with 's' in groups starting with 'p'
./quake test 'p*:s*'List tests without running - Use --dry-run to see which tests would be executed:
# List all available test groups and tests
./quake test --dry-run
# List tests in a specific group
./quake test probe --dry-run
# List tests matching a pattern
./quake test 'n*:*peer*' --dry-runConfigure RPC timeout for tests (default is 1 second):
./quake test --rpc-timeout 5sprobe - Basic connectivity and sync validation:
connectivity- Verifies all nodes are reachable via RPC and returns their current block heightsync- Checks that all nodes have completed syncing (not currently syncing)
net - Network peer validation:
peer_count- Ensures all nodes have at least one peer connectioncl_persistent_peers- Verifies that persistent peers defined in the manifest are actually connected
infra - Substrate-level readiness checks (excluded from the default quake test run):
-
latency_emulation- Verifies thetc netemrules inside each node's CL and EL containers match the manifest. Whenlatency_emulation = true, cross-checks each peer's expected delay againstAWS_LATENCY_MATRIXwithin ±10% tolerance. Whenlatency_emulation = false, asserts nonetemqdiscs exist (catches stale rules left over from a prior latency run).Known limitation: the check probes only the container's primary interface (
eth0). Bridge nodes attached to multiple subnets (sentries, relayers) and local compose containers connected tohost-accesscarry additionaleth*interfaces that also receivetc netemrules; the check does not verify those. A node passing the check today guarantees the matrix on its primary interface only. Broken or stale rules on secondary interfaces are not caught.
The infra group is intended as a pre-experiment readiness gate on a running
testnet, not a CI signal. It shell-execs into every container, so cost scales
with node count; especially relevant on remote testnets where each probe
involves an SSH hop.
Tests automatically register themselves using the #[quake_test] macro. To add new tests:
-
Create a new test file in
crates/quake/src/tests/(e.g.,foobar.rs) or add to an existing file:use tracing::debug; use super::{quake_test, in_parallel, CheckResult, RpcClientFactory, TestOutcome, TestResult}; use crate::testnet::Testnet; /// Test description #[quake_test(group = "foobar", name = "my_test")] fn my_test<'a>( testnet: &'a Testnet, factory: &'a RpcClientFactory, ) -> TestResult<'a> { Box::pin(async move { debug!("Running my test..."); // Example: Check all nodes are reachable let node_urls = testnet.nodes_metadata.all_execution_urls(); let results = in_parallel(&node_urls, factory, |client| async move { client.get_block_number().await }) .await; // Use structured test results let mut outcome = TestOutcome::new(); for (name, url, result) in results { match result { Ok(block_number) => { outcome.add_check(CheckResult::success( name, format!("{} (block #{})", url, block_number), )); } Err(e) => { outcome.add_check(CheckResult::failure( name, format!("{} - Error: {}", url, e), )); } } } outcome.with_summary("All nodes checked").into_result() }) }
-
Add module declaration in
crates/quake/src/tests/mod.rs:mod foobar;
-
Run your new test:
./quake test foobar:my_test
Key points:
- Tests register automatically via
#[quake_test(group = "...", name = "...")]macro - Each
group:namecombination must be unique (enforced at compile time) - Tests use async functions that return
Pin<Box<dyn Future<Output = Result<()>>>>wrapped inBox::pin(async move { ... }) - The
RpcClientFactoryprovides RPC clients with consistent timeout configuration - Use the
in_parallel()helper for concurrent RPC operations across nodes - Use
TestOutcomeandCheckResultfor structured test reporting with consistent formatting - All test output goes to stdout with
✓for success and✗for failure
Tests network resilience under chaos conditions with transaction load and validator set changes.
./scripts/scenarios/nightly-chaos-testing.sh [scenario] [spam_duration_in_seconds] [tx_rate]Our CI workflow runs this script every night with:
./scripts/scenarios/nightly-chaos-testing.sh crates/quake/scenarios/nightly-chaos-testing.toml 3600 1000Which loads the network for 3600s (1 hour) at a 1000tx/s rate, while continuously running chaos testing.