Testing — Developer Quick Reference
One-page orientation. See CI Pipeline for the automation details.
TL;DR
- Writing code: run
cargo test -p <crate-you-touched>as you go. - Committing: nothing happens. No hook tests a commit.
- Per batch of work: run
just standard(Tier 1) yourself. Nothing starts it for you. - Pushing: Tier 0 runs automatically and blocks on fail.
- Daily: Tier 3 (02:00) runs via systemd. Tier 2 is on demand — nothing starts it for you.
- Suite overview:
just status.
Prerequisites (system tools)
The interop and integration tests shell out to real external tools, so these system packages must be installed (all apt-installable on Debian):
- docker — the Tier 2 scenario corpora (run by periculum).
- python3 + Python RNS (
rns) — thernsd_interoptests drive a real Pythonrnsd/rnstatusas the compatibility reference. - socat — bridges a virtual serial pty pair so the serial-family interfaces (KISS, AX.25, Pipe) can be interop-tested against a real Python peer.
- nomadnet (optional) — the real NomadNet node used by the on-demand lnomad
acceptance (
scripts/lnomad_nomadnet_acceptance.sh, see The lnomad acceptance). Not part of any tier; install withpip install nomadnet. - i2pd (optional) — provides the SAM bridge on 127.0.0.1:7656 that the
I2PInterfacelive tests use. The default suite coversI2PInterfacewith an in-process mock SAM bridge, so i2pd is not needed to go green; it only gates the#[ignore]d live tests (cargo test -p leviculum-std i2pd_live -- --ignored). Enable the bridge withsam.enabled = truein/etc/i2pd/i2pd.confand start thei2pdservice. - cargo-fuzz + nightly (optional) — drive the wire-format parser fuzz
harness under
leviculum-core/fuzzandleviculum-std/fuzz(see Fuzzing the wire parsers). Not part of any tier; install withcargo install cargo-fuzz && rustup toolchain install nightly. - just, cargo, flock, notify-send — build/CI plumbing.
scripts/install-ci.sh checks for these at setup and prints the sudo apt install hint for any that are missing. Whenever a test starts depending on a
new tool, add it BOTH here and to that check list so a fresh machine can be set
up from scratch.
The four tiers
| Tier | When | Command | Time | Scope |
|---|---|---|---|---|
| 0 | on git push (hook) | just fast | ~3 min | fmt + clippy + workspace lib tests |
| 1 | on demand, once per batch1 | just standard | ~15 min (40 min cold2) | Tier 0 + core/tests + ffi + proxy + rnsd_interop |
| 2 | on demand: systemctl --user start leviculum-ci-tier2.service3 | just extensive | 30–90 min | Tier 1 + periculum conformance/ + regression/ |
| 3 | 02:00 daily (systemd timer) | just nightly | 2–6 h | Tier 2 + LNode flash-from-HEAD + periculum hardware/ |
Each tier includes every lower tier, so a green nightly proves the whole stack.
just guards is not a fifth tier: it is the coder pass's standing gate —
run beside cargo fmt, cargo clippy -D warnings and
cargo test --workspace after every batch — and it is the subset of Tier 0's
guards, censuses and selftests that costs under ten seconds each, about 20 s
warm in total. fast and standard stay the landing gate's, unchanged: a
green guards is an early verdict on part of Tier 0, never a substitute for
it. It exists because the recipes in it are the ones a coder never ran and
the landing gate did, which cost two landings an hour each on 2026-09-26
(check-supervised-spawns, check-source-invariant-census).
The fuzz crates' committed Cargo.lock files are a guards precondition
too (check-fuzz-lock, 0.1 s, Codeberg #442), because the land gate up to
bc07cee3 went red in fuzz-regress on a lock that 8042e34a's new path
dependency had made stale, with every fuzz target green.
That the list really is a subset is checked and not promised:
check-guards-subset (in guards itself) reads just --dump and refuses a
guards member fast never reaches, a member the two lists run in different
orders, and a member of fast that is in neither guards nor the ledger of
not-in-guards: reasons in the Justfile.
Results go to ~/.local/state/leviculum-ci/last-results.txt:
GREEN = passed, RED = failed, SKIPPED = deferred because another
test held the lock (see "Concurrent runs" below). No tier raises a
notify-send alarm today — read the ledger, or just status. See
CI Pipeline.
While writing code
Fast feedback. Run only what you changed:
cargo test -p leviculum-core --lib # touched core lib code
cargo test -p leviculum-std # touched std
cargo clippy -p leviculum-core # clippy for one crate
cargo fmt # apply formatter (not --check)
End of a batch of work:
just standard # Tier 1 (~15 min)
This is the "15-minute-budget" check that CLAUDE.md expects after every task, and typing it is the only thing that runs it.
Never run a full scenario corpus casually:
# DON'T do this without a reason — the containers and the USB handles
# collide with anything else running scenarios on the box.
periculum run ../periculum/conformance
If you must, use just extensive, which builds the binaries the nodes
mount first and is lock-protected.
Before pushing
Nothing to type. git push triggers .githooks/pre-push, which lints
the Woodpecker pipelines (.githooks/pre-push:204) and then runs
just fast (Tier 0, .githooks/pre-push:207). A red Tier 0 aborts the
push — fix, commit, and push again.
Three cheap guards run before those minutes are spent. Two are about
what reaches a public forge: only master and tags go to Codeberg
(.githooks/pre-push:31), and no commit carrying CLAUDE.md,
.mcp.json or .claude/ goes there at all. The third is about the
gates themselves — they test the working tree, so the working tree has
to be what is being pushed. A tree with uncommitted tracked changes
(.githooks/pre-push:144), or a push that would move master to
anything but HEAD (.githooks/pre-push:162), is refused before the
first gate starts: otherwise the verdict describes code that is not
being pushed, in either direction. Untracked files are exempt; they are
in no commit.
To push a sha this tree is not standing on, run
scripts/push-clean.sh <sha> [<remote>]. It keeps a clone outside the
tree ($LEV_PUSH_TREE, default ~/.cache/leviculum/push-tree), checks
the sha out there detached, initialises the submodules, points that
clone's remote at the URL this repository uses for it, and — the part a
hand-written git clone silently omits — sets core.hooksPath, so the
gates actually run on what is being pushed. The clone keeps its build
cache between runs.
just fast runs the guards and that script against scratch repositories
(just prepush-guard, ~0.3 s), so a guard broken by an edit is caught by
the next push instead of by the push it wrongly refuses. One of its cases
pushes from a clone with core.hooksPath unset and requires the proof to
fail there: a push that arrives is no evidence a hook ran.
The rest of the hook: it used to also block on Tier 2 staleness, at
5 commits/8 h (warn) and 10 commits/24 h (block); the block was
unsatisfiable and was removed on 2026-08-07, along with the
git push --no-verify habit it taught. See
CI Pipeline.
After committing
Nothing. There is no post-commit hook — deliberately, since
2026-08-07 (footnote 1 above). Tier 1 is just standard, typed once
per batch.
Checking state
just status # last result per tier
just logs # tail most recent Tier 1 log
cat ~/.local/state/leviculum-ci/last-results.txt # full history
LoRa hardware tests
Tier 3 only. Requires two Heltec T114 boards + two RNode radios connected via USB. Manual runs:
just flash # flashes ALL attached T114s; touch-free
# since the Bug #13 firmware change.
# Double-tap RESET only if the runner
# prompts you (crashed-firmware fallback).
just flash-one /dev/ttyACM3 # flash one specific T114 (A/B testing)
just nightly # full Tier 3 run
A single LoRa scenario in isolation:
periculum run ../periculum/hardware/<name>.toml
Hardware scenarios are not gated behind a flag: they live in
hardware/, and periculum decides from the scenario itself whether
this bench can serve it. One that binds a board the bench does not
hold reports SKIPPED_INFRA naming what was missing, never RED.
Radio duty-cycle lock is OFF by default in tests
The harness writes airtime_limit_long = 0 into every generated
radio interface (single RNode, multi-vport RNode, serial LNode), so
the firmware duty-cycle airtime lock never engages mid-run. Without
this, the driver's lawful-by-default ETSI cap (#55) silently stops a
saturating sender once its rolling-hour airtime hits 10 %: the modem
stops radiating while still accepting frames, which reads from above
as an intermittent resource stall. A test that itself exercises the
duty-cycle lock opts back in explicitly:
[radio]
frequency = 869463000
airtime_limit_long = 10 # percent; arms the 10 % ETSI cap
or per subinterface under [[nodes.x.rnode_interfaces]], or for a
one-off run with LORA_AIRTIME_LIMIT_LONG=10.
Concurrent runs
Only one scenario run can be in flight at a time — Docker names and USB handles would otherwise collide. A second invocation exits in under a second:
[leviculum] Another integration test is already running.
[leviculum] Current holder:
[leviculum] pid=12345
[leviculum] started=2026-04-14T02:01:33
[leviculum] pkg=periculum
[leviculum] binary=periculum
[leviculum] cwd=/path/to/leviculum
[leviculum] Wait for it to finish or stop that process, then retry.
A scheduled Tier 2 / Tier 3 that hits this case logs SKIPPED,
not RED, and sends a normal (not critical) notification. No
action needed — the next scheduled slot runs normally. In
practice this means: if you're doing late-night hardware work and
the 02:00 nightly fires, it silently defers. You don't need to
stop it.
Unit tests in the leviculum crates (leviculum-core,
leviculum-std, leviculum-ffi, leviculum-proxy,
leviculum-cli) run in parallel with a held scenario lock — they
never touch containers or boards.
Installing / updating the CI
just install-ci
Idempotent. Installs git hooks, systemd user units, state dirs,
separate cargo target dir, and the pinned tools that live in venvs the
runner owns — esptool, and the Python Reticulum periculum's host
type = "python" nodes run
(~/.local/state/leviculum-ci/rns-<pin>, built by
../periculum/scripts/install-python-runtime.sh; without it those
cells skip with reason=image_runtime_missing). Safe to re-run after
pulling.
The lnomad acceptance
lnomad is the terminal NomadNet browser. Its end-to-end acceptance drives a
real NomadNet node rather than a mock: NomadNet runs as the shared Reticulum
instance and as a node server hosting a known index.mu; lnomad --print
fetches and renders that page over the shared-instance path, and the script
asserts the rendered output contains the known content.
../periculum/periculum/assets/scripts/lnomad_nomadnet_acceptance.sh
It prints ACCEPT-PASS / LNOMAD-ACCEPT-COMPLETE and exits 0 on success, or
ACCEPT-FAIL: <reason> and non-zero otherwise. It creates an isolated RNS +
NomadNet config under a temp dir and always cleans up (kills nomadnet, removes
the temp dir) on exit.
This is an on-demand acceptance — it is NOT wired into any tier, because it
needs nomadnet (and its Python RNS) installed and takes ~30 s of real
announce/link setup. Requirements and overrides:
- nomadnet on PATH (or point
NOMADNETat the executable). - python3 + RNS on PATH (or point
PYat the interpreter). - The musl
lnomadrelease binary; the script builds it (cargo build --release -p lnomad) if it is missing. Override its location withLNOMAD, or the cargo target dir withCARGO_TARGET_DIR. - Tune timing with
NN_SETTLE(nomadnet startup, default 25 s) andLNOMAD_TIMEOUT(fetch timeout, default 40 s).
Fuzzing the wire parsers
The functions that parse UNTRUSTED bytes off the wire (packet, resource
advertisement, discovery announce app-data, announce field-slicer, IFAC,
HDLC/KISS deframers, I2P SAM reply lines) have a coverage-guided fuzz harness
(cargo-fuzz / libFuzzer). A parser that panics, overflows, hangs, or OOMs on
malformed input is a remote DoS, so each target asserts graceful Err/None.
There are two detached fuzz crates, one per library crate that owns a parser:
leviculum-core/fuzz—packet_unpack,resource_advertisement_unpack,discovery_announce,announce_from_packet,ifac_verify,hdlc_deframe,kiss_deframe.leviculum-std/fuzz—sam_parse(the I2P SAM reply-line parser and the base64/destination decoders it feeds).
announce_from_packet (the ReceivedAnnounce::from_packet field-slicer) and
sam_parse were added in Codeberg #108 alongside the #23 targets. Both reach
crate-internal parsers through a #[cfg(fuzzing)]-gated fuzz module in each
crate, so no fuzz-only surface leaks into the normal public API.
Running them
just fuzz runs every target in both crates; scripts/run-fuzz.sh is what it
calls. Until Codeberg #290 there was no recipe, no CI step and no schedule, so
the targets were run by nobody — and three defects found by hand in September
2026 sit exactly on top of three of them: unbounded msgpack recursion (#263)
and a wrapping bin32 length (#267) in resource_advertisement_unpack, an
uncapped HDLC accumulator (#271) in hdlc_deframe. A length field, a nesting
depth and an unbounded accumulator are what a fuzzer finds in minutes.
just fuzz # every target, 60 s each
just fuzz --seconds 900 # the budget a scheduled run wants
just fuzz hdlc_deframe # one target by name
just fuzz --list # what would run, without building
Exit codes separate the three outcomes, because a run that could not happen
must never look like a clean one: 0 every target ran its budget and found
nothing, 1 a crash (the input is kept, with its hash and a hexdump), 2
it could not run — missing nightly toolchain, missing cargo-fuzz, or a
fuzz_targets/*.rs that the crate manifest does not register as a [[bin]]
and that therefore no run reaches.
Still NOT part of any tier: it needs nightly + cargo-fuzz and the glibc host target (the workspace defaults to musl, which ASan does not want), and even a short run costs minutes — 345 s for all eight targets at 30 s each (measured 2026-09-19, warm registry; 95 s of that is the leviculum-std ASan build).
The corpus persists outside the checkout
The working corpus and any crash input live under
~/.local/state/leviculum-fuzz (LEVICULUM_FUZZ_STATE), not in the fuzz
crates:
~/.local/state/leviculum-fuzz/corpus/<crate>/<target>/ inputs libFuzzer kept
~/.local/state/leviculum-fuzz/findings/<crate>/<target>/ crash inputs
~/.local/state/leviculum-fuzz/logs/<timestamp>/ full per-target output
This is the difference between fuzzing and re-fuzzing. A corpus inside the tree would be thrown away by construction: the nightly runs on a fresh clone it deletes when green, so every scheduled run would restart from the checked-in seeds and re-explore the same shallow paths, and its value would plateau on night two. The first run above left 1837 inputs across the eight targets; the next run starts from them.
The runner is a one-line hook for a scheduled job — bash scripts/run-fuzz.sh --seconds <budget>, exit 1 on a crash — and its FUZZ_SUMMARY line is
deliberately not spelled SUMMARY, so a nightly that tallies scenario results
out of ^SUMMARY lines cannot silently fold fuzz counts into them.
Keeping the runner honest
just fuzz-selftest (on the Tier 0 push path, ~15 s,
scripts/test-run-fuzz.sh) injects the failures into a throwaway fuzz crate
and asserts what the runner concluded: a target that crashes, a target that
does not, a corpus that has to survive between runs, an unregistered target
file, and a host without cargo-fuzz. The last one is the case that must never
look green. It skips with a named reason where nightly or cargo-fuzz is absent,
so the push path does not inherit the toolchain requirement.
Any crash the fuzzer finds is fixed at the root AND pinned by a deterministic
regression unit test in the normal suite, so it stays fixed without the fuzzer.
See leviculum-core/fuzz/README.md and leviculum-std/fuzz/README.md for the
target lists and exposure ranking.
Golden rules
- Tests are never flaky. A failure is a real bug — diagnose and fix at the root, don't retry until green.
- Don't commit while tests are red.
#[ignore]is only for hardware-dependent tests. For CPU-expensive non-hardware tests, use a Cargo feature flag.
The ignored-test census
An #[ignore]d test is run by nothing. Codeberg #189 found one such
suite that had been broken for a month with no gate anywhere to say so,
so the size of that bucket is pinned per test unit in
scripts/ignored-counts.txt and checked by
scripts/check-ignored-counts.py at the end of just standard. The
census is exhaustive: every test executable in the workspace plus every
package's doc-tests, with units absent from the pin file expected to
have zero, so a new test binary is covered without an entry.
A new #[ignore] therefore fails Tier 1 until you either route the test
into a tier or raise its number in the pin file — a one-line diff, on
purpose, with the reason belonging in the commit message. python3 scripts/check-ignored-counts.py --print dumps the current census in
pin-file format.
Routing an ignored test by name (rather than lifting the ignore) is what
scripts/run-status-parity.sh does for the three status_parity tests,
which need to run serially. A test filter that matches nothing exits 0,
so any such script must also assert how many tests actually ran.
-
A
post-commithook detachedscripts/run-tier1.shafter every commit until 2026-08-07. A commit is not a unit anybody wants tested — a WIP commit, an amend and a commit mid-refactor each started the same forty-minute docker run — and there is nothing an author can do about a red gate that lands twenty minutes after the commit it judges. Removed; see CI Pipeline and the hook rule in Checks That Are Actually Checks. ↩ ↩2 -
just standardtyped by hand builds in the repo's owntarget/. The separateCARGO_TARGET_DIRat~/.cache/leviculum-ci-target— so IDE builds and CI builds don't fight over the same incremental cache — belongs toscripts/run-tier1.sh, which nothing has started since the post-commit hook went. Either way the first run against a cold target dir compiles the workspace from scratch (~40 min); subsequent runs are incremental (~15 min). ↩ -
Tier 2 had a 12:30/18:30 timer until 2026-06-12, when it was retired in favour of on-demand runs (
scripts/install-ci.shstep 9, which also deletes any timer a previous install left behind). This page went on advertising the timer, and.githooks/pre-pushwent on telling people to "wait for the next scheduled run" until 2026-08-07 — see CI Pipeline.scripts/ci-status.shprints how long it has been since a Tier 2 run was recorded. ↩