Testing — Developer Quick Reference

One-page orientation. See CI Pipeline for the automation details.

TL;DR

  • Writing code: run cargo test -p <crate-you-touched> as you go.
  • Committing: nothing happens. No hook tests a commit.
  • Per batch of work: run just standard (Tier 1) yourself. Nothing starts it for you.
  • Pushing: Tier 0 runs automatically and blocks on fail.
  • Daily: Tier 3 (02:00) runs via systemd. Tier 2 is on demand — nothing starts it for you.
  • Suite overview: just status.

Prerequisites (system tools)

The interop and integration tests shell out to real external tools, so these system packages must be installed (all apt-installable on Debian):

  • docker — the Tier 2 scenario corpora (run by periculum).
  • python3 + Python RNS (rns) — the rnsd_interop tests drive a real Python rnsd/rnstatus as the compatibility reference.
  • socat — bridges a virtual serial pty pair so the serial-family interfaces (KISS, AX.25, Pipe) can be interop-tested against a real Python peer.
  • nomadnet (optional) — the real NomadNet node used by the on-demand lnomad acceptance (scripts/lnomad_nomadnet_acceptance.sh, see The lnomad acceptance). Not part of any tier; install with pip install nomadnet.
  • i2pd (optional) — provides the SAM bridge on 127.0.0.1:7656 that the I2PInterface live tests use. The default suite covers I2PInterface with an in-process mock SAM bridge, so i2pd is not needed to go green; it only gates the #[ignore]d live tests (cargo test -p leviculum-std i2pd_live -- --ignored). Enable the bridge with sam.enabled = true in /etc/i2pd/i2pd.conf and start the i2pd service.
  • cargo-fuzz + nightly (optional) — drive the wire-format parser fuzz harness under leviculum-core/fuzz and leviculum-std/fuzz (see Fuzzing the wire parsers). Not part of any tier; install with cargo install cargo-fuzz && rustup toolchain install nightly.
  • just, cargo, flock, notify-send — build/CI plumbing.

scripts/install-ci.sh checks for these at setup and prints the sudo apt install hint for any that are missing. Whenever a test starts depending on a new tool, add it BOTH here and to that check list so a fresh machine can be set up from scratch.

The four tiers

TierWhenCommandTimeScope
0on git push (hook)just fast~3 minfmt + clippy + workspace lib tests
1on demand, once per batch1just standard~15 min (40 min cold2)Tier 0 + core/tests + ffi + proxy + rnsd_interop
2on demand: systemctl --user start leviculum-ci-tier2.service3just extensive30–90 minTier 1 + periculum conformance/ + regression/
302:00 daily (systemd timer)just nightly2–6 hTier 2 + LNode flash-from-HEAD + periculum hardware/

Each tier includes every lower tier, so a green nightly proves the whole stack.

just guards is not a fifth tier: it is the coder pass's standing gate — run beside cargo fmt, cargo clippy -D warnings and cargo test --workspace after every batch — and it is the subset of Tier 0's guards, censuses and selftests that costs under ten seconds each, about 20 s warm in total. fast and standard stay the landing gate's, unchanged: a green guards is an early verdict on part of Tier 0, never a substitute for it. It exists because the recipes in it are the ones a coder never ran and the landing gate did, which cost two landings an hour each on 2026-09-26 (check-supervised-spawns, check-source-invariant-census). The fuzz crates' committed Cargo.lock files are a guards precondition too (check-fuzz-lock, 0.1 s, Codeberg #442), because the land gate up to bc07cee3 went red in fuzz-regress on a lock that 8042e34a's new path dependency had made stale, with every fuzz target green.

That the list really is a subset is checked and not promised: check-guards-subset (in guards itself) reads just --dump and refuses a guards member fast never reaches, a member the two lists run in different orders, and a member of fast that is in neither guards nor the ledger of not-in-guards: reasons in the Justfile.

Results go to ~/.local/state/leviculum-ci/last-results.txt: GREEN = passed, RED = failed, SKIPPED = deferred because another test held the lock (see "Concurrent runs" below). No tier raises a notify-send alarm today — read the ledger, or just status. See CI Pipeline.

While writing code

Fast feedback. Run only what you changed:

cargo test -p leviculum-core --lib   # touched core lib code
cargo test -p leviculum-std          # touched std
cargo clippy -p leviculum-core       # clippy for one crate
cargo fmt                             # apply formatter (not --check)

End of a batch of work:

just standard                        # Tier 1 (~15 min)

This is the "15-minute-budget" check that CLAUDE.md expects after every task, and typing it is the only thing that runs it.

Never run a full scenario corpus casually:

# DON'T do this without a reason — the containers and the USB handles
# collide with anything else running scenarios on the box.
periculum run ../periculum/conformance

If you must, use just extensive, which builds the binaries the nodes mount first and is lock-protected.

Before pushing

Nothing to type. git push triggers .githooks/pre-push, which lints the Woodpecker pipelines (.githooks/pre-push:204) and then runs just fast (Tier 0, .githooks/pre-push:207). A red Tier 0 aborts the push — fix, commit, and push again.

Three cheap guards run before those minutes are spent. Two are about what reaches a public forge: only master and tags go to Codeberg (.githooks/pre-push:31), and no commit carrying CLAUDE.md, .mcp.json or .claude/ goes there at all. The third is about the gates themselves — they test the working tree, so the working tree has to be what is being pushed. A tree with uncommitted tracked changes (.githooks/pre-push:144), or a push that would move master to anything but HEAD (.githooks/pre-push:162), is refused before the first gate starts: otherwise the verdict describes code that is not being pushed, in either direction. Untracked files are exempt; they are in no commit.

To push a sha this tree is not standing on, run scripts/push-clean.sh <sha> [<remote>]. It keeps a clone outside the tree ($LEV_PUSH_TREE, default ~/.cache/leviculum/push-tree), checks the sha out there detached, initialises the submodules, points that clone's remote at the URL this repository uses for it, and — the part a hand-written git clone silently omits — sets core.hooksPath, so the gates actually run on what is being pushed. The clone keeps its build cache between runs.

just fast runs the guards and that script against scratch repositories (just prepush-guard, ~0.3 s), so a guard broken by an edit is caught by the next push instead of by the push it wrongly refuses. One of its cases pushes from a clone with core.hooksPath unset and requires the proof to fail there: a push that arrives is no evidence a hook ran.

The rest of the hook: it used to also block on Tier 2 staleness, at 5 commits/8 h (warn) and 10 commits/24 h (block); the block was unsatisfiable and was removed on 2026-08-07, along with the git push --no-verify habit it taught. See CI Pipeline.

After committing

Nothing. There is no post-commit hook — deliberately, since 2026-08-07 (footnote 1 above). Tier 1 is just standard, typed once per batch.

Checking state

just status                                  # last result per tier
just logs                                    # tail most recent Tier 1 log
cat ~/.local/state/leviculum-ci/last-results.txt   # full history

LoRa hardware tests

Tier 3 only. Requires two Heltec T114 boards + two RNode radios connected via USB. Manual runs:

just flash                        # flashes ALL attached T114s; touch-free
                                  #   since the Bug #13 firmware change.
                                  #   Double-tap RESET only if the runner
                                  #   prompts you (crashed-firmware fallback).
just flash-one /dev/ttyACM3       # flash one specific T114 (A/B testing)
just nightly                      # full Tier 3 run

A single LoRa scenario in isolation:

periculum run ../periculum/hardware/<name>.toml

Hardware scenarios are not gated behind a flag: they live in hardware/, and periculum decides from the scenario itself whether this bench can serve it. One that binds a board the bench does not hold reports SKIPPED_INFRA naming what was missing, never RED.

Radio duty-cycle lock is OFF by default in tests

The harness writes airtime_limit_long = 0 into every generated radio interface (single RNode, multi-vport RNode, serial LNode), so the firmware duty-cycle airtime lock never engages mid-run. Without this, the driver's lawful-by-default ETSI cap (#55) silently stops a saturating sender once its rolling-hour airtime hits 10 %: the modem stops radiating while still accepting frames, which reads from above as an intermittent resource stall. A test that itself exercises the duty-cycle lock opts back in explicitly:

[radio]
frequency = 869463000
airtime_limit_long = 10   # percent; arms the 10 % ETSI cap

or per subinterface under [[nodes.x.rnode_interfaces]], or for a one-off run with LORA_AIRTIME_LIMIT_LONG=10.

Concurrent runs

Only one scenario run can be in flight at a time — Docker names and USB handles would otherwise collide. A second invocation exits in under a second:

[leviculum] Another integration test is already running.
[leviculum] Current holder:
[leviculum]   pid=12345
[leviculum]   started=2026-04-14T02:01:33
[leviculum]   pkg=periculum
[leviculum]   binary=periculum
[leviculum]   cwd=/path/to/leviculum
[leviculum] Wait for it to finish or stop that process, then retry.

A scheduled Tier 2 / Tier 3 that hits this case logs SKIPPED, not RED, and sends a normal (not critical) notification. No action needed — the next scheduled slot runs normally. In practice this means: if you're doing late-night hardware work and the 02:00 nightly fires, it silently defers. You don't need to stop it.

Unit tests in the leviculum crates (leviculum-core, leviculum-std, leviculum-ffi, leviculum-proxy, leviculum-cli) run in parallel with a held scenario lock — they never touch containers or boards.

Installing / updating the CI

just install-ci

Idempotent. Installs git hooks, systemd user units, state dirs, separate cargo target dir, and the pinned tools that live in venvs the runner owns — esptool, and the Python Reticulum periculum's host type = "python" nodes run (~/.local/state/leviculum-ci/rns-<pin>, built by ../periculum/scripts/install-python-runtime.sh; without it those cells skip with reason=image_runtime_missing). Safe to re-run after pulling.

The lnomad acceptance

lnomad is the terminal NomadNet browser. Its end-to-end acceptance drives a real NomadNet node rather than a mock: NomadNet runs as the shared Reticulum instance and as a node server hosting a known index.mu; lnomad --print fetches and renders that page over the shared-instance path, and the script asserts the rendered output contains the known content.

../periculum/periculum/assets/scripts/lnomad_nomadnet_acceptance.sh

It prints ACCEPT-PASS / LNOMAD-ACCEPT-COMPLETE and exits 0 on success, or ACCEPT-FAIL: <reason> and non-zero otherwise. It creates an isolated RNS + NomadNet config under a temp dir and always cleans up (kills nomadnet, removes the temp dir) on exit.

This is an on-demand acceptance — it is NOT wired into any tier, because it needs nomadnet (and its Python RNS) installed and takes ~30 s of real announce/link setup. Requirements and overrides:

  • nomadnet on PATH (or point NOMADNET at the executable).
  • python3 + RNS on PATH (or point PY at the interpreter).
  • The musl lnomad release binary; the script builds it (cargo build --release -p lnomad) if it is missing. Override its location with LNOMAD, or the cargo target dir with CARGO_TARGET_DIR.
  • Tune timing with NN_SETTLE (nomadnet startup, default 25 s) and LNOMAD_TIMEOUT (fetch timeout, default 40 s).

Fuzzing the wire parsers

The functions that parse UNTRUSTED bytes off the wire (packet, resource advertisement, discovery announce app-data, announce field-slicer, IFAC, HDLC/KISS deframers, I2P SAM reply lines) have a coverage-guided fuzz harness (cargo-fuzz / libFuzzer). A parser that panics, overflows, hangs, or OOMs on malformed input is a remote DoS, so each target asserts graceful Err/None.

There are two detached fuzz crates, one per library crate that owns a parser:

  • leviculum-core/fuzz — packet_unpack, resource_advertisement_unpack, discovery_announce, announce_from_packet, ifac_verify, hdlc_deframe, kiss_deframe.
  • leviculum-std/fuzz — sam_parse (the I2P SAM reply-line parser and the base64/destination decoders it feeds).

announce_from_packet (the ReceivedAnnounce::from_packet field-slicer) and sam_parse were added in Codeberg #108 alongside the #23 targets. Both reach crate-internal parsers through a #[cfg(fuzzing)]-gated fuzz module in each crate, so no fuzz-only surface leaks into the normal public API.

Running them

just fuzz runs every target in both crates; scripts/run-fuzz.sh is what it calls. Until Codeberg #290 there was no recipe, no CI step and no schedule, so the targets were run by nobody — and three defects found by hand in September 2026 sit exactly on top of three of them: unbounded msgpack recursion (#263) and a wrapping bin32 length (#267) in resource_advertisement_unpack, an uncapped HDLC accumulator (#271) in hdlc_deframe. A length field, a nesting depth and an unbounded accumulator are what a fuzzer finds in minutes.

just fuzz                    # every target, 60 s each
just fuzz --seconds 900      # the budget a scheduled run wants
just fuzz hdlc_deframe       # one target by name
just fuzz --list             # what would run, without building

Exit codes separate the three outcomes, because a run that could not happen must never look like a clean one: 0 every target ran its budget and found nothing, 1 a crash (the input is kept, with its hash and a hexdump), 2 it could not run — missing nightly toolchain, missing cargo-fuzz, or a fuzz_targets/*.rs that the crate manifest does not register as a [[bin]] and that therefore no run reaches.

Still NOT part of any tier: it needs nightly + cargo-fuzz and the glibc host target (the workspace defaults to musl, which ASan does not want), and even a short run costs minutes — 345 s for all eight targets at 30 s each (measured 2026-09-19, warm registry; 95 s of that is the leviculum-std ASan build).

The corpus persists outside the checkout

The working corpus and any crash input live under ~/.local/state/leviculum-fuzz (LEVICULUM_FUZZ_STATE), not in the fuzz crates:

~/.local/state/leviculum-fuzz/corpus/<crate>/<target>/    inputs libFuzzer kept
~/.local/state/leviculum-fuzz/findings/<crate>/<target>/  crash inputs
~/.local/state/leviculum-fuzz/logs/<timestamp>/           full per-target output

This is the difference between fuzzing and re-fuzzing. A corpus inside the tree would be thrown away by construction: the nightly runs on a fresh clone it deletes when green, so every scheduled run would restart from the checked-in seeds and re-explore the same shallow paths, and its value would plateau on night two. The first run above left 1837 inputs across the eight targets; the next run starts from them.

The runner is a one-line hook for a scheduled job — bash scripts/run-fuzz.sh --seconds <budget>, exit 1 on a crash — and its FUZZ_SUMMARY line is deliberately not spelled SUMMARY, so a nightly that tallies scenario results out of ^SUMMARY lines cannot silently fold fuzz counts into them.

Keeping the runner honest

just fuzz-selftest (on the Tier 0 push path, ~15 s, scripts/test-run-fuzz.sh) injects the failures into a throwaway fuzz crate and asserts what the runner concluded: a target that crashes, a target that does not, a corpus that has to survive between runs, an unregistered target file, and a host without cargo-fuzz. The last one is the case that must never look green. It skips with a named reason where nightly or cargo-fuzz is absent, so the push path does not inherit the toolchain requirement.

Any crash the fuzzer finds is fixed at the root AND pinned by a deterministic regression unit test in the normal suite, so it stays fixed without the fuzzer. See leviculum-core/fuzz/README.md and leviculum-std/fuzz/README.md for the target lists and exposure ranking.

Golden rules

  • Tests are never flaky. A failure is a real bug — diagnose and fix at the root, don't retry until green.
  • Don't commit while tests are red.
  • #[ignore] is only for hardware-dependent tests. For CPU-expensive non-hardware tests, use a Cargo feature flag.

The ignored-test census

An #[ignore]d test is run by nothing. Codeberg #189 found one such suite that had been broken for a month with no gate anywhere to say so, so the size of that bucket is pinned per test unit in scripts/ignored-counts.txt and checked by scripts/check-ignored-counts.py at the end of just standard. The census is exhaustive: every test executable in the workspace plus every package's doc-tests, with units absent from the pin file expected to have zero, so a new test binary is covered without an entry.

A new #[ignore] therefore fails Tier 1 until you either route the test into a tier or raise its number in the pin file — a one-line diff, on purpose, with the reason belonging in the commit message. python3 scripts/check-ignored-counts.py --print dumps the current census in pin-file format.

Routing an ignored test by name (rather than lifting the ignore) is what scripts/run-status-parity.sh does for the three status_parity tests, which need to run serially. A test filter that matches nothing exits 0, so any such script must also assert how many tests actually ran.


  1. A post-commit hook detached scripts/run-tier1.sh after every commit until 2026-08-07. A commit is not a unit anybody wants tested — a WIP commit, an amend and a commit mid-refactor each started the same forty-minute docker run — and there is nothing an author can do about a red gate that lands twenty minutes after the commit it judges. Removed; see CI Pipeline and the hook rule in Checks That Are Actually Checks. ↩ ↩2

  2. just standard typed by hand builds in the repo's own target/. The separate CARGO_TARGET_DIR at ~/.cache/leviculum-ci-target — so IDE builds and CI builds don't fight over the same incremental cache — belongs to scripts/run-tier1.sh, which nothing has started since the post-commit hook went. Either way the first run against a cold target dir compiles the workspace from scratch (~40 min); subsequent runs are incremental (~15 min). ↩

  3. Tier 2 had a 12:30/18:30 timer until 2026-06-12, when it was retired in favour of on-demand runs (scripts/install-ci.sh step 9, which also deletes any timer a previous install left behind). This page went on advertising the timer, and .githooks/pre-push went on telling people to "wait for the next scheduled run" until 2026-08-07 — see CI Pipeline. scripts/ci-status.sh prints how long it has been since a Tier 2 run was recorded. ↩