Leviculum
Leviculum is a Rust implementation of the Reticulum network stack. It is wire-compatible with the Python reference implementation and runs on Linux, macOS, and embedded devices.
What is Reticulum?
Reticulum is a networking stack for building resilient, encrypted mesh networks over any transport medium. It works over LoRa radios, TCP, UDP, serial links, or anything that can carry bytes. Every node gets a cryptographic identity. Every connection is end-to-end encrypted. No servers, no accounts, no infrastructure required.
What does leviculum do?
Leviculum provides the same functionality as Python Reticulum but compiled to native code. The lnsd daemon is a drop-in replacement for rnsd. The lncp file transfer tool replaces rncp. Python CLI tools like rnstatus, rnpath, and rnprobe work against a running lnsd without modification.
The protocol core (leviculum-core) compiles as no_std with only alloc, so it runs on microcontrollers. The same code powers the Linux daemon, a future Android app, and embedded firmware.
Who this manual is for
- Daemon users running
lnsdalongside or instead ofrnsd: start with the lnsd quickstart and the man pages. - Developers embedding or extending the stack: read the Concepts part — the Architecture overview plus Interface Isolation, Python-RNS Compatibility, Identity and Forward Secrecy, and Storage and Embedding.
- Firmware flashers putting Leviculum on nRF52 boards: see the firmware section and the RNode protocol reference.
The Concepts part explains the non-obvious design ideas; the appendix carries the authoritative Reticulum and LXMF specifications.
Tools
Leviculum ships four binaries:
- lnsd -- the Reticulum network daemon
- lnstest -- test and diagnostics tool: integration self-test, diagnostic bundles, identity management, and interactive sessions
- lncp -- standalone file transfer utility (compatible with Python
rncp) - lnstatus -- network status tool (compatible with Python
rnstatus)
Architecture Overview
This is the entry point to the Concepts part of the manual. It covers the sans-IO core, the crate split, the driver event loop, and the platform-abstraction traits — the mechanics that the four concept pages build on:
- Interface Isolation — why only the interface knows its medium's quirks.
- Python-RNS Compatibility — wire/semantic compatibility and the drop-in daemon, vs. internal parity (not a goal).
- Identity and Forward Secrecy — dual keypairs, derived destinations, ratchets.
- Storage and Embedding — the
Clock/Storage/Interfacetraits that let one core run on a host or a microcontroller.
The crate split
The protocol logic lives in one no_std crate; everything platform-
specific wraps around it:
| Crate | Role |
|---|---|
leviculum-core | All protocol logic, #![no_std] + alloc, zero async (leviculum-core/src/lib.rs:59). |
leviculum-std | Host driver: tokio event loop, interfaces, FileStorage, RPC, config. |
leviculum-nrf | Embedded driver: Embassy event loop on nRF52 (cross-compiled, outside the host workspace). |
leviculum-ffi | C ABI over the core for other-language bindings. |
leviculum-cli | The lnsd / lnstest / lncp binaries. |
The application boundary is NodeCore: feed it bytes via
handle_packet / handle_timeout and drain a
TickOutput { actions, events }. The core decides what to send; the
driver decides how and when to put it on the wire. See
Storage and Embedding for the
injected Clock/Storage/Interface traits that make this portable.
Sans-I/O Core
┌─────────────────────────────────┐
│ leviculum-core │
│ │
handle_packet() ──►│ NodeCore<R, C, S> │──► TickOutput {
(iface_id, data) │ ├── Transport (routing) │ actions: Vec<Action>,
│ ├── Links + Channels │ events: Vec<NodeEvent>,
handle_timeout() ─►│ └── Destinations │ }
│ │
next_deadline() ──►│ Returns: Option<u64> │
└─────────────────────────────────┘
Action::SendPacket { iface, data, peer } — send to one interface,
optionally naming the peer
behind it the bytes are for
Action::Broadcast { data, exclude } — send to all interfaces (except one)
Driver Event Loop
The leviculum-std driver has 6 select! branches:
#![allow(unused)] fn main() { loop { select! { // 1. Packet from any interface (iface_id, data) = registry.recv_any() => { output = core.handle_packet(iface_id, &data); post_dispatch(output); } // 2. External action (connect, send, announce) output = action_dispatch_rx.recv() => { post_dispatch(output); } // 3. Timer fires _ = sleep_until(next_poll) => { output = core.handle_timeout(); post_dispatch(output); } // 4. Shutdown _ = shutdown.changed() => break // 5. New interface (TCP accept, local client connect) handle = new_interface_rx.recv() => { registry.register(handle); output = core.handle_interface_up(iface_idx); post_dispatch(output); } // 6. Periodic storage flush (crash protection, hourly) _ = sleep_until(next_flush) => { core.storage_mut().flush(); } } } }
Post-dispatch (after every core call)
dispatch_actions(&mut ifaces, &output.actions)— routes Actions to interfaces (protocol logic in core)- React to errors —
BufferFull: log.Disconnected: callhandle_interface_down() - Forward
output.eventsto the application - Schedule
handle_timeout()fromoutput.next_deadline_ms
Interface Trait
#![allow(unused)] fn main() { pub trait Interface { fn id(&self) -> InterfaceId; fn name(&self) -> &str; fn mtu(&self) -> usize; fn is_online(&self) -> bool; fn try_send(&mut self, data: &[u8]) -> Result<(), InterfaceError>; } }
Send-only. Receive is driver-specific (tokio: mpsc::poll_recv, Embassy:
interrupt DMA, bare-metal: poll FIFO). try_send is fire-and-forget:
Reticulum is best-effort, higher layers retransmit.
dispatch_actions() lives in core (not the driver) because action routing
(broadcast exclusion, interface selection) is protocol knowledge.
In leviculum-std, InterfaceHandle wraps tokio::sync::mpsc::Sender
behind the trait. An embedded driver implements it directly on a radio struct.
Core processes packets with zero delay. Collision avoidance (jitter, CSMA) is the interface's responsibility — fast interfaces (TCP) transmit immediately, slow interfaces (LoRa) apply send-side jitter. This is the interface-isolation rule in code.
Writing a Driver
1. Create interface objects
Implement Interface on your outbound channel. Register with your own
bookkeeping. Core references interfaces by InterfaceId only.
2. Run the event loop
Minimum 3 branches: receive, timer, shutdown. Feed everything through the post-dispatch sequence above.
3. Handle the receive path
Driver-specific. On complete packet: core.handle_packet(iface_id, &data)
→ post-dispatch. On disconnect: core.handle_interface_down(iface_id).
Packet Flow
Incoming
Interface → deframe → mpsc → recv_any() → handle_packet()
→ Transport::process_incoming() → TickOutput
→ dispatch_actions() → interfaces → wire
→ events → application
Outgoing
Application → connect/send/announce → TickOutput (via action_dispatch)
→ dispatch_actions() → interfaces → wire
Local Client (Shared Instance)
lnstest/lncp → Unix socket → LocalInterface (HDLC)
→ handle_packet() with is_local_client=true
→ local_client_known_dests updated (6h TTL)
RPC (rnstatus, rnpath, rnprobe)
Python CLI → Unix socket → RPC server (multiprocessing.connection, pickle)
→ handlers query NodeCore state or trigger probe
→ pickle response → CLI
The shared-instance socket and this RPC channel are what make lnsd a
drop-in for rnsd; see
Python-RNS Compatibility.
IPC platform support
The shared-instance data channel and the RPC control channel use abstract Unix sockets on Linux, filesystem Unix sockets on macOS/BSD, and TCP loopback on Windows (mirroring Python-RNS's AF_INET fallback). Linux is the tested path and is the one exercised by our CI; macOS/Windows IPC is community-supported and not exercised by our CI.
Storage Trait
For the conceptual rationale (one core, host or embedded backend) see Storage and Embedding; for the per-method deep dive see Storage Trait Split Analysis.
Type-safe methods organized by collection:
| Collection | Key methods |
|---|---|
| Packet dedup | has_packet_hash, add_packet_hash |
| Path table | get_path, set_path, remove_path, expire_paths |
| Reverse table | get_reverse, set_reverse, remove_reverse |
| Link table | get_link_entry, set_link_entry, remove_link_entry |
| Announce table | get_announce, set_announce, remove_announce |
| Announce cache | get_announce_cache, set_announce_cache |
| Receipts | get_receipt, set_receipt, remove_receipt |
| Ratchets | load_ratchet, store_ratchet, list_ratchet_keys |
| Cleanup | expire_* per collection |
Shared types in storage_types.rs: PathEntry, ReverseEntry, LinkEntry,
AnnounceEntry, PacketReceipt.
Implementations: NoStorage (no-op), MemoryStorage (BTreeMap, host/tests),
EmbeddedStorage (heapless FnvIndexMap, fixed capacity, used by leviculum-nrf),
FileStorage (wraps MemoryStorage + disk).
FileStorage Persistence
| File | Format | Strategy | Contents |
|---|---|---|---|
known_destinations | msgpack map | Batch flush (hourly + shutdown) | Identity → destination |
packet_hashlist | msgpack array | Batch flush | 32-byte dedup hashes |
ratchets/{hash} | msgpack map | Write-through | Receiver ratchet keys |
ratchetkeys/{hash} | signed msgpack | Write-through | Sender ratchet private keys |
Non-persistent collections (paths, reverses, links, announces, receipts) are RAM-only and rebuilt from network on restart.
Logging
Sentence-style messages with inline context. Good:
Destination <81b22f60> is now 4 hops away via <ecc35451> on iface 1
Answering path request for <4c0c6c7f> on iface 1, path is known
Bad:
path updated dest=81b22f60 hops=4
Use HexShort for hashes. Always explain drop reasons ("rate limited",
"duplicate packet", "no path known").
| Component | What | Level |
|---|---|---|
| transport process_incoming | Packet dispatch, drop reasons | trace! |
| transport handle_announce | Path updates, rebroadcast decisions | debug! |
| transport forward_packet | Forwarding decisions | debug! |
| node/link_management | Link lifecycle, RTT retry | debug! |
| driver | Startup, interface registration | info! |
| interfaces | Connection events, I/O errors | info!/warn! |
Interface Isolation
The single most important architectural rule in Leviculum:
Only the interface knows the quirks of its carrier medium. The core, the transport, and the daemon are media-agnostic.
A packet is a packet. At the boundary where the core hands bytes to an interface, there is no distinction between an announce, a link request, a data packet, or a resource chunk. They are all just bytes.
What "media-agnostic core" means
leviculum-core decides what to send, to which interface, and —
on an interface that carries several peers — for which peer (the
peer hint on Action::SendPacket, Codeberg #376: an identity, never
a link, a handle or an address, so the interface still owns the map
from peer to link). It never decides when to put a frame on the
wire, never spaces transmissions, and never reasons about
contention. The core processes
every packet with zero delay and emits an Action::SendPacket or
Action::Broadcast immediately (see Architecture).
Because the core is the same code on a Linux daemon, an Android app,
and an nRF52 firmware image, it cannot afford to know whether the
medium underneath is a fibre-fast TCP socket or a half-duplex LoRa
radio whose airtime budget is measured in minutes. Medium awareness
lives entirely on the far side of the
Interface trait.
What an interface is allowed to know
A LoRa interface knows it cannot transmit and receive at the same
time. It knows its RadioSettings (bandwidth, spreading factor,
coding rate) and therefore the airtime cost of any given frame. It
holds packets back, applies its own randomised pre-TX jitter on top of
the RNode firmware's CSMA, and refuses new frames when its airtime
budget is exhausted. Concretely:
- Send-side jitter — packets are queued, not sent immediately; a
frame that acquires an idle channel first serves a randomised wait,
DIFS plus a contention window sized from the radio parameters, so two
nodes do not re-collide (
leviculum-std/src/interfaces/rnode.rs:2742-2773, where the TX loop arms it; the ceiling it reports iscompute_jitter_max_ms,leviculum-std/src/interfaces/rnode.rs:293-303). - CSMA — radio-level carrier sensing is handled by the RNode
firmware on top; the interface hands the modem one frame at a time so
that CSMA runs for every frame, and none of this reaches the core
(
leviculum-std/src/interfaces/rnode.rs:2264-2273). - Airtime backpressure — a per-interface credit bucket charges
every send by its airtime cost and signals
BufferFullrather than flooding the serial queue (leviculum-std/src/interfaces/airtime.rs:1). This explicitly "never leaks intoleviculum-core, so theno_stdcore stays free of host-side backpressure concerns" (same file).
A TCP interface has none of this. It just writes bytes
(leviculum-std/src/interfaces/tcp.rs).
Why the rule is hard, not advisory
The rule binds anyone writing a fix. If a proposed fix for a
collision, contention, or duplex problem introduces an awareness flag
or counter in transport.rs, the node/ modules, or the daemon
("is a link in flight?", "am I forwarding a link request?"), it is at
the wrong layer. Such a fix must be redirected into the interface.
Interface implementations are therefore free to diverge from Python-Reticulum's thin serial-writer style — that divergence is exactly where medium-specific intelligence belongs, and it satisfies the project's deviation rule as long as wire and semantic compatibility are preserved.
Consequences
- The same routing logic runs unchanged over LoRa, TCP, UDP, serial, and the in-process local socket. That includes relaying a packet back out of the interface it arrived on — the same-interface relay decision is taken in the media-agnostic core; the interface is not involved.
- New media are added by implementing one trait, not by threading medium-specific cases through the protocol core.
- Collision-avoidance bugs are debugged in one place — the interface — instead of being smeared across six stack layers.
The one place the medium is named
InterfaceKind (leviculum-core/src/traits.rs) names the carrier —
Tcp, Rnode, Serial, and so on — and that looks like an exception
to the rule. It is not: the kind is reported, never acted on. It
exists so rnstatus can print the Python-RNS interface class name and
so a status consumer can group interfaces by transport instead of by
their peer label.
A match on it outside traits.rs may produce a string, a number or a
status field, and nothing else. As of 2026-07-30 there are exactly two
consumers: transport.rs's sparse-map bookkeeping (where Unknown
means "no entry") and rpc/handlers.rs::interface_type. A third that
decides what the stack does — a longer timeout for LoRa, a skipped
step on serial — is the wrong-layer fix this page describes; widen the
Interface trait so the interface answers the question itself.
This rule is deliberately not machine-checked. A guard could only see
a syntactic comparison against a variant, which is not the shape the
violation takes: an exhaustive match kind returning a timeout reads
identically to one returning a label. Both existing consumers would
need an exemption, so on today's three call sites the exemption list
would be longer than the finding, and it would grow with every
legitimate status field. The rule is stated here and on the enum
instead.
See also: Storage and Embedding for the parallel isolation of persistence and time, and the RNode protocol page for the LoRa carrier details an interface must handle.
An interface that holds several peers
A TCP listener, an AutoInterface, a BLE radio and an I2P endpoint all have the same shape: one configured section, many peers behind it. Each one has to answer the same question, and the answer decides how much the rest of the stack has to know about the carrier:
When the transport wants these bytes to reach one of the peers behind this interface, how does it say which one?
This page records the answers we run today, the answer the references run, what one more interface actually costs measured on the firmware, and which model we take forward. It is a design document, not a status page: what is open belongs on the tracker.
The rule it has to live under is Interface isolation — only the interface knows its medium. Nothing below weakens that; the whole argument is about where the peer-to-link map lives, and in every option it lives on the interface's side of the boundary.
What we do today, and it is not one thing
TCP server: a child interface per connection
The listener binds and accepts; every accepted connection becomes its
own InterfaceHandle with its own id, drawn from the shared counter
(spawn_tcp_server, leviculum-std/src/interfaces/tcp.rs:365). The
child is built from the already-connected stream
(spawn_tcp_interface_from_stream,
leviculum-std/src/interfaces/tcp.rs:432), inherits IFAC, mode and
ingress control from the listener, and is handed to the event loop,
which registers it in the routing map like any other interface
(registry, leviculum-std/src/driver/mod.rs:4917).
The listener itself is deliberately not in that map. It carries no
packets, so it would be a send target that cannot send; it is recorded
in the reporting inventory instead, where the child registers its
display identity and its parent link (add_spawned,
leviculum-std/src/interfaces/tcp.rs:419; listener_id,
leviculum-std/src/interfaces/tcp.rs:427). The split between the
routing map and the reporting inventory is the point of that module
(interface_names, leviculum-std/src/interfaces/inventory.rs:7).
Teardown runs through the ordinary disconnect path: the event loop
notices the channel closed, calls handle_interface_down
(leviculum-std/src/driver/mod.rs:4571) to cull the routing entries,
and the child's byte counters are folded into its parent's departed
totals so the listener's reported traffic does not shrink when a client
leaves (remove_spawned, leviculum-std/src/driver/mod.rs:4086).
AutoInterface, I2P, shared instance: the same shape
- Every discovered AutoInterface peer becomes a separate handle
(
spawn_auto_interface,leviculum-std/src/interfaces/auto_interface/orchestrator.rs:168; the per-peer handle atInterfaceHandle,leviculum-std/src/interfaces/auto_interface/orchestrator.rs:805). - Every accepted I2P stream becomes a handle
(
new_interface_tx,leviculum-std/src/interfaces/i2p/mod.rs:495). - Every accepted shared-instance IPC client becomes a handle
(
add_spawned,leviculum-std/src/interfaces/local.rs:319).
Four multi-peer carriers, one model: a child interface per peer.
BLE: one interface, many links, a peer hint
The fifth is different. One configured BLEInterface section is one
Reticulum interface and one broadcast domain (BLEInterface,
leviculum-std/src/interfaces/ble/mod.rs:5). Every outbound packet
goes through one planner that decides which links get a copy
(plan_tx_to, leviculum-std/src/interfaces/ble/links.rs:898, driven
from send_packet,
leviculum-std/src/interfaces/ble/mod.rs:1029). Since Codeberg #376 the
core supplies the addressee: a path entry carries the identity it was
learned from (via_peer, leviculum-core/src/storage_types.rs:48) and
that identity rides out with the packet. With a hint the planner picks
one link; without one (an announce, a path request) it floods, which is
what a broadcast domain owes its peers.
Since Codeberg #422 there is a third statement, and it is the one that
makes a board a relay between two of its own links: a broadcast the
core is re-sending carries the peer it ARRIVED from
(exclude_peer, leviculum-core/src/transport.rs:524), and the
interface serves every link but that one
(try_send_excluding_peer, leviculum-core/src/traits.rs:347). The
default implementation sends nothing, which is exactly what excluding
the whole interface did, so every single-link carrier is unchanged and
so is lnsd's BLE interface, whose peripheral role notifies its
subscribed centrals as one group and cannot address a subset of them.
The firmware implements it (try_send_excluding_peer,
leviculum-nrf/src/ble/mod.rs:698), where the decision is a pure
function of the registry (TxAim,
leviculum-nrf/ble-tx/src/registry.rs:828). Without it a path request
from the phone died at the board: one InterfaceId covered both links,
so excluding the arrival interface silenced the neighbour board that
was the only node able to answer.
The firmware runs the same shape with fixed ids: serial 0, LoRa 1, BLE
2, set once at startup (set_interface_name,
leviculum-nrf/src/bin/t114.rs:286) and hardcoded in the interface
itself (BleInterface, leviculum-nrf/src/ble/mod.rs:630), with the
announce gate naming the same constant (BLE_IFACE,
leviculum-nrf/src/announce.rs:85). The fan-out is a task that maps
the hint onto a per-link queue (tx_fanout_task,
leviculum-nrf/src/ble/mod.rs:489; LINK_OUT,
leviculum-nrf/src/ble/mod.rs:453).
The receive side is already peer-aware on both stacks. The board
reports which peer a packet came from and when a peer appears or
disappears (handle_packet_from_peer,
leviculum-nrf/src/bin/t114.rs:1153;
handle_interface_peer_lost, leviculum-nrf/src/bin/t114.rs:1181;
handle_interface_peer_up, leviculum-nrf/src/bin/t114.rs:1193; the
same three in handle_packet_from_peer,
leviculum-nrf/src/bin/rak4631.rs:1150), the core stamps the peer onto
the path entry it installs, and a peer loss culls exactly the paths
through it (drop_paths_via_peer,
leviculum-core/src/transport.rs:5434). So the identity-shaped
addressing already exists end to end; the only open question is whether
the send side spends an interface object on it.
How far the two stacks already drift under one model
Both stacks run the peer-hint model for BLE today, so what follows is
not the cost of two models — it is the baseline drift between two
implementations of one model, which is the floor any split model
would build on top of. lnsd's central task always reports CentralGone
when it ends,
including when the dial never connected at all (CentralGone,
leviculum-std/src/interfaces/ble/bluez.rs:400), and the orchestrator
restarts the strict scan phase on that event (CentralGone,
leviculum-std/src/interfaces/ble/mod.rs:827, into note_reset,
leviculum-std/src/interfaces/ble/links.rs:1470). The firmware restarts
its strict phase only when a new neighbour links or the strict rule finds
a target (note_phase_event, leviculum-nrf/src/ble/columba.rs:1685,
under the reset rule PhaseEvent::resets_strict in
leviculum-nrf/ble-tx/src/admission.rs, narrowed by #504 so a known
peer's reconnect and a teardown no longer reset it); a dial that timed
out records at most a dead end and leaves the clock running
(note_dead_end, leviculum-nrf/src/ble/columba.rs:1760).
Same protocol, same shared constant, different behaviour after a failed
dial: lnsd owes another full 30 s strict bound, the board does not.
Neither is obviously wrong. The point is that nobody decided it — it
fell out of the two stacks having different event vocabularies
(CentralGone fires for a dial that never connected; conn_link_down
cannot). That happens under one shared model. Option B below would give
the two stacks different structures as well, and the drift rate is
what it would multiply.
The reference, as a source of ideas
Python-RNS spawns a child interface per connection in five places:
spawned_interfaces (TCPInterface.py:632),
spawned_interfaces (AutoInterface.py:590),
spawned_interfaces (I2PInterface.py:998),
spawned_interfaces (BackboneInterface.py:129) and
spawned_interfaces (WeaveInterface.py:990). The child is appended
to RNS.Transport.interfaces and is from then on an ordinary send
target.
What that buys:
- The transport addresses an interface. There is no hint, no second addressing concept, no interface that means different things depending on an extra argument.
- Per-peer statistics fall out: each child has its own
rxb/txb, andrnstatusshows a row per peer. - Teardown is one code path — detach the child, and everything keyed on its id goes with it.
What it costs: an interface object, its state, and its registration per connection. Python has no BLE interface at all, so on the carrier this page is actually about, the reference offers no precedent.
The nearest thing that does is the ble-reticulum package Columba
uses, which is our wire counterpart. It spawns one child per peer
(BLEPeerInterface,
ble-reticulum/src/ble_reticulum/BLEInterface.py:2380; at commit
07d9413, 2026-01-18, in the sibling checkout). Two things about that
child are worth more than the precedent itself:
- The child is a shim, not an interface. Its
process_outgoing(ble-reticulum/src/ble_reticulum/BLEInterface.py:2437) fetches the fragmenter from its parent, fragments, and hands each fragment back to the parent's shared driver keyed by address (ble-reticulum/src/ble_reticulum/BLEInterface.py:2474). Every piece of per-medium machinery stays on the parent. The child holds an address and two counters. - On the peripheral side the addressing is not real. The GATT
server's
send_notification(ble-reticulum/src/ble_reticulum/BLEGATTServer.py:537) takes acentral_address, and then writes the value to the one TX characteristic (set_value,ble-reticulum/src/ble_reticulum/BLEGATTServer.py:575), which notifies every subscribed central. The address argument only selects which counter to increment.
So the reference implementation of a per-peer interface, on this exact
carrier, does not actually address one peer where the carrier cannot.
That is not a criticism of it — the Columba wire spec has one notify
characteristic, and no software layer can conjure a second one. It is
the reason "the reference spawns children, so we should" is not an
argument here. Our own planner is explicit about the same limit
(plan_tx_to, leviculum-std/src/interfaces/ble/links.rs:898).
The firmware is the exception, and it cuts the other way: the SoftDevice's notification takes a connection handle, so a board can address one peripheral-role link. lnsd, on BlueZ, cannot.
The numbers
Per-interface state in Transport and NodeCore
NodeCore holds no interface-keyed collection of its own; every
per-interface field lives in Transport, and there are 24 of them
(interface_announce_caps, leviculum-core/src/transport.rs:2761
through own_tunnel_ids,
leviculum-core/src/transport.rs:3099 — the BTreeMap<usize, _> and
BTreeSet<usize> fields in that block).
Method, and why not size_of. Summing size_of over those 24
value types would be the wrong number by a wide margin in both
directions: a BTreeMap allocates in nodes of up to 11 entries, so the
first interface pays for a whole node in every map and the next ten pay
nothing, and several of the values are themselves growable
(interface_held_announces, leviculum-core/src/transport.rs:2982, is
a map of maps). What the 96 KiB firmware pool actually sees is
allocator traffic, so that is what was measured: a counting
GlobalAlloc around System — the harness already in the tree as
CountingAlloc (leviculum-core/tests/heap_leak.rs:56) — reporting
net live bytes (allocated minus freed), sampled around building one
Transport with N interfaces registered through the setter sequence
the firmware bins run, then driving 40 rounds of announce RX across all
of them.
Positive control in every run: the transport must hold 3 of 3 peer
paths at the end, or the measurement is discarded as vacuous. That
control earned itself immediately — the first version of this harness
reported a flat 0 B for every N, because it drove NodeCore with
NoStorage and no announce was ever accepted.
The harness itself is deliberately not committed. It would have to
live in leviculum-core/tests/, and fast runs cargo test --workspace --lib, so nothing would ever execute it — a test that runs nowhere is
the exact Guarantee-B failure
Checks that are actually checks is about. It
is ~180 lines and the recipe above is enough to rebuild it: the
allocator from heap_leak.rs, Transport::new with enable_transport: true and a long path_expiry_secs, the seven setters, and
clear_packet_hashes per round so the dedup cache does not mask the
per-interface growth.
Run on i686-unknown-linux-musl, not the host default: the board is a
32-bit-pointer target and every one of those maps is pointer-heavy, so
x86-64 overstates it. How much depends on what is being counted — a
third on the pointer-dominated first registration (1 363 B vs 891 B),
about 5 % on the warm 3 → 4 step (1 411 B vs 1 347 B), 10 % on an
11-interface transport (22 228 B vs 20 136 B). The 32-bit column is the
one quoted below.
| Step | Live-heap delta (i686) |
|---|---|
| registration only, 1st interface into an empty node | 891 B |
| registration only, each further interface | 3 B (the name String) |
| 3 → 4 interfaces, warm | 1 347 B |
| 4 → 5, 5 → 6, 6 → 7, warm | 579 B each |
Registration is nearly free; the cost appears when the interface
carries traffic and the lazily-created maps get their entry. Take
1 347 B for the first extra interface and ~600 B for each one after.
Reproduced identically across two runs, with the one exception of a
single 6 → 7 step where a node split landed differently (939 B) —
which is the BTreeMap node granularity showing, not noise in the
method.
What the firmware can afford
Three measured budgets, all from the T114 on the rig, all post-#372:
| Budget | Measured | Headroom |
|---|---|---|
Heap, 96 KiB pool (HEAP_SIZE, leviculum-nrf/heap-budget/src/lib.rs:54) | worst watermark 65 044 B of 98 304 (rig-run/proof-372-t114.log, 2026-09-08); typical 56 000-57 000 | 33 260 B at the worst point |
Stack, flip-link region below .data | min_free=72 280 of a 104 464 B region, peak_used=32 184 (rig-run/proof-dup-t114.log, 2026-09-10) | ~70 KiB never touched |
| SoftDevice RAM ceiling | 928 B of margin (leviculum-nrf/memory.x:166) | not the relevant budget, see below |
Three BLE children cost 1 347 + 2 × 579 = 2 505 B of heap, 4 041 B if
every step happens to split a node. Against 33 260 B free at the worst
watermark ever observed that is 7-12 %, and against the ~42 000 B free
at the typical watermark it is 6-10 %. The heap affords it.
The 928 B SoftDevice margin does not bound this, and it is worth
being explicit because the number is small enough to look alarming.
That margin sizes the SoftDevice's own RAM requirement, which scales
with conn_count: 15 272 B at two connections, 23 968 B at four
(leviculum-nrf/memory.x:166), and #372 paid for that by moving the app
RAM floor up 8 576 B. A Reticulum interface object is application heap;
it does not appear in sd_ble_enable's requirement at all. Spawning
three children over the same four BLE connections costs the SoftDevice
nothing. A fifth BLE connection would cost about another 4 300 B (the
measured two-to-four slope) and blow the 928 B margin — but that is
equally true today with one interface, and the boot assert catches it
either way.
Stack is likewise not per-interface: the send loop iterates, it does not
recurse. What does scale with the interface count is the broadcast
fan-out — an announce emits one action per entry in the routing map
(interface_names, leviculum-core/src/transport.rs:12342), each
carrying a cloned packet. With three BLE children an announce would
allocate three ~500 B action buffers where today it allocates one that
tx_fanout_task clones per link (leviculum-nrf/src/ble/mod.rs:489).
Same peak, moved one layer up.
Which machinery is per medium, and which is per link
This is what decides whether a child can stand on its own or needs a parent to lean on. Measured against the tree, not assumed:
| Machinery | Per | Where it belongs |
|---|---|---|
Airtime credit bucket (AirtimeCredit, leviculum-std/src/interfaces/airtime.rs:23) | medium — one radio, one duty cycle | parent |
Pre-TX jitter / CSMA deference (compute_jitter_max_ms, leviculum-std/src/interfaces/rnode.rs:297) | medium — contention is on the air | parent |
Announce cap and egress slot (interface_announce_caps, leviculum-core/src/transport.rs:2761; interface_next_slot_ms, leviculum-core/src/transport.rs:3019) | medium — it rations a shared resource | parent (splitting it per link multiplies the budget by the link count) |
Max-airtime backchannel (interface_max_airtime_ms, leviculum-core/src/transport.rs:3045) | medium | parent |
Advertising and scanning (reconcile_advertising, leviculum-std/src/interfaces/ble/mod.rs:958; ScanScheduler, leviculum-std/src/interfaces/ble/links.rs:1426) | medium — one adapter | parent |
| IFAC | medium — it is a property of the configured section | parent |
BLE inter-packet gap (LinkPacer, leviculum-std/src/interfaces/ble/links.rs:1540) | link, except on the shared notify pipe where one pacer serves every subscriber (leviculum-std/src/interfaces/ble/mod.rs:369) | child, mostly |
| Negotiated MTU and fragmentation state | link | child |
| Keepalive and expiry timers | link | child |
| Byte counters | link | child |
Six of the ten rows are per medium, and the four that are not are the
small ones. That is the finding: on a shared-carrier
medium almost everything that makes an interface an interface is per
medium. A child would own an MTU, a pacer, two timers and two counters,
and would have to reach the parent for everything else — which is
precisely the shape ble-reticulum's child ended up in.
The options
A — children everywhere. BLE spawns an interface per link on both stacks, matching TCP, AutoInterface, I2P and the shared instance.
B — children on lnsd, peer hint on the board. The daemon can afford interface objects; the firmware keeps three compile-time ids.
C — peer hint everywhere. One interface per medium; the transport passes "for this peer"; the interface maps peer to link. This is what Codeberg #365 and #376 built, and what runs today.
D — peer hint everywhere, per-peer rows in the reporting
inventory. C, plus the one thing A gives away for free: the
inventory already models a parent with spawned children and merges a
departed child's bytes into its parent (add_spawned,
leviculum-std/src/interfaces/inventory.rs:184; remove_spawned,
leviculum-std/src/interfaces/inventory.rs:190), and it is
driver-owned and deliberately outside the routing map. A BLE peer
appearing and disappearing already crosses the driver boundary as a
peer event, so it can create and retire an inventory row without ever
becoming a send target.
| A | B | C | D | |
|---|---|---|---|---|
| Correct addressing on a central-role link | yes | yes | yes | yes |
| Correct addressing on a BlueZ peripheral link | no — one notify characteristic, every subscriber gets it | no | no, and says so | no, and says so |
| Correct addressing on a SoftDevice peripheral link | yes | yes | yes (per-slot queues) | yes |
| Firmware heap, 3 children | +2.5 to 4.0 KiB | 0 | 0 | 0 |
| Per-medium policies | must be hoisted to a parent object or duplicated per link | hoisted on one stack only | untouched | untouched |
| Per-peer statistics | free | on lnsd only | absent today | yes, in the inventory |
| Adding a new multi-peer medium | write a parent + a child + the hoisting | pick a stack, then both | implement plan_tx_to | implement plan_tx_to, emit peer events |
| Model count for one protocol | 1 | 2 | 1 | 1 |
A peer reachable on two links at once. Under A there are two child
interfaces and therefore two path entries with different interface
indices; the transport picks by hop count and the loser is a live
standby, and losing one link culls only its own paths. That is the
cleanest behaviour of the four, and it is a real scenario: a rotated-
address reconnect holds two links to one identity for a moment. Under
C and D there is one interface and one path entry; the planner takes
the first link it finds for that identity
(plan_tx_to, leviculum-std/src/interfaces/ble/links.rs:898), and a
peer loss is reported only when the last link for that identity dies
(knows_identity, leviculum-std/src/interfaces/ble/links.rs:503).
The observable difference is which of two equally good links carries
the next packet — the transport cannot express a preference it has no
information to form. Under B, whichever of the two the stack in
question runs.
What an address rotation looks like, from each side. The observed
trigger is a Columba phone rotating its resolvable private address
about every ~90 s (review notes columba-befunde.md §5; the
2026-09-12 field capture on #360 shows the same ~90 s cycle). The
phone abandons its old connection when it rotates — from our side that
link simply goes silent mid-keepalive-interval — and reappears under
an address no table can associate with it, because an advertisement
carries no identity. The two sides of the same event:
- Seen from the board or lnsd: the old link's inbound frames stop;
seconds later the same 16-byte identity handshakes on a NEW
connection — either because the phone dialled us, or because our own
scanner dialled the unrecognizable new address. Which link survives
is
judge_duplicate(leviculum-nrf/ble-tx/src/registry.rs, shared by lnsd): an old link its peer has stopped keepaliving forLINK_ABANDONED_MS(30 s) loses to the newcomer (BLE_LINK_REPLACED … rule=abandoned); otherwise the decision is the PEER's own,preferred_ble_role— a port of Columba'spreferredBleRole— and we keep whichever connection it keeps. The losing link is disconnected by us at the decision, in either role. - Seen from the phone: it runs the same function on the same four inputs, which is the whole point: a rule both ends compute identically is the only one that cannot leave the pair linkless. The brief two-links-one-peer window above is the hand-over moment.
Before #360 round 1 the board refused every outgoing-origin duplicate, which kept the abandoned link and left the rotating phone linkless for the 45 s expiry of every ~90 s cycle. Round 1 then read the old link's PAYLOAD recency, which displaced links the phone was still keepaliving and produced the mirror-image hole on 2026-09-12 21:41 — each side closing the link the other had kept. Round 2 removed our own opinion from the general case and copies the peer's instead; see Bluetooth interfaces for the branches.
The recommendation
D: keep the peer hint on both stacks, and recover per-peer visibility
in the reporting inventory rather than in the routing map. The
deciding argument is not the RAM — 2.5 KiB of a 33 KiB worst-case
margin is affordable, so option A is not blocked by the firmware and we
should stop saying it is. It is that a child interface on a shared
carrier is not an interface: six of the ten mechanisms above are
per medium, so every child would delegate straight back to a parent,
and the reference implementation of exactly this idea on exactly this
carrier ended up as a shim holding an address and two counters, with a
peripheral-side central_address that only picks a counter. We would
pay an object, a registration, a teardown path and a second addressing
concept to buy an addressing capability the carrier does not have. The
hint costs one Option<[u8; 16]> on an action we already emit, works
identically on both stacks, and is honest about the peripheral-side
limit instead of papering over it. What A genuinely buys — a row per
peer in rnstatus — is a reporting concern, and the inventory is
already the place where reporting rows live without being send targets.
Option B is rejected outright. The two stacks already drift under one
shared model — the failed-dial scan phase above is this month's
example — and giving them different structures as well buys nothing
the numbers ask for: the firmware is not the constrained party here,
which was B's whole premise.
What would change this, and what to measure again
The recommendation is a reading of today's mechanisms, not a permanent verdict. Three triggers, each with the measurement that settles it:
- A carrier arrives where per-link addressing is real on both roles and per-link policy is genuine — per-link airtime, per-link congestion control, per-link IFAC. Then a child owns something and A wins on its merits. Measure: how many of the ten rows above move from "medium" to "link" for that carrier.
- The transport gains a reason to prefer one of two links to the same peer — a per-link quality or cost signal. Today it has none, which is why C's "first link found" is not a shortcoming. Measure: whether a per-link metric changes the chosen route on the rig at all.
rnstatusparity againstrnsdneeds per-peer BLE rows. That is D's second half, and it is an inventory change, not an architecture change. Measure: theinterface_statsrow set from both daemons on the same topology, which the drop-in property makes a single-driver comparison.
What would not change it is a future firmware with more RAM. The argument above is about where the mechanisms live, and that is the same on a board with 96 KiB of heap and on a server with 96 GB.
One correction this page owes
Bluetooth interfaces records an earlier
"Stage B" plan to model ble-reticulum as "one ordinary byte only
interface per connected peer (the Columba BLEPeerInterface shape)".
That is option A, written before #365 and #376 built the peer hint and
before anyone read what the Columba child actually does on the
peripheral side. This page supersedes that paragraph; the surrounding
argument there — that the broadcast plane and the link plane are two
interfaces and not one clever hybrid — is unaffected and still holds.
The Core Lock Budget
The async driver owns exactly one NodeCore behind one mutex
(leviculum-std/src/driver/mod.rs:1321). Every packet the node decrypts,
routes, forwards or emits passes through it. It is the narrowest point
in the stack, and the rule that follows from that is:
No caller holds the core lock across CPU-heavy work. Work that scales with payload size runs off the lock, between a cheap capture and a cheap commit.
This is not a style preference. It was measured.
The measurement that set the rule
NodeCore::send_resource used to run its whole build — bz2 compress,
bulk token encrypt, full/map hashing — inside the mutex. For a 1 MiB
incompressible payload with compression on, that was a 141 ms hold
in release. On a 20-link inbound flood it cost a 32 % inbound
throughput stall while round-sized sends ran; ~0 % after the fix
(Codeberg #152).
The fix is the three-phase shape, and it is the house idiom for anything with the same profile:
| phase | lock | code |
|---|---|---|
NodeCore::resource_send_params | brief | leviculum-core/src/node/mod.rs:1417 |
resource::prepare_resource_send | none | leviculum-core/src/resource/outgoing.rs:89 |
NodeCore::commit_resource_send | brief | leviculum-core/src/node/mod.rs:1452 |
Commit re-validates what could have changed while the build ran
unlocked: link gone, a transfer raced in, or the link re-keyed (#66) —
the last returns the retryable ResourceError::LinkStateChanged and
the caller rebuilds once. The std driver calls the three phases itself
(leviculum-std/src/driver/mod.rs:3602).
NodeCore::send_resource still exists as the composed single call
(leviculum-core/src/node/mod.rs:1543) because no_std and FFI callers
have no lock to hold and no second thread to starve. It is the
composed form that is dangerous behind the driver, not the code it
composes.
The numbers, restated for callers
Measured with leviculum-core/compression on (as leviculum-std
builds it), comparing the composed call against the locked portion of
the phased path.
The conditions, because a timing without them is an anecdote: profile
release, target x86_64-unknown-linux-musl, an AMD Ryzen 9 7950X.
Each cell is the median of five timed runs after one discarded
warm-up, and each run builds a freshly linked pair — a second
Resource to the same link cannot build at all
(ResourceError::TransferInProgress), so re-using one would not repeat
the measurement. Both harnesses are #[ignore]d and print every number
below:
measure_send_lock_costs (leviculum-lxmf/src/node.rs:2607) for the
send tables, measure_deferred_tick_costs
(leviculum-lxmf/tests/direct_delivery_attempts.rs:1560) for the tick
table.
Every column names the bytes it was given, because the cost being
reported is a compressor's and a compressor's cost is a property of its
input. The three classes are generated by
incompressible (leviculum-lxmf/tests/common/payloads.rs:32),
compressible (leviculum-lxmf/tests/common/payloads.rs:79) and
degenerate (leviculum-lxmf/tests/common/payloads.rs:104), and both
harnesses below take them from the same array so neither can drift onto
different bytes under the same name:
| class | what it is | bz2 at 1 MiB |
|---|---|---|
incompressible | seeded xorshift bytes | 1.00x — grows 0.5 % |
compressible | dictionary words with sentence breaks | 9.0x |
degenerate | vec![0x5a; n] | 21 845x |
| payload | incompressible | compressible | degenerate | phased (locked) |
|---|---|---|---|---|
| 16 KiB | 1.67 ms | 0.91 ms | 0.14 ms | 2 µs |
| 256 KiB | 15.8 ms | 11.0 ms | 1.70 ms | 3 µs |
| 1 MiB (segment 1) | 64.6 ms | 48.7 ms | 7.61 ms | 2–11 µs |
The first three columns are the composed call under the lock; the last is the locked half of the phased path, which does not vary with the payload because no payload passes through it.
Incompressible data is the worse case, not compressible data. An earlier revision of this page said the opposite — "bz2 does more work when it succeeds" — and the measurement above refuses it at every size: compressible text costs 0.54x to 0.75x of incompressible bytes of the same length. The mechanism is that the Burrows-Wheeler sort is paid in full either way, and a run that succeeds then has less output left to code, not more. Plan for the incompressible number: it is both the larger one and the one an attachment actually hits, since anything already compressed looks incompressible to bz2.
Without the compression feature — the embedded default — the 1 MiB
build drops to 6.4 ms, which is why the same code is tolerable on an
nRF52 and intolerable behind the driver. That figure is from the #152
pass and was not re-measured here.
The segment boundary sits under the 1 MiB row
RESOURCE_MAX_EFFICIENT_SIZE is 1 048 575 bytes
(leviculum-core/src/resource/mod.rs:69) and the split is decided on
the packed, uncompressed length
(leviculum-core/src/resource/outgoing.rs:106), so a 1 MiB body is
above it in every payload class — the compression ratio does not move
the boundary. The 1 MiB row therefore measures segment 1 of a
two-segment transfer, not a whole one. Segments 2..N are built on the
receive path, under the caller's lock, and the phased path does not
cover them.
That also bounds who can reach the row at all. Python LXMF refuses an
incoming delivery Resource larger than DELIVERY_LIMIT × 1000, which
is 1 000 000 bytes
(delivery_resource_advertised, reference/LXMF/LXMF/LXMRouter.py:1977)
— below the segment boundary. So no LXMF transfer that a Python peer
would accept is ever a two-segment one, and the segment path is reached
only by a Rust-to-Rust transfer or a non-LXMF core Resource user.
What the payload class was worth, as a number
This matters beyond bookkeeping, because a table measured on
degenerate was once proposed as a correction to this page. On the
same machine and profile, degenerate reads 8.5x below
incompressible at 1 MiB and 12x below it at 16 KiB: bz2's
run-length front end collapses a single repeated byte before the
Burrows-Wheeler transform ever runs, so the number that comes out is
the cost of compressing almost nothing.
Two other candidate explanations for a table reading low were tested and do not hold.
A missing warm-up is not one. The harness prints the run it discards
next to the median it keeps, so this is checkable rather than assumed:
every cell above 1 ms has its cold run within 2 % of its median, and
the largest gap anywhere is 11 % — on degenerate at 16 KiB, the
cheapest cell in the table at 154 µs cold against 139 µs. An n=1
harness on this path is imprecise; it is not biased low.
A build-profile difference is not one either. This run reproduces the
figures the page carried before it — 1.7, 16.4 and 65.2 ms — to within
4 % on the incompressible column, so those were release-profile
numbers taken on comparable hardware, and the column they belong to is
incompressible.
Costs that do not justify phasing, measured the same way: packing a 1 MiB LXMF message is 0.8 ms, and unpacking one with signature verification is 3.2 ms. Inbound verification cannot be phased away — the bytes are already in hand — and at that magnitude it does not need to be.
The adapter-side number: LxmfRouter::tick
Measured during the #196 design pass, on the same machine and profile:
one LxmfRouter::tick with 8 due 256 KiB messages holds 126.6 ms in
a single uninterrupted borrow. It is the composed-send_resource cost
of the table above, multiplied by a queue depth an adapter reaches
routinely — the router builds each due message in turn, and nothing
between them yields.
That figure is consistent with the table above: eight times the
incompressible 256 KiB build is 127 ms.
LxmfRouter can hand the build out instead, under
RouterConfig::defer_resource_builds
(leviculum-lxmf/src/router.rs:124). What the tick then costs, for one
due message, measured the same way:
| payload | deferred tick | composed tick (incompressible) |
|---|---|---|
| 16 KiB | 13.6 µs | 1.63 ms |
| 256 KiB | 300 µs | 15.9 ms |
| 1 MiB | 1.11 ms | 66.6 ms |
The deferred column is flat across all three payload classes — 13.6 to
14.6 µs at 16 KiB, 282 to 300 µs at 256 KiB — which is the mechanism
showing through: a deferring tick copies bytes and does not compress
them, so the class it was handed cannot matter. What is left in it is
the Message clone, and that is why the deferred column still grows
with size at all.
One due message rather than the eight above, because a second Resource to the same link is refused before it builds, so an eight-message comparison would not be comparing like with like. The composed column carries the router's own tick work on top of the build, which is why it reads a little above the send table at the same size; the gap is under 3 % at every row.
It is recorded here rather than in the issue because this page is where
a caller looks for it, and because a number that lives only in a tracker
cannot be cited from the tree: PROCESSOR_TICK_BUDGET
(leviculum-std/src/driver/processor.rs:181) is set against this
measurement, and until it was written down the only number behind a
public constant could not be traced at all.
What this binds
Any protocol adapter layered on the core. leviculum-lxmf is the
current instance and leviculum-lxst will be the next. An adapter that
offers only a monolithic submit call forces its host either to hold the
lock for the build or to fork the driver. Adapters that are expected to
run behind the async driver expose the phase split; the composed form
stays for the embedded caller.
Anything the driver runs inside its event loop. The loop's
dispatch_output (leviculum-std/src/driver/mod.rs:5247) routes
actions to interfaces and forwards events. Work done there blocks not
just the lock but interface I/O dispatch — strictly worse than the
mutex case. The in-loop /status responder
(leviculum-std/src/driver/remote_mgmt.rs:82) is the reference for how
much is acceptable there: take the lock, build a small bundle, hand
back a TickOutput, return.
Two things the loop's callees may never do. They may not .await,
and they may not call back into the driver's public async API: those
methods end in action_dispatch_tx.send(output).await on a bounded
channel that the same loop drains, so a full channel deadlocks the
node.
Those are one rule, not two, and knowing which way round matters when
you have to enforce it. The second is a consequence of the first: an
async fn called and not awaited builds a future and drops it, sends
nothing and blocks nothing. The deadlock needs the bounded-channel send
to complete, and only .await can complete it. So a callee expressed
as a synchronous fn has both prohibitions closed at once, which is
what the in-driver core processor (#196) is built on — see
leviculum-std/src/driver/processor.rs. It follows that a runtime
guard on the async API would be the wrong shape: there is nothing to
guard until an .await that cannot be written.
The residue is re-entrancy, not the async API
Both prohibitions above are special cases of a plainer one, and stating
them first got the emphasis wrong for two commits. The loop calls its
callees with the core mutex held, and that mutex is a non-reentrant
std::sync::Mutex. Any path from a callee back to it hangs the node
immediately — first call, no load required.
The async route is one such path and not the instructive one. Consider
the block_on case the previous wording named as the whole residue: a
callee that smuggles a PacketSender and blocks on send does not
reach the bounded channel at all, because PacketSender::send
(leviculum-std/src/driver/sender.rs:92-107) takes the core lock in a
block and releases it before its .await. It deadlocks one line
earlier, on the mutex.
And no block_on is needed. The number is 58, and it is a number
rather than a phrase on purpose: scripts/check-core-lock-census.py
rebuilds the list of public methods that lock the core out of the
sources on every just fast and pins it, name by name, in a
checked-in file (TOTAL, scripts/core-lock-census.txt:30). Two
earlier revisions of this page estimated the size of this set in words
and were low by nearly half — which is what an estimate nothing can
check is worth.
53 of them are on ReticulumNode — plain synchronous pub fns that
open by locking the core, of which has_path
(leviculum-std/src/driver/mod.rs:3085) is
self.inner.lock_recover().has_path(dest_hash) and entirely typical.
The other five are on PacketSender and LinkHandle, which matters
more than the count suggests: those are the two handles a callee is
most likely to have been handed in the first place.
A callee holding an Arc<ReticulumNode> deadlocks the node on its first
invocation, in ordinary safe synchronous code, with no .await, no
channel, and nothing a compile-fail fixture can catch.
So the rule for anything the loop calls is: hold no handle to the node
you run inside. The &mut StdNodeCore the seam hands out locks
nothing and is the whole intended surface. Everything else belongs on
the far side of a channel.
What holds this up is not the type system. It is that registration
happens on the builder, before the node exists, so a callee cannot be
constructed holding a node handle — injecting one afterwards takes a
deliberate OnceLock or Weak and a reference cycle. That is a
construction-order barrier, and it is why the hazard stays theoretical
in practice. It is not a guarantee, and this page should not be read as
offering one.
The one call the seam hands out that this page forbids
NodeCore::send_resource is pub
(leviculum-core/src/node/mod.rs:1747) and therefore reachable on the
&mut StdNodeCore a processor hook holds. It is the 141 ms composed
call this page opens with — one line, in consumer code, behind the
driver and under the lock. PROCESSOR_TICK_BUDGET reports it 141 ms
after the fact and cannot prevent it, and no fixture can refuse it: it
compiles, because for the no_std and FFI callers it is the correct API.
A hook that has to send a resource uses the three-phase form the driver
itself uses — resource_send_params, prepare_resource_send off the
lock, commit_resource_send — or hands the send to the application side
of a channel. This is named again in
leviculum-std/src/driver/processor.rs where a consumer will meet it.
Work that is already off the core by construction
Proof-of-work stamp generation borrows only its executor, never the
router and never the core
(leviculum-lxmf/src/router/stamp_runtime.rs:26). The router emits a
pending-stamp event, the application computes the stamp on whatever
schedule it likes, and hands the result back. This matters because a
peer chooses the stamp cost: an announced cost of 254 is legal and
effectively unfinishable (#185). If that search could ever run under
the core lock, any peer could stop the node by announcing a number.
It cannot, and no seam added later may make it possible.
That pattern — emit a request, compute detached, submit the result — is the general answer whenever the cost of a step is not ours to bound.
One correction to the paragraph above, because its phrasing is wider
than what holds. It is exactly true of the peer-priced search, which
is the one that matters: generate_with is an async fn, so no
synchronous callee of the event loop can drive it to completion at all.
It is not true that the tree contains no synchronous proof-of-work.
leviculum-core::discovery::stamp::generate_stamp
(leviculum-core/src/discovery/stamp.rs:176) is a public synchronous
brute-force loop taking a caller-supplied cost, and nothing stops a
loop callee from calling it. It is not a DoS vector today because no
peer picks its number: the only caller is the discovery announcer,
which runs it once per discoverable interface during
ReticulumNode::start() — off the loop — at the locally fixed
DEFAULT_STAMP_VALUE. The invariant to keep is therefore "no
peer-chosen cost is ever ground synchronously", and the async
signature is what enforces it.
That fixed cost is not static across RNS versions: #328 raised it from 14 to 16 to stay visible to RNS 1.5.0 listeners, and each extra bit doubles the search. Measured on the coder host (release build, x86-64), one mint went from 39 ms mean / 198 ms max to 191 ms mean / 737 ms max over 24 samples. It is paid once per discoverable interface at wiring time and the result is reused for every re-announce, so this is startup latency, not a per-announce or per-loop cost. The budget argument is unchanged; the number it is measured against is four to five times larger.
A diagnostic write is inside the budget too (#418)
The budget above is about CPU: work whose cost scales with a payload. The miauhaus soak found the other half, and it is worse, because nothing about the call site looks expensive.
Over 49 days and 397 023 881 events, the node's 10 s PATH_TABLE
liveness heartbeat missed at least one beat 2 928 times out of
421 605 intervals, with a tail to 37.0 s. During each of those the
daemon emitted nothing at all — no packet, no announce, not the
heartbeat — on a node that otherwise logs 30 to 150 events a second.
Both long stalls pulled out of the raw log have the same shape: the
hole sits between the ANN_RX of one announce and the PATH_ADD of
that same destination, a span in which nothing can take seconds.
The emission can. Until #418 the event-log layer wrote each line with
a blocking write(2), flushed, under a process-global mutex, on the
thread that emitted it — and the event loop emits while it holds the
core mutex (apply_inbound,
leviculum-std/src/driver/mod.rs:4456). A write(2) to a USB disk
under writeback throttling blocks for seconds, so the loop stopped,
and everything that wanted the core queued behind it. That is why the
symptom was total silence rather than a missing log line.
The rule that follows is the CPU rule's sibling:
No caller holds the core lock across an I/O call whose latency belongs to a device. A diagnostic that can stop the transport is a worse bug than the missing diagnostic.
The event log now inverts the trade (FileSink,
leviculum-std/src/event_log.rs:1332): the emitting thread does a
bounded enqueue and returns, one writer thread owns the file, and an
overrun drops lines and says so with EVENT_LOG_DROPPED rather than
blocking the mesh. LEVICULUM_EVENT_LOG_SYNC=1 restores the old
behaviour for anyone who would rather block than lose a line, and it
is what the mvr's positive-control arm runs.
What measures this
Three events, all threshold-gated so a healthy node emits none of them, and deliberately at different altitudes so they disagree informatively:
| event | where | says |
|---|---|---|
ANN_SLOW | handle_announce (leviculum-core/src/transport.rs:6104) | announce handling itself took ≥ 100 ms |
CORE_STALL | spawn_core_stall_watchdog (leviculum-std/src/driver/mod.rs:4247) | an outside thread waited ≥ 250 ms for the core lock |
EVENT_LOG_WRITE_SLOW | writer_loop (leviculum-std/src/event_log.rs:1426) | one batch write to the log file took ≥ 50 ms |
CORE_STALL without ANN_SLOW means the loop was stopped by
something other than announce handling; EVENT_LOG_WRITE_SLOW
alongside it names the disk. The watchdog measures the wait for the
core mutex, never the hold, which is exactly "how long the loop spent
not polling" without an Instant in a dozen select! arms.
A durable store append is inside the budget too (#384)
The #418 rule above is about a diagnostic, and a diagnostic can be dropped. The propagation node found the same shape in work that cannot be: storing a message somebody sent us.
lnpnd ran for 32 minutes in the public mesh on miauhaus (2026-09-21) and
emitted 73 CORE_PROCESSOR_OVER_BUDGET warnings out of 324 log lines,
54 of them under hook="on_event" and 19 under hook="on_tick", with nine
over 100 ms. lnsd on the same host, same uptime, same traffic: none. The
worst one names its own cause in the line above it:
10:25:12 lnpnd: sync in peer 535d9c5db65bfcd4 transferred 105 (30240 B): ok
10:25:12 CORE_PROCESSOR_OVER_BUDGET hook="on_tick" elapsed_us=239016 budget_us=5000 events=0
A quarter of a second of core lock, having processed events=0. The time
was not in event work; it was in the store. FilePropagationStore::append
is a write, an fsync, a rename and a second fsync
(leviculum-std/src/file_propagation_store.rs, "Power-cut safety, and why
this store fsyncs"), and the engine stored the whole inbound batch inside
the one hook that classified it.
The attribution is a measurement, not a reading of the code. 105 appends of
288-byte bodies cost 124 µs into a memory store and 127-189 ms into
the file store on the coder host's ext4, over six runs — 1.1 to 1.5 ms each,
worst single append 6.4 ms. Three orders of magnitude, on the same verb with
the same bytes: the cost is the device, not the book-keeping. append_cost
(leviculum-std/src/file_propagation_store.rs) prints both lines on
whatever filesystem it is pointed at, and pointing it at a tmpfs — where
fsync never reaches a device — reads 60x cheaper and proves nothing.
Two things follow, and only one of them is a fix.
The store cannot be the layer that fixes it. The fsyncs are what makes
"persist before you prove" true
(docs/src/concepts/propagation-node-on-a-board.md §3): dropping them
trades a message for a millisecond. The cost is the device's, and the
device's cost is not negotiable from this side.
The hook is. lnpnd's engine now queues validated payloads and stores
PERSIST_PER_HOOK of them per hook (lnpnd/src/engine.rs), asking the
driver to come straight back for the rest; a peer's batch is resumable
across hooks (PeeringRuntime::advance_sync_resource, lnpnd/src/peering.rs)
and still reports itself as one round. The batch still costs what it costs
in wall-clock — the disk did not get faster — but the core lock is released
between messages, so the node keeps answering while it absorbs a sync.
A cost that cannot be made cheap and cannot be dropped is still a cost that may not be paid all at once under the lock. Bound the work per hook and come back.
What this does not buy: one append can exceed the budget on its own —
6.4 ms measured, against a 5 ms budget — so CORE_PROCESSOR_OVER_BUDGET
can still fire on a slow disk. What is gone is the multiplier, which is the
part that scaled with someone else's batch size.
A diagnostic's volume is a cost too, and the cadence is the wrong lever
The same log gave the other half of the lesson. Of those 397 023 881
events, 214 594 631 — 54 % — are PATH_TABLE_ENTRY, the per-path
snapshot of the path table; in the last megabyte of the live tail it is
77 %. The file is 61 GiB and grows ~2.6 GB a day.
That share is what remains after a fix. The dump used to fire every
10 s; 5ab938a4 (2026-07-13) moved it to every five minutes and cut
its volume thirtyfold. It did not settle the problem, because a
cadence is the wrong lever for this cost:
One snapshot costs one line per path. The miauhaus path table is 22 362 entries — the number is in the heartbeat itself,
PATH_TABLE node=miauhaus size=22362. Five minutes apart, that is still ~6.4 million lines a day, and it grows with the mesh, not with anything the code chose.
Against that cost, nothing consumes the lines. The soak's analyze.py
puts them on a count-only fast path and never reads dst, hops,
iface, next_hop or expires_in_ms; periculum does not reference
the event at all. And the one number counting them yields — how many
paths there are — is already emitted every 10 s as PATH_TABLE size=.
So the dump is now off by default
(TransportConfig::path_entries_dump, config key path_entries_dump
under [reticulum]), reachable for a debugging session that wants the
per-path fields. Nothing else moved: the 10 s PATH_TABLE heartbeat
still carries liveness and the count, and PATH_ADD still records
every insertion, so the path table's history survives the silence.
The rule, sibling to the one above:
A diagnostic whose volume scales with mesh state, not with a rate the code picks, is not a diagnostic you can leave on. Gate it, and keep the cheap scalar that answers the question people actually ask of it.
The Self-Deadlock Tripwire
A deadlock is the worst failure this daemon produces. A panic leaves a stack and an exit code; a dropped packet leaves a counter; a wrong answer leaves a log line somebody can grep. A self-deadlock leaves nothing at all. The node stops, keeps its listening socket open, and reads as healthy to every supervisor watching it.
This page records what detects one, what it costs, which decisions were made against project precedent, and what it does not reach.
The hazard
93ba351 shipped [CoreProcessor], a seam that runs consumer code inside
the driver's tick. Both entry points — run_event_tap and the timer branch's
run_tick — call the consumer's hook with the core std::sync::Mutex guard
live, because handing the hook &mut StdNodeCore is the entire point of the
seam.
std::sync::Mutex is not reentrant. So any handle the consumer smuggled in
that re-locks the core parks the driver's event loop on a lock only that same
loop can release. ReticulumNode::has_path
(leviculum-std/src/driver/mod.rs:3084-3086) does it in one line, and it is
one of roughly forty synchronous pub fns on ReticulumNode shaped exactly
like it. No .await, no unsafe, no channel — nothing a compiler or a
compile-fail fixture can see.
The seam's own defence is a construction-order barrier, not a guarantee: a
processor is registered on the builder, before the node exists, so it cannot
be built holding a handle to the node it will run inside. Getting one takes a
deliberate OnceLock/Weak cycle. That is why the hazard is theoretical
rather than routine — and it is exactly what a future set_core_processor on
a live node would give up.
Two things that were tried on paper and are not the answer
try_lock that reports instead of blocking is strictly worse. From the
same thread try_lock returns WouldBlock, so has_path would answer
false for a destination that has a path. A loud deadlock becomes a silent
wrong answer inside routing, which is the worse failure for Priority 1. It
would also mean turning roughly forty public accessors fallible, a breaking
change for every consumer, to defend against a hazard reachable only from
inside the seam.
Type-system prevention does not exist. The processor is
Box<dyn CoreProcessor> and its struct is opaque by construction. Every bound
available ('static, Send) is already applied and none can express "holds
no Arc<ReticulumNode>".
Detection is the remaining avenue.
The mechanism
MutexRecover::lock_recover (leviculum-std/src/sync_ext.rs) is the single
choke point: every std::sync::Mutex acquisition in leviculum-std goes
through it, 149 call sites against one implementation.
Before it blocks, it records the mutex's address in a thread-local set. An
address already present means this thread is about to wait for a lock only
this thread can release, which never wakes. The guard it returns
(TrackedGuard) removes the address on Drop.
Three properties are worth naming because each is a defect if it is wrong:
- Removal is by address, not by position. Guards are values and nothing forces them to drop in acquisition order. Popping the top would evict the wrong entry and leave a stale address behind — a false positive on the next acquisition of a mutex that was released long ago.
- The set is per thread. Another thread holding the mutex is an ordinary
wait, not a self-deadlock.
TrackedGuardis!Sendfor the same reason aMutexGuardis, which is also what makes the thread-local sound: a registration can never be read from a thread other than the one that made it. - An unwind must clean up. A hook that panics for an unrelated reason unwinds through the guard, and if the address survived that, the next acquisition on that thread would report a re-entry that is not one. The tripwire's worst failure mode is a false positive in production, so this has its own test.
Depth is capped at 32 held mutexes per thread — the crate's deepest measured
nesting is 4 — and exceeding it is reported as LOCK_DEPTH_OVERFLOW rather
than absorbed. Past that point a negative answer means nothing, and a check
that has quietly stopped checking is the failure mode
Checks That Are Actually Checks exists to remove.
The measurement that chose the shape
Two shapes were on the table. The narrow one — a static CORE_LOCK_OWNER: AtomicU64 written by the tap and the timer branch, read in lock_recover
under #[cfg(debug_assertions)] — is scoped to the one lock and the two
callers we currently suspect, and does nothing in release, which is where the
deadlock actually bites. The broad one is the thread-local address set above,
which covers every mutex and every caller and needs nobody to have guessed
right.
The question that settles it is what the broad shape costs on a real workload.
Per acquisition (release, opt-level=3, uncontended mutex, 20 M
iterations, best of 5, four-core host):
| shape | ns/acquisition | delta |
|---|---|---|
| no tripwire (previous behaviour) | 3.024 | — |
| narrow: one relaxed atomic load | 3.021 | +0.00 |
| broad: thread-local set, depth 1 | 3.482 | +0.458 |
| broad: thread-local set, depth 4 | 3.867 | +0.843 |
In an unoptimised build the same three are 24.6 / 26.5 / 67.1 ns, because nothing inlines; that is a test-run cost, not a shipped one.
Acquisition rate on a real workload. The TCP-hub load test
(leviculum-std/tests/rnsd_interop/loadtest_tcp_hub_tests.rs) driving the
real lnsd binary at 128 steady connections, one packet per connection every
15 ms for 40 s, 377,572 packets forwarded at 100 % delivery: 2,061,220
acquisitions in 42.56 s = 48,428/s, or 5.46 acquisitions per forwarded
packet.
So the prediction is 2,061,220 × 0.458 ns = 0.94 ms of CPU across the whole 42.6 s run — 0.0055 % of the hub's 17.1 s of CPU time, and 2.5 ns against the 45.4 µs of CPU the hub spends per forwarded packet, about one part in eighteen thousand.
And the end-to-end A/B agrees, which is the point of doing both. Two
release lnsd binaries differing only in the tripwire, alternated over the
same load:
| hub CPU per run | mean | |
|---|---|---|
| no tripwire | 18.07, 17.54, 15.83 s | 17.147 s |
| tripwire | 17.33, 16.61, 17.72, 16.90 s | 17.140 s |
Delta −0.04 %, against a run-to-run spread of ±6 % on the same binary. Delivery was 100.0000 % on every run of both. A null A/B on its own would only say the effect is below the noise floor; paired with the prediction it says why — the effect is three orders of magnitude below it, and no amount of extra runs would resolve it.
It is noise. The broad shape wins, on the criterion set before the numbers were taken.
Panic, not log-and-hang — and in release too
Both decisions follow from the same observation, and neither inherits from the
poison precedent above it in sync_ext.rs.
Why not log-and-hang. Project policy prefers a degraded daemon to a
crashing one, and MutexRecover's poison recovery argues exactly that. But
that argument does not transfer, because there the daemon really is degraded
and still serving: the panicking task has unwound, every other task still
runs, and the node keeps forwarding. Here the thread about to block is
normally the driver's event loop, and once it stops the node forwards nothing,
answers no path request and maintains no link — while holding its listening
socket open. That is not degraded operation. It is a stopped node wearing a
healthy face, and Priority 1 is packet delivery, of which this delivers zero.
Panicking is also, for the hazard this exists for, the cheapest possible
recovery — because the seam already catches it. run_event_tap and run_tick
wrap every hook in catch_unwind. The unwind releases the outer core guard
(poisoned, which lock_recover then recovers), the driver logs
CORE_PROCESSOR_PANICKED, detaches the offending processor permanently, emits
NodeEvent::CoreProcessorPanicked on the application's control plane, and
carries on serving the mesh without it. That is precisely the graceful
degradation the policy asks for, and log-and-hang forfeits it. Away from the
seam, a report here means an ordinary lock-order bug in our own code on a path
some test exercises, and failing at the defect beats hanging at it.
Why release carries it. Debug-only was the cheap answer and it does
nothing where the deadlock actually bites: a #[cfg(debug_assertions)] check
is absent from every binary an operator runs. The only argument for compiling
it out is cost, and cost is 0.0055 % of a busy hub's CPU. There is nothing to
trade.
The standing canaries
Per Checks That Are Actually Checks: a tripwire
that has silently stopped tripping satisfies "no deadlock detected" forever,
so the demonstration is permanent rather than one-time. Both halves live in
leviculum-std/src/sync_ext.rs and run in just fast
(cargo test --workspace --lib):
canary_re_entering_one_mutex_trips— a re-entry the detector is meant to see, asserting the report names both the failure and the mutex.canary_distinct_mutexes_nest_without_tripping— legitimately nested different mutexes, which must not fire. Without it the positive canary is satisfied by a tripwire that reports everything, which would take the daemon down on its first tick.
The acceptance test is leviculum-std/tests/mvr/core_lock_reentrancy.rs: a
registered CoreProcessor calling has_path on the node it runs inside, plus
a negative control that is the same node with a processor that does not
re-enter. That test hung forever before this landed, which is why it
carries two independent bounds — a tokio::time::timeout on a runtime the
node does not own, and Drop for ReticulumNode, which polls for at most 400 ms
and then calls Runtime::shutdown_background. Verified by disabling the
tripwire by hand: the test fails in 15 s with a message naming the deadlock,
and the binary exits normally.
What this does not reach
- Deadlocks between two threads and two mutexes. A holds M1 wanting M2 while B holds M2 wanting M1 is a lock-order inversion, and neither thread's own set contains what it is waiting for. Detecting that needs a wait-for graph across threads, which is a different mechanism and a different cost.
- Mutexes not taken through
lock_recover.tokio::sync::Mutex, anyRwLock, any.lock().unwrap()that skipped the trait, and every mutex outsideleviculum-std. The choke point is what makes this cheap and it is also its boundary. - Blocking that is not a mutex. A hook that blocks on a channel, a socket
or a
block_onstalls the loop just as completely and is invisible here. The core lock budget is the surface for that, and it reports after the fact rather than preventing. - Depth beyond 32 on one thread, which is reported and then untracked.
See also
- The core lock budget — what a hook may cost while it holds the lock, as opposed to whether it may re-take it.
- Checks That Are Actually Checks — the standing-canary rule this page applies.
Bluetooth interfaces
Reticulum uses Bluetooth in three distinct ways. They are not variants of one interface, they are three separate carrier protocols with different connection models, different peers, and different scaling behaviour. This chapter names them, maps them to what Python-RNS and the Columba app call the same things, and records the design decisions behind them.
The actionable status and open work for each lives on Codeberg, not here. This chapter is the durable concept; the tracker is the source of truth for what is done.
The three protocols at a glance
| Name | What it is | Connection model | The peer is | Scales |
|---|---|---|---|---|
| RNode over BLE | Drive a dumb RNode radio over a transparent BLE link | Connection oriented (GATT) | a radio, not a node | n/a |
ble-reticulum | A nearby device is a full Reticulum node, one link per peer | Connection oriented, one GATT link per peer | a Reticulum node | no, 3 to 4 reliable links |
ble-leviculum | Reticulum broadcasts ride BLE 5 extended advertising | Connectionless | many nodes in range | yes |
Naming and lineage
Two of these already exist in the wider ecosystem, so we adopt their names to keep wire compatibility obvious. The third is our own invention.
| Our protocol | Python-RNS calls it | Columba calls it |
|---|---|---|
| RNode over BLE | BLEConnection inside RNodeInterface, ble://, via bleak | BluetoothLeConnection, RNodeInterface[BLE] |
ble-reticulum | no direct equivalent (closest is the new WeaveInterface / WDCL, but different) | the ble-reticulum Python package: BLEInterface plus BLEPeerInterface, wire spec "Protocol v2.2" |
ble-leviculum | none, genuinely new | none |
ble-leviculum is a BLE 5 connectionless broadcast carrier for Reticulum
packets. It is leviculum originated and is not an upstream RNS standard. We
chose a name in our own namespace, not ble5-reticulum, on purpose: this
protocol interoperates with nobody yet, and the *-reticulum namespace is not
ours to reserve. The name itself marks it as ours, which sets it apart from the
two protocols above whose names we adopted because we must match their wire.
The no_std layering
Every Bluetooth interface splits into two layers, and the split follows the interface isolation rule (see Interface isolation).
- Carrier logic, no_std. Framing, fragmentation and reassembly, the
protocol state machine. This belongs in
leviculum-core, which is no_std and already builds forthumbv6m. BLE framing already lives inleviculum-core/src/framing/ble.rs. Keeping the carrier logic no_std means the same code runs on the nRF firmware and on the host. - Platform binding. The radio and OS specific glue. On nRF this is the
SoftDevice glue in
leviculum-nrf(no_std). Onlnsdthis is a Linux BLE stack, candidatebluerover BlueZ via DBus (necessarily std). On a phone it is the OS BLE API.
The goal is no_std carrier logic wherever possible so it runs on embedded
devices. One known exception: the RNode over BLE byte channel seam currently
lives in leviculum-std with tokio traits, so it is std only. A no_std variant
over embedded-io-async would be needed for on device use, tracked separately.
RNode over BLE
BLE is used purely as a cable. The far end is an RNode radio that speaks the RNode KISS protocol; leviculum still drives detection, configuration and the radio lifecycle. The peer is not a Reticulum node.
The enabling work is a generic byte channel seam: drive the RNode lifecycle over any duplex byte channel instead of a serial port path, so a process that never sees a serial device (Android USB host, BLE GATT, iOS BLE) can still run an RNode. Compatibility is unaffected, the serial path is unchanged and the wire format does not change.
ble-reticulum (BLE 4 link mesh)
Each nearby device is a full Reticulum node. A node opens one connection
oriented GATT link per peer, acting as both peripheral (GATT server) and
central (scan and connect). Columba implements this as the ble-reticulum
Python package with a BLEInterface for protocol handling and one
BLEPeerInterface per connected peer.
To interoperate with Columba we must match its wire spec, "Protocol v2.2": a
fixed service UUID 37145b00-442d-4a94-917f-8f42c5da28e3, RX and TX and
Identity characteristics, and the connection handshake. The Identity
characteristic carries a stable Reticulum transport identity hash so peers can
be tracked across the BLE MAC address rotation that phones perform for privacy.
The hard limit of this protocol is the number of simultaneous links. Columba
caps at MAX_CONNECTIONS = 7 and Android allows about 8 BLE connections total
across all apps; in practice 3 to 4 links are reliable. This protocol therefore
does not scale to a dense mesh, which is the motivation for ble-leviculum.
Incoming link slots and the slot-contention policy
An LNode accepts three incoming (peripheral-role) links and initiates
one outgoing (central-role) link (#372), and since #432 lnsd has the
same shape: max_connections (default 4) still means the total, and
inside it one slot is this node's own dial and max_connections - 1
are incoming. Before #432 the budget was undivided, and a node that
filled it in either direction both stopped advertising and stopped
dialling. That is how the rig's regression/ble_room_10 --seed 2 room
ended ABSORBING: nine boards formed eighteen links among themselves,
all nine went dark, and the tenth — holding the room's lowest address,
so the sort said it must dial everyone and nobody may dial it —
arrived 4 s late to a room with nothing on the air. No link ever died
to free a slot, and the #375 fallback verdict each of the nine held on
its advertisement died on their own full-table gate. The split makes
the two questions independent: may-dial reads the OUTGOING slot (an
lnsd with three incoming links still dials), on-air reads the INCOMING
slots (an lnsd that has spent its dial still advertises and still
accepts). Those same nine boards then land eight dials instead of
eighteen and every one of them is still advertising when the tenth
arrives. Admission does not widen for it: admit refuses a surplus
link by role, so nothing on the air is over-promised. Three incoming
slots dissolve the
single-slot field failure where two boards beside a phone paired with
each other first and the phone could only reach the board whose one
slot was still free: with slots to spare, a neighbour board and a
phone link to the same relay simultaneously.
The SoftDevice RAM cost of this configuration (conn_count = 4,
periph 3, central 1) is measured, not extrapolated: the S140 wants an
app RAM base of 0x20005DA0 (23 968 B), 928 B under the linked
ceiling in leviculum-nrf/memory.x (rig T114 SD_RAM_FLOOR,
2026-09-08). The margin is deliberately small — the requirement is a
fixed, deterministic property of this exact configuration and the
boot-time SD_RAM_FLOOR assert refuses any config that outgrows it.
There is deliberately NO preference of a phone over a board for the last free slot. With more slots than nearby peers the policy would decide nothing, and deciding it well needs information a connect-time policy does not have (which peer carries traffic the mesh needs). It becomes worth revisiting when a deployment has more adjacent boards than a relay has slots — the boards can then occupy every slot before a phone arrives, which is the single-slot failure again, one layer up. The firmware keeps advertising while any slot is free and goes silent when full, so a scanner not seeing the relay is the honest signal of that state.
Initiation direction, the fallback, cycles and duplicate links
Who initiates is the v2.2 address sort: the lower BLE address dials, the higher one advertises and waits (with the v0.3.0 capability override for peripheral-only peers). The sort alone does not reliably connect a room of boards (#375): a full board stops advertising, so the highest addresses can run out of permitted targets and sit scanning forever while the sort forbids them to dial anyone lower. A Monte Carlo of random arrival orders puts ten boards at a 21 % chance of a disconnected BLE graph under the pure sort.
The fallback closes the stranded-board gap: a central task that has
scanned for SCAN_FALLBACK_AFTER_MS (30 s) without a single initiate
verdict, and without a connection event in either role, accepts any
advertising Reticulum peer (BLE_SCAN_FALLBACK marks the switch, once
per strict phase, and the verdict logs as rule=initiate_fallback in
BLE_SCAN_DECISION). Full boards do not advertise, so a fallback dial
only ever lands on a free slot. The rule stays pure and host tested in
leviculum-nrf/ble-tx (should_initiate with a ScanMode input);
the firmware measures the time and hands the mode in.
The clock is eager: it runs whenever the outgoing slot is free, live
links notwithstanding, so a board that holds links but keeps losing
the sort can still dial a third party and merge two components. Two
guards make that safe. A live connection's address is excluded before
it can leave the scanner (the registry's addr_linked): the Core Spec
permits one connection per address pair (Vol 6 Part B §4.5), so such a
dial could only time out — the rig once showed exactly that, a doomed
5 s dial at an already-linked peer every ~20 s, forever. And a
fallback target that does not even connect goes into a dead-end table
for two minutes (BLE_DIAL_DEAD_END, sized to the RPA rotation
timescale), so a vanished advertiser is not re-dialled every backoff.
The interim alternative — suspending the clock while any link is live
— was shipped briefly and measurably cost merges: 28 and 78 all-linked
splits per 1000 arrival orders at 10 and 20 boards, against eager's
zero, because a component whose boards all hold some link can never
initiate a cross-component dial.
Which eligible advertiser gets dialled is not first-heard-wins: the
scanner collects one bounded window (BLE_SCAN_WINDOW) and dials the
best eligible candidate, via a CandidateTable shared by the firmware,
lnsd and the simulation. Best is, in order: strict verdicts before
fallback verdicts, then the peer with the MOST free incoming slots,
then the lowest address. First-heard-wins is what produced the
saturated-cycle lock — fallback dials closing cycles inside their own
component until nobody scans — measured at 48 of 1000 arrival orders
for twenty boards; the window's choice removes it entirely.
The free-slot count is the middle term, and boards put it on the air
themselves: capability bits 1-2 of the v0.3.0 record carry how many of
the three incoming slots are still free, bit 3 says the count is
present at all (leviculum-nrf/ble-tx/src/adv.rs). Without it a
searching board picks blind between a peer with three free slots and
one with its last one free, so several searchers elect the same board
and all but one are refused. In the arrival-order simulation the
preference leaves connectivity untouched (0 of 1000 orders
disconnected at both sizes, as before) and halves the boards that end
saturated: 976 to 498 at ten boards, 2538 to 914 at twenty.
Three properties keep it honest:
- It is a hint, never a permission. Nothing in the duplicate or refusal path reads it. A board that advertised a free slot and has none by the time the connection lands refuses exactly as before.
- It can understate, never overstate. The advertisement is rebuilt at each advertising start, and an incoming link can only land on the board that is currently advertising — which then stops and lets the next free task advertise the new, lower count. A slot freed while another task is mid-advertisement stays unannounced until that advertisement resolves, so the air can lag behind a board that got emptier, never behind one that got fuller. No running advertisement is stopped to rewrite it, so no advertising interval is dropped.
- Silence is not zero. A peer that advertises no count — an older board, a phone, another implementation — is ranked as if all its slots were free, the same "assume full capability" the v0.3.0 §3.2 rule applies to the capability flags, so it keeps exactly its pre-#375 standing. Bit 3 exists precisely so that an older record with only bit 0 set cannot be read as "zero slots free".
lnsd publishes the count too, since #432. Until the asymmetric cap it
did not, and the reason was not modesty: its capacity was one budget
shared by both roles, so a number in the boards' units would have been
a different quantity wearing the same bits. With one outgoing slot and
max_connections - 1 incoming, lnsd's free-incoming count IS a board's
free-incoming count, and it goes on the air from one shared builder
(links.rs's capability_record against the firmware's, the same two
calls) — so a peer can tell an lnsd record from a board's by the count
in it and by nothing else. BlueZ cannot edit a live record, so a
changed count is a deregister and a register; the published value is
remembered beside the handle and compared, so that pair is paid when a
link goes up or comes down and at no other moment. lnsd understates the
same way a board does, for the same reason: the link is admitted before
the record naming its slot leaves the air. And it now skips a peer
advertising zero — a peer whose last incoming slot is gone cannot
accept our dial, so the dial could only end in a refusal or, since a
full node goes dark and a dark node cannot even refuse, in the 20 s
setup timeout. That is the one place the count gates instead of ranking,
it gates only our own spending of a dial, and silence is still not
zero.
Two consequences of dialling against the sort are deliberate:
-
Cycles in the BLE graph are harmless. Reticulum treats every interface as a lossy broadcast domain and deduplicates packets at the transport, so a packet arriving over two paths costs one discarded duplicate, not a loop; announce rebroadcast is suppressed the same way. Connectivity is what the graph owes the mesh, minimal edge count is not.
-
A second link to an already linked identity is decided by the rule the PEER runs. It adds no reachability, burns one of three incoming slots and the airtime of a connect, so both roles resolve the identity at connect — read from the Identity characteristic as central, presented in the handshake as peripheral — and then apply one rule,
judge_duplicateinleviculum-nrf/ble-tx/src/registry.rs, which lnsd calls too. Three branches, and therule=token on every duplicate line says which one fired:rule=abandoned— the old link has delivered NOTHING, payload and keepalives alike, forLINK_ABANDONED_MS(30 s, two keepalive intervals). Its peer has walked away from it; the newcomer wins (BLE_LINK_REPLACED). This is the one case the peer's own arbitration never sees, so nothing it decides can contradict us.rule=same_role— both connections carry the same role, which a rotated address makes possible in either direction (the pre-dial exclusion is address-keyed, this rule identity-keyed). The peer then holds both in one role and has no central-vs-peripheral choice to make, so we decide alone: our own second dial is refused (it reaches nothing the live link does not), the peer's second dial wins (a node that dials again is done with the connection it has).rule=columba_mtu/rule=columba_identity— the general case:preferred_ble_role, a port of Columba's ownpreferredBleRole(centralMtu, peripheralMtu, localIdentity, peerIdentity), evaluated from the PEER's perspective. It keeps the role with the larger usable MTU and breaks a tie on identity order (localIdentity < peerIdentitykeeps central), and we keep whichever connection it keeps. "Usable MTU" is the peer's own conversion, bounds included:usableValueLength(rawAttMtu) = (rawAttMtu - 3).coerceIn(20, 512)(columba/rns-host/src/main/kotlin/network/columba/app/rns/host/ble/model/BleConstants.kt:87-88). The ceiling is not cosmetic — at the top of the range ATT 517 and ATT 515 both read 512, so a pair holding those two connections is a TIE the identity order decides; reading the 517 as 514 would decide it by MTU instead, and the two sides would keep different links.
Whichever branch fires, the LOSING link is disconnected by us at the decision, in either role — never left to the expiry.
Why the peer's function and not a rule of our own (#360 round 2): the only duplicate rule that never leaves the pair linkless is one both sides compute identically from the same inputs. On 2026-09-12 at 21:41 the field T114 and a Pixel running Columba 2.2.4-beta each deduplicated in the OPPOSITE direction within one second. The board kept the newer connection because the old one's last payload was 15 185 ms old; the phone kept its own central because its ledger still read the new connection's MTU at the floor (
centralPeerMtus[address] ?: BleConstants.MIN_USABLE_MTU, soMTU=20although the ATT exchange had long settled at 517). Each side then closed the link the other had kept, and the phone's cancel did not even drop the ACL: our outgoing link stayed up for 45 s collectingRejecting pre-identity writerefusals until our own expiry. 45 s of dead air per rotation cycle — the same hole round 1 closed from the other direction.What that capture showed at the floor was a FRESH connection, one the phone had not bookkept yet, and the rule is built on it being transient: our own dial enters the peer's comparison at
MIN_USABLE_MTUand is refused pre-handshake, while every other number in the comparison is the ATT MTU we negotiated, assumed to be the number the peer holds too. #377 puts that assumption in doubt from the peripheral end: on the desk 2026-09-09 a Columba peer listed a link the BOARD had dialled at "MTU 20 bytes" long after the exchange had settled, i.e. a peripheral-role ledger stuck at the floor for the life of the link. If its arbitration reads that number, a pair whose two connections negotiated the same MTU is a tie to us — settled by identity order, keeping the link we dialled — and an MTU decision to the peer, keeping the link it dialled, and the pair ends up linkless again in the identity order where those differ. Unconfirmed on the device: #377 waits for the capture #376 is taking, and whether our rule should model a peer's bookkeeping bug is that issue's decision, not a silent change of input here. The arithmetic of the divergence is asserted ina_peer_reading_its_peripheral_link_at_the_floor_decides_the_pair_the_other_way(leviculum-nrf/ble-tx/src/registry.rs), so a confirmation has one place to land.Why the liveness test is any-frame and not payload: Columba sends a 1-byte keepalive every 15 s on every connection it holds, so "any traffic within two intervals" is what distinguishes a link its peer still holds from one it has abandoned. Payload silence is what an IDLE phone looks like — it is not evidence of anything, and round 1 reading it is precisely what displaced a link the phone was actively keepaliving.
old_data_silence_msis still measured and logged, so a round-1-vs-round-2 capture stays greppable, and consulted by nothing.A link that has really stopped answering is removed by the expiry, not by a replacement:
last_heard_mscounts every inbound frame including the peer's keepalive, and a link that delivers neither payload nor keepalive forLINK_TIMEOUT_MS(45 s, three missed keepalives) is torn down by lnsd'sLinkTable::expireand by the firmware session'slink_silentarm — whether or not anybody dials the identity (BLE_LINK_EXPIRE role=<r> slot=<n> conn=<h> silence_ms=<n>). The replacement handles only the case the expiry is too slow for: the peer is here, on a new connection, asking to be reachable now. A REFUSED address goes into the dead-end table for its TTL so the scanner does not immediately re-offer it; a replaced link's queued packets move to the new link first.Every duplicate line carries
rule=,origin=(the role map the arbitration is computed over),old_mtu=/new_mtu=(the usable MTUs as compared, in the peer's ledger),old_silence_ms=(the any-frame measurement the abandonment test reads) andold_data_silence_ms=— round 1's input, reported for comparison,neverfor a link that carried none.What the expiry bound itself replaced is a measurement (#382). Over 14.1 h beside a Columba phone (
ble-accept-rns/lnsd.log, 2026-08-30) the gaps between received non-keepalive packets from a peer that was demonstrably present throughout ran to a median of 51 s, a 90th percentile of 182 s and a maximum of 5590 s; 502 links outlived 45 s with no payload at all. Payload silence is what an idle phone looks like, which is why round 2 removed it from the rule entirely. The decisions are pinned bya_live_old_links_own_dial_is_refused_before_the_handshake,an_abandoned_rotated_link_is_displaced_by_our_dial,payload_recency_no_longer_flips_the_verdict,keepalives_keep_a_link_out_of_the_abandoned_branch,a_same_role_pair_is_decided_without_the_peers_arbitrationandpreferred_ble_role_is_kotlinblebridge_44in the registry, by the two-sidedregistry::duelscenarios that assert the pair is never linkless for longer than 2 s, and byan_idle_link_is_alive_and_our_redundant_dial_is_refused,payload_recency_no_longer_flips_admission_in_either_role,our_dial_against_the_peers_live_central_link_is_refusedandthe_peers_second_dial_displaces_its_own_first_linkinleviculum-std/src/interfaces/ble/links.rs.
The simulation that motivated the fallback is a host test:
leviculum-nrf/ble-tx/tests/graph_formation.rs replays random arrival
orders through the real rule with the real slot limits. It reproduces
the issue's strict-rule disconnection rates exactly, asserts that no
fallback configuration ever leaves a board linkless, pins the quiet
suspension's measured cost as the record, and holds the shipped
configuration — eager clock, lowest-eligible window — to zero
disconnected orders at both sizes.
Three implementations exist in tree: the shared carrier logic
(leviculum-core/src/framing/ble.rs, plus the advertisement/decision
logic in leviculum-nrf/ble-tx), the firmware's dual-role
implementation (leviculum-nrf/src/ble/), and lnsd's dual-role BlueZ
interface (leviculum-std/src/interfaces/ble/, type = BLEInterface,
via bluer). The lnsd interface treats all live BLE links as one
broadcast domain behind one Reticulum interface — which is also the
only shape BlueZ supports on the peripheral side, where a GATT
notification reaches every subscribed central at once.
lnsd as peripheral: one ordered socket per central
When a peer dials lnsd, lnsd is the GATT server and every frame the
peer sends is a write to the RX characteristic. lnsd takes those
writes over BlueZ's AcquireWrite, not WriteValue (#436): the
characteristic declares write-without-response with bluer's
CharacteristicWriteMethod::Io, BlueZ asks for a socket on the first
write of each connection, and one reader task per central
(bluez::periph_reader) reads that SEQPACKET socket and hands each
value to the interface's event loop in the order the kernel received
it. Through WriteValue every write was its own D-Bus call, and bluer
runs each on its own tokio task, so two fragments a board wrote back
to back could reach the defragmenter END before START; the
2026-10-06 ble_lxmf_delivery red lost ten whole packets that way
with every fragment delivered.
Which writes take the socket is BlueZ's routing, not ours. The
GattCharacteristic API document (bluez 5.82,
doc/org.bluez.GattCharacteristic.rst, AcquireWrite) says that once
the fd is acquired WriteValue is locked, and the server side of
bluetoothd implements exactly that per connection: in
src/gatt-database.c chrc_write_cb looks up the connection's
acquired socket first and sends any write it finds one for down it,
a Write Request with response included, and acknowledges the ATT
request itself. Only a Prepare Write request is answered before the
socket is consulted, and only a failed AcquireWrite falls back to
WriteValue, which bluer refuses under Io. So the identity handshake, which the boards and
lnsd's own central write WITH response, also arrives on the socket:
it is the write that triggers the AcquireWrite, and bluetoothd
forwards it once the socket exists. The socket's end is the
central's disconnect (client_io_new registers a disconnect callback
on the ATT channel that closes it); the reader reports it as
PeriphGone and the peripheral link is released there, as a
central-role link is when its session ends, instead of waiting for
the expiry sweep.
The park-and-replay of an inverted fragment pair (#434, Link::oo_tail
in links.rs) stays behind this as the last line of defence: free on
an ordered stream, and a counted loss rather than a silent one should
a carrier ever reorder again.
Notifications are flow controlled, not fired and forgotten
The GATT server sends a packet as a sequence of notifications, one per
fragment. The SoftDevice queues those per connection, and the queue is one
entry deep by default (BLE_GATTS_HVN_TX_QUEUE_SIZE_DEFAULT = 1 on S140).
The second sd_ble_gatts_hvx of a packet is therefore refused with
NRF_ERROR_RESOURCES until the first one has actually gone out over the air.
A loop that pushes every fragment back to back and ignores the return value consequently delivers fragment 0 and drops the rest, silently, on both sides: the peer sits on an assembly that never completes and the node believes it transmitted. That is what the interface did until Codeberg #264, and it is why only single-fragment traffic — anything below one fragment payload, 177 bytes at the default MTU — ever arrived. An announce did not.
The rule that replaces it: a fragment is offered again after the queue drains, and a fragment that cannot be sent is reported, never discarded.
- The wait is on the SoftDevice's own
BLE_GATTS_EVT_HVN_TX_COMPLETE, not on a guessed interval.nrf-softdevicesurfaces it asgatt_server::Server::on_notify_tx_complete, whose default implementation throws the event away and which the#[gatt_server]macro does not generate — so the server type implementsServerby hand. - The wait is bounded, so a peer that stops listening cannot wedge the outbound task. On expiry the packet is abandoned like any other failure.
- Every abandonment emits
BLE_TX_DROP(see Structured event logs) and bumps a counter. A dropped fragment is never again indistinguishable from a sent one.
The decision itself — retry, abort, report, and the exactly-once ordering —
is pure and lives in leviculum-nrf/ble-tx, unit-tested on the host against a
scripted notification sink; the firmware only performs the actions. This is
the same split as the GNSS and telemetry policies, for the same reason: the
interesting states are queue-full-then-drains, queue-full-then-times-out and
hard-error-mid-packet, and none of them are reachable on demand with a real
phone in the loop.
Pacing on a link
Two packets whose fragments leave back to back on one connection can cost the receiver the first packet: the 2026-09-09 desk measurement (#376) showed a Columba phone one hop from two boards losing exactly the first of two fragmented packets arriving back to back, on both boards' links, reproducibly — and receiving both once the sender left 100 ms between the packets.
Every pump therefore serves a per-link inter-packet gap, measured
from the last fragment of one packet to the first fragment of the next
on the same connection: the compiled default is 100 ms
(leviculum-ble-tx's DEFAULT_TX_GAP_MS). The number is the measured
value, not a derived one — one to two connection intervals (30 to
50 ms) may suffice but was not measured. The knob stays for
measurement: lnflash --set-ble-tx-gap <ms> overrides the default on a
running board (0 disables the gap entirely), is never persisted, and
a reset restores the default. The first packet of a connection is never
deferred, an idle link pays nothing, and keepalives are neither paced
nor slide the window. Every actual wait logs
BLE_TX_GAP conn=<h> waited_ms=<n>. lnsd's Columba interface serves
the same default on its notify pipe and each central link — it is the
phone stand-in on the rig and must behave like a board toward a real
phone.
Related, from the same desk session: the peripheral pump drains
nothing before the peer can receive. The first notify on a fresh
connection, sent before the central had written the TX CCCD, fails
inside the SoftDevice (sd_error code=13313,
BLE_ERROR_GATTS_SYS_ATTR_MISSING) and the packet dies. The pump now
holds the drain until both the CCCD subscription and the identity
handshake have happened — packets queued before that wait, they are not
dropped — and logs BLE_TX_HELD conn=<h> reason=not-subscribed once
per connection when it actually held one. Policy host-tested in
leviculum-ble-tx's hold module.
ble-leviculum (BLE 5 broadcast mesh)
Reticulum broadcasts are sent as real BLE 5 connectionless extended advertisements, so any number of devices in range form a mesh without per peer links. This sidesteps the 3 to 4 link ceiling entirely.
To the core this is just another lossy broadcast medium, the same model LoRa already uses, so the existing robustness logic applies. A BLE 5 broadcast interface is a normal lossy broadcast Interface; the connectionless and size limited nature is a carrier quirk handled inside the interface.
Feasibility was confirmed by a read only spike (see
docs/ble5-broadcast-protocol3-spike.md in the repository). Key results, valid
for nrf-softdevice rev 5949a5b and SoftDevice S140 v7.0.0:
- Connectionless extended advertising is supported, including the pure
broadcast type
EXTENDED_NONCONNECTABLE_NONSCANNABLE_UNDIRECTED. - One advertisement carries at most 255 bytes, about 245 usable after framing.
- The receive side (extended advertising scan) is supported but gated behind
the
ble-centralfeature, currently off. - Periodic advertising is absent in S140 7.0.0. It is optional; repeated extended advertising suffices for a broadcast mesh.
The consequence is fragmentation. Small packets such as announces fit in one advertisement. A full 500 byte Reticulum packet (the MTU) does not and must be fragmented across two advertisements and reassembled. Because broadcast is lossy, a fragmented large packet only arrives if both fragments do; reliable large transfers use links over a connection oriented path, not broadcast, so this is acceptable.
Coded PHY: automatic range extension, first-class
Coded PHY (S=2/S=8) is a first-class part of the ble-leviculum design, not an
optional add-on. The reason is the range and rate frontier: room scale BLE at
1M on one end, km scale LoRa at kbit/s on the other, and nothing in between.
Coded PHY at S=8 trades 1/8 rate for roughly 4x range and fills exactly that
empty middle, about 200 to 800 m at a ~100 kbit/s class rate.
The architecture keeps it automatic. 1M extended advertising is the universal
floor: every device transmits and scans it. Controllers that support LE Coded
(runtime feature detection: the LE Coded feature bit; on Android
isLeCodedPhySupported()) additionally dual advertise the same payload on a
coded primary advertising chain and scan both PHYs (scan_phys = 1M | Coded,
the controller time shares the scan windows). No user configuration, and no
parallel meshes: coded capable nodes bridge by construction because 1M always
stays on. Double reception of the same payload is absorbed by the normal
Reticulum packet hash dedup.
Costs to tune, stated here and not solved here. Coded TX is about 8x airtime at S=8, so coded repeats are rate limited, for example one coded transmission per N 1M intervals. Splitting the scan budget across two PHYs lengthens discovery latency. Android background scan limits apply.
Capability is unevenly distributed, and the design accounts for that honestly. Recent Android flagships largely support Coded PHY, often both scan and advertise; midrange devices are mixed, sometimes scan only; iOS exposes no Coded PHY at all; cheap BT5 USB dongles often omit it. The nRF52840 on our boards and test dongles supports it fully. This asymmetry is why the design makes coded an automatic bonus above the 1M floor, never a requirement.
Combining the protocols
The broadcast and connection oriented protocols are not mutually exclusive. The useful combination is broadcast for reach (announces, discovery, small packets to everyone, no connection limit) and a connection for reliable directed bulk. This maps onto Reticulum's own layering: announces are best effort broadcast, links and resources are reliable and directed.
The constraint on how to combine them comes from the interface boundary. An
Interface is sent only bytes: try_send(&[u8])
(leviculum-core/src/traits.rs:280, plus a prioritized variant that adds a
priority hint) takes a packet buffer, no destination and no next hop. The next hop and the choice of interface live one
layer up in the transport. So an interface cannot decide "open a connection
because this packet is for node X" without reading the destination out of the
packet bytes, which is the link awareness the interface isolation rule forbids.
The Columba maintainer raised the same objection on the original proposal (see
the discussion linked below).
The clean way to combine them is therefore to keep the broadcast versus connection decision in the transport, which is allowed to be path aware, and to run two dumb interfaces rather than one clever one. Two staged options:
- Stage A, broadcast only. Run
ble-leviculumalone. Reliability for large or important traffic comes from Reticulum's existing link and resource layers riding on top of the lossy broadcast, exactly as they already do over LoRa. The interface stays a puretry_send(bytes)broadcast pipe, fully isolated, unbounded in scale, with a minimal failure surface. Build this first and measure throughput. - Stage B, broadcast plus per peer connections. If Stage A throughput is
not enough, run
ble-leviculumandble-reticulumtogether. Modelble-reticulumas one ordinary byte only interface per connected peer (the ColumbaBLEPeerInterfaceshape): each GATT link is a normal interface that sends the bytes it is given over its one connection, and the transport routes over the set of interfaces normally. Which peers to connect is a neighbour and discovery policy with an idle timeout to free connection slots, not a per packet trigger. The hybrid benefit then emerges from running both planes at once and letting the transport choose, with no clever single interface and no change to the byte only interface boundary.
Power shapes the deployment. Continuous advertising and scanning is costly on
phones, so powered nodes (lnsd, stationary RTNodes) run the broadcast plane,
while phones connect sparingly over ble-reticulum to a nearby powered relay.
True on demand, opening a connection because a packet needs to reach node X,
belongs in the transport, which knows the next hop, via a control path beyond
try_send. That changes the media agnostic interface boundary and is deferred
until measurement shows Stage A and Stage B are not enough.
This analysis follows a proposal and debate in the Columba project, discussion 880, a hybrid broadcast and on demand connection model. The points above record why a single hybrid interface is not the chosen path here.
Capability matrix
| Protocol | no_std carrier | nRF | lnsd | Phone | Interop with |
|---|---|---|---|---|---|
| RNode over BLE | seam is std today | planned | planned | n/a | Python-RNS, Columba |
ble-reticulum | in core + ble-tx | yes | yes (BLEInterface, via bluer) | Columba | Columba "Protocol v2.2" |
ble-leviculum | planned in core | feasible, spike done | via bluer | hardware dependent | leviculum only |
Decisions
- Broadcast instead of more BLE 4 links. The connection oriented model caps
at 3 to 4 reliable links, which does not scale to a dense mesh. BLE 5
connectionless advertising removes the ceiling, hence
ble-leviculum. - Fragmentation for full size packets. One extended advertisement holds 255
bytes on S140 7.0.0, the Reticulum MTU is 500, so the
ble-leviculuminterface fragments and reassembles. This is carrier logic, it lives in the interface, the core stays unaware. - Name in our own namespace.
ble-leviculum, notble5-reticulum, because the protocol is our unilateral invention, interoperates with nobody yet, and the*-reticulumnamespace is controlled upstream. - no_std carrier logic. So the same protocol code runs on embedded nRF and on the host. Only the radio and OS binding is platform specific.
- Combine by two interfaces plus transport, not one hybrid interface. The interface boundary is bytes only, so a single interface that opened connections per destination would need link awareness, which the isolation rule forbids. Keep the broadcast versus connection choice in the transport.
- Stage broadcast first, then measure. Build
ble-leviculumalone, let Reticulum's link and resource layers provide reliability over it, and only add per peer connections if measured throughput requires it. - Coded PHY is first class and automatic. It fills the empty middle of the range and rate frontier between 1M BLE and LoRa. 1M stays the universal floor; coded capable nodes dual advertise and dual scan on top of it, so the mesh never partitions and no user configures anything.
- Idle timeout governs connection management, not packet routing. Closing idle connections to free slots is fine. Triggering a connection open from a per packet destination is not.
See also
- Interface isolation
docs/ble5-broadcast-protocol3-spike.md, the ble-leviculum feasibility spike- Columba discussion 880, the hybrid broadcast and connection proposal: https://github.com/torlando-tech/columba/discussions/880
Media profiles
An LNode meshes over LoRa and BLE at once, by default, from the moment it boots. That is the right behaviour for a field node and the wrong one for a measurement: a packet delivered over the other medium masks a loss on the medium under test, so every single-medium number a dual-carrier node produces is falsifiable. "LoRa PDR was 94 %" is not a statement about LoRa if BLE was carrying the same mesh.
A media profile is the node's declaration of which carriers it meshes over. It is declared (a host says so), applied (the firmware honours it at boot and at runtime), and proven (the board says out loud, every boot, what it is on). All three are needed: a declaration nothing applies is a comment, and an application nothing proves is a hope.
The default is both carriers on
Absence of a profile changes nothing. A board with no record on its flash page, a board whose record is corrupt, a board whose record names a carrier this firmware does not know — all of them come up on both carriers, which is what every fielded board is already doing. There is no firmware update on this path that can take a board off the mesh.
That is why the record has a magic and a checksum rather than being a bare flag byte: a zeroed page read as flags would say "both carriers off", which is a silent way to lose a node.
Setting one
$ lnflash --set-media lora=on,ble=off
/dev/ttyACM0: media profile set — running lora=on ble=off, configured lora=on ble=off.
It survives resets; the board's own [MEDIA] line on if00 says the same.
lora= and/or ble=, comma or space separated, on/off
(true/false, yes/no, 1/0 also accepted), keys and values
case-insensitive. A carrier not named keeps the board's own setting —
lnflash reads the board back first and applies the spec on top, the
same read-modify-write contract --set-tx-power holds for the radio, and
for the same reason: a host that substitutes its own default for the
field it did not mean to touch changes what it claimed not to change.
--set-media with no value only reads:
$ lnflash --set-media
/dev/ttyACM0: running lora=on ble=off, configured lora=on ble=off.
What "off" means on the board
At boot, a carrier that is off is never started. The LoRa task that resets, configures and keys the SX1262 is not spawned, so the chip is never brought up; the Columba tasks are not spawned, so there is no advertisement, no scan, no GATT service and no connection. Nothing is transmitted and nothing is received.
The SoftDevice is still enabled with ble=off, deliberately. It is not
only the BLE stack: sd_flash_write is the one legal way to write
internal flash once it is enabled, and both persistence store tasks ride
on it. A board that could not persist its own profile could not be put
back on BLE — the one state this feature must never be able to reach.
At runtime, switching a carrier off stops it carrying Reticulum traffic in both directions immediately: the interface drops what the core hands it and the binary's receive arm drops what the medium hands up.
For BLE the runtime off also takes the carrier off the air: the Columba tasks disconnect every live link, central and peripheral role — a connected phone sees the board go, exactly as if it had left range — and stop advertising and scanning. Each dropped link unwinds through the same per-link teardown as range loss, so the core receives the same peer-lost report and culls its paths identically.
For LoRa it does not: the LoRa task keeps listening (nothing is transmitted, and what it hears is dropped before the core sees it). LoRa radio silence needs the boot path, which is why the acceptance for LoRa silence is set-then-reset and not a runtime set.
What "on" means, and when it needs a reset
Switching a carrier back on is immediate if it came up at boot — its tasks are still there, gated, and for BLE they resume advertising and scanning at once. A carrier that did not come up has no task to un-gate, and an embassy task cannot be spawned from nothing after the fact, so it cannot start before the next reset.
The board says which case it is in rather than acking either way. Both media frames are answered with a report carrying two profiles:
running— what the board is carrying traffic on right now.configured— what a reset would come up with.
They differ exactly when a carrier is waiting for a reset, and lnflash
renders that difference as a sentence naming the carrier. An ack would
have said "done" to a request the board cannot honour yet, and a
measurement run reading that ack would believe it had a BLE link that
does not exist.
Because "configured on, not running" is a terminal statement — the host prints "Reset the board" for it — the board must not be able to say it while it is merely still booting. USB comes up before the carriers by design, so the serial task answers frames during a window in which nothing has been spawned yet; in that window the board reports the declared profile as running, and narrows it to what really started as soon as both spawn decisions are made. Without that, a set-then-reset script that connects the moment the tty appears reads "did not come up" off a board that came up perfectly (seen on the rig, #255).
The proof line
Every boot, on the debug CDC, before anything can have moved:
[MEDIA] lora=on ble=off src=flash t=1183
The two carrier fields are what the board is running; src= is
flash for a profile read off the page and default for the both-on
fallback. It is re-emitted with the [FW_BUILD] banner every five
seconds, so a capture attached after the boot window still reads the
carriers off the board rather than off an operator's memory, and a
runtime change shows up within five seconds.
The shape is an interface. Assertions read it; it is frozen in
docs/src/structured-event-logs.md and must not drift.
Where the pieces live
| Piece | Where |
|---|---|
| Wire format, both frames and the report | leviculum-core/src/envelope.rs (TYPE_MEDIA_PROFILE, TYPE_MEDIA_QUERY, TYPE_MEDIA_REPORT) |
| Flash record | leviculum-core/src/media_profile_store.rs ("LMED") |
| Page layout and the store task | leviculum-nrf/src/telemetry.rs (+0x200 on BoardConfig::telemetry_flash_page) |
| Runtime state, the gate and the banner | leviculum-nrf/src/media.rs |
| Running/configured state machine, boot window, drop-run reporting (host tests) | leviculum-nrf/media-state/ |
| Boot spawn decisions | leviculum-nrf/src/bin/{t114,rak4631}.rs, leviculum-nrf/src/ble/mod.rs |
| Host flag | lnflash/src/media.rs, lnflash/src/flow.rs |
What this is not
It is not a power-saving feature and not a way to run a node on one carrier in the field — a node with a carrier off is a node that cannot be reached over it, which for a mesh is a fault, not a setting. It exists so that a measurement can say which medium it measured. Deployments run both carriers; that is the default and it stays the default.
How far one firmware build reaches
Every board we support costs a build, a bundle entry, a row in the test matrix and a place in everyone's head. The policy that keeps that cost from growing with the hardware catalogue:
A firmware build serves a whole family of boards. A build for one hardware configuration is only justified where universality is unreachable, and the burden of proof lies with the specialised build.
This is not an aspiration. It is already how the builds we have behave, and it was decided twice before it was written down.
Why a family and not a board
Boards differ in dozens of ways and almost none of them matter. What decides whether firmware runs at all is the SX1262 wiring: seven pins, plus the TCXO voltage and whether DIO2 owns the antenna switch. Every other difference is peripheral in the literal sense.
Those seven pins are usually not a property of the product. They are a property of whatever part carries the radio. When the radio ships together with the MCU as one module, every carrier board built around that module inherits identical wiring, and a product family of a dozen devices collapses to a single set of pins. The RAK4630 is the clean case: twelve different carriers in Meshtastic's tree, from the Pocket V2 to a solar repeater to an Ethernet gateway, all repeat the same seven numbers because the RAK4630 datasheet fixes them.
The unit of support is therefore the pinout family, and the concrete membership lists live in Supported boards.
What universality requires of the firmware
A single image only serves a family if everything a carrier adds either announces itself or costs nothing when absent. That is a design constraint on peripheral handling, not a hope:
- Probe where the bus allows it. The RAK baseboard display is found
by an I2C address probe (
ack_probe,leviculum-nrf/src/display.rs:55); when nothing answers, the task logs and exits (DetectedKind::None,leviculum-nrf/src/display.rs:161-167). - Fail into the harmless state. The user button is configured
Pull::Up(leviculum-nrf/src/button.rs:37), so an absent button reads as not pressed rather than as noise. - Park rather than spin. The GNSS task awaits a UART that simply stays silent when no receiver is fitted.
- Publish nowhere. The battery task samples a pin that floats on a bare module, but its only subscriber is the display task that is not running, so no wrong reading escapes.
Measured on this tree, carrying all of that costs 47.6 KiB of flash and 2.5 KiB of RAM over the stripped build. The RAM figure is affordable by construction rather than by luck: the same image already runs on a populated carrier with the same chip and the same memory, and a carrier board adds peripherals, never RAM.
Where a bus cannot be probed, writing blind is acceptable only for a known board. The T114 drives its ST7789 panel blind because MISO is not connected and detection is physically impossible, which is safe because we know what else is on those pins. The same reasoning does not transfer to an unfamiliar board, where the identical pins may carry something that must not be driven.
Runtime detection is therefore the lever, and a build-time feature is the fallback. Every peripheral moved from a feature flag to a probe removes a reason for a second build. Codeberg #240 does this for GNSS presence.
When a specialised build is justified
Three conditions, any one of which is sufficient:
- The radio wiring differs. No amount of runtime detection recovers
from pins that are simply elsewhere. The third build,
solarnode(Codeberg #233), is this case and nothing more interesting: the Wio-SX1262 puts all seven pins elsewhere and adds an eighth, a host RX-enable the other two families do not have. - A radio parameter is board-specific rather than family-specific and is compiled in. The Elecrow ThinkNode M1 matches all seven T114 pins but runs its TCXO at 3.3 V against the family's 1.8 V. Either the value becomes data, or the board needs its own build.
- A peripheral is dangerous when mishandled. A board with an external power amplifier needs its enable line driven correctly. Silence is not a safe default there, unlike a missing display.
Convenience, code tidiness, and "it would be cleaner to separate them" are not on this list.
A frequency region is not on it either. The compiled default is the
eu868 community profile, but the radio configuration is data, chosen at
flash time: lnflash ends every flash with a preset menu — eu868,
us915, au915, or a custom five-number entry — and stores the choice on
the board (see Flashing an LNode, "The radio
configuration belongs to the flash"). The presets are the settings each
regional Reticulum community has converged on; a default, not legal
advice.
Two device classes, one policy
Everything above was written about nRF52840 boards, because for a long
time those were the only boards we had firmware for. leviculum-esp
adds a second class — ESP32 and ESP32-S3 — and the policy does not
change: one build per pinout family still holds, and a board file
still describes a wiring rather than a product.
What the second class does add is a layer the first one never had to name out loud, because with one SoC family there was nothing to separate it from. The split is:
| Layer | Holds | Lives in |
|---|---|---|
| Class | the init order, the clock, which peripheral is the debug port, how the build stamp is formatted and emitted, the panic behaviour, the heap | leviculum-esp/src/lib.rs and the binary |
| Board | which GPIO carries which net, the radio's seven pins, the LED and its polarity, the battery divider, whether a GNSS receiver is fitted, the supply-enable lines | leviculum-esp/src/boards/<board>.rs |
| Protocol | every SX1262 opcode sequence, register bracket and timing calculation | leviculum-core::sx126x, shared with the nRF class |
The third row is the one that earns the split. The sequences that decide
what goes on the air are not per class and not per board; they were
lifted into leviculum-core precisely so a host test could run them, and
a second device class must not become a second copy of them. What a class
crate implements is the two traits leviculum_core::sx126x reaches the
hardware through — CommandBus and RegisterBus — and nothing above
them.
Adding the second board of a class is therefore: one
boards/<name>.rs with its pins and its BoardConfig, one
src/bin/<name>.rs that hands those pins to the shared init, one
[[bin]] stanza, one Justfile recipe. No new idiom, and nothing in
lib.rs moves. If adding a board does require moving something in
lib.rs, that is the signal that the thing being moved was a board fact
sitting in the class layer.
Where the classes differ, and why that is not a policy exception. An
ESP32-S3 board has no UF2 bootloader and no mass-storage volume, so
nothing in the ESP class corresponds to the Board-ID discussion below;
identification there happens over the serial protocol the ROM speaks, and
the same question — is the identifier bound to the same unit as the
wiring — has to be asked again on its own terms. The class boundary is
about where code lives. It is not a second policy.
The axis the policy does not have: the transceiver family
Everything above varies two things, the class and the pinout family, and holds a third fixed without ever saying so: every board in this book carries an SX1262. The unit of support is a set of seven pins because seven pins is all that differs once the part itself is settled. A board with a different transceiver does not sit anywhere on that scale, and the Seeed SenseCAP Card Tracker T1000-E (Codeberg #406) is the first one we own: the nRF52840 we already build for, a bootloader and a USB path we already flash through, and a Semtech LR1110 where every board above has an SX1262.
The class table above puts every opcode sequence in one row. That row is where the cost lands, and it is not spread evenly across the code that looks radio-shaped. Measured on this tree, 2026-09-25:
| What | Lines | Reaches a second radio family |
|---|---|---|
leviculum-core/src/sx126x.rs | 2461 | No: command set, register map, timing arithmetic |
leviculum-nrf/src/sx1262.rs | 1601 | The SPI and pin glue partly, the opcodes not at all |
leviculum-nrf/rx-arming/src/lib.rs | 3879 | Mostly yes, see below |
leviculum-nrf/channel-access/src/lib.rs | 745 | Mostly yes, see below |
The two pure crates are the cheaper half, and that is worth stating because their prose names the chip on nearly every page and reads like the expensive half. Both were lifted out of the driver so a host test could drive them, and being liftable is the same property as being portable:
rx-armingreaches a radio only through two traits,RxPort(leviculum-nrf/rx-arming/src/lib.rs:167) andRxWindowProbe(leviculum-nrf/rx-arming/src/lib.rs:706). The driver implements both,RxPort(leviculum-nrf/src/sx1262.rs:1450) andRxWindowProbe(leviculum-nrf/src/sx1262.rs:1480), and a test fake implementsRxPort(leviculum-nrf/rx-arming/src/lib.rs:1838) beside it. A second family writes a third implementation; it does not fork the crate.channel-accesstouches no radio at all. The caller reports what its own channel-activity detection said —cad_clear(leviculum-nrf/channel-access/src/lib.rs:323),cad_busy(leviculum-nrf/channel-access/src/lib.rs:333),cad_error(leviculum-nrf/channel-access/src/lib.rs:350) — and the jitter slot is derived from bandwidth, spreading factor and coding rate (jitter_slot_ms,leviculum-nrf/channel-access/src/lib.rs:116), which are properties of the modulation rather than of the part. Only the module's own text names the SX1262 (leviculum-nrf/channel-access/src/lib.rs:17).
Where the seam would have to open, if it opens. RxLatch
(leviculum-nrf/rx-arming/src/lib.rs:367) is three named LoRa interrupt
bits — preamble, header, RxDone — read back without being cleared, and
the two bounds computed from them are forwarded to
leviculum_core::sx126x::tx_defer_ms and
leviculum_core::sx126x::false_preamble_ms. A part that latches those
three the same way slots in behind the existing traits. A part that does
not forces the traits themselves open, and choosing between widening them
and carrying a second implementation is a design decision, not a port.
That decision is not taken here, and nothing in this tree has measured
an LR11xx part against these traits. Until one has, the honest estimate
is the driver in full and the two crates untouched.
The limit that bites
Universality reaches exactly as far as the identification does. A build may serve twelve carriers, but something has to decide that the board in front of it is one of those twelve, and that decision is made from the bootloader (see Flashing an LNode). Where the bootloader identifier is bound to the same unit as the wiring, the two line up and the family is safe end to end. Where a vendor shares one identifier across models with different wiring, a correct universal build can still be written onto a board it does not fit.
So the reach of a build and the precision of its identification have to be argued together. A family is only as wide as the narrowest of the two.
The firmware's host-test seam
A decision lives in a host-testable crate. The nRF binary calls it.
leviculum-nrf builds for thumbv7em-none-eabihf, sits outside the root
workspace with its own .cargo config, and runs no test of its own. Any
statement about firmware behaviour that is written inside it is therefore
provable only by flashing a board. This page says where the line between
the two sides runs, why it is not a matter of taste, and which decisions
are still on the wrong side of it.
Why the rule is about the bench
The rig is one bench. It is shared between the conformance corpus, the land gates and every manual measurement, and a hardware run costs hours — so anything provable only there competes with everything else provable there. A host assertion costs a second and runs on every push.
The pattern was found under time pressure rather than designed. Codeberg
#402 stayed open for days over a board's announce cap registration that
had in fact been correct since 594dd3f8, because nothing could show it.
What closed it was moving the step that carried the meaning
(AnnounceCapBitrate::sync_phy) into leviculum-core, where a host test
makes the same call the firmware makes. That is the pattern; this page is
it stated as policy instead of as one lucky fix.
What counts as a decision
A decision is anything whose wrongness is a behaviour, not a wiring fault: a cadence, a threshold, an ordering, a predicate, an arithmetic budget, a byte-exact line other tools grep. If a sentence about the code can be written as "under X it does Y", it is a decision, and that sentence belongs in a test.
The far side — what legitimately stays in leviculum-nrf/src — is
everything whose argument is a pin, a register, a SoftDevice syscall or
an embassy_time::Instant: peripheral access, task spawning, the boot
order, the I/O half of a driver.
The seam between them is a value type. The crate holds the decision and
the vocabulary it is expressed in; the firmware supplies the I/O and the
clock and does what it is told. leviculum-rx-arming is the sharpest
example already in the tree — it holds the order in which the receive
path re-arms and hands a frame up, and src/sx1262.rs supplies a chip to
drive (stand_down_for_rx, leviculum-nrf/src/sx1262.rs:965).
Where the seam runs today
The leviculum-nrf workspace has 32 members besides the firmware crate,
every one of them pure and host-testable: screen, sd-policy,
gnss-time, gnss-presence, gnss-init, telemetry-policy, ble-tx,
announce-policy, queue-budget, log-line, tx-spacing, rx-arming,
persist-ack, boot-trace, boot-count, channel-access,
media-state, record-log, pn-store, store-spike, qspi-bitbang,
battery-scale, settle-budget, drop-budget, heap-budget,
node-name, sync-batch, upload-proof, mute-lease, qspi-selftest,
qspi-boot, usb-policy (members, leviculum-nrf/Cargo.toml:14). Before
heap-budget and node-name joined them they carried 806 host assertions across 61 test targets —
measured 2026-09-25 by the host-triple lines of lint-nrf
(Justfile:75). mute-lease is the
newest and the rule's own case twice over: the deadline on a host's
transmit mute (Codeberg #410) is a decision that would have been
unassertable inside src/lora.rs, and the LORA_MUTE_EXPIRED line that
announces it went into log-line beside the two mute lines it closes
out, rather than being spelled at its one call site.
leviculum-core is the other half of the seam and counts the same way: a
decision that is not board-specific belongs there, where lnsd runs the
identical code. leviculum-announce-policy is deliberately shared with
the daemon for exactly that reason, so the cadence the desk measures on a
board is the cadence the daemon runs.
That number — how many firmware decisions can be asserted without a board — is the measure this page is judged by. It goes up when a decision moves, and it is the only thing that does.
The far side is not a choice
Whether leviculum-nrf should gain a host test target of its own is
settled by the compiler, not by preference. Both BSP features route
through the softdevice aggregator, lib.rs refuses a build with no BSP
selected, and the SoftDevice bindings do not compile for a host triple:
$ cd leviculum-nrf
$ cargo check -p leviculum-nrf --features bsp-t114 \
--target x86_64-unknown-linux-gnu
error: invalid register `r0`: unknown register
...
error: could not compile `nrf-softdevice-s140` (lib) due to 548 previous errors
(measured 2026-09-25). There is no feature combination that both links
and builds for the host, so #[test] inside leviculum-nrf has nowhere
to run. Consistently, leviculum-nrf/src contains zero #[test] and no
tests/ directory today.
So: "cannot move" is the definition of hardware-only. A decision that has not been moved is not hardware-only, it is untested. The question to ask of any firmware behaviour is never "can the firmware crate test this?" — it cannot test anything — but "what is the value type that carries this decision, and what is left over once it is gone?"
How the rule is gated
lint-nrf runs clippy and the tests of every workspace member except the
firmware crate, on the host triple, as --workspace --exclude leviculum-nrf. It is spelled that way rather than as a list of -p
flags so that adding a seam crate to members is the whole act of
gating it. The list it replaced lived in three places, and its prose copy
had already lost leviculum-upload-proof within a day of that crate
landing. A positive control confirmed the failure mode is silent: a
member carrying a deliberately red test failed the workspace form with
exit 101 and passed the -p list with exit 0, because the list did not
name it.
New code follows the rule. Old code moves when it is touched anyway — a bug fix in a stranded decision is the moment to move it, not a reason to defer.
Still stranded
The four areas that prompted this page are already seamed, and so is most of what the list below used to name; it is worth saying which, because the list below is what is actually left:
| Decision | Crate | Since |
|---|---|---|
| Announce cadence and per-peer limit | leviculum-announce-policy | 787ce002, 2026-09-09 |
| LoRa channel access (jitter, CAD retry) | leviculum-channel-access | 11c532f0, 2026-09-01 |
| Receive re-arm / hand-off order | leviculum-rx-arming | 6d289255, 2026-08-26 |
| Media flags, running vs configured | leviculum-media-state | 8bbce725, 2026-09-01 |
| Heap budget: endpoint links, boot and live serve cap | leviculum-heap-budget | 03e99edf, 2026-10-05 |
| Node name, derived defaults and the BLE-pending flag | leviculum-node-name | 8128a0d2, 2026-10-06 |
| Front-end position: never both switch paths at once | leviculum-rx-arming (front_end) | 1a1ef641, 2026-10-06 |
| The 1200-baud touch predicate | leviculum-usb-policy | 604b5fae, 2026-10-06 |
| Control-envelope answer windows and the deferral bound they cover | leviculum-usb-policy | 604b5fae, 2026-10-06 |
| The duty-hold notice: one frame per hold edge, and the state behind a media report | leviculum-usb-policy (DutyHoldNotice) | leviculum#501, 2026-10-07 |
The answer windows moved as constants, not as functions of the modulation. What the modulation moves is the delay they must cover, one maximum-size frame's airtime (728 ms at the SF8/125 kHz default, 4756 ms at SF10/62.5 kHz); the windows answer to the hosts' fixed waits (lnsd's 2 s per attempt), so a slow profile is answered busy and acked on the retry rather than waited out. The crate asserts both halves, and the envelope-inside-lnsd relation is a build-time assert the firmware build inherits.
What has no host assertion, in the order it is cheap to move:
1. The [TRANSPORT] ticker. The re-arm deliberately drops missed
periods so a busy loop does not then emit a burst of catch-up lines
(poll, leviculum-nrf/src/transport_stats.rs:106), and the line is
byte-exact because capture consumers grep it (log,
leviculum-nrf/src/transport_stats.rs:114). leviculum-log-line already
exists for the second half.
See also: Checks that are actually checks for why a stated rule without a gate does not hold, and Evidence and honesty in testing.
Flashing an LNode
Every board we flash is an nRF52840 carrying a factory UF2 bootloader. That bootloader, not the firmware on top of it, is the part that decides what a flashing tool can do. The rule that follows:
The application is not the board's identity. The bootloader is. Anything a tool needs to know before writing, it learns from the bootloader, never from the firmware currently running.
A T114 may arrive carrying Meshtastic, Meshcore, microReticulum, RNode firmware or our own. Each picks its own USB identity, and a crashed one picks none. The bootloader underneath is the same in all five cases, answers on a fixed USB ID, and publishes what it is in a text file.
The three states a board can be in
Application. Our firmware enumerates 1209:0001 (T114) or
1209:0002 (RAK4631), from usb_vid/usb_pid in
leviculum-nrf/src/boards/t114.rs:172-173 and
leviculum-nrf/src/boards/rak4631.rs:187. Two CDC ports: interface 00
is the debug log, interface 02 the Reticulum transport. Both IDs are
squatted pid.codes test IDs, flagged as a TODO at
leviculum-nrf/src/usb.rs:119. Heltec stock firmware uses 239a:8071.
Meshtastic on the SenseCAP Solar Node uses 2886:0059, and calls itself
"XIAO-BOOT" while doing so. Nothing stops an application from naming
itself after a bootloader, which is the sharpest available argument for
the rule above: the product string is application data, not evidence.
Bootloader (UF2/DFU). A different USB ID entirely, which is why an
application-ID match can never be true while a board sits in DFU
(leviculum-nrf/tools/uf2-runner.sh:279). Measured on the rig:
| board | bootloader USB ID | mass-storage label |
|---|---|---|
| T114 | 239a:0071 "Adafruit HT-n5262" | HT-n5262 |
| RAK4631 | 239a:0029 "Adafruit WisBlock RAK4631" | RAK4631 |
| XIAO nRF52840 | 2886:0044 "Seeed XIAO nRF52840" | XIAO-BOOT |
The last row is the Solar Node, measured 2026-08-17. Note that its bootloader is the one row whose product string does not contain the word BOOT, while its application's does.
Dark. Crashed firmware never enumerates at all. No touch reaches it; only a physical double-tap does.
Getting into the bootloader
Two mechanisms, and only one of them is ours to control.
1200-baud touch. The host opens a CDC port at exactly 1200 baud.
Our firmware answers the resulting SET_LINE_CODING in
leviculum-nrf/src/usb.rs:188, writes DFU_MAGIC_UF2_RESET (0x57) to
GPREGRET at 0x4000_051C and resets. The bootloader reads that
retained register on the next boot and stays in mass-storage mode.
Measured latency from stty ... 1200 to the bootloader appearing on
USB: 5 s on the T114, 3 s on the RAK.
Double-tap RESET. The bootloader's own mechanism, independent of any firmware: it sets a RAM flag at startup and stays in DFU if a second reset arrives while the flag is still live.
The touch only exists if the running firmware implements it. Ours does.
Stock Meshtastic does not, which is why a first flash away from
Meshtastic needs the manual double-tap (Justfile:1719); for that case
Meshtastic offers its own admin command, wrapped as just dfu-rak4631.
For Meshcore, microReticulum and RNode firmware on nRF we have not
measured it.
This is the load-bearing limit for any "fully automatic" tool. There is no universal software trigger. The double-tap is the only mechanism that works regardless of what is running, and it needs a human. A tool can be fully automatic for boards already carrying our firmware, which is every re-flash, and must fall back to one clearly announced key press otherwise.
And what that key press is differs per board. The WisMesh Pocket V2
has no externally accessible RESET at all: the contact is reachable only
through a hidden pinhole beside the USB socket, double-tapped with a
needle (Recovery, "the hidden-pinhole
caveat"). Telling its owner to press RESET twice sends them looking for a
button that does not exist, which is a worse failure than saying nothing
— they conclude the board is dead. So the wording is a board fact like
any other and lives in lnflash/catalogue.toml
([board.<name>.flashing.double_tap]), not in a branch around the
prompt; a board
that says nothing there gets the ordinary wording. just flash-rak4631
carries the same line into the developer runner's prompt through
LEVICULUM_DOUBLE_TAP_HINT (Codeberg #261).
RNode firmware on the T114 lands a board in exactly that state
Mark's official RNode build for the Heltec T114
(rnode_firmware_heltec_t114.zip, v1.85/1.86) is an app-only Nordic DFU
package. Its application vector table reads SP 0x20040000 and reset
vector 0x00051819, so the image is linked for a flash base around
0x51000. This board's factory bootloader with S140 7.3.0 runs
applications at 0x27000, which is the base our own image is linked for
as well (leviculum-nrf/memory.x:16). rnodeconf pushes the app-only
package through the factory bootloader (adafruit-nrfutil dfu serial --package … -t 1200), so it lands at 0x27000; the bootloader jumps
there, reads a reset vector pointing into unprogrammed flash, and
hard-faults before USB comes up. The same dark board as the SoftDevice
mismatch below, from a different cause. rnodeconf itself calls the
T114 target experimental and AS-IS.
The operational consequence is the one this page keeps arriving at:
nothing is running to answer the 1200-baud touch, and a board that never
enumerates is not even a candidate for a tool to find, so it takes a
double-tap. After that nothing further is special — lnflash writes our
application at 0x27000 over whatever was resident and it boots. Where
the previous firmware was linked does not affect where the next one is
written. Measured on 2026-08-10 on serial 183004F712B4A7FE:
rnodeconf v1.86 reported "Device programmed" and then could not reopen
the port, the board did not re-enumerate, a double-tap brought the
HT-n5262 bootloader back, and CURRENT.UF2 showed the RNode vector
table with reset 0x51819 sitting at 0x27000. lnflash then flashed
over it and the board came up as 1209:0001 "leviculum T114".
What the bootloader tells you
Mounting the mass-storage drive gives three files: INFO_UF2.TXT,
INDEX.HTM and CURRENT.UF2. The first is the entire basis for
deciding whether a board can be flashed. Read from the rig, verbatim
(CRLF line endings):
UF2 Bootloader 0.9.0-2-g836c8dc-dirty lib/nrfx (v2.0.0) lib/tinyusb (0.12.0-145-g9775e7691) lib/uf2 (remotes/origin/configupdate-9-gadbb8c7)
Model: HT-n5262
Board-ID: HT-n5262
Date: Jul 9 2024
SoftDevice: S140 7.3.0
UF2 Bootloader 0.4.3
Model: WisBlock RAK4631 Board
Board-ID: WisBlock-RAK4631-Board
Date: May 20 2023
Ver: 0.4.3
SoftDevice: S140 7.3.0
The SoftDevice: line is generated at runtime by the bootloader from
the SoftDevice actually installed, so it answers the version question
directly rather than by inference. The T114's Board-ID is exactly
HT-n5262, not a substring of something longer, which retroactively
justifies the substring match at
leviculum-nrf/tools/uf2-volumes.sh:258.
Note the T114 bootloader build date matches Heltec's published
HT-n5262-bootloader-20240709.hex, so that board still carries its
factory bootloader.
What a UF2 is allowed to write
The bootloader validates every block's target address before writing.
From src/usb/uf2/uf2cfg.h upstream:
#define USER_FLASH_START MBR_SIZE // skip MBR included in SD hex
#define USER_FLASH_END (BOOTLOADER_REGION_START - DFU_APP_DATA_RESERVED)
in_app_space() accepts USER_FLASH_START <= addr < USER_FLASH_END;
blocks below the window are skipped silently while still reporting
success, and blocks at or above it are rejected outright.
Both bounds are measurable without reading a single constant, because
the bootloader generates CURRENT.UF2 from exactly that window. On both
rig boards it covers 0x1000 to 0x000EA000. So USER_FLASH_START is
0x1000, one MBR page, and USER_FLASH_END is 0xEA000. Nordic's own
nrf_mbr.h agrees: #define MBR_SIZE (0x1000), carried through into
the nrf-softdevice-s140 bindings as 4096.
The consequence is the important part: the writable window opens
directly above the MBR and includes the whole SoftDevice region. An
image converted from a SoftDevice hex has its MBR blocks declined and
the rest installed, which is precisely what the skip MBR included in SD hex comment describes. A SoftDevice can therefore be replaced through
the ordinary mass-storage path, without touching the bootloader.
leviculum-nrf/memory.x used to contradict this, computing safe
application space as 0xEC000 - 0x27000 (788 KiB) — 8 KiB the
bootloader would have refused to write. It now links the application
against the bootloader's own window minus the record store's region and
the boot-record page, 0xD9000 - 0x27000 = 0xB2000 (712 KiB); the
image is ~660 KiB, so the change costs nothing today and the gate prints
the remaining gap on every push (scripts/check-nrf-store-gap.sh).
Everything at or above 0xEA000 survives every UF2 flash, because the
bootloader declines those blocks. All three persistence pages live there:
identity at 0xEC000, the radio configuration at 0xEB000 and the
telemetry target / fixed position / media profile at 0xEA000
(leviculum-nrf/src/boards/t114.rs, radio_store.rs, telemetry.rs). A
user's chosen frequency therefore survives a firmware update as well as a
reset.
The record store (#384, 0xDA000–0xEA000, 16 pages) survives one too,
but for a different reason, and the difference matters because it is the
weaker guarantee of the two. The store sits inside the writable
window, so USER_FLASH_END does not protect it. What protects it is that
the bootloader erases only the pages it writes: flash_nrf5x_write
buffers one page and flash_nrf5x_flush (upstream
src/flash_nrf5x.c) erases and programs exactly that page, and only when
its content differs. Our .uf2 carries blocks from 0x27000 to the end
of the image and none above it, so no page of the store is ever a target
and no erase reaches one. That holds as long as the image stops below
the store, which memory.x's ASSERTs make a link error and the gate
above reports as a number.
The boot record (#380, 0xD9000, one page) is the same case one page
further down, and it is where the image now stops: it counts the boots a
field board made while nobody was watching (boot_count.rs), so it has
to outlive both a power loss and a firmware update. It is protected by
the same argument as the store and by no other — our .uf2 carries no
block for it — and it is deliberately the thing directly above the
image, so the linker's "will not fit in region FLASH" is the first
thing an over-grown image hits. A foreign image, or a tool that writes
blocks up here, erases the page; the record's magic is what makes the
bytes it leaves read as foreign rather than as a count.
Family IDs seen in practice:
| family | meaning | writes |
|---|---|---|
0xADA52840 | nRF52840 application | application region |
0x239A0071 / 0x239A0029 | board-specific application | application region |
0xD663823C | bootloader self-update | MBR, 0xF4000 region, 0xFD800, UICR 0x10001000 |
Only the last one touches the bootloader itself. A flashing tool must
never emit it. With the bootloader intact, an interrupted flash is
recoverable: the board comes back into DFU and can be rewritten. Replace
the bootloader and a failure needs SWD to undo. There is a precedent for
the general risk: a bad application build once left the RAK4631 in a
HardFault boot loop that required opening the case to reach the internal
reset button (commit 43d25830).
The SoftDevice
Our firmware links against S140 v7.x and places its application at
0x27000 (FLASH ORIGIN, leviculum-nrf/memory.x:84). A board
carrying S140 6.1.1
puts the boundary at 0x26000 instead, so the version is not cosmetic.
The bindings we compile against are generated from S140 7.0.1
(SD_VERSION = 7000001, which decodes as major 7, minor 0, bugfix 1 —
not 7.0.0 as docs/ble5-broadcast-protocol3-spike.md:39 states). The
boards run 7.3.0. That works because Nordic keeps the ABI stable within
a major version, which is what makes >=7.0.1, <8.0.0 the honest
version constraint rather than a guess.
Reading the version without trusting the bootloader
INFO_UF2.TXT is the convenient source, but it depends on the
bootloader choosing to emit the line. The SoftDevice states its own
version in flash, independently: nrf_sdm.h puts the info struct at
offset 0x2000 above the MBR, with the version word 0x14 into it, so
the absolute address is 0x3014. The encoding is
major * 1000000 + minor * 1000 + bugfix.
That address sits inside the range CURRENT.UF2 dumps, so it can be
read without writing anything. Verified 2026-08-09 on both rig boards:
the word reads 7003000, decoding to S140 7.3.0 and agreeing exactly
with what each board's INFO_UF2.TXT claims. A tool that cross-checks
the two is immune to a bootloader too old to report the line at all.
The image is not part of this repo's source. The crate dependency
(leviculum-nrf/Cargo.toml:254) supplies Rust bindings, not the blob.
The authoritative copy is Nordic's own distribution, downloaded
2026-08-10 to ~/coding/s140_nrf52_730/, containing
s140_nrf52_7.3.0_softdevice.hex (md5
29013ba2d0507c25f62dffa96b6c67af), the API headers, release notes and
s140_nrf52_7.3.0_license-agreement.txt. Parsed, the hex covers
0x0-0xB00 (MBR) and 0x1000-0x26498 (the SoftDevice itself,
149 KiB), which lands exactly below our 0x27000 origin once
page-aligned.
For a while the only copy on our machines was inside a Meshtastic checkout. That copy is authentic — byte-identical after stripping CR, both files 9726 lines — but it had been converted to LF, whereas Nordic ships CRLF. Two consequences: an Intel-HEX parser must handle both, and the vendored image should be Nordic's original so that no build path leads through somebody else's repository.
Nordic's headers also settle two constants this page derived by other
means. nrf_mbr.h defines MBR_SIZE (0x1000), and nrf_sdm.h gives
SD_MAJOR_VERSION 7, SD_MINOR_VERSION 3, SD_BUGFIX_VERSION 0 —
encoding to exactly the 7003000 both rig boards report from 0x3014.
The mismatch is a soft brick, not a dead board
Flashing our application onto a board still carrying 6.1.1 produces a
device that goes dark: the old SoftDevice forwards to 0x26000, finds
no vector table there because our image starts a page higher, and
crashes before USB initialises. No CDC ports, no bootloader drive,
nothing on the bus. It looks hardware-dead and is not. The bootloader
region at 0xF4000 is never touched by an application flash, so a
physical double-tap always brings the UF2 drive back. The 1200-baud
touch is useless here, because the application never runs far enough to
answer it.
Confirmed on 2026-08-08 on the T114 with serial 183004F712B4A7FE,
which had been written off as bricked for weeks. Its INFO_UF2.TXT read
SoftDevice: S140 6.1.1. That board was never part of the 2026-05
spike, which names only DEC9947DAD9D2869, so it is direct evidence for
the factory state: factory T114 boards ship S140 6.1.1, as
leviculum-nrf/memory.x:5 claims.
The SoftDevice carve-out: the RAK4631 states the constraint and ships no remedy
The version constraint is the same on both boards, >=7.0.1, <8.0.0, and
for the same reason: our image is based at 0x27000 and cannot start on
a 6.1.1 boundary. What differs is what the bundle can do about a
violation.
For the T114 the remedy is measured end to end — a genuine 6.1.1 board
was repaired unattended on 2026-08-10, recorded under "Verified on
hardware" below. For the RAK4631 (Codeberg #261) it is unmeasured.
Our only RAK, serial DEC9947DAD9D2869, has carried 7.3.0 since we first
flashed it, so we have never seen the failing state on that board and
cannot say what a factory Pocket V2 ships. memory.x:5 claims both
boards leave the factory on 6.1.1, but that line predates the T114
measurement that confirmed it for the T114 alone.
So the bundle carries no SoftDevice for the RAK4631, deliberately. A
precondition without a remedy is allowed by construction — Requires and
Remedy are separate tables, and flow::resolve answers an unmet
constraint with no remedy by refusing:
rak4631 needs SoftDevice >=7.0.1, <8.0.0 and the board has 6.1.1
(bootloader and flash agree), but this bundle carries no remedy for
that. Nothing was written.
That is the honest outcome for a case nobody has seen: "I cannot fix this, here is why" beats writing anyway, and it beats shipping a repair path whose only evidence is that the same blob works on a different board. The opposite arrangement is what the loader refuses outright — a remedy with no precondition to trigger it will not load.
What would close this. One factory or Meshtastic-stock Pocket V2 read
through lnflash --dry-run under sudo, which prints the SoftDevice:
line off INFO_UF2.TXT without writing anything. If it reads 6.1.1, the
same vendored S140 7.3.0 hex serves this board too — it is an nRF52840
blob, not a board-specific one — and the carve-out becomes one line in
scripts/lnflash-bundle.sh's board list plus the payload pair beside it.
Until somebody reads one, the gap stays visible here rather than assumed
closed.
Installing the SoftDevice
adafruit-nrfutil's serial DFU protocol does not work against the
Heltec bootloader; the 2026-05 spike got "Timed out waiting for
acknowledgement" (f517a172). Do not retry it. The mass-storage path
works, and needs no SWD probe:
python3 ~/coding/meshtastic/bin/uf2conv.py -f 0xADA52840 -c \
-o s140_7.3.0.uf2 \
~/coding/meshtastic/bin/s140_nrf52_7.3.0_softdevice.hex
Copy the result onto the mounted bootloader drive and sync. The
application flashed earlier boots immediately afterwards; it was intact
all along, only the SoftDevice beneath it was wrong.
The conversion is deterministic, so its output can be checked before anything is written. Reproduced 2026-08-09:
| property | value |
|---|---|
| output size | 311 296 bytes, 608 blocks |
| family | 0xADA52840 |
| ranges | 0x0-0xB00 and 0x1000-0x26500 |
| highest byte touched | 0x26500, below the 0x27000 app base |
The last row is the one that matters: the update cannot reach the application, which is why the app survives it.
Eleven of those 608 blocks are never written. They carry the MBR
below 0x1000 and the bootloader declines them, silently and with a
success return. The block counter still sees all 608 arrive and reboots
on the last one, so nothing about the transfer looks unusual. This is
harmless, since the MBR is already present and identical, but it means
an application-family UF2 can never replace an MBR, and a report that
counts copied blocks is not evidence that all of them landed.
Why we cannot simply ship it
The SoftDevice is under Nordic's five-clause BSD variant, and two of those clauses decide the architecture of any flashing tool we build. This is a reading of the licence text, not legal advice.
Clause 2 permits redistribution in binary form, provided the copyright notice, the conditions and the disclaimer travel with the distribution. Clause 4 restricts use to Nordic silicon, which our case satisfies. Clause 3 is trivially satisfiable. So handing the blob to a user is allowed, as long as the licence goes with it.
The obstacle is the combination with our own licence. Clause 4 limits what the software may be used for, and clause 5 forbids modification, decompilation and disassembly outright. AGPL-3.0 grants every recipient the right to use and modify the whole work for any purpose, and permits no additional restrictions of that kind. A blob carrying clauses 4 and 5 therefore cannot become part of one combined work with AGPL code.
The practical consequence: the SoftDevice must not be linked into an
lnflash binary via include_bytes!. That would make it part of the
executable and put the two licences in direct conflict. Shipping it
alongside as a separate file, with Nordic's own licence file next to it,
is ordinary aggregation and does not have that problem.
Decided (2026-08-09): we ship it, as a separate file with its licence beside it, sourced from Nordic's own distribution rather than a third-party checkout. Our own firmware images travel the same way, even though being ours they could be embedded. One payload layout beats a split where some images live inside the binary and others outside, and it is what makes the bundle below the extension point for new boards. The binary stays a single static executable; it just is not the only file.
Ship Nordic's own licence file, not a copy of the text. The
distribution includes s140_nrf52_7.3.0_license-agreement.txt, and it
differs from the widely circulated LICENSE-NORDIC in exactly one line:
its notice reads Copyright (c) 2007 - 2020, Nordic Semiconductor ASA
where the other says only Copyright (c) Nordic Semiconductor ASA.
Clause 2 obliges us to reproduce the above copyright notice, so the
file that travels with the blob is the one Nordic shipped alongside it.
It also names its own product and version, which the circulated variant
does not. That is what lnflash/payload/t114/ vendors, next to Nordic's
CRLF original of the hex.
Note also that Meshtastic vendors the blob under GPL-3 without any
accompanying Nordic notice, so their practice is not the precedent it
was taken for; it fails clause 2 on its face. The nrf-softdevice
project is the counter-example worth copying: it is MIT/Apache licensed
and places a LICENSE-NORDIC in every crate that carries Nordic
material. That project has no copyleft conflict to solve, so it
demonstrates correct attribution, not that the AGPL question goes away.
One further detail. Converting the hex to UF2 does not alter a byte, only the container, so it is hard to read as the "modification" clause 5 prohibits; still, distributing the untouched hex and converting at runtime avoids the question entirely.
CURRENT.UF2 as a backup
The bootloader exposes the installed flash as CURRENT.UF2, 1.9 MB
covering 0x1000-0xEA000 under the board-specific family ID. Filtering
it to blocks at or above 0x27000 and renumbering blockNo/numBlocks
yields a restorable application image. Verified on both rig boards on
2026-08-09: read, filtered to 3120 of 3728 blocks, written back, and
each board returned with its original serial and firmware
([FW_BUILD] git_sha=bb7c4f64 on the T114), PANIC_COUNT total=0.
That makes a flash reversible for the user who wants their previous firmware back, at the cost of one file copy before writing. The SoftDevice portion of the dump is not needed for restore and is filtered out; keeping it would only re-write identical bytes.
It also identifies what is installed, without running it. The dump is
the application region, so the strings in it are the application's. Read
off the Solar Node on 2026-08-17 it gave Meshtastic 2.7.15, build
567b8ea, build target seeed_solar_node, and an occupancy of 92.4 per
cent that distinguishes a programmed board from a blank one. Two uses
follow. A tool can name what it is about to overwrite instead of
reporting that it found "a board", and where the Board-ID is ambiguous
the foreign firmware's own build target often names the carrier that the
bootloader does not. The second use is inference from a third party's
build strings and belongs in a prompt to the user, never in a silent
decision to write.
Practical details that bite
The USB serial number may change between modes, and whether it does is
board-specific. The T114 reports 183004F712B4A7FE as an application
and 12B4A7FE183004F7 in the bootloader: the two 32-bit words are
swapped. The Solar Node reports 40E37463CA8A59DF in both, unchanged
(measured 2026-08-17). Anything correlating a device across app to
bootloader to app must therefore accept both forms rather than assume
either. That is what same_serial (lnflash/src/usb.rs:140) does: it
tests equality first and the swap only as an alternative, so a board
that keeps its serial is matched as readily as one that swaps
it. The older runner is unaffected because it
only compares serials in application mode
(leviculum-nrf/tools/uf2-runner.sh:287).
Writing needs root. The mass-storage device appears as /dev/sdX
owned root:disk. Automounting assumes a desktop stack that a headless
host does not have. A single self-contained binary can read USB identity
from sysfs and issue the touch through termios without any external
tool, but it cannot write the drive unprivileged.
A successful write ends in a kernel error. The bootloader reboots
the moment the final UF2 block lands, while the filesystem still wants
to flush metadata, producing device offline error ... lost async page write. This is the normal completion path, not a failure
(leviculum-nrf/tools/uf2-runner.sh:234).
A copy returning 0 does not mean the flash took. Verify that the
application re-enumerated and that the bootloader drive is gone
(leviculum-nrf/tools/uf2-runner.sh:308). Stronger still, read the
periodic [FW_BUILD] banner off the debug port and compare the git SHA,
as scripts/flash-lnodes-from-head.sh:113 does.
More than one board can be in its bootloader at once, and the wrong
one is usually first. The volumes are anonymous mass storage; only
Board-ID distinguishes them. A tool that takes the first UF2 volume it
finds and stops looking will be handed the same wrong volume on every
retry, refuse it every time, and never examine the board it was asked to
flash. find_uf2_drive therefore enumerates every candidate and
poll_matching_drive selects across the whole set
(leviculum-nrf/tools/uf2-volumes.sh).
A mount you make is a mount you owe back. Mounting a volume to read
its Board-ID and then leaving it behind because it was the wrong board
is worse than not looking: the leaked mount shadows every later board in
the search path and survives until somebody unmounts by hand. Volumes
this tooling mounts go under /run/leviculum-uf2/<device> — one mount
point per device, never the single shared /mnt, so a foreign volume
cannot occupy the only slot — and are released on every exit path.
A give-up message must report, not assert. The runner used to end
every failure with "app never re-enumerated" even when no volume for the
board had ever been found, naming a symptom that had not been reached.
It now names the volumes it saw and their Board-IDs.
Structuring a flashing tool
Everything above is mechanism. What follows is the shape a tool takes if it has to survive more boards than the two we support today.
Start with how wide the field actually is. The Meshtastic tree carries 162 variants, and they collapse onto very few flashing mechanisms:
| chip family | variants | how it is flashed |
|---|---|---|
| nRF52840 | 50 | UF2 mass storage |
| RP2040 / RP2350 | 12 | UF2 mass storage |
| ESP32 / S3 / C3 / C6 / S2 | 93 | ESP ROM bootloader over serial |
| STM32 | 5 | its own path |
Two transports cover 155 of the 160 flashable variants, and we
already own both: the UF2 path in leviculum-nrf/tools/uf2-runner.sh
and the ESP path behind Justfile:774, which drives esptool. The
work is not building 162 things. It is separating two mechanisms
cleanly and turning everything else into data.
Four axes, not one "board"
Treating a board as one indivisible unit is the design mistake to avoid. A board is four independent answers, and a new device rarely changes all four:
Identify — what is attached? For UF2 boards the truth is the
Board-ID in INFO_UF2.TXT; for ESP32 it is the chip identity the ROM
bootloader reports. Never the USB ID of the running application, which
belongs to whatever firmware happens to be installed.
That truth is authoritative but not always sufficient, and the condition under which it is sufficient can be stated exactly:
A
Board-IDcarries a write decision only where it is bound to the same physical unit as the radio wiring. Where the two are bound to different units, a match is a hint.
Three real bindings, all measured in 2026-08:
- Coupled. On the RAK4630 the SX1262 wiring and the bootloader both
belong to the module. Twelve different carriers report
WisBlock-RAK4631-Boardand share seven identical pin numbers. The key is exact, and one image serves all of them. - Decoupled by the vendor. Heltec records the same bootloader product
string
HT-n5262for the Mesh Node T114, for MeshSolar and for the Mesh Pocket. The first two share our wiring; the Mesh Pocket puts CS onP0.26and BUSY onP0.15. The identifier belongs to a bootloader shared across models while the wiring belongs to the model. - Decoupled by construction. The Seeed XIAO is an MCU module with the
radio outside it, so a SenseCAP Solar Node and a DIY XIAO with
different radio wiring both report
nRF52840-SeeedXiao-v1.
In both decoupled cases a manifest entry keyed on info_uf2_board_id
alone is not a decision. Such a board needs a second discriminator or an
explicit question naming the model, and the honest failure is to stop and
ask rather than to write the more likely of two images. The cost of
getting this wrong is not a failed flash but a board driving the wrong
pins, which on hardware carrying a power amplifier is a repair rather
than a retry.
The cheaper answer, where it is available, is to make the ambiguity stop
mattering. The RAK4631 looks like the same problem, since the bare
module and the Pocket V2 share a Board-ID and have separate builds,
but the baseboard build degrades cleanly on a bare module: the display
is found by an I2C probe and its task exits when nothing answers
(leviculum-nrf/src/display.rs:161-167), the button pin is pulled up so
it never reads as pressed (leviculum-nrf/src/button.rs:37), the GNSS
task waits on a UART that stays silent, and the battery task publishes
into a watch channel whose only subscriber is the display that is not
running. One image therefore covers both, at 47.6 KiB of flash and
2.5 KiB of RAM that the bare module does not use, and the RAM cost is
already proven affordable because the same image runs on a Pocket V2
with the same chip and the same memory. Prefer that over asking a
question, and reserve the discriminator for boards whose peripherals
genuinely cannot be probed.
Enter — how does it reach a programmable state? 1200-baud touch, physical double-tap, a DTR/RTS sequence on ESP32, BOOTSEL on RP2040.
Transport — how do the bytes get in? The two above.
Verify — did it take? Re-enumeration plus the [FW_BUILD] banner
with a matching git SHA.
Crossing all four sit preconditions. The SoftDevice version is the only one today; a bootloader minimum version would be the next. A precondition must be data that names its own remedy, never a special case in code.
Separated this way, a new nRF or RP2040 board is data entry, and a new chip family costs exactly one new transport.
The bundle is the extension point
Since third-party blobs cannot be linked in anyway, the payload lives beside the binary and the manifest describes it:
lnflash # board-agnostic binary
firmware/
manifest.toml # index, checksums, licences
t114/
leviculum-t114-0.8.0.uf2
s140_nrf52_7.3.0_softdevice.hex
s140_nrf52_7.3.0_license-agreement.txt
rak4631/
leviculum-rak4631-0.8.0.uf2
The second board arrived in Codeberg #261 and cost exactly what this
structure promised: a catalogue entry, an image, and a line in the board
list scripts/lnflash-bundle.sh walks. No Rust changed except the
per-board wording of the one prompt a human has to act on — see
"Getting into the bootloader" below.
Note what the RAK directory does not contain. The SoftDevice remedy is per board and this bundle carries none for that one; see "The SoftDevice carve-out" below for why, and what a board that needs it is told.
The data is split in two by what it describes. Board facts — USB
IDs, Board-ID, flash geometry, the SoftDevice constraint — are
properties of the hardware and do not change when a release is cut, so
they live in lnflash/catalogue.toml, compiled into the binary:
[board.t114]
family = "nrf52840"
candidate_usb = ["1209:0001", "239a:8071"]
[board.t114.flashing]
transport = "uf2-msc"
entry = ["touch-1200", "double-tap"]
identify = { info_uf2_board_id = "HT-n5262", bootloader_usb = ["239a:0071"] }
requires.softdevice = ">=7.0.1, <8.0.0"
The entry is itself split, along the same seam (Codeberg #233). The
top level is what talking to a running board needs, and it is one
field: the USB IDs its firmware answers on. Everything a write needs
sits under flashing, and that table is optional. A board with none is
control-only — --watch, --announce, --set-time and the --radio-*
flags reach it exactly as they reach any other board, while no bundle may
carry an image for it, --board <name> is refused before the bus is
read, and a flash session that meets it on the bus says so and does not
even reboot it.
That is not a lesser kind of support; it is the honest kind for a board
whose Board-ID is not an identity. The two halves rest on different
evidence: a control frame reaches a board that is up and identifying
itself, while a write rests on what a bootloader publishes. The Solar
Node is the first such entry, and the reason is data rather than a
comment — a control-only board must state not_flashable, which is the
sentence the user is refused with, and stating both halves or neither
fails to load. The refusal is therefore impossible to lose to an edit
that widens the entry by accident.
Release facts — which images this tarball carries, what they hash to, and which compiler produced them — are the bundle's:
[bundle]
version = "0.8.0"
rustc = "rustc 1.97.1 (8bab26f4f 2026-07-14)"
[board.t114.app]
file = "t114/leviculum-t114-0.8.0.uf2"
sha256 = "..."
[board.t114.remedy.softdevice]
file = "t114/s140_nrf52_7.3.0_softdevice.hex"
license = "t114/s140_nrf52_7.3.0_license-agreement.txt"
convert = "hex-to-uf2"
The split is Codeberg #342. Before it, both halves were in the bundle
manifest, and run() loaded it before dispatching — so --set-time
and --set-telemetry, which read nothing but the USB IDs, refused to
start without a firmware bundle on disk. Activation is configuration,
not firmware (#236/#238); somebody pointing a node they already own at
an LXMF address was being sent to hunt for an image they had no use
for. The configure-only sessions now take the catalogue and never
locate a bundle at all; the flashing paths locate one and still fail
with no bundle found, naming every place they looked.
A bundle built before the split still loads: its board-fact sections
are ignored, and the catalogue in the binary reading them is the more
trustworthy of the two copies anyway, since binary and bundle ship
together. A bundle built before rustc was recorded loads too, and
lnflash says the compiler is unrecorded rather than refusing an image
it can still verify against its checksum.
rustc is the compiler that produced both the images and the flasher
beside them, and it is rustc --version as the bundle build ran it
rather than the channel rust-toolchain.toml names (Codeberg #305).
Whichever image a board is running, the first question behind "these
two boards behave differently" is which compiler built each of them:
this is embedded code with a measured 3264-byte stack-frame margin, and
codegen differences between compiler versions move frame sizes.
scripts/lnflash-bundle.sh asks both workspaces and refuses to write a
manifest claiming one compiler for a bundle built by two.
A new board still needs no new binary in the sense that matters — it is
data entry, in the catalogue plus one image. The license field is not
bureaucracy: it makes shipping a third-party blob without its licence
impossible by construction, which is exactly the mistake described
above. Board names stay identical to the firmware-side ones in
leviculum-nrf/src/boards/mod.rs:39, so that two namespaces never
diverge.
Identify in two stages, write only after
Before entering the bootloader we know only "some USB device". The
reliable identity exists only afterwards. The order is therefore:
find candidates, enter, confirm identity there, check
preconditions, check the checksum, and only then write. No write may
rest on a guessed identity. Commit 362c1c2d records why: a T114 image
once landed on a RAK4631 during bring-up. Several devices on the bus
must each be resolved individually rather than assuming "the one UF2
drive".
Confirm by reading back, do not infer
The stage above resolves identity before the write. Afterwards there is a second, separate question — which board actually received it — and a UF2 mass-storage volume gives no help with it at all: the volume carries no board serial, so a runner that finds one has no way from the volume alone to say whose it is.
uf2-runner.sh used to answer by pairing the volume with a candidate
from its own USB enumeration, which yields a board of the right type
and not the board that owns the volume. With two T114s attached, one in
DFU and one running, it wrote to the one in DFU and reported the other.
Measured twice on the rig, 2026-08-23 and 2026-08-24, naming the
opposite board each night — the answer follows enumeration order, which
is neither stable nor related to which board was in the bootloader
(Codeberg #343). The same gap made flash CONFIRMED mean "a board of
this type re-enumerated" rather than "the named board runs the named
image".
The fix is a read-back, and it turns attribution into a measurement:
- The firmware carries
leviculum_nrf::FW_BUILD_STAMP, one contiguous literalgit_sha=<sha> dirty=<bool>, and prints it after[FW_BUILD]at boot and every five seconds thereafter. - The runner greps that same literal out of the flat image it is
about to write (
tools/fw-readback.sh,fw_image_stamp). Not out ofgit rev-parse: that describes the working tree at the moment of the question rather than the bytes going to the board, and it cannot express a dirty tree at all, so two different images built from one commit would both answer with that commit. - After the copy it opens the candidate's debug port and requires the stamp to match. If the named board is carrying something else, the other attached candidates are asked, and the one that answers with the image is the one the summary names.
Three outcomes, kept apart on purpose, because they need different things done to them:
| Outcome | Meaning | Reported as |
|---|---|---|
| match | the named board answers with this image | flash CONFIRMED — serial=… reports …, read back from the board |
| mismatch | it answers with a different image | flash NOT CONFIRMED — serial=… reports <its stamp>, the image that was written is <ours> |
| no answer | no debug port, or nothing on it | flash UNCONFIRMED — serial=… did not answer on its debug port …; that is not the same as carrying the wrong image |
A board can legitimately fail to answer — crashed firmware, a port that
never appears — and silence must never be reported as wrong firmware,
nor as right firmware. When the write cannot be bound to any board at
all, the runner says exactly that (flash UNATTRIBUTED — … the runner does not know which board it wrote) and exits non-zero. A guess in that
position is what produced the ticket.
Two constraints the read-back has to respect, both long established on
the rig: the debug CDC transmits only with DTR and RTS asserted, so
a port opened without them is silent for reasons that have nothing to
do with its firmware; and the by-id symlinks are the stable
handle (-if00 debug, -if02 transport), because a /dev/ttyACM*
number is a position and moves between enumerations — trusting a
position for an identity is the defect itself.
tools/test-fw-readback.sh (just nrf-fw-readback) drives all of this
against stubbed boards, so it runs with no hardware.
The line has to be one the board said after the reset
Reading a window and keeping the last [FW_BUILD] in it is not the same
question as "what is running now". The port's input queue was filled
before the question was asked, and a line in it is an answer to a
question nobody asked — on 2026-09-09 that made lnflash report
1 of 2 board(s) confirmed running the firmware in this bundle.
1-1 (rak4631): not confirmed — Some(WrongBuild { saw: "ead0bce", expected: "daa8b8e" })
about a board whose own debug port said git_sha=daa8b8e seconds
later. ead0bce was the build that had been running before the flash.
A completely successful flash exited non-zero, which stops any script
that chains on it (Codeberg #378).
The confirmation therefore establishes a boundary the tool itself
creates, and only accepts what comes after it
(lnflash/src/verify.rs, fresh_banner):
- The board must have left the bus. If the bootloader is still
there when the wait times out, the board never rebooted into what was
written, and nothing a port says next is about the new image. That is
Absent, not a build claim. - The port's input queue is flushed on open — one
tcflush, the same one every control transaction does — so no line from the previous session can be read as an answer. - The debug port is resolved to its
by-idpath and the open is proved against the board's bus identity before a byte is read (entry::wait_for_interface_tty,flow::open_debug). A bare/dev/ttyACMnumber is a position, and on a multi-board run the board it moves to is the one still carrying the firmware this flash replaced. - Only a complete line counts. A half-read
git_sha=daa8b8eparses asdaa8and would be reported as a different build — a failure manufactured out of a partial read. - A line the firmware is quoting is not a line the firmware claims.
At boot our images re-emit the last ~2 KiB of the previous boot's
log out of retained RAM, each line wrapped in
[PERSISTENT_LOG](leviculum-nrf/src/bin/t114.rs, thepersistent_logblock). That replay includes the previous firmware's own[FW_BUILD]banner, so a board that has just booted into a new image says the old image's sha within milliseconds of the port opening — after the reset, after the flush, and still not about itself. The wrapper is what disqualifies such a line; its arrival time is not consulted. - A banner that is not the one we wrote does not end the read. The read is told which sha it is waiting for: a matching banner is an answer and stops it, anything else is kept as evidence while the read goes on. The board repeats its banner every five seconds, so the expected build gets another chance for as long as the budget lasts, and a stale line — from the replay or from anywhere else — gets exactly one. Only the last non-matching line, once the budget is gone, is reported as "the write did not take". A genuinely failed flash therefore costs the whole 15 s budget; the alternative is believing the first line on the port, which is what #372 did.
- The budget is three banner periods, 15 s. The firmware emits one
every 5 s, so a healthy board answers inside the first; the margin
covers a board whose banner task ticks just before the port opens and
a line the flush cut in half. A board still silent after that is
silent, and silence is
unknown.
This is the second false negative of the same shape and the reason the
rules above are structural rather than temporal. On 2026-09-27 the rig
flashed a T114 and a RAK from one bundle
(/home/lew/rig-run/boot-proof-flash.log, section === flash 01eb398b 2026-09-27T20:37:14):
3-2.3.4.4: back as leviculum T114 [1209:0001]
3-2.3.4.4: the board reports git_sha=abaea121f, not 01eb398b7, read as a
[FW_BUILD] banner line on …_T114_183004F712B4A7FE-if00
(/dev/ttyACM0), 0.0 s after that port was flushed.
The write did not take.
lnflash rc=1, 1 of 2 board(s) confirmed
The board's own capture has [FW_BUILD] git_sha=01eb398b7 … t=7448
92 seconds later: the write had taken. What lnflash read was the new
firmware quoting the old one (Codeberg #372).
The exit code carries the same three-way split, because "not confirmed" and "failed" need different things done about them:
| exit | meaning |
|---|---|
| 0 | every board was written and named the build in this bundle |
| 1 | the flash failed: a board did not come back, or named a different build, or nothing was written |
| 2 | every board took the write, none contradicted it, and at least one could not be read back |
Naming the old sha as if it were current is the defect. A confirmation
that cannot decide says unknown and never names a sha it did not
read.
Every build claim names where it was read
The rules above are what the confirmation does; they are not, by
themselves, evidence that it did it. On 2026-09-11 the same report came
back on a two-board run under a capture reader holding if00 on both
boards:
1-1: back as leviculum RAK4631 [1209:0002]
1-1: the board reports git_sha=b9b4a9c3, not de6e74ed. The write did not take.
1-2: running git_sha=de6e74ed. Done.
That line names a sha and nothing else, so the first question it raises — was that line read on this board's own port at all? — needed two capture files and a hand correlation, and ended undecided. A claim the operator cannot check is not much better than no claim.
So every build claim carries its provenance (verify::Source), and the
verdict prints it:
1-1: running git_sha=de6e74ed, read as a [FW_BUILD] banner line on
/dev/serial/by-id/usb-leviculum_RAK4631_DEC9947DAD9D2869-if00
(/dev/ttyACM3), 4.2 s after that port was flushed, the line
stamped t=7448 ms of board uptime. Done.
Four facts, each answering a question the bare sha left open: which
mechanism decided it (the banner read after the reset — the control
envelope carries no build query, so there is only one), which path was
opened and which node the fd was proved against, how long after the
flush the line arrived, and the t=<ms> the line stamped itself with.
The delay is the load-bearing number: a banner from a board that has
just booted arrives seconds in, so a sha delivered in the first
milliseconds was already in flight and the claim deserves that doubt.
The uptime stamp is what separates the two ways that can happen — a
fresh boot stamps its first periodic banner near 5000 ms, while a line
quoted out of retained RAM carries the stamp of a session that had been
up for hours (#372). Both are reported, neither is judged: a rule that
refused a line for its stamp would be guessing at how fast a board
boots.
just nrf-shellcheck (Codeberg #345) is the static half of the same
coverage: shellcheck -x over leviculum-nrf/tools/*.sh and
scripts/flash-lnodes-from-head.sh, which sources them. The scripts
carried # shellcheck directives long before anything ran them, and an
SC2034 and an SC2015 sat in the runner until #341 and #343 happened
to remove them. It runs from the repo root because the source=
directives name repo-relative paths.
The radio configuration belongs to the flash
A board that has just been written runs the compiled eu_medium
profile, and until the firmware learned to remember a configuration
(radio_store.rs) there was nowhere else for one to live: every host
that bound the board had to send the frequency again, and a standalone
LNode with no host had no way to be on anything else.
With the flash page in place the honest moment to choose is the flash
itself, once. lnflash therefore ends its sequence with a fifth step:
after the [FW_BUILD] banner confirms the write, it asks "Flash
default radio settings? [Y/n]". Enter takes the eu868 preset; "n"
opens a preset menu — eu868, us915, au915, custom — where custom is
the five-number field-by-field path. The choice goes to the board's
transport CDC; since #238 lnflash opens with a capability probe and
sends the configuration inside the control envelope
(docs/src/firmware/usb-control-envelope.md), falling back to the
legacy magic-prefixed frame lnsd still uses
(leviculum_core::rnode::build_radio_config_frame, HDLC-framed,
answered by RADIO_CONFIG_ACK) when the probe goes unanswered. Non-interactively,
--radio-preset <eu868|us915|au915> names a preset outright; it
cannot be combined with the --radio-* value flags (two ways to state
one configuration).
The presets are community profiles, not conformance claims
The names look like regulatory bands, which is why the menu spells out what they actually are: the settings each regional Reticulum community has converged on (the Reticulum wiki's "Popular RNode Settings"), with the regulatory situation documented next to them rather than implied by the name.
| preset | freq (Hz) | BW (Hz) | SF | CR | txpower |
|---|---|---|---|---|---|
| eu868 | 869463000 | 125000 | 8 | 4/5 | 22 dBm |
| us915 | 914875000 | 125000 | 8 | 4/5 | 22 dBm |
| au915 | 925875000 | 250000 | 9 | 4/5 | 22 dBm |
eu868 is the ReticulumNet consensus channel and the compiled firmware default. The channel sits in the band ERC 70-03 Annex 1 designates as h1.7 (869.4-869.65 MHz, 500 mW e.r.p.), so 22 dBm conducted stays lawful up to roughly 7 dBi of antenna gain, and the derived long-term airtime lock arms the ETSI 10 % duty cycle.
us915 follows the US community, and the tool prints a note when it is chosen because power conformance is not the whole story: FCC 15.247(a)(2) requires at least 500 kHz of occupied bandwidth for non-hopping digital systems, which a fixed-frequency 125 kHz node does not meet. The preset ships the profile the entire known US Reticulum scene runs; whether that is lawful for a given deployment is the operator's call, and the note says so at selection time, not in a footnote.
au915 matches the Western Sydney and Brisbane communities. Under the ACMA LIPD class licence (915-928 MHz, digital modulation, 1 W EIRP, no minimum bandwidth) 22 dBm plus a typical antenna sits far under the limit — sourced from two agreeing secondary references, as the primary ACMA text could not be retrieved.
eu433 is decided but deferred: ERC 70-03 allows 10 mW e.r.p. at 433.05-434.79 MHz, i.e. 10 dBm, and the SX1262 driver's lowest PA profile is 14 dBm. Offering the preset today would transmit 4 dB over the limit, so asking for it is refused with that reason until the driver can produce 10 dBm.
Three details are not obvious:
The step cannot fail the flash. It runs after the firmware is on the board and confirmed. A board that does not answer is a board running the compiled default, which is a warning and a re-run, not a failed flash.
Two of the seven wire fields are not the user's to state. The
preamble is derived from the PHY the way the RNode firmware derives it
(derive_preamble_symbols); a preamble belonging to a different SF
mis-prices airtime on both sides. The long-term airtime lock is sent at
the value the firmware would have derived for the chosen frequency,
because the frame's own presence (lt_alock_present) switches that
derivation off — sending zero would persist "no duty-cycle limit" onto
a board whose operator only picked a frequency.
Validation happens before the board is touched. --radio-sf 3 and
a bandwidth the SX1262 has no register code for are refused at the
command line. Left to the board, an unparseable frame is silently
dropped and looks exactly like a dead port.
An unavailable preset refuses the same way: --radio-preset eu433
stops the run with the 10 dBm reason before any board is enumerated,
rather than shipping a board 4 dB over the limit.
The real bottleneck is not the tool
A manifest invites the belief that the whole palette is a matter of
configuration lines. It is not. Our firmware supports exactly two
boards today, bsp-t114 and bsp-rak4631. A manifest entry without a
matching firmware build is an empty promise, and a LoRa board needs more
than an entry: pin mapping, TCXO voltage, SPI frequency and maximum
transmit power all live in BoardConfig and have to be right per device
and measured.
The structure should therefore follow the firmware side's growth rather than anticipate it. The value appears immediately and independently of it: a tool that identifies a board reliably, checks the SoftDevice precondition, and says "I do not know this board" instead of writing to it is what is missing today.
Deliberately not
No plugin system with shared libraries; it contradicts the statically linked binary. No scripting language in the manifest — such fields become a programming language within a year; when declarative data is not enough, the answer is a new transport in Rust. And no fetching firmware from the network: it contradicts "no infrastructure" and adds an attack surface to a tool that overwrites other people's devices with root privileges.
The structure proves itself on the second board, not the first. Building the UF2 transport for the T114 alone, but already split along the four axes and driven by the manifest, is only a claim until the RAK4631 — same transport, different board ID, different bootloader — runs through without a code change.
Verified on hardware
The factory path was exercised end to end on 2026-08-10, on the T114 with
serial 183004F712B4A7FE, using the tarball rather than the repo build.
A genuine factory state was reconstructed rather than waited for: the
6.1.1 SoftDevice was extracted from Meshtastic's combined hex and
trimmed to the window 0x1000-0x27000, which drops 11 MBR blocks and
135 blocks that would have written bootloader, bootloader settings and
UICR. Without that trim the bootloader rejects those blocks and the
whole copy fails; had it accepted them, it would have been the brick
path. The board then went dark exactly as predicted, and did not stay in
DFU — a physical double-tap was required, confirming the 2026-08-08
observation.
From there the tool ran unattended: it read SoftDevice 6.1.1 with
bootloader and flash agreeing, found >=7.0.1, <8.0.0 violated,
installed the SoftDevice (608 blocks, 11 declined), and then — the part
worth naming — the board rebooted into the old application, which now
booted because its base finally matched the installed SoftDevice. The
tool touched it back into the bootloader, re-read the version as 7.3.0,
and wrote the current application. That re-entry is the step a naive
implementation gets wrong by answering "wait for the bootloader" with the
pre-reboot sysfs entry.
That the application runs at all is the independent proof that the
SoftDevice is 7.3.0: an image based at 0x27000 cannot start on 6.1.1.
The board was confirmed afterwards over the debug port at
git_sha=d82ccfc with LoRa cycling normally.
Also confirmed in the same session: a board already sitting in its bootloader is handled without a redundant touch, several devices on one bus are resolved individually, and a RAK4631 on the same hub is neither offered nor written to, because it had no catalogue entry.
That last clause is now history rather than behaviour. Codeberg #261 gave the RAK4631 its entry, so a Pocket V2 on the same hub is found, hinted at its own board, brought into its own bootloader and written its own image — and a bundle that carries no image for it says so in those words rather than falling silent. The property the 2026-08-10 session actually demonstrated is the one that survives: each device is resolved individually, and no write rests on an identity the bootloader did not publish. The structure's own claim — "the second board proves it" — is what #261 collected on: the RAK reached hardware readiness through a catalogue entry, an image, and one board-list line, with no change to the transport, the identify step or the write path.
Open questions
- Does the touch handler exist in Meshcore, microReticulum or RNode firmware on nRF? Unmeasured; assume no, and fall back to the double-tap prompt.
- Nothing further on the SoftDevice. Provenance, version constraint and the two independent ways to read the installed version are settled above.
- Whether boards leaving the factory today still carry 6.1.1 is
unknown; the measured board is one unit from one batch. A tool must
read
INFO_UF2.TXTand decide, never assume a version.
OTA stage 1: entering BLE DFU on a command from the mesh
Status: concept, not scheduled.
Every firmware update this project has ever shipped ended with a hand on the board: a cable, or a needle in a pinhole (Flashing an LNode). That is fine for a board on a desk and wrong for the one on the roof, in the hedge, or two valleys away. Stage 1 splits the problem along the seam where it is cheapest to split it:
The decision travels the mesh. The bytes do not.
A command over Reticulum — any carrier, any hop count — puts the board into the bootloader's BLE DFU mode. Somebody then walks up to within Bluetooth range with a phone and pushes the image. Nothing about the image itself crosses the radio. OTA stage 2 is the document for when nobody can walk up at all, and it costs orders of magnitude more.
What the bootloader already does for us
All three board types carry a factory Adafruit nRF52 bootloader, and
that bootloader — not our firmware — owns DFU. Besides the UF2
mass-storage mode the flash runner uses, it has an over-the-air mode
speaking Nordic legacy DFU over BLE. Both modes are selected the same
way: a magic value in GPREGRET, the retained register that survives a
soft reset, read by the bootloader on the next boot.
We already drive one of them. The 1200-baud touch writes
DFU_MAGIC_UF2_RESET (leviculum-nrf/src/usb.rs:226) and resets, and
the bootloader comes up as a mass-storage volume. The OTA mode is a
different constant written to the same register at the same point by the
same code path.
The tree already knows the OTA mode exists, from a direction that had
nothing to do with flashing: memory.x reasons about it because in OTA
mode the bootloader enables the SoftDevice itself, whose RAM then reaches
up over the retained band (the RETAINED comment,
leviculum-nrf/memory.x:142, and Recovery,
step 2).
Open, and it must be closed by reading rather than remembering: the
exact magic byte. INFO_UF2.TXT names the bootloader build each board
carries (0.9.0 on the T114, 0.4.3 on the RAK — see
Flashing an LNode, "What the bootloader tells you"),
and the constant belongs to that build's source, not to anybody's memory
of the Adafruit tree. Writing a wrong magic is harmless — the bootloader
ignores what it does not recognise and the board boots normally — which
makes this cheap to settle on the bench and unacceptable to guess in the
document.
The control frame, and why it cannot be the one we have
We already have a control plane to the firmware: the #238 envelope, one
framing for every host command, with TYPE_RESET among them
(leviculum-core/src/envelope.rs:79). It is unauthenticated, and
deliberately so — read its header
(leviculum-core/src/envelope.rs, "Framed control envelope for the LNode
USB channel"): every argument in it is about framing and about not
colliding with Reticulum packets, and none is about who is allowed to
send this, because the answer was "whoever holds the cable". A frame
that arrives over the mesh has no cable behind it, so the envelope's
trust model does not travel with its wire format.
The pattern that does travel is already in this tree, on the propagation
node's control destination: the request runs over a Link, the remote's
identity hash is taken from that link, and it is checked against an
allow list before any verb is answered — with a typed refusal, not
silence, when it is absent or unlisted (answer_control,
lnpnd/src/engine.rs:988; the check itself at
lnpnd/src/engine.rs:1001). The list is operator configuration
(control_allowed, lnpnd/src/config.rs:532) and the node's own
identity is always on it. Codeberg #384 part 4 built that because the
reference has it; stage 1 needs exactly the same shape for a different
verb.
What a board needs on top of the daemon's version:
- The list has to survive a reset, so it is a flash record, and it
belongs in the band the bootloader declines to overwrite — the same
band the telemetry target, the radio config and the identity live in
(
leviculum-nrf/memory.x, the0xEA000–0xEC000pages). The telemetry target is the closest existing model for the record shape: magic, version, checksum, and a blank page decoding to "nobody" rather than to a garbage identity (leviculum-core/src/telemetry_target_store.rs). A board whose allow list decodes to "nobody" accepts no remote DFU command at all, which is the correct default for a board that was never configured. - A replay of a captured frame must not work. A Link is not a replayable object — it is established, identified on, and forward secret (Cryptographic identity and forward secrecy) — so requiring the command to arrive on an identified link already carries most of this. What it does not carry is the operator's own mistake: the same command sent twice. That is harmless here (a board already in DFU is not made worse by being told again) and it is worth stating rather than discovering.
- Nothing about the carrier. Whether the command arrived over LoRa, BLE or TCP, and over how many hops, is not the board's business (Interface isolation). A command that only works at one hop is a command that does not solve this page's problem.
What the board does on the command
Validate, acknowledge, write the magic, reset. In that order, and the order is the whole of it.
The acknowledgement has to leave before the reset, because in this scenario there is no second channel: on the cable the operator sees the port disappear and the volume appear, and over the mesh the ack is the only evidence the command was ever received. A board that resets first and acks never is indistinguishable, from the far end, from a board that never heard anything.
The write itself has a hot-path constraint that the UF2 path already
pays and that a mesh path pays in full: between the GPREGRET write and
sys_reset() there must be no await, no allocation and no logging, and
under a live SoftDevice the write must go through the SoC syscalls
rather than the register, or it records a bogus MWU panic per attempt
(leviculum-nrf/src/usb.rs:226-254, Codeberg #249). A mesh-borne
command always arrives with the SoftDevice enabled, so only the syscall
branch of that code is ever exercised — the branch the cable case takes
least often.
What the bootloader advertises, and what disappears
While the bootloader is in OTA mode, our firmware is not running. Our
BLE identity is gone, the Columba service is gone
(leviculum-nrf/src/ble/columba.rs:122), the LoRa interface is gone,
the board is off the mesh. To a phone it is a different device with a
different name, and to the mesh it has vanished. That is not a
side-effect to be engineered away; it is what DFU is, and it is the
reason the timeout question below is the important one on this page.
What it advertises — the service UUID, the device name, whether the address is resolvable — is a measurement, not a recollection, and this project has a sniffer and the tooling to take it (Bluetooth interfaces); nRF Connect on a phone answers it in one scan without any of that. Nothing downstream should be designed against a remembered UUID.
The open question: does DFU mode ever come back on its own
Open. It has two possible answers and they differ in how dangerous stage 1 is:
- It times out and boots the application again. Then a command sent
to a board nobody reaches costs one reboot and one gap in the mesh,
and the boot counter records it (
BOOT_COUNT n=… reset_reason=…, Codeberg #380). Stage 1 is safe to use on any board, including one that cannot be reached physically. - It waits indefinitely for a connection. Then a command sent by mistake — or one whose operator's plan changed — takes the board off the mesh until somebody physically resets it. On the roof that is a ladder; in the hedge it is a walk; two valleys away it is the end of that node until spring.
How to settle it, and it is one afternoon on the bench: put one
board into OTA mode, connect nothing, and watch the debug port. If the
application comes back, BOOT_TRACE and BOOT_COUNT say so on the next
boot and the elapsed time is the timeout
(leviculum-nrf/src/boot_count.rs:122, Codeberg #380). Hold it for at
least an hour before concluding there is no timeout; a short watch that
sees nothing has measured nothing. The bootloader's own source for the
build on the board is the second reading, and the two together are the
answer.
Until that measurement exists, the honest scope of stage 1 is boards somebody can still reach on foot — which is most of them, and already worth having, because reaching a board on foot with a phone is much cheaper than reaching it with a laptop and a cable.
Failure modes, and why an interrupted transfer is survivable
| What goes wrong | What the board is left as | Recovery |
|---|---|---|
| Command lost in the mesh | Running firmware, nothing happened | Resend |
| Command refused (not on the allow list) | Running firmware, typed refusal returned | Fix the list over the cable |
| Board enters DFU, nobody arrives | Bootloader, off the mesh | The open question above |
| Transfer starts and is interrupted | Bootloader, incomplete image staged | Reconnect and push again |
| Image is for the wrong board | Bootloader, refused or written and dark | Double tap, UF2, as today |
The fourth row is the one that makes stage 1 worth doing at all. A DFU
that dies halfway leaves the bootloader intact, because nothing in
this path ever writes the bootloader region: it sits above the
bootloader's own USER_FLASH_END and every mechanism here declines
writes there (Flashing an LNode, "What a UF2 is
allowed to write"). An interrupted transfer therefore produces a board
sitting in DFU waiting to be told again — not a dark board. A bad UF2
write, by contrast, produces a board that boots into nothing and
enumerates nothing, which is the state that costs a pinhole and a
needle.
What stays physical, unchanged. A board whose application has crashed answers no command, over any carrier, because nothing is running to answer it. The double tap remains the only trigger that works regardless of what is on the board, it needs a human, and on a Pocket V2 it needs a needle (Recovery, the hidden-pinhole caveat). Stage 1 does not shrink that set; it shrinks the set of routine updates that fall into it.
lnflash, and the phone that needs nothing from us
A phone running nRF Connect speaks Nordic legacy DFU today. It needs one
.zip package from us and no code at all, which means stage 1's useful
half — the control frame — is the only thing that has to be built before
the first remote update is possible.
Teaching lnflash the same trick is a separate and much larger decision.
It would need a BLE central stack (BlueZ through bluer or btleplug),
the legacy DFU control-point and packet characteristics, the init packet
and the package format — and, more expensively, a live D-Bus and a
running bluetoothd underneath. lnflash today is a single static
binary that shells out to nothing and calls syscalls rather than
programs, and that is a stated property of the bundle, not an accident
(lnflash/Cargo.toml, the libc dependency's comment;
Flashing an LNode). Adding a BLE central spends
that property. It may still be the right trade one day — for a fleet, a
phone in the loop is the bottleneck — but it is not free and must not
be smuggled in as an implementation detail of this page.
What stage 1 does not solve
Somebody has to stand within Bluetooth range of the board. Where that is impossible, the image itself has to travel the mesh, and OTA stage 2 measures what that costs — in hours of channel time, in hardware two of our three boards do not have, and in the one piece of code in this system whose bug is an unrecoverable board.
OTA stage 2: the image itself over the mesh
Status: concept, not scheduled. Conditional on hardware two of our three boards do not have.
OTA stage 1 moves the decision over the mesh and leaves the bytes to somebody standing within Bluetooth range. This page is for the board where nobody can stand: the image travels the mesh too, as a signed Reticulum Resource to the board's control destination, is staged on external flash, and is written into the application window by a copier that runs from RAM. A golden image and the boot counter undo it when the new image does not come back.
Two numbers decide almost everything below, and neither is negotiable: the application window is smaller than two images, and a 620 KB transfer costs hours of a channel everybody else is also using.
Why internal flash cannot do it
The nRF52840 has 1 MB of internal flash and it is already spoken for.
The SoftDevice sits at the bottom, the bootloader at the top, and what
is left is the application window: origin 0x27000, length 0xB2000 —
729 088 B, 712 KiB (FLASH, leviculum-nrf/memory.x:84). Above it the
boot record takes a page (BOOT, leviculum-nrf/memory.x:108, Codeberg
#380) and the record store sixteen more (STORE,
leviculum-nrf/memory.x:91, Codeberg #384), then the three persistence
pages.
The image that lnflash wrote to a T114 on 2026-09-24 was 2 419 UF2
blocks of 256 payload bytes: 619 264 B, 605 KiB. A staged second copy
beside the running one needs 1 238 528 B, 1 210 KiB. That does not fit in
712 KiB, and it does not fit even if the boot record and the record store
are given up and the whole writable window above the SoftDevice is
claimed: 0x27000 to 0xEA000 is 798 720 B, 780 KiB
(Flashing an LNode, "What a UF2 is allowed to
write") — still 440 KB short of two banks.
So there is no A/B bank internally, and the alternative — erase the
running image and write the new one from a transfer held in RAM — is not
an alternative at all. It fails on two counts: the board has 96 KiB of
heap (HEAP_SIZE, leviculum-nrf/heap-budget/src/lib.rs:54) against a 605 KiB
image, and a power cut during the write leaves a board with no image and
no copy of one. Staging on flash the copier does not erase is not an
optimisation; it is the property that makes the operation survivable.
Which boards are in, and on what evidence
T114 and RAK4631: out. Neither carries QSPI flash. That is measured,
not read off a header — three units answered nothing to 05h, 9Fh,
90h or the datasheet reset while every pin followed our drive, and both
board files state qspi_part: None
(leviculum-nrf/src/boards/t114.rs:184,
leviculum-nrf/src/boards/rak4631.rs:199). The EXTERNAL_FLASH_DEVICES
lines in the vendor variant headers that once suggested otherwise sit
under comments denying the part — RAK's own reads "No onboard flash"
(leviculum-nrf/src/boards/rak4631.rs:152-153) — and
leviculum-nrf/src/qspi.rs carries the whole account. For these two
boards stage 1 is the entire answer, and that is why the two stages are
separate documents rather than two halves of one.
SolarNode: conditional, and the condition is a boot line. The XIAO
nRF52840 module's CONFIG.qspi_part is Some(&crate::qspi::P25Q16H)
(leviculum-nrf/src/boards/solarnode.rs:402), 2 MB
(P25Q16H, leviculum-nrf/src/qspi.rs:435). But that part number comes
from Seeed's variant headers and the wiring from Seeed's schematic, and
the schematic for the plain XIAO v1.1 draws the same footprint marked
DNP — do not populate. So the firmware does not assert the part, it
asks: identify_at_boot (leviculum-nrf/src/qspi.rs:463) reads the
JEDEC id once at boot and prints [QSPI] JEDEC … state=ok or does not.
Our SolarNode prints that line, and the part has been driven. On
2026-09-24 the unit answered [QSPI] JEDEC id=85:60:15 … match=1 state=ok (/home/lew/rig-run/solarnode-dfu/dfu-test.log). On
2026-10-02 the destructive self-test (qspi-selftest, single-line bus
at 32 MHz) ran over the whole 2 MB of the PUYA P25Q16H
(/home/lew/rig-run/solarnode-qspi/qspi-selftest-20261002T134004Z.log):
all 512 sectors erased clean in 8.8 s per pass, slowest sector 17 ms;
2 MB programmed in 13.4 s, 152 KiB/s with single-line page program;
2 MB read back in 0.55 s. The read-back is red: 11 445 mismatching
bytes against the first pattern and 7 012 against the second
(RESULT pass=0 reason=mismatch), while every erase verified all-0xFF.
The error is the 32 MHz read, not the part and not the program
(Codeberg #435). On 2026-10-04 the self-test wrote each pattern at one
clock and read it five times at each
(/home/lew/rig-run/solarnode-qspi/qspi-selftest-20261004T211914Z.log):
written at 32 MHz and read at 8 MHz, and written and read at 8 MHz, not
one byte was wrong; read at 32 MHz the same data came back wrong in
1 644 555 and 1 743 635 of 2 097 152 bytes, 99.8 % of them different from
one read to the next, on all eight bit positions, with 0-to-1 flips 130
to 170 times more frequent than 1-to-0. Data the reads cannot agree on
is not data the cells hold, and the 2026-10-02 red was the same
artefact. The part allows the clock: the P25Q16H datasheet gives
104 MHz for FAST_READ (0Bh), the opcode the firmware uses. What does
not hold is the nRF52840's input sampling delay as embassy-nrf sets
it, IFTIMING.RXDELAY = 2: 31.25 ns after the SCK edge.
The read sweep of 2026-10-05 measured the eye
(/home/lew/rig-run/solarnode-qspi/qspi-selftest-20261004T230630Z.log,
build 8c52c797). Pattern B, programmed at 8 MHz, read five times per
chunk at the 8 MHz control and at 16 and 32 MHz under every RXDELAY;
wrong bytes of 2 097 152:
| SCK | RXDELAY | wrong | unstable | stable | reading |
|---|---|---|---|---|---|
| 8 MHz | 2 | 0 | 0 | 0 | control |
| 16 MHz | 0, 1, 2 | 0 | 0 | 0 | clean, three steps wide |
| 16 MHz | 3 | 1 969 638 | 1 956 751 | 12 887 | edge |
| 16 MHz | 4, 5, 6 | 2 089 026 | 0 | 2 089 026 | stable wrong, sampled one bit late |
| 16 MHz | 7 | 2 090 781 | 1 977 264 | 113 517 | edge |
| 32 MHz | 0 | 2 088 639 | 1 896 695 | 191 944 | edge |
| 32 MHz | 1 | 0 | 0 | 0 | clean, one step wide |
| 32 MHz | 2 | 1 915 318 | 1 900 523 | 14 795 | edge, the embassy-nrf default |
| 32 MHz | 3 | 2 089 026 | 0 | 2 089 026 | stable wrong, one bit late |
| 32 MHz | 4, 5, 6 | about 2 090 000 | mixed | wrong |
"One bit late" is lost1 = gained1 = 20 979 770: the whole image
shifted by one bit. RXDELAY 7 at 32 MHz is not in the capture, which
closed after point 15. At 16 MHz the clean eye is three RXDELAY steps
wide (0 to 31 ns), at 32 MHz one step (15.6 ns), and the default delay
sits on the falling edge of the 32 MHz one: that is the whole of #435.
The firmware therefore reads the part at 16 MHz, RXDELAY 1, the
middle of the only clean run at least three steps wide, one constant
(P25Q16H_BUS, leviculum-nrf/src/qspi.rs). 32 MHz at RXDELAY 1 is
faster and has no margin on either side, and temperature and supply move
a 15 ns eye by more than that. The self-test's SWEEPBEST line applies
the same rule (MIN_CLEAN_RUN, leviculum-nrf/qspi-selftest/src/lib.rs)
and names a faster clock with only a narrower eye as margin=too-narrow.
At 16 MHz a full 2 MB read takes about 1.05 s instead of 0.55 s; for a
605 KiB staged image that is 0.3 s per verifying read.
The self-test's RESULT line judges the part, not the bus (Codeberg
#413): it is green iff both 8 MHz read sets match the pattern, every
erase verified all-0xFF, and SWEEPBEST found a wide window, and names
the failing one as reason=program, not-erased or no-wide-window.
The 32 MHz read sets stay in the plan, since they measure the default
timing's eye, and their bytes are reported as read_side_mismatches=.
The second sweep run of 2026-10-05
(/home/lew/rig-run/solarnode-qspi/qspi-selftest-20261005T005318Z.log)
printed RESULT pass=0 reason=mismatch mismatches=3523349 on a part
that held both patterns; under this rule it reads RESULT pass=1 reason=ok mismatches=0 read_side_mismatches=3523349 not_ff=0.
What this means for stage 2: staging plus golden fits the part with room to spare (below), and programming is not the constraint, since a 605 KiB image takes about 4 s to write and about 2.6 s to erase its 152 sectors. The constraint is the transfer, about 4.9 h per image per hop at our default PHY under the 10 % duty cycle (below). And no image is trusted from this flash until its read-back is green, and the read-back is green only at a bus timing the self-test has shown clean: a part that hands back other bytes than it was given turns every signature check over the staged copy into a coin toss, and a golden image read back wrong is a rollback to something nobody built.
The budget on a 2 MB part, and the other claimant
Staging plus golden is 1 238 528 B of 2 097 152 — 59 % of the part —
leaving about 838 KiB. That is enough, and it is not so much that the
region layout can be left implicit, because the same part is already
wanted by something else: the record log, the message store of Codeberg
#384, mounts over the whole part today
(log_store, leviculum-nrf/src/qspi.rs:1054, read-only and formatting
nothing, precisely because that decision had not been taken). Two
claimants and one part means one region map, decided once, in one place —
not two mounts that each believe they own sector 0. Whichever batch
first writes to that part owns the map, and stage 2 must not be that
batch by accident.
Erase granularity is 4 KiB (ERASE_SIZE, leviculum-nrf/src/qspi.rs),
so staging an image is 152 sector erases. Against the part's rated
endurance that is free; it is the pattern that matters, not the count —
a staging region rewritten in place from sector 0 every time wears one
end of the part and nothing else, which is the same argument the record
log already makes for itself
(An LXMF propagation node on a board).
The image arrives as a Resource and is never held
The image is a Reticulum Resource on a Link to the board's control
destination — stock Reticulum, the same mechanism rncp and lncp use.
The board writes each part to the staging region as it arrives and holds
none of the image in RAM.
That is not a preference. The board has 96 KiB of heap and the cost of
holding a copy of anything is measured: a 5 427 B propagation response
held 51 662 B live before Codeberg #384 B1 and 29 842 B after it, and
before either, a board died on a 5 446 B allocation between serving a
request and answering it — PN_GET … bytes=5376 as the last line of one
boot and PANIC_PMRT … "memory allocation of 5446 bytes failed" on the
next (leviculum-nrf/heap-budget/src/lib.rs:177-202, pinned by
leviculum-std/tests/mvr/pn_serve_peak_outgrows_the_board_heap.rs). A
path that holds one whole copy of 5 KB killed a board; 605 KB is not a
question of tuning.
Two consequences follow and both are design constraints, not details:
- The transfer must be resumable across a reboot. Five hours of channel time (below) is far longer than the interval at which a field board reboots for its own reasons. The staged region and its header are the resume state; a transfer that starts from zero after every reset never completes on a board that reboots at all.
- The staging region is untrusted until the whole image verifies. Partial contents are exactly what an interrupted transfer leaves, and they must be indistinguishable from garbage to everything downstream.
Header and signature: refuse before the first erase
Ahead of the staged copy sits a header, and it is checked in full before the copier touches the application window:
| Field | Why it is there |
|---|---|
| magic + format version | A staging region holding a foreign or half-written thing reads as foreign, the same argument the boot record makes for its own magic (leviculum-nrf/src/boot_count.rs) |
| board family | An image for another pinout family bricks this board. The flash runner already refuses a wrong-SoftDevice board and a wrong UF2 volume before writing (nrf-sd-guard, Justfile:234); this is the same refusal, without an operator to read it |
| image length | Bounds the copy, and is what "the transfer is complete" is decided against |
| version | What the board says it is running, and what a rollback is a rollback from |
| hash over the image | Catches the interrupted transfer and the bad sector |
| Ed25519 signature over the header and the hash | Catches everything else |
ed25519_dalek is already a dependency of the core and builds for the
firmware target (leviculum-core/src/identity.rs:63-65), so verification
costs a public key compiled into the image and one pass over the staged
bytes. That pass is cheap against the five hours the transfer took.
Why a signature, when stage 1's allow list already says who may command. The allow list authenticates a peer on a live link. The staged image outlives that link, the reboot, and possibly the operator's key: what the copier reads at 3 a.m. after a power cut has no link behind it and no peer to ask. The allow list says who may command; the signature says what may run. Neither substitutes for the other.
And the order is load-bearing. A signature checked after the erase is not a check — by then the board has nothing to fall back to but the golden image, which turns a refusable mistake into a rollback. Verify, then erase.
The copier, and what it must never depend on
One routine, running from RAM with the SoftDevice disabled: erase
0x27000 for the image's length, write from the staging region, verify,
mark done, reset.
It must not depend on:
- the image it is erasing — no call into it, no vector table in it, no panic handler in it, no string in it;
- the SoftDevice, which owns flash timing while it is enabled and
which the boot record already steps around by writing before
Softdevice::enable(leviculum-nrf/src/boot_count.rs); - the heap, USB, BLE, the radio, the log path — anything that might route through a region being erased or an allocator that might fail;
- interrupts it did not itself arm;
- the transfer that produced the staged copy, which finished hours or reboots ago.
It must be idempotent under a power cut, because that is the failure
it exists to survive: a cut halfway through leaves a partly written
application window, and the only correct behaviour on the next boot is to
copy again from the same source. That means "a copy is in progress" is
itself a flash record, written before the first erase, in a page the
copier never erases — the boot record's neighbourhood
(BOOT, leviculum-nrf/memory.x:108) is the shape, for the same reason
the boot record chose it: it must survive a power loss, which is the case
it exists for.
This is the one piece of code in the system whose bug is an
unrecoverable board in a place nobody can reach. It should be the
smallest, dullest, most heavily host-tested thing we own — the same
disposition leviculum_boot_count took, where a host test can cut the
power at every word boundary (leviculum-nrf/src/boot_count.rs).
Golden image and the rollback trigger
The golden image is the last one that came back healthy, kept on the same
part. The trigger for restoring it is the boot counter that landed on
2026-09-24 (record_at_boot, leviculum-nrf/src/boot_count.rs:122,
Codeberg #380): one 16-byte record per boot with the raw RESETREAS and
whether retained RAM survived, appended to the BOOT page, one erase per
256 boots. It exists because a Pocket V2 restarted twice on a field walk
and every boot read reset_reason=0x00000000 with prev_magic=absent —
a power loss takes retained RAM with it, and a rollback trigger that
lives in RAM is a rollback trigger that a sagging battery switches off.
The rule's shape: the copier marks the new image unproven; the application clears the mark once it reaches a defined healthy point; N consecutive boots with the mark still set restores the golden image.
Two parts of that are decisions this page deliberately does not take, and naming them is more useful than guessing them:
- What "healthy" means. It must be later than "reached
main" — a board that boots and cannot bring up its radio is precisely the failure that needs rolling back, and it reachesmainevery time. It must not be "heard a peer", or a quiet mesh looks like a broken image and the board rolls back a perfectly good update because nobody was talking. - What N is. Too small and one unlucky brown-out undoes a good update; too large and a boot-looping board spends a day looping before it heals.
The counter is also the only evidence anybody will ever get from a board
in a hedge: BOOT_COUNT n=… reset_reason=… retained=… since_erase=…,
read on the next visit or over the mesh, says what the board did while
nobody was watching.
What it costs the channel, in hours
Measured on 2026-09-24, lora_lncp_push_to_python_50kb on the rig:
five 50 KiB pushes at 869.525 MHz, SF7, BW 62.5 kHz, CR 4:5, 212–333 s
of wall clock each with the airtime lock deliberately switched off for
the cell, and about 165 s of transmitter airtime each — 109 parts of
491 B at roughly 1.50 s of airtime apiece.
Scaling to the 619 264 B image is a factor of 12.1:
| at the cell's PHY (SF7/BW62.5) | at our default PHY (SF8/BW125) | |
|---|---|---|
| coded rate | 2 734 bit/s (342 B/s) | 3 125 bit/s (391 B/s) |
| transmitter airtime for one image | ~2 000 s (33 min) | ~1 750 s (29 min) |
| wall clock with no duty limit | 43–67 min | ~40–60 min |
| wall clock under the lawful cap | ~5.5 h | ~4.9 h |
869.463 MHz — our default carrier — sits in the 869.4–869.65 MHz
sub-band, whose lawful duty cycle is 10 %
(etsi_eu868_duty_cycle, leviculum-core/src/rnode.rs:1525;
Regulatory airtime). Ten percent of an hour is
360 s of airtime, so 1 750–2 000 s of airtime is five hours of wall clock
however fast the modem is. The PHY does not change the answer; the
regulation does.
And it is worse than "five hours", because that budget is not spare capacity:
- The sending node spends its entire lawful hour on the image, for five hours. It forwards nobody's traffic while it does. A transport node updating a neighbour stops being a transport node for an afternoon.
- Every relay on the path pays it again, once per hop, and pays it serially. A two-hop image is ten hours of two nodes' budgets.
- Everyone else on the channel pays too. The band is occupied ten percent of the time, continuously, for hours; every other node's pre-transmit window finds a busy channel more often and backs off (The randomised pre-transmit window).
Compression halves it at best. The measurement above is deliberately
incompressible — /dev/urandom, so that 50 KiB of payload is 50 KiB on
the air and the durations compare — while a real ARM firmware image does
compress, though less than a halving: xz -9e takes the built image
from 619 124 B to 379 608 B, 61 %, so about 4.9 h per hop becomes about
3 h. That is also not small enough for an internal A/B on the boards
without external flash. It changes the scheduling of the operation and
not its nature.
The arithmetic forces the conclusion, and the conclusion is the point of the page: stage 2 is for the board nobody can reach at all. An update over the mesh is an event the mesh is told about in advance, planned around, and done once. A fleet update over LoRa is not a thing that happens. Wherever somebody can get within Bluetooth range, stage 1 does the same job for the cost of one control frame.
Python-RNS compatibility: untouched
Nothing here is on the wire between stacks. The transfer is a Resource on a Link to a destination resolved the ordinary way, which is stock Reticulum in both directions — a Python peer could be the sender without knowing what it is sending. The header, the signature, the staging layout, the copier and the rollback rule are all behind our own control destination, visible to nothing but us, in the same sense the propagation node's control destination is (Python-RNS compatibility). No new packet type, no announce semantics, no change to any field a Python node reads.
Python-RNS Compatibility
Leviculum is built to live in the same mesh as Python Reticulum
(rnsd) and to be a drop-in replacement for the daemon and its
tooling. Compatibility is pursued at two distinct levels, and one
thing that is not pursued at all.
Level 1: wire and semantic compatibility
The protocol the two stacks speak must be identical on the air. The exact bytes of identities, destinations, announces, packets, and links are fixed by the Reticulum specification; the message format layered on top is fixed by the LXMF specification. Leviculum implements those formats so that a Python peer cannot tell a Leviculum neighbour from another Python node.
Semantic compatibility goes beyond byte layout: behaviours a Python peer expects from a neighbour — answering path requests, rebroadcast decisions, link lifecycle, ratchet handling — must still be delivered. Where the precise expected behaviour matters and is subtle, it is captured as a source-of-truth reference; the broadcast path is documented in Broadcast: Python-RNS parity reference, which records what Python does for every broadcast mechanism so the Rust core can match it.
Semantic compatibility is decided field by field: what a peer decides from a value we generate is part of the contract, even when the byte layout is right. The audit method and testing rule for that are in Wire Field Semantics.
It cuts both ways. The reference's accept set is as much of the
contract as its output set, and it is wider — Python's decoder takes
forms Python's encoder never emits. A mesh has more than two
implementations in it (reticulum-kt, microReticulum, hand-rolled
senders), and any of them may pick a legal encoding Python happens not
to use. Testing only against rnsd cannot see this: it never produces
the form we refuse. Being stricter than the reference on the read
path is a compatibility defect of the same standing as emitting a
wrong value, and it fails silently — see
Wire Field Semantics.
Level 2: drop-in daemon and tooling
lnsd shares two interfaces with Python's rnsd:
- The shared-instance IPC socket. A running daemon exposes a local
control/data channel that client tools connect to. Leviculum speaks
the same protocol, so Python's
rnstatus,rnpath,rnprobe, andrncpdrive a runninglnsdwithout modification, and the Leviculum toolslnstestandlncpdrive a runningrnsdjust the same. The RPC control channel that backsrnstatus/rnpath/rnprobeis implemented inleviculum-std/src/rpc/(it speaks Python'smultiprocessing.connectionframing with pickle payloads, seerpc/connection.rsandrpc/pickle.rs). A client on that socket —lxmf-node,lnmsg,lnomad, everyln*tool — keeps its own known-destination table rather than delegating it, so a peer it has used stays recallable forKNOWN_DEST_USED_LINGER_MS(leviculum-core/src/constants.rs:176) after the daemon stops offering a path to it, instead of being forgotten on the next sweep (Codeberg #389). - The config-file format.
lnsdparses the same INI-style config thatrnsduses (leviculum-std/src/config.rs,leviculum-std/src/ini_config.rs). Even keys Leviculum does not act on are parsed for compatibility — for exampleshared_instance_typeandshared_instance_socketare read and honoured per RNS 1.3.x semantics so an existingrnsdconfig works unchanged (leviculum-std/src/config.rs:76-82).
This drop-in property is a deliberate design goal, not an accident.
It is also what makes honest A/B testing possible: the test harness
points the same client binary (e.g. lnstest selftest) at either
daemon, never a parallel per-stack driver. A parallel driver would
smuggle configuration differences into what claims to be a stack
comparison.
What is explicitly not a goal: internal parity
Compatibility is not the same as parity.
- Compatibility — our stacks interoperate at the wire and semantic level.
- Parity — our internals mirror Python's (same algorithms, same retry timings, same state-machine structure).
Leviculum needs the first, not the second. The historical parity
documents under docs/src/architecture-*-python-parity.md are
reference material for getting behaviour right, not commitments to
maintain identical internals.
The deviation rule
A deviation from Python-RNS's implementation is acceptable if and only if all three of the following hold:
- Wire-format compatibility is preserved.
- Semantic compatibility is preserved (behaviours Python peers expect from a neighbour are still delivered).
- The deviation measurably improves robustness or mesh delivery.
"Because Python does it differently" is not, on its own, an objection; only "this breaks wire or semantic compatibility" is. The interface-isolation design — interfaces applying their own jitter, CSMA, and airtime budgeting — is a deliberate deviation that satisfies this rule.
A deviation that is not written down is indistinguishable from a bug. Each one is pinned here with the reference line it departs from, so the next reader can check the claim instead of re-deriving it.
Pinned deviation: a pathless never-used destination does not linger
The reference's known-destination sweep spares a pathless entry on three
grounds: an application pinned it, it was used within
DESTINATION_TIMEOUT * 1.25, or it was never used but announced within
UNUSED_DESTINATION_LINGER — 6 minutes
(reference/Reticulum/RNS/Identity.py:349-352, the two timeouts at
Transport.py:91-92). We implement the first two and not the third
(leviculum-core/src/memory_storage.rs:1226): our announce cache stores
the raw announce, not the moment we heard it, so there is no age to
compare against. The third arm only ever protects a destination nothing
has asked about, so dropping it a few minutes early costs a path request
rather than a fact, and clause 3 of the rule is satisfied by the smaller
resident set. EmbeddedStorage goes one further and implements only the
first (leviculum-core/src/embedded_storage.rs:1103), because a board
tracks no use-state at all.
Pinned deviation: ingress-control default on dial-out links
The reference gives every interface ingress_control = True
(Interface.py:112), overridable per interface by the config key of
the same name (Reticulum.py:768-769, applied at Reticulum.py:910).
Leviculum defaults it off on dial-out point-to-point links —
TCPClientInterface, BackboneClientInterface, UDPInterface, and an
I2PInterface without connectable — and leaves it on everywhere
else, including every listener
(ingress_control_default_for_type, leviculum-std/src/config.rs:947).
Against the rule: the flag decides only whether we hold incoming
announces, so no wire byte and no behaviour a peer observes changes
(1 and 2). It gains us the announces the limiter would otherwise hold
silently on a link carrying one known peer's startup burst — the
mechanism behind the Codeberg #44 flake, on our receive side (3).
An operator who wants the reference behaviour writes
ingress_control = yes on the interface.
The default is a role distinction, not a medium one: an interface
that accepts connections from arbitrary unknown peers is exactly the
announce-storm surface the limiter exists for, so a listener keeps the
reference default. What a listener resolves is inherited by every
connection it accepts, as in the reference
(TCPInterface.py:582, I2PInterface.py:951,
BackboneInterface.py:409). Shared-instance IPC clients are never
ingress-limited on either stack — the reference hard-wires
should_ingress_limit to False for them
(LocalInterface.py:137-138) — so that is not a deviation.
Pinned deviation: an absent txpower is the board maximum
The reference resolves an omitted txpower key to 0 dBm
(RNodeInterface.py:153: int(c["txpower"]) if "txpower" in c else 0).
Leviculum resolves it to the board maximum, capped by the lawful
e.r.p. limit for the configured frequency: 22 dBm — the ceiling of
the SX1262 high-power PA and the highest value an RNode-firmware board
takes before clamping — or the sub-band's limit from ERC 70-03,
whichever is lower (rnode::resolve_tx_power and
DEFAULT_TX_POWER_DBM, leviculum-core/src/rnode.rs:742, capped by
lawful_erp_dbm, applied in both interface builders and in the
SerialInterface LNode path). The standalone LNode firmware's
compiled profile carries the uncapped board maximum
(RadioConfig::eu_medium, leviculum-nrf/src/lora.rs:414-460), which
is the capped resolution's own result at that profile's 869.463 MHz.
Against the rule: TX power is a local modem setting. It is never on the
wire, and no peer — Python or otherwise — learns or expects anything
about a neighbour's transmit power, so (1) and (2) are untouched. What
it gains (3) is the whole failure mode: 0 dBm is 1 mW, and a 1 mW node
has no symptom at the node. It boots, configures, transmits, logs
nothing unusual, and is simply not heard. 22 dBm is 158 mW. An operator
who wants 0 dBm writes txpower = 0 and gets 0 — the resolution keeps
None and Some(0) distinct, the same way an explicit
airtime_limit_long = 0 beats the derived lawful default.
Because the request is not preceded by a capability probe, a board
whose maximum is lower answers by clamping and echoing the clamped
value (RNode_Firmware/RNode_Firmware.ino:861-879 — 17 dBm on an
SX127x, PA_MAX_OUTPUT on an SX1262 with an external PA). Confirmation
is otherwise an exact match on both stacks (ours at
leviculum-std/src/interfaces/rnode.rs:1874, the reference at
RNodeInterface.py:677), so the derived default — and only the derived
default — accepts a confirmation below what it asked for, logs the
board's ceiling, and runs. An explicitly configured power keeps the
strict check: a board that cannot deliver a value the operator chose
must say so. A confirmation above the request is a mismatch either
way.
Regulatory note (EU 863-870 MHz). 27 dBm e.r.p. is permitted only
in 869.4-869.65 MHz (ERC Recommendation 70-03, Annex 1, sub-band
h1.7); every other listed European sub-band allows at most 25 mW
e.r.p. = 14 dBm. That is why the derived default is capped by
frequency (rnode::lawful_erp_dbm): a node on a community frequency
like 867.2 MHz (Rotterdam/Duffel), 867.5 (UK), 868.0 (Bern) or 868.2
(Madrid) resolves an absent txpower to 14 dBm, not 22. An explicit
txpower wins even above the cap — the operator may hold a licence
or sit in another jurisdiction — and the excess is logged. Inside
869.4-869.65 MHz, 22 dBm conducted stays under the 500 mW limit up
to roughly 7 dBi of antenna gain (22 + 7 - 2.15 dBd ≈ 26.9 dBm
e.r.p.); above that the operator has to set txpower down
explicitly, and that residual is documentation, not a runtime
warning: the stack does not know what antenna is attached, and a
warning it cannot condition on anything is a warning operators learn
to ignore. A carrier overlapping one of the narrowband alarm bands
between the wideband sub-bands is named in a warning at interface
build (rnode::erp_band_gap) and then transmitted: a 125 kHz signal
cannot meet their ≤ 25 kHz channel spacing on any power, but the
judgement is the operator's, not ours — see No radio configuration
is refused.
Not a deviation: a class constant is not the value on the wire
An interface's HW_MTU in the reference is a class attribute that
looks like the answer and is not it. TCPInterface.HW_MTU = 262144
(TCPInterface.py:42) is only what the interface carries into
interface_post_init (Reticulum.py:879), which immediately runs
optimise_mtu (Interface.py:198) over the interface bitrate for
every interface with AUTOCONFIGURE_MTU set — TCP among them. With
TCPServerInterface.BITRATE_GUESS = 10 Mbps
(TCPInterface.py:453) the derivation lands on 8192 up to and
including Reticulum 1.5.0, and on 16384 from 1.5.2 on, where the
thresholds became >=. The class value never reaches the wire on
either.
We read the class constant instead of the derivation until Codeberg #355. The measurement that closed it drove the same two Python clients, attached to the shared instances of two relays, over one TCP hop between the relays, swapping only which daemon the relays were:
| relay 1 | relay 2 | negotiated link MTU | largest single packet |
|---|---|---|---|
rnsd 1.3.5 | rnsd 1.3.5 | 8192 | 8111 |
rnsd 1.5.2 | rnsd 1.5.2 | 16384 | 16303 |
lnsd (before) | lnsd (before) | 262144 | 262063 |
lnsd (after) | lnsd (after) | 16384 | 16303 |
So the constant acted on the wire: a Python client on our shared
instance negotiated a link MTU 32x larger than the same client gets
from rnsd, and the frames actually crossed the TCP hop at that size.
A Python peer on the path clamps a too-large signalled MTU down to
its own on the hop it carries — the mixed rows of the same measurement
settle on 8192 — so nothing broke as long as one was there to do it.
That conditional is the semantic-compatibility risk: a Reticulum 1.5.x
peer's receive path rejects a frame longer than its own HW_MTU, so any
route change onto such a peer silently drops the traffic. Speed is
Priority 2 and does not buy that.
The lesson generalises past MTU: before adopting a reference class
attribute as a value we signal, check whether the reference derives it
at interface post-init, and whether it signals it at all.
UDPInterface was the sibling case, closed by Codeberg #357. It sets
self.HW_MTU = 1064 (UDPInterface.py:74) and leaves the base
class's AUTOCONFIGURE_MTU = False and FIXED_MTU = False
(Interface.py:93-94) alone, and every gate that puts an MTU on the
wire reads those flags rather than the value:
- the initiator asks
Transport.next_hop_interface_hw_mtu, which returnsNonefor such an interface (Transport.py:2682-2683), so it signalsRNS.Reticulum.MTU(Link.py:310-314); - a relay forwarding onto such a next hop truncates the link request
by
LINK_MTU_SIZE(Transport.py:1599-1602); - a receiver clamps against
RNS.Reticulum.MTUrather thanHW_MTU(Transport.py:2101-2104).
All three land on 500. 1064 is what the interface's own read path
accepts off the wire, never what it negotiates. We signalled it until
#357, which is why our own interop suite carried two UDP numbers: 500
for every link with a Python end on it and 1064 between two Rust
ends. HW_MTU in InterfaceInfo now means the value the interface
signals, so a UDP interface carries None there and reports nothing
under the mtu stats key — a key that postdates our pinned 1.3.5
reference, where Reticulum 1.5.x reports 1064.
The reference gates several more interfaces off the same way —
PipeInterface, KISSInterface, AX25KISSInterface,
SerialInterface, I2PInterface, RNodeInterface and
RNodeMultiInterface all set an instance HW_MTU without either
flag — and we still signal ours on each. Those are not #357: LoRa in
particular has a Priority-1 argument for keeping the link inside one
508-byte frame that the UDP case has no counterpart to, so each wants
its own measurement.
Same-interface relay on shared media
Path-directed transport forwarding transmits on the next-hop interface even when it equals the receiving interface. Same-interface relay is NOT suppressed — it is how multi-hop works on one shared medium. In the fundamental single-channel LoRa topology A-B-C (A↔B and B↔C in range, A↮C), B's only route to A is back out of the very interface C's packet arrived on; a relay that declines that hop kills every data flow the announce flood just made possible.
The reference behaves this way on every forwarding path:
- Path-table data and link requests: the outbound interface IS the
receiving-side path interface —
outbound_interface(Transport.py:1583) — andTransport.transmit(Transport.py:1635) sends there with no receiving-interface guard. - Link-table data: when next-hop and receiving interface are equal,
"direction doesn't matter, and we simply repeat the packet"
(
Transport.py:1651). - Proofs: an LRPROOF goes back out of the interface the link request
arrived on (
Transport.py:2197), a data proof out of the reverse table's receiving interface (Transport.py:2263) — on one shared channel, each is the interface the proof itself arrived on.
Loop-freedom never came from interface suppression. It comes from
transport_id addressing (only the addressed relay processes a Type2
transport packet), the hop-count limit, and packet-hash dedup —
has_packet_hash (leviculum-core/src/transport.rs:3956) drops a
repeated copy, add_packet_hash
(leviculum-core/src/transport.rs:4010) records it.
The forwarding decision lives in the media-agnostic core
(forward_on_interface_from,
leviculum-core/src/transport.rs:8677). Whether the relayed echo
needs TX spacing on a half-duplex channel is the interface's business
— see Interface Isolation.
Pinned by test_pkt_journey_same_interface_relay_forward (one relay,
one interface, forward asserted with iface_out == iface_in) and the
three-node test_shared_medium_multihop_data_forward (announce flood
A→C via B, then data C→A delivered through B's single shared
interface), both in leviculum-core/src/transport.rs.
Client Tools
Leviculum ships command-line tools next to the daemon, and every tool we will ever ship is governed by the same small set of rules. This page records them for whoever adds the next tool: the counterpart rule, the drop-in contract, the explicit permission to exceed the reference, the case where a tool with no Python equivalent is the right deliverable, the one situation in which our own tool must not be used, and the marking rule for tools that can forge traffic.
The daemon-level half of this story — shared-instance IPC and config compatibility, and why drop-in is a design goal rather than an accident — is in Python-RNS Compatibility. This page extends the same property from the daemon to the clients.
A counterpart for every reference tool
The reference stack ships a family of utilities under
reference/Reticulum/RNS/Utilities/. The rule is one ln* counterpart
per reference tool. The honest current state, as of 2026-09:
| Reference tool | Counterpart | State |
|---|---|---|
rnsd | lnsd | Shipped. Drop-in at IPC and config level; see Python-RNS Compatibility. |
rnstatus | lnstatus | Shipped. Local-mode output is byte-parity-pinned against the reference by the 2×2 matrix status_parity_matrix_2x2 (status_parity_tests.rs:1394), the reported inventory by status_inventory_parity_across_daemons (status_parity_tests.rs:2089); Periculum wiring is Codeberg #174. |
rncp | lncp | Shipped (send, fetch, listen). |
rnprobe | lnprobe | Shipped. Same command line, output, and exit codes; proven against both lnsd and rnsd over the shared instance. |
rnpath | lnpath | Shipped for the path-query verb — query, wait, drop — with the reference arguments, output and exit codes for those three. The table and rate views, the blackhole verbs and remote management get no flag rather than a differing one; lnstatus --tables covers the first. Periculum wiring (manifest plus bridge parser) is still owed, as it is for lnprobe (Codeberg #173). |
rnid | — | Missing, not yet filed. |
rnx | — | Missing, not yet filed. |
rnsh | — | Missing, not yet filed. |
rnir | — | Missing, not yet filed. |
rnpkg | — | Missing, not yet filed. |
rnodeconf | — | Missing, not yet filed. (just flash/just flash-rnode cover our own rig's flashing needs, but they are build tooling, not a counterpart.) |
The vendored 1.3.5 tree also carries rngit as a utility subpackage;
the counterpart rule covers it the same way, and it is likewise
missing and unfiled. "Not yet filed" rows are gaps in the issue
tracker, not decisions to skip the tool — file the issue when work on
one begins.
Alongside the counterparts we ship tools with no reference equivalent:
lnstest (test and diagnostics driver, see below), lnomad (NomadNet
browser; its reference counterpart is the NomadNet application rather
than an RNS utility), and lblogd (blog daemon). The counterpart rule
does not restrict these; the drop-in rule below does not apply to them
because there is no reference surface to be compatible with.
Drop-in first
With the reference tool's arguments, our tool behaves like the reference tool. Same flags mean the same thing, same exit codes, and the output is something a script written for the Python tool can still parse. This is the daemon-level drop-in property extended to the clients, and it is what lets a comparison harness point one driver at either stack.
The worked example is lnstatus: the renderer consumes the same
interface_stats dict a Python rnsd exposes, and feeding an
identical stats dict into lnstatus and rnstatus yields
byte-identical output (lnstatus_render.rs:1-15, pinned by the 2×2
parity matrix, which drives both clients against both daemons after
byte-identical controlled traffic, status_parity_tests.rs:1-30).
lnstatus mirrors the reference flag surface — -a, -A, -P,
-l, -j/--json, -m/--monitor, -s/--sort, the name filter — with
the reference meanings, because those flags and formats are the
compatible surface a user's muscle memory and a user's scripts depend
on.
Drop-in is judged against the vendored reference
(reference/Reticulum), which is the same source of truth the
protocol work measures against. When the reference tool's own output
changes between versions, the bridge parsers in Periculum absorb it —
compatibility of meaning, not a frozen byte format, per
Wire Field Semantics.
Drop-in is about the answer, not just the query
A client tool asks a daemon a question. Drop-in means the answer describes the same world, not merely that the query succeeded — and an answer about a different world is the hardest kind of difference to notice, because nothing fails.
Codeberg #177 was exactly that. rnstatus against rnsd listed three
interfaces; against lnsd it listed none, and the query returned cleanly
both times. The cause was structural: our interface_stats was assembled
from Transport's routing map — the interfaces the core can send packets
on — while a Python rnsd reports RNS.Transport.interfaces, everything
Reticulum runs (Reticulum.py:1334). Listeners carry no packets, so they
were in no collection at all on our side, and their absence was the symptom.
The reporting inventory now lives in the driver
(leviculum-std/src/interfaces/inventory.rs) and interface_stats reports
the union: transport's routable interfaces plus the listeners the daemon
runs. Transport stays free of listener rows, which it would otherwise try
to send on.
Codeberg #190 was the same shape one field down. A radio row's bitrate
answered with the per-medium BITRATE_GUESS — 10 Mbit/s, TCP's — because the
key was filled only from a configured bitrate, and a radio configures none.
Python fills it from the interface itself, which for an RNode is the on-air
rate derived from the radio settings (RNodeInterface.updateBitrate,
RNodeInterface.py:693-696). The precedence is now Python's, in Python's order:
a configured bitrate first (if configured_bitrate: interface.bitrate = configured_bitrate, Reticulum.py:887), else the interface's own rate for its
medium, else the guess. Nothing failed while it was wrong; a client asking how
fast the air is simply got an answer about a different medium.
What appears, and under what name
The name is the interface's identity to a script, so each row reproduces the
reference __str__ exactly:
| Row | Name | short_name | type | Reference |
|---|---|---|---|---|
| shared-instance server | Shared Instance[rns/<instance>] | Reticulum | LocalServerInterface | LocalInterface.py:391, 496-498 |
| an accepted IPC client | LocalInterface[rns/<instance>] | <n>@\0rns/<instance> | LocalClientInterface | LocalInterface.py:372-374, 441 |
| a TCP listener | TCPServerInterface[<section>/<ip>:<port>] | <section> | TCPServerInterface | TCPInterface.py:666-672 |
| a connection it accepted | TCPInterface[Client on <section>/<ip>:<port>] | Client on <section> | TCPClientInterface | TCPInterface.py:443-449, 577 |
<section> is the config section name ([[My TCP Server]]), which is
Python's interface.name; it is carried on InterfaceConfig::name because
flattening the parsed config used to drop it. A spawned row also carries
parent_interface_name / parent_interface_hash pointing at its listener
(Reticulum.py:1342-1344), and hash is the full 32-byte
Identity.full_hash(str(interface)) on both stacks, so a script may key
interfaces by hash across daemons. A listener reports clients (its live
spawned count) and its children's byte totals, including those of children
that have since disconnected — the reference gets the latter for free by
incrementing the parent counter alongside the child's
(TCPInterface.py:306-308).
Pinned deviations
Each is a decision, not an accident, and each has a test that fails if it drifts:
- A listener's frequency fields are the sum of its live children's, where
the reference keeps a deque on the listener itself
(
TCPInterface.py:634-644). Identical at rest (both read exactly 0, which is what the frozen comparisons assert) and equal in aggregate under load; they can differ while a child that contributed samples has already disconnected. - Interfaces other than the four rows above still report our internal
name (
tcp_client_0,rnode_0,auto/eth0/…) rather than the reference'sTCPInterface[<section>/<host>:<port>]family. A script that keys on those names still sees an unfamiliar identity; the drop-in gap is narrowed, not closed. - Config interfaces are ordered by section name, not config-file order: the parsed config is a map keyed by section name, so file order is not recoverable. Deterministic run to run, which HashMap iteration was not.
- Extra and missing keys: ours adds
announce_queueandpeers, the reference addsautoconnect_source. Pinned exactly byassert_daemon_stats_parity. tx_jitter_max(Codeberg #190): the ceiling, in seconds, of the randomised pre-TX delay an interface draws against before a frame goes on the air. The reference has no equivalent — Python'sRNodeInterfaceleaves medium access to the RNode firmware's CSMA and holds no such attribute, so nothing inget_interface_stats(Reticulum.py:1326-1470) reports it. Purely additive:rnstatusreads every field by name (ifstat["name"], rnstatus.py:391;if "<key>" in ifstatfor the optional ones) and never enumerates an interface dict, so an unknown key is not read. Emitted only where the concept applies, the way Python gatesairtime_shortand friends onhasattr, which is why it does not appear in the TCP-only rows the parity matrix compares.- An IPC client's
short_nameindex mirrors the reference's live-client count at accept time (LocalInterface.py:441/355), so the labels of two daemons agree only when their clients connected and left in the same order.
Exceeding is allowed and wanted
Additive flags and richer output are welcome. The constraint is only that the reference-compatible surface stays reference-compatible: an extra flag must not change what an existing flag does, and extra output must not break what a Python-tool script would parse.
Current examples: lnstatus --instance_name
(leviculum-cli/src/lnstatus.rs) selects a shared instance
by name where the reference tool only reads it from the config file,
and lnstatus --tables (the second half of Codeberg #174) exposes
internal tables rnstatus cannot show at all.
Both are additive: run lnstatus with exactly rnstatus's
arguments and you get rnstatus's behaviour.
The shape an additive dump takes
--tables is the worked example, and four of its decisions generalise
to the next one.
Additive key, not an envelope. The tables go into the -j object
under one new key rather than wrapping it, so the stats dict stays the
top-level object. Everything that parses lnstatus -j today — Periculum's
parse_status scans for the line whose object carries interfaces —
keeps working untouched, and -j without the flag is what it always was.
Wrapping would have been tidier and would have broken every existing
reader.
Reference names where the reference has one, its own vocabulary where
it does not. Python serves exactly one of these tables over RPC
(get_path_table, Reticulum.py:1516-1538), so those six keys and their
units are taken verbatim and our one addition sits beside them. The other
tables Python holds but never exposes; it names their fields only by list
index (IDX_RT_*, IDX_LT_*, IDX_AT_*, IDX_TT_*,
Transport.py:3556-3586), so the string keys are ours, spelled after those
constants. The additive keys are safe against a Python reader for the same
reason tx_jitter_max is: every Python consumer of an RPC response reads
it by name and none enumerates it.
Naming collides once, and it is worth knowing about: link_table in the
dump is Transport.link_table, the links this node relays. The
pre-existing link_table RPC (lnstest diag) is the links this node
terminates, and appears in the dump as local_links. The reference name
won the contested word because the reference has a table by that name; the
inventory the reference has no table for at all took the qualified one.
Absent is not empty. A daemon that implements the query answers with
the key present and its tables possibly empty. A daemon that does not — a
Python rnsd, or an lnsd from before the flag — makes the client omit
the key, print why on stderr, and exit 0. Presence therefore distinguishes
"cannot answer" from "nothing there". Nulling the key, or defaulting it to
empty lists, would have made every assertion about an empty table pass
silently against a daemon that cannot answer it — the read-side tolerance
question of Codeberg #183, one layer up. Python's rpc_loop matches no arm
for an unknown command and falls through to conn.close()
(Reticulum.py:1213-1260), so the absence surfaces as a fast transport error
rather than a hang; that is pinned against a real rnsd in
reverse_rpc_interop_tests.
The expensive half is opt-in, and absence covers it too. Sizes and rows
are different questions with costs three orders of magnitude apart: a size is
a len(), a row makes the daemon build a dictionary with a string key per
field, and a field node holds 43 000 of them in one table. The dump therefore
always answers the cheap question and answers the expensive one only when a
request names the table (--table-rows, Codeberg #028) — so the common poll,
"how full is it", stops moving the daemon's resident set by tens of megabytes.
This is where the rule above earns its second use: a table whose rows were not
asked for is absent, because the one thing it must not be is present and
empty, which already means something else. The response says in rows_for
which tables it carries rows for, so the reader never has to infer it from
which keys turned up.
One honesty note, because it is easy to get wrong in the other
direction: -j/--json, -m/--monitor and the announce/path-request/
link statistics flags are not exceedances — the reference
rnstatus has all of them (rnstatus.py:685-706), and ours mirror
them under the drop-in rule. Claiming reference-mirrored features as
our extensions would misstate where the compatible surface ends;
check the reference before calling a flag additive.
New tools where testing gains from them
lnstest exists because Periculum needed a driver the Python tool
set does not offer — deterministic selftest phases (delivery,
ratchet, link) with machine-parseable summary lines
(leviculum-cli/src/lnstest.rs:1-4). That is the precedent: when a
test cannot be written because no tool can express it, the tool is
the deliverable.
The currently open examples — not a closed list — are:
- Codeberg #175: wire-level tools, a packet injector and a decoder with no Python equivalent.
- Codeberg #176: a structured event tap, so Periculum can assert on events instead of scraping container logs.
A new tool of this kind has no reference surface, so the drop-in rule does not bind it; the evidence rules do. In particular a test tool's output is a diagnostic indicator: it must measure the production path and be observed in both states before its green is believed.
The rule a comparison must not break
A stack comparison drives the same client against both daemons. That is the whole point of the drop-in property, and it cuts both ways: our own client may not replace the reference tool in the very tests that measure against the reference. Substituting "our better tool" on one side smuggles config, cadence and timeout differences into a result that claims to be about the stacks — the parallel-driver failure described in Evidence and Honesty.
So: lnstest selftest pointed at either daemon is a valid A/B.
lnstatus against lnsd compared with rnstatus against rnsd is
not a stack comparison — it varies two things at once. The 2×2 parity
matrix (status_parity_tests.rs:5-30) is the shape that untangles
this: each client against each daemon, so client-render parity and
daemon-stats parity are separated instead of conflated. This is the
one place where "use our better tool" is wrong.
A tool that can forge is marked as such
Anything that emits crafted frames — the planned packet injector of Codeberg #175 first among them — refuses to run without an explicit flag acknowledging that it forges traffic, and names itself in its output so a capture containing forged frames is attributable. A crafted frame in a mesh is indistinguishable from a real one by design; the honesty has to live in the tool. No shipped tool forges today; this rule binds the first one that does.
Adding the next tool: the checklist
- Name and scope.
ln*counterpart of one reference tool, or a new testing tool per the precedent above. Check the tracker first (#173 covers probe and path query). - Drop-in surface. Implement the reference tool's flags with the reference tool's meanings and exit codes. Divergence from the reference's internals is fine under the deviation rule; divergence of the compatible surface is not.
- Parity evidence. Pin the drop-in claim with a test that could
fail — the 2×2 matrix of
status_parity_tests.rsis the model. - Periculum wiring. Add a client manifest
(
periculum/periculum/adapters/clients/*.toml) and an output parser in the bridge (periculum/periculum/src/bridge.rs), so scenarios can drive the tool and assert on its output. - Docs. Guide page, man page, and both in
docs/src/SUMMARY.md. - Forgery marking, if the tool can emit crafted frames.
See also
- Python-RNS Compatibility — the daemon-level drop-in property this page extends, and the deviation rule.
- Evidence and Honesty in Testing — the A/B discipline the comparison rule enforces.
- Wire Field Semantics — compatibility as meaning, not bytes.
Cryptographic Identity and Forward Secrecy
Every node and every endpoint in a Reticulum network is identified by cryptography, not by an address handed out by infrastructure. This page explains the conceptual model. For the exact byte layouts, defer to the Reticulum specification (its Identity, Destination, and Announce sections).
Identities are dual keypairs
A Reticulum identity holds two keypairs, used for two different
jobs (leviculum-core/src/identity.rs:55):
- X25519 — for key agreement (ECDH). This is how two parties derive a shared secret to encrypt traffic to each other.
- Ed25519 — for digital signatures. This is how a node proves an announce or a packet genuinely came from the holder of the identity.
An identity may be full (it holds the private halves and can decrypt
and sign) or public-only (it holds just the public keys, learned
from someone else's announce, and can only encrypt and verify). In the
source this is the difference between the Option-wrapped private
fields and the always-present public fields
(leviculum-core/src/identity.rs:58).
Destinations are derived addresses
You do not pick a Reticulum address; you derive one. A
Destination is an addressable
endpoint whose 16-byte hash is computed from an application name, a
set of aspects, and (for most types) an identity
(leviculum-core/src/destination.rs:1). Because the address is a hash
of stable inputs, it is reproducible and self-authenticating: anyone
who knows the inputs computes the same address, and the identity bound
into it proves ownership.
A destination also carries a type (SINGLE, GROUP, PLAIN,
LINK) that selects its encryption behaviour, and a direction
(IN, OUT) that selects whether it can receive or send
(leviculum-core/src/destination.rs:6-7).
Announces carry the public keys
A node makes itself reachable by broadcasting an announce: a signed notification that carries the destination's public keys out into the mesh. Peers that receive it learn the destination's address and the keys needed to encrypt to it, and Transport learns a path back. The exact announce wire format is specified in the Reticulum spec.
Ratchets: forward secrecy without a link
End-to-end encryption protects traffic in flight, but if a long-lived
identity key is ever compromised, an attacker who recorded past
ciphertext could decrypt it. Ratchets close that window for
packets sent to SINGLE destinations without first establishing a
Link (leviculum-core/src/ratchet.rs:1).
The mechanism, conceptually:
- A destination enables ratchets and generates an initial X25519 keypair.
- It includes the current ratchet public key in its announces.
- Senders encrypt to the ratchet public key, not the long-term identity key.
- The destination rotates its ratchet keypair periodically (default ~30 minutes).
- Old ratchets are retained for a while so late-arriving packets still decrypt (default 512 retained), then discarded.
Because the rotating key is short-lived and the private half is thrown away after rotation, compromising the long-term identity does not expose traffic encrypted to expired ratchets. That is forward secrecy.
Persisting ratchet keys across restarts is the job of the
Storage trait (the ratchets/ and
ratchetkeys/ collections, see
Architecture). Links —
the other path to forward secrecy, via an ephemeral session handshake —
are a separate mechanism; see the
Reticulum specification.
Where to read the exact bytes
This page stays conceptual on purpose. The authoritative definitions
of identity serialisation, destination hashing, announce structure,
and ratchet encoding are in the
Reticulum specification. The
Rust types above (identity.rs, destination.rs, ratchet.rs)
implement that specification.
Storage and Embedding
leviculum-core is #![no_std] with only alloc
(leviculum-core/src/lib.rs:59-70). It contains no I/O, no clock, no
filesystem, and no async runtime. That is what lets the exact same
protocol code run on a Linux daemon, a future Android app, and a
bare-metal nRF52 firmware image. The bridge to the outside world is a
small set of traits the core depends on but does not implement.
Three injected dependencies
The core declares its platform needs as traits in
leviculum-core/src/traits.rs and takes implementations from the
driver:
Clock(traits.rs:419) — supplies the monotonicnow_ms()and, only where the platform has a real wall clock,wall_unix_secs()(defaultNone). The core never calls a system clock; time is handed in.now_ms()is a timer, not a calendar — on the host it counts milliseconds since process start (leviculum-std/src/clock.rs:45), on the nRF52 it is the Embassy timer (leviculum-nrf/src/clock.rs:8). Which wire fields need calendar time instead, and where a clockless node gets it, is the subject of Time and Clocks.Storage(traits.rs:500) — supplies persistence and lookup for every collection the protocol maintains.flush()defaults to a no-op (traits.rs:856) so a RAM-only backend needs to implement nothing extra.Interface(traits.rs:242) — supplies framing and the wire (see Interface Isolation and the Interface trait).
Randomness is injected the same way, as an explicit
rng: &mut impl CryptoRngCore parameter rather than a global
(leviculum-core/src/lib.rs, "Platform Dependencies").
The Storage trait
Rather than a generic key/value blob store, Storage exposes
type-safe methods grouped by collection — packet-dedup hashes, the
path table, the reverse table, link/announce tables, receipts, and
ratchets — with typed entries from storage_types.rs. The full method
inventory is tabulated in
Architecture.
This shape was a deliberate decision. The deep analysis of every method — who calls it, how often, and whether it matters on an embedded target — is in Storage Trait Split Analysis. Read that page before changing the trait surface.
Three backends, one core
The same NodeCore is parameterised over its Storage
implementation, so embedding is a matter of choosing a backend
(leviculum-core/src/node/mod.rs:409,
NodeCore<R: CryptoRngCore, C: Clock, S: Storage>):
| Backend | Where | Behaviour |
|---|---|---|
NoStorage | tiny / stateless | no-op |
MemoryStorage | host / tests | BTreeMap, RAM only (inner store of FileStorage) |
EmbeddedStorage | embedded (nRF52) | heapless::FnvIndexMap, fixed capacity, no allocator for maps |
FileStorage | host (leviculum-std) | wraps MemoryStorage + disk |
FileStorage persists only what must survive a restart — known
destinations, the packet dedup hashlist, and ratchet keys — and keeps
the rest (paths, reverses, links, announces, receipts) in RAM,
rebuilt from the network on restart. The file formats and flush
strategy are in
Architecture.
What the split buys you
- Host vs. embedded from one source tree.
leviculum-stdbuilds a tokio driver around the core;leviculum-nrfbuilds an Embassy driver around the same core (leviculum-nrf/src/bin/t114.rs,leviculum-nrf/src/bin/rak4631.rs, both#![no_std]and both constructing the core viaNodeCoreBuilder). - Testability. Because time and storage are injected, the core is
driven deterministically in tests — feed bytes and a fixed clock,
drain the
TickOutput, assert. This is the basis of the minimal-reproducer tests underleviculum-std/tests/mvr/. - No host concerns in the core. Backpressure, airtime budgeting,
and serial queueing live host-side in
leviculum-stdand never leak into theno_stdcore (leviculum-std/src/interfaces/airtime.rs:1).
See Architecture for the sans-IO core diagram and the driver event loop that pumps these traits.
Time and Clocks
How a node keeps calendar time, why it must always be able to, and how a wrong calendar heals. This applies to every firmware and every platform port, present and future. Issues come and go; this concept stays.
This page is the binding spec of the anchor model. It replaces the earlier doctrine "no verified clock, no authorship": the rule is now that every instance always authors, stamped with its best honest estimate and never ahead of that estimate — and that foreign garbage timestamps must never break an instance. "Never ahead" is exactly what the mechanisms deliver, no more: every fallback anchor lies in the past, so an uncertain calendar errs backwards by construction. It is not a claim to detect a wrong source: an anchor that is wrong but inside the sanity window — including a plausible forward-wrong one — is adopted and stamped as-is, a named residual, time-bounded by the healing loop. The sections below define the estimate, the anchor's provenance rank, the one filter every time source passes, and the loop that heals a wrong calendar. Testing the model closes the page: the rule-by-tier matrix that binds every one of these rules to a named test cell, Periculum included.
The two poisons
Calendar failures hurt in exactly two directions, and they are not symmetric.
Too old, self-consistently. A Reticulum peer orders
same-destination paths by the announce emission timestamp. Python-RNS
stamps int(time.time()) — epoch seconds — into the announce random
hash (reference/Reticulum/RNS/Destination.py:282) and replaces a
stored path only when a new announce carries a newer emission than
the stored one (announce_emitted > path_timebase,
reference/Reticulum/RNS/Transport.py:1772; the worse-hop branch
runs the same newer-wins comparison against a separately computed
field, announce_emitted > path_announce_emitted,
Transport.py:1809). The value
we stamp is therefore compared on other machines, across our reboots,
against values our earlier selves emitted. A clock that is merely
self-consistent — process uptime — restarts from zero on every reboot
and loses that comparison forever: the node keeps announcing, and no
peer ever updates its path entry again (Codeberg #155). Obtaining
usable time is a protocol obligation, not a platform convenience.
#155 is one instance of a general failure class — a generated field
that is self-consistent between our writer and our reader and means
something else to a peer; the class and the audit method for it are
in Wire Field Semantics.
Too far in the future. The reverse error is worse, because it poisons others, silently and permanently. Receivers advance monotonic cursors over the stamps they ingest — the worked case is the telemetry collector cursor (Telemetry), which one future-stamped row raises past every honest later reading, forever, with no refusal logged anywhere. A past-stamped message, by contrast, sorts too far back where it is displayed directly: visible, attributable, self-limiting. One reader must be priced honestly rather than folded into that: a collector serves only rows above a requester's cursor, so a past-stamped row below an already-synced cursor is not "sorted backwards" — it is invisible to that requester, indistinguishable from loss, until the sender's calendar heals (see the one honest cost).
These two poisons shape the whole model. A node must always be able to author — an instance that falls silent for lack of a clock fails the switch-on-and-it-works requirement exactly when it is needed — and when its calendar is uncertain, the error must land in the benign direction: backwards, never forwards.
Two clocks, strictly separated
A node runs two clocks with disjoint jobs, and no time correction ever crosses from one to the other.
The stopwatch is the monotonic tick counter, Clock::now_ms
(leviculum-core/src/traits.rs:369). All timeout and deadline
arithmetic stays on it — retries, link timeouts, announce cadences.
No anchor change, GNSS fix, or healing step ever stretches or
shrinks a protocol timer.
The calendar clock is an estimate, not a measurement: an anchor (unix seconds, obtained from some source) plus the stopwatch time elapsed since that anchor was seated. Two properties follow by construction:
- It never stands still. Between anchors it advances with the stopwatch, so repeated identical stamps — the #217 class of same-timestamp ID collisions — cannot come back.
- It moves in jumps only when a better anchor re-seats it. An anchor change is an explicit, recordable event with a source, not a drift.
The clockless arm of the implementation already has this shape: a
learned floor advanced by the monotonic clock, inside
Transport::emission_secs (leviculum-core/src/transport.rs:5014).
On std platforms the same estimate answers from the OS: the
SystemClock implementation of wall_unix_secs
(leviculum-std/src/clock.rs:49) reads SystemTime, with the
anchor-keeping delegated to the OS and its NTP discipline.
One value, one producer
Transport::emission_secs (leviculum-core/src/transport.rs:5014)
is the single point that turns the calendar estimate into the
unix-seconds value for wire fields that peers compare across our
process lifetimes: announce emission timestamps, built by
generate_random_hash (leviculum-core/src/announce.rs:156), and
request timestamps
(leviculum-core/src/node/mod.rs:1477, :1154, Codeberg #164). Any
new wire field with cross-lifetime semantics draws from it too —
never from the monotonic Clock::now_ms, which is a timer, not a
calendar.
Crates layered on NodeCore reach the same producer, they do not
take a parameter. NodeCore::emission_secs exposes it; LXMF's three
cross-lifetime fields — the message timestamp, the ticket expiry and
the propagation upload timestamp — resolve through it inside the
router (leviculum-lxmf/src/router.rs, Codeberg #182). They used to
arrive as a now_unix: f64 argument on enqueue, tick,
handle_event and issue_ticket_field, which is the #155 shape with
the defect moved into the caller: nothing about the signature stops a
clockless node from passing uptime seconds. An API that cannot be
called wrongly beats a doc comment warning about it. Pinned at
leviculum-lxmf/tests/wall_clock_producer.rs, structurally (no public
router entry point takes an f64) as well as by value.
One producer, two resolutions
The field decides the resolution, not the clock. emission_secs
returns whole seconds for fields that are whole seconds on the wire —
the 5-byte announce emission timestamp above all. emission_micros
returns the same instant in microseconds, and
NodeCore::emission_secs_f64 divides it back into fractional unix
seconds for fields that are floats. They are the same producer with
the same source ranking:
emission_micros(now) / 1_000_000 == emission_secs(now) on every arm.
The float resolution exists because LXMF hashes the message timestamp
into the message ID. At whole seconds, two identical messages created
inside one second are one ID, and the second is refused as a duplicate
of the first — a message the reference, writing time.time()
(reference/LXMF/LXMF/LXMessage.py:357), would have sent (Codeberg
#217).
Microseconds, not milliseconds. The collision is between two calls
in one code path, not two user actions. Two consecutive
LxmfRouter::create_message calls measure ~115 µs apart — each signs
an Ed25519 message — so a millisecond value collides on every such
pair; that was measured at 20 pairs out of 20 before this unit was
chosen. Microseconds is also what the reference effectively produces:
time.time() is an f64 of unix seconds, resolving to ~0.24 µs at
present-day timestamps.
On the clockless arm the sub-second part comes from our own monotonic
clock, and Clock::now_ms is the only monotonic source there. So
emission_micros returns floor_secs * 1_000_000 plus the
milliseconds elapsed since the floor was set, scaled up: monotonic,
separating two instants one millisecond apart, and honest about the
resolution the platform actually has. It is not a claim to know the
wall-clock microsecond, and nothing reads it as one. Platforms whose
Clock implements only wall_unix_secs get the trait default —
secs * 1_000_000 — and keep working unchanged, with no precision to
gain and none invented.
An age is not a timestamp, and the producer steps
Everything above is about instants a peer compares. A second kind of consumer measures ages against the same value — how long ago did we see this ID, how long ago did this peer announce its stamp cost — and for that consumer the calendar clock has a property the instant consumers do not care about: it moves in jumps.
A clockless node stamps from uptime seconds until its first re-anchor, so the step at that moment is the full distance from a few seconds to present-day unix time — about 1.7e9 seconds, at once. Any window aged by subtracting a stored calendar value from the current one expires in that single step, however short the window's real age. Nothing grew old; the ruler changed.
The rule, therefore: age on the stopwatch, stamp on the calendar.
A cache that must survive process restarts has to persist calendar
values — the stopwatch epoch does not outlive the process — so it
stores the calendar stamp and absorbs the step instead. The LXMF
router does exactly that: it keeps the (emission_secs, now_ms) pair
of the last tick and, when the calendar advances by more than the
stopwatch says it should have, shifts every stored stamp by the
difference (LxmfRouter::anchor_wall_clock,
leviculum-lxmf/src/router.rs:1927, Codeberg #186). The ages the
cache encodes are then preserved exactly across the re-anchor, and
real elapsed time still expires entries.
Two things are deliberately NOT absorbed. A stamp is never moved past
the current calendar value, so a checkpoint restored from a life with
a real clock onto a node still on uptime seconds cannot become
unexpirable. And expiries a peer wrote from its own clock — an LXMF
ticket's expires_unix — are absolute instants, not ages measured
here: a node whose calendar has just become real should start
honouring them, not carry them along.
The one refusal left: a field the peer discards in silence
Authorship is never refused — that is the headline rule of this page. One narrow refusal survives it, and it blocks no message.
The LXMF ticket expiry is compared on the peer's clock: a peer
keeps a ticket only while time.time() < expires on its own machine
(reference/LXMF/LXMF/LXMRouter.py:1854) and says nothing when it
does not. A backwards-biased expiry from an unhealed calendar is
therefore already expired on arrival — issuing it is emitting a field
the peer silently discards. LxmfRouter::issue_ticket_field
(leviculum-lxmf/src/router.rs:669) returns
RouterError::NoWallClock while the calendar is not a plausible wall
clock, rather than issue one: a named error is a diagnosis; a
discarded ticket is a mystery that surfaces months later as "replies
from this peer are slow".
"Plausible" here is a question about the anchor's
provenance rank, not
about its value: a birth-anchored (rank 5) calendar holds a value
that passes the sanity window and is refused anyway, because its
estimate is recognisably behind real time and every expiry it
computes is already in the past on every healed peer. The gate is
NodeCore::has_plausible_wall_clock
(leviculum-core/src/node/mod.rs:3608), and since Codeberg #247 it
asks the rank: anchor_rank() < BIRTH_ANCHOR_RANK.
Fields the peer decides nothing on are always emitted. The LXMF
message timestamp (displayed and sorted, LXMessage.py:357
unvalidated) and the propagation upload timestamp (bound and dropped,
LXMRouter.py:2238-2240) both flow regardless of clock state — a
backwards-biased stamp mis-sorts, and mis-sorting is the accepted
cost, not a failure. Withholding a message because our clock is
uncertain would be the far worse failure.
The refusal is always on writing, never on reading: a peer's ticket is remembered and used regardless of the state of our own clock.
The sanity window
Every candidate anchor, from every source — RTC, GNSS, host injection, network-learned time — passes one filter. No source gets a special case, and no source bypasses it.
- Lower bound: the build timestamp. The firmware or binary build time is compiled in: free, always present, and incorruptible by any runtime input. Real time is always after it, so any source claiming a moment before it — a 1999 RTC with a dead backup cell — is deterministically garbage, not merely suspicious.
- Upper bound: a generous margin above the best known anchor, on the order of decades above the build floor. It only has to separate values a real clock could hold from values none can; its exact size is a tunable practice parameter (see Practice parameters), not dogma.
A source that fails the window is refused as an anchor — never "corrected" — and the calendar keeps running on the best anchor it has. GNSS gets no bypass: it is simply the highest-ranked source inside the same filter, trusted by default and rejected when implausible. A receiver subtly shifted within the window (a spoofed or faulty fix that still looks plausible) is an accepted residual risk, time-bounded by the healing loop.
The window answers exactly one question: may this value seat an anchor at all. Every other predicate in the model — is first adoption still unbounded, may a ticket be issued, does the calendar count as a plausible wall clock — keys on the anchor's provenance rank, not on its value clearing the window. The next section is that rule; skipping it re-introduces two regressions by accident.
Implementation status. The lower bound is the build timestamp
(Codeberg #247): leviculum-core/build.rs embeds it, BUILD_UNIX_SECS
(constants.rs:621) carries it, and EMISSION_SANITY_FLOOR_SECS
(constants.rs:654) is the bound the filter applies — the build stamp,
floored by the old fixed date EMISSION_PLAUSIBLE_MIN_SECS
(constants.rs:628, 2020) so a bogus SOURCE_DATE_EPOCH cannot lower
it. The upper bound is still the fixed date
EMISSION_LEARN_CEILING_SECS (constants.rs:611, 2200-01-01). One
filter enforces both on learning, on host injection and on GNSS. The
stamp is precise enough by construction: real time is always after the
moment the binary was built, and a stale stamp only widens the window.
Anchor provenance is first-class state
An anchor is a pair: the value it seated and the rank of the source that seated it (the arms of the next section). The calendar keeps both, and the model's predicates split cleanly over the two:
- The sanity window bounds anchor admission — a question about the value, asked once, on the way in.
- Adoption, healing and ticket predicates key on the anchor's rank — never on whether its value happens to clear the plausibility floor.
The distinction is load-bearing, because the build floor (arm 5) sits at the plausibility floor by construction: every firmware is built after 2020, so a birth anchor passes every value test from the moment arm 5 is plumbed. Two predicates in the tree then misfire if they stay keyed on the value:
- Unbounded first adoption.
learn_emission_timebase(leviculum-core/src/transport.rs:5201) selected the unbounded branch bycurrent < EMISSION_PLAUSIBLE_MIN_SECSuntil #247. A birth-anchored cold node clears that test, so its first credible announce would have fallen into the bounded branch and the node would have crawled to real time at one day per announce — the exact #161 §1 regression this page forbids — instead of healing in one step. It now asksanchor_rank() == BIRTH_ANCHOR_RANK. - The ticket refusal.
NodeCore::has_plausible_wall_clock(leviculum-core/src/node/mod.rs:3608) became vacuously true at the build floor: the refusal would never fire again, and a birth-anchored node would issue tickets whose expiry is already in the past on every healed peer — the silently-discarded field the refusal exists to prevent. It asks the rank too.
The binding rule: while the calendar is anchored at rank 5 (birth), first adoption is unbounded, tickets are refused, and the calendar does not count as a plausible wall clock — regardless of the anchor's value. Stated the other way round: a build-floor anchor passes the sanity window for stamping — the node authors, per the stamping rule below — but it is never "plausible" for tickets or for capping adoption.
Both switches landed in the same change as the build floor (#247),
which is what that requirement meant: a port that plumbs the stamp
without moving the predicates ships both regressions at once. The
rank is TimeSource::rank (leviculum-core/src/transport.rs:1579)
and BIRTH_ANCHOR_RANK (leviculum-core/src/transport.rs:1651) is
the constant every predicate compares against.
The source ranking
The calendar takes its anchor from the best source available, and a higher-ranked source re-seats an anchor from a lower-ranked one — the arms 1 to 4 of the old chain survive here as ranks. The order is by how hard the source is to fool, not by precision: a GNSS fix is a live measurement, a host injection is an explicit claim by an operator, an RTC is whatever it was last set to, and network-learned time is arbitrary input from anyone in radio range. Every arm passes the same sanity window. For each: what it costs, when it is unavailable, what it guarantees.
Rustdoc debt, paid in #247. Two doc comments in the tree stated a different order and were corrected by the issue that implemented this ranking: the rustdoc of
set_wall_time_unix_secs(leviculum-core/src/node/mod.rs:936, and on the transport attransport.rs:4038) said a platform wall clock always takes precedence over an injection — the reverse of arms 2 and 3 — and itsNodeCore::emission_secs(leviculum-core/src/node/mod.rs:3593) rustdoc listed the chain as "platform wall clock, learned announce timebase, host injection, uptime". This page is the spec; both now say so, and both name the one place the implementation still deviates from the order — it consults the platform clock first, which no platform can currently observe because none offers a second arm alongside it.
Arm 1: GNSS
Where the board has a receiver — today the WisMesh Pocket V2's
u-blox ZOE-M8Q (leviculum-nrf/src/gnss.rs). The NMEA RMC sentence
carries UTC date and time in every fix.
- Cost: receiver power, a sky view, and cold-start acquisition time (seconds to minutes).
- Unavailable: indoors or shadowed — a node without sky view never gets a fix, so GNSS seeds the calendar, it never replaces the ranking below it.
- Guarantees: UTC to well under a second, far beyond the one-second wire granularity. Trusted by default; a fix outside the sanity window is refused like any other source.
- Status: the firmware parses RMC but consumes only position and
validity (
GnssFix,leviculum-nrf/src/baseboard.rs:26, has no time field yet). Seeding the calendar from RMC is tracked as a firmware implementation issue, not here. See GNSS specifics below.
Arm 2: Host injection
Node::set_wall_time_unix_secs (leviculum-core/src/node/mod.rs:934
→ leviculum-core/src/transport.rs:5125), for deployments where a clockless node has a
host that does know wall time — e.g. a control frame on the LNode
serial channel (the radio-config envelope of
leviculum-core/src/rnode.rs).
- Cost: one control-channel frame; requires a host that itself has a trustworthy clock.
- Unavailable: standalone nodes with no host attached.
- Guarantees: host-clock quality, sanity-gated: values outside
[EMISSION_PLAUSIBLE_MIN_SECS, EMISSION_LEARN_CEILING_SECS]are refused (leviculum-core/src/transport.rs:5133), because an injection claims to know wall time, so a value no real clock can hold is self-refuting. Pinned attest_implausibly_low_wall_time_injection_is_refused(transport.rs:23649) andtest_absurd_wall_time_injection_is_refused(transport.rs:24112). - Status: wired (#238): the control envelope's wall-time frame
(
docs/src/firmware/usb-control-envelope.md) carries a u64 of unix seconds from the host to the seam; the seam's bool picks the enveloped ack or the namedvalue refusedanswer, an accepted seed logs[TIME_SEED] source=host, andlnflash --set-timeis the speaker.
Arm 3: Platform clock passing sanity
Clock::wall_unix_secs (leviculum-core/src/traits.rs:440). On std
platforms SystemClock answers from SystemTime
(leviculum-std/src/clock.rs:50) — effectively the OS's NTP-managed
clock. On a board it is a battery-backed RTC, which under this model
also carries anchors back: see
the healing loop for the write-back.
- Cost: none.
- Unavailable: on MCUs without an RTC. The LNode's
EmbassyClock(leviculum-nrf/src/clock.rs) keeps the trait default ofNone— that is correct, not a gap: returning uptime-derived values fromwall_unix_secswould be lying to the transport. - Guarantees: whatever the platform clock guarantees — NTP
quality on a host, last-set-plus-drift on an RTC. An RTC counts
only when its value passes the sanity window; below the build
floor it is a dead cell, not a time source. Note the trait
contract: this is not a timer source; all timeout and deadline
arithmetic stays on the monotonic
now_ms(traits.rs:369). - Status: the implementation consults this arm first when it
answers (
emission_secs,leviculum-core/src/transport.rs:5014). No current platform offers both a platform clock and GNSS or injection, so the difference in order has no behavioural effect today; a port that has both follows this ranking. The arm reports itself asTimeSource::PlatformClockonly while its value passes the window (time_source,leviculum-core/src/transport.rs:5135); below the build floor it is a dead cell and the calendar stays birth-anchored, which is what refuses tickets on a node with a dead RTC.
Arm 4: Network-learned
The sourceless fallback: learn_emission_timebase
(transport.rs:5201) adopts the highest emission timestamp seen in
any signature-valid announce as the calendar anchor, then advances it
with the monotonic clock (transport.rs:4494). This includes the
node's own pre-restart announce echoing back from a neighbour —
learning deliberately runs before the own-destination echo drop, so a
rebooted node re-seeds past exactly the value its next announce must
exceed — pinned at
test_own_announce_echo_reseeds_timebase_before_echo_drop
(transport.rs:23353).
- Cost: nothing — no hardware, no host.
- Unavailable: on a mesh where no participant has a clock, or before the first plausible traffic arrives.
- Guarantees: only as good as radio-range neighbours, and
validate()proves only that the announce signs itself — the field is arbitrary input from anyone in range. Hence every hardening rule below and the re-anchor rules of the healing loop. A calendar seated by this arm is anchored from traffic, unconfirmed: the evidence may itself be another unhealed node's birth clock (see the one honest cost). The learned anchor is in-memory only on RTC-less boards: after a reboot the node starts from the build floor again until live traffic re-seeds it. Adoption records the transition: the source becomesTimeSource::Overheard, rank 4, and a higher-ranked anchor that is merely being advanced keeps its own rank.
Arm 5: The build floor
The birth state of every instance with no better source: the calendar anchors at the build timestamp, advanced by uptime. This replaces raw uptime seconds as the bottom of the ranking, and it is a valid state, not a defect: the stamp is recognisably old, but unique, monotonic, and non-toxic — it can never poison a peer's cursor, because it errs backwards. Nobody stays silent, nobody poisons; the cost is confined to the one honest cost below. The rank is the part that matters beyond the value: a rank-5 anchor stamps, but it never counts as a plausible wall clock, never issues tickets, and never caps adoption (provenance rank).
- Cost: none; the build timestamp is compiled in.
- Unavailable: never — that is the point.
- Guarantees: uniqueness and monotonicity within the boot, a value that is always in the past, and a floor the sanity window can trust.
- Status: implemented in the core (#247). An instance with no
source anchors at
BUILD_UNIX_SECS(leviculum-core/src/constants.rs:621) advanced by uptime, which retires the raw-uptime state every cross-restart comparison lost (#155). A port inherits it with the core and adds nothing; what a port still owes is the arms above it.
Stamping: always author, never ahead of the estimate
Every instance always stamps outgoing fields with its best honest estimate of now — the calendar clock, whatever its anchor. The directional guarantee is stated at its real strength: we never stamp ahead of our own estimate, and every fallback anchor lies in the past, so an uncertain calendar errs backwards by construction. That is where the safety comes from: the poison was never wrong time but future time (the cursor mechanism in the two poisons is why — permanent, silent, hurts others), while past time hurts only presentation. What the mechanisms do not deliver is detection of a wrong source: an anchor that is wrong but inside the sanity window — a plausible forward-shifted GNSS fix, a mis-set host clock — is adopted and stamped as-is. That residual is named here rather than implied away, and it is time-bounded by the healing loop. A source-less instance anchors at build-time plus uptime and authors anyway.
The old rule — no verified clock, no LXMF authorship — is withdrawn. It bought cursor safety by silencing exactly the instances a mesh is for, and the same safety is now had cheaper: backwards-biased stamping at the writer, ingress clamping at the reader.
The one honest cost
Between cold start and healing, outgoing messages carry recognisably old stamps, and the cost is priced per path. Where a stamp is displayed directly, it sorts backwards — a Sideband conversation shows messages out of order until the calendar heals: visible, attributable. On the collector path the same stamp costs more, and the real price is stated rather than rounded down: a collector serves only rows above a requester's cursor, so a birth-stamped row sits below every already-synced requester's cursor — not sorted backwards but invisible to that requester, indistinguishable from loss, until the tracker heals and stamps climb past the cursor (Telemetry).
How the cost ends is graded by what healed the calendar:
- Contact with arms-1–3-quality time — a GNSS fix, a host injection, a plausible platform clock — ends it outright.
- Traffic healing (arm 4) ends it provisionally. A foreign stamp that passes our sanity window may itself be another unhealed node's birth clock: a build-floor stamp from a node built after us (or comparably) clears our own floor. The calendar is then anchored from traffic, unconfirmed — better than birth, not yet known-good — and stays in that state until arms-1–3-quality contact confirms or corrects it. Two clockless nodes healing from each other are an echo chamber, and nothing on the wire can fully prevent it: the stamps are signature-valid and in-window. A named residual, not a solved problem.
We state the cost rather than hide it, because the alternatives are the two poisons: silence or invented time.
Ingress: clamp for semantics, keep for display
The mirror image of backwards-biased stamping protects us from everyone else. An incoming stamp beyond the local plausible-now is clamped to receive time for ordering and cursor semantics: indexing and above all cursor advancement — in our collector (#239) and in any future propagation node. The original stamp is kept alongside as local display information, so nothing is destroyed and a viewer on this node can still show what the sender claimed.
Local plausible-now, defined. The basis is the local calendar
estimate at receive time — Transport::emission_secs
(leviculum-core/src/transport.rs:5014) — plus a bounded forward
tolerance for honest clock skew between sender and receiver. A stamp
at or below basis-plus-tolerance passes as-is; above it, it is
clamped to receive time. The tolerance is a practice parameter (see
Practice parameters): its job is
to keep two honestly-synced clocks from clamping each other, not to
admit the future.
Three boundaries keep the clamp from doing damage of its own:
- Dedup keys are not clamped. Deduplication runs on content or transient ID, never on the timestamp, so clamping can neither make two distinct readings collide nor let a replay through. The clamp covers ordering and cursor semantics only.
- What is served is the clamped value. A telemetry stream row
has exactly one timestamp slot (the row form in
Telemetry), and a collector serves what it
indexed. Keep-for-display is local-only — and the limit of the
defence is stated with it: the row's
packed_telemetrypayload still carries the sender's raw claim in its ownSID_TIME, so a downstream reader that parses the payload sees the claim. Cursors — ours and every requester's — advance over the clamped value regardless. - The clamp is armed only by a healed calendar. Clamping "to receive time" presumes the receiver knows what time it is. An unhealed collector — birth-anchored, or below arms-1–3 quality with no traffic re-anchor yet — would clamp every honest current stamp down to its own ancient notion of now and blackhole the mesh's telemetry into a years-old index. While unhealed, it takes in-window sender stamps as-is; those same stamps are simultaneously its healing evidence (arm 4). Rows ingested before healing keep their index stamps — there is no re-index — and that cost is part of the honest cold-start story above.
The reason for the clamp is the cursor mechanism: a cursor that advances over a raw foreign stamp hands every sender a lever to starve it. Clamped, the worst a garbage stamp can do is index as "arrived now" — wrong by presentation, harmless by mechanism. Foreign garbage can no longer break us, which is the other half of the always-author rule: authorship without ingress protection would just move the poison one hop.
Scope: this clamp lives at the application layer — what we index, serve, and advance cursors over. It does not touch announce path ordering, which stays raw for reference parity; see the non-behaviours below.
Rules that hold regardless of source
The wire field is 40 bits; every producer saturates
The announce timestamp field holds
8 * RANDOM_HASH_TIMESTAMP_SIZE = 40 bits. A larger value would
silently drop its high bits on the wire and sort below every stored
path entry — the node instantly loses path replacement everywhere.
EMISSION_TIMESTAMP_MAX_SECS
(leviculum-core/src/constants.rs:602) caps it, enforced at the
point of resolution (transport.rs:4504) and again at the wire
producer (announce.rs:167), so truncation is unrepresentable
regardless of which source produced the value. Incident: Codeberg
#160. Pinned at test_emission_secs_saturates_at_wire_field_max
(transport.rs:24136).
The timebase never moves backwards, and adoption is windowed
Within arm 4, an older emission never regresses the anchor
(emitted_secs <= current, leviculum-core/src/transport.rs:5223),
and adoption is bounded by the sanity window: values above
EMISSION_LEARN_CEILING_SECS (constants.rs:611, 2200-01-01) cannot
come from a real clock and are refused outright
(leviculum-core/src/transport.rs:5208); the lower bound is the build floor
EMISSION_SANITY_FLOOR_SECS (leviculum-core/src/constants.rs:654),
which real time is always after, so a peer's uptime seconds and a
dead RTC's 1999 are refused as anchors by the same filter (#247).
Incidents: #160, #161. Pinned at
test_clockless_node_learns_emission_timebase_from_announce
(transport.rs:23249) and
test_timebase_floor_cannot_pass_learn_ceiling
(transport.rs:24051).
The per-announce no-backwards guard is not contradicted by the
healing loop: a backwards re-anchor is a deliberate event that
requires arms-1–2 evidence and respects the emitted high-water mark
(see the healing loop) — never one announce
dragging the anchor down.
The FIRST adoption is unbounded while anchored at rank 5
While the calendar is still on its birth anchor — rank 5 — adoption
is deliberately unbounded (learn_emission_timebase,
leviculum-core/src/transport.rs:5201): a node starting at the
build floor must climb to real unix time in one step.
First-plausible-wins is correct here, and only here: a birth
anchor has nothing worth defending, and instant recovery beats
attack resistance for it. The predicate is the anchor's rank,
not its value: a birth anchor's value clears the plausibility floor
now that the build floor is plumbed, and a value test would route a
cold node into the bounded branch below
(provenance rank). The
negative that catches exactly that misfire is
test_birth_anchor_value_clears_the_window_but_not_the_rank
(leviculum-core/src/transport.rs:23747).
Capping this was a real regression (#161 §1): with the bounded
advance applied to an implausibly low anchor, recovery crawled at
one day per announce — about 20 602 announces, ~429 days at a
30-minute LoRa cadence — where a single credible announce used to
recover the node instantly. Do not re-introduce that cap. The
no-backwards guard above keeps the unbounded branch from being
abused downwards. Pinned at
test_clockless_first_timebase_adoption_is_unbounded
(transport.rs:23445) and
test_clockless_timebase_advance_is_bounded_after_first_adoption
(transport.rs:23394).
Advance past rank 5 is bounded — per announce, not per peer
Once the calendar is no longer birth-anchored, one announce may
advance it by at
most EMISSION_LEARN_MAX_ADVANCE_SECS (constants.rs:693, one
day), so a peer whose clock is decades wrong cannot capture the
calendar in one announce. State the protection level honestly:
learning runs before the per-destination announce rate limit and the
rebroadcast dedup, so N announces advance the floor by N × cap
regardless of how many identities or destinations they came from.
The real cap on the walk rate is announces-per-second on the air —
nothing identity-shaped. This measured reality is pinned at
test_timebase_walk_is_capped_per_announce_not_per_identity
(transport.rs:23977); the walk terminates at the learn ceiling,
pinned at test_timebase_floor_cannot_pass_learn_ceiling
(transport.rs:24051). No durable defence is claimed from the
healing loop's median: a median over free
identities resists a broken peer, not an attacker. What actually
bounds a hostile forward walk is the ceiling and the airtime it
costs; what undoes one afterwards is arms-1–2 evidence, because the
healing loop deliberately refuses to move a calendar backwards on
traffic alone.
We do not validate our own clock, or incoming timestamps
Two deliberate non-behaviours, both reference parity. The anchor model does not soften them — it validates anchors on the way into the calendar, never emissions on the way out or announces on the way into the path table:
- Our own calendar estimate is emitted verbatim. Python fills the
field from
time.time()unvalidated (Destination.py:282); bounding, substituting, or withholding our value at emission time would desynchronise us from a network that does not validate. Under the anchor model an "implausible own clock" collapses to "no anchor better than the build floor" — and that state emits too, per the stamping rule. The once-per-process operator warning inTransport::announce_emission_secs(transport.rs:5089) remains the only reaction to an implausible value — never an altered emission. Pinned attest_own_wall_clock_is_not_plausibility_bounded_on_emission(transport.rs:24297) andtest_implausible_own_wall_clock_warns_once_and_leaves_emission_unchanged(transport.rs:24345). - Incoming emission timestamps are not plausibility-checked on
path acceptance. Ordering is per-destination comparison only,
exactly
announce_emitted > path_timebase(Transport.py:1772; the worse-hop branch compares against its own stored field atTransport.py:1809). A clockless peer's uptime-seconds announce must enter the path table (that is how a #155 node is reachable at all), and an absurdly high emission must win the newer-emission comparison. Python peers accept both; filtering would only desynchronise our path tables from every other node's view of the same announces. Pinned attest_incoming_emission_not_plausibility_checked_on_acceptance(transport.rs:24416). The ingress clamp operates strictly above this layer — on what we index and serve, never on what we route.
The sanity window exists for anchor adoption — learning, host injection, GNSS, RTC — never for emission or path acceptance.
The healing loop
The protocol self-heals wherever the evidence for healing exists; where it does not, this section names the residual instead of claiming one. The loop, binding as spec:
- Collect. Plausible foreign times are gathered from live traffic the node already receives — LXMF message stamps, propagation announces. Plausible means: passes the sanity window.
- Re-anchor — by rank and by direction.
- While anchored at rank 5, a single plausible source re-anchors the calendar. This is the unbounded first adoption above, restated as the healing rule: a birth anchor has nothing worth defending. Unless the source was arms 1–3, the result is anchored from traffic, unconfirmed.
- A calendar already anchored to real time is never re-anchored by a single sender. A forward correction requires gross deviation from the median of several distinct senders; cohort size and deviation threshold are practice parameters.
- A gross backwards correction additionally requires arms-1–2 evidence — a GNSS fix or a host injection — never traffic alone. Announce identities are free (this page already concedes that for the walk cap), so a traffic median in the past is exactly what an attacker can fabricate; a backwards path open to traffic would be a remote lever for dragging any healed calendar down and silencing its announces. Closing it costs a residual, named under "sparse meshes" below.
- The emitted high-water rule, binding on every re-anchor:
the calendar is never re-anchored below the highest value this
identity has ever emitted — not even by arms-1–2 evidence; the
correction floors at the high-water mark. Peers order our
announces by emission timestamp (
announce_emitted > path_timebase,reference/Reticulum/RNS/Transport.py:1772), so dropping below our own emitted high-water silences our announces mesh-wide until the calendar climbs past it again — and re-stamping a range we already stamped would revive the #217 same-stamp class this page claims cannot come back. Where storage exists, the high-water mark is persisted across boots. An RTC-less, storage-less node cannot persist it, and that cost is named: such a node re-enters the same build-plus-uptime stamp range on every boot until re-seeded. The practical mitigation is already in the model — arm 4 hears the node's own pre-restart announces echo back and re-seeds past them, pinned attest_own_announce_echo_reseeds_timebase_before_echo_drop(transport.rs:23353).
- Write back to the RTC. Every better anchor is also written into a present RTC, so the hardware clock itself heals and the next boot starts from the healed value instead of the build floor. The write-back is what turns arm 3 from "whatever it was last set to" into "whatever we last verified".
What the median is, honestly. A median over distinct senders resists a broken peer: one wrong clock in an honest cohort cannot move it. It does not resist an attacker — identities are free, and a cohort of them is one attacker with a loop. The attack-facing guarantees come from the other rules: the ceiling and the advance cap bound a forward walk, the arms-1–2 requirement closes the backwards lever, and the high-water rule caps what any accepted correction may do to our own emissions.
Sparse meshes, honestly. With one neighbour no cohort exists, so a calendar that is grossly forward — and no longer at rank 5 — does not heal from traffic at all: the median that would justify a correction cannot form, and traffic alone may never pull backwards anyway. The residual is stated rather than papered over: such a node heals through a GNSS fix, a host injection, or a reflash — not from listening. Until then it keeps authoring, and the learn ceiling keeps it from walking further.
The cold-start story then reads: a device with nothing stamps from its birth anchor; the first plausible contact re-anchors it — to known-good time when the contact was arms 1–3, to traffic-unconfirmed when it was a foreign stamp, which may itself be another unhealed node's birth clock (the echo-chamber residual of the one honest cost); the RTC (where present) keeps it across power cycles; and a calendar later walked wrong is corrected forward by its cohort's median, backwards only on arms-1–2 evidence. No operator action at any step — switch on and it works, with the residuals stated.
Implementation status. Arm 4's single-announce learning
(learn_emission_timebase, transport.rs:5201) implements the
rank-5 re-anchor today. The median re-anchor, the collection of
LXMF-stamp evidence, the high-water persistence, and the RTC
write-back are spec, tracked as implementation issues per platform.
GNSS specifics for a board bring-up
- Use NMEA UTC, never raw GPS time. GPS system time does not observe leap seconds and is currently 18 s ahead of UTC. The receiver applies the broadcast UTC offset before it builds the RMC sentence, so RMC date + time is UTC — take it from there. Getting this wrong is silent: an 18-second skew breaks nothing visibly and is indistinguishable from clock drift in the field.
- Acquire, seed, let the receiver sleep. One fix seeds the calendar; the monotonic clock carries it forward. Crystal drift (tens of ppm — under half an hour per year of isolation) is irrelevant at one-second wire granularity. Keeping the receiver powered buys nothing for time.
- Seed through the sanity window. Route the fix through the same filter as every other source so a garbage fix cannot wedge the calendar; a valid RMC should pass it trivially. A subtly shifted fix inside the window is the accepted residual risk named above, time-bounded by the healing loop.
- GNSS never replaces the ranking. A node without sky view never gets a fix. Arms 2–5 must behave exactly as if no receiver were fitted.
Record the source
A node should be able to state, at any moment, where its notion of
time came from: GNSS, host injection, platform clock, learned from
traffic (whose, and when, and whether still
unconfirmed), a median re-anchor (over which
cohort), or the build floor. This is more than a breadcrumb: the
anchor's rank is live state the model's predicates key on
(provenance rank), so a
node that cannot answer it cannot even decide whether it may issue a
ticket. Diagnosis of a path-ordering problem starts with
"what did this node think the time was, and who told it" — without
provenance, a wrong timestamp in a peer's path table cannot be
attributed to a dead RTC, a lying neighbour, or a boot-order race.
The healing loop raises the stakes: a re-anchor is a calendar jump,
and an unattributed jump is indistinguishable from a bug. The core
answers the first half: Transport::time_source
(leviculum-core/src/transport.rs:5135) names the arm — including
the platform clock, which answers for itself — and
NodeCore::anchor_rank (leviculum-core/src/node/mod.rs:3625) is
the number the predicates use. The rest — which neighbour, when, and
the cohort behind a median re-anchor — is still only the
once-per-process implausible-own-clock warning
(announce_emission_secs, leviculum-core/src/transport.rs:5089);
exposing the origin (status RPC, control-channel query) is part of
implementing this concept on each platform.
Practice parameters, not dogma
The model fixes mechanisms; these values tune them. Each is a practice parameter: chosen to work, changed by measurement and a reasoned commit, never load-bearing for the model itself.
- The upper sanity margin above the best known anchor — order of
decades; today the fixed date in
EMISSION_LEARN_CEILING_SECS(constants.rs:611). - The lower bound is the build timestamp and not a practice
parameter at all;
EMISSION_PLAUSIBLE_MIN_SECS(constants.rs:628) survives only as the guard under it, for a build stamp no firmware was ever built at. - The per-announce advance cap
EMISSION_LEARN_MAX_ADVANCE_SECS(constants.rs:693). - The healing cohort: how many distinct senders form a median, and how large a deviation counts as gross.
- The local plausible-now tolerance: the bounded forward skew
allowance added to the local calendar estimate at the
ingress clamp —
basis is
Transport::emission_secsat receive time; the tolerance absorbs honest sender/receiver skew, nothing more.
Testing the model
Every binding rule on this page maps to a named test cell in a named tier, and this section is that map. It is a concept-level test contract: each implementing issue lands with its cells from this matrix, in the same change, never as a follow-up, and a reviewer ticks the touched rules off against this section before merge. Cells that pin behaviour already implemented are cited; cells for spec-only mechanisms are named here and land with the issue that implements them — for those, named-but-not-landed is the intended state until the issue lands, not accumulated debt. This page names cells; it writes no scenario files and no test code — the implementing issues fill the names in.
The tiers
Four tiers, in the order a failure should be caught:
- unit / mvr. Deterministic, one to two nodes, under five
seconds, structured event logs from all sides — the mvr
constraints of the project's protocol-debugging discipline.
Canonical mvr home:
leviculum-std/tests/mvr/; unit pins live beside the code they pin. Everything decidable on one node's state machine is decided here. - workspace integration. The workspace test suite: multi-crate seams, restart persistence, structural API pins.
- periculum conformance. Docker, against real Python-RNS,
under the drop-in discipline: the SAME driver (
lns selftest,rnstatus, the same client binary) pointed atlnsdandrnsd— never parallel per-stack drivers, which smuggle config differences into what claims to be a stack comparison. The project's interop rule applies in full: every cross-stack rule gets a positive AND a negative cell against real Python. - periculum hardware. The rig, and only where the radio or a real peripheral — a GNSS receiver, an RTC — is itself the subject. Time logic decidable without hardware is decided in a lower tier; this tier verifies the peripheral wiring, not the model.
The matrix
A dash means the rule has no cell in that tier by design, not that one is missing. Names not yet in the tree are the binding intent; the implementing issue may adjust a file name, but it updates this matrix in the same commit.
| # | Rule | unit / mvr | integration | conformance | hardware |
|---|---|---|---|---|---|
| 1 | Stopwatch/calendar separation | calendar_jump_timer_isolation | — | — | — |
| 2 | Sanity window, per source | window trio per arm (arms 2 and 4 pinned) | — | — | — |
| 3 | Anchor rank predicates | rank-5 pair + build-floor negative | wall_clock_producer.rs | — | — |
| 4 | Stamping asymmetry | never-ahead property + monotonic pair | — | — | — |
| 5 | Emitted high-water | re-anchor floor | restart persistence | time_high_water_path_retention | — |
| 6 | Healing loop | median / single-sender / backwards / Sybil / RTC write-back | — | — | — |
| 7 | Echo chamber | — | — | time_echo_chamber_third_peer | — |
| 8 | Ingress clamp + collector | clamp / dedup / unarmed | — | time_collector_clamp_cursor, time_collector_unhealed_heals | — |
| 9 | Authorship interop | — | — | time_cold_authorship_python_receiver, time_python_future_stamps | — |
| 10 | Peripheral seeding and write-back | — | — | — | time_gnss_seed_rig, time_rtc_writeback_rig |
The cells, rule by rule
1. Stopwatch/calendar separation
(Two clocks). mvr
calendar_jump_timer_isolation: a wall-time injection lands in
the middle of a live link, in both directions — a decades-forward
jump and a backwards re-seat — and no relative timer moves. Retry
cadence, keepalive interval and path expiry are observed unchanged
across the jump on the structured event timeline; the test fails
if any timer stretches or shrinks. This is the separation rule's
negative cell and its only cell: the rule forbids exactly one
thing, and this watches for it.
2. Sanity window, per source
(The sanity window). One unit trio per
source arm: a below-build-floor value is refused (the 1999-RTC
dead-cell shape), an above-margin value is refused (a 2500-RTC),
an in-window value is accepted. Arm 1 is
test_refused_gnss_seed_reports_false_and_keeps_source
(leviculum-core/src/transport.rs:23706), which carries the
no-bypass assertion: it drives the same seam the host-injection
cells drive, so a GNSS special case would have to be written into
the shared filter to pass it. Arm 2 is
test_implausibly_low_wall_time_injection_is_refused
(transport.rs:23649) and
test_absurd_wall_time_injection_is_refused
(transport.rs:24112). Arm 3 is
test_platform_clock_outside_the_window_is_not_a_time_source
(leviculum-core/src/transport.rs:23826). Arm 4's floor is
test_implausibly_low_floor_recovers_in_one_adoption
(leviculum-core/src/transport.rs:23494) and its ceiling
test_timebase_floor_cannot_pass_learn_ceiling
(transport.rs:24051).
3. Anchor rank predicates
(provenance rank).
Unit, landed with the build-floor plumbing (#247): at rank 5 the
unbounded first adoption fires
(test_birth_anchor_adoption_is_unbounded_despite_a_plausible_value,
leviculum-core/src/transport.rs:23783) AND tickets are refused
(clockless_node_refuses_to_issue_a_ticket_a_peer_would_discard,
leviculum-lxmf/tests/wall_clock_producer.rs:237, which asserts
the rank and that the birth value clears the old floor); after an
anchor from a better source the per-announce cap binds AND tickets
are issued (same two cells' second halves, plus
test_clockless_timebase_advance_is_bounded_after_first_adoption,
transport.rs:23394); and the negative that catches the value-test
misfire is
test_birth_anchor_value_clears_the_window_but_not_the_rank
(leviculum-core/src/transport.rs:23747) — the build-floor value
alone never satisfies the plausibility predicate, even though it
clears the sanity window. The older
test_clockless_first_timebase_adoption_is_unbounded
(transport.rs:23445) keeps its assertions and now asserts the
rank alongside them. Integration:
leviculum-lxmf/tests/wall_clock_producer.rs remains the
structural pin that no router entry point takes a caller-supplied
wall clock.
4. Stamping asymmetry
(Stamping).
Unit, landed with #247:
test_stamps_never_run_ahead_of_the_estimate_across_anchor_changes
(leviculum-core/src/transport.rs:23878) walks a sequence of
anchor changes — birth, uptime advance, traffic adoption, a host
re-anchor, saturation — and asserts after every step that the
stamp is the estimate rather than ahead of it, that the
microsecond producer describes the same instant, and that the value
never repeats; the birth step asserts the clockless stamp is
exactly build floor plus uptime. Wire saturation is also pinned at
test_emission_secs_saturates_at_wire_field_max
(transport.rs:24136).
5. Emitted high-water (healing loop).
Unit: no re-anchor — including an arms-1–2-quality backwards
correction — takes the calendar below the highest value this
identity has emitted; the correction floors at the high-water mark
(the below-floor attempt is the negative cell). Integration:
persistence of the mark across a restart where storage exists; the
RTC-less mitigation is pinned at
test_own_announce_echo_reseeds_timebase_before_echo_drop
(transport.rs:23353). The announce-timebase consequence — a peer
keeps the newest path — is shown against a Python-shaped peer:
conformance time_high_water_path_retention, where a re-anchored
lnsd keeps its announces ordered above its own high-water and
the Python peer's path entry keeps updating.
6. Healing loop (healing loop). mvr, one
cell per branch of the re-anchor rule: a median over a cohort of
distinct senders re-anchors a grossly-forward calendar; a single
sender does NOT re-anchor a calendar past rank 5 (negative; the
rank-5 single-source adoption is rule 3's cell); a gross backwards
correction via traffic alone is refused, including a fabricated
traffic median (negative). The Sybil-style bound —
N identities from one neighbour advance the calendar no further
than the documented N × cap — is pinned at
test_timebase_walk_is_capped_per_announce_not_per_identity
(transport.rs:23977) and stays a negative cell: the assertion is
the bound, not a defence the model does not claim. RTC write-back
on re-anchor is a unit cell against a mock RTC; the real
peripheral is rule 10.
7. Echo chamber (the one honest cost).
Conformance time_echo_chamber_third_peer: two cold, build-floor
nodes adopt each other's birth clocks without harm — traffic
flows, nothing poisons — and BOTH converge once a third peer with
real time appears. The provenance transitions (birth → anchored
from traffic, unconfirmed → healed) are observable in the
structured event log of both nodes; a run that converges without
showing the intermediate state fails the cell.
8. Ingress clamp + collector
(Ingress,
Telemetry, Codeberg #239). Unit: a future-stamped
row is clamped for cursor and index and served with the clamped
value; dedup still catches a twice-delivered row despite clamping
(dedup keys are not clamped); an UNHEALED collector does not clamp
(the arming rule's negative). Conformance
time_collector_clamp_cursor: the cursor of a synced requester is
never poisoned by a future-stamped row — the cell reproduces the
pre-fix starvation shape and asserts it gone. Conformance
time_collector_unhealed_heals: an unhealed collector takes
in-window stamps as-is and heals from the very traffic it is
collecting.
9. End-to-end authorship interop. Conformance against real
Python, drop-in discipline both ways.
time_cold_authorship_python_receiver: a cold-start lnsd
authors LXMF to a Python receiver; the message is accepted and
readable — it sorts old, so the cell asserts delivery, never sort
order — and after healing, stamps are current. The reverse cell,
time_python_future_stamps: Python sends future-stamped traffic
at our stack; nothing breaks and nothing poisons — cursors are
asserted not to advance past local plausible-now, and path tables
stay intact.
10. Hardware tier — only where the subject is real.
time_gnss_seed_rig: GNSS time seeding on the rig boards, lands
with the GNSS issues (#69/#70). time_rtc_writeback_rig: RTC
write-back on a board with an RTC, lands with the write-back
implementation. Both are marked "lands with the implementing
issue", not pre-existing debt: the rows exist precisely so those
issues cannot land without their cells.
Meta-rules
- Every implementing issue carries its cells from this matrix. The issue is not done until its cells are green in the tier named here, in the same change.
- A rule without a named cell is a spec hole. The fix is a new cell in this matrix — added to it, and to the port checklist where the rule touches a port duty — before the rule is implemented, not after.
- Negative cells are mandatory wherever a rule refuses something. A refusal without a red-path test is a refusal nobody has seen fire; every "refused", "never" and "does not" in this page has its negative cell above.
Checklist for a new firmware port
- Inventory the arms. Which of the five can this platform offer? (GNSS receiver? an attached host? OS/RTC clock?)
- Implement
Clock::wall_unix_secsonly if the platform has a real wall clock. ReturningNoneis correct and engages the ranking. Never return an uptime-derived value from it. - Inherit the build floor; do not reinvent it. The build
timestamp is the sanity floor and the birth anchor, and the core
carries both (
BUILD_UNIX_SECS,leviculum-core/src/constants.rs:621, fromleviculum-core/build.rs), together with the rank-keyed adoption and ticket predicates that must move with it (provenance rank). A port that substitutes its own value test for either one regresses cold-start healing and ticket refusal in one step. A packager that wants a reproducible stamp setsSOURCE_DATE_EPOCH. - Wire every available better source. GNSS: seed from RMC UTC
through the sanity window. Attached host: implement the
control-channel frame that calls
set_wall_time_unix_secs. - Write healed anchors back to the RTC, where one exists.
- Do not touch announce learning. It comes with the core for free. Do not disable it, and do not "improve" it with local filtering of incoming timestamps — that is a semantic deviation from the reference (see the non-behaviours above).
- Never stamp cross-lifetime wire fields from
now_ms. New fields go throughTransport::emission_secs. - Expose the time source for diagnosis (see Record the source).
- Leave the pins green. The tests cited throughout this document are the contract; a correct port never needs to change them. The full rule-by-tier map is Testing the model — a port's own issues land with their cells from it.
Wire Field Semantics
Every field we write into a wire structure carries a meaning that some peer acts on. This page records the failure mode where the meaning is wrong while all our own tests stay green, the audit method that finds it, and the testing rule that keeps it found. It applies to every field we will ever add to any protocol we implement — Reticulum, LXMF, or our own.
The failure mode: self-consistent and wrong
The dangerous defect is not a malformed field. It is a field whose value is self-consistent between our writer and our reader, and means something else to a peer. Our writer produces it, our reader consumes it, both apply the same rule — the same wrong rule — and every test that exercises both halves of the misunderstanding passes. Interop tests pass too, as long as they assert that exchanges succeed rather than that generated values mean what the reference takes them to mean.
Codeberg #155 is the worked example. The 5-byte announce emission
timestamp is defined by the reference as unix seconds, and a peer
orders same-destination paths by it. We stamped process uptime. Our
own reader applied the ordering rule correctly to our own wrong
values, so a pure leviculum mesh was blind to the defect; 318 interop
tests against real Python missed it, because the exchanges all
succeeded. The damage lived only in foreign path tables: entries that
could never win the newer-emission comparison again, worsening with
every restart because every reboot restarted the value from zero. A
live rnsd was found carrying 660 of 10 983 path entries with
non-timestamp values. The full story of where wall time comes from is
in Time and Clocks; this page is about the
class, not the instance.
The audit method: four questions per generated field
For every field we generate (not fields we merely echo back), answer four questions, each with a citation:
- What does a peer DECIDE from it? Ordering, acceptance, expiry, deduplication, routing — the rule the value feeds, not the byte layout it sits in. Ask this one first: it partitions the surface. A field no peer decides anything from needs only a shape check, and the answer tells you which adverse conditions in question 4 are worth constructing.
- What does the reference put there? File and line into
reference/Reticulum(andreference/LXMFwhere applicable). - What does the reference DECLINE to put there, and why? Some guards live only in the writer, and their absence produces a well-formed field that harms the reader.
- Does our value satisfy that rule under adverse conditions? Process restart, absence of a clock, long uptime, the field at its representable limit, a peer that has been up much longer or much shorter than we have.
Question 1 is the one our old tests never asked. A field whose encoding round-trips perfectly can still fail the decision rule — the #155 timestamp round-tripped for months.
Question 3: the refusal is part of the contract
Codeberg #181 is the worked example, and it is a different shape from #155: not a wrong value, but a missing refusal to send a value. Questions 1 and 2 both pass on it. What a peer decides is clear (mine a stamp at the announced cost) and what the reference puts there is the configured cost — we wrote the same field, with the same meaning, from the same source. Only question 3 finds it.
LXMRouter.get_announce_app_data (LXMRouter.py:1033-1052) starts
from stamp_cost = None and overwrites it only when
0 < cost < 255. The reader applies no bound of its own: the
announced cost is stored unvalidated
(update_stamp_cost, LXMRouter.py:1027-1032) and passed straight to
LXStamper.generate_stamp (LXMessage.py:320), whose search loop
(LXStamper.py:199) runs until a digest meets 1 << 256-cost. At 255
that never happens. We announced whatever u8 the caller passed, so
one announce from us could wedge every Python peer's outbound queue
for our destination — with nothing in their logs naming us.
Two rules generalise from it:
- A guard in the writer implies no guard in the reader. When the reference validates on write, look for the matching check on read. If it is not there, the write-side guard is load-bearing, and omitting it is not a cosmetic deviation.
- The same refusal usually appears twice. #181's window sits both
at the emit boundary and one layer earlier in
set_inbound_stamp_cost(LXMRouter.py:378-393), where the refusal is visible in the return value. Mirroring both is what lets a caller learn, without weakening the boundary that actually protects peers.
Symmetry is worth asking about but is not automatic: our read side
now drops an announced 255 (leviculum-lxmf/src/router.rs:1181-1227)
although the reference does not, because that deviation is invisible
on the wire and to any conforming peer, and removes an unbounded loop
reachable from the network.
The frame is a field: how often we send it
The four questions are asked of a value inside a frame, but they apply unchanged to the emission itself — whether a frame goes out at all, and how many times. A peer decides from that too: a second copy of an announce it already holds is absorbed by its packet hashlist and costs it only airtime, which on a shared medium is airtime nobody else can use, and which no counter on either side reports.
Codeberg #192 is the worked example. Answering a path request, we
inserted the response into the announce table with retries = 0 —
the value the reference uses for a received announce, which it
means to rebroadcast twice (Transport.py:1867). The reference
inserts a path response with retries = PATHFINDER_R
(Transport.py:2970) and completes the entry at
retries > PATHFINDER_R (Transport.py:585-587): one transmission,
not two. Every field in both frames was correct; the second frame
should not have existed. It was found by decoding what each daemon
transmitted under one byte-identical traffic script and comparing the
two multisets frame by frame (status_parity_tests.rs, TX frame
census), not by comparing byte totals — a percentage says something
diverged, a census says what.
The mirror question: what do we refuse to read?
The four questions above are asked of fields we generate. They cannot find the mirror defect, which is being stricter than the reference on the read path: refusing a value the reference accepts. Nothing in a generated-field audit reaches it, and interop testing against Python does not either, because Python only ever produces the form we already accept. The defect surfaces only against a third implementation — reticulum-kt, microReticulum, a hand-rolled encoder — and it surfaces as silence: the message is dropped, and the sender sees a peer that never answers.
Codeberg #183 is the worked example. LXMF writes time.time(), so
payload[0] from a Python peer is always msgpack float64, and our
decoder demanded the 0xcb marker. The reference performs no type
check at all — timestamp = unpacked_payload[0] (LXMessage.py:766)
— so an integer second delivers on Python and was refused by us. The
same audit found the second half: the reference hashes the payload
bytes it received when there is no stamp (packed_payload, :753,
:762), while we re-encoded canonically before hashing, so even a
timestamp we decoded correctly would have failed its own signature.
Two rules generalise:
- The read side has a contract too, and it is the reference's accept set, not its output set. What Python's writer emits is a subset of what Python's reader takes. Auditing only against the writer measures the wrong boundary. Read the reference's decoder and enumerate what it lets through.
- Where the reference's reader keeps received bytes, keep them. Re-deriving a value that a signature or hash covers substitutes our encoder's opinion for the sender's bytes. It is invisible while every peer encodes as we do, and silent when one does not.
Refuse on write, accept on read
The two sides are not symmetric, and treating them as one rule is what produces an inconsistent codebase. Refusing to emit a value costs no peer anything: nothing conforming expects it, so the refusal is wire-invisible. Refusing to accept one costs the sender its message. So:
- A value we cannot bound the effect of goes in the writer's refusal
set. Codeberg #184 put the non-finite message timestamps there: a
NaNcompares False against everything, so it orders arbitrarily at any peer that sorts by it, and we cannot cite what a client does with it because the decision rule lives outside the reference. - The same value is still accepted on read, because a peer that sent it has already made its choice and dropping the message adds nothing.
- The exception is a value that becomes a bound on our own
behaviour — a ticket expiry we store, a snapshot field we restore.
There the reader refuses too, because accepting it hands an
unbounded quantity to a comparison that governs our resource use
(
Ticket::from_field_value,leviculum-lxmf/src/ticket.rs).
Working the method
- Grep the reference for the field's read sites before writing the test. The decision rule is in the reader, not the writer, and it is routinely in a different file from the one that emits the field.
- Extend
gen_vectors.pyrather than hand-writing expected bytes. Expected values then come from the reference's own emitter and its own decoder. Hand-written bytes encode the auditor's belief about the reference, which is the thing under test. - Check the reference submodule's actual HEAD before auditing
against it. Auditing against a remembered version produces
confident findings about code that is not what we ship. The pinned
commits are asserted by
leviculum-lxmf/tests/reference_lock.rs.
The testing rule: pin the meaning, recomposed independently
A field is verified only when a test pins its meaning, not its encoding — and the test must recompose the expected value independently, never by calling the same helper the writer uses. A test that shares the writer's helper does not test the writer; it tests that the helper equals itself, and stays green when writer and reader drift together.
The worked example of getting this right is the announce-signature
pin from the #159 audit,
announce_signature_covers_reference_byte_order_on_the_wire
(leviculum-core/src/destination.rs:2250). It takes the raw wire
bytes of a packed announce, rebuilds the signed data in the exact
order the reference composes it (Destination.py:297-298:
hash + public_key + name_hash + random_hash + ratchet [+ app_data]),
and verifies with raw Ed25519 against the key half at payload bytes
32..64 — then proves the pin bites by showing that dropping the
destination hash from the front makes verification fail. The
signature tests that existed before it could not have caught a drift:
they called verify_signature, which shares build_signed_data
(leviculum-core/src/announce.rs:108) with the writer, so a writer
and reader that both composed the wrong bytes would have verified
each other forever — exactly the #155 class, one layer up.
Where a reference value is computable offline, pin it as a known-
answer test with the reference's own output (the name-hash and
destination-hash KATs in the same audit tranche,
destination.rs:1902).
The standing instrument: lndecode
Recomposing independently once per test is the rule; lndecode is
that rule built once and reusable. It parses a raw frame into JSON
from offsets re-derived out of reference/Reticulum alone, and its
library links no leviculum-* crate at all — the independence is in
the dependency list (lndecode/Cargo.toml), not in a promise, so a
future edit cannot quietly route it back through the writer's
helpers. On an announce it recomputes the identity hash, the
destination hash and the Ed25519 signature from the wire bytes, which
is the #159 pin's method applied to any frame instead of one fixture.
Two properties matter for the audit. It reports rather than refuses:
a hop count above PATHFINDER_M, an emission timestamp holding
uptime seconds (the #155 shape), a link request signalling an MTU of
3 all decode completely and land in a warnings array — a decoder
that rejected adversarial frames would be useless on exactly the
traffic worth reading. And it answers the signature question twice,
because the permissive Ed25519 verifier the mesh applies accepts an
all-zero key and an all-zero signature for some messages:
signature_valid is what a peer decides, signature_strict_valid is
whether that decision means anything.
Its own agreement with the writer is asserted in
lndecode/tests/agrees_with_the_writer.rs, on packets leviculum-core
produced rather than on hand-built bytes — an oracle nobody calibrates
is just a second opinion.
Deliberate non-behaviours get pins too
When we intentionally do not do something — usually because the
reference does not and doing it would desynchronise us — that
non-behaviour is itself a semantic contract, and it gets a pinned
test with the reference citation in the test's doc comment. A later
"improvement" then breaks a test whose comment explains why the
missing behaviour is deliberate, instead of silently shipping a
semantic deviation. The worked examples are the two time
non-behaviours — we emit our own wall clock verbatim and we do not
plausibility-check incoming emission timestamps — pinned with their
Destination.py/Transport.py citations; see
Time and Clocks.
Where the audit stands
The systematic sweep over every generated field is Codeberg #159 —
the issue, not this page, is the source of truth for its state. As of
2026-08-03 all four tranches are done: the announce layer (tranche 1,
pins in leviculum-core/src/destination.rs and transport.rs), the
link and resource layers (tranche 2, pins in
leviculum-core/src/node/mvr_generated_field_pins.rs and the resource
modules), the transport layer (tranche 3, pins in the same two files),
and LXMF (tranche 4, pins in
leviculum-lxmf/tests/generated_field_pins.rs and
leviculum-lxmf/tests/wall_clock_producer.rs).
Tranche 2 found two fields that failed the audit and were fixed
red-first: the request timestamp carried process uptime (#164) and the
resource advertisement sent a content hash where the reference sends
the salted per-transfer hash (#165). Tranche 3 found four routing
defects (#168, #169, #170, #172). Tranche 4 found that the announced
LXMF stamp cost was not clamped to the reference's 0 < cost < 255
window (#181, fixed red-first, and the origin of question 3 above),
and that the LXMF crate resolved none of its wall-clock wire fields
through Transport::emission_secs — it took them from a caller
parameter instead (#182, fixed: the router now resolves them from the
NodeCore it holds, and refuses to issue a ticket whose expiry it
knows a peer will discard).
Working tranche 4 also turned the method around and asked the mirror
question above, which produced two more: we refused every
payload[0] that was not float64 and re-hashed the payload instead
of keeping the received bytes (#183, fixed red-first, and the origin
of that section), and Message::create signed NaN and ±Inf
timestamps while a dozen other sites in the same crate refused them
(#184, resolved by refusing on write and continuing to accept on
read). Pins for both are in
leviculum-lxmf/tests/foreign_payload_encodings.rs and
generated_field_pins.rs, backed by VEC-MSG-FOREIGN-* vectors that
record the reference decoder's own verdict.
The recurring lesson across all four: the offenders were timestamps and identifier-derivation order, never framing. Nothing that round-trips was ever wrong; everything that a peer compared against a value from another machine was worth checking — and, from #181, everything the reference deliberately declines to send, and from #183, everything the reference declines to require.
See also
- Time and Clocks — the #155 instance in full: where wall time comes from and the hardening around it.
- Python-RNS Compatibility — why the reference's decision rules, not its internals, are the contract.
Regulatory Airtime
Unlicensed LoRa bands are shared under duty-cycle rules. This page records where the limit is enforced, what a node does when nobody configured one, why no radio setting is ever refused for a regulatory reason, what it takes to switch the limit off, and one measurement pitfall. It is a durable rule for every radio firmware we write, present and future.
Enforcement belongs in the firmware, not the host
The firmware is the only place that knows what actually went on the air: retransmissions, preambles, frames queued by a host that has since crashed — none of that is visible from above. A host-side budget can shape traffic, but only the modem firmware can enforce a duty cycle, because only it stands between the queue and the antenna.
The RNode firmware is the model: it accounts every transmitted
frame's airtime into rolling bins, raises airtime_lock when the
short- or long-term limit is exceeded
(reference/RNode_Firmware/RNode_Firmware.ino:1673-1675), and gates
the transmit queue on it — if (!airtime_lock && queue_height > 0)
(RNode_Firmware.ino:1624). The limits arrive from the host as
CMD_ST_ALOCK / CMD_LT_ALOCK (Framing.h:36-37), but the
enforcement never leaves the device.
Our LNode firmware enforces the same way: AirtimeTracker
(leviculum-core/src/rnode.rs:1832) mirrors the RNode ledger, and
the nRF TX path holds a queued frame instead of keying the radio
while the tracker is locked (is_locked,
leviculum-nrf/src/lora.rs:1902-1989), continuing to listen so RX is
not starved — until the frame has waited so long that keying it would
be pointless, which is
the age rule below.
The host-side airtime credit bucket
(leviculum-std/src/interfaces/airtime.rs, see
Interface Isolation) is backpressure, not
regulation: it keeps the serial queue from absorbing minutes of
backlog. It is a comfort for the stack, not a legal control, and
nothing may treat it as one.
Lawful by default
A node that is not told otherwise obeys the band it is on. When no
airtime_limit_long is configured, the host derives the lawful
long-term limit from the TX frequency (resolve_lt_alock,
leviculum-std/src/driver/mod.rs:513-547) and sends it to the modem; a
standalone LNode whose host never sent one derives it in the firmware
from its own frequency (firmware_default_lt_alock,
leviculum-core/src/rnode.rs:1638). Both read the same table,
etsi_eu868_duty_cycle (leviculum-core/src/rnode.rs:1524), which
carries the EU 863-870 MHz sub-bands with their 0.1 % / 1 % / 10 %
duty cycles and the 433.05-434.79 MHz band at 10 %. An explicit
configured value always wins — including an explicit 0, which the
firmware reads as unlimited.
A cap that cannot be read back is not a cap anyone can check. The
firmware states the settings it applied and the limits it loaded into
the tracker on the boot-critical log path — the one that bypasses the
debug port's runtime drain gate (airtime_limits,
leviculum-nrf/log-line/src/facts.rs:421) — and states them again on
every runtime reconfiguration. Until 2026-08 both were ordinary
runtime lines: a board that came up before a reader attached dropped
them with everything else, so the two facts a compliance question is
actually about were the two that could never be obtained from a
running board. Neither is recoverable any other way — the settings
live in the radio's registers and the cap in the airtime tracker, and
nothing reads either back out.
The limits line is unconditional, and it names an origin per limit.
It used to be emitted only when the firmware had derived the cap
itself, which left the more dangerous case silent: a host that sent
an explicit 0 switched the cap off and produced no line at all, so
the cap in force had to be inferred from an absence, and a board
legitimately unlimited on a shielded bench read exactly like one
unlimited in the field. It now carries both limits with the raw u16
and a human rendering (lt_cap=unlimited versus lt_cap=0.10% — the
one confusion on this line with a legal consequence), whether each
came from the host or was derived, and the lawful cap the frequency
alone would give, so a host's choice can be weighed against the band
without looking a sub-band up in this page. grep AIRTIME on a fresh
boot answers "under what cap is this board transmitting, and who
chose it".
Every row of the table has been verified against the standard text: ERC Recommendation 70-03, Annex 1, sub-bands h1.3-h1.9 for 863-870 MHz and the 433.05-434.79 MHz entry of the same annex. The duty cycle is also the only compliance route open to a fixed-frequency LNode: every sub-band's requirement reads "≤ x % duty cycle or LBT+AFA", and AFA — adaptive frequency agility, changing channel — is impossible here by construction.
One honesty note, deliberate: the table covers only the bands above. Other bands (US 902-928, AU/NZ, ...) have no citable source in this tree, so they get no auto-limit and a warning that says so — a limit invented from memory would read as authoritative to exactly the operator who most needs it not to be. Supply the citation and the table grows.
TX power follows the same lawful-by-default shape (resolve_tx_power
capped by lawful_erp_dbm, leviculum-core/src/rnode.rs:1570): an
absent txpower asks for the board maximum, capped by the sub-band's
e.r.p. limit — 25 mW everywhere in the European SRD spectrum except
500 mW in 869.4-869.65 MHz and 10 mW in 433.05-434.79 MHz. An
explicit txpower wins even above the cap (the operator may hold a
licence or know the jurisdiction); the excess is logged. The
narrowband bands between the wideband sub-bands (868.6-868.7 MHz
and its four siblings, alarms, ≤ 25 kHz channel spacing) fit no LoRa
bandwidth this stack configures, so a carrier that overlaps one is
warned about by name at interface build (erp_band_gap,
leviculum-core/src/rnode.rs:1513) — falling through to "no known
limit, board maximum" without a word would be the most permissive
outcome exactly where the operator most needs to be told. The
carrier is then honoured; see No radio configuration is
refused below.
Python-Reticulum does not do lawful-by-default; the cap only shapes local TX and is invisible to receivers, so this is a Priority-1 enhancement under the deviation rule.
No radio configuration is refused
Every radio setting this stack is given is honoured. A setting that looks unlawful for a region is warned about, loudly, by name — and then applied. That is project policy, decided 2026-08-16, and it supersedes the hard band-gap error this page used to describe.
Two reasons, and the second is the stronger one:
- The jurisdiction is not knowable from here. The same carrier is lawful under a licence, in another region, on an amateur allocation, or in a shielded chamber with dummy loads. A check that reads a frequency cannot tell those apart from an unlawful deployment, so it would refuse the lawful cases too.
- The operator is the responsible party. In the EU it is the operator, not the software author, who answers for compliant operation. Software that refuses a setting takes on a responsibility it does not hold, and hands the operator a daemon that will not start instead of the information they need. Our job is to make the consequence impossible to miss, not to make the choice.
The warning is emitted at WARN, never at debug: a decision narrated below the default log level is the silent substitution this policy exists to prevent.
What stays a refusal is anything with no regulatory content in it — the SX1262's 150-960 MHz tuning range, the ten bandwidths the modem has a register code for, the 0..=37 dBm field of the RNode wire protocol, the SF and CR ranges shared with Python-RNS, and the SoftDevice version guard that keeps a flash from bricking a board. Those are arithmetic and device protection, not paternalism: they describe what the hardware can be asked for at all, and honouring them is not a judgement about anybody's licence.
Prose alone has drifted twice here — the code once, this page once —
so both halves are mechanical now. The code is pinned behaviourally
by no_radio_configuration_is_refused_for_a_regulatory_reason
(leviculum-std/src/driver/interface_build/mod.rs:690), which drives
the known regulatory edge cases through the config-building entry
point and asserts each one builds and warns at WARN, with a second
half pinning the capability refusals so the first cannot be satisfied
by deleting every check. This page is pinned by
the_book_describes_the_band_gap_as_a_warning_never_a_refusal
(leviculum-std/tests/doc_radio_policy.rs:198).
Disabling is an operator act, not a test convenience
Switching the limit off is sometimes legitimate — a shielded bench with dummy loads, a throughput scenario that cannot measure what it exists to measure at 1 % duty. But it is an operator decision with a paper trail, never a default and never a convenience:
- It requires a written justification. Periculum's
[disable_airtime_lock]section refuses to parse without one (periculum/src/topology.rs:258,DisableAirtimeLockDef). - Every run that had the limit off must say so — in its terminal
output (the airtime banner prints the rendered limit per frequency,
whichever route produced it) and in its result document (the
measurement cell records the policy, the rendered limit, and the
lawful limit for that frequency,
AirtimeContextinpericulum/src/bench.rs, with the source — scenario or rig — recorded in the results).
Of the three possible outcomes, a silent green under a lifted limit is the worst:
- Red under the lawful limit is honest: the design exceeds the band's budget, and the result says exactly that.
- Green with a declared lifted limit is honest: it measures the stack, not the law, and every reader can see which.
- Silent green under a lifted limit is a lie with a green checkmark: it reads as evidence that the system works lawfully when it never once ran under the law. It also poisons comparisons — a figure taken with the lock off next to one taken with it on is a comparison of the lock, not of the stack — and it ships that lie forward into every document that cites the run.
Where the bench-level switch lives, and why
The blanket switch for a whole bench lives in the Periculum rig
profile (rig.toml, periculum/src/rig.rs) — site data, not
scenario data. Containment is a property of the site: whether the
bench is shielded and on dummy loads is true of THIS rig, not of a
scenario file that travels between benches and operators. A scenario
that must not run unlimited even on such a bench can carry
[require_airtime_lock], which wins. The mechanics are Periculum's
to document; the durable rule here is only the split: scenario files
describe the experiment, the rig profile describes the site, and the
airtime carve-out belongs to the site.
The measurement pitfall: reading the meter restarts it
The duty-cycle history lives in RAM, and on ESP32 targets the RNode
firmware's startRadio() zeroes it: it calls init_channel_stats()
(RNode_Firmware.ino:523), which clears the airtime bins and both
utilisation figures (reference/RNode_Firmware/Utilities.h:1858). So
a diagnostic that starts (or restarts) the radio in order to read the
airtime counters measures nothing — the act of taking the reading
destroyed the reading. We hit this in practice.
The general lesson is not radio-specific: a diagnostic must not disturb what it measures, and a diagnostic that can must be checked for it before its numbers are believed. See Evidence and Honesty in Testing.
A hold ages traffic; it does not thin it
The gate above holds the frame at the head of the queue and re-checks. That is FIFO under a hold, and the consequence is worth stating plainly for anyone reading a LoRa capture: a duty-cycle hold does not thin traffic, it ages it. Each dip of the ledger below the cap admits exactly one frame, which re-pins the lock, so under sustained load the queue drains at the cap's rate with its order intact and the frames that reach the air are as old as the standing backlog. Nothing is lost and the airtime stays lawful — but what the lawful airtime carries is history.
Measured on the WisMesh Pocket V2 during the field test of
2026-09-27 (docs/measurements/2026-09-27-field-test-lora-chain-columba.md,
outside the book because it is a measurement record, not a rule):
pinned at a 10 % long-term cap, one frame left every 10–15 s and the
three relayed link requests in the window were keyed 145.6 s, 137.7 s
and 144.0 s after the stack handed them over. A forwarded link
request is routable only until the relay's link-table entry expires,
(hops + path_hops + 2) × 6 s — 30 s in that topology — so every
proof came back to an entry that had died about 115 s earlier.
So the interface drops what it has held too long, at the point where it waited (Codeberg #433):
- The rule. At dequeue, while
AirtimeTracker::is_lockedis true, a frame whose age since the interface accepted it exceedsleviculum_queue_budget::HOLD_MAX_AGE_MS(18 s) is thrown away instead of keyed. The 18 s is derived, and the derivation is in that constant's own doc comment: it is the shortest entry any relay grants,(1 + 0 + 2) × 6 sfor a one-hop request to a destination the relay reaches directly, so a frame older than that is dead on arrival in every topology; it sits above the 15 s short-term airtime window, so a lock that engaged on short-term airtime alone never loses a frame it was about to release, and below the measured field topology's 30 s entry deadline minus the measured 1.0 s return leg. - Both stacks.
lnsddriving an RNode cannot read the modem's lock, so its send queue applies the same constant against what it can see: the CMD_READY gate held shut, with frames waiting, for longer than a full frame's airtime plus one re-query interval explains (DutyHolds, in the same crate). A frame past the cap whose wait overlapped such a hold is dropped at dequeue withLORA_TX_STALE iface=<name> age_ms=<n> len=<n> held_ms=<n>and counted astx_stale_drops(also insidetx_queue_drops), whichlnstatusshows asTX stale. Frames the modem already holds age inside it, out of the host's reach. - Only under the lock. A frame that waited for CSMA, for an acquisition-jitter draw or behind a burst gap is not stale in this sense, whatever its age: those waits are the interface's own and end by themselves. Only the regulatory lock holds a frame for minutes.
- Type-blind. The interface reads the age and never the packet, so a link request, an announce and a resource chunk of the same age get the same verdict (Interface Isolation). What makes the drop acceptable is not a judgement about the packet but the cadences above it: a link request is reissued every 6 s per hop, an announce on its own cadence, LXMF at its own layer — a frame older than 18 s behind a lock has been superseded or written off by its sender already.
- Counted, never silent. Every drop raises
LORA_QUEUE_DROP reason=stale age_ms=<n> bytes=<n> total=<n>under the[LORA]prefix, rate-limited like the core's[DROP]lines with aLORA_QUEUE_DROP suppressed=<n> window_ms=<n>line for what a clipped window held back, and the running total is thelora_stale=field on the periodic[TRANSPORT]line. The counter is how the fix is read on the air:lrproof_no_linkon a relay should fall toward zero for through-traffic aslora_stalerises.
What this does not do is create airtime. The cap spends the same
milliseconds it spent before; the change is that they carry current
traffic instead of fossils. A board that is permanently over its cap
is still a board with too much to say, and the honest reading of a
climbing lora_stale= is a load problem, not a solved one.
A relay in duty hold advertises no route — and the hold provably lifts
The same arithmetic binds the control plane (leviculum#493, decided on #433). A relay whose interface is in duty hold cannot carry a link setup under any stack's clocks — initiator 20 s, forwarding entry 24 s ours and 12 s Python's, against two held frames of up to 36 s — so an announce it relays there advertises a service it cannot render, and every initiator behind it spends 20 s per attempt learning that. The rule has two halves, and the second is the guard on the first:
- While the hold is in force, the interface advertises nothing.
The interface reports its hold to the core as a boolean (
duty_hold, mirrored like the online flag — the core computes nothing), and the core's announce admission refuses rebroadcasts and path responses alike on a held interface, logged asANN_TX_SUPPRESSED … closed=duty_hold. The refusal is final, not queued: the destination re-announces on its own cadence, and a queue released at lift time would key routes exactly as old as the hold. The node's own announces are exempt — the #402 announce cap governs them, priced against the duty budget by #401 rule 5. Python is the precedent for the half that refuses: an interface over itsannounce_caprelays no announce throughTransport.outbound()at that moment (Transport.py:1252-1294) — though Python queues what we refuse, which is the one place we deviate, per the stale-route argument above. - The hold ends when the rolling window frees budget, visibly. The
firmware states the edge once —
[LORA] hold lifted lt_ms=<n> held_for_ms=<n> dropped_stale=<n>beside the per-turn[LORA_AIRTIME_LOCK] … holdingline — and lnsd emitsDUTY_HOLD iface=<name> state=held|lifted lt_ms=<n>on the same two edges. The lift is pinned by host tests on the real ledger: flat at the cap at t0 it holds, nothing inside the rolling hour frees it, and the first bin that leaves the hour drops the long-term sum below the cap and readsduty_holdfalse again, also when the loop was parked in a receive window across that edge (leviculum-nrf/queue-budget/src/tests.rs, and the end-to-end mvr inleviculum-std/tests/mvr/). That holds because the ledger retires every bin a call skipped, not only the bin after the current one as the reference does (RNode_Firmware.ino:688,:698): a bin no call landed on would otherwise keep an hour-old charge and hold the lock on air spent in the previous hour (leviculum#495). The wire sees none of this.
Three topologies set that boolean, and each from the layer that owns the carrier:
- The board as a node. The firmware's own core reads the lock gate's
flag (
LORA_DUTY_HOLD) on every main-loop wake and mirrors it onto its LoRa interface. - lnsd driving an RNode. lnsd's RNode interface watches the modem's
CMD_READYflow control, and a gate closure that qualifies as a hold (DutyHolds::holding) raises the flag. - lnsd driving an LNode over its serial protocol (leviculum#501). The
board's lock gate holds, but the host's interface is a plain serial
port with no flow control, so the board tells it: a
DUTY_HOLDframe on the USB control envelope on each edge, and the current state behind every media report, which lnsd asks for each time the port attaches. lnsd's serial interface sets the flag from it, logs the sameDUTY_HOLDline as the RNode path, and clears it when the port goes down. Before this, a relay whose modem was locked for 261 s of a 310 s run measuredduty_hold=never ann_suppressed=0on the host (497): the board knew, and the protocol between board and host had no word for it. The frame is USB protocol between our board and our host (docs/src/firmware/usb-control-envelope.md), never on the air.
A propagation node leaves half its budget to forwarding
The hold above is the end state; the second half of the same decision
(leviculum#494, Lew 2026-10-07) keeps a propagation node from walking
into it on its own traffic. The field relay of 2026-09-27 was a
propagation node: its own sync rounds filled the lawful budget, the
hold followed, and the link setups of everyone behind it died there
(#433). So a board running the propagation role originates nothing
once the rolling hour has spent half its long-term cap
(OWN_TRAFFIC_SHARE_PERMILLE, 500, in
leviculum-nrf/queue-budget/src/lib.rs). The split is by origin, and
it is decided where the traffic is originated, not where it is keyed:
the LoRa queue still treats every frame alike, and what the node
forwards as a transport node never meets the share, so forwarding
always has the other half.
The engine asks before it starts a sync round, before it identifies
and offers on the round's link, before it sends the transfer the peer
asked for, before an active delivery, and before it accepts an inbound
offer, which past the share is answered with the reference's own
postponement, ERROR_THROTTLED, on which a stock peer waits and offers
again. Answers to clients (/get, upload proofs) are not its own
initiative and are not gated; its announces are priced against the
budget by #401 rule 5. A refusal is stated once, on its edge, as
PN_YIELD reason=duty_share lt_ms=<n> cap_ms=<n> share=<permille> site=<first site>, and every site retries on its own cadence. The
used figure is the LoRa task's ledger, published once per loop turn.
The firmware ledger is not a cross-session account
The same fact has a second consequence, and it is the one that decides where an hour-scale budget lives. Because a radio start clears the bins, the firmware's long-term figure covers airtime since the last radio start, not the rolling hour. It is a lower bound, and the bound is zero exactly when the question is worth asking: a harness that reboots a board to give a test a defined starting state (Periculum does, before every scenario that binds one) has zeroed it, and the daemon under test zeroes it again when it brings the radio up. An offline radio's history survives in RAM and cannot be read out-of-band at all — the only way to make the firmware emit it is the call that clears it first.
So: enforcement belongs to the firmware, but the hour-scale account belongs to whoever drives the radio. The board is the only thing that can refuse to transmit, and the only thing that cannot tell you what it transmitted an hour ago. Anything that needs to know — a test harness spacing its runs, a scheduler shaping traffic — keeps its own ledger and states plainly that the figure is modelled, not measured, and a floor rather than a total. A reset makes the board forget what it radiated; it does not make the airtime unspent.
See also
- Interface Isolation — why airtime backpressure is host-side and per-interface while airtime enforcement is firmware-side.
- Python-RNS Compatibility — the deviation rule that lawful-by-default satisfies.
- Evidence and Honesty in Testing — the wider discipline behind "say so in the output".
lnmsg: a terminal LXMF messenger
Status: design record. All ten open questions are decided.
This began as a draft for argument in which nothing was decided. Every significant decision was presented as options with trade-offs and a stated preference, and the preference was an opening position rather than a conclusion. Between 2026-08-08 and 2026-08-10 all ten questions were settled, so the document graduates from discussion to record and the remaining work is issues and batches.
The rejected alternatives are kept throughout, on all five pages. They are the why: a decision recorded without the options it beat is an assertion, and the next person to ask "why not SQLite / why not modal keys / why not a daemon" has to re-derive the argument from nothing.
- Architecture — the library it stands on, the driver seam that nearly forces the shape, process architecture, the TEA split, scriptability, and the library gaps.
- The user interface — the input model, and what honesty about delivery looks like.
- Conversation storage — SQLite, the schema, retention.
- The mailbox and who you talk to — propagation-node policy, trust, contacts and names.
1. What the program is
A terminal client for reading and sending LXMF messages over Reticulum,
connected to a running lnsd or rnsd shared instance, the same way
lnomad connects.
It must talk to a propagation node. A propagation node is the mailbox that holds messages addressed to a client that was not reachable. Without it, a laptop that is closed for eight hours simply does not receive mail, and the program is a toy. With it, the program is usable on hardware that is off most of the time, which is the actual deployment.
Scope explicitly excluded: hosting a propagation node. The library side
of hosting has since been written: leviculum-lxmf carries both ends of
the client to node exchange and the node role that stores uploads and
answers /get (leviculum-lxmf/src/lib.rs:16-21, leviculum#384 part 1),
and the node to node direction, peer sync and the /offer path, in its
peering module (leviculum-lxmf/src/lib.rs:23-34, leviculum#384 part 2).
The crate still performs no I/O, so a host owns links, resources and
stamp-validation scheduling, and the two client modules lnmsg is built on
keep hosting out in their own headers
(leviculum-lxmf/src/router/propagation_runtime.rs:3-5,
leviculum-lxmf/src/propagation_client.rs:5-8). Hosting is a separate
program and a separate argument — and it now exists: the channel design record
(Public channels over LXMF) reopens hosting,
because the channel retrieval side only works on nodes we run.
The name
The house convention is one ln* counterpart per reference tool
(Client tools, "A counterpart for every reference
tool"), and the existing family is lnsd, lnstatus, lncp, lnomad.
There is no single reference tool to be a counterpart of: LXMF messaging in
the reference world lives inside NomadNet and Sideband, not in a standalone
utility.
Decided 2026-08-10: lnmsg. Rejected: lnmail promises mail while the
UI is chat (the input-model decision); lnchat promises
chat while the protocol delivers mailbox behaviour on slow paths.
"Message" is the only word the protocol can always honour, and lnmsg send
in a cron job explains itself.
2. The ten decisions
| # | Question | Verdict | Where |
|---|---|---|---|
| 1 | Single process or daemon plus client | C built as A first: one process now, the router and store behind an interface so daemon mode is a wiring change. The lnsd-resident variant is rejected twice over. The core/frontend boundary is binding from day one, so a GUI is a frontend swap. | Architecture |
| 2 | Modal, always-insert, focus-follows-pane, or prefix | C, focus-follows-pane, with a command palette and no modes. Enter in compose sends — the quiet keyboard's one named exception — Alt-Enter newline, empty-buffer guard, send_on_enter switch for email style. Ctrl-C clears to a recoverable draft and never quits. | UI |
| 3 | SQLite or a pure-Rust store | SQLite, via rusqlite with bundled. The musl static build was tried and works, FTS5 included. Identity-scoped schema, two timestamps, raw msgpack fields, attachments out of line. | Storage |
| 4 | Default sync interval, adaptive or not | Asymmetric adaptivity: fifteen minutes as a hard never-exceeded upper bound, faster (about two minutes, decaying) after activity. Silent lengthening rejected as a breach of trust. | Mailbox |
| 5 | Trust: three states or four, derived "suspicious"? | Exactly three — known, unknown, blocked. Trust follows from having named someone. No suspicious state: name collisions and name changes are normal mesh life, handled by display disambiguation. | Mailbox |
| 6 | Delivery ledger: visible, keypress, or debug? | One keypress away in the per-message detail view, plus a one-cell coloured state glyph per message, all from one shared state table, with an ASCII fallback. | UI |
| 7 | Does lnmsg send block? | No. Exit 0 means "queued cleanly" and claims nothing more. A success prints nothing (amended 2026-08-21: the ID came off stdout; it stays reachable as LNMSG_ENQUEUED id=… in the event log). --wait opts into blocking with timeout-is-not-failure exit codes; lnmsg status <id> answers later. | Architecture |
| 8 | Retention policy | Text forever, attachments under a total-bytes budget (about 500 MB by default, oldest evicted, the message row keeps name, hash and size and says so). Ledger rows share the message's retention. A per-conversation age cap is deliberately deferred. | Storage |
| 9 | The name | lnmsg. | above |
| 10 | Close library gaps first or discover by building? | Triage. Gaps 1, 2, 3 and 8 — announce names, the propagated-awaiting-collection state, file-backed storage, Display on errors — close in one library wave before lnmsg starts. The other eight become issues met in build order. | Architecture |
3. Scale and honesty about hardware
This runs on machines from a Raspberry Pi upward, which sets a few hard numbers.
Memory. The whole model plus the visible window, not the whole history.
lnomad's eager full-page layout (lnomad/src/tui.rs:910-934) is
acceptable for a page and not for a conversation. Attachments in a
byte-budgeted cache, not a count-budgeted one
(lnomad/src/image_cache.rs:11-14).
Disk. SD cards die from writes. This argues against NomadNet's
file-per-message plus a full .index rewrite on every change, and for a
store that appends. It also argues for batching persist() rather than
calling it on every PersistenceRequested, at the cost of a bounded window
of loss.
CPU. Proof-of-work is the one unbounded computation in the system. On a Pi it must run off the core lock and at low priority, and the UI must be able to say "this will take a while" with a number that was measured on that class of hardware.
The core lock is the real budget. 5 ms per hook call
(PROCESSOR_TICK_BUDGET, leviculum-std/src/driver/processor.rs:181),
against 3.2 ms to verify a 1 MiB message's signature per
The core lock budget. A messenger doing anything
expensive inside the hook stalls the node's inbound path for every other
client of the shared instance. Everything that is not the router's own
state machine belongs on the far side of a queue.
Airtime is the scarcest resource of all. Every design choice on these pages that trades bytes for clarity — a sync poll, an announce, a read receipt we are not going to invent — is spending the one thing that cannot be bought back.
4. Compatibility constraints
Non-negotiable, per the project's first priority.
- No wire-format changes. Everything on these pages is expressible in LXMF as it is. The delivery ledger, the collision warning, the cost estimate and the mailbox screen are all local views over data the protocol already carries.
- No new fields. If threading or replies are wanted, they use
FIELD_THREAD (0x08),FIELD_REPLY_TO (0x30)andFIELD_REPLY_QUOTE (0x31)with the reference's semantics (reference/LXMF/LXMF/LXMF.py:15,:23-24), not a private encoding. - Unknown fields round-trip, as the library already guarantees
(
leviculum-lxmf/src/message.rs:5-8). A reply composed by us to a message from a newer client must not silently drop what we did not understand. - Names on the wire are bytes, and we sanitise for display without altering what we forward.
- Interoperability is tested, not assumed. Per the project's test
discipline, this program needs interop tests against real Python LXMF
peers and a real propagation node, positive and negative, before it is
called finished.
leviculum-lxmf-nodeexists precisely to make such A/B comparisons possible with one driver (leviculum-lxmf-node/src/lib.rs:12-18), and the same trick applies here.
5. Provenance
Prior art, condensed
Two programs were read end to end before any of this was decided, and most of the arguments on the other four pages are arguments with one of them.
NomadNet. Storage is one directory per conversation named by the peer
hash and one file per message named by the LXM hash, with unread and
failed side-car counters and a msgpack .index that caches timestamp,
state, title and the full content of every message
(NomadNet's Conversation.py, lines 62-71, 120, 236, 93-109 and 944-960).
Attachments are stripped out into storage/attachments/<hash>/ with a
manifest and sanitised names, which is good. Everything else about the
storage is not: scan_storage() does a full listdir and constructs a
ConversationMessage for every file on every call, and
update_message_widgets() rebuilds every widget and replaces the whole
listbox on every change (Conversations.py, lines 2254-2291). NomadNet
commit 7bc6911 exists specifically because calling the conversation list
on every announce caused file-descriptor starvation during announce storms.
There is no pruning, no paging and no search anywhere.
Its chat pane is a two-column layout with the editor in the frame footer
and initial focus on it. Composition is an inline MessageEdit with a
readline mixin; there is no $EDITOR integration, and send is Ctrl-D.
Every command is Ctrl-modified, so plain typing always reaches the edit
box, at the cost of an exhausted and context-overloaded Ctrl namespace:
Ctrl-X is "delete conversation" in the list and "clear history" in the
body, Ctrl-U is "ingest URI" and "purge failed", Ctrl-P is "my QR" and
"paper message". Delivery state is glyph and colour only: the distinction
between SENT, DELIVERED and propagated-SENT genuinely exists and gets
three different styles, but it is never stated in words, and propagated
shares a style with paper messages. Two honest touches: signature failure
is rendered in plain English ("Unknown Origin", "Invalid Signature"), and
on load a message stuck mid-flight that is no longer in the router's
pending set is forced to FAILED.
Identity handling is where NomadNet is genuinely ahead of everything else:
four trust levels with real UI consequences, and, uniquely, duplicate
display names raise WARNING. Notification is a terminal bell, literally
sys.stdout.write("\a").
columba. A Room database, fifteen entities, and two schema decisions
worth taking. Every row is scoped by identityHash with composite primary
and foreign keys, so multiple local identities coexist with cascading
isolation. And receivedAt (local clock) is stored separately from
timestamp (sender clock), with sorting on
COALESCE(receivedAt, timestamp) and an explicit comment that this is
immune to sender clock skew. NomadNet has the same problem and solves it
worse, as a user-toggled sort mode on Ctrl-O. Paging3 with a DESC query
and a reversed layout is the right answer to NomadNet's rebuild-everything
problem.
Its propagation handling adds three things NomadNet lacks: a
user-configurable sync interval, real failover to an alternative relay when
one dies, and a structured SyncResult / SyncProgress sum type where
manual syncs report loudly and periodic ones stay silent. The relay is
modelled as a contact with an isMyRelay flag, which is a nice way to make
the relationship visible. Contact status is a persisted enum — ACTIVE,
PENDING_IDENTITY, UNRESOLVED — rather than a live key lookup, and name
resolution has a documented precedence: in-memory cache, then the user's
own nickname, then the announced name, then the stored conversation name.
User nickname beating announced name is exactly right.
What columba gets wrong: it has no trust model at all. A search for
trustLevel / isTrusted across its app, domain and data modules returns
nothing outside tests, there is no duplicate-name detection, and
consequently no trust gate on relay auto-selection. It will happily adopt
the nearest stranger's relay. Its delivery display also collapses sent
and propagated into the same single check mark, which is precisely the
distinction that matters most on a delay-tolerant network. Its per-message
detail screen, showing delivery method with a one-sentence explanation plus
hop count, interface, RSSI and SNR, is the best idea in either program and
is trivial to do in a terminal.
Neither program prunes history.
The citation convention on these pages
Claims about existing code carry file:line. Where something could not be
verified, it says so.
In-tree citations were re-pinned against this tree when the concept was
promoted (2026-08-10); the draft had been written days earlier and the
leviculum-lxmf crate had moved substantially under it. The citation guard
(leviculum-std/tests/doc_citations.rs) resolves every one of them.
Claims about the vendored references use the same form under reference/,
for example reference/LXMF/LXMF/LXMRouter.py:38, which is what every
other page in this book does.
NomadNet and columba are not vendored here and are not in this repo, so
the guard cannot resolve a citation into either and would report one as a
file that had been deleted or renamed. Their line references are therefore
written with the file name backticked and the lines outside it — NomadNet's
Conversation.py, lines 62-71 — which says exactly as much and does not
claim a path this tree holds. Every such reference names NomadNet or
columba in the same sentence, so a reader always knows which tree to open.
lnmsg: architecture
Part of the lnmsg design record. This page carries the ground
truth the design stands on — lnomad and leviculum-lxmf — the driver
seam that nearly forces the architecture, the process and event-loop
decisions, scriptability, and the triage of what the library does not yet
expose.
1. lnomad, and one correction to the premise
lnomad is described as supporting "Emacs keybindings, vi keybindings,
Firefox keybindings and Firefox mouse behaviour, all at the same time".
That is the observable behaviour, but it is not implemented as four
schemes. It is one keymap with three resolution mechanisms, and only one of
them is table-driven.
What is table-driven. SCROLL_KEYS (lnomad/src/tui.rs:3159-3291) is
a static table of ScrollKey { keys, desc, chords } where each
ScrollChord { code, mods, cmd } carries a modifier class
(ScrollMods::{Any, Plain, Ctrl, Alt}, lnomad/src/tui.rs:3128-3135). One
row carries the vi and the emacs and the arrow spelling of the same motion
side by side:
#![allow(unused)] fn main() { keys: "j / k ↓ / ↑ Ctrl-n / Ctrl-p", // lnomad/src/tui.rs:3161 desc: "scroll a line", }
Resolution is a linear scan, first match wins
(key_to_scroll, lnomad/src/tui.rs:3293-3311). The table is read by both
the key handler and the help overlay (lnomad/src/tui.rs:4669-4680), and
the doc comment says that is deliberate: "the SINGLE source of truth read
by BOTH" (lnomad/src/tui.rs:3147-3151).
What is not. Everything else is a hand-written if-chain in
update_browse_key (lnomad/src/tui.rs:1810-1957): roughly twenty
sequential if key.code == ... { return ...; } statements. There is no
binding map, no user-configurable keymap, no keybinding config file. The
help overlay's non-scroll groups are a second, unlinked static list
(lnomad/src/tui.rs:4687-4788) that can silently drift from the handler.
How the conflicts are actually resolved. Three layers, in this order.
- Global escapes, before any mode dispatch
(
update_key,lnomad/src/tui.rs:1566-1606): any key dismisses the toast;Ctrl-Cquits from anywhere; an open help overlay swallows everything; an open places panel takes over. - Mode gating.
Mode::{Browse, Address, Hint, Search, Field}(lnomad/src/tui.rs:292-311), each with its own handler. Text modes forward unclaimed keys to atui_input::Inputeditor.Mode::Fielduses a whitelist rather than a catch-all so that "field editing never leaks into browse hotkeys" (lnomad/src/tui.rs:1614, whitelist at:1643-1652). - Modifier discrimination. Nearly every single-letter binding is guarded
&& !ctrl && !alt, which is what letsf(hint mode) andCtrl-f(page down),d(places) andCtrl-d(half page),n(next match) andCtrl-n(line down),g(top) andCtrl-g(cancel) all coexist.
The ordering is load-bearing. In browse mode key_to_scroll is consulted
last (lnomad/src/tui.rs:1951-1954), so single-letter commands claim
their keys first and j/k/Space reach the scroll table only because
nothing above claims them. In the places panel the order is inverted
(lnomad/src/tui.rs:2145) with a comment explaining why: there Ctrl-d
must be a half-page motion, not the d that closes the panel.
There is one further principle worth carrying over verbatim. Bare r is
deliberately left unbound, because "a mesh reload is expensive and single
letters are reserved for cheap local actions"
(lnomad/src/tui.rs:1930-1933); reload requires R, Ctrl-R or F5.
That is a cost-aware keymap, and a messenger sends over the same radios.
Architecture. Elm-style with an explicit effect list, all in
lnomad/src/tui.rs: Model at :680-838 (#[derive(Clone, Debug, Default)] at :679), AppEvent at :1197-1264, Effect at :592-646,
update(&mut Model, AppEvent) -> Vec<Effect> at :1271-1381,
view(&Model, &ImageStore, &mut Frame) at :3532-3571, and a single
effect interpreter run_effects at :5783-5889. update mutates rather
than returning a new model, and effects are plain data, not closures. That
combination is what makes the 245 in-file unit tests possible: build a
Model, feed a synthetic AppEvent::Key, assert on the model and on the
returned Vec<Effect>, with no IO anywhere (lnomad/src/tui.rs:6228
onwards; helpers at :6238-6251). The view is tested against
ratatui::backend::TestBackend (:6232, :7055), and the --print path
has byte-identical golden files (lnomad/tests/render_golden.rs:18-37).
Unsolicited inbound events already exist. This matters more than
anything else for a messenger, and the answer is yes.
AppEvent::NodeDiscovered (lnomad/src/tui.rs:1264) arrives from
announces with no user action. The chain is: an announce sink installed
before the session is shared so nothing is missed at startup
(set_announce_sink, lnomad/src/fetch.rs:225-227, wiring comment at
lnomad/src/tui.rs:5995-5998), a non-blocking unbounded send on every
recorded announce (note_announce, lnomad/src/fetch.rs:369-389), a
dedicated background task parked on the shared session in 250 ms lock
slices (spawn_discovery, lnomad/src/tui.rs:5721-5762), and a
tokio::select! arm that folds the result into the model
(lnomad/src/tui.rs:6154-6156). The main loop has five arms
(lnomad/src/tui.rs:6058-6161), and the timer arm is conditionally enabled
(, if animate at :6157, driven by needs_tick() at :1018-1021) so an
idle browser does not wake eight times a second.
Persistence. Three small files under
${XDG_CONFIG_HOME:-~/.config}/lnomad/: bookmarks.toml, identify.toml,
and a binary identity. Everything else is RAM. The write path is
fs::write with no atomic rename, no fsync, and errors deliberately
ignored (lnomad/src/bookmarks.rs:124-130, effect handler at
lnomad/src/tui.rs:5879-5885), and load treats corrupt exactly like
missing (lnomad/src/bookmarks.rs:116-121). For bookmarks that is a
defensible trade. For a message store it is data loss.
load_or_create (lnomad/src/identity.rs:39-53) silently mints a fresh
identity when the stored one fails to decode. For a browser, whose identity
is disposable, that is right. For a messenger, whose identity is the
user's address, silently replacing it breaks every contact's address book
with no warning. That default must be inverted.
Two caches worth copying. The page cache stores the parsed document
rather than the laid-out page, because layout depends on width and theme
(lnomad/src/page_cache.rs:10-13). The image cache is bounded by bytes
rather than count, and the reasoning generalises directly to attachments:
"a cache of 'the last fifty pictures' says nothing about how much memory a
browser is holding, and pictures differ in size by three orders of
magnitude" (lnomad/src/image_cache.rs:11-14).
Rendering. One layout core, two sinks. layout_blocks
(lnomad/src/render.rs:192-217) produces Vec<RLine> where RLine is a
vector of StyledChar { ch, st, link, field }
(lnomad/src/render.rs:340-361): already wrapped, aligned and indented,
one RLine per output row, every cell carrying its resolved style and its
owning link index. That IR feeds either to_ratatui_text
(lnomad/src/tui.rs:5143-5162) for the TUI or emit_ansi
(lnomad/src/render.rs:224-231) for --print. Scrolling is a slice, not a
widget scroll, because the page is pre-wrapped
(lnomad/src/tui.rs:3782-3789), and there is one scroll rule shared by
every scrollable window (scrolled, lnomad/src/tui.rs:274-289).
Two rendering caveats. Wrapping compares cur.len() > width, i.e.
character count rather than display width (wrap,
lnomad/src/render.rs:848-875), which will overflow on CJK and emoji. No
test covering that was found, so whether it is a known limitation or an
oversight is unclear. And the whole page is laid out eagerly on every
relayout (lnomad/src/tui.rs:910-934), including on every keystroke in a
form field (:1657-1660).
Scriptability, and the absence of settings. --print fetches, renders
and prints once (print_once, lnomad/src/browser.rs:133-141). Output is
raw ANSI page text and nothing else: no link markers, no legend, and with
--no-color links are indistinguishable from body text
(lnomad/src/render.rs:143-146). There is no JSON output anywhere in the
crate: serde_json is not a dependency. Non-interactive detection is
automatic: interactive = !args.print && stdin().is_terminal() && stdout().is_terminal() (lnomad/src/main.rs:167-168), so piping never
blocks on the UI. Exit codes: 0 success, 1 operational failure, 2 argument
or URL error (lnomad/src/main.rs:174, :188, :222, :238).
lnomad has no settings file at all. --config points at the Reticulum
config directory; the lnomad/ directory holds only data. The theme is
auto-detected via OSC 11 before raw mode is entered
(lnomad/src/tui.rs:5956-5963) and toggled at runtime with t; theme
colours are hard-coded (lnomad/src/theme.rs:112-193).
The handoff that already exists. lnomad recognises lxmf@<hash>
links, and because it has no composer it copies the address to the
clipboard and says so in a toast (follow_link,
lnomad/src/tui.rs:2768-2777). The messenger is the natural target of that
handoff, and wiring the two together is an explicit goal.
2. leviculum-lxmf: what it gives and what it does not
Three layers, all sans-IO: NodeCore (Reticulum transport, owned by the
app), LxmfNode (leviculum-lxmf/src/node.rs:375, the lxmf.delivery
destination adapter), and LxmfRouter
(leviculum-lxmf/src/router.rs:461, the queue, retry scheduler, stamp and
ticket policy, dedup caches and propagation client). The application builds
on LxmfRouter and owns both it and the core; the router never owns the
core, every method takes it as a parameter.
Note that LxmfRouter, RouterEvent, RouterOutput, RouterConfig and
MessageState are not re-exported at the crate root — the crate root
exports only BuiltResource, DeliveryStampRequest, InboundStampRequest,
PendingResourceBuild and PropagationStampRequest from that module
(leviculum-lxmf/src/lib.rs:97-100) — so they are reachable as
leviculum_lxmf::router::* only.
Events are return values, not a channel
#![allow(unused)] fn main() { #[must_use] pub struct RouterOutput { // leviculum-lxmf/src/router.rs:300-303 pub core: TickOutput, pub events: Vec<RouterEvent>, } }
Every router method that can produce work returns this. There is no
callback and no channel. The library never drops an event, but it never
retains one either: if the application drops a RouterOutput, those events
are gone. #[must_use] on both RouterOutput and TickOutput
(leviculum-core/src/transport.rs:534) is the only safety net, and
TickOutput's own doc says dropping it "silently loses outbound packets
and application events" (leviculum-core/src/transport.rs:498-500).
There is a re-entrancy obligation that is easy to miss and fatal to get
wrong: RouterOutput.core.events contains NodeEvents that must be fed
back into router.handle_event(), recursively, until the worklist
drains. This is exactly Codeberg #204's subject. leviculum-lxmf-node
implements it with a bounded worklist and says why the bound is the
consumer's choice (MAX_ABSORB_ROUNDS,
leviculum-lxmf-node/src/processor.rs:82-92, absorb at :392-394). A
client that forgets this will silently never see incoming messages.
RouterEvent
Fifteen variants (leviculum-lxmf/src/router.rs:294-342):
MessageQueued, MessageState { message_id, state }, MessageReceived,
InboundRejected, DirectLinkEstablished, Duplicate,
InvalidSignature, InvalidStamp, ResourceBuildPending, StampPending,
InboundStampPending, PropagationStampPending, PropagationSyncState,
PropagationSyncComplete, PersistenceRequested.
What is missing is as informative as what is there. There is no announce
event: LxmfNodeEvent::PeerAnnounced carries the destination hash only,
with app data discarded (leviculum-lxmf/src/node.rs:133-135,
:746-754), and handle_node_event does not forward it at all — it falls
into _ => {} (leviculum-lxmf/src/router.rs:1369). The router does
decode the delivery announce but keeps only stamp_cost and
compression_supported, discarding the display name
(leviculum-lxmf/src/router.rs:1209-1219). Display-name learning is
entirely the client's job, from raw NodeEvent::AnnounceReceived.
Sending does arrive, on the event every verdict travels on:
RouterEvent::MessageState, from all three sites that enter the state —
the composed send (leviculum-lxmf/src/router.rs:1832-1837), the
built-transfer commit (leviculum-lxmf/src/router.rs:1033-1038) and the
upload the transport reports through UploadSubmitted
(leviculum-lxmf/src/router/propagation_runtime.rs:354-366) — and on the
transition only: a submission onto an entry already in that state reports
nothing. For a direct delivery it is the only thing between being accepted
and being answered. For a propagated one it names the message on a link
PropagationSyncState was already narrating, which narrates the link and
cannot say what is on it. An opportunistic message reports it and then
goes quiet until the verdict: the router moves it on to Sent in the same
tick, and that transition is reported nowhere.
There is still no event for Outbound and none for progress: the router
folds LxmfNodeEvent::Progress into OutboundEntry::progress without
emitting anything (leviculum-lxmf/src/router.rs:1421-1433), so progress
must be polled through outbound()
(leviculum-lxmf/src/router.rs:701).
MessageState and what it honestly means
#![allow(unused)] fn main() { pub enum MessageState { // leviculum-lxmf/src/router.rs:62-71 Generating = 0x00, Outbound = 0x01, Sending = 0x02, Sent = 0x04, Delivered = 0x08, Rejected = 0xfd, Cancelled = 0xfe, Failed = 0xff, } }
Discriminants are the Python LXMessage constants. Four traps:
Generatingis dead. It is only ever produced by snapshot decoding (leviculum-lxmf/src/router.rs:2356); nothing assigns it.Sentmeans two different things and never applies to direct delivery. For opportunistic messages it means the packet was handed to Reticulum unproven, and the message is still queued and still retryable (leviculum-lxmf/src/router.rs:1327-1341). For propagated messages it means the propagation node accepted the upload, and the entry is deleted (leviculum-lxmf/src/router/propagation_runtime.rs:386-393). Direct delivery goesOutbound -> Sending -> Delivered | Rejected | Failedand never passes throughSent, because theSubmittedhandler matches onlyDeliveryMethod::Opportunistic(leviculum-lxmf/src/router.rs:1403-1407).Deliveredis a Reticulum transport proof, not an application receipt. It comes fromPacketDeliveryConfirmed/LinkDeliveryConfirmed(leviculum-lxmf/src/node.rs:1282-1304) or fromResourceCompleted { is_sender: true }(leviculum-lxmf/src/node.rs:1098-1109). It proves the bytes arrived at the destination identity. It does not prove an LXMF client parsed them and it certainly does not prove a human read them. There is no read-receipt field in LXMF at all (leviculum-lxmf/src/constants.rs:52-78).Rejectedis ambiguous. It means either "the receiver cancelled the Resource transfer" (leviculum-lxmf/src/router.rs:1285-1298) or "the propagation node refused the upload for an insufficient stamp" (leviculum-lxmf/src/router/propagation_runtime.rs:418-430), and the event alone cannot distinguish them.
And one omission that shapes the whole UI: there is no "propagated but
not yet collected" state. A propagated message reaches Sent, its queue
entry is removed, and from then on it is indistinguishable from a message
that vanished.
Terminal states remove the entry from the outbound map
(remove_outbound, leviculum-lxmf/src/router.rs:873-876; call sites at
:888, :1334, :1369 and five in the propagation runtime, among them
leviculum-lxmf/src/router/propagation_runtime.rs:895). If the client does
not capture the Message at enqueue time it cannot render its own sent
message afterwards, and it cannot offer a retry button.
MAX_DELIVERY_ATTEMPTS is 5 (leviculum-lxmf/src/router.rs:47).
Propagation: what the router does, and what it refuses to do
Setup requires the client to mint a second lxmf.propagation destination
via PropagationTransport::destination
(leviculum-lxmf/src/propagation_client.rs:282-292),
register it, and hand it to enable_propagation_client
(leviculum-lxmf/src/router.rs:603); the transport identity must equal the
router's or you get RouterError::IdentityMismatch
(leviculum-lxmf/src/router.rs:608-610).
Node discovery is automatic from announces (remember_announce,
leviculum-lxmf/src/propagation_client.rs:384-400, driven from the
announce arm at :733-742), and the decoded announce carries enabled,
transfer_limit_kb, sync_limit_kb, stamp_cost, peering_cost and
metadata (PropagationNodeAnnounce,
leviculum-lxmf/src/propagation.rs:513-525), all of which are directly
displayable. select_outbound_propagation_node with None auto-ranks by
route, hops, peering cost and stamp cost
(leviculum-lxmf/src/router/propagation_runtime.rs:1161-1196).
Once a sync starts, everything is automatic: path request, link, identify,
list request, want/have partitioning, download, acknowledge and purge
(begin_list_request,
leviculum-lxmf/src/router/propagation_runtime.rs:459-551). The observable
state machine is PropagationClientState
(leviculum-lxmf/src/router/propagation_runtime.rs:60-75),
wire-compatible with Python's PR_* constants: Idle, PathRequested,
LinkEstablishing, LinkEstablished, RequestSent, Receiving,
ResponseReceived, Complete, NoPath, LinkFailed, TransferFailed,
NoIdentity, NoAccess, Failed. ResponseReceived is never assigned in
practice. There is automatic failover to another reachable node when the
selected one loses its route
(leviculum-lxmf/src/router/propagation_runtime.rs:825-846).
What the router will not do:
- It never schedules a sync.
request_messages_from_propagation_node(leviculum-lxmf/src/router/propagation_runtime.rs:1365) must be called by the application every time.PropagationClientConfighas three fields and none of them is an interval (leviculum-lxmf/src/router/propagation_runtime.rs:35-45), andnext_deadline()returnsNonein every state exceptPathRequested(leviculum-lxmf/src/router/propagation_runtime.rs:1142-1148). - It does not persist known propagation nodes. They live in an
in-memory map (
known_nodes,leviculum-lxmf/src/propagation_client.rs:267) and are absent from the router snapshot (snapshot,leviculum-lxmf/src/router.rs:2069-2086). The client must persist and replay them viarestore_known_propagation_node(leviculum-lxmf/src/router/propagation_runtime.rs:1317). The selected node is not snapshotted either. - It does not clamp the transfer limit against the node's advertised
one. The download request carries the local
delivery_transfer_limit_kb(default 1000) regardless of what the node announced (leviculum-lxmf/src/router/propagation_runtime.rs:533-538).
Default retain_synced_on_node is false
(leviculum-lxmf/src/router/propagation_runtime.rs:50), meaning the client
tells the node to purge what it has collected. That is a user-visible
policy decision disguised as a config default, and
the mailbox page argues it should be surfaced.
The reference holds messages for MESSAGE_EXPIRY = 30*24*60*60, i.e.
thirty days (reference/LXMF/LXMF/LXMRouter.py:38).
Storage is a bare key/value trait
#![allow(unused)] fn main() { pub trait LxmfStorage { // leviculum-lxmf/src/storage.rs:18-26 fn load(&self, key: &[u8]) -> Result<Option<Vec<u8>>, StorageError>; fn store(&mut self, key: &[u8], value: &[u8]) -> Result<(), StorageError>; fn remove(&mut self, key: &[u8]) -> Result<(), StorageError>; fn keys(&self, prefix: &[u8]) -> Result<Vec<Vec<u8>>, StorageError>; fn flush(&mut self) -> Result<(), StorageError> { Ok(()) } } }
There is no conversation, thread, contact or history concept in it. Two
implementations exist, both in that file: MemoryLxmfStorage
(leviculum-lxmf/src/storage.rs:42) and NoLxmfStorage
(leviculum-lxmf/src/storage.rs:116). The file-backed one is FileLxmfStorage
(leviculum-std/src/file_lxmf_store.rs:27), in the std crate because the LXMF
crate is no_std.
The router writes exactly one key, b"lxmf/router-state"
(ROUTER_STATE_KEY, leviculum-lxmf/src/router.rs:64), holding the
outbound queue, delivered and processed ID windows, stamp costs, tickets
and the ignore set (leviculum-lxmf/src/router.rs:2069-2086). A client
should stay off the lxmf/ prefix and is otherwise free.
Restore resets every queued message to Outbound with
next_attempt_ms = 0 and progress = 0.01
(leviculum-lxmf/src/router.rs:2054-2057), because in-flight correlation
is expressed in a process-local monotonic clock that does not survive a
restart. A UI therefore cannot show a stable "sending" progress across
restarts, and must not pretend to.
Features a UI could surface
- Attachments (
leviculum-lxmf/src/attachments.rs): files, one image, one audio clip, asMessageAttachments::into_fields()(leviculum-lxmf/src/attachments.rs:57) /from_fields()(leviculum-lxmf/src/attachments.rs:86). Attachments are inline bytes in the message, so anything with a real attachment exceeds the packet MDU and forces link or Resource delivery (representation,leviculum-lxmf/src/node.rs:577-605). - Paper messages (
leviculum-lxmf/src/paper.rs): a message encrypted to a destination and rendered as anlxm://base64 URI (to_uri,leviculum-lxmf/src/paper.rs:172), capped atPAPER_MDU = 2210bytes (leviculum-lxmf/src/constants.rs:38). Ingest viarouter.ingest_paper(uri)(ingest_paper,leviculum-lxmf/src/router/paper_runtime.rs:17). No QR generation exists; that is the client's job. - Tickets (
leviculum-lxmf/src/ticket.rs): a 16-byte secret you issue to a contact so their future messages skip proof-of-work. Mostly invisible and automatic: received tickets are remembered from any signature-valid inbound message —remember_verified_ticket(leviculum-lxmf/src/router.rs:1512) — and applied when a message is enqueued (leviculum-lxmf/src/router.rs:820). Expiry 21 days, renew at 14, minimum one day between issuances to the same peer (leviculum-lxmf/src/constants.rs:40-43).issue_ticket_fieldrefuses withRouterError::NoWallClockwhen the node's clock is implausible (leviculum-lxmf/src/router.rs:681-682), and can also legitimately returnOk((None, _))when rate-limited (leviculum-lxmf/src/router.rs:699). A UI has to distinguish "granted", "not yet, try tomorrow" and "cannot, no clock". - Stamps (
leviculum-lxmf/src/stamp.rs): proof-of-work over the message ID, cost being required leading zero bits, so expected work is 2^cost hashes plus a workblock expansion of 3000 rounds (WORKBLOCK_EXPAND_ROUNDS,leviculum-lxmf/src/constants.rs:45). Costs above about 40 bits are described in-tree as "already unreachable in practice" (leviculum-lxmf/src/router.rs:1145-1146). No wall-clock benchmark exists in the crate and none was run for this document, so any UI estimate of mining time must be measured first, not guessed. There is no cancellation and no deadline:generateloops until it succeeds (leviculum-lxmf/src/stamp.rs:356-367), andStampError::Cancelledexists but is never constructed (leviculum-lxmf/src/stamp.rs:25).
Fields with constants but no codec
leviculum-lxmf/src/constants.rs:52-78 declares the full LXMF field set
including FIELD_THREAD (0x08), FIELD_RENDERER (0x0F),
FIELD_REPLY_TO (0x30), FIELD_REPLY_QUOTE (0x31),
FIELD_REACTION (0x40) and FIELD_COMMENT (0x41), but only files, image
and audio have typed codecs. Unknown fields round-trip byte-for-byte
(leviculum-lxmf/src/message.rs:5-8), so nothing is lost, but a client
wanting replies, threads, reactions or renderer-aware display must
hand-roll the msgpack via the exported msgpack module.
RENDERER_MICRON = 0x01 (reference/LXMF/LXMF/LXMF.py:100) is interesting
here: leviculum-micron already parses micron into a document model
(leviculum-micron/src/lib.rs:25-27) and lnomad already renders that
model. A messenger in this workspace can honour FIELD_RENDERER almost for
free, which no other terminal LXMF client does.
3. The driver seam, which nearly forces the architecture
An LXMF client cannot be fed from leviculum-std's public event stream.
The reason is documented at the seam itself: the tap sits on
output.events inside dispatch_output, before the event sink
classifies, and seven of the event types LXMF needs, including
PacketReceived and LinkDataReceived, are EventClass::Data and
therefore droppable under load. A processor fed from take_event_receiver
"would silently lose inbound messages with nothing underneath to retransmit
them" (leviculum-std/src/driver/processor.rs:191-199, "Where the events
come from").
So the messenger must register a CoreProcessor
(leviculum-std/src/driver/processor.rs:274-302) on the builder, and the
LXMF router lives inside the driver's tick, under the core mutex. That
carries hard obligations:
- Both hooks run with a non-reentrant mutex held. The processor may not own
a handle to the node it runs inside; roughly forty synchronous
pub fns onReticulumNodeopen with a lock and one of them in a hook body deadlocks the node in ordinary safe code. - Every side effect must be a non-blocking queue push.
leviculum-lxmf-nodedoes exactly this: stdout lines, stderr lines, proof-of-work jobs and shutdown are all channel sends (leviculum-lxmf-node/src/processor.rs:15-30). PROCESSOR_TICK_BUDGETis 5 ms per hook call (leviculum-std/src/driver/processor.rs:181), reported rather than enforced. Message packing costs about 0.8 ms and unpacking with signature verification about 3.2 ms for 1 MiB, per The core lock budget.NodeCore::send_resource(leviculum-core/src/node/mod.rs:1747) must not be called from a hook: 141 ms under the lock for 1 MiB.- The processor needs its own periodic slot to drain its command queue,
because an event tap can never initiate anything.
leviculum-lxmf-nodeuses 200 ms (POLL_INTERVAL_MS,leviculum-lxmf-node/src/processor.rs:80).
This is a strong constraint and a gift at the same time: it means the "model" that talks to the network is a synchronous state machine with a queue on either side, which is exactly the shape that tests well.
4. Decision: process architecture
Options
A. One process. TUI plus an in-driver CoreProcessor. The binary
builds a ReticulumNode as a shared-instance client with
core_processor(...) installed, exactly as leviculum-lxmf-node does
(leviculum-lxmf-node/src/main.rs:384-390). The processor owns the
LxmfRouter; the TUI owns the model. They talk over two unbounded
channels.
For: one binary, one config, no IPC to design, matches lnomad's
deployment shape. Against: mail is only received while the TUI is
running. Closing the terminal stops collecting.
B. Two processes. A headless daemon plus a thin TUI client. A lnmsgd
holds the router and the store and exposes a local socket; the TUI is a
view onto it. For: mail arrives while the UI is closed, several front
ends can attach, and the store has one writer. Against: an entire IPC
protocol, a second daemon on a system that already runs lnsd, and a
second thing to package and supervise.
C. One binary, two modes. lnmsg with a --daemon flag, and the TUI
attaching to a running daemon if there is one and otherwise running the
router itself. For: option A's simplicity on day one, option B's
availability when the user asks for it. Against: two code paths for every
operation, and the temptation to test only one.
Decision (2026-08-08)
C, built as A first. Start with a single process, but put the router
and the store behind an interface from the beginning so that the daemon
mode is a wiring change rather than a rewrite. Whether the daemon mode is
ever built is decided empirically: if syncing with a propagation node on
start plus every N minutes proves sufficient in the mesh we care about, A
alone stays. The reference retention default of thirty days
(reference/LXMF/LXMF/LXMRouter.py:38) suggests it might. This aligns with
the standing decision that propagation nodes, not client uptime, are the
answer to offline delivery.
The lnsd-resident variant — daemon mode as a CoreProcessor registered
inside lnsd — is rejected, twice over: it would put LXMF knowledge
into the transport daemon, which the Codeberg #196 seam was explicitly
designed to avoid, and it contradicts the standing rule that client
programs do not merge into lnsd (there will be more clients than this
one).
Requirement: the core must not know it has a terminal
Decided 2026-08-08, and binding for lnomad too: the messenger will grow
other frontends on other platforms later — a GUI is expected — and that
must be a frontend swap, not a rework.
Concretely, the crate splits into two layers with a hard boundary:
lnmsg-core(or a module boundary with the same discipline until a crate split is warranted): the model,update, effects, the store, the router glue, sync scheduling, trust, delivery bookkeeping. This layer never imports crossterm, ratatui, or any terminal type. Everything in it is driven byAppEventin andEffectout, and is testable headless.- The TUI frontend: rendering, key mapping, terminal lifecycle. It
translates terminal events into
AppEvents and draws the model. A GUI frontend later is a second translator and a second renderer over the same core — no change to the core's types.
The TEA split below is what makes this cheap: the discipline is not a new
architecture, it is refusing to let the existing one leak. The test for the
boundary is mechanical and should exist from day one: the core compiles
without the TUI dependency tree (feature gate or crate split), and the
headless test suite drives complete user stories through AppEvents alone.
For lnomad the same requirement holds as a future refactor: its TEA split
already keeps the model headless-testable, but model, update and view live
in one 11,873-line file with crossterm types reachable throughout. When
lnomad next gets substantial work, the same core/frontend boundary is
carved there. Tracked as its own issue, not as part of this program.
5. Decision: the event loop and the TEA split
lnomad's split survives contact with a messenger with one change.
The shape that follows from the driver seam is three layers, not two:
crossterm events ──┐
router events ──┼──> AppEvent ──> update(&mut Model) ──> Vec<Effect>
timer ──┘ │
v
run_effects
│
Command queue ────────────┘
│
v
CoreProcessor::on_tick / on_event
(LxmfRouter, under the core lock)
│
RouterEvent queue
│
└──> AppEvent
The processor is not part of the TEA model. It is a second, synchronous
state machine on the far side of two queues, and it is testable on its own
terms without a terminal, exactly as leviculum-lxmf-node is.
Three specific things lnomad does that must change:
- Bottom-anchored scrolling with a pinned flag.
lnomad'sscrollis the index of the top visible line (lnomad/src/tui.rs:274-289). A message list wants a "pinned to bottom" boolean so an inbound message appends without yanking the viewport out from under a user who has scrolled up. NomadNet gets this wrong: it resets to the bottom on every refresh (NomadNet'sConversations.py, line 2287). - Windowed layout.
lnomadre-lays out the whole page on every relayout (lnomad/src/tui.rs:910-934). A ten-thousand-message conversation must not do that, and a compose buffer must not trigger it per keystroke. Lay out the visible window plus a margin, and cache per message keyed by(message_id, width, theme). - The timer must run.
lnomaddisables its tick when idle (lnomad/src/tui.rs:6157). A messenger has relative timestamps, a sync schedule and retry deadlines. A one-second tick when there is anything pending, and a slower one otherwise, driven bynext_deadline()(leviculum-lxmf/src/router.rs:2000).
Things to carry over unchanged: the generation counter for stale-result
rejection (spawn_fetch, lnomad/src/tui.rs:5305-5346), the tick-counted
toast whose expiry is a pure function and therefore unit-testable without
real time passing (Toast, lnomad/src/tui.rs:661-676, test at :7271),
the TerminalGuard RAII plus panic hook that restores the terminal before
the backtrace prints (lnomad/src/tui.rs:5229-5273), and OSC 52 for the
clipboard so copy works over SSH with no X11 dependency (osc52,
lnomad/src/tui.rs:2519-2551).
One thing to fix from day one: lnomad is 11,873 lines in
src/tui.rs. A messenger has strictly more state. Split
model.rs / event.rs / update/ / view/ / shell.rs before the first
thousand lines, not after the tenth.
6. Decision: scriptability
lnomad's --print prints rendered ANSI and nothing machine-readable
(lnomad/src/render.rs:143-146); there is no JSON anywhere in the crate.
For a browser that is defensible. For a messenger it is a missed
opportunity: "send me a message when the backup finishes" is a real use and
needs no UI at all.
Non-interactive subcommands from the start, following lnomad's automatic
non-tty detection (lnomad/src/main.rs:167-168) and its exit-code
convention (0, 1, 2):
lnmsg send <address> [--title T] [--from NAME] [--attach F] [--via direct|propagated] [-]
lnmsg read [--conversation A] [--since T] [--unread] [--json]
lnmsg sync [--json]
lnmsg contacts [--json]
lnmsg paper <address> - # emit an lxm:// URI
lnmsg ingest <lxm://...>
with --json producing one object per line so jq works, and the exit
code distinguishing "sent" from "queued but not confirmed", which a script
genuinely needs to know.
Decision (2026-08-10): send returns immediately with the message ID
on stdout; exit 0 means "queued cleanly" and claims nothing more, so it
never lies. lnmsg status <id> answers at any time (state, ledger,
--json). --wait opts into blocking until the delivery proof, with a
configurable timeout, and its exit codes distinguish delivered /
still-pending-at-timeout / terminally-failed — a timeout is not reported as
a failure, because the message may still arrive. Rationale: the common case
is a script that must not hang, and enqueueing is the only operation whose
success is knowable immediately; everything after it is a history, not a
result.
Amended (2026-08-21): the message ID comes off stdout. A successful
lnmsg send now prints nothing at all and exits 0; errors keep going to
stderr. The ID is not interesting to the person running the command, and
saying nothing on success is the ordinary Unix contract — a cron job that
mails its output should mail nothing when the send worked. The rest of this
decision is untouched: exit 0 still means "queued cleanly" and claims
nothing about delivery, and that is now the entire success signal, which is
why the exit code is what the tests assert. The ID does not become
unobtainable, because lnmsg status <id> needs it: LNMSG_ENQUEUED … id=…
carries it into the structured event log, which LEVICULUM_EVENT_LOG=<path>
turns on and which is written by an unfiltered layer, so the line arrives even
at the warn default (leviculum-std/src/event_log.rs:644-651). No
--print-id flag was added: nothing consumes the ID today, and an option
added against a hypothetical user is an option nobody tests.
Decision (2026-08-21): the sender's name. The delivery announce carried
the literal lnmsg, which names the tool rather than the person, so every
recipient saw the same sender for every operator on every host. The default
is now the account name, resolved in this order: getpwuid(getuid()) first,
then $USER, then $LOGNAME, then lnmsg as a last resort. The password
database comes first deliberately — the first real consumer is a health
monitor started from cron, whose environment has no $USER at all, and a
name that is right interactively and wrong from cron would be discovered
late and by a machine. Resolution never fails: a status line that does not
go out is worse than one from an oddly-named sender.
--from NAME, and LNMSG_DISPLAY_NAME for the cron case, override it,
with the flag winning. They exist because a bare account name is ambiguous
when the same user runs the monitor on several machines — but what goes in
them is the operator's choice, not a policy of ours: no automatic hostname
suffix and no templating. An empty or whitespace-only override is exit 2,
not a silent fall back to the default, since it was set on purpose.
lnmsg/src/display_name.rs holds the order; LNMSG_SENDER from=… source=…
records which step answered, which is what separates a cron run that fell
through to the last resort from an interactive one that read $USER.
7. Structured event log
Structured event logs and the project's
debugging discipline call for EVENT_NAME key=val t=<ms> lines. A
messenger that can be started with a log file, and whose every protocol
transition appears in it, is debuggable in the field in a way that no
terminal-scrollback client is. lnomad has no tracing dependency at all.
This one is cheap and is not treated as speculative.
8. What the library does not expose
Naming these is useful because each is a candidate issue.
Decision (2026-08-10) on sequencing: triage, not either extreme. Gaps
1, 2, 3 and 8 are closed in one library wave before lnmsg starts,
because the decided design cannot be built honestly without them: gap 2
blocks the mailbox glyph and the truthful delivery display outright, gap 1
blocks the naming-based trust model, gap 3 is shared infrastructure every
client rewrites, and gap 8 is small and stops two clients wording the same
errors differently. The remaining eight are filed as issues and met in
build order — the Codeberg #196 precedent (the library's biggest gap was
found by building a real consumer) argues for letting lnmsg discover the
gaps nobody has named yet, but waiting to "discover" a gap that is already
understood is delay, not empiricism. Gaps 10 and 11 are already covered by
the queued #204/#202/#203 batch; gap 7 shares its core-side prerequisite
with the S2 test-infrastructure question from the #212 work.
Closed 2026-08-11. The four are done: RouterEvent::PeerAnnounced (1),
MessageState::AwaitingCollection (2), FileLxmfStorage in leviculum-std
(3), and Display plus core::error::Error on the error types (8). The
entries below are left as written — they are the record of what was missing,
not a list of open work.
- No display name reaches the application.
LxmfNodeEvent::PeerAnnouncedcarries the destination hash only (leviculum-lxmf/src/node.rs:133-135), the router drops the name after reading the stamp cost (leviculum-lxmf/src/router.rs:1209-1219), andRouterEventhas no announce variant. Every client will re-implement announce filtering andDeliveryAnnounce::decode. ARouterEvent::PeerAnnounced { destination, announce }would remove that duplication. - No "propagated, awaiting collection" state. A propagated message
reaches
Sentand its queue entry is deleted (leviculum-lxmf/src/router/propagation_runtime.rs:386-393), so the client cannot distinguish "in a mailbox" from "gone" without keeping its own shadow record. This is the single biggest obstacle to an honest delivery display. - No file-backed
LxmfStorage. Two implementations exist, both in-memory or null (leviculum-lxmf/src/storage.rs:42,leviculum-lxmf/src/storage.rs:116). Every host application writes the same one. - No periodic sync scheduler and no interval config.
PropagationClientConfighas three fields (leviculum-lxmf/src/router/propagation_runtime.rs:35-45). Arguably correct for a sans-IO crate, but it means every client invents its own policy. - Known propagation nodes and the selection are not in the snapshot
(
leviculum-lxmf/src/router.rs:2069-2086), so every client writes its own persistence and replay. - No stamp cancellation or deadline.
generateloops until success (leviculum-lxmf/src/stamp.rs:356-367) andStampError::Cancelledis declared but never constructed (leviculum-lxmf/src/stamp.rs:25). A user who starts a message to a high-cost peer and changes their mind has no way out. - No inbound Resource cancellation, stated as deliberate pending core
support (
leviculum-lxmf/src/node.rs:518-519). A user receiving a large attachment they do not want can only watch. - Most error types are
Debugonly.RouterError(leviculum-lxmf/src/router.rs:359),LxmfNodeError(leviculum-lxmf/src/node.rs:257),PropagationTransportError(leviculum-lxmf/src/propagation_client.rs:144),MessageError(leviculum-lxmf/src/message.rs:40) andStorageErrorhave noDisplay. Every user-facing string is the client's to write, and two clients will word them differently. - No typed codecs for reply, thread, reaction or renderer fields
(
leviculum-lxmf/src/constants.rs:60-78), so each client hand-rolls msgpack for the same wire structures. This is a compatibility risk more than an ergonomics one. - Codeberg #203 (
StampExecutor::generatereturns a!Sendfuture) applies to us as it applied toleviculum-lxmf-node, which worked around it with a dedicated thread running a current-thread runtime (leviculum-lxmf-node/src/main.rs:430-481). We will make the same workaround. - Codeberg #204 (a hook owns the events its own core calls return) is a documentation gap we will hit on day one. The bounded re-feed loop is not optional.
- Codeberg #186 (LXMF caches age on wall-clock time and are wiped by a timebase jump) matters more for a laptop that suspends than for a daemon that runs continuously, and should be checked against the suspend-resume path before it is dismissed.
lnmsg: the user interface
Part of the lnmsg design record. Two decisions live here: how keys are dispatched, and what an honest delivery display looks like. They share one discipline — a single table read by everything that renders from it — and that discipline is the reason both are on one page.
1. Decision: the input model
This is the hard one, and the browser's answer does not transfer. In a browser, plain letters are free because there is nothing to type into most of the time. In a chat client, the single most common action is typing prose.
Options
A. Modal, vi-style. Normal mode for commands, insert mode for
composing, i to enter, Esc to leave.
For: every key stays available for commands; scales to any number of
bindings; vi users are instantly at home; it is the only option where
j/k mean what they mean in lnomad.
Against: it is the single biggest complaint non-vi users have about
terminal software. A user who types a message, presses Esc out of habit,
and then types "quit" has issued four commands. Mode errors in a messenger
are worse than in an editor because the consequence can be sending
something.
B. Always-insert, everything Ctrl-modified. NomadNet's answer
(NomadNet's Conversations.py, lines 68-80). Focus starts in the compose
box, plain typing always composes, every command is Ctrl-something.
For: zero mode errors; a user who has never read the manual can still
type and send. Against: the Ctrl namespace is about twenty-six slots wide
and readline already claims a dozen of them. NomadNet ran out and started
overloading by focus, so Ctrl-X means two different things depending on
invisible state.
C. Focus-follows-pane. The conversation list and the message list are
command panes with lnomad-style bindings; the compose box is a text pane
where plain keys type. Tab moves focus.
For: no modes to learn, because the mode is visible as which pane has the
cursor; single-letter commands survive in the panes where they make sense;
it matches how tmux, mutt and every mail client behave.
Against: the same key does different things in different panes, which is
a mode by another name, just one you can see. And Tab becomes precious.
D. Prefix key. Everything types; a prefix (Ctrl-Space, Ctrl-A, or
,) introduces a command.
For: unlimited namespace, no modes, tmux and screen users know it.
Against: two keystrokes for everything, including scrolling, which is the
operation you do most.
Decision (2026-08-08)
C, with a command palette (D) layered on it, and no modes.
Concretely:
- Three panes: conversations (left), messages (right, upper), compose
(right, lower).
TabandShift-Tabcycle; a click focuses. - In the conversation and message panes,
lnomad's keymap applies unchanged: theSCROLL_KEYStable verbatim (lnomad/src/tui.rs:3159-3291), soj/k,Ctrl-n/Ctrl-p, arrows,Ctrl-f/Ctrl-b,Ctrl-v/Alt-v,Ctrl-d/Ctrl-u,g/G,Home/Endand the wheel all work, in all four idioms, for free. - In the compose pane, plain keys type. Only
Ctrl-andAlt-chords are commands, plus the scroll table'sCtrlandAltrows, which do not collide with readline because they are page motions and readline's are line motions.Entersends: LXMF is email on the wire but chat in every deployed client — Sideband, NomadNet, MeshChat and columba all render conversations, nobody uses the title field — so users arrive with chat expectations, and Enter-to-send is the universal chat convention. Newline isAlt-Entereverywhere and additionallyShift-Enterwhere the terminal speaks the keyboard enhancement protocol (legacy terminals cannot distinguishShift-EnterorCtrl-EnterfromEnter— same byte — which is also whyCtrl-Enter-to-send was dropped; crossterm'sPushKeyboardEnhancementFlagsenables the modern protocol where available).Enteron an empty buffer does nothing. The keymap table carries asend_on_enterswitch; off flips to email style (Enternewline,Ctrl-Dsend) for long-form writers.Ctrl-Don an empty buffer must not send and must not quit. - A command palette on
:in a command pane andCtrl-Pin the compose pane, with fuzzy matching over named commands. This is what makes the scheme learnable: every command has a name, the palette lists them all, and each entry shows its key if it has one. NomadNet's static shortcut bar is the failure mode to avoid. fhint mode fromlnomad(hints,lnomad/src/tui.rs:1155-1195), extended to conversations. Its best property is that it matches either the hint label or a substring of the target's text (hint_matches,lnomad/src/tui.rs:1182-1195), softhen typing part of a contact's name jumps to that conversation. That is a better contact switcher than anything in either prior-art program.
What must not be copied: lnomad quits on Ctrl-C from any mode,
before the mode dispatch (lnomad/src/tui.rs:1575-1578). In a messenger
that is a half-written message thrown away by a reflex. Ctrl-C in the
compose pane clears the buffer to a recoverable draft; quitting needs
Ctrl-Q or the palette.
The keymap must be one table. lnomad applied single-source-of-truth
discipline to scrolling and nowhere else, and its help overlay's other
groups are hand-typed strings that can drift
(lnomad/src/tui.rs:4687-4788). Every binding here is one table read by
the key handler, the help overlay, the palette and the footer hints. That
table is also the natural place to hang user configuration later, which
lnomad has none of.
What would change this: if the compose box turns out to be where users spend nearly all their time, B (always-insert, Ctrl for everything) is simpler and has no invisible state at all, at the cost of losing single-letter commands entirely. The way to find out is to build C and count how often focus is in the compose pane.
The quiet keyboard, with its one exception (decided 2026-08-08). The
lnomad rule that single letters are reserved for cheap local actions
(lnomad/src/tui.rs:1930-1933) generalises here to: no single keystroke
may put bytes on the air — except Enter in the compose pane. The
exception is principled, not a leak: the deliberate act was typing the
message into a deliberately focused pane, and Enter completes that act.
Everywhere else the rule is absolute: no key in a command pane transmits,
syncing needs a chord or the palette, announcing needs a chord or the
palette. Everything free is free.
2. Decision: what honesty about delivery looks like
Most chat UIs lie by simplification: one check for sent, two for delivered, and everything ambiguous rounded to the friendlier reading. LXMF has real states and the library reports them, so there is no excuse.
The states as they actually are, from the library survey:
| Situation | Library state | What is actually true |
|---|---|---|
| queued, waiting for a route or a retry | Outbound | nothing has been transmitted |
| computing proof-of-work | Outbound + StampPending | nothing has been transmitted, and it may take a while |
| link being built, or bytes going out | Sending | in flight, no confirmation |
| opportunistic packet handed to Reticulum | Sent | transmitted once, unproven, still retryable |
| propagation node accepted the upload | Sent, entry deleted | it is in a mailbox; the recipient may never collect it |
| transport proof received | Delivered | the bytes reached the destination identity |
| receiver cancelled the resource | Rejected | they refused it, or their client did |
| node refused the upload (bad stamp) | Rejected | the mailbox refused it, not the recipient |
| five attempts exhausted | Failed | give up, keep the text |
What a truthful UI shows
Different marks for different truths, and words in the detail view. The three that must never be collapsed:
- handed to the network, unproven (opportunistic
Sent), - left in a mailbox for later (propagated
Sent), - arrived at the destination (
Delivered).
columba collapses the first two into one check mark; NomadNet gives them different colours but never words. Three visually distinct marks, and, crucially, a per-message detail view (columba's best idea) showing delivery method with a one-sentence explanation, attempt count, hop count, and the error string when there is one. In a terminal that is a key press on a focused message.
Never claim a read receipt. LXMF has no such field
(leviculum-lxmf/src/constants.rs:52-78). Any UI element that suggests one
is a lie in the protocol's own terms.
Say what Delivered means, once. It means the bytes reached the
destination identity, not that a human saw them. The detail view is where
that sentence lives.
Distinguish the two Rejecteds by correlating with the message's
method before the entry disappears. "Your mailbox refused this message" and
"the recipient's client refused this message" are different problems with
different fixes.
Handle the restart discontinuity honestly. restore resets every
queued message to Outbound (leviculum-lxmf/src/router.rs:2054-2057). So
after a restart, a message that was "sending" is "queued" again, and the UI
must show that rather than a frozen progress bar. NomadNet's equivalent,
forcing a stale mid-flight message to FAILED on load (NomadNet's
Conversation.py, lines 455-467), is at least honest, though our library
gives us the better option of an honest retry.
Never let a late failure demote a success. columba added exactly this
guard after a real bug. Delivered must be terminal in the UI even if
something arrives afterwards claiming otherwise.
Decision (2026-08-10): the ledger and the glyph
The ledger is one keypress away, inside the per-message detail view — part of the normal program, not a debug flag, but costing no space in the conversation. It is also where the airtime number lives. The ledger rows are stored beside the message in SQLite (a handful of roughly 50-byte rows per message) and share the message's retention (storage).
Decided with it: one coloured glyph per message in the conversation view, occupying a single cell, carrying the delivery state at a glance. The glyph is the compact face of the same state machine the detail view explains; both render from one shared state-mapping table (the keymap discipline above, applied again — the glyph, the detail view and the ledger can never disagree).
The glyph language, chosen so the colours carry the meaning even before the shapes are learned:
| state | glyph | colour | motion |
|---|---|---|---|
| queued, nothing transmitted | ○ | dim gray | static |
| computing proof-of-work | braille spinner ⠋⠙⠸⠴⠦⠇ | dim yellow | animated on the tick |
| in flight (link building, bytes out) | braille spinner | yellow | animated on the tick |
handed to the network, unproven (opportunistic Sent) | ◇ | amber | static |
left in a mailbox (propagated Sent) | ⌂ | blue | static |
arrived (Delivered) | ✓ | green | static, terminal |
refused (Rejected, either kind) | ✗ | red | static |
given up (Failed) | ✗ | dim red | static |
The rules the table encodes: green and a check mark appear only on
proof — the mailbox state is a blue house, deliberately not a second
check, because "in a mailbox the recipient may never open" must not read as
progress toward delivered; amber ◇ (hollow) against green ✓ (solid)
mirrors unproven-versus-proven; animation means "the machine is working
right now" and nothing else, driven by the
event-loop tick so it freezes honestly if the
program hangs. An ASCII fallback set (. * o ^ v x) ships behind the same
table for terminals that mangle the glyphs.
The timeline itself:
14:02:11 queued
14:02:11 no route known, asked the network
14:02:19 route found, 3 hops
14:02:20 link established
14:02:21 sent, 412 bytes
14:02:26 delivery proof received
This is close to free, since the library already emits every one of those transitions, and it turns "why is this taking so long" from a support question into something the user can read.
3. Proposals, not requirements
Everything in this section is a proposal rather than a decision, and several of these will not survive contact with a real user.
Cost before commitment
Show what a message will cost before it is sent, next to the send action:
412 bytes ~7 s airtime at the slowest hop no stamp required
The pieces exist. leviculum_core::rnode::airtime_ms
(leviculum-core/src/rnode.rs:1061) and packet_airtime_ms
(leviculum-core/src/rnode.rs:1655) are public, interfaces report a
bitrate (leviculum-std/src/interfaces/mod.rs:598-600) computed from
spreading factor, coding rate and bandwidth
(compute_bitrate, leviculum-std/src/interfaces/rnode.rs:3185), and
fetch_remote_status (leviculum-std/src/remote_status.rs:192) retrieves
the interface list from the daemon, which is how lnstatus works. Note two
honesty constraints: fetch_remote_status needs the management authkey,
and the status surface reports bitrate but not the raw radio parameters,
so an estimate from bytes * 8 / bitrate is the best available and must be
labelled as an estimate.
When the peer advertises a stamp cost, the estimate must include the
proof-of-work, and that number has to be measured first: the crate
contains no benchmarks, and the cost model (2^cost hashes plus a 3000-round
workblock, leviculum-lxmf/src/constants.rs:45,
leviculum-lxmf/src/stamp.rs:458-466) predicts scaling but not
milliseconds on a Pi.
Offline as a state, not a failure
An offline-first client should look deliberate rather than broken. Two concrete moves:
- A single posture line that says what the program can do right now: "3 peers reachable directly, mailbox 4m ago, 2 messages waiting to send". Not a red error banner; a statement of fact.
- Queued messages shown in the conversation, in place, greyed, with
their reason ("waiting for a route", "computing proof-of-work, about a
minute"). A message the user wrote should never vanish into a queue they
cannot see. The library gives progress and next-attempt time per entry
(
OutboundEntry,leviculum-lxmf/src/router/outbound.rs:58-72).
The mouse as a first-class citizen
lnomad enables mouse capture unconditionally with no toggle
(lnomad/src/tui.rs:5215-5218) and handles neither drag nor selection
(lnomad/src/tui.rs:1374). In a browser that costs you the terminal's own
copy-paste; in a messenger, where copying message text is a constant, it is
worse.
Proposal: handle drag selection ourselves over the message IR, so selecting
text across wrapped lines and across message boundaries works and copies
via OSC 52 (lnomad/src/tui.rs:2519-2551, which works over SSH). Plus a
--no-mouse flag and a runtime toggle for people who want the terminal's
own selection back. The StyledChar IR already carries per-cell ownership
(lnomad/src/render.rs:340-354), so the hit-testing is a small extension
of visible_links rather than new machinery.
Terminal QR for paper messages and for your own address
PaperMessage::to_uri() produces an lxm:// URI
(leviculum-lxmf/src/paper.rs:172) and the crate stops there. A QR code
rendered in Unicode half blocks is a well-trodden trick, and lnomad
already has the half-block ladder for images. That gives an air-gapped send
path: compose, render, photograph, and the recipient scans it. Also useful
for showing your own address to someone sitting next to you, which NomadNet
does (Ctrl-P in the conversation list).
Constraint: PAPER_MDU is 2210 bytes
(leviculum-lxmf/src/constants.rs:38), which is near the practical limit of
what a QR code can hold and certainly beyond what a phone camera reads off
a terminal at normal font sizes. The UI must say when a message is too big
to be a QR and offer the URI as text instead.
A "what changed while I was away" view
On start, after the first sync, one screen summarising what arrived, grouped by conversation, with the option to mark all read or step through them. Every mail client has this; no Reticulum client does. It fits the usage pattern exactly, because the whole point of the propagation node is that the user was away.
Micron in messages
If FIELD_RENDERER says micron (reference/LXMF/LXMF/LXMF.py:100), render
it with leviculum-micron (leviculum-micron/src/lib.rs:25-27) and
lnomad's renderer. If it says markdown, we already depend on
pulldown-cmark at workspace level (Cargo.toml:82). This is a capability
the workspace has and nobody has spent, and it costs a match statement.
Sending micron is the more interesting half: a compose box with a preview toggle, in a client whose sister program is a micron browser.
lnmsg: conversation storage
Part of the lnmsg design record.
The library stores one key (ROUTER_STATE_KEY, b"lxmf/router-state",
leviculum-lxmf/src/router.rs:64) and hands each received message to the
application exactly once. All history is the client's problem.
Options
A. NomadNet's shape: a directory per conversation, a file per message.
For: trivially crash-safe per message if each write is
temp-plus-rename; human-inspectable; deleting a conversation is rm -rf;
no dependency. Against: one inode per message forever; listing a
conversation is O(n) listdir; no search without reading everything;
NomadNet needed a .index sidecar that duplicates every message body on
disk (NomadNet's Conversation.py, lines 944-960) and still starves file
descriptors under announce storms (NomadNet commit 7bc6911). On a
Raspberry Pi with an SD card this is the worst option.
B. An append-only log per conversation plus a separate index. Messages appended as length-prefixed msgpack; an index mapping message ID to offset; compaction on deletion. For: appends are one write and one fsync; sequential reads are fast; the format is simple enough to recover by hand. Against: you are writing a small database, including the index, the compaction and the crash-consistency argument between them.
C. SQLite. One file, one messages table with the columns columba
already proved out. For: paging, search (FTS5), indices, transactions and
crash safety all solved by someone else; a 2 GB history is unremarkable;
sqlite3 on the command line is the debugging tool. Against: rusqlite
with bundled compiles SQLite's C into the binary. lnomad treats a C
library in the path of the musl-static .deb as disqualifying
(lnomad/Cargo.toml, the ratatui-image comment), though that was about a
pkg-config probe for a shared library rather than a vendored static
one. Checked empirically 2026-08-08: rusqlite with bundled compiles
clean against x86_64-unknown-linux-musl on the project toolchain —
statically linked binary, SQLite 3.46 embedded, FTS5 verified working by
query, 2.6 MB total, 25 s build. The lnomad disqualifier was a
pkg-config probe for a shared library and does not apply to the vendored
static build.
D. A pure-Rust embedded store (redb, sled, fjall). For: no C
toolchain, ACID, ordered keys, so range scans give paging for free.
Against: no query language and no full-text search, so search is
hand-rolled; another dependency to trust with the user's mail.
Decision (2026-08-10)
C — SQLite, via rusqlite with the bundled feature. The packaging
question was answered by the build probe above, so the fallback to D is
retired. The schema below, with its three commitments (identity-scoped
rows, two timestamps sorted on COALESCE(received_at, timestamp), raw
msgpack fields), is the starting point; attachments live out of line as
content-addressed files. Retention is settled further down.
The schema to start from, taking columba's two good decisions:
messages(
id BLOB, identity BLOB, conversation BLOB,
direction INT, state INT, method INT, verification INT,
timestamp REAL, -- the sender's clock, from the wire
received_at REAL, -- our clock, when we saw it
title BLOB, content BLOB, fields BLOB, -- fields as raw msgpack
error TEXT,
PRIMARY KEY (id, identity))
with an index on (conversation, identity, COALESCE(received_at, timestamp))
and one on (conversation, identity, direction, read).
Three reasons for the shape:
- Every row scoped by local identity. A user may hold several addresses; columba's composite keys make that free, and retrofitting it later means a migration.
- Two timestamps, sort on
COALESCE(received_at, timestamp). The wire timestamp is the sender's clock, and nothing makes another node's clock trustworthy. Sorting on it puts a peer with a wrong clock at the top or bottom of your history forever. NomadNet exposes this as a user-facing sort toggle, which is not a fix. (Sub-second precision is not the problem it once was on our side: the emission timestamp carries it, as the reference'stime.time()does, because at whole-second granularity two identical messages created inside one second collapse to one ID —leviculum-lxmf/src/router.rs:531-535, Codeberg #217. Precision and skew are different failures, and only the second one is a sorting question.) fieldsstored as raw msgpack, not exploded into columns. Unknown fields round-trip byte-for-byte in the library (leviculum-lxmf/src/message.rs:5-8) and must round-trip here too, or a reply to a message from a newer client loses information.
Attachments out of line, as NomadNet does (its Conversation.py, lines
752-812): content-addressed files under an attachments/ directory, with
the row carrying names and hashes. Blobs in the database make the database
the size of the blobs, and a 2 MB voice message has no business in a row
you page through.
Retention, decided 2026-08-10
Neither prior-art program prunes, and both will therefore eventually fail on small hardware. The decision:
- Message text is kept forever. A million messages are a few hundred megabytes, SQLite territory, and the searchable archive is precisely the value the FTS decision bought.
- Attachments live under a total-bytes budget (default about 500 MB,
configurable, numbers stated in the manual), evicting oldest first — the
same pattern as
lnomad's byte-budgeted image cache (lnomad/src/image_cache.rs:11-14). - Evicting an attachment never touches the message row. The message
keeps the attachment's name, hash and size and renders "attachment
(2.1 MB), evicted under the storage budget on
". The history does not lie, it just gets lighter. - Ledger rows (the delivery ledger) share the message's retention, so for text they live forever; no extra rule.
- A per-conversation opt-in age cap remains open as a possible later addition and was deliberately not built now.
On the 2 GB Raspberry Pi question specifically
With option C, a 2 GB history is roughly ten million short messages, and the operations that matter are "open the last screenful of a conversation" and "search". Both are index lookups and neither touches the bulk. With option A, opening a conversation reads every file in it. That asymmetry is the whole argument.
Durability rule, from lnomad's counter-example
Nothing in this program may use fs::write on a file it cannot afford to
lose, and nothing may treat a corrupt load as an empty load
(lnomad/src/bookmarks.rs:116-130). Bookmarks can be silently forgotten.
Mail cannot.
lnmsg: the mailbox, and who you talk to
Part of the lnmsg design record. Two decisions live here: when and how the client talks to a propagation node, and how it names and trusts the people on the other end. They belong together because the trust model is what gates the mailbox choice.
1. Decision: propagation node interaction
This is where a naive design produces a client that silently loses mail,
and the library has arranged things so that the naive design is the
default: nothing syncs unless the application asks
(request_messages_from_propagation_node,
leviculum-lxmf/src/router/propagation_runtime.rs:1365, and
next_deadline() returns None outside PathRequested,
leviculum-lxmf/src/router/propagation_runtime.rs:1142-1148).
When to sync
Options: on demand only; on a timer; on start plus timer; on announce of the selected node; adaptive.
NomadNet uses a six-hour timer with a limit of eight messages and no sync
on start (NomadNet's NomadNetworkApp.py, lines 148-150 and 456-471). Six
hours is a very long time for something calling itself a messenger. columba
makes the interval configurable.
Decision (2026-08-10): sync on start, on resume from suspend, on a manual key, opportunistically when the selected node announces (free evidence that it is reachable right now), and on a timer with asymmetric adaptivity:
- The configured interval — default fifteen minutes — is a hard upper bound that is never exceeded. That is the promise the user can rely on: mail is at worst one interval old, always.
- After activity (a sync that returned something, or an outbound send), the client syncs more often for a while — on the order of every two minutes — decaying back to the bound. This is the chat-feel half of adaptivity (the input-model decision), applied to the mailbox.
- The dangerous half — silently lengthening the interval when syncs come back empty — does not exist. Full adaptivity was rejected because an interval that quietly stretches is exactly the "why didn't I get that message for two hours" machine, and unpredictability in a messenger is a breach of trust.
Never sync while a direct link to the peer is up and working, because that spends airtime to learn nothing.
The bound is configurable, its cost is stated in the manual in airtime rather than in seconds (on LoRa a fifteen-minute poll is not cheap), and the status line always shows when the next sync is due — the current cadence is visible, never inferred.
What the user sees
The library hands over a fourteen-state machine
(PropagationClientState,
leviculum-lxmf/src/router/propagation_runtime.rs:60-75), progress as an
f32, transfer size, and a result of { received, duplicates }
(PropagationSyncResult,
leviculum-lxmf/src/router/propagation_runtime.rs:78-83). That is more
than enough to be honest.
A permanent one-line status, taking NomadNet's best idea (its
Conversations.py, lines 517-548) and refusing its modal dialog. Something
like:
mailbox a1b2c3d4 Node-Name 3 hops last sync 4m ago, 2 new next in 11m
and during a sync the same line becomes the progress display, naming the state in words: "asking the network where the node is", "connecting", "asking what it has", "downloading 4 of 7". columba's plain-English state descriptions are better than NomadNet's terse ones and both are better than a bare progress bar.
When there is no reachable node, the line must say which of the several different failures happened, because they need different fixes:
| Library state | What the user must be told |
|---|---|
| no node selected | "no mailbox chosen"; offer the picker |
NoPath (leviculum-lxmf/src/router/propagation_runtime.rs:69) | "cannot find a route to the mailbox"; it may come back |
LinkFailed (leviculum-lxmf/src/router/propagation_runtime.rs:70) | "the mailbox did not answer" |
NoAccess (leviculum-lxmf/src/router/propagation_runtime.rs:73) | "the mailbox refused you"; this one will not fix itself |
NoIdentity (leviculum-lxmf/src/router/propagation_runtime.rs:72) | "the mailbox does not know your key" |
TransferFailed (leviculum-lxmf/src/router/propagation_runtime.rs:71) | "the transfer broke"; will retry |
A single "sync failed" for all six is the lie this section exists to prevent.
The non-interactive slice's selection order (decided 2026-09-12)
The shipped lnmsg send/lnmsg fetch have no picker to offer, so the
CLI resolves the node in a fixed order: the --pn flag, then the
propagation_node key in ${LNMSG_HOME}/config (the persisted
default), then the most recently announced node heard while attached.
Recency rather than the library's route/cost ranking, because for a
short-lived CLI the node that just announced is the one whose
reachability is evidence rather than cache; the reference
LXMRouter ships no autoselection at all — its clients pick, each
with their own rule (NomadNet by hops among trusted, columba by hops).
Nothing is auto-adopted silently in the TUI sense: a cron job's
operator wrote --pn or the config key, and the announced fallback is
for the interactive shell where the operator reads the LNMSG_PN
line. What is persisted where: the identity at ${LNMSG_HOME}/identity,
the node default in ${LNMSG_HOME}/config, the cross-run seen-message
ids in ${LNMSG_HOME}/seen; the selection itself is per-run and never
written back.
Which node, and the trust question
NomadNet auto-selects the fewest-hops node whose trust level is
TRUSTED (its NomadNetworkApp.py, lines 607-631). columba auto-selects
the fewest-hops node, full stop. The library's own auto-selection ranks by
route, hops, peering cost and stamp cost
(select_outbound_propagation_node,
leviculum-lxmf/src/router/propagation_runtime.rs:1161-1196) with no trust
input at all, because it has no notion of trust.
Your mailbox sees the envelope of every message sent to you: who sent it and when, even though it cannot read the content. Handing that to whoever happens to be nearest is a real privacy decision, and neither prior-art program presents it as one.
Never auto-adopt silently. On first run, and whenever the selected node
becomes unreachable, present a picker with the candidates, their hop
counts, their advertised limits and costs (PropagationNodeAnnounce,
leviculum-lxmf/src/propagation.rs:513-525), and require one keystroke to
accept. Automatic failover between nodes the user has already approved is
fine and the library already does it
(leviculum-lxmf/src/router/propagation_runtime.rs:825-846); automatic
adoption of a stranger is not.
Note that NomadNet's trust propagation makes this worse: trusting a person
auto-trusts their node (its Directory.py, lines 198-202), which makes it
eligible as your mailbox. Do not inherit that.
The purge default
retain_synced_on_node defaults to false
(leviculum-lxmf/src/router/propagation_runtime.rs:50), so by default the
client tells the node to delete what it has collected. That is the right
default for privacy and for the node operator's disk, and it is the wrong
default for a user who runs two clients on the same identity, because the
first one to sync takes the mail. This must be a visible setting with the
consequence spelled out, not a config-file default nobody reads.
What the client must implement itself
- The sync schedule (there is none in the library).
- Persistence of known propagation nodes and of the selection, since
neither is in the router snapshot (
snapshot,leviculum-lxmf/src/router.rs:2069-2086); replay viarestore_known_propagation_node(leviculum-lxmf/src/router/propagation_runtime.rs:1317). - Re-selection after restart.
- Proof-of-work for
PropagationStampPending, off the core lock. - Calling
persist()onPersistenceRequested.
The mailbox as a visible relationship (proposal)
The propagation node is currently magic in every client: something chosen
for you, syncing on a schedule you did not set, holding mail you cannot
see. Make it a first-class object in the UI, with its own screen: who it
is, how many hops away, what it advertises, when you last spoke to it, what
it is holding for you if it will say, and whether it is purging what you
collect. columba's trick of modelling the relay as a contact with an
isMyRelay flag is a cheap way to get there.
The honest version of this includes telling the user what the mailbox learns about them: the envelope of every message they receive. No prior program says this out loud.
2. Decision: identity, contacts, and names
The address is the identity. A 16-byte destination hash, rendered as 32 hex characters, is the only thing that is true about a peer.
The delivery announce carries a display name as arbitrary bytes
(DeliveryAnnounce, leviculum-lxmf/src/announce.rs:39-44), and
display_name() strips NUL and trims but does nothing else
(leviculum-lxmf/src/announce.rs:164-168). The announce is signed: the
Reticulum announce signature covers the app data
(leviculum-core/src/announce.rs:104, :213: "the signature covers
destination_hash + public_key + name_hash + random_hash + [ratchet] + app_data"). So a verified announce proves that the holder of that key
chose that name. It proves nothing about uniqueness, and there is no naming
authority in Reticulum. Two identities can both announce "Lew", and one of
them can be doing it on purpose.
Two further facts a UI must not paper over. Message::verification can be
Unverified when the source identity has never been announced to us
(leviculum-lxmf/src/message.rs:229-231), and such messages are
delivered to the application anyway
(leviculum-lxmf/src/router.rs:1474-1477). And the router discards the
display name from announces entirely, so the client must maintain its own
hash-to-name map from raw NodeEvent::AnnounceReceived.
Options for naming
A. Announced name only. What most chat UIs do. Simple, and impersonation is trivial.
B. Local petname only. Nothing is displayed until the user names the contact; strangers show as a hash prefix. Maximally safe, maximally tedious, and hostile to the case where someone new writes to you.
C. Petname wins, announced name shown as provenance. columba's precedence chain (nickname, then announced, then stored), with the announced name still visible somewhere.
Decision
C, plus NomadNet's collision check, plus a hash that never fully disappears.
- Display: petname if set, otherwise the announced name, and in both
cases a short hash suffix. NomadNet suppresses the hash for trusted
peers (its
Directory.py, lines 277-297); we shorten it rather than suppress it, because a four-character hash costs nothing and makes "wait, that is not the Lew I know" possible at a glance. - Collision warning, the one genuinely novel thing in the prior art:
if an announced name is already claimed by a different address, mark it
(NomadNet's
Directory.py, lines 306-320). Extend it to a name-change warning: if an address you have talked to announces a different name than last time, say so once in the conversation. That is cheap, and it is the actual impersonation vector. - Sanitise names. They are arbitrary bytes off the wire. Strip control
characters, normalise, cap the display width, and refuse to let a name
contain something that renders as a checkmark or as another contact's
name. NomadNet does this because micron markup in a name would otherwise
render (its
util.py,strip_modifiers). - Unverified messages must look different. NomadNet's "Unknown Origin"
/ "Invalid Signature" plain-English rendering (its
Conversation.py, lines 620-630) is right; a coloured glyph alone is not. - Refuse to compose to an address whose key is unknown, with the
reason and a "ask the network" action, as NomadNet does (its
Conversations.py, lines 2186-2204). The library will tell you: our own helper checkscore.storage().get_identity(&peer)before composing and reports which call was skipped rather than timing out later (leviculum-lxmf-node/src/processor.rs:1160-1169). - Contact status as persisted state, columba's
ACTIVE/PENDING_IDENTITY/UNRESOLVED, rather than a live lookup, so the list can be rendered without touching the network.
Trust levels: how many?
NomadNet has four — WARNING, UNTRUSTED, UNKNOWN, TRUSTED (its
Directory.py, lines 410-413); columba has none. Four levels with a radio
group is more ceremony than most users will perform.
Decision (2026-08-10): exactly three states that a user reaches by accident and understands: known (I named this contact), unknown (I have not), and blocked. Trust is not a thing to configure but a thing that follows from having named someone, which is an action people take anyway.
A derived fourth state ("suspicious", computed from name collisions and name changes) was considered and rejected: on a mesh, two identities sharing a display name and a contact renaming themselves are normal occurrences, not indicators of attack, and a state that cries wolf on normal behaviour trains the user to ignore it. The signals themselves are not discarded — they are a rendering matter, not a trust matter: when two contacts share a display name the list must disambiguate them (short hash suffix), and identity, not name, is always what messages are keyed by. The columba defect this section opened with was the missing trust anchor for relay auto-selection, and three states cover that: only known contacts qualify.
The identity file itself must not behave like lnomad's
(load_or_create, lnomad/src/identity.rs:39-53, silently regenerating on
a decode failure). A corrupt identity is a refusal to start with a clear
message, because minting a new one silently changes the user's address.
Public channels over LXMF
Status: design record. Not yet implemented.
Meshtastic and MeshCore both have public channels, and they are the reason a lot of people pick those stacks. You flash a board, type a channel name, and you are talking to whoever is in range. LXMF has nothing equivalent. It has mailboxes, and a mailbox is a conversation with one person.
This document describes how to add channels without changing anything in the existing Reticulum or LXMF infrastructure, why that is possible at all, and where the sharp edges are.
Two crates are involved, neither of which exists yet: a propagation node
server, and channel support in lnmsg. The lnmsg design record currently
states that hosting a propagation node is out of scope. That changes here:
the channel feature only works if we host nodes, because the retrieval
side is ours.
1. Why the obvious approach does not work
Reticulum has a broadcast packet form and a symmetric group destination type, and neither carries a channel.
Destination.announce() refuses anything that is not a SINGLE destination
(Destination.py:251-252), so a GROUP or PLAIN destination can never be
announced, and without an announce there is no path table entry.
Transport.outbound skips the path lookup for both types outright
(Transport.py:1121), and any receiving node drops such a packet once
hops > 1 (Transport.py:1354-1373, mirrored in
leviculum-core/src/transport.rs:3233-3265). The reach of a Reticulum
broadcast is exactly one hop plus locally attached clients.
For PLAIN the manual gives the reason: "To be transportable over multiple
hops in Reticulum, information must be encrypted, since Reticulum uses
the per-packet encryption to verify routing paths and keep them alive"
(docs/source/understanding.rst:112-114). For GROUP the same manual says
only that packets are "not currently" carried over multiple hops,
"although a planned upgrade to Reticulum will allow globally reachable
group destinations" (understanding.rst:118-120). That sentence has
stood unchanged since 2022-04-28.
The other flooding primitive, the announce, is capped at two percent of
interface bandwidth by ANNOUNCE_CAP (Reticulum.py:114), with rate
penalties on top. It is not a carrier for chat.
2. The carrier that does exist
Propagation nodes flood among themselves over the full multi-hop transport, because peer-to-peer sync runs over ordinary links between SINGLE destinations. An earlier draft of this document called that flooding "blind", and overstated it: a node never looks inside what it stores, but it does not accept unconditionally. Three properties make the carrier usable, and two toll gates stand in front of it.
A node does not inspect what it stores. lxmf_propagation
(LXMRouter.py:2487-2518) requires only that the data is at least
LXMF_OVERHEAD bytes long (LXMessage.py:63) and that its transient ID,
the SHA-256 of the bytes, is new. It then takes the first 16 bytes as the
destination hash, writes the message to disk and queues it for
distribution. There is no signature check, no decryption, and no check
that the destination has ever been announced or exists.
Distribution is unfiltered. flush_peer_distribution_queue
(LXMRouter.py:2472-2485) offers every new message to every peer except
the one it came from, with no criterion whatsoever.
Peering is automatic — within three limits. A node peers with any
node whose announce it hears, as long as autopeer is set and the hop
distance is within autopeer_maxdepth (Handlers.py:81-83, and again on
inbound sync at LXMRouter.py:2365). Three gates bound it. A peer whose
announced peering cost exceeds the local max_peering_cost is refused
(LXMRouter.py:2005-2010; default maximum 26, LXMRouter.py:50-51). A
node that already holds MAX_PEERS peers — twenty by default
(LXMRouter.py:43) — refuses further ones (LXMRouter.py:2032). And an
inbound sync offer must present a peering key: a proof of work over the
two node identities at the announced peering cost, 18 bits by default,
checked by validate_peering_key (LXMRouter.py:2300-2312). An offer
without a valid key is rejected with ERROR_INVALID_KEY, and a party
without one may deliver at most one message per transfer
(LXMRouter.py:2382-2385).
Every stored message has paid for admission. Both ingest paths — the
link-packet path and the resource path — run validate_pn_stamps
(LXMRouter.py:2242-2243, again at LXMRouter.py:2401-2402) before
lxmf_propagation ever sees a byte. A propagation stamp is a proof of
work over the message, carried as 32 trailing bytes. The minimum accepted
cost is the node's propagation_stamp_cost minus its flexibility — 16
minus 3 with the defaults (LXMRouter.py:52-54), and the cost knob is
clamped so it can never be configured below PROPAGATION_COST_MIN = 13
(LXMRouter.py:137). An unstamped or understamped message is dropped; a
transfer containing one tears the link down and throttles the sender for
PN_STAMP_THROTTLE = 180 seconds (LXMRouter.py:63, applied at
LXMRouter.py:2449-2452).
So a channel message injected anywhere reaches every propagation node in
the connected network — provided it carries a stamp the ingest node
accepts, which with default configurations means minting 16 bits of work
per post, and never less than 13. Existing Python nodes are the
transport; our nodes are the access points. Nothing upstream has to
change for the distribution half to work, but a node of ours that wants
to distribute has to implement all of the above: the announce format
position by position, peering-key minting and validation, stamp minting
and validation, transfer and sync limits, and the throttle behaviour. A
node built without the peering key gets ERROR_INVALID_KEY on every sync
offer it makes and is limited to one message per transfer; a node that
skips stamp validation accepts what its peers will not. Either way the
distribution half collapses.
Writing does not even require one of our nodes. A client with no presented identity may deliver one stamped message per transfer to any node, and that message is flooded normally. Only reading requires us.
3. What we must not build
The retrieval endpoint has to work without an identity check, or channels do not work at all: many people read the same address, and none of them owns it. The naive form is "client asks for a destination hash, node returns what it holds".
That must not be built. Through peer sync our node also stores every foreign mailbox message in the network. An endpoint that answers for an arbitrary destination hash becomes a traffic-analysis oracle over the whole mesh: ask about any address and learn how many messages are pending, how large they are, and when they arrived. The contents stay encrypted, and it still destroys exactly the property Reticulum exists to provide.
The rule: a channel endpoint may only ever serve entries the node has
positively classified as channel entries. Mailbox entries are reachable
solely through the identity-bound /get path, unchanged from Python:
message_get_request (LXMRouter.py:1482-1484) refuses a client that
presents no identity, and serves only entries addressed to the
destination derived from that identity. Section 5 is what makes the
classification sound rather than a matter of trust.
What the rule does not give: any operator of any propagation node — Python or ours — sees the destination hash of every stored entry in plaintext, and can watch per-address volume and timing locally. The rule closes the remote oracle, the one that answers strangers about addresses they merely name. It cannot close the local view, because carrying traffic means seeing it. Section 4 states what that costs each channel type.
4. Addressing
Both channel types are ordinary 16-byte destination hashes, derived with
the standard Reticulum construction, Destination.hash
(Destination.py:116-130), so that stock tooling can compute them
unmodified.
Open channels derive from the name alone:
name_hash = SHA-256("lxmf.channel." + name)[:10]
dest_hash = SHA-256(name_hash)[:16]
which is Destination.hash(None, "lxmf", "channel", name). Anyone who
knows the name can compute the address, which is the point.
A key-stretching alternative was considered and rejected: derive the address through an expensive KDF over the name instead of one cheap hash, so that an observer holding a stored entry's address cannot cheaply dictionary-test candidate names against it. It buys nothing here. Every open-channel message carries its name in cleartext (section 5), because the node must recompute the address from the name to classify the entry at all — so the name is public the moment the first post exists, and hardening the derivation only slows down legitimate clients on weak hardware. For closed groups the question does not arise, because their addresses do not derive from names.
Closed groups. This section replaces an earlier design that review killed, and the replacement is a proposal, not a settled decision — it changes the shape of the group secret and needs its own review before anything is built.
The earlier draft derived the group address from a symmetric group key and authenticated readers by an HMAC under a second symmetric key. One sentence on why that was unsound, so the scar stays visible: the node never holds any group's key, so it could verify neither the HMAC nor the claimed binding between address and key — the "authentication" reduced to knowledge of a 16-byte address, which is exactly the oracle section 3 forbids.
The proposal: a group is an ordinary Reticulum identity — an asymmetric keypair — whose full key material is shared among the members as the group secret, alongside a symmetric content key for the payloads. The address derives from the group identity with the standard construction and a fixed aspect set:
dest_hash = Destination.hash(group_identity, "lxmf", "group")
Destination.hash accepts an identity (Destination.py:122-124) or 16
bytes of raw hash material (Destination.py:125-126); the proposal uses
the identity form. No name enters the derivation, so a node holding the
group's public key can recompute the address — and verify that key and
address belong together — without ever learning a name or a secret. That
recomputation is what makes closed-group authentication and
classification verifiable (sections 5 and 7) where the HMAC design was
not.
What a closed group hides, stated honestly rather than generously: the name and the contents. Not the existence, and not the traffic. The address stands in plaintext as the first 16 bytes of every stored entry on every node that syncs it, and the envelope that makes classification possible (section 5) marks the entry as closed-group traffic and carries the group's public key. After the first post, any node operator in the network can observe that this group exists, how much it posts and when. Without the group secret the address cannot be derived in advance and the group cannot be found by name — but "unguessable and undiscoverable", as the earlier draft had it, overstated the property. Members who need their group's existence hidden from node operators need a different tool.
Both channel namespaces are disjoint from lxmf.delivery by the name
hash that enters the construction, so a channel or group hash can never
collide with a mailbox hash short of a SHA-256 collision.
5. Message format, and how a node recognises a channel message
A channel message is an ordinary LXMF message whose destination hash is
the channel address. The packed form is destination(16) || source(16) || signature(64) || payload, where the payload is a msgpack array
(LXMessage.py:382-386). The signature covers destination, source,
payload and the message hash (LXMessage.py:364-368 builds the hashed
part, LXMessage.py:375-378 signs it).
For propagation, Python encrypts everything after the destination hash
and appends the propagation stamp: destination || encrypt(rest) || stamp (LXMessage.py:430-435). An open channel message differs in
exactly one step: the rest is not encrypted, so the stored body is
readable. A receiving Python node cannot tell the difference, because it
never decrypts either one. On disk, a propagation node stores the entry
with the validated stamp re-appended (LXMRouter.py:2512-2515) and
strips it again when serving (LXMRouter.py:1549); the stamp is 32 bytes
(STAMP_SIZE, LXStamper.py:15).
Three reserved custom fields carry the channel data (LXMF.py:44-46):
FIELD_CUSTOM_TYPE (0xFB): the discriminator and format version.FIELD_CUSTOM_META (0xFD): the channel name, and the sender's public key.FIELD_CUSTOM_DATA (0xFC): reserved for later use.
Classification of open-channel entries. The classifier takes a stored
entry, strips the trailing 32-byte stamp, and parses the region after
byte 96 — after destination, source and signature, not after the
destination alone — as msgpack. An earlier draft parsed at the wrong
offset and claimed that encrypted mailbox traffic "fails this
immediately, being ciphertext". Both halves were wrong: a literal
implementation would have failed on every legitimate entry, and
ciphertext does not reliably fail a msgpack parse — any byte in
0x00-0x7f is a complete, valid positive-fixint document, and larger
accidental structures parse too. Parsing is a cheap prefilter, nothing
more. The checks, in order:
- The entry is at least
LXMF_OVERHEADplus stamp bytes long. - The region after byte 96 parses as a msgpack array of four or five elements that consumes the region exactly — trailing garbage fails.
- The first element is a plausible timestamp, the fourth is a field map
carrying our
FIELD_CUSTOM_TYPEdiscriminator and aFIELD_CUSTOM_METAwith a name and a sender key. - The authoritative test: recompute the address from the name in
FIELD_CUSTOM_META(section 4) and require it to equal the destination hash the entry is stored under.
Steps 1-3 can, with residual probability, be satisfied by ciphertext. Step 4 cannot be satisfied by accident: the entry must contain a name that hashes to its own address. For a mailbox entry to be misclassified, its ciphertext would have to embed a name whose channel-namespace hash equals the mailbox's identity-derived hash — a second preimage across disjoint namespaces. The realistic false-positive is an entry deliberately constructed to pass, and that is not a false positive at all: an entry addressed to hash(name) that names itself correctly is a channel post, possibly with garbage content, which an open channel admits by definition. Misclassification can waste channel storage; it cannot expose mailbox entries.
Classification of closed-group entries cannot work that way — the payload is ciphertext and carries no name, and no amount of parsing classifies ciphertext positively. The proposal (with section 4's caveat) is an explicit envelope outside the ciphertext:
destination(16) || tag+version || group_pubkey || group_signature
|| ciphertext
The node classifies by recomputing the address from the embedded public key — fixed aspects, no name needed — and requiring equality with the entry's destination, then verifying the group signature over destination and ciphertext against that key. Both checks need no secret. The signature additionally means non-members cannot inject entries into a closed group, which open channels by design cannot promise. The price is the metadata stated in section 4: the envelope is what makes the entry observably closed-group traffic, linkable by its public key. That trade — classifiable but visible, versus hidden but unservable under section 3's rule — is the crux Lew should weigh when reviewing this proposal.
What the signature proves. Every post carries a signature, and the
sender's public key travels in FIELD_CUSTOM_META because
unpack_from_bytes (LXMessage.py:747) resolves the sender through
RNS.Identity.recall (LXMessage.py:776) and leaves the signature
unverified when that returns nothing (LXMessage.py:809-816), which in a
public channel is the common case. Verifying the in-message key against
the source hash makes a channel self-supporting rather than dependent on
whether an announce happened to arrive. But this must not be oversold, as
an earlier draft did with "cryptographically attributable": the key is
attacker-chosen material that travels in the message, and the source
hash derives from it. Verification therefore proves key continuity — two
posts verified against the same key were made by the same key holder —
and nothing about who that holder is. An impersonator mints a fresh
keypair, copies a display name, and every one of their posts verifies.
The remedy is the same as for look-alike channel names: the client
displays a short hash prefix of the author identity beside the display
name, so two authors who read alike are visibly distinct. Key continuity
plus visible key prefixes is the honest offer; it is more than nothing,
and less than identity.
The key costs 64 bytes; on narrow links a client may send it only for the first message per author per time window.
6. The directory, and creating a channel
A node cannot invert a hash, so it cannot enumerate channels from stored entries alone. It does not have to: for open channels the name travels inside the message and is verified on arrival, so every node learns of a channel the moment its first message passes through. The directory builds itself out of traffic, network-wide, with no registry, no gossip protocol and no configuration.
Closed groups never appear in any directory. Precisely: the node knows a classified group's address, its public key and its traffic volume — it can enumerate that a group exists — but it never learns a name, and the directory lists names.
Creating a channel is therefore not an operation. A client picks a name and posts. The channel exists, and it appears in the directory of every node that sees the message. There is nothing to register and no one to ask.
The cost is that names are unowned, so squatting and homoglyph confusion ("general" against "genera1") are possible. An earlier draft called this "the same trade Meshtastic makes", and that equivalence is false: a Meshtastic channel name collides only within RF range, while this namespace is worldwide — there is exactly one "general" for the entire connected network, and whoever posts to it first shapes it for everyone. The global namespace makes squatting strictly worse than in the systems this feature borrows from, and two mitigations follow. Directory entries must be ranked and bounded by activity, not merely accumulated, or the listing becomes a spam surface. And the client must display a short hash prefix beside the name, so two channels that read alike are still visibly distinct.
7. The retrieval protocol
Three request handlers are registered on the lxmf.propagation
destination alongside the existing /offer and /get. Python nodes do
not know them and will fail the request, which doubles as the fallback
when discovery is stale.
/channel/list returns the directory: name, hash prefix, an activity
figure and the timestamp of the most recent post. Open channels only,
paged and bounded.
/channel/get takes a channel name, not a hash, plus an optional
cursor. The node derives the hash itself and serves only entries
classified per section 5. Passing the name rather than the address is
what keeps section 3's rule enforceable for open channels: there is no
name that produces a mailbox address, so the endpoint cannot be aimed at
one.
/channel/auth admits a closed-group reader, under section 4's proposal,
by proof of key possession. The client presents the group public key; the
node recomputes the address from it and, if it holds classified entries
for that address, issues a random nonce; the client returns a signature
over the nonce and the link ID under the group key; the node verifies
against the presented key and serves that one address for the lifetime of
the link. At no point does a group secret reach the node, and knowledge
of an address alone gets nothing: the address must be derived from a
key the client demonstrably holds, which is the verifiable binding the
earlier HMAC design lacked (section 4). Since closed-group entries are
also positively classified by their envelope, section 3's rule holds on
this path with or without the auth step; what auth adds is that
non-members cannot remotely harvest a group's ciphertext and traffic
pattern through our own endpoint. The local view of a node operator is
out of scope for auth, as section 3 states.
Retrieval never deletes. Python's /get treats the client's "have" list
as a purge instruction (LXMRouter.py:1509-1514), which is right for a
mailbox and fatal for a channel, since the first reader would empty it
for everyone. Channel entries on our nodes expire on their own TTL and a
per-channel ring buffer instead.
8. Discovery
Position 6 of the propagation node announce is a metadata dict, and
pn_announce_data_is_valid checks only that it is a dict
(LXMF.py:224-244). Unknown keys pass validation untouched; the
receiving router stores the dict on the peer object (LXMRouter.py:2018)
and otherwise ignores what it does not know. LXMF reserves
PN_META_CUSTOM = 0xFF (LXMF.py:138) for exactly this.
Our node advertises channel support under a namespaced key there:
protocol version, which channel types it serves, directory size, and the
minimum stamp cost it requires for channel posts. Because propagation
node announces travel over the ordinary announce mechanism, every client
in the mesh sees them. lnmsg collects them, filters on the capability
key, and selects by hops_to. That is the automatic discovery the
feature needs, and it costs nothing upstream.
One caveat, from the LXMF source itself: the metadata fields "may be
highly unstable in allocation and availability until the version 1.0.0
release, so use at your own risk until then, and expect changes!"
(LXMF.py:128-131). The mechanism is sound, the field numbering is not
guaranteed. Our entry therefore carries its own version field and the
client must tolerate absence and garbage.
9. Storage, cost and fairness
Two directions of cost, and both need a deliberate answer.
Inbound. A node participating in the peer network receives everything, not only channels. Ours will store and forward the entire LXMF propagation traffic of the reachable network. That figure should be measured on a real node before the server is built, since it decides the hardware floor.
Outbound. Our channel messages come to rest on every Python node in
the network, for up to the thirty-day MESSAGE_EXPIRY
(LXMRouter.py:38, applied from receive time in clean_message_store,
LXMRouter.py:1144-1163), and nobody ever collects a channel post, so
until expiry only storage pressure removes them. Under pressure the
eviction pass weighs every entry by get_weight
(LXMRouter.py:1056-1067) — priority weight times age times size — and
evicts the heaviest first (LXMRouter.py:1188-1196). The stamp value is
stored with each entry and eviction never consults it. An earlier draft
claimed the opposite, citing this very code, and built its politeness
story on posts carrying deliberately low stamps so channels would be
"first evicted". That claim is withdrawn: no such lever exists. What the
weighting actually does is evict big-and-old first and spare
operator-prioritised destinations, so an active channel competes with
other people's mail on exactly those terms — it does displace mail on
foreign nodes under pressure. The honest levers are the remaining ones:
keep posts small (size is a linear factor in the weight), keep volume
moderate, and keep the footprint visible and measured, which is section
10's bargain rather than a technical trick.
Stamps are admission control, not eviction control. Every post must carry the proof of work section 2 describes — at least 13 bits against any conforming node, 16 to clear every default configuration — and our clients mint to the announced cost of the node they post through. That is a mandatory floor the network enforces, not a knob of ours. Above that floor, a channel may declare its own minimum stamp cost as a spam barrier, and here the earlier draft's contradiction has to be resolved rather than papered over: one stamp value cannot be pulled low for eviction politeness and high for spam defence at once. The resolution is that the politeness direction never existed (see above), so the stamp is pulled in one direction only — upward, as a cost on posting. Its enforcement is also honestly narrower than "the network": our nodes refuse to serve entries below the channel's declared minimum and our clients refuse to display them, but a Python node stores and floods anything at or above its own floor regardless. A per-channel stamp floor filters what readers see through us; it does not keep spam off the carrier.
Channel entries on our own nodes live in a separate quota from mailbox entries, so that a busy channel can never crowd out the mailbox function we also promise to provide.
The flood is global and the use case is local. This is a real mismatch and it gets recorded as a weighed decision, not smuggled past. The mechanism puts every post of every channel onto every propagation node in the connected network for up to thirty days, while the motivating use case — "type a name, talk to whoever is around" — is local chat. Alternatives were considered. Carrying channels only on our own nodes avoids imposing on foreign operators but gives up the free transport that is this design's entire reason to exist, and is kept as the degraded mode in section 10 rather than the default. Regionally scoped names ("hb.general") reduce reader collision but change nothing about propagation — the carrier does not read names, every post still floods globally. A sender-chosen TTL shorter than thirty days does not exist in the protocol: expiry is the storing node's constant, counted from receive time, and nothing in the message can lower it. So the trade is accepted for v1 with open eyes: global flood is the price of zero-infrastructure distribution, it is bounded by the stamp floor and the smallness of chat messages, and the inbound measurement above doubles as the check on whether the price is as low as this paragraph assumes. If the measured numbers say otherwise, the decision gets revisited, not defended.
10. The compatibility bargain
Everything above rests on Python nodes accepting without inspection and flooding without filtering, gated only by stamps and peering keys. That is a factual property of the current implementation, not a promise anyone made us. A single upstream commit restricting storage to destinations that have been announced would end the free-transport half, and from a node operator's perspective that would be an entirely reasonable change.
This cannot be secured technically, only socially. We say plainly what we are doing, we make it identifiable in the announce metadata, and we keep our footprint on foreign nodes visibly small. Doing it quietly and being found out later closes the route permanently.
The degradation plan, decided now rather than improvised later. If upstream starts classifying or rejecting channel entries — requiring announced destinations, parsing payloads, or filtering our discriminator — the design degrades to an overlay instead of breaking. Our nodes are full propagation nodes that peer with each other statically, not only by autopeer, so channel distribution continues over our own peerings with the same sync protocol; clients already post and read through our nodes, so nothing changes for them. What is lost is exactly the free transport: reach shrinks from every propagation node to the connected set of our nodes plus whatever Python nodes still carry unclassified traffic. The mailbox function is untouched either way, because our nodes are conforming propagation nodes first. To notice the change when it happens rather than months later, a canary: periodically post a test channel message through a Python node and confirm arrival at one of ours over a Python-only path; when that stops, the assumption behind sections 2 and 9 has expired and the overlay mode becomes the documented default.
The conformance obligation is the ordinary one: our node has to be a correct propagation node first and a channel node second. Stamp validation, peering keys, sync limits, throttling and announce format all have to match Python exactly, or autopeering will not happen and none of this runs. Channels are an addition to a compatible node, never a deviation from one.
11. Open questions
- The inbound storage and bandwidth figure for full peer participation (section 9) is unmeasured. It should be measured before the server is designed, not after — it decides both the hardware floor and whether section 9's global-flood trade is as cheap as assumed.
- The closed-group proposal in sections 4, 5 and 7 needs its own review: whether the shared-keypair group secret is acceptable, and whether the envelope's metadata cost (visible group existence, linkable public key) is a price the use case can pay.
- Whether closed groups need key rotation, and what happens to the address when a member leaves, given that the address is derived from the group keypair.
- Directory ranking: what activity measure, over what window, and how large a listing before it needs paging or filtering.
- Whether channel messages should carry a
FIELD_THREADequivalent, so replies can be shown threaded rather than flat.
Telemetry
How a node reports where it is and how it is doing, and why almost none of that format is ours to decide. This applies to every firmware and every platform port, present and future.
The goal
A node — a tracker in a rucksack, a solar relay on a roof, a handheld with a display — produces readings that belong somewhere else: a position, a battery charge, a link quality. Telemetry is the path from that reading to a map pin or a database row, over the same mesh the node already speaks.
The receiving end is not ours. Sideband and Columba display telemetry today, from any peer, without knowing what produced it. A node that emits the format they already read is useful on the day it ships; a node that emits its own format is useful when someone writes a viewer for it. So the format is adopted first and the node conforms to it, not the reverse.
Where the format comes from, and how far it is settled
Telemetry rides inside a normal LXMF message, in the fields dictionary
of the encrypted payload — not in a destination of its own, not in a
side protocol. It therefore inherits everything LXMF already provides:
end-to-end encryption, the three delivery methods, propagation nodes,
and receipts. A peer that does not know the field sees a message with no
text, which is the correct failure — subject to the empty-content rule
below.
The field numbers are LXMF's (leviculum-lxmf/src/constants.rs:52-78).
Three matter here: FIELD_TELEMETRY (0x02) carries one node's readings,
FIELD_TELEMETRY_STREAM (0x03) carries many nodes' readings collected
by a third party, and FIELD_COMMANDS (0x09) carries the request that
asks for the second one.
The content of FIELD_TELEMETRY is Sideband's, defined by its
Telemeter class in sense.py: a msgpack map from sensor ID to that
sensor's packed value. Sideband is the origin of this format and, for
most of it, the only implementation. Columba reimplements two sensor
IDs of the twenty-four — time and location — in TelemeterCodec.kt,
naming sense.py as its own reference. That is corroboration for those
two, not a second independent implementation of the format. For every
other sensor we encode, we are the second implementation, and the
accept-set rule of Python-RNS
compatibility applies with full force: we
must read what the origin emits, not only what we would have written.
How settled the format is differs by part, and the difference decides how much freedom a reader has:
| Part | Status |
|---|---|
FIELD_TELEMETRY, time and location sensors | Two implementations agree. Settled. |
| Other sensor IDs | One implementation. Follow sense.py exactly. |
The collector exchange (FIELD_COMMANDS / FIELD_TELEMETRY_STREAM) | Two implementations that disagree today. See the wire table below. |
Why citations to Sideband and Columba carry no line numbers. This is a new rule, stated here rather than inherited: those repositories are not in this tree and not pinned by it, so a
path:linecitation into them cannot be checked by the citation guard and cannot be trusted to still point at what it claims. They are cited by symbol, which survives the edit that moves a line, plus the upstream commit the reading was taken from — Sideband2000d81, Columba0930293— so a reader can reconstruct the line. Checks that are actually checks argues the converse case for in-tree references and does not cover out-of-tree ones; this is the gap being filled, not an application of that page.
Which viewer wins a disagreement
Columba is the app this project targets: it is the one whose users we serve first and the one we can realistically patch. That ordering decides what we build for. It does not decide what is correct on the wire, and the two questions are answered in this order:
- Does one form break the other implementation? Then that form is a bug, not a tie, however confidently it is implemented. Emit the form both accept and report the bug to whoever emits the other.
- Are both forms legitimate readings of the format? Then prefer the one Columba displays.
- Would our output only be understood by a build carrying our own patch? Then it is not permitted, whatever step 2 said.
Step 3 is the operative test, and it is the one that can be checked before the change lands rather than after: run the output against an unpatched build of both viewers. Three concrete temptations it rules out, each of which would work and each of which would leave our node broken against every implementation but the patched one — a receiver taught to accept an unverifiable signature so a node can skip announcing its delivery destination; a receiver taught to ignore timestamps entirely so a node can skip keeping a calendar at all; a receiver taught to reinterpret a placeholder coordinate so a node can transmit without a fix.
Rules that hold regardless of platform
Telemetry always stamps — never ahead of the calendar estimate
SID_TIME and the location sensor's last_update are real UTC seconds,
and they are load-bearing at the receiver in three separate places:
- Storage keys on them. Sideband dedups on exact
(source, ts)equality and otherwise inserts, so a wrong timestamp is stored rather than rejected — a node emitting a wrong time appears in the viewer, wrongly dated. - Collection filters on them. A collector serves only rows strictly newer than the requester's stored cursor.
- The cursor advances over every row seen, saved or not. So one row timestamped in the future raises a requester's cursor past it permanently: every honest later reading from that collector is "not newer" forever, and nothing anywhere logs a refusal.
That third mechanism is why stamping is directional. A future stamp is poison — permanent, silent, and it hurts every subscriber of a collector, not the sender. A past stamp merely sorts too far back: visible, attributable, harmless. An earlier version of this page drew the conclusion that a node without a verified clock must not report at all ("arms 1–3 or silence"); that rule is withdrawn. It silenced exactly the switch-on-and-it-works trackers this feature exists for, and it defended the cursor at the wrong end.
The binding rule comes from the anchor model of Time and clocks: a telemetry node always stamps with its best honest calendar estimate, and never ahead of it — every fallback anchor lies in the past, so an uncertain calendar errs backwards by construction; a source that is wrong but plausible is stamped as-is, the residual the anchor model names. A node whose calendar is still at the build floor reports recognisably old readings, and the cost is priced per path. Where the reading is displayed directly, it sorts to the back of the viewer: visible, attributable. Where it travels through a collector, the real price is higher: a collector serves only rows above a requester's cursor, so a birth-stamped row sits below every already-synced requester's cursor — not sorted backwards but invisible to that requester, indistinguishable from loss, until the tracker heals. The cost ends at the node's first plausible contact — outright at arms-1–3-quality time, provisionally when healed from traffic (the one honest cost).
The cursor is defended at the reader instead: our collector (#239)
clamps inbound stamps beyond its local plausible-now to receive time
for indexing and cursor advancement, keeping the original stamp as
local display information (ingress
clamping).
The clamp is armed only while the collector's own calendar is healed:
an unhealed collector clamping "to receive time" would drag every
honest current stamp down to its own ancient notion of now and
blackhole them below every synced cursor. Until it heals, it takes
in-window sender stamps as-is — the same stamps double as its
healing evidence — and rows ingested before healing keep their index
stamps; there is no re-index. Dedup keys are never clamped: dedup
runs on content and transient ID, the clamp covers ordering and
cursor semantics only. And what a collector serves is the clamped
value — a stream row has one timestamp slot — while keep-for-display
stays local; note the stated limit that the row's packed_telemetry
still carries the sender's raw SID_TIME claim to any downstream
parser. Foreign future stamps then cannot starve anyone's cursor
through us.
A node still has to know which arm anchored its calendar — the provenance requirement of Time and clocks — not as a gate on reporting but as the diagnosis surface: "why does this tracker report from 2026-01-01" must be answerable from the node itself.
No fix, no position — and absence has one encoding
A sensor with no reading is absent from the map. It is never present with a placeholder: no zero coordinates, no last-known value restamped as current, no accuracy invented to fill the field. A viewer cannot distinguish a placeholder from a measurement, and 0°/0° is a real place in the Gulf of Guinea.
The format does not decide how absence is spelled, so we do. Sideband
packs every active sensor, and a sensor with no data packs as its SID
mapped to None — key present, value empty. That is two encodings of
absence in one format, and the choice is ours to fix:
- We emit the omission. A sensor without a reading contributes no
key.
Telemeter.from_packedinstantiates only the sensors present in the map, so an omitted key round-trips cleanly. - We accept both on read. A SID mapped to
Noneis a sensor without a reading, not a malformed message. This is the accept-set half of Python-RNS compatibility: the origin emits a form we would not write, and refusing it would be our defect.
The encoder is multi-sensor from the first line
The unit is a Telemeter with n sensors, not a position packet with extras bolted on later. Position, battery, temperature and link quality are the same code path with different sensor IDs, and every board that measures anything gets to report it without a second design.
The justification is the format's own extension mechanism, not our future plans: a viewer ignores a sensor ID it does not know, so emitting a reading no current viewer displays costs one map entry and breaks nothing. A sensor left out of the encoder, by contrast, needs a new design to add later.
Followed through on the boards: a node reports everything it can
measure, so both LNode builds carry the nRF52's die temperature as a
Temperature reading (SID 0x07) whether or not a baseboard is fitted.
The die thermometer belongs to the SoftDevice, so it is read through
sd_temp_get and never off the TEMP registers; its 0.25 °C step is
already exact at the two decimals Sideband rounds to, so the value goes
on the wire unrounded. Battery charge stays feature-gated — it needs a
baseboard that has a gauge — and an absent sensor contributes no key.
The same rule adds the physical link (SID 0x05), which is the one
sensor that describes the mesh rather than the box. It is the last
frame the radio received, not a mean over a window and not the best of
one: a mean on a mesh averages over whichever neighbours happened to
transmit, so it falls when a distant node joins and reads on a viewer as
the near link degrading. The last frame is also what the references
report under this name — RNS.Link.rssi/.snr keep the last received
packet's figures (RNS/Link.py), and RNodeInterface's r_stat_rssi
belong to the frame the stat bytes arrived with — so our rssi and a
Python-RNS peer's rssi are the same quantity. q is Reticulum's own
SNR-to-quality map (RNodeInterface.Q_SNR_*, whose floor drops 2 dB per
spreading factor), not a scale of ours; where the PHY has none defined
the slot goes out nil and the two measurements beside it still go.
A board that has heard nothing for longer than the fastest reporting
cadence sends no physical-link sensor at all, which is the general
absence rule applied to a reading that would otherwise never expire: a
stale rssi is the one number in the set that says the opposite of the
truth. The rule and its bound live in
leviculum-nrf/telemetry-policy/src/link.rs, where a host test can
reach them.
Two sensors of the type stay empty on every board we build, and for a
reason worth writing down rather than a gap. charging inside the
battery sensor is None because a voltage divider cannot tell charging
from discharging and no board brings a charger status line to the MCU —
upstream's own answer on the RAK4631 (NRF_APM) reads the nRF52's USB
VBUS state, which is "USB is plugged in", not "the pack is gaining".
power_production is empty for the same kind of reason: not one carrier
exposes a panel current, and on the Solar Node in particular every XIAO
pad is accounted for in boards/solarnode.rs with nothing left over
(Codeberg #233).
A reporting message carries no text
A telemetry message sets content and title to empty. This is a hard
wire requirement, not tidiness: Sideband suppresses the notification for
a telemetry-bearing message only when both are empty, so a reporting
node that fills in either one notifies its recipient once per reporting
interval, forever, and the feature is indistinguishable from spam.
Delivery is opportunistic unless a port shows otherwise
Of LXMF's three methods, a reporting node uses the opportunistic single packet by default. A direct delivery pays a link setup — three round trips before any payload — for a payload of roughly fifty bytes, and a propagated delivery pays a node round trip and adds a delay that makes a position stale. A port that needs delivery confirmation, or that reports to a target it can only reach through a propagation node, may choose differently, and states why.
Cadence is policy; airtime is the interface's business
A telemetry producer states when it wants to report: a minimum interval, a minimum distance moved, a maximum interval as a heartbeat so that "stationary" stays distinguishable from "dead", an accuracy threshold, and a settle time so a cold start does not spend the channel on a drifting first fix.
It does not state when the radio may transmit, and it holds no airtime figure of its own. Duty cycle, spacing and back-off belong to the interface — see Interface isolation and Regulatory airtime, which also settles how any such figure is to be described: modelled, and a floor rather than a total, because the board's own meter clears when it is read.
The configuration surface itself is settled in #236: setting a target is the on-switch, and the tracker and station profiles bundle the cadence defaults.
Activation is configuration, not firmware
Setting or changing the telemetry configuration never requires
rewriting firmware. lnflash gains a config-only session — the same
post-flash serial configuration channel, entered without a UF2 write —
and the #238 control envelope and #235 remote management make the same
configuration changeable at runtime later. The reason is operational: a
node already running in the field must be adoptable into telemetry, and
retirable from it, where it hangs.
One input, and it is the address. Most users configure nothing beyond the target, so the target is the only thing configuration may require: profiles bundle the cadence and one of them is the default (station — a node that does not move is the common case, and a tracker misconfigured as a station still proves it is alive on the heartbeat, where the converse merely costs airtime). Everything else is an expert flag underneath a preset, in the same shape as the radio menu.
Sending the position is the switch for sending everything
Reports go out iff both a target is configured and a position source is — a fixed position set, or a GNSS receiver built into the firmware and not switched off (Lew, 2026-08-30). The reading is intent, not possession: a receiver that has never seen sky still counts, because the four use cases behind this feature — tracker, sensor node, mobile transport node, quasi-modem — are all nodes whose operator meant to say where they are, and a tracker that goes silent in a garage is indistinguishable from a dead one. Such a node keeps reporting on the heartbeat with the position absent and its other sensors fresh, which is the designed behaviour and not a degraded mode. GNSS is optional throughout; a fixed position is settable on any board, with or without a receiver.
What the rule refuses is the unconfigured node: a target and no
answer of any kind to "where am I". It sends nothing — not a heartbeat,
not a battery reading — and says why, state=no-position-source on the
same surface and the same cadence rule as awaiting-key. Position-less
telemetry is deliberately not a mode: a node that reports readings
a collector cannot place is a row on a map with nowhere to put it, and
the operator who wanted one had only to set a position. Setting one, or
flashing a build that has a receiver, leaves the state at once — it is
a runtime path, like applying a target, and needs no reboot.
The destination hash alone is enough
A user knows the LXMF address. Requiring the public key alongside it would make the common case the hard one, so the key is optional wherever a target is set: hash-only is expected, and the node resolves the key itself over the air — a path request is answered with the destination's announce, and the announce carries the identity.
This creates a state that did not exist when a target implied a key: the
node holds a perfectly valid target it cannot yet encrypt to. That state
is stated, not waited out silently. A reporting node says which of
four it is in — off, no-position-source, awaiting-key, ready —
on the same surface as its clock provenance, and a target that never
resolves is then a visible condition rather than an absence of packets.
no-position-source outranks awaiting-key in that line: a node that
will not report anyway must not spend airtime chasing a key it has no
use for.
The consequence for the previous rule is that the immediate report fires on key arrival, which for a hash-only target is later than the moment the target was set. A frame that carries a key skips that wait, which is the whole benefit of carrying one.
Setting a target emits one immediate report
When a telemetry target becomes usable — set or changed with a key, or set by hash and then resolved — the node sends one report at once, regardless of the configured cadence. Success must be observable within seconds: a station profile on an hourly heartbeat would otherwise leave the operator without any confirmation for up to an hour. The immediate report follows every other rule in this document — no fix, no position; empty content and title.
The report is owed until it is actually sent, not until it was first due: a node that has the key but no path yet retries rather than counting an attempt it could not make. And a report that a target cannot receive is a report nobody can verify, so the node announces its own delivery destination before sending — the announce is what puts our public key in the receiver's hands.
A request from the target triggers one report
Sideband lets a peer ask for telemetry on demand: an LXMF message whose
FIELD_COMMANDS field (reference/LXMF/LXMF/LXMF.py:16) carries the
TELEMETRY_REQUEST command with a timebase. A reporting node answers
one such request with an immediate report — the same report a target
write arms, subject to the same rules (owed until sent, announce first,
attempt floor). That gives the operator a position on demand instead of
at the profile's cadence, and gives a field test a way to force the
whole chain — path request, answer, data — at a known moment.
The gate is authentication, not reachability: the request must name the configured target as its source and carry a signature that verifies against the target's identity. The target is the only allowed sender for now; a general allow list is a later step and will follow the collector rule above — explicit, empty by default, set only off-radio. A request from anyone else is answered with nothing but a debug line, because an unauthenticated trigger for a radio transmission is a remote airtime primitive.
Requests are rate-limited to one accepted request per min_interval_ms
of the active profile; a request inside the window is logged and
dropped, not queued. The request's timebase is read but does not change
the answer: a node keeps no history, so the current reading is the
answer to every timebase.
Sideband's own source is not vendored under reference/, so the
command id (0x01) and the value shape (a list of command maps,
[{0x01: <timebase>}]) are a documented assumption — recorded at
COMMAND_TELEMETRY_REQUEST (leviculum-lxmf/src/telemetry.rs:678) —
to be verified against a captured Columba/Sideband request and then
pinned as a fixture.
On the host side, lnsd neither registers an lxmf.delivery
destination nor runs a telemetry policy today, so there is nothing for
a request to trigger there; the screen in leviculum-lxmf is the seam
a host-side consumer will reuse when that changes.
Fan-out is the expensive shape; collection is the cheaper one
Telemetry to n recipients is n individually addressed and individually encrypted messages. Mutual reporting in a group of n is therefore n(n−1) messages per interval — thirty a minute for six people at a one-minute cadence.
Routing the same group through a collector is n reports plus, for those who want to see the others, n request/response pairs: eighteen rather than thirty for the same six. The saving is real but it is roughly a third, not fivefold, and it is not free in kind — a stream response carrying several sources will exceed one packet and become a link plus a resource transfer, where a report is a single packet. Both figures belong in a port's own arithmetic; neither is a licence to skip it.
There is no third shape. Reticulum has no multi-hop broadcast at all —
see Public channels over
LXMF, which
settles this at the protocol level rather than by what two apps happen to
implement. LXMF's FIELD_GROUP is conversation metadata and not a
delivery mechanism; it would not produce a broadcast even if every
viewer read it.
Relaying someone else's readings needs permission
A collector redistributes positions of people who are not asking for
that redistribution. It therefore answers requests only from an explicit
allow-list, empty by default, set only through a path that is not the
radio: on a host that is the config file, as remote_management_allowed
is (leviculum-std/src/config.rs:119, empty by default at
leviculum-std/src/config.rs:375); on a firmware with no filesystem it
is the local control channel, which is the same requirement in a
different envelope.
The same applies to precision. Where a deployment wants a node's position blurred, the blurring happens at the producer, before the reading is packed — a receiver's promise to round a coordinate is not privacy.
A port may ship send-only
Emitting telemetry needs an encoder, a target and a cadence. Consuming it needs the inbound LXMF path, a peer table and something to show. A port may implement the first without the second, and nothing in this document may be read as requiring both: a tracker that cannot receive is a complete node for its purpose.
The collector exchange
Two implementations exist and they disagree, so this table is normative for us rather than descriptive of them.
| Element | Form | Note |
|---|---|---|
| Request | FIELD_COMMANDS = list of single-entry maps | Sideband's Commands.TELEMETRY_REQUEST is key 0x01 |
| Request argument | [epoch_seconds, is_collector_request] | An absolute UTC epoch, not an interval or an age |
| Legacy request | {0x01: epoch_seconds} — bare scalar | Must be accepted on read; the origin still emits it and infers is_collector_request = true |
| Response | FIELD_TELEMETRY_STREAM = list of rows | |
| Row | [source_hash, timestamp_seconds, packed_telemetry, appearance] | Always four elements, None in the fourth when there is no appearance |
The four-element rule is the one that costs something. Sideband indexes the fourth element unconditionally, so a three-element row raises inside its ingest and aborts processing of the whole message, not just that row — while Columba's native collector emits three elements when appearance is absent, and its own Python-backed collector emits four. Under the ordering above this is step 1, not step 2: one form breaks a conforming receiver, so we emit four and the three-element form is a bug to report.
A collector additionally never serves a row whose timestamp lies in the future relative to its own healed clock. Serving one permanently starves the cursor of every requester that sees it. Under ingress clamping such a row cannot enter the index while the clamp is armed — the stamp is clamped to receive time on ingest — so for a healed collector this serving rule is defence in depth. It is the primary barrier for exactly one population: rows ingested before the collector's own calendar healed, which were taken as-is and keep their index stamps (there is no re-index).
What a reporting node owes
- An announced delivery destination. A receiver verifies the LXMF signature against the sender's public key, which it can only have from an announce. A node that reports announces its delivery destination, with a display name in the announce data — otherwise the reading is unverifiable and the pin, if it appears at all, is a hex string.
- No proof for a message it cannot keep. That announce is read by
every peer as the claim that messages sent to this hash will be
received, and a board cannot keep it: it has no inbox and no message
store, and the only reader of what arrives is the telemetry
reporter, which keeps a request and discards the rest. So
the board's delivery destination carries
ProofStrategy::None(LxmfNode::delivery_destination_without_inbox, whose only caller isregister_delivery_destinationinleviculum-nrf/src/telemetry.rs) whilelntd, which does have an inbox, keepsdelivery_destinationand itsProofStrategy::All. Withholding a proof claims nothing; theProofStrategy::Alla board used to inherit hadNodeCoresign one for every arrival, so a Python peer's LXMF marked a message DELIVERED and the bytes were dropped without a line. A discard is now always named ([TELEMETRY] discarded ... reason=). - No link either. The same destination used to ACCEPT an inbound
link —
accepts_linksistrueon a freshDestination, as Python-RNS has it — and then serve nobody:LinkDataReceivedis read only by the propagation role's own destination (leviculum-nrf/src/pn.rs), a link's resource strategy defaults toAcceptNone, and the link path never reaches the[TELEMETRY] discardedline above. A link is a session, so that swallowed a conversation rather than a packet, on every nRF board — the three binaries register this destination whether a telemetry target is configured or not.delivery_destination_without_inboxnow switchesaccepts_linksoff as well: the peer gets no link proof, which is the same thing the link cap (max_links) already sends it when the table is full, and its establishment timeout handles it. Destinations that DO serve links —lntdand every other caller ofLxmfNode::register, and the board's own propagation destination — are untouched. - The target's public key, before anything else. Encrypting to a
destination is impossible without it. There is no broadcast around
this: the reference's transmit-on-all-interfaces branch
(
reference/Reticulum/RNS/Transport.py:1177-1182, our equivalentsend_on_all_interfaces,leviculum-core/src/transport.rs:4292) applies to a packet that already exists, and building one required the key. So a port either preconfigures the target identity or waits until it has heard the target announce. - A path, or a request for one.
send_to_destination(leviculum-core/src/transport.rs:4408) fails without a path entry. The primitive for obtaining one isrequest_path(leviculum-core/src/node/mod.rs:3405); a node with the key but no path asks and waits rather than giving up. - An out-of-band trust step at the receiver, in the operator's hands. Sideband can be configured to ingest telemetry only from trusted peers, and both viewers gate collector requests on an explicit allow-list. When that step has not been taken, a correctly reporting node is silently invisible, and it looks exactly like packet loss. Documentation of a reporting port says so; a diagnostic that cannot distinguish the two cases is worth building before the third support question arrives.
What a receiving node owes
lntd (leviculum-cli/src/lntd.rs) is the endpoint side of the same
contract: one announced LXMF delivery destination, whose hash is what
lnflash --set-telemetry writes into a board. It attaches to a running
shared instance the way lnmsg and lnpnd do, and every report it
accepts becomes one row in a SQLite file.
The rule it is built around is the inverse of the decoder's tolerance.
Telemeter.from_packed skips a sensor ID it has no class for, and our
decoder skips it too — right for a codec, wrong for an archive. So the
raw Telemeter blob is stored on every row, beside the columns for
the sensors we did understand. A field test that silently dropped the
one field nobody had implemented yet would have nothing to go back to,
and Sideband's sensor set is still growing. A blob that does not decode
at all is likewise a row that says so rather than a message on the
floor; the only thing that produces no row is a Telemeter map with zero
entries, which is what the producer rule above says never to send.
Two further consequences of "lose nothing", both tested in
leviculum-cli/src/lntd_store.rs:
- The LXMF message id is a UNIQUE column. A restart across a write neither loses the row already committed nor doubles it when the same report is delivered again.
- The database is the only durable state, and the hooks never touch
it. Rows cross a channel to a thread that owns the connection, so a
synchronous = FULLfsync never happens under the core lock. Shutdown drops the sender and waits for that thread, so a report already handed over is written before the process returns.
There is no query surface and no viewer. WAL mode is what makes the
operator's own sqlite3 session safe against the daemon that is still
writing.
What lxmf-node does with a telemetry message
lxmf-node (leviculum-lxmf-node/src/telemetry.rs) is the same
contract without the database, because the helper is what a field host
runs when it needs a collector today rather than a deployed lntd. A
reporting message has an empty body, so the lxmf_msg_received line it
has always emitted says nothing about it beyond who sent it. Every
FIELD_TELEMETRY value, and every row of a FIELD_TELEMETRY_STREAM,
therefore gets a second line — EVENT lxmf_telemetry_received src= via= time= lat= lon= alt= speed= bearing= accuracy= fix_time= battery_pct= battery_charging= battery_temp_c= rssi= snr= link_q= temp_c= power_w= fields_hex=, one key per sensor this build decodes, none where the
sensor had no reading so the key set is the same on every message —
and one JSON row appended to <LXMF_STORAGE>/telemetry.jsonl, which is
what survives a restart of the helper and a rotation of its log. src
is who measured and via who delivered; they differ exactly when the
reading came out of a collector's stream. A blob that does not decode
becomes lxmf_telemetry_undecodable with the same fields_hex, and
both sinks carry the raw packed bytes on every reading, decoded or not,
for the reason the section above gives: the decoder's tolerance is an
archive's data loss.
It also sends one, which is ours alone: send_telemetry <hex> <telemetry_field_hex> puts a packed Telemeter blob on the wire as an
empty-bodied FIELD_TELEMETRY message and acks it as lxmf_msg_sent … fields=telemetry. The Python helper has no twin for it, as it has none
for the pn_* verbs, and the verb is additive, so a scenario that never
says the word drives either helper unchanged. The blob is decoded once
as a gate and then travels verbatim — re-encoding it would shorten a
sensor this build has no arm for — which is what lets a two-helper
loopback compare the hex a driver typed against the fields_hex the
receiver prints (leviculum-lxmf-node/tests/telemetry_loopback.rs).
The extension ladder
The format is fixed by implementations we do not control, so extending it is a cost with a blast radius, not a design choice. Every addition climbs from the bottom and stops at the first rung that works.
- An existing sensor ID already carries it. Sideband defines twenty-four, well beyond position: battery, temperature, pressure, physical link, power production and consumption, and free-text information. A reading that fits one of them is not an extension, and emitting it costs nothing even where no viewer shows it yet.
FIELD_CUSTOM_META(0xFD) carries it — under someone else's keys. This slot is not free. LXMF reserves it for private use, and Columba has claimed it: it carries an unnamespaced map with the keyscease,expires,approxRadiusandts, and a truthyceasemakes Columba delete the sender's whole track. There is one value per message. So a node targeting Columba may emit Columba's semantics here — a bounded sharing session that expires itself is the case that earns it — and may not put its own vocabulary in the same map. An extension of our own does not go on this rung.- A new field number is genuinely required. Then it belongs upstream in LXMF, proposed as such, and not shipped into one client ahead of that. A field number minted by us and understood by one app is a fork of the format with a friendlier name.
Two conditions apply at every rung. An extension must be demonstrated between two implementations that are both ours before it is offered to anyone else — what gets proposed is then a working feature rather than an idea, and the cost of being wrong stays inside this project. And the failure mode on a viewer that does not participate must be written down: an extension whose effect on an old build has not been stated has not been designed.
Non-goals
- A second format. Not even for our own daemon-to-daemon path: one encoder, one decoder, one thing to get right.
- Mirroring a viewer's internals. Session bookkeeping and collector scheduling are each app's business. We match the wire, not the state machine — the same distinction Python-RNS compatibility draws between compatibility and parity.
- Telemetry as a transport diagnostic. What a node reports about itself is not how the mesh is measured; that is periculum's job and the status surfaces'.
Checklist for a port, or for a new sensor
- Does every stamp come from the calendar clock of Time and clocks — best honest estimate, never ahead of it — and can the node say which arm anchored it? No clock state blocks reporting; an unanswerable "where did this time come from" does block shipping.
- Does every sensor have a "no reading" state that omits its key?
- Does an existing sensor ID fit? Climb the ladder from rung one.
- Is the cadence stated as policy alone, with no airtime figure inside the telemetry module?
- Are
contentandtitleempty on every reporting message? - Does the node announce a delivery destination with a name, and does
it have a defined answer for a target it has never heard? The
defined answer is
awaiting-keyplus a path request, reported on the node's own status surface — not silence, and not a refusal to accept the target. - If it collects for others: allow-list empty by default, set off the radio, four-element rows, inbound stamps clamped on ingest once the own calendar is healed (taken as-is before that), no future timestamps served?
- What does a viewer that lacks this sensor show? Answer before emitting.
Items 1 and 7 are test-bound beyond this page: the rule-by-tier matrix in Testing the model names the cells. For a collector, the ingress-clamp row — clamp, cursor safety, dedup, unhealed behaviour, Codeberg #239 — and the authorship-interop row are the ones an implementation must land green, with its implementing issue.
An LXMF propagation node on the boards' internal flash
This page used to cost the store on a QSPI NOR part — 1 MB on the Pocket
V2, 2 MB on the T114. Neither board carries one: both beliefs came from
an EXTERNAL_FLASH_DEVICES line that is a template default on both
vendors, and three units answered nothing to a JEDEC read (081522b2; the
evidence sits in the board files, CONFIG,
leviculum-nrf/src/boards/t114.rs:171 and CONFIG,
leviculum-nrf/src/boards/rak4631.rs:186). So the store went where there
is flash: 16 pages of the nRF52840's own flash between the firmware
image and the persistence pages (STORE, leviculum-nrf/memory.x:91,
landed in 81fcb46e; the log format chosen in 59c36129). Every capacity,
endurance and scan figure below is recomputed for that region, and the
region is two orders of magnitude smaller than the part this page first
costed: 176 field-sized messages, not 5 544. If a board or an add-on
ever brings a QSPI part, the arithmetic is the same arithmetic with a
bigger region and a tenfold larger erase budget; nothing below assumes
one.
Codeberg #384 asks whether such a store should hold an LXMF propagation node, so the mesh has somewhere to put a message when the recipient is not reachable. The walk that prompted it had a link built from one phone to another across two of our nodes and a hill, which is exactly the topology where a store matters.
This page establishes what the role obliges us to, measures what it costs on the region we now have, sets out the options, and recommends one. It is a design document. Nothing here is a status page; what is open belongs on the tracker.
The framing binds the whole argument. The reference is a source of ideas, never a blueprint. What binds us is wire and semantic compatibility: a Python or Sideband peer must be able to use our node as a propagation node without knowing it is small. Everything else is ours to design, and a board in a pocket is not a server in a basement.
1. What the role obliges us to
Established against reference/LXMF at 1.1.0.
The destinations, and the verbs on them
A propagation node owns one inbound SINGLE destination,
lxmf.propagation, created from the router identity
(propagation_destination, LXMRouter.py:190). Two request handlers
hang off it, and they are the entire public protocol:
| Verb | Path | Who calls it | Handler |
|---|---|---|---|
| offer | /offer | another propagation node | offer_request (LXMRouter.py:2266) |
| get | /get | a client (Sideband, lnmsg, rnsd) | message_get_request (LXMRouter.py:1482) |
The paths are constants on the peer
(OFFER_REQUEST_PATH, LXMPeer.py:14; MESSAGE_GET_PATH,
LXMPeer.py:15). A third destination, lxmf.propagation.control,
carries the operator verbs /pn/get/stats, /pn/peer/sync and
/pn/peer/unpeer (STATS_GET_PATH, LXMRouter.py:89;
SYNC_REQUEST_PATH, LXMRouter.py:90; UNPEER_REQUEST_PATH,
LXMRouter.py:91) and is behind an allow-list, so it is not part of
what a stranger can drive.
Messages move in three shapes, and only three:
- A client uploads one message. A single link packet carrying
[timestamp, [lxmf_data || stamp]](propagation_packet,LXMRouter.py:2234). No peering key is needed for a single message. The node proves the packet — and it proves it after storing, not before (packet.prove,LXMRouter.py:2255). - A peer offers a batch.
/offercarries[peering_key, [transient_id, …]]; the node answersTrue(want all),False(want none), or the sublist it wants (offer_request,LXMRouter.py:2266). The bodies then follow as one Reticulum Resource. - A client drains its mailbox.
/getwith both fieldsNonereturns the list of transient IDs held for that client's delivery destination; a second/getwith[wants, haves, limit]returns the bodies and deletes everything inhaves(message_get_request,LXMRouter.py:1482). The client sends the purge only after it has taken local delivery (message_get_response,LXMRouter.py:1607).
A transient ID is SHA-256(lxmf_data) where lxmf_data is
destination_hash || destination-encrypted payload. We already
implement the client half of this exchange (MESSAGE_GET_PATH,
leviculum-lxmf/src/propagation.rs:25).
What is protocol, and what is that implementation's bookkeeping
Per message the reference keeps seven fields
(propagation_entries, LXMRouter.py:2518): destination hash,
file path, receive timestamp, size, handled peers, unhandled peers,
stamp value.
Of those, three are protocol: the destination hash (it decides who
may /get the message), the transient ID (the key of every exchange),
and the bytes themselves. The stamp value is protocol-adjacent — a
peer drops messages whose stamp value is below its requirement
(sync, LXMPeer.py:267) — but a node that requires nothing needs
only to remember zero. The receive timestamp is local policy: it feeds
expiry and the cull weight (clean_message_store, LXMRouter.py:1144),
and no peer ever sees it. The file path and the two peer lists are pure
bookkeeping of that design.
Per peer, to_bytes (LXMPeer.py:138) persists twenty-odd fields.
Only four of them are visible on the wire in any form: the peer's
destination hash, its peering key, its announced limits, and its
announced costs. Everything else — link establishment rate, sync
transfer rate, rx/tx byte counters, offered/outgoing/incoming counts,
backoff state — is statistics. And two of them are the problem:
handled_ids and unhandled_ids, a pair of 32-byte-per-message sets
per peer, because every message the node accepts is enqueued into
every other peer's unhandled set
(flush_peer_distribution_queue, LXMRouter.py:2472).
That distinction decides what we may drop. We may drop all of the statistics and both peer sets. We may not drop the destination hash, the transient ID, or the bytes.
What it advertises, and whether a peer believes it
The propagation announce is a seven-element msgpack list
(get_propagation_node_app_data, LXMRouter.py:324):
| # | Field | Reference default |
|---|---|---|
| 0 | legacy PN support | False |
| 1 | node timebase | now |
| 2 | propagation node state | True |
| 3 | per-transfer limit, kilobytes | 256 (PROPAGATION_LIMIT, LXMRouter.py:55) |
| 4 | per-sync limit, kilobytes | 10240 (SYNC_LIMIT, LXMRouter.py:59) |
| 5 | [stamp cost, flexibility, peering cost] | [16, 3, 18] (PROPAGATION_COST, LXMRouter.py:54; PEERING_COST, LXMRouter.py:50) |
| 6 | metadata dict | name |
A node can honestly announce a small capacity, and peers respect
it. Field 3 is enforced by the offering peer: a message larger than
our advertised transfer limit is dropped from its queue for us and
marked handled, so it is never retried (sync, LXMPeer.py:267).
Field 4 is enforced by us: an inbound resource larger than the
advertised sync limit is refused before it transfers
(propagation_resource_advertised, LXMRouter.py:2206). Field 5 is
read by both, and a client mines its stamp to the cost we name.
Two caveats, and they matter.
- For a client, field 3 is advisory. Nothing on the client side checks a node's transfer limit before uploading; the only enforcement is our refusal of the resource, which the client reports as a failed sync rather than as "too big".
- We would be the first to advertise cheap. The reference clamps
its own configured cost up to
PROPAGATION_COST_MIN(LXMRouter.py:52, applied atLXMRouter.py:137), so a Python node never announces below 13. Announcing 0 is wire-legal and semantically honoured — a peer's accepted cost ismax(0, our_cost − flexibility)— but it is a policy nobody else in the mesh runs, and it hands away the only spam brake the protocol has.
The announce also has a switch: field 2 false makes every Python router
unpeer us on the next announce (LXMFPropagationAnnounceHandler,
Handlers.py:35). That is the clean way to leave the role, and it is
also the reason a node that serves only static peers is invisible to
clients: the reference computes field 2 as "propagation node and
not static-only".
What a peer expects when a node forgets. This is the crux
There is no verb for it. Say it plainly, because the design has to be built around the absence.
The error space is rich and none of it means "I dropped it":
ERROR_NO_IDENTITY, ERROR_NO_ACCESS, ERROR_INVALID_KEY,
ERROR_INVALID_DATA, ERROR_INVALID_STAMP, ERROR_THROTTLED
(LXMPeer.py:29), ERROR_NOT_FOUND (LXMPeer.py:30),
ERROR_TIMEOUT. ERROR_NOT_FOUND is defined and never returned by
either request handler.
What actually happens when a message is gone:
- On
/getlist, it is simply not in the returned list. The client cannot tell "never arrived" from "arrived and was dropped". - On
/getfetch, a wanted ID that is no longer in the store is skipped silently and the response is shorter than the request (message_get_request,LXMRouter.py:1482). No error, no gap marker. - On the peer side the same thing happens in reverse: the offering
peer discovers on its next sync that an ID it had queued is gone
from its own store and quietly drops it (
sync,LXMPeer.py:267).
And the reference already forgets, routinely and silently: messages
expire after 30 days (MESSAGE_EXPIRY, LXMRouter.py:38) and, when
the store exceeds its configured limit, entries are culled by a weight
of age × size until enough bytes are free (clean_message_store,
LXMRouter.py:1144). Peers vanish after 14 days unreachable
(MAX_UNREACHABLE, LXMPeer.py:39).
Two conclusions follow, and they point in opposite directions.
Forgetting is normal, so a small node forgets faster, not differently. There is no promise in the protocol that we would be breaking. Sideband's own retry behaviour already has to cope with a node that dropped something.
But acceptance is proven and retention is not. The node proves the
upload packet (packet.prove, LXMRouter.py:2255), so the sender is
told "accepted" and is never told "and then discarded". A store that
accepts more than it can plausibly hold converts a proof of acceptance
into a lie by omission. So our design must avoid promising: accept
less rather than accept and drop, and make the advertised limits small
enough that the acceptance is honest.
The one honest back-pressure verb that does exist is
ERROR_THROTTLED, and a peer handles it correctly by deferring its
next sync (LXMPeer.py:421). "Not now" is expressible. "Not ever" is
not.
2. The numbers
Where the store lives
One file decides it. leviculum-nrf/memory.x carves STORE out of the
top of the application window and exports __srecord_store /
__erecord_store; region (leviculum-nrf/src/record_store.rs:219)
reads those two symbols, and nothing else in the tree knows the
addresses.
| Address | Length | What |
|---|---|---|
0x00000 | 4 KiB | MBR |
0x01000 | 152 KiB | SoftDevice S140 v7.3.0 |
0x27000 | 0xB3000, 716 KiB | FLASH — the firmware image window |
0xDA000 | 0x10000, 64 KiB, 16 pages | STORE — the record log |
0xEA000 | 4 KiB | telemetry target / fixed position / media profile |
0xEB000 | 4 KiB | radio configuration |
0xEC000 | 4 KiB | identity |
0xED000 | 28 KiB | Heltec license/version data, T114 only |
0xF4000 | — | bootloader |
Source for every row: the map at the head of memory.x (FLASH,
leviculum-nrf/memory.x:78; STORE, leviculum-nrf/memory.x:91).
The gap between image end and store start, as
scripts/check-nrf-store-gap.sh reports it on every just fast — this
run, on the tree at 81fcb46e:
[store-gap] t114 image ends 0xa55a0, store 0xda000..0xea000 (16 pages), gap 215648 B (210 KiB)
[store-gap] rak4631 image ends 0xa6c08, store 0xda000..0xea000 (16 pages), gap 209912 B (204 KiB)
The gate measures the PT_LOAD segments the .uf2 is built from, not the
sections the linker charged to FLASH, and it reads the region's bounds
from the symbols the firmware itself mounts. An image that grew into the
region would be a link error before it could be a lost store: four
ASSERTs in memory.x hold the edges (ASSERT,
leviculum-nrf/memory.x:291) — the image stops below the boot-record
page (#380), that page stops below the store, the store stops at
USER_FLASH_END, and the store is a whole number of 4 KiB pages.
Why the region survives a UF2 update — and what is not yet proven.
The store sits inside the bootloader's writable window, so
USER_FLASH_END (0xEA000) does not protect it the way it protects the
three persistence pages above. What protects it is that the Adafruit
bootloader erases only the pages it writes: flash_nrf5x_write buffers
one page and flash_nrf5x_flush (upstream src/flash_nrf5x.c) erases
and programs exactly that page, and only when its content differs. Our
.uf2 carries blocks from 0x27000 to the end of the image and none
above it, so no page of the store is ever a block's target and no erase
reaches one (docs/src/concepts/lnode-flashing.md, §What a UF2 is
allowed to write).
That is an argument from the bootloader's source, and it is not yet a
board proof: nobody has written records to a board, flashed a new
.uf2 over it and remounted. Until that run exists, treat "the store
survives a firmware update" as expected rather than as established. It
is in the batch list in §4.
What our messages actually weigh
Measured, not assumed. Source: the two field logs from the 2026-09-09
walk, /home/lew/rig-run/feld-archiv/pocket-lauf10.log and
t114-lauf10.log, 12.69 h of wall clock each (10:52 to 23:33).
Method: for each [TELEMETRY] report target= line, take the packet the
node emitted within the next 14 log lines, excluding the two lengths
that are the nodes' own lxmf.delivery announces (181 and 183). 284
messages, which agrees with the 288 report lines to within the four
that straddle a log boundary.
| On-wire packet (Type 1) | Count |
|---|---|
| 227 B | 6 |
| 259 B | 80 |
| 275 B | 198 |
Median 275 B, worst case 275 B, minimum 227 B. A Type 1 header is 19 B
(HEADER_MINSIZE, leviculum-core/src/constants.rs:69), and the
propagation form re-prepends the 16-byte destination hash, so
lxmf_data is the packet length minus 3: median 272 B, range 224 to
272 B. With the 32-byte propagation stamp appended, the stored
object is 304 B median, 256 to 304 B over the run.
Cross-check against the relayed form: the same message crossing a hop
was logged as [LORA] TX split 291 bytes (254+37) with a Type 2
header, and 291 − 35 = 256 = 275 − 19. The two framings agree.
No text messages were present. The walk carried telemetry only, so
this distribution is a telemetry distribution and nothing else. A
Sideband text message is LXMF_OVERHEAD = 112 B
(LXMF_OVERHEAD, LXMessage.py:63) plus the RNS encryption overhead
plus the text, so a one-line message lands in the same 250 to 350 B
band; anything with an image or an audio field is one to two orders of
magnitude larger and is exactly what the advertised transfer limit
exists to refuse.
How many fit
From the record log as built, not from the costing this page first did.
The header is still the 42 bytes that costing tabulated; what changed
with the part is the region, the page header, and that every offset is a
multiple of a word because sd_flash_write takes a length in words.
| Field | Bytes |
|---|---|
| body length, u16 LE | 2 |
| key — the transient ID | 32 |
| timestamp, u32 LE | 4 |
| tag — the stamp value | 1 |
flags: 0xFF uncommitted, 0xFE live, 0xFC purged | 1 |
| CRC-16 over the header and the body | 2 |
header total (HEADER_LEN) | 42 |
| body | len |
| padding to a multiple of 4 | 0 to 3 |
(HEADER_LEN, leviculum-nrf/record-log/src/lib.rs:218; the layout is
tabulated at leviculum-nrf/record-log/src/lib.rs:108-119.) The
destination hash is not a field: it is the first 16 bytes of the body,
as it is in the reference, which reads it back from the head of its file
(LXMRouter.py:2498).
Each page carries a 12-byte header written once per erase
(SECTOR_HEADER_LEN, leviculum-nrf/record-log/src/lib.rs:220), which
leaves 4 084 B of the 4 096 for records (SECTOR_PAYLOAD,
leviculum-nrf/record-log/src/lib.rs:225). A record never straddles a
page.
At the measured median body of 304 B the stride is
align_up(42 + 304) = 348 B (record_stride,
leviculum-nrf/record-log/src/lib.rs:258), so 4 084 / 348 = 11 records
to a page with 256 B of tail (6.3 %).
| Body | Stride | Per page | In the 16-page region |
|---|---|---|---|
| 304 B — the field median | 348 B | 11 | 176 |
| 256 B — the smallest the walk produced | 300 B | 13 | 208 |
4 042 B — the largest the format allows (MAX_BODY, leviculum-nrf/record-log/src/lib.rs:227) | 4 084 B | 1 | 16 |
The 176 is the count at the brim. Reclaim is round-robin — when the active page cannot fit the next record the next page is erased and becomes active — so a store in steady state holds between 166 (just after a reclaim: fifteen full pages and one record) and 176.
Where the 64 KiB goes at that fill: 53 504 B of message body, 7 392 B of record headers, 352 B of record padding, 192 B of page headers and 4 096 B of per-page tail. 52 KiB of the 64 is message.
For scale in the other direction: one message at the reference's default
per-transfer limit of 256 kB (PROPAGATION_LIMIT, LXMRouter.py:55) is
63 times the largest body this format can hold at all. That is the
argument for announcing a small field 3, and it is now an argument about
a hard bound rather than a preference — see What we announce below.
Endurance
The budget got an order of magnitude worse with the part: 10 000 erase
cycles per page on the nRF52840, against 100 000 on the NOR parts this
page first costed (nRF52840 Product Specification, NVMC chapter, quoted
at leviculum-nrf/record-log/src/lib.rs:37).
Duty, measured: 284 messages from two moving trackers over 12.69 h = 22.4 messages/hour = 196 224 a year. At 11 field-sized records to a page, that is 17 838 page erases a year, and where they land is the whole design:
| Erases/year | Budget | Life | |
|---|---|---|---|
| Spread over the 16 pages | 1 115 per page | 10 000 | 9 years |
| Spread over 68 pages (the whole window, for scale) | 262 per page | 10 000 | 38 years |
| One fixed metadata page | 196 224 | 10 000 | 18.6 days |
(The table is the spike's, recomputed for the region as landed:
leviculum-nrf/record-log/src/lib.rs:55-59.)
Read the last row twice. The message data is not the endurance risk; a fixed metadata page is. A store that keeps its head pointer, its index or its sequence counter at a fixed address and rewrites it on every accepted record spends its entire budget in eighteen days at the duty we actually measured in the field. On the external part the same line read six months, which is long enough to sound survivable. It is why this format has no superblock, no index page and no head pointer: everything the log needs to mount itself is recovered by reading the page headers, and a page header is written exactly once per erase of the page it heads.
16 pages is a size choice, and it has a price. 68 pages — the rest
of the application window — would have bought 38 years and cost the
image 212 KiB of headroom it may want for BLE and LXMF; the gate above
says 210 KiB of gap is what remains at 16. Widening the region upward is
impossible, because 0xEA000 is the bootloader's USER_FLASH_END, and
widening it downward moves every record, so a later resize is a
reformat. That is the deliberate price of the smaller default, and
mount (leviculum-nrf/src/record_store.rs:518) already treats a
region that is not ours as unformatted rather than as corrupt, so the
reformat is a boot line and not an incident.
The 9 years is at the field walk's telemetry duty and scales with it: ten times that duty is 11 months, a hundred times is 33 days. Those are the numbers to re-run when a real message mix exists rather than a telemetry one.
Writing next to the radio
With the SoftDevice enabled the NVMC is Restricted: only
sd_flash_write / sd_flash_page_erase may touch this flash (S140 SDS,
Hardware peripherals), which is also where a word becomes the only
program unit and two writes per word between erases the only budget
(leviculum-nrf/record-log/src/lib.rs:37-41). A page erase is 85 ms, a
word write 41 µs (same source).
Per median record the log does three program runs — the header up to the commit word, then the CRC and the body from the far side of it, then the commit word itself — 87 words in all, 3.6 ms of NVMC time, and one 85 ms erase every eleventh record.
But NVMC time is not the cost that matters here. The SoftDevice
schedules flash work between radio events and fails the operation
outright when it finds no gap (S140 SDS, Flash API timing), so a refusal
is a statement about the next few milliseconds of radio traffic and not
about the part. The store answers it with four attempts and a doubling
delay — 50, 100, 200 ms (FLASH_ATTEMPTS,
leviculum-nrf/src/record_store.rs:206) — and prints every refusal and
a running count on its debug port.
A refusal part way through an append seals the page: the rest of it
is given up, because programming over bytes already down would have to
raise bits. The spike's sweep puts a number on how often that is the
outcome — of the twenty places a refused operation can land inside one
append, three leave the page usable and seventeen give up the rest of it
(sealed, leviculum-nrf/store-spike/tests/record_log.rs:360). What
sealing costs in wear is the endurance arithmetic with fewer records to
a page: a board that sealed on every append would erase a page per
message, 196 224 erases a year over 16 pages, and spend the nine years
in 10 months. That is the upper bound, not an expectation; it is
also the reason the refusal counter is on the debug port rather than
silent.
What the erase storm does to BLE and LoRa is to be measured on the
rig, and this page will not predict it. The instrument is already in
the firmware: STORE_STORM (TYPE_STORE_STORM,
leviculum-core/src/envelope.rs:264) appends N synthetic records of a
given size, bounded at 1 000 records of 1 024 B
(STORE_STORM_MAX_BYTES, leviculum-core/src/envelope.rs:286) and
tagged so a later batch can purge exactly those (TAG_BENCH,
leviculum-nrf/src/record_store.rs:87); lnflash --store-storm COUNT[,BYTES] sends it (--store-storm, lnflash/src/main.rs:410).
The numbers owed are a connected phone's throughput and a LoRa link's
delivery rate across a storm, measured against the same run without one.
Scan, and what it costs in RAM
Reads do not go through the SoftDevice's flash scheduler at all. The
internal flash is memory-mapped and the read is a memcpy that cannot
fail or be refused (read, leviculum-nrf/src/record_store.rs:335) —
unlike a write or an erase, it never waits for a gap between radio
events. That single fact removes the RAM index the QSPI costing needed:
a lookup is a scan, and a scan is free of the radio.
The RAM it would have competed with, measured on the Pocket at the end of the 12.69 h field run:
[HEAP] used=58612 free=39692 watermark=58996 size=98304
96 KiB of heap, 58 996 B at the high-water mark, so 39 308 B of free heap in the worst observed moment. Against a 176-message region:
| Bytes | Against 39 308 B free | |
|---|---|---|
| A reference-shaped full index: 32 B key + 16 B destination + 4 B offset + 2 B size + 4 B timestamp + 1 B stamp = 59 B an entry | 10 384 | fits |
The reference's per-message peer sets, 20 peers × 2 sets × 16 B a hash (MAX_PEERS, LXMRouter.py:43) | 112 640 | does not fit |
So the arithmetic that killed the RAM index on a 2 MB part no longer kills it on 64 KiB: a full index of this region would fit in a quarter of the free heap. It is still not worth having — the heap has other claimants and the scan that replaces it is cheap — but the honest statement is "unnecessary", not "impossible". What remains impossible is the second row: the peer sets scale with peers, and we control neither how many peer with us nor, therefore, that number.
What a full scan costs. for_each
(leviculum-nrf/record-log/src/lib.rs:541) walks every page, reads each
record's 42-byte header and then its body, because probe_record
(leviculum-nrf/record-log/src/lib.rs:901) checks the CRC over both. A
full region is therefore one pass over at most 64 KiB.
This is arithmetic, not a measurement, and the assumption is stated:
the read is a memcpy from mapped flash, so the work is the bit-serial
CRC-16 at eight shift-and-test steps a byte (crc16_update,
leviculum-nrf/record-log/src/lib.rs:270) — 524 288 steps for the whole
region, and at one to four cycles a step on the Cortex-M4 at 64 MHz that
is 8 to 33 ms. The measured number is owed and nearly free, because
the mount already performs exactly this scan and reports what it found
(count, leviculum-nrf/record-log/src/lib.rs:406).
Either way the conclusion is the same and it is not close: a /get list
request costing tens of milliseconds of CPU, and nothing of the radio
scheduler, is not a design constraint. The on-flash directory is the
design.
Draining it: LoRa and BLE
LoRa. Measured, from the field log: a telemetry message crossing a
hop is 291 bytes on the wire, split into 254 + 37 byte frames, and the
firmware reported op=tx duration_ms=903..905 for 13 of the 15 such
transmissions in the run (SF8, BW 125 kHz, CR 4:5, 18-symbol preamble).
Our airtime model agrees: computed 543 ms for the 184-byte frame the
same log reports as airtime_ms=544 (airtime_ms_with_preamble,
leviculum-core/src/rnode.rs:1061; preamble from
derive_preamble_symbols, leviculum-core/src/rnode.rs:917, which
floors at 18, LORA_PREAMBLE_SYMBOLS_MIN,
leviculum-core/src/rnode.rs:852).
At 904 ms per message and the 10 % duty-cycle cap the firmware enforces
([LORA_AIRTIME_LOCK] lt=1000 lt_cap=10.00% in the same run), the full
region is 176 × 904 ms = 2.7 minutes of pure airtime, 27 minutes of
wall clock — with zero retransmissions, zero link setup and no other
traffic on the channel.
That reverses a conclusion the QSPI costing drew. A full 2 MB store was 13.9 h of wall clock at the duty cap, which is a museum; a 64 KiB store is drainable in half an hour. In a walk-past it is still only the delta between two nodes that moves, but the whole store is no longer out of reach, and that makes the single-message upload path (§3) a usable way for two boards to meet rather than a consolation prize.
BLE. The negotiated MTU is bounded by measurement rather than
assumed: the SoftDevice is configured with an ATT MTU ceiling of 256
(CONN_GATT, leviculum-nrf/src/ble/mod.rs:795), but the field log
shows a 183-byte packet fragmenting into 2 and a 275-byte packet also
into 2, which brackets the payload per fragment to 138 to 182 bytes and
the MTU to 146 to 190 — consistent with the 185 default
(DEFAULT_MTU, leviculum-core/src/framing/ble.rs:89;
payload_per_fragment, leviculum-core/src/framing/ble.rs:107). At
177 bytes per fragment a 304-byte message is 2 notifications, so the
full region is 352 notifications.
The sustained notification rate is still not measured and this page will not invent it. The field run carried sparse traffic — the tightest observed spacing is two packets in the same millisecond, which is a burst, not a rate. The shape is all that can be said: at 10 notifications/s the full region is 35 seconds. BLE is not the binding constraint on a store this size, and the measurement is owed rather than critical. It is named in §4.
What accepting a message costs in CPU
A propagation stamp is validated by expanding a 1 000-round workblock
from the transient ID and hashing it with the stamp
(WORKBLOCK_EXPAND_ROUNDS_PN, LXStamper.py:13; stamp_workblock,
LXStamper.py:49; validate_pn_stamp, LXStamper.py:84). Each round
is one SHA-256 over the salt input plus one HKDF-SHA256 producing 256
bytes: one extract HMAC and eight expand HMACs, four SHA-256
compressions each. 37 compressions per round, 37 000 for the
workblock, plus 4 000 for the final digest over the 250 KiB workblock:
41 000 SHA-256 compressions, 2.62 MB hashed, per message.
Two things follow.
The reference materialises the 250 KiB workblock in RAM. We do not
have to, and already do not. workblock_hasher
(leviculum-lxmf/src/stamp.rs:310) streams the HKDF blocks straight
into the digest and keeps one 256-byte block. The RAM objection to
stamp validation is already solved in our tree; only the CPU cost
remains.
At an advertised cost of 0 the cost is not incurred at all. Our
validator short-circuits before the workblock when the cost is zero
(validate_stamp, leviculum-lxmf/src/stamp.rs:373). The firmware ran
this way for delivery stamps until the role landed, with the LXMF
dependency pulled without pow; the role brought pow in, because the
board announces stamp cost 13 and has to validate what it accepts, through
the streaming validator above (leviculum-nrf/Cargo.toml:25-33).
The expensive case is peering out to a Python node, which requires
mining a key at that node's advertised peering cost, default 18, over a
25-round workblock (WORKBLOCK_EXPAND_ROUNDS_PEERING,
LXStamper.py:14; generate_peering_key, LXMPeer.py:242). With the
precomputed-digest-state trick our miner already uses, that is 925
compressions for the workblock plus about two per trial over 2^18
expected trials: 525 000 compressions, 33.6 MB hashed, once per peer,
and the result is persistable. The reference's own miner rehashes the
6.4 KB workblock every trial and so hashes 1.7 GB for the same key;
this is a legitimate deviation under the deviation rule, since the
stamp produced is byte-identical.
Converting compressions to seconds needs a SHA-256 throughput on the nRF52840 at 64 MHz that we have not measured. For orientation only, at 20 / 40 / 60 cycles per byte the stamp validation is 0.8 / 1.6 / 2.5 s and the peering key is 10 / 21 / 32 s. The measurement is owed (§4); the conclusion that survives any plausible value is that per-message stamp validation at a nonzero cost is seconds of the only core we have, and a peering key is a one-off we can afford.
What we announce
Field 3 of the propagation announce, the per-transfer limit, is parsed
with int() (propagation_transfer_limit,
reference/LXMF/LXMF/Handlers.py:61), so the only values that exist on
the wire are whole kilobytes. The offering peer enforces it against
lxm_size + 16 and reads a kilobyte as 1 000 bytes
(propagation_transfer_limit, reference/LXMF/LXMF/LXMPeer.py:370),
where lxm_size is the stored object — lxmf_data with the stamp
appended, which is exactly what our record body holds
(propagation_entries, LXMRouter.py:2518).
Our hard bound is one page: a record never straddles one, so a body
above MAX_BODY = 4 042 B cannot be stored at all. Against the
reference's arithmetic that bounds field 3:
| Announced field 3 | Largest lxm_size a peer will offer | Fits a page? |
|---|---|---|
| 3 | 2 984 B | yes, 1 058 B spare |
| 4 | 3 984 B | yes, 58 B spare |
| 5 | 4 984 B | no — 942 B over |
Announce 4. It is the largest whole kilobyte whose worst case still fits the page a record may not straddle, and it is thirteen times the measured field median. Announcing the reference's 256 would be the failure §1 names: a proof of acceptance the store cannot honour.
Field 4, the per-sync limit, bounds one resource rather than one message, and its bound is the region. 176 messages is 53 504 B of body, so a sync allowed to carry more than that laps the log inside a single transfer and overwrites its own earlier records. Announce 32 — about a hundred median messages, well under a lap, and about three times what five minutes of a LoRa walk-past can carry at the duty cap.
Field 5, the stamp cost, stays open: it is the one field whose right value depends on a measurement we do not have (SHA-256 throughput on the board, above), and both the cost and the reason for it belong in §4.
3. The options
Four, and the fourth is doing nothing on the board.
| A. Full propagation node | B. Bounded node, honest limits | C. Courier for recently-seen peers | D. Nothing on the board; lnsd carries it | |
|---|---|---|---|---|
| What Sideband sees | a normal propagation node | a normal propagation node with small limits | nothing; not a PN | the PC's node, if in range |
Announces lxmf.propagation | yes | yes | no | n/a |
| RAM at capacity | 10 KiB index + 113 KiB of peer sets at 20 peers | an on-flash directory, scanned; no index | same as B | 0 |
| Peers | autopeer, up to 20 | autopeer, capped low | none | as configured |
| Stamp cost advertised | 16 | 0, or a low nonzero once measured | n/a | 16 |
| Two boards meet, no phone | works, if both can peer | works | works, but only between our own boards | does not work |
| Board switched off mid-transfer | client retries; nothing lost | client retries; nothing lost | our own protocol, our own problem | n/a |
| Verdict | impossible | viable | not compatible | insufficient |
A, the full node, is still out, but the smaller region moved which
argument does it. On the 2 MB costing the RAM index alone was eight
times the whole heap; on 176 messages it is 10 KiB against 39 KiB free,
so message count no longer decides anything. Peer count does. The
handled/unhandled sets are held per message as lists of peer hashes
(propagation_entries, LXMRouter.py:2518), which is 112 640 B at
176 messages and the reference's 20 peers, and there is no cap we
control on who peers with us: any Python router within four hops that
hears our announce peers automatically (AUTOPEER_MAXDEPTH,
LXMRouter.py:45; LXMFPropagationAnnounceHandler, Handlers.py:35).
A number we do not control is not a budget.
The count that is untouched by the region shrinking is the one that
actually kills A: every message we accept is enqueued for every peer
(flush_peer_distribution_queue, LXMRouter.py:2472), so a store
filled from a phone over BLE in seconds would be re-offered over a
10 %-duty LoRa link to everyone in range, at 904 ms a message. That is
not a tuning problem.
B, the bounded node, is the only option that satisfies the framing.
Everything it needs is already expressible in the announce: a small
field 3 and field 4 that peers and clients honour, a stamp cost we
choose, and a max_peers of our own. It costs the spam brake — a node
advertising cost 0 can be filled by anyone — which the small transfer
limit and the size cull bound but do not remove. It is the only option
where a Sideband user gets the thing they expect without knowing the
node is small.
C, the courier, is out on compatibility, not on cost. Holding
messages only for destinations we have recently seen is a good policy
and would fit the RAM budget comfortably. But there is no verb for it:
a node that does not announce lxmf.propagation is invisible to
Sideband, and a node that announces it and then behaves as a courier is
lying about field 2. C is a policy inside B, not an alternative to
it — and as a policy inside B it is exactly the right one for the
bounded case.
D is what we do today and it is insufficient for the case that prompted the issue. Two boards on a hill with no PC in range have no store between them.
The case with no phone present
Explicitly, because it is the operator's case. Under B, two boards that meet with no phone can exchange messages by two paths, and the cheap one is worth naming:
- Full peering. Both announce as propagation nodes, autopeer within four hops, mine a peering key at each other's cost (which, since we choose our own, can be low between our own boards), and sync over a Link and a Resource. Correct, and bounded by the 10 % duty cycle: the delta, not the store.
- Single-message client upload.
propagation_packet(LXMRouter.py:2234) accepts one message per link packet with no peering key at all. Two boards can hand each other one message at a time with no peering, no Resource, and no mining. For a walk-past on LoRa, where 904 ms of airtime per message is the real budget, this is the path that matches the medium.
Under A the same is true but the store re-offer makes it unusable. Under C it works only between our own boards. Under D it does not work.
Switched off mid-transfer
The protocol is safe against our disappearance at every point, and the reason is worth recording because it constrains our implementation:
- Mid-upload, the client's packet is proven only after the message
is stored (
packet.prove,LXMRouter.py:2255). If we die first, the client gets no proof and retries. This makes "persist before you prove" a rule, not a preference — a proof written before the record is durable converts a power cut into a lost message. - Mid-
/get, the node deletes only on the client's explicithavespurge, and the client sends that purge only after local delivery (message_get_response,LXMRouter.py:1607). If we die during the transfer, nothing is deleted and the client repeats the exchange. - Mid-sync with a peer, the offering peer marks nothing handled
until the transfer concludes, and a failed request tears the link
down and backs off (
request_failed,LXMPeer.py:395).
The one thing that is not safe is a store whose own recovery is
unsound. A power cut in the middle of an append must leave a store that
reopens with every completed record and no partial one, and the log as
built delivers that with a one-word commit rather than with a
probability: a record counts as present only if its flags byte reads
live or purged, that byte sits in a single word programmed last, and a
word cannot be half-written (FLAG_LIVE,
leviculum-nrf/record-log/src/lib.rs:238). Any cut before that word
leaves the record invisible, deterministically; the CRC is then
catching a dropped bit rather than standing in for a commit protocol.
4. The recommendation
Option B, and the store it needs now exists. The region is 176
field-sized messages with 9 years of page-erase budget at the measured
field duty, provided nothing is ever written to a fixed page — a fixed
metadata page dies in 18.6 days, and the format has none. Reads never
touch the SoftDevice's flash scheduler, so a lookup is a scan of tens of
milliseconds and there is no RAM index to fit. A full region drains over
LoRa in half an hour of wall clock at the duty cap, which is the
difference between a store and a museum. None of those conclusions is
about LXMF; all of them are about the store, which is why the store came
first and is why it is worth having whether or not the propagation node
ever exists — lnmsg's mailbox and telemetry retention want the same
16 pages.
What stands between here and the role is not arithmetic. It is three things nobody has measured on a board: what an erase storm does to BLE and LoRa while it runs, whether the region really survives a UF2, and the two rates that decide the stamp cost and the phone drain.
The sequence, as it now stands
- Store region and mount — done. The log format was chosen against
the internal flash in 59c36129 and given its region, its linker
symbols and its boot-time mount in 81fcb46e. Nothing stores messages
in it: no LXMF, nothing announced. The gap gate prints the remaining
headroom for both bins on every
just fast. - The erase storm under BLE and LoRa load, on the rig — owed, and it
is the next batch. The instrument is in the firmware already
(
STORE_STORM, above). Acceptance: a phone connected over BLE and a LoRa link under traffic, each run twice — once with a storm of field-sized records and once without — reported as throughput and delivery rate with the event volumes on both sides, not as pass or fail. A storm that costs the radio nothing measurable and a storm that costs it everything are both results; a run that cannot tell them apart is not. - UF2 survival, on a board — owed, and cheap. Write records, flash
a
.uf2built from a different commit, remount, and assert the same record count and a byte-exact digest of the region. The bootloader source says it must survive (§2); this is the run that makes it established rather than expected. It belongs with the next firmware flash on the rig, not in a batch of its own. - Then the role, on top of a store that has been measured. The
accept path with the limits this page recommends (field 3 = 4, field
4 = 32),
/getlist and fetch answered from a scan rather than an index,/offerwith a peer cap of our own, and no per-peer sets anywhere. Acceptance: a Sideband client and a Pythonrnsdboth use the board as their propagation node without knowing it is small, and a power cut during an upload leaves a store that reopens with every completed record and no partial one.
Steps 2 and 3 are measurements, step 4 is the feature, and the order is not negotiable: a propagation node that lands before the storm is measured is a node whose failure mode is a radio that stutters when somebody sends a message.
The open questions, none of them closed by this page
- Whether the region survives a UF2 on a board. Argued from the bootloader's source, not yet run. Step 3 above.
- What an erase storm costs BLE and LoRa. The one number that could still make a propagation node on the board a bad idea. Step 2 above.
- SHA-256 throughput on the nRF52840 at 64 MHz. Decides whether we can advertise a nonzero stamp cost and keep the only spam brake the protocol has. At 20 / 40 / 60 cycles a byte, validating one stamp is 0.8 / 1.6 / 2.5 s of the only core we have; the spread is too wide to decide on.
- The sustained BLE notification rate. Decides whether a phone can drain 352 notifications in a usable time. Expected to be comfortable, unmeasured.
- The message-size distribution beyond telemetry. Every body figure on this page comes from 284 telemetry messages. A run carrying real Sideband text, and a phone that sends an image, will move the median and will show how often field 3 actually bites.
- Whether autopeering can be bounded in practice. The reference
peers with anyone within four hops (
AUTOPEER_MAXDEPTH,LXMRouter.py:45). B assumes a cap we enforce ourselves keeps the peer-set arithmetic survivable; three Python routers in range would show whether it does. - Whether 16 pages is the right size. 68 would buy 38 years and cost
the image 212 KiB of headroom. The number lives in
memory.xand the trade is argued there; changing it later is a reformat, whichmounthandles as an unformatted region.
What this page corrected, and what it still owes
Codeberg #384 observed that a search of leviculum-nrf finds no QSPI.
It was right, and for a reason neither board file admitted at the time:
a vendor variant header's EXTERNAL_FLASH_DEVICES line was read as a
statement that a part is fitted, when on both vendors it is a template
default under a comment denying one. Neither board answered a JEDEC read
on any unit we own; both sets of pin aliases are gone and the reasons
are in the board files (CONFIG,
leviculum-nrf/src/boards/t114.rs:171; CONFIG,
leviculum-nrf/src/boards/rak4631.rs:186).
The larger correction was this page's own premise. It costed a store on two parts that do not exist, and the recommendation rested on figures that were an order of magnitude too generous in capacity and an order of magnitude too generous in erase budget. The protocol half needed no change, which is the useful lesson: the analysis that was about LXMF survived the part being wrong, and everything that was about a datasheet did not.
What it still owes is a board. Three of the numbers above are arithmetic or datasheet figures — the scan time, the erase storm's cost, the UF2 survival — and a rig run replaces each of them with a measurement.
5. Peering: the design part 2 built
Peering is the core of the role — a node that does not peer is a
mailbox, not a mesh (Lead decision, 2026-09-11). This section was the
binding design for part 2 of leviculum#384 and is now updated to what
part 2 built: first what the reference actually keeps and exchanges,
measured against the pinned tree (795fdaa), then our design inside this
page's constraints, with every number re-derived from the code as
landed. The protocol half is leviculum-lxmf/src/peering.rs
(no_std + alloc, behind the PeerStore trait the board implements in
part 3); the host glue is lnpnd/src/peering.rs.
What the reference keeps, per peer and per message
LXMPeer.to_bytes (LXMPeer.py:138-175) persists, per peer: the
destination hash, the peering key and its value, the peering timebase,
alive flag, last-heard, sync strategy, metadata, the announced transfer
and sync limits, the announced stamp cost, flexibility and peering
cost, the last sync attempt, and six statistics counters (link
establishment rate, sync transfer rate, offered / outgoing / incoming,
rx/tx bytes) — plus the two sets the §1 analysis flagged:
handled_ids and unhandled_ids. Those two are not stored on the
peer at all at runtime: they live per message, as lists of peer
hashes in propagation_entries[4] and [5]
(LXMRouter.py:2518; membership filtered per peer in
LXMPeer.handled_messages, LXMPeer.py:574-588), 16 bytes per peer
per message, filled by flush_peer_distribution_queue
(LXMRouter.py:2472) which enqueues every accepted message for every
peer. At this store's 176 messages and the reference's 20-peer default
that is the 112 640 B that §2 measured against 39 308 B of free heap:
the one reference structure we cannot carry.
A sync round exchanges three things (LXMPeer.sync,
LXMPeer.py:267-390): a /offer request carrying
[peering_key, [transient_id, …]] with the ids filtered by the peer's
minimum stamp value and packed under its announced limits
(:334-385); the response True / False / wanted-sublist
(offer_request, LXMRouter.py:2266-2329, which answers out of its
own propagation_entries membership); then one Reticulum Resource
whose body is msgpack([timestamp, [lxmf_data ‖ stamp, …]])
(:457-468). Only on the concluded transfer are the sent ids moved
handled (resource_concluded, LXMPeer.py:492-517); ids the peer
declined were moved handled already at the response (offer_response,
LXMPeer.py:443-448). An id purged from the store before its offer is
silently dropped at the next sync (:348-352) — forgetting needs no
verb between peers either.
Our peer record, and the cap
Per peer we keep what is wire-visible plus the minimum liveness state,
and nothing statistical. As built (Peer / PeerRecord,
leviculum-lxmf/src/peering.rs), the packed persistable record weighs:
| Field | Bytes |
|---|---|
| destination hash | 16 |
identity hash (the peering-key material's first half, LXMPeer.py:258) | 16 |
| peering key + value | 34 |
| announced limits (transfer, sync) | 8 |
| announced costs (stamp, flexibility, peering) | 3 |
| peering timebase, last heard | 12 |
| cursor into the store sequence | 8 |
| static flag | 1 |
| per peer | 98, call it 104 aligned |
The design's 80 grew to ~104 in the build: the identity hash joined
the record (mining material must survive a restart or the key is
useless), and the cursor widened to the u64 the host store's sequence
uses (the board packs the same pair into 6 bytes, below). Re-derived
against the same budget: during one sync (one at a time on the board,
lnpnd/src/peering.rs holds one round in flight) the offer list is
bounded at 6 144 B (OFFER_BYTES_LIMIT) — 34 B per encoded id, so
at most 179 ids per round, re-offering the rest next round. Against
§2's worst observed free heap of 39 308 B, the same 8 KiB peering
slice gives 104·N + 6 144 ≤ 8 192, N ≤ 19. That 6 144 B is a RAM
ceiling, never the budget on its own: an offer is sized by the link
it is handed to, offer_budget_for_mdu(link.mdu()), because
NodeCore::send_request refuses a body larger than the link MDU and
that refusal is local, silent and permanent (the store only grows, so
the next round is refused the same way). Over LoRa the MDU is 431 B
and one request names ten ids; a link with a larger negotiated MTU
uses more, up to the ceiling. Both engines therefore plan the offer
when the link is up, not when the round is scheduled — until then
there is no MDU to size against, and the plan made at scheduling time
answers only "is anything above the cursor offerable at all". Board cap: 16 peers
still holds, now with less margin; host config default: 20, the
reference's own MAX_PEERS (LXMRouter.py:43), settable as
max_peers. The full-table policy is deterministic and documented on
DeclineReason::TableFull: first heard wins, a full table declines
new candidates (the reference's own behaviour, LXMRouter.py:2032),
and slots free only by the 14-day unreachability cull
(MAX_UNREACHABLE, LXMPeer.py:39), an unpeer, or the peer leaving
the role.
The host persists its table in one msgpack file
(FilePeerStore, leviculum-std/src/file_peer_store.rs), as the
reference does (LXMRouter.py:599-631). The board's PeerStore
implementation is part 3's: the trait demands only upsert-by-key
(append a new tagged record, purge the old), full-scan load, and
nothing rewritten in place — no fixed page, which is §2's endurance
rule. Mined peering keys ride the same record, so the grind happens
once per peer, not per reboot.
The cursor, instead of per-peer sets
The record log is append-ordered: pages carry a monotone sequence
written once per erase (SECTOR_HEADER_LEN header,
leviculum-nrf/record-log/src/lib.rs:220), records within a page are
ordered by offset. A store position is therefore the pair
(page_sequence: u32, offset: u16), and each peer holds one cursor:
everything at or below it has been offered and concluded. As built,
the store trait carries this as StoredMessage::sequence: u64
(leviculum-lxmf/src/propagation_store.rs): the board maps
page_sequence << 16 | offset into it, the host store assigns a
monotone append counter persisted in its file names
(leviculum-std/src/file_propagation_store.rs), so cursors survive a
host restart too. A sync offers every live id newer than the cursor
(one for_each scan, §2 prices it at 8-33 ms; build_offer,
leviculum-lxmf/src/peering.rs); on the concluded transfer — or on a
"want none" response — the cursor advances to the plan's target
(resource_concluded is the reference's own only-on-conclusion rule,
LXMPeer.py:492-517). That replaces both per-peer sets with one
integer per peer, and it cannot lose messages: a message is either at
or below a concluded cursor (offered once), evicted (absent
everywhere, the reference's own behaviour at :348-352), or ahead of
the cursor (offered next round).
Three cursor semantics the build pinned down, tested in
leviculum-lxmf/src/peering_tests.rs:
- Permanent skips advance the cursor. An entry whose stamp value
is below the peer's minimum (
LXMPeer.py:340) or whose size exceeds the peer's per-message limit (:370-373) is stepped past for good — exactly the ids the reference marks handled without sending. - Resumable stops do not. The peer's per-sync limit and the offer's byte budget (the link's MDU, under the 6 144 B ceiling) end the round without advancing past what they excluded; the walk is in append order, so nothing above the target was withheld for a resumable reason. (The reference offers weight-sorted and keeps scanning past a sync-limit hit; ours stops there — a selection-order deviation with no wire effect, and the property that lets a single integer replace the sets.)
- A stale cursor is a bounded full re-offer. A cursor naming a
reclaimed page (board) or a reset store (host) orders below
everything live, so the next round re-offers everything — one link's
worth of ids per round, the rest in the rounds after it — and the
peer answers "want none" for what it holds. The
conformance cells drive this path explicitly (
lxmf_pn_reoffer).
What a Python peer observes: offers that may include ids it already
holds — including messages it itself sent us, since a cursor cannot
encode the reference's from_peer exclusion
(flush_peer_distribution_queue, LXMRouter.py:2484). That is
wire-legal and self-limiting: offer_request answers out of its own
store membership (:2317-2318) and declines them, and the cost is
offer-list bytes, not message bodies. The round-robin page reclaim
also means a cursor's page can be erased and reused while the cursor
still names the old sequence; page sequences are monotone, so a
cursor pointing into a reclaimed page simply reads as "older than
everything live" and the next offer is a full offer — the reboot case
again, bounded the same way.
One accept-path consequence, decided in the design and built as
decided: part 1 stored stamp value 0 for messages accepted at cost 0
(the validator short-circuits, where the reference computes the true
value even at cost 0, LXStamper.py:95). The offering side drops ids
whose stored value is below the peer's minimum (LXMPeer.py:340),
and a default Python peer's minimum is 16 − 3 = 13, so a store full of
value-0 records would offer that peer nothing. Part 2 therefore
computes the true stamp value at accept time whenever any known peer
requires more than 0 (PropagationNode::set_compute_stamp_value,
driven from the peer table; the measuring validator is
CooperativeStamper::measure_stamp) — 41 000 SHA-256 compressions per
message, §2's orientation says 0.8-2.5 s on the board, free on the
host — and keeps the shortcut otherwise. The record tag is written
once at append, so the decision is per-message at accept, not
retrofittable; a store accepted cheap stays cheap until it turns over
(at most 30 days). Note the practical consequence the chain cell ran
into: a true value of a free stamp is small (geometric, expected ~1
bit), so computing it honestly does not make a cost-0 store
propagatable through a default stock node — a node that wants its
store to travel through default peers must announce a stamp cost whose
minimum clears theirs (the cells use 16).
Peering with a Python node: the price of its key
Outbound peering requires mining a key at the peer's announced peering
cost over the 25-round peering workblock
(WORKBLOCK_EXPAND_ROUNDS_PEERING, LXStamper.py:14;
generate_peering_key, LXMPeer.py:242-265). The reference announces
18 by default and accepts configuration up to 26 (PEERING_COST,
MAX_PEERING_COST, LXMRouter.py:50-51). With our precomputed-
digest-state miner (§2): 925 compressions for the workblock plus ~2
per trial, expected 2^cost trials —
| Peer's cost | Compressions | On the board (at §2's 20/40/60 cycles/byte orientation) |
|---|---|---|
| 18 | ~5.3 × 10^5 | 10 s / 21 s / 32 s |
| 26 | ~1.3 × 10^8 | 45 min / 89 min / 134 min |
One-off per peer and persistable — a cost-26 Python neighbour would
otherwise cost the better part of an hour of the board's single core
per reboot. Part 2 therefore persists the mined key inside the peer
record itself (the PeerStore boundary above): on the host that is
the peer file, on the board a ~104 B tagged record — append-only, no
fixed page, a negligible tenant of the region. The host mines on a
worker thread, as the reference does (LXMPeer.py:285-286), never
under the core lock; costs above remote_peering_cost_max (default
26, MAX_PEERING_COST, LXMRouter.py:51) are refused at the table,
so the grind is bounded by configuration. The SHA-256 throughput
figure that pins this table's real column is §4's owed measurement,
still owed here.
Our own announced peering cost defaults to 0, the same policy as the
stamp cost and this time without even a reference counter-argument:
the PROPAGATION_COST_MIN clamp applies to the propagation cost only
(LXMRouter.py:137); the peering cost is passed through unclamped, so
0 is a value the reference itself can be configured to and validates
trivially (validate_peering_key with target 0 accepts any key,
LXStamper.py:73-82).
Two falsy-zero quirks of announcing cost 0, both observed against the reference and both ours to route around:
- Client side (part 1's interop run):
get_outbound_propagation_costtreats 0 as falsy (LXMRouter.py:429), re-requests the path, logs "stamp cost still unavailable" — and then proceeds correctly, mining a free stamp and uploading. Cost 0 is honoured on the wire; the reference client just grumbles first. - Peer side, and this one is a dead end:
LXMPeer.peering_key_readyshort-circuits false on a falsy peering cost (LXMPeer.py:228), so a stock node's sync toward a cost-0 peer postpones forever on "peering key has not been generated yet" — the key IS generated, the readiness check just never looks at it. A stock lxmd can never sync toward a node announcing peering cost 0. Our own outbound side deviates (any key is ready at cost 0,Peer::peering_key_ready,leviculum-lxmf/src/peering.rs— wire format untouched, the validator side accepts any key at target 0,LXStamper.py:79-82), so rust-to-rust peering at 0 works; a node that wants stock peers to sync to it announces at least 1, which is what the conformance chain cells do and why. Upstream is not told (standing policy); the workaround is a one-bit cost.
Evidence and Honesty in Testing
A mesh stack fails in ways that are easy to explain away: radios, timers, schedulers, six layers of asynchrony. The only defence is a set of rules about what counts as evidence and what counts as closure. These rules are engineering discipline, not process — they bind anyone contributing a fix, a test, or a measurement.
A check you have never seen fail is not a check
A green result is evidence only if the check could have been red. Two defects in this codebase's history make the point:
- The sx1262 RX-extend guard that could never fire. The guard
against truncating an in-flight slow-SF frame tested the
PreambleDetectedandHeaderValidIRQ flags — but the IRQ latch mask passed toSetDioIrqParamshad disabled exactly those bits, so the condition was unsatisfiable for the guard's whole life (#144, fixed in 26ce3a0). It survived because the firmware crate cross-compiles and had no host test target: nothing had ever run the guard at all, let alone watched it fire. The fix moved the mask and the extend decision intoleviculum-core/src/sx126x.rsas pure functions precisely so a host test could make them fail. - The benchmark step that always passed. Periculum's
execute_benchmarkused to end in an unconditionalOk(()): it drove probe loops, printed a throughput table, and returned green whatever the table said. Six of ten recorded runs carried zero packets and reported GREEN. From the outside, a step that measures and asserts nothing is indistinguishable from a step that measures and passes (periculum/src/assertions.rs).
The rule that follows: when you add a check — a guard, an assertion, a test — make it fail once, on the real failing condition, before you believe its pass. A test that has only ever been green proves only that it compiles.
A diagnostic indicator is trusted only after both states
Before reasoning from any indicator — a log line, a counter, a status field — verify two things:
- It measures the production path. Check in the code that the indicator reads the same lookup, the same state, the same branch the production behaviour depends on — not a parallel reimplementation that can drift.
- You have observed it in both states. An indicator you have only seen in one state might be stuck there. Drive the condition both ways and watch it follow.
And a diagnostic must not disturb what it measures. The canonical in-tree example is the airtime meter that a radio restart zeroes — taking the reading destroyed the reading (see Regulatory Airtime).
Symptom pairs are not mechanisms
"X happens and Y happens" is a correlation, not a diagnosis. A named mechanism has a causal chain that ends at a file and line, with the failing condition reproduced — you can point at the code and say "this branch, under this input, produces this observation, and here is the run where it did." Until then you have a hypothesis, and hypotheses get tested, not implemented: write the test that would confirm or refute it, and if it is refuted, move to the next one. Do not ship a fix for a mechanism that measurement has not shown to be the actual cause.
Minimal reproduction before the fix
When a bug's mechanism is not obvious on sight, write a minimal reproducing test first — one that reproduces exactly the failure mode and nothing more — and only then the fix. Writing the fix first lets you stop at "seems to work". "I already know the fix" is not a reason to skip; that is precisely where a test prevents wishful thinking. (Trivial fixes — a typo, an obvious null dereference — are exempt, and an existing test reliably made red by the bug counts as the minimal test.)
A green minimal test alone is not closure. The minimal test characterises one mechanism in isolation; the real context may contain more. Close a bug only when both the minimal test and the full end-to-end scenario are green — this codebase has seen isolated tests go green while the hardware stayed red.
Reference-first for compatibility-bound behaviour
When the failing behaviour is something we match to a reference — Python-RNS for protocol mechanics, the RNode firmware for LoRa CSMA — measure the reference on the same failing scenario before committing to a fix direction. Otherwise you cannot tell a bug in our stack from a property of the protocol.
"Same scenario" is strict: every input equal except the stack under test. Exploit the drop-in compatibility (Python-RNS Compatibility) — the harness points the same client code at either daemon, never a parallel per-stack driver, which would smuggle configuration differences (cadences, phases, timeouts) into what claims to be a stack comparison. And before interpreting any A/B result, count the event volumes on both sides: if they differ by more than a few percent, the comparison is invalid and any downstream timing analysis is meaningless — fix the test, not the hypothesis.
No noise framings
"Edge of window", "environmental variance", "probably a flake" are not diagnoses; they are the absence of one. In a controlled lab there is no noise floor to hide behind: any benchmark below 100 % packet delivery is a bug, and a flaky test is a deterministic bug with a flaky symptom. Re-running until green is forbidden — a failure is root-caused, fixed at the root, and the fix verified to address the cause rather than mask it. A pre-existing failure is still a failure; broken tests are not accumulated as known issues.
The same honesty extends to results reporting: a run that skipped a gate says so, a measurement taken under non-default conditions names them (see Regulatory Airtime), and a bound that does not cover something says what it dropped. Silent gaps read as "covered".
See also
- Wire Field Semantics — the field-level testing rule: pin meanings, recomposed independently.
- Python-RNS Compatibility — the drop-in property that makes honest A/B comparisons possible.
- Regulatory Airtime — declared deviations and non-disturbing diagnostics, applied to radio law.
Checks That Are Actually Checks
A check makes two promises: that it ran, and that it could have failed. A citation makes a third: that it still points at what it claims. None is self-evident, and this codebase has broken all three — silently, and for months at a time.
This page records the three mechanical guarantees that make those promises verifiable, what they cost, and what they do not reach.
The underlying rule is in Evidence and Honesty: a check you have never seen fail is not a check. That page tells a person what to do. This one is about the cases where the person forgot.
The incidents, and which guarantee would have caught each
| Defect | Caught by |
|---|---|
execute_benchmark ended in an unconditional Ok(()) (periculum/src/assertions.rs); 6 of 10 recorded runs carried zero packets and reported GREEN | neither — a scenario step, not a Rust test |
| the sx1262 RX-extend guard tested IRQ flags its own latch mask had disabled (#144, fixed 26ce3a0) | neither — no test at all, and exclude (Cargo.toml:59) puts leviculum-nrf outside the workspace |
status_parity #[ignore]d with a reason naming a procedure no script implements; never executed by any gate (#189) | B |
14 further ignored tests in rnsd_interop executed by nothing (#189) | B |
| scenario steps across the corpus produced a delivery figure no step asserted; GREEN at 70-90 % (#188, Periculum #25) | neither — scenario steps |
| the status-parity volume guard compared one interface, so a whole-inventory divergence stayed green (#177) | A, only if the author's negative control covers the whole inventory rather than the one interface they compared |
drifted file:line citations — six across five concept documents in the 2026-07 manual audit (leviculum-std/tests/doc_citations.rs:9), sixteen across the whole book on the guard's first automated run | C |
reference/LXMF sat twelve commits behind its gitlink for five weeks; every LXMF citation meant something other than it said | C, and the red reference_lock test that should have said so was itself unobserved — a B failure masking a C failure |
a Co-Authored-By: naming a model reached a periculum commit on 2026-08-07, against a rule the same author had cited correctly hours earlier (#205) | neither — a commit message, which all three explicitly do not reach |
PROCESSOR_TICK_BUDGET justified the only number in a public API constant with "the number comes off docs/…/core-lock-budget.md" and then named 126.6 ms; that figure occurred exactly once in the tree, in that comment (#200) | C, only since the figure check below — a prose attribution carries no line and no identifier, so the resolver never saw it |
just standard held a decided red for two hours, alive and silent, because a test that aborted in a destructor leaked a daemon holding the gate's stdout pipe (2026-08-07) | none of the three — every one of them reports, and a gate that never terminates reports nothing at all. See A gate must pass, fail, or say it gave up |
seven orphaned scripts/test_daemon.py processes alive at once on 2026-08-07, the oldest over four hours, from several different runs — every one of them left by a test whose Drop was written correctly and did not run | none of the three, and nothing else either: an orphan makes no gate red, so the only thing that ever reported it was somebody running pgrep by hand. Now B, via the census and the SIGKILL proof under A harness that spawns a process must ensure it dies with the harness |
The last row is worth reading twice: the guarantees are not independent. A rotted citation had a test attached, and that test ran nowhere. Guarantees that only report are worth what their observation is worth.
Guarantee A — every pin carries its own negative control
A pin is a test that fixes a claim we rely on: a wire-field semantic, a deliberate deviation, a measured protection level, a chosen non-behaviour.
The rule: a pin must contain an assertion that fails when the claim it
pins is broken, in the same test, on every run. The pattern is in the
tree — announce_signature_covers_reference_byte_order_on_the_wire
(leviculum-core/src/destination.rs:2250) verifies its signature, then
drops the destination hash from the front and asserts that verification
now fails.
The gate checks presence. Only the audit checks efficacy.
A registry gives the gate a citation. The strongest thing it can check is that the cited assertion still exists where it says — drift detection. It cannot see a vacuous control: one that verifies against a random key, one behind an early return, one in a branch the test never takes. Efficacy is checkable only by mutating the pin's subject and confirming the pin dies.
A is therefore not in force until the audit runs. The registry gate will land first because it is cheap; a tree that has only the registry has drift detection on negative controls and nothing more, and should not be described as having Guarantee A.
Why mutation cannot be the per-batch gate
-j used to be unsafe here, and ports were the reason. Storage never
was: the suites take theirs from tempfile::tempdir(), and
cargo-mutants gives each job its own copy of the tree anyway. Both
the mvr and rnsd_interop suites drew listeners from a counter over a
fixed band, 61000-65000, and that counter was per-process: each test
binary started at the same base and walked the same numbers, so two
concurrent processes raced in the alloc → bind handoff window. Measured
on two concurrent runs of the mvr binary under strace -e trace=bind:
10 ports bound by both processes and ~80 EADDRINUSE binds in the band
per run, none of which went red, because the probe loop retried.
That is fixed. next_port_candidate
(leviculum-std/tests/support/port_alloc.rs) draws from one counter per
host, kept in a file and bumped under flock, so no two processes are
handed the same number — the same measurement after the change reports
0 and 0. tests/port_alloc_multiprocess.rs pins it with four concurrent
worker processes and a negative control that must collide.
What is left is cost, not correctness: cold baseline builds and hang-mutant timeouts, and the rebuild term below.
Measured on a 32-core host (musl, warm; the CI host has four cores,
where every figure below is worse): incremental rebuild after a content
change in leviculum-core/src/transport.rs ~1.65 s; the downstream
leviculum-std --test mvr binary ~1.8 s; leviculum-core --lib
build-and-run 4.3 s. Floor ~2 s per mutant. cargo-mutants generates
roughly 3-15 mutants per subject (its documented behaviour, not measured
here). At an assumed 200 pins that is 600-3000 mutants:
realistically over an hour, against a ~15 min batch budget. The
dominant term is the rebuild, which scales with the codebase, not with
the number of pins.
Two reach limits. cargo-mutants replaces a function body with a
guessed value, swaps binary operators, deletes unary ones, deletes match
arms a wildcard still covers, replaces match guards with true and
false, and deletes fields from struct literals that have a base
expression. It does not substitute literals and does not mutate consts —
where this project's semantics live: PATHFINDER_RETRIES
(leviculum-core/src/constants.rs:157). The #192 defect was
retries: 0 where PATHFINDER_RETRIES belonged, and the fixed site
writes PATHFINDER_RETRIES (leviculum-core/src/transport.rs:11309)
into a literal that spells every field out — so not even the
field-deletion operator reaches it, and no operator substitutes one
const for another.
What is a pin
Marked in the test by a doc-comment tag, and listed in
scripts/pins.txt with the location of its negative control. The gate
checks both directions: every tag has an entry and every entry has a
tag. Registration alone would make pinhood circular — the gate would
check that everything in the registry is in the registry, and a
pin-worthy test nobody registers would be silently absent.
The location is checked by reusing the window rule from
leviculum-std/tests/doc_citations.rs, so a moved control is caught the
same way a moved doc citation is. That is Guarantee C doing A's work,
which is the point of having all three on one page.
Guarantee B — every test is executed by some gate
Half of it exists: scripts/check-ignored-counts.py enumerates tests
per unit and pins the ignored count. That stops the bucket growing
silently; it says nothing about whether anything in the bucket runs.
The other half:
- Every gate emits a manifest of what it executed, by test name,
per binary, derived from run output, not from
cargo test --list— a list records intent. The hazard this must catch is a by-name selector that matches nothing:cargo test <filter>runs zero tests and exits 0, so a gate that selects by name reads green whether or not it measured anything.scripts/run-status-parity.shalready closes that hole by hand, pinningEXPECTED=3and parsing the summary line back — which is one gate's worth of what a manifest gives every gate. - A check reports every test that exists and appears in no manifest, by name.
Enumerate at runtime, pin only counts
Item 2 must cover all tests, not only pins and declared exceptions.
The 14 unrun rnsd_interop tests were ordinary tests; a report scoped
to pins would not have seen them, and the per-unit count would have
stayed right the whole time — which is the argument against counts,
reintroduced.
The existing set is enumerated at run time, the way
check-ignored-counts.py already does it. Nothing needs a checked-in
list of ~3600 names, which would invite a --bless flag that silently
blesses the deletion it exists to catch.
What is pinned is counts per unit, and only for deletion
detection: a test that vanishes leaves both the manifest and the
runtime enumeration, so only a baseline catches it. Counts suffice for
that, and the file stays the size of ignored-counts.txt.
Note the per-configuration caveat: leviculum-ffi is gnu-only and
outside default-members, leviculum-nrf outside the workspace, so the
canonical configuration for the counts must be named.
Rollout has an ordering constraint: counts and reporting cannot be enforced until every gate emits manifests, or everything reads as unrun. Manifests first, enforcement second.
Step 1 is built: scripts/run-with-manifest.py wraps a gate's test
command, parses the names out of the run and writes one JSON manifest per
gate under ~/.local/state/leviculum-ci/test-manifests/, next to the
other CI run state rather than in the tree or in target/ — the tiers
run with their own CARGO_TARGET_DIR, so a manifest under target/
would split the union into one per tier.
Coverage was the problem the manifests then made visible: fourteen gates
emitted, and their union covered 3361 of the 3719 tests the workspace
held. The 358 outside it were 31 #[ignore]d and 327 ordinary tests
executed by no gate at all — leviculum-cli and leviculum-micron
entirely, most of leviculum-lxmf, the jl/jldiff suites and every
lnomad test (#194).
Naming the missing 327 is what produced the gap: naming is a list, and a
list loses something again with every new test file. The fix inverts it.
just complete, in the extensive tier, runs
cargo test --workspace --all-targets --no-fail-fast and
cargo test --workspace --doc --no-fail-fast, selects nothing by name,
and therefore covers everything by construction. Tiers define latency,
not coverage, and leaving the complete run is what has to be declared.
Sixteen gates emit today and the union covers 3690 of 3721; the 31
outside it are #[ignore]d, and no ordinary test is outside it.
Two spellings matter and neither is optional. --all-targets makes cargo
drop doctests, so --doc is a second invocation. Without --no-fail-fast
cargo stops after the first red binary, and the manifest would then record
a prefix of the workspace while reading like the whole of it.
Step 2 must normalise libtest's run-time name suffixes before it can
report anything: a no_run doctest is listed plainly and run as
<name> - compile, exactly as a #[should_panic] test is run as
<name> - should panic. Six doctests read as uncovered in the first
measurement for that reason alone, having in fact executed.
Declared exceptions expire by execution, not by date
A test may legitimately run in no automatic gate: a rig test needing hardware, a soak run before a release. The mechanism does not forbid it; it requires the fact to be declared, with who runs it and where.
An expiry date is gateable but is the wrong quantity — it goes red on a day unrelated to any change, and the cheapest fix is bumping it a year, a one-line diff indistinguishable from maintenance. Instead: an exception is stale when no manifest has recorded that test executing for longer than N. It uses the manifests already being built, cannot be satisfied by editing a number, and folds the exception list into the staleness bound rather than leaving it a separate off switch.
The same bound applies to manifests themselves: a union with no age limit counts a gate retired weeks ago.
The repo tried exactly this bound for tier 2, and how it failed is worth
recording, because the failure was not in the idea. pre-push blocked
when the newest tier2 GREEN line in the CI ledger was over 24 hours or
10 commits old. The only writer of that line was
scripts/run-tier2.sh; the timer that started it was retired on
2026-06-12; and the remedy the block printed — just extensive — drives
periculum directly and writes no line at all. So the bound became
unsatisfiable on the day the timer went, and stayed that way for 46 days
and 502 commits, every one of which reached master through
--no-verify. That flag disables the whole hook, lint and Tier 0 and
the trailer guard with it. The block was removed on 2026-08-07 rather
than repaired.
A staleness bound measures its writer, not its subject. Age a
manifest out only against a signal that something is still scheduled to
emit, and let the remedy name that emitter rather than a recipe which
merely looks equivalent. A bound whose remedy cannot clear it does not
fail open or closed — it fails into --no-verify, and takes the checks
that worked with it.
Two reach limits: tier-3 hardware manifests and the release-only
hardware corpus (periculum list hardware prints the live count; it
grows) are produced on another host on a per-release cadence, so
either their transport is specified or the guarantee is scoped to
host-runnable gates. And nextest cannot run doctests, of
which scripts/ignored-counts.txt tracks two units (leviculum-core --doc and leviculum-std --doc) — whichever runner is used, the
doctest gap is explicit.
Guarantee C — every citation still points at what it claims
Our method rests on citations: a pinned deviation means nothing without the reference line it deviates from. When a citation rots, the test still passes and the sentence still reads — it simply stops being true.
leviculum-std/tests/doc_citations.rs is the working exemplar and the
proof that this rots fast. The practice of checking the reference
before auditing against it is already stated in
Wire Field Semantics; what follows is the
mechanism, not a restatement of the method.
1. The submodule check is O(1), and that is the whole incident
The five-week LXMF drift was not hundreds of citations going wrong
individually. It was one fact: the checked-out submodule disagreed
with its gitlink. The cheap catch is a per-batch assertion that
git submodule status reports no +/- for the four vendored
references — a few lines of shell, in a gate rather than in a #[test]
that Guarantee B can fail to observe.
That, and nothing more, would have caught the incident on day one.
2. Re-verification on a bump is a separate, larger problem
Binding every citation to a submodule commit and failing until they are
re-verified is a different mechanism, and it did not cause the incident.
If it is built, it must be incremental or it will not be used: on a
bump, git diff --name-only old..new inside the submodule bounds the
work to citations into changed files — usually a handful, not hundreds.
This is what makes gap 3 a prerequisite rather than an afterthought.
A citation written as ``ident (path:line) can be re-located
mechanically at the new commit and its line updated. A bare
Transport.py:2970 can only be re-verified by a person reading it. So
converting reference citations to the identifier-adjacent form is what
turns a submodule bump from hundreds of manual re-reads into a
mechanical re-resolve.
3. Coverage: source is uncovered, and most citations are weak
The guard reads docs/src/**. A grep for reference citations
(reference/…, Transport.py:, LXMRouter.py:, Destination.py:)
across leviculum-core, leviculum-lxmf and leviculum-std finds
421 in their src/ trees, 497 counting their tests/ directories —
against 804 covered in the book — and the uncovered ones are
load-bearing, because they are what a pinned deviation cites.
Extending the glob to Rust source is cheap and should be done, but be clear what it buys: for a bare citation it is existence-and-length checking, so it catches renames and deletions and not drift inside a file that stays long enough.
Counts from the guard's own output at the commit that landed this page: 804 total, 72 identifier-checked, 731 bare, 1 external. The bare majority is an editorial problem — the fix is to write citations in the identifier-adjacent form, and it cannot be mechanised without rewriting prose. The honest response is to publish the ratio on every run so nobody reads a green guard as full coverage, and to convert opportunistically (#167 is the standing example).
That last sentence was half wrong, and section 5 below is what replaced it. The identifier form cannot be mechanised — but the identifier is not the only thing about a citation that survives a move, and the other thing needs no prose written at all.
4. A prose attribution is not a citation, and was not checked
leviculum-std/tests/doc_citations.rs resolves path:line and checks
identifier proximity. A sentence that attributes a number to a document
by name carries neither, so it sailed through — and that is not a rare
shape. PROCESSOR_TICK_BUDGET
(leviculum-std/src/driver/processor.rs:181) was justified with "The
number comes off docs/src/concepts/core-lock-budget.md" and then named
126.6 ms. The figure occurred exactly once in the whole tree: in that
comment. The measurement was real — taken in the #196 design pass — but
the page it was attributed to did not contain it, so the only number
behind a public constant could not be traced by anyone but its author.
The check: a doc-comment paragraph naming a document under docs/ and
quoting a decimal figure must have that figure occur in that document.
figure_attributions in the same file, with the reach limits written
where a reader hits them.
Two decisions are worth lifting out of the code.
Paragraph scope, not sentence. The defect attributed across a sentence boundary — the document named in one sentence, "The failure mode it names is 126.6 ms" two sentences later — so a sentence-scoped trigger would have missed the case it exists for. Reconstructed and run: it is reported at paragraph scope and invisible at sentence scope.
Decimal figures only, which is where the precision comes from. In the paragraph the defect lived in, "126.6 ms", "3.2 ms" and "0.8 ms" are the page's figures; "5 ms" is the constant being defined and "~25x" is arithmetic done in the comment. Checking every number would have reported three of its own numbers alongside the one real finding — on the very comment the check exists for. What it gives up is integers: "the page names 141 ms" is unchecked, and a wrong round number is as believable as a wrong precise one. That is the largest known gap and it is stated rather than closed, because a guard with false positives gets switched off, and a switched-off guard is worse than none.
Two paragraphs in the tree trigger it and five figures are checked. Both numbers are printed on every run for the same reason the citation counts are: with a trigger this narrow, "no failures" and "the trigger stopped firing" are otherwise the same output.
5. The bare half is checkable against the tree's own history
The counts above are a coverage ratio, and a ratio does not say how many of the uncovered citations are actually wrong. Measured, on 2026-09-21: of 3228 citations in the two corpora, 2464 carry no identifier, and 424 of those pointed at the wrong line — 541 individual line numbers. Before that measurement the guard had been reporting zero for as long as it had existed, and the only competing figure was the 126 an aborted merge sweep happened to touch.
Nothing in the tree as it stands can check a bare file.rs at line 810. The
file exists and has 810 lines, and it goes on having them after the
cited code slides to 883. But the text of the cited line survives a
move exactly the way an identifier does, and unlike an identifier it is
already there — no citation has to be rewritten to acquire one. The
tree's own history is where it is kept:
git blamethe line the citation sits on gives the commit that last wrote it — the newest moment anyone can be assumed to have looked at the citation.- The cited file as of that commit, at the cited line, is the anchor.
For
reference/<submodule>citations that means the submodule's own history at the gitlink this tree pinned back then, so a citation into Python-RNS is checked against the reference we actually pinned. - If the cited line still holds that text, the citation still points at what it pointed at.
- If not, and that text now sits at exactly one other line, the number is wrong and the guard says by how much.
The unique match in step 4 fails a run with a repair. That is the load-bearing choice: the cited text is demonstrably in the file at a line the citation does not name, so the report needs no judgement about what the citing sentence meant, and a guard with false positives gets switched off.
An anchor that has vanished fails a run too, but without a number: the code was rewritten in place, and what the sentence should point at now is for a reader to decide, so the report says the text is gone and sends the reader to the sentence. Until 2026-10-05 this verdict was computed and then filed as undecidable whenever no endpoint of the citation was fresh or moved, which is every single-line citation. The status line read "0 rewritten in place" on every run, and 33 cited lines whose text was gone sat in the undecidable count, green; 16 of them pointed into the Python reference after it moved to 1.3.5.
Everything else is counted in the open and decides nothing. The run prints the undecidable count split by reason, so this census is on every run rather than in a one-off instrumentation; on 2026-10-05, after the 33 were fixed:
| Undecidable because | doc | source | Could a rule decide it? |
|---|---|---|---|
| the cited line was blank when it was cited | 48 | 147 | Only when the citation's worded endpoints place it; a blank line alone, never |
| the cited file was absent at the citing commit | 30 | 3 | Yes: every sampled case is the reticulum-* to leviculum-* crate rename, so following the rename to read the old blob would decide them |
| the cited text is on several lines | 16 | 11 | Some: a wider context window than one line either side; a wordless line (}) alone, never |
| the cited file was shorter than the citation when cited | 7 | 4 | Decidable as an error, not as drift: the citation was wrong when it was written |
| the cited line carries no word to anchor to | 2 | 0 | Never on its own, by design (below) |
| the citing line is not committed | 0 | 0 | Once it is committed |
| Sum | 103 | 165 |
Two refinements earned their place by being needed on the real corpus, and both place endpoints by evidence rather than by inference:
- A line that is ambiguous alone is often unique with its
neighbours.
else:is forty lines inTransport.py;else:with the line above and below it is usually one. - A range whose endpoints agree on one displacement can carry the
endpoint that placed nothing.
Transport.pylines 1722-1764 end on a]the reference now has four of, and exactly one of them is a line from where the opening line's +145 puts it.
And one anti-refinement, which the corpus also demanded: an anchor
carrying no word never establishes drift. Nobody cites a docstring
delimiter on purpose. Identity.py lines 84 and 383 pointed at a stray one the
day it was written and at the right constant today; "repairing" it to
where that delimiter went would have broken a correct citation. A
wordless anchor can still be carried along by a displacement the rest of
the citation has established — that is what keeps the ] case
repairable — but on its own it says nothing.
What this cannot see, stated rather than hidden. The baseline is the citing line's own last edit, so a citation that was already wrong when it was written passes, and reflowing a paragraph re-baselines every citation in it. A citation whose sentence went wrong while the cited line stayed put is invisible here as it is to every other check on this page. The number is a floor, not a census.
The identifier-anchored class gets the same anchor, since
2026-10-06. Both classes run through one anchoring pass and are
reported apart (bare-citation anchors, identifier-citation anchors),
so a count can still be compared with a run from before. For a named
citation the identifier decides nothing on its own: it only breaks the
tie when the anchored text now sits on several lines, picking the
occurrence nearest a line that carries the name, and only within
WINDOW of one.
Until then that class was checked by proximity alone: the identifier
anywhere within WINDOW (8) lines of the cited span, or enclosing it,
held. A shift of up to eight lines was green by design, and so was a span
that never covered what it claimed as long as the name was close by.
Turning the anchor on found 108 named citations that tolerance had
hidden (129 moved endpoints in the book, 2 in source, 2 rewritten in
place, measured 2026-10-05), and the day it was measured showed the cost
live: the pass that measured it moved the Placed::Proved line this
page cites, and the guard stayed green, because the enum variant Proved
sat three lines from the stale number. The proximity check still runs;
it is now the second check a named citation passes, not the only one.
The fixer repairs a named citation on the same evidence it acts on for a bare one, and on nothing weaker. All three must hold:
- Every endpoint moved by one common shift. A range whose ends moved by different amounts grew or shrank, and whether it still covers what it meant is a question for a reader.
- Every endpoint's text is unique in the file now, alone or with
its two neighbours. An endpoint that only the shift or the
identifier's tie-break placed is inferred from the rest of the
citation, and the inference carries forward whatever offset the
citation had when it was written. The
select4range in the embedded developer guide is the case: it opened on a))four lines above the call it meant, and the shift its unique end established would have kept that error exactly. - The identifier sits inside the repaired span. The text moving is what says the line moved; the name being in the new span is what says the citation's subject moved with it.
Landing the rule, the fixer rewrote 83 citations in one run (the 108, less those it refused, plus the ones the rule's own edit to the guard displaced) and refused 28, which were read by hand one at a time and committed apart so their diff can be read alone. A refused citation is reported with the number the anchor would give it and the condition that failed, never with a bare "should read".
Four standing controls keep the verdicts honest, each a fixture pair
under leviculum-std/tests/fixtures/ (citation_anchor_target.rs.in
and citation_anchor_doc.md.in, one bare and one identifier-anchored
citation) committed into a scratch repository and then perturbed:
anchor_control_a_moved_line_is_named_with_its_new_number (the bare
citation is reported moved with its new number, the identifier citation
holds at +1), anchor_control_a_rewritten_line_is_named_rewritten_not_undecidable
(reported rewritten in place, with no number offered),
anchor_control_an_identifier_moved_past_the_window_is_drift (the
identifier one line past the window is drift, with the distance named),
and anchor_control_an_identifier_citation_moved_three_lines_is_named_and_repaired
(the named citation's line moves by three, inside the window: the
proximity check still holds it, the anchor names the new line, and the
fixer rewrites it). an_identifier_citation_is_renumbered_only_on_proof
refuses each of the fixer's three conditions on its own.
Because the finder knows where the text went, it can also put the
citation there: LEVICULUM_CITATION_FIX=1 rewrites the repairable ones
in place. Finder and fixer are the same code deliberately — a separate
fixing script would be a second implementation of the anchor rule, and
the first symptom of the two disagreeing is a repair pointed at the
wrong line. 445 of the 2026-09-21 findings were repaired that way; the
remaining 21 were spans whose other end a human had to locate, and eight
of those left the bare class entirely by acquiring the identifier they
should have had.
Until 2026-10-06 that fixer reached the bare class only; it now reaches a named citation as well, on the three conditions above. What repairs the identifier-anchored classes when the anchor cannot, without re-anchoring them onto their identifiers, is section 7.
The fixer has one trap worth naming, because it sprang once: this guard is itself in the corpus it guards, so a fixture string naming a git hook and a line number in its own tests is a citation as far as the scan is concerned, and got "repaired" to the line that text had moved to — breaking the assertion beneath it. Fixture citations in that file are now built rather than spelled out, which is what the canary fixtures already did and for the same reason.
The standing canary is a miniature git repository: two bare citations,
one of which a second commit makes wrong by moving the code under it.
Both directions are asserted, because every failure mode of a check made
of subprocess calls — blame returning nothing, a cat-file batch
desynchronising — produces no findings, which reads exactly like a clean
tree.
6. A reference table names its subject, and was read as naming nothing
Sections 1-5 divide every citation into two classes by one question: does an identifier sit immediately before it? That question has a third answer, and it is the densest citation shape in the book. A reference table writes
| `fn has_path(&self, dest_hash: &DestinationHash) -> bool` — `driver/mod.rs:NNNN` | Whether a path is known |
The row names its subject as plainly as a citation can, and the
adjacency rule read it as naming nothing, because the name is inside a
signature several words from the citation. docs/src/developer/rust-api-spec.md
is 113 such citations on its own — the largest single bare cluster in
the corpus — and a 17-row ReticulumNode method table in it had aged
past a thousand lines under a green guard (Codeberg #307).
So: inside a table row, a fn NAME( in a backticked span names the next
citation on that row. A row rather than a cell, because the corpus
writes the signature and the citation in one cell and in adjacent cells
about equally, and where the | falls between them is a typesetting
choice. What replaces adjacency as the guarantee of an unambiguous
pairing is that a citation between the signature and this one takes
the signature for itself — the row's second citation stays bare, exactly
as it was.
Section 5's history anchor does not reach this case and could not: the method table's citing lines were last written in a restructuring commit older than the file they cite at its current path, so there is no blob to anchor against and every one of them counted as undecidable. A name the row already carries needs no history at all.
The measurement, on the corpus the day it landed: 76 book citations
moved from bare to drift-checked, and 59 of the 59 signature rows in
rust-api-spec.md were pointing somewhere other than the item they
name — 41 of them far enough out to fail the guard outright, the other
18 inside the 8-line window with the budget already spent. Each was
re-derived from the definition in the file the section names, not
shifted by the distance the guard reported; a citation that was already
wrong and gets shifted to a new wrong number is harder to spot than one
that is obviously stale.
What it still does not reach, in the same file: the prose citations
around those tables (Defined at …, the enum-variant rows whose name is
in another cell, the newtype accessors written as a signature followed
by a parenthesised citation). Those remain bare, and section 5's anchor
is what covers them.
7. The repair follows the diff, not the nearest name
Section 5's fixer repairs a citation by its own anchor: the text of the line it names says where that line went. Until 2026-10-06 that anchor existed for the bare class only, and it still repairs a named citation only on the three conditions section 5 lists. What the proximity check prints for the identifier-anchored classes is the nearest occurrence of the cited identifier — and that is exactly the number a repair must not take.
The reason is the trap the order-257 coder recorded, which is worth
quoting rather than paraphrasing: a 16-line doc comment inserted at
line 950 of transport.rs reddened 121 citations below it — 38 bare
ones the LEVICULUM_CITATION_FIX mode rewrote, and 83 identifier-
anchored ones it cannot fix, which the coder repointed by hand from the
git diff -U0 line map, because the guard reports "identifier now at
3569" for a citation that deliberately points three lines ABOVE its
identifier (3566 after the move), so following the report would have
re-anchored 83 citations onto their identifiers and destroyed the
offsets their authors chose. A manual step with a known failure mode,
performed about 120 times in one pass, is a tool that has not been
written yet. Three passes in one day paid that tax; one of them for a
single citation.
So the map is the diff and the proof is the text.
LEVICULUM_CITATION_FIX_BASE=<rev>names the state the citations were right about. The default isHEADwhen the tree has uncommitted changes to tracked files — the insertion is still in the working tree — andHEAD~1when it has none, the insertion being the commit just made. Untracked files do not count as dirty: a stray log beside the tree says nothing about where a cited line was. Which base was used is printed on every run, because a map read against the wrong base is this mode's one failure mode.git diff -U0 <base> -- <cited file>is a line map. Every hunk that ends above the cited line displaces it by that hunk's own length change, and nothing else does.-U0is what makes this true: with context lines a hunk's bounds say nothing about which lines actually changed. A pure insertion is written-N,0and lands after old line N, so it displaces every line past N and contains none — reading its start as a contained line would report the line immediately above an insertion as replaced by it, which is the commonest shape there is.- A cited line inside a hunk was not displaced, it was replaced. Its text is not somewhere else, it is gone. That one is reported with the line the hunk now starts at and never rewritten: what replaced text meant is a question only a reader can answer.
- The rewrite happens only when the mapped line now holds the text the
cited line held in
<base>. That is what makes the author's offset provably preserved — the citation follows its own line, whatever that line pointed at and however far the identifier has moved. Otherwise the citation stays red and both candidates are printed, the mapped line and the nearest-identifier line, for a reader to choose between.
Step 4 is the whole safety of the mode, and it is worth being exact about what it can and cannot catch. A correctly parsed diff against the right base cannot fail it: the map is exact by construction. It fails when the premise does — a base at which the cited file did not exist, a line past the end of that copy, a path inside a reference submodule whose lines this tree's diff does not move, or a base against which nothing under the citation moved at all. In each of those the citation is left red rather than renumbered from a map that has stopped describing the tree. All four are the same sentence: a number this cannot prove is a number it does not write.
The fixers share their finder and their byte surgery for the reason section 5 gives, and they run in one invocation rather than in parallel: both rewrite the same citing files, and a bare repair changes the length of the citation it rewrites, so the identifier pass rescans before it places anything. While a fix run is rewriting the corpus the three resolving guards skip — loudly, naming the variable that silenced them, because an environment variable that quietly turns a guard green is the shape this page exists to remove. A fix run applies no verdict; the verdict is the next run without it.
The measurement that landed it is the 257 corpus itself, replayed. A
copy of the tree at the commit before 257 with 257's code and 257's
stale citations — the exact state that coder faced — repairs to 38 bare
and 83 identifier-anchored citations, which is that pass's split to the
citation, and the resulting tree is byte-identical to the one the coder
produced by hand across all 119 changed lines. On the tree this landed
in, the live case was a citation into memory.x whose line the
preceding commit had pushed 38 lines down: the line map places it at
291, where the ASSERT it names now is, while the nearest match in the
guard's report was line 289 — a comment that merely mentions ASSERTs.
Two lines, silently, from following the report instead of the diff.
Its fixture repository is in the guard's own tests: a citation three lines above its identifier with an insertion above both must move by the insertion and not onto the identifier, a citation whose line was replaced must stay red with both candidates named, and a citation into a file the base has no copy of must not be touched. Each has an injected-drift control, because two of the three verdicts are "leave it alone", and a fixer that has stopped repairing anything leaves everything alone.
7a. The fixer runs once, and it runs last
Once per pass, after the last edit to any file it cites into, and
immediately before the gates. The mode could not notice that it was
repairing its own earlier output, and what it did when handed that
output was not to give up: it applied the same displacement a second
time and printed repaired by the line map. It refuses that case now,
by the premise test at the end of this section — but the rule stays,
because a refusal still leaves the citation red and a reader has to
unpick it.
Step 4 above, read from the other side, is the mechanism. The proof is
Placed::Proved (leviculum-std/tests/doc_citations.rs:3450) — the
mapped line's text now against the text the cited number held in
the base. The premise of the whole map is therefore that every number
in the corpus is one the base was right about, and a citation an
earlier run rewrote is not. Its number names a line in the post-edit
tree, so the base text it is compared against is whatever unrelated
line sat at that number before the insertion; a pure insertion
displaces every line below it by the same amount, so that unrelated
line is still exactly that far down, the comparison passes, and the
citation moves a second time.
Measured on this tree 2026-09-27 by replaying order 339's shape: HEAD
2242a470, 85 lines inserted at line 191 of
leviculum-core/src/transport.rs, the fixer, then three more lines at
the same place, which is 339's "off by three". The citation under test
sits at line 327 of docs/src/concepts/time-and-clocks.md, written
seven lines above the rank it names. None of the numbers in the table
is in the citation form, for the reason section 5 gives: they record
where a citation was, and a later fix run would helpfully rewrite
them.
| Run | the citation reads | correct | lines to rank |
|---|---|---|---|
| before | 1389 | 1389 | 7 |
| 1, after the 85-line insertion | 1474 | 1474 | 7 |
| 2, after three further lines | 1562 | 1477 | 78 |
| 3, no further edit | 1650 | 1477 | 166 |
| revert the rewrites, insert 88, run once | 1477 | 1477 | 7 |
Runs 2 and 3 each reported 2 red citation(s), 2 repaired by the line map, 0 left for a reader. It neither converges nor warns. The verdict
run afterwards is still red, so nothing lands on a lie — but every
attempt to help moves the citation another 88 lines from its subject,
which is why 339 could not repair its two survivors and reverted
instead. That pass recorded the mechanism as "the map pass leaves a
rewritten citation alone"; the measurement says the opposite, and the
opposite is the worse of the two.
The bare pass does not double-apply. It goes quiet instead: its anchor is read from the tree's own history, and a doc line the fixer has just rewritten is uncommitted, so there is no commit at which to ask what that line said when it was written. Run 1 reported 21 doc and 13 source citations moved and rewrote 31 of them; run 2, with those same citations now three lines stale, reported 0 moved and its undecidable counts up from 79 to 98 and from 17 to 29 — 19 and 12, which is the 31 it had repaired. Undecidable is not a failure, so for the bare half a second run turns red into silence.
Recovery is the one 339 used: revert the rewrites, rebuild the tree as
the base commit plus the source edits, run the fixer once. Saving the
source hunks and git checkout -- . is enough. No base revision
repairs a doubly-rewritten citation, because the state its number was
right about is a working tree nobody committed.
Committing the rewrites is the other way out, and the same statement. With run 1's output committed, a run after three further lines moved 1477 to 1480 and the bare pass was back to 21 moved and 31 rewritten. So the rule is not really "run it last": it is the citations must be right about a commit, and running the fixer last is the cheap way to be sure they are.
Could it detect its own earlier run? Not by the test that suggests
itself. "The mapped line does not hold the base text, but the line at
the cited number does" fires on neither measured case. At 2242a470
line 1474 of transport.rs is pub(crate) packets_sent: u64,; the
mapped line 1562 holds exactly that, which is why the repair was
applied again; and line 1474 of the working tree holds the age_secs
doc comment instead. Both halves are false, and the half that would
have to be true is the one the code already proves. It is unsound in
the other direction too: wherever trimmed lines repeat — }, a
blank, a bare /// — the two texts match by coincidence.
The test that does fire is the premise rather than the result: the
cited line must resolve its own identifier in the base. At 2242a470
nothing within eight lines of 1474 contains rank, and nothing within
eight lines of the second survivor's 478 contains TickOutput, while
their pre-run numbers 1389 and 393 both resolve. Replayed over the
three logs with that rule, the enclosing-block fallback aside: 0 of the
83 doc repairs the single clean run made would have been refused, and 2
of 2 in each double run.
So it prints a refusal in place of repaired by the line map, not a
hint: left red — line 1474 does not resolve rank at 2242a470
either, so this citation was never right about the base and the line
map cannot carry it. If an earlier fix run rewrote it, revert the
rewrites and run once.
Built 2026-09-27. The proximity test had lived inline in the checker,
closed over the current file's lines; it is now ident_resolves
(leviculum-std/tests/doc_citations.rs:695), a function over the lines
it is handed, and place_citation asks it a second time against the
base copy of the file (leviculum-std/tests/doc_citations.rs:3662). A
map repair is emitted only where both halves hold: the map proves the
move, and the base resolved the citation's own identifier at the
citation's own number. The refusal above is the other branch.
What it costs is a real refusal, paid knowingly. A citation already
stale at the base — the twenty Justfile numbers 339 found — is no
longer repaired by the map pass. Correctly so: a map against that base
says nothing about drift older than it. The bare pass still repairs
those where it can, from the tree's own history, which is what repaired
them in 339.
The fixture is the fourth verdict of
a_repair_follows_the_line_map_and_not_the_nearest_name, beside the
three the guard already had: a citation written nineteen lines above
the twice_shifted it names, displaced by the same insertion as the
others. Its move is proved by the line map — asserted directly, so
that a case 4 which passed because the map had broken for some
unrelated reason would fail — and it is refused all the same, with the
message above. The clean single-run repair in the same tree is the
control, and is still made.
8. A green guard has to say how much of the corpus it read
Sections 5-7 made the guard see more. What it still did not do was say
how much it had not seen, and that omission is the whole of Codeberg
#307 as it was actually experienced: the guard reported the file
carrying the 1035-line-stale has_path row as fine. It was not lying —
it had checked everything it could check. What it never said was that
this was 912 of 1188 citations.
So the status line now names three groups rather than one, because they are three different claims:
doc citations: 1684 total; 814 checked against the symbol they name
(within 8 lines; their cited line anchored by its text, as a bare one
is) and holding; 0 named but not holding (reported below);
870 could not be checked by name (862 bare: anchored by the text of the
cited line in bare_citations_still_point_at_the_text_they_cited,
8 external: not in this workspace, 0 skipped: into a reference/
submodule that is not checked out)
and when the third group is more than half the corpus, a warning line follows it saying so with the number. Measured 2026-09-26, all three corpora:
| Corpus | Total | Checked against a symbol, holding | Could not be checked | Warns |
|---|---|---|---|---|
doc (docs/src/**) | 1661 | 809 | 852 (51 %) | yes |
| source (the four crates) | 2013 | 182 | 1831 (90 %) | yes |
script (scripts/*.sh, Justfile) | 8 | 1 | 7 (87 %) | yes |
A warning and not a failure, deliberately. The source corpus is 90 %
bare by nature — a // comment cites a line without writing the name
beside it — and a guard that fails on that gets switched off within a
week, which is worse than a guard that counts out loud. The number is
the point: it is the size of the blind spot, and it is now in front of
whoever reads a green run. Section 5's history anchor is the partial
backstop for that class and prints its own split (on 2026-10-05, 1205
of the book's bare cited lines still hold their text, 103 undecidable;
2820 and 165 in source), so "could not be checked by name" is not the
same as "not checked at all", and the line now says which one a corpus
got: the book and the crates are anchored, the scripts are not, and
their bare citations are still existence and length only. Until
2026-10-05 the line said "existence and length only" for all three,
which had stopped being true of two of them when the anchor landed.
Since 2026-10-06 the named group says it is anchored too, and the
warning no longer calls it "checked by proximity only".
The same pass gave the drift report the distance it had been leaving to the reader. It printed the cited line and the nearest occurrence of the identifier and stopped there; now it prints how far apart they are, because that is the number #307 was argued from. "1035 lines from the cited span" distinguishes a citation a refactor slid past from one nobody has read in a year, and a reader should not have to subtract two four-digit line numbers to learn which one they are looking at.
What C cannot reach
Issue comments, commit messages and batch reports carry hundreds of
file:line claims that nothing checks — and given how much of this
project's reasoning is recorded there rather than in the tree, that is
the largest uncovered surface of the three guarantees. A just cite
helper that emits a verified citation would reduce fabrication at the
point of writing; nothing can verify it after the fact.
One line shape on that surface, and only one
scripts/check-commit-trailers.sh is the first mechanical check on a
commit message in either repository. It is worth being exact about how
little it does: it checks one line shape, not one claim. A message
may cite a file that does not exist, attribute a measurement to a page
that never carried it, and describe a fix it did not make; none of that
is reachable from here, and the paragraph above still stands whole.
What it does reach is #205, which was not a claim going wrong but a rule
losing to a default. periculum/CONTRIBUTING.md:43 said "no AI
trailers. Commit under your real name and a reachable e-mail" while the
assistant harnesses used here instruct their agents to append exactly
that trailer to every commit. A rule in that position, with nothing
behind it, is not half-remembered — it is reliably broken, and its
violation is invisible unless a person reads every message before every
push. Which is how the one on 2026-08-07 was caught, and is not a
mechanism.
The check is a forge step
(.woodpecker/commit-trailers.yml, and the commit-trailers step in
periculum's .woodpecker.yml) over every commit since a pinned
baseline. .githooks/commit-msg runs the same script at commit time and
just fast runs it over the outgoing range, but neither is the
enforcement: a fresh clone has no hooks, which is the whole reason the
gate is where it is.
It matches only at column 0. That is not a compromise, it is the
definition: column 0 is where git's interpret-trailers and every forge
harvest a trailer, so it is where the default writes and the only place
the line is doing anything. It also leaves the one escape a message
sometimes needs — this guard's own commit message quotes the offending
trailer — namely the indentation git already uses for quoted material.
The alternative was git's own trailer block, the last paragraph; that
was rejected because the default's Generated with <tool> line sits in
its own paragraph above it, so the rule would have covered half the
default while reading as covering all of it.
A leading space defeats the check. That is a bound, not a hole: this stands against a tool's default, and nothing message-shaped stands against a person who has decided to misattribute authorship.
Sixteen commits below leviculum's baseline carry such a trailer, from
before the rule had anything behind it. They are recorded in
scripts/commit-trailer-baseline.txt rather than rewritten out of
published history, and their count is recomputed and compared on every
run — which is what stops the baseline being the off switch the expiry
dates above are criticised for being. Moving it forward to silence a
fresh failure moves a violation across that line and changes the number.
The rule was narrower than the check, and the gap cost a day
As first written the guard scanned by message text alone. The rule the
repositories actually hold is narrower: a machine-authorship trailer is
a violation on our commits and not on anyone else's — we do not edit,
and do not refuse, a message somebody outside the project wrote. The
commit-msg hook in Lew's checkout had keyed on the author e-mail since
2026-06-13 and said so in a comment; the guard did not, and could not,
because nothing in the tree stated the policy. Merging PR #201, an
external contribution whose commit carries such a trailer, then turned
just fast, the pre-push hook and the forge check permanently red on a
tree with nothing wrong with it. The brief for #205 argued the rule
entirely in terms of our own harness default and never mentioned external
contributors; the guard did exactly what it was told.
This is the failure mode the page's opening promise misses. A check that could have failed and did run can still be red for a reason the rule does not hold — and a gate that is red on a clean tree gets switched off or bypassed, which costs more than the gate was ever worth. A check also has to be able to go green on every tree the rule permits.
Authorship is the discriminator and git log --format=%ae is the whole
mechanism. Two things follow, both of which are this page's own
arguments applied one level down:
- The exemption is counted, not trusted.
foreignin the baseline file pins how many commits above the baseline carry a trailer under an author that is not ours, recomputed every run. Uncounted, an exemption granted by class is an off switch anyone can reach by setting an author e-mail; counted, a new external contribution lands with its message intact and still moves a number in a diff, which is the outcome we wanted — we want to know when it happens, we just do not want to rewrite somebody else's message. - Who counts as ours is a list in the reviewed file, not a constant in
the script. It has to be a set:
lp@lew-palm.deis what we use at a terminal, but Codeberg stamps a web-UI edit with its own noreply address, and three commits in leviculum's history carry it. A single-address discriminator would have handed the foreign exemption to the project lead's own commits, silently — the exact failure the guard exists to prevent, arriving by accident rather than by intent. There is no computed control on that list, because who is inside the project is a declaration and not a fact the script can derive; the control is that the list sits next to the counts, where adding a line is a diff.
A gate must pass, fail, or say it gave up
There is a fourth way for a gate to end, and it is worse than any red:
still running, verdict already determined, nobody told. On 2026-08-07
just standard sat for two hours in exactly that state. It was found
by noticing that its log's mtime was two hours old.
The mechanism, end to end:
python_accepts_ratcheted_c_announce(leviculum-ffi,--test ffi_interop) panicked in a destructor during cleanup — "thread caused non-unwinding panic. aborting." — and took SIGABRT.- An abort skips unwinding, so the
Dropthat would have killed thescripts/test_daemon.pythat test had spawned never ran. - The orphaned daemon had inherited cargo's stdout. Confirmed by walking
/proc/*/fd: PID 960389 held fd 2 on the same pipe. cargoexited and became a zombie.- The wrapper read
for raw in proc.stdout:. That loop ends on EOF of the pipe, not on exit of the child.proc.wait()sat behind it and was never reached. One surviving write end held the gate open.
Killing the orphan by hand finished the run immediately, printing the
exit code 101 it had held for two hours.
The class is wider than the test that triggered it: a gate that waits
for a pipe to close instead of for its child to exit can be held open by
any leaked grandchild. Every harness here that spawns an external
process is exposed — the FFI tests spawn Python daemons and C binaries,
the interop tests spawn lnsd and rnsd, others spawn docker. So the
fix belongs in the wrapper, which is one place, and not in each spawn
site, which is many and will grow.
Three properties, in scripts/run-with-manifest.py:
- The verdict waits on the child, not on the pipe.
proc.wait()runs on the main thread and a reader thread drains stdout, so nothing the wrapper decides depends on EOF ever arriving. This is the property that would have ended the two hours by itself, and the only one that still holds when the other two fail. - Orphans die with the gate. The child is spawned with
start_new_session=True, so it leads its own process group, and the whole group is killed once the child has exited — SIGTERM, then SIGKILL after a short grace, because what leaks here is daemons holding sockets, tempdirs and sometimes a serial port. The pipe then closes on its own and the tail of the output is drained normally. The cost on a clean run is one/procscan that finds nothing. - A hard timeout that reports. Default 1800 s per wrapped command,
overridable with
--timeoutper gate andLEVICULUM_GATE_TIMEOUTglobally. On expiry the gate exits 124 —timeout(1)'s code — naming itself, how long it waited and what was still alive. The per-line flush that already existed makes the partial log true for free.
And the reporting rule, which is not decoration: if the gate had to
kill survivors, it says so, by pid and command line. A leaked daemon is
a bug in the test that leaked it, and a gate that cleans up silently
hides the bug it just worked around. The manifest carries the same facts
(timed_out, killed, survived_sigkill) so a nightly that kept only
the JSON can still see it.
What none of this reaches: an orphan that calls setsid() has left the
group and survives the kill. It cannot hold the gate open — property 1
does not care — but it is still alive afterwards, and the reader thread
has to be abandoned with the tail of the log unwritten. Both facts are
printed rather than swallowed. Interrupting the wrapper is also no longer
free: start_new_session detaches the child from the terminal's
foreground group, so INT/TERM/HUP are relayed by hand, or a Ctrl-C would
trade the hang at the end for an escape at the start.
A harness that spawns a process must ensure it dies with the harness
The wrapper is a backstop, not an excuse. The rule for the spawn sites:
A harness that spawns a long-lived external process must ensure it dies with the harness, however the harness dies. Cleanup code is a convenience; the kernel is the guarantee.
Relying on Rust's Drop to kill a child satisfies the "cleanly" half and
nothing else: an abort skips unwinding, and so does a SIGKILL of the test
binary. The Linux answer is PR_SET_PDEATHSIG on the child, set after
the fork and before the exec, so the kernel signals it when its parent
dies for any reason. Drop then becomes the polite path rather than the
only one.
The receipt, from the afternoon this was written: seven orphaned
scripts/test_daemon.py processes alive at once, the oldest over four
hours, from several different runs — so the leak was the normal case
and not the exceptional one. One of them held a pipe open and hung
just standard for two hours with its verdict already decided.
leviculum_std::process::spawn_supervised is the mechanism, and it takes
its Command by value, so that a supervised spawn and a bare one do
not look alike at a call site. Four things it has to get right, each of
which has bitten somebody:
-
PR_SET_PDEATHSIGis per-task, not per-process. The kernel stores it on the child'stask_structand delivers it fromforget_original_parent(), which runs when the forking task exits — not when that task's process exits. A tokio worker or aspawn_blockingthread finishing mid-test would therefore kill the daemon under the test, which turns the fix into a flake generator. So every supervised spawn is forked from one dedicated thread that never exits, and the only event that ends that thread is the process ending. Nothing else in the design substitutes for this: thegetppid()check below reports the parent's thread group leader, so a forking thread exiting while its process lives leaves it unchanged and the check sees nothing wrong.The same fact bites the measurement, not only the mechanism.
copy_process()clearspdeath_signalfor every new task, threads included, and libtest runs a test body on a spawned thread even under--test-threads=1— so a canary written as a#[test]reads its own parent-death signal as 0 while its process's main thread carriesSIGKILL. That is why the probe below is afn mainand not a test. -
The race between
forkandprctl. If the parent dies inside that window the signal is already missed and the child runs on forever. After setting the flag the child re-readsgetppid()and_exits if it no longer names the process that spawned it. -
It does not reach grandchildren, and the remedy is not
setsid. A supervised child deliberately stays in its parent's process group, sorun-with-manifest.py's group kill still reaches everything below it — an orphan that hassetsid-ed is the one thing that wrapper names as out of reach. Where a supervised process starts its own long-lived children, that is a separate link needing the same treatment at its own site. -
The signal is
SIGKILL, and the reason is the state it fires in.PDEATHSIGis delivered only once the parent is already dead, so there is nobody left to wait for a polite exit and nobody to escalate if the child declines. A catchable signal there is the same "cleanup that usually runs" the mechanism exists to replace, and the daemon whose graceful shutdown is being trusted is the one whose graceful shutdown hung a gate for two hours. The polite path is still tried first, by the owningDrop, and those destructors already end inkill()— so this is the same signal, moved to where it cannot be skipped. The cost is the child's own last wishes:test_daemon.pyremoves itsmkdtempconfig directory in afinally:block, and underSIGKILLthat directory stays. Ports, sockets, ptys andflocks are released by the kernel on death, so the loss is a few KiB under/tmpin a run whose parent has already crashed.
And the other half of the same rule, in the destructors. A Drop
that owns an external process must reach the kill on every path.
PyDaemon::drop (leviculum-ffi/tests/support/python_daemon.rs) did
fallible I/O first — query("shutdown"), which expected a JSON
response and got the empty body a mid-shutdown daemon returns — and a
panic in a destructor that is itself running during unwinding is a
non-unwinding panic, so the process aborted before reaching
child.kill() two lines down. That is a second, independent reason the
same daemon leaked. The shape to write is: try the polite shutdown,
ignore every error it can produce, then kill unconditionally. It pairs
with PDEATHSIG rather than replacing it — the kernel covers "the parent
died", this covers "the parent lived and its cleanup threw".
Which sites. Every spawn whose process outlives the call that made
it: the Python TestDaemon and its socat pty pair, PyDaemon, the C
lnsd/lncp/levcat programs in the FFI suite, lnsd and the vendored
rnsd in the mvr, load-test, reverse-RPC and status-parity harnesses,
the instance-conflict holder — and, outside the tests, the
PipeInterface bridge program, which is the one long-lived process the
shipped daemon starts. Deliberately not covered: everything spawned and
awaited inside one call — cc, git, stty, rnstatus, the jl /
jldiff filters, the event-log-helper and port-allocator workers.
Those are counted rather than argued about; see the gate below.
Standing canaries
Every gate on this page carries a permanent pair, checked before it reports anything else:
- A, registry gate: a tagged test deliberately absent from
scripts/pins.txtthat must be reported, and a correctly registered one that must not be. - A, mutation audit: a deliberately weak pin that must survive and a
strong one that must die. Additionally: an unresolvable declared
subject is a hard error,
mutants_generated > 0is asserted per pin, and the equivalent-mutant allowlist carries per-entry justifications and is pinned likescripts/ignored-counts.txt. - B, manifest writer: a fixture of libtest output, parsed before every
wrapped run, holding tests that must land in the manifest — including a
#[should_panic]one, which libtest names<test> - should panicand--listnames plainly (ano_rundoctest is the same shape, run as<test> - compile) — and lines that must not: an ignored test, an indented look-alike a test could print. The counts are then reconciled against libtest's own summary line, so a manifest that disagrees with the run fails the gate. - B, manifest check: a test deliberately in no gate that must always be reported, and one in a gate that must never be.
- B, gate termination: a child that leaks a grandchild holding the
gate's stdout and then exits — the wrapper must report the child's exit
status promptly and must name what it killed — paired with a child
that leaks nothing, which must report no kill at all, because a kill
claimed on every green run is noise and noise is how a real one goes
unread. Two further arms: a child that never exits, which must produce
the named timeout rather than a wait, and an orphan that escapes the
process group with its own
setsid(), which must not delay the verdict past the drain grace. This canary is bounded from outside the thing it tests — a watchdog in the canary kills the grandchild by an argv marker after twenty seconds, so a regression fails in twenty seconds instead of wedging the suite the way the incident wedged the run. Walking into that trap while fixing it would be poor form. It costs about 0.85 s per wrapped gate, which is the price of the only check that can see a wrapper that has stopped terminating: every other gate on this page reports nothing at all when that happens, including the parser canary in the same script, which passes happily while the run it belongs to never ends. - B, supervised spawns: two of them, because the property has two
halves that fail independently. The census
(
scripts/check-supervised-spawns.py) classifies two fixtures before it reports anything about the tree — one holding a bare spawn in each of the four shapes the tree writes them in, all of which must be reported, and one holding a supervised call, a runtimespawn, a string and a comment that mention the words, none of which may be. Without the first arm a classifier that has stopped matching reports a clean tree forever; without the second it reports every line in it. The behaviour (leviculum-std/tests/supervised_spawn.rs) SIGKILLs a parent and requires its child to be gone, paired with the same experiment on a child spawned bare, which must still be alive after the same deadline — otherwise "the child is gone" is satisfied by a child that never started. Both arms are bounded and fail loudly rather than waiting, which is the mistake the incident above was about, and the child's ownPR_GET_PDEATHSIGis checked first as the cheap form: a refactor that drops thepre_execis named in milliseconds. - C, citation guard: a deliberately drifted citation that must be reported. The existing guard has floor asserts against parser rot; the canary is the stronger form.
- C, figure attribution: a fixture page and a fixture doc comment attributing two figures to it — one on the page, one not — plus a version string that must not be read as a figure and an attribution to a page that is not in the tree. Exactly two must be reported. This one is not optional in the ordinary way: #200 was fixed hours before the check was written, so the corpus has no failing case left and a broken trigger would report zero forever against a tree that reads clean.
- C, commit-trailer guard: a message carrying the trailer that must be rejected and one that only quotes it, indented, that must not — both run before the guard reports anything else. leviculum's pinned count of sixteen below-baseline violations is the same canary in stronger form, covering both directions over 1275 real commits on every run; periculum's count is zero and covers only the direction that cannot rot, which is why the pair is in the script rather than only in the baseline file.
- C, commit-trailer authorship arm: the ours/foreign split needs its
own pair, because "no violation found" is what a guard that has stopped
noticing foreign trailers reports forever, and equally what one that has
started calling everything foreign reports forever. Two independent
failures, two checks. The plumbing is verified against git itself on
HEAD — a
--formatthat lost%aewould leave every commit classified as foreign and go green on our own violations from then on, and checking it against a constant in the script would only prove the script agrees with itself. The classification is verified on a synthetic set carrying one hit per declared identity plus one from outside: on a tree whose history happens to hold no foreign hit an inverted comparison would also read green, and a comparison tightened until it stopped recognising the forge's noreply address would be silent without the per-identity cases. What no canary reaches is an identity deleted from the baseline file, since that file is the only statement of who we are — that one is caught by reading the diff, which is why the list lives where a reviewer already looks.
A one-time demonstration at implementation time is not enough. A gate that stops matching — a glob that no longer resolves, a parser that returns nothing — is green forever, which is the defect this page exists to remove.
Composition
A pin is exempt from no gate, and may not appear in the exception
list. Nothing else here makes that true: a pin that is #[ignore]d
satisfies A — its negative control is present and correct — while no
gate observes its green.
What may live in a git hook
A hook is the most tempting place to put a check and the worst place to get it wrong, because it is the one gate that has an off switch every author already knows. So the admission test is narrow:
A git hook may only contain checks that are fast, deterministic, and that fail for a reason the author can fix at that moment. Anything else belongs in a scheduled run or an explicit command.
Three conditions, and a check has to pass all three. Two worked examples from 2026-08-07, both of which were in hooks and neither of which should have been:
- The tier-2 staleness block (
pre-push) failed all three. It was slow by construction — the remedy it named was a 30-90 minute docker run. It was not deterministic in the sense that matters: its verdict depended on a ledger line written by a process nothing scheduled, so the same tree pushed on two days gave two answers for a reason unrelated to the tree. And the author could not clear it at all, at any moment, becausejust extensive— the remedy it printed — does not write the line it read. It blocked for 46 days and 502 commits. The full telling is under Guarantee B, above. post-commitfailed the first and the third. It detachedscripts/run-tier1.sh—just standardunder docker, fifteen minutes warm and forty cold — after every commit. A commit cannot wait forty minutes, and a commit is not a unit anyone wanted tested in the first place: WIP commits, amends and commits mid-refactor each started a run, which is why the runner carried a dirty-flag loop to coalesce them — machinery repairing a granularity that was wrong to begin with. The third condition is the decisive one: when that gate came back red twenty minutes later, there was nothing the author could do about it at the moment of committing, which is the only moment a hook has. It was removed on 2026-08-07 and Tier 1 became an explicitjust standard, once per batch.
The condition that keeps getting skipped is the third one, so state it
positively: a hook fires at a moment the author is still holding the
thing being judged. That is the whole value of the position — a red
fast names a file the author has open. A check whose result arrives
after that moment has passed, or whose remedy is somewhere other than the
work in hand, is not cheaper in a hook; it is only louder.
The corollary, which is the part that bites
A hook that is ever unsatisfiable trains people to bypass the whole
hook. There is no partial override. --no-verify is one flag for all
of pre-push, so the 502 commits that walked past the tier-2 block also
walked past the pipeline lint, Tier 0, mvr and the commit-trailer guard —
checks that were working, that were fast, and that nobody had any
complaint about. The unsatisfiable check did not merely fail to protect
anything. It switched off the ones that did, and then went on being
green-adjacent in the ledger while it did so.
Two consequences worth writing down:
- The cost of a bad hook is paid by the good ones. A check's admission to a hook is therefore not a decision about that check alone, and "it can't hurt to also verify X here" is false as stated.
- Bypassing becomes the habit, not the exception. After the first
few
--no-verifys the flag stops being a considered override and becomes how one pushes. Removing the offending check does not undo that by itself; the habit outlives it, which is why the removal is worth recording where people read rather than only in the diff.
The same reasoning is why .githooks/commit-msg and the Tier 0 half of
.githooks/pre-push stay: milliseconds and ~3 minutes respectively,
deterministic given the tree, and each fails naming a file the author can
open. It is also why neither of them is the enforcement — a fresh clone
has no hooks at all. The forge check is the gate; the hook is the same
rule delivered early, at the moment it is cheapest to obey.
What none of the three reaches
- Whether the check is the right check. A pin can be executed, carry a negative control, cite a live line, and still assert the wrong thing. That is what the reference-first and independent-recomposition rules in Wire Field Semantics are for.
- A negative control is author-chosen at both ends. It proves the pin bites on one break the author thought of — not on the break that will happen. Mutation supplies an adversary who is not the author, which is the other reason the audit is not optional.
- A control written to pass regardless — verifying against a random
key, asserting
is_err()on something unrelated. It satisfies the registry gate; only the audit can see it. - Whether a test asserts anything at all. A body of
let _ = f();can be executed and coupled to its subject. - Prose. Issue comments, commit messages, reports. One line shape in a commit message is now checked (#205, above); nothing a message claims is, and that is the surface that matters.
- Scenario corpora. A Periculum step that asserts nothing is the same defect in another language; its analogue is the delivery bar (Periculum #25).
- Firmware.
leviculum-nrfis excluded from the workspace (Cargo.toml:39) and cross-compiles. All three stop there, and the sx1262 incident lives on the far side.
Where this stands
Codeberg is the source of truth for what is built. At the time of
writing: C covers docs/src/**, the Rust sources of leviculum-core,
leviculum-lxmf, leviculum-lxmf-node and leviculum-std, and (since
2026-09-23) the gate scripts under scripts/ and the Justfile, and its
submodule check runs first in just fast; its bump path is unbuilt.
Both halves of it repair since 2026-09-25: the bare class against the
tree's own history, the identifier-anchored classes against the line map
of a named base (section 7), and both only when run exactly once, after
the last source edit of the pass (section 7a). Since 2026-10-06 both
classes are anchored by the text of their cited line, and the history
fixer repairs a named citation too when the anchor proves where it went
(section 5); Codeberg #307 closed on that.
Two shapes it refused to see until 2026-09-23, both found by reading
rather than by a red gate. A backwards line spec (a-b with a > b)
resolved like any other range, because the length check only looks at the
larger endpoint — while repaired() declines to rewrite one, so the
citation became unrepairable the moment its anchors moved and the drift
report had nothing to offer. It is now refused where it is written. And a
citation into the Justfile was existence-checked only: the corpus
that could have named a recipe did not include shell scripts, and a
substring search for a recipe name is not a definition check — the
mention of just standard in a comment sat one line from the wrong cited
line, inside WINDOW, while the recipe was 140 lines away. A Justfile
citation that names a recipe (the just <recipe> spelling) now has to
land on the recipe's header. Prose attributions in Rust doc
comments are checked for decimal figures and for nothing else. Its
commit-trailer step runs on every push to either repository, and checks
one line shape and no claim. B emits manifests from every
host gate that runs tests, and just complete in the extensive tier
runs the whole workspace by construction, so no ordinary test is outside
the union; the check that reads that union, and the staleness bound that
ages manifests out, are unbuilt. Its wrapper terminates on its child
rather than on the pipe, kills the child's process group and reports what
it killed, and gives up at 1800 s with a named failure. The spawn-site
rule above is built and audited: every long-lived spawn goes through
spawn_supervised, the eleven bare spawns that remain are pinned per
file in scripts/supervised-spawn-counts.txt with a reason each, and
both halves run in just fast. What it does not reach is a process a
supervised child starts for itself — a separate link, covered only by the
wrapper's group kill — and any platform that is not Linux, where the
helper compiles to the Drop path and says so. A is unbuilt.
All three are subject to the rule they enforce. The standing canaries above are the demonstration made permanent, because a one-time one decays.
See also
- Evidence and Honesty — the rule this page mechanises, and the incidents behind it.
- Wire Field Semantics — what a pin must assert, which is a different question from whether it can fail; and the practice of checking the reference before auditing against it, which Guarantee C mechanises.
The randomised pre-transmit window
Two nodes released by the same event reach their radios at the same instant. Carrier sense cannot separate them, because both probe a channel on which neither has keyed yet. The only thing that can is a randomised wait drawn before the probe, and the only question worth arguing about is what that wait is made of.
This page is the study Codeberg #347 asked for: five questions, each
with a number, the reason for it, and the artifact that settled it. It
is written after the fact. The window landed while the study was still
open, in a_directed_packet_is_jittered_on_acquisition_and_free_in_a_burst
(leviculum-std/src/interfaces/rnode.rs:5985)
and the firmware policy behind it, so four of the five questions are
answered by code rather than by argument. The fifth is not, and is
stated as open at the end.
Question 6 was asked separately, as Codeberg #40 against the reference firmware, and it is the same question about the other end of the window: not what the wait is made of but where it starts. It is answered here because the answer is made of the same five artifacts.
What we build it out of
| Term | Value | Where |
|---|---|---|
| Slot | 12 symbol times, clamped to [24, 100] ms, floor 6 ms above 30 kbps | csma_slot_ms (leviculum-core/src/rnode.rs:1005), reached as jitter_slot_ms (leviculum-nrf/channel-access/src/lib.rs:116) |
| DIFS | 2 slots (SIFS is 0) | JITTER_DIFS_SLOTS (leviculum-nrf/channel-access/src/lib.rs:84) |
| Contention window | uniform over 0..=13 slots | JITTER_CW_SLOTS (leviculum-nrf/channel-access/src/lib.rs:88) |
| Owed when | once per channel acquisition, never per packet | channel_released (leviculum-nrf/channel-access/src/lib.rs:262) |
| Discharged by | listening it through, not by being asked for it | jitter_spent (leviculum-nrf/channel-access/src/lib.rs:305) |
1. What is the window sized from?
From symbol time, not from milliseconds. A slot is 12 symbol times, so it tracks the modulation the way the frames it separates do. What that yields, at the PHYs the corpus actually runs:
| PHY | Slot | DIFS + widest draw | Mean wait |
|---|---|---|---|
| SF7/125 kHz | 24 ms | 48..360 ms | 204 ms |
| SF8/125 kHz (project default) | 24 ms | 48..360 ms | 204 ms |
| SF9/125 kHz | 49 ms | 98..735 ms | 416 ms |
| SF10/125 kHz | 98 ms | 196..1470 ms | 833 ms |
| SF12/125 kHz | 100 ms | 200..1500 ms | 850 ms |
| SF5/500 kHz | 6 ms | 12..90 ms | 51 ms |
The table is pinned, not quoted:
the_widest_acquisition_wait_is_a_function_of_the_modulation
(leviculum-nrf/channel-access/src/lib.rs:403).
And there is one slot, not two. The slot is a compatibility figure:
it is the unit a neighbour running the reference firmware counts its own
DIFS and contention window in, so a node whose slot is ten times its
neighbours' either talks over them or starves behind them. The derivation
therefore lives with the airtime and preamble arithmetic in
leviculum-core (csma_slot_ms, leviculum-core/src/rnode.rs:1005) and
is pinned there against a literal float transcription of the reference's
own three lines over every PHY the reference admits
(csma_slot_matches_reference_float,
leviculum-core/src/rnode.rs:3235).
Until Codeberg #147 the firmware had a second one. The CAD retry gate
counted its backoff in max(24, airtime(500)/10) — a tenth of a
500-byte airtime, with a floor and no ceiling — which at BW125/CR4:8 is
120 ms at SF7 and 2687 ms at SF12, 5x and 27x the slot above. It multiplied
into the gate's doubling window, so a board could owe 63 of those slots
on one busy channel where its RNode neighbours owe at most 58 of theirs.
The gate now counts in the same slot the draw does
(leviculum-nrf/src/lora.rs:2130), and the before/after figures are
pinned in csma_slot_is_no_longer_a_tenth_of_a_500_byte_airtime
(leviculum-core/src/rnode.rs:3321).
A millisecond constant would have been wrong in both directions. At SF12 the clamp is what binds and 12 symbol times would be 393 ms, so the ceiling is doing real work; at SF5/500 kHz the floor is what binds and the raw figure is under a millisecond. Between them the value moves by a factor of four.
What confirms this is the right relation rather than a transcription of the reference. #344 measured the reference's own inter-frame gap off the air: never closer than ~80 ms, median 205 ms over 81 gaps. The model above predicts a median of DIFS plus the median draw, 48 + 156 = 204 ms, and a floor of DIFS alone, 48 ms. The median agrees to one millisecond; the floor is a bound the observation respects rather than a prediction it confirms. Our firmware put two packets 15 ms apart before the window existed, which is below the model's floor by a factor of three, and that is the defect the window closed.
2. Does it widen under load?
No, and the reason is that we already widen on something better.
The reference keys four bands to a measured airtime average and walks
the window from 0..14 slots up to 45..59 (update_csma_parameters,
reference/RNode_Firmware/RNode_Firmware.ino:1603). We mirror band 1
only. Under sustained contention our CAD retry gate doubles its own
contention window per busy probe, from CAD_CW_INITIAL to
CAD_CW_MAX (leviculum-nrf/channel-access/src/lib.rs:80), which
reacts to a channel observed busy rather than to an airtime average
computed over the last several seconds. Adding the band escalation on
top would widen the window twice for the same congestion.
This is a deviation under the project's deviation rule: the wire format is untouched, a peer expects no particular window from a neighbour, and reacting to the observed channel is what Priority 1 asks for.
One number in this area is not settled, and it is an off-by-one in the
reference rather than in us. update_csma_parameters assigns
cw_min/cw_max only when the band changes
(reference/RNode_Firmware/RNode_Firmware.ino:1616). A board boots in
band 1 with cw_max declared as CSMA_CW_PER_BAND_WINDOWS, i.e. 15
(reference/RNode_Firmware/Config.h:127), and only an excursion into
band 2 and back rewrites it to band * 15 - 1, i.e. 14. So the
reference has two band-1 windows depending on its history: 15 equally
likely draws as booted, 14 after an excursion. We mirror the second.
The two predict a median gap of 216 ms and 204 ms; the bench measured
205. That is suggestive and not decisive at n=81, and changing
JITTER_CW_SLOTS is a radio-behaviour change, so it stays open.
3. Every access, or only contended ones?
Every acquisition, and no packet inside a burst.
"Acquisition" is the unit, not "packet": a frame that continues a burst
we are already transmitting owes nothing, because the frame before it
served the wait. The wait comes back when the channel is handed back,
which the transmit path does after its post-TX listening window
(leviculum-nrf/src/lora.rs:2298).
Asking for the wait does not discharge it. The wait is spent listening
and the listen returns early on a reception, so a wait cut short by an
incoming frame has de-tiled nothing: the frame that ended it released
every other waiting node at the same instant. Only listening it through
counts (acquisition_jitter_ms,
leviculum-nrf/channel-access/src/lib.rs:289).
What it costs. At the project default PHY the mean cost is 204 ms per acquisition and the worst case 360 ms. A link setup is three acquisitions on each side, so the window adds roughly 0.6 s of median latency to a link establishment at SF8/125 kHz and up to 1.1 s in the tail. At SF10 those become 2.5 s and 4.4 s. That is the price, and it is paid against a collision whose cost is a whole frame plus a retransmission timeout.
There is no "the channel was clear, skip it" shortcut, and there must not be one: the case the window exists for is precisely the one where the channel is clear for both senders.
There is also no packet-type shortcut. A high-priority frame at the head of an idle queue used to key the radio outright, which meant only announces were ever jittered. That is type-awareness in collision avoidance and it is the thing the interface-isolation rule forbids; see Interface isolation. Both halves of it are pinned now, the answering half and the queue jumper's.
Since 2026-09-25, a burst frame does owe a draw — a different one.
The sentence above is about the ACQUISITION wait, and it still holds: a
frame that continues a burst serves no acquisition. What it does owe is
the post-handover hold, and that hold now carries a fresh random term of
its own (tx_hold, leviculum-std/src/interfaces/rnode.rs:1062):
owed = airtime(the packets the modem keys for the frame) + DIFS + (cw_max - 1) x slot
hold = owed + rand(0 ..= TX_HOLD_SPREAD_SLOTS x slot)
The owed part is the guarantee that keeps the modem's queue at one
frame, and the spread only ever sits on top of it. The reason for the
spread is that owed is a constant, identical on every node running the
same PHY: trace 245 (2026-09-25) watched held_ms=863 on both daemons of
an A/B pair at once, and once two ends' key-ups fell inside the modem's
~40 ms carrier-sense rise time of each other, every following frame pair
collided again — three collisions exactly 880 ms apart at staggers 0, 1
and 2 ms, and a 50 KB transfer that retried one part window 155 times
until it timed out. No acquisition draw reaches those frames, because
they are burst continuations and deferral releases. A constant hold is a
metronome; the spread is what stops two metronomes agreeing.
Four slots, because the quantity it has to clear is the modem's
carrier-sense rise time: on SX127x dcd is a live read of
SIG_DETECT|SIG_SYNCED (reference/RNode_Firmware/sx127x.cpp:197-204
@1.85), so it cannot rise before the peer's preamble has been detected,
and trace 228 measured that blind window at ~40 ms at SF7/BW62.5. Four
slots span 0 to 96 ms there against the 80 ms owed. It is counted in
contention slots and not in milliseconds for the same reason every other
term here is: a slot is 12 symbol times, and a typed millisecond would be
true of one carrier only.
The airtime term is the modem's packets, not the host's frame
(Codeberg #430, since 2026-09-26). The host hands the modem one frame
over KISS; the firmware writes a one-byte header in front of it and, past
254 payload bytes, splits it, flushing a packet every time its byte
counter reaches 255 and beginning the next one with the header byte again
(transmit, reference/RNode_Firmware/RNode_Firmware.ino:716-751, the
written == 255 && isSplitPacket(header) arm at line 729). Every packet
therefore pays its own preamble, and add_airtime(written) charges each
one separately. A 508-byte handover — the interface's HW_MTU, and
exactly 2 x 254 — is three packets of 255, 255 and 1 bytes: the last
flush lands on the last payload byte and the closing endPacket() keys
the header alone. 511 bytes and three preambles on the air for 508 bytes
handed over.
Until #430 the hold priced that as one 508-byte packet, so it was below the modem's busy time on every frame over 254 bytes, at the MTU by two preambles and three header bytes:
| PHY | airtime(508) as one packet | as the modem keys it | reported turnaround |
|---|---|---|---|
| SF7/62.5 kHz | 1557 ms | 1713 ms | 2013 -> 2169 ms |
| SF8/125 kHz | 1373 ms | 1529 ms | 1829 -> 1985 ms |
| SF10/125 kHz | 4426 ms | 5045 ms | 6288 -> 6907 ms |
| SF12/125 kHz | 17703 ms | 19852 ms | 19603 -> 21752 ms |
The shortfall could not be absorbed anywhere: every other term of the
hold is spent on something else, the firmware sends no TX-done, and the
spread deliberately sits on top of owed rather than inside it. The
arithmetic is rnode::host_frame_airtime_ms, one function both the hold
and the modem's airtime-ledger expectation are priced from, so the two
cannot disagree about what a handover cost.
The interface reports the top of the band — owed + spread — as its
per-frame turnaround, so the receiver's resource part timeout and the
selftest's drain window size on a hold no draw can step over.
4. What does it do to the ack window, the burst yield, and link setup?
The pre-review audit named the post-TX receive window as the one thing a transmit window could break, and it was right: at one PHY it did. That is Codeberg #423, and it is fixed here rather than described.
The window the transmit path opens after a transmission is one full
single-frame reply airtime plus the peer's turnaround
(post_tx_rx_window_ms, leviculum-nrf/src/lora.rs:865, deciding in
leviculum-channel-access where the peer's window is defined). The right
comparison is against the peer's time to key up, not to finish its
frame: the receiver stops its timeout on preamble detect and then runs
to packet completion regardless of length
(SET_STOP_RX_TIMER_ON_PREAMBLE, leviculum-nrf/src/sx1262.rs:760), so
a reply that starts inside the window is heard whole even when it ends
outside it.
Until #423 the turnaround term budgeted the peer's DIFS and nothing else
— and in the wrong slot at that, two slots of the CAD gate's backoff
slot, max(24, airtime(500)/10), rather than of the 12-symbol slot a
contention window is actually drawn in. It was written before the peer
had a window to draw, and it stayed that way through #149 and #347. That
backoff slot is itself gone since #147 (see §1): there is now one slot on
the board, so the two terms cannot disagree again.
| PHY | Post-TX window | before #423 | Peer's widest wait | Covered |
|---|---|---|---|---|
| SF7/125 kHz | 876 ms | 666 ms | 360 ms | yes |
| SF8/125 kHz | 1188 ms | 1094 ms | 360 ms | yes |
| SF9/125 kHz | 2127 ms | 1866 ms | 735 ms | yes |
| SF10/125 kHz | 3948 ms | 3338 ms | 1470 ms | yes |
| SF12/125 kHz | 10000 ms (clamped) | 10000 ms | 1500 ms | yes |
| SF7/250 kHz | 680 ms | 396 ms | 360 ms | yes |
| SF7/500 kHz | 582 ms | 270 ms | 360 ms | yes, was no |
| SF5/500 kHz | 231 ms | 189 ms | 90 ms | yes |
The old window was airtime-derived and the wait is slot-derived, so they diverged exactly where the slot's 24 ms floor stops tracking a shrinking airtime: at wide bandwidths and low spreading factors. At SF7/500 kHz the listening window closed 90 ms before the peer could be expected to have keyed, and the only reason the other rows held is that the airtime term was large enough to absorb a missing contention window by accident — SF7/250 kHz held by 36 ms.
Missing the reply there needed the peer to draw high and us to draw low in the same exchange, since our own next acquisition owes its jitter immediately afterwards and that wait is also spent listening: the composite listen was 270 + 48..360 ms. That is a probability, not a guarantee, and a probability is not what an ack window should rest on.
What the term is now is the peer's whole wait —
widest_acquisition_wait_ms, the same DIFS-plus-window this crate hands
our own transmit path — so the coverage is structural rather than
arithmetical luck: the quantity the window has to cover is a summand of
the window, and no PHY, preamble override or clamp can take it back out.
The window is computed in post_tx_rx_window_ms
(leviculum-nrf/channel-access/src/lib.rs:169) and its rows are pinned
where they are computed, from leviculum-core's own airtime rather than
transcribed: the_post_tx_window_is_one_reply_plus_the_peers_whole_turnaround
(leviculum-nrf/channel-access/src/lib.rs:526) for the table, and
the_post_tx_window_covers_the_peers_keyup_at_every_modulation
(leviculum-nrf/channel-access/src/lib.rs:552) for the property it is one
sample of, over five bandwidths x SF5..SF12 x CR4/5..4/8. The pre-#423
arithmetic is kept as an assertion of its own,
budgeting_only_the_peers_difs_closed_the_window_before_it_could_key_up
(leviculum-nrf/channel-access/src/lib.rs:490), so a revert of the term
goes red at SF7/500 kHz rather than silently.
The cost is the second column against the third: every PHY but SF12/125
kHz, where the ceiling already bound, listens longer now — by 210 ms at
the bench PHY and 610 ms at SF10/125 kHz. It is spent
only when the channel stays silent through the whole window — rx_window
returns on the first frame — and it is spent at a yield, where the peer's
turn is the point. A window that ends before the peer's turn can start is
not a cheaper window, it is a missed reply and a retransmission timeout.
burst_should_yield (leviculum-core/src/rnode.rs:1813) is unaffected:
it bounds a burst by frame count and accumulated airtime, and the window
is spent before the burst starts rather than inside it.
5. Do we listen during the wait?
Yes, and it is the whole reason the wait is safe to impose. A wait
that went deaf would trade a collision for a missed frame, which is the
same loss at the layer that counts. The transmit path arms the receiver
for the drawn duration and reports back what it actually listened
through, and a reception that cuts the wait short leaves the debt
standing (leviculum-nrf/src/lora.rs:2024).
6. Does the window's floor matter?
No, and that is worth stating because it is the remedy an observer
reaches for first. Codeberg #40 recorded the reference's light-traffic
window as cw_min = 0 (CSMA_CW_MIN,
reference/RNode_Firmware/Config.h:109) and proposed raising the floor to
two or three slots "so a competing node's preamble falls into the other's
CAD window". Our draw mirrors that window, so the proposal reads as a
proposal about us too. Three things about it do not hold, and none of them
needs a new measurement.
A term both ends owe cancels out of what separates them. Two nodes
released by the same event are separated by the difference of their
waits, not by either wait, so a constant added to every wait moves the
whole distribution of waits and leaves the distribution of separations
exactly where it was. That is not an argument but a pin, and it was
already made for a different proposal of the same shape: a deferral of one
frame airtime on every answer, refused in
direction_3_a_deferral_that_is_a_constant_cancels_and_buys_nothing
(leviculum-std/tests/mvr/two_responders_overlap_inside_one_airtime.rs:503),
which asserts the deferred census equal to the baseline one count for
count and says in place that a constant common to both cancels the way
DIFS already does. A floor is that constant, spelled in slots instead of
airtimes.
What a floor does buy is latency: at SF10/125 kHz two more slots cost
every acquisition another 196 ms out of the budget priced in question 3.
Its one second-order effect points the same way rather than the other: a
longer wait is longer exposed to a third node's frame arriving inside
it, and a wait cut short that way re-anchors both ends on that frame at
the same instant — the same mechanism that makes deliberate carrier sense
a losing direction here, priced in
direction_4_carrier_sense_re_anchors_the_pair_and_doubles_the_odds
(leviculum-std/tests/mvr/two_responders_overlap_inside_one_airtime.rs:551).
Our pre-TX wait has never been able to be zero anyway. The draw's own
floor is zero, but a floor of zero draws is not a floor of zero
milliseconds: every acquisition also owes DIFS unconditionally, two slots
(JITTER_DIFS_SLOTS, leviculum-nrf/channel-access/src/lib.rs:84), which
is the narrowest column of question 1's table — 48 ms at the bench PHY,
12 ms at SF5/500 kHz, 200 ms at SF12
(the_widest_acquisition_wait_is_a_function_of_the_modulation,
leviculum-nrf/channel-access/src/lib.rs:403), and pinned again through
the draw itself for a thousand seeds in
boot_owes_jitter_and_the_draw_is_difs_plus_a_bounded_window
(leviculum-nrf/channel-access/src/lib.rs:593). The reference is no
different: tx_queue_handler
(reference/RNode_Firmware/RNode_Firmware.ino:1623) waits difs_ms
(reference/RNode_Firmware/Config.h:119), also two slots, and waits it
while sensing: a medium that goes busy clears difs_wait_start, so the
DIFS restarts from the top, while the contention countdown only freezes
(cw_wait_passed survives and is reset at the flush, not at the
interruption). cw_min = 0 means the contention term can be zero, not
that a node transmits the instant it is handed a frame.
And the colliding set is not "both drew zero". #40 priced the risk at
1/15², which is the chance of that one pair. Two ends that draw the same
value, whichever value it is, end their waits in the same slot, so the
colliding set is every equal pair and its size is one over the number of
draws: 1/15 in the reference as booted and 1/14 after an excursion
(question 2), and 1/14 for us. That figure is the one the arms are
measured against —
FrameClass (leviculum-std/src/interfaces/rnode.rs:563) takes a
same-class pair to 1/56 over arm 3's fourteen counts — and to 29/784,
about 1/27, over arm 4's seven, which is what a shorter span costs — and
states in the same place that the count alone leaves it at 1/14, i.e. that
a change which does not increase the number of distinguishable outcomes
buys nothing. A floor does not increase it.
Width, quantisation against the frame, and per-identity pinning do, and
which of those wins is the open A/B, not this.
None of it is reachable from the host in any case, which is what #40
concluded and still holds: the reference's whole command set carries one
CSMA opcode and it is a read-only stat, CMD_STAT_CSMA
(reference/RNode_Firmware/Framing.h:48); the nearest thing to a setter,
CMD_DIS_IA (reference/RNode_Firmware/Framing.h:67), switches
interference avoidance off and never touches the window. The only window
we can change is our own.
What is still open
- The post-TX receive window's turnaround term now budgets the peer's whole wait (Codeberg #423, above), and what is left open is the term beside it: whether the reply airtime belongs in the window at all. Nothing about coverage needs it — the peer's key-up is what the window has to reach — so dropping it would take SF12/125 kHz from the 10 s clamp to 1600 ms and SF10 from 3948 to 1570. That is a large change to the listening duty cycle at slow PHYs, it interacts with the peer-turn yield the window is doubled into (#23 Bug B) and with the receiver's resource part timeout, and it is a radio-behaviour change that wants a rig measurement with the other medium switched off. #423's own acceptance is the same rig scenario: SF7/500 kHz, LoRa under test, Bluetooth off on the boards.
JITTER_CW_SLOTSmirrors the reference's post-excursion band-1 window, 14 draws, where a freshly booted reference uses 15. See question 2.CSMA_DIFS_MSandCSMA_MAX_CW_MS(leviculum-core/src/rnode.rs:1041) are millisecond constants pinned to a 24 ms slot and are therefore wrong at every SF above 8. Their only consumer iscompute_spacing_ms(leviculum-core/src/rnode.rs:1216), which has no caller: the host interface prices the same shape from the modem's reported slot instead (tx_hold,leviculum-std/src/interfaces/rnode.rs:1062). Nothing is broken by them today and something would be by the next caller.
Self-hosted infrastructure
A plan to give Leviculum and Periculum a second home on our own server,
workhorse.de (public name leviculum.network), reachable over both the
clearweb and the Reticulum network, running permanently in parallel
with Codeberg rather than replacing it.
This was written as a design record for later implementation. Part of it has since been built — the clearweb download path, in the nightly this project already runs. The server itself is not provisioned. The section below says exactly which part is in the tree, so that the rest of this record can go on being read as what it was: a plan.
What of this is already built
Written 2026-08-17, when none of it existed. On 2026-09-18 the sending half of the clearweb download path landed (63140e56, fixed by 7a5b7940), under the nightly pipeline's own issues rather than this one, so these parts of the record are now in the tree:
scripts/publish-site.sh— the sender. It tarsdist/plus the build id and pipes it over ssh to the receiver. The host key is pinned and there is no accept-on-first-use, because an unattended job cannot recognise a host it has never seen. Unconfigured — none ofSITE_SSH_TARGET,SITE_SSH_KEY,SITE_SSH_HOST_KEYset — it prints a banner naming what is standing still and exits 0; partly configured is a mistake and fails.packaging/site/lev-receive-nightly— the receiver, meant to run on the server as the forced command of an ssh key. It verifies the whole upload before a byte of it is reachable under a public URL, publishes to/var/www/leviculum/releases/nightly/<build-id>/, and movesnightly/latestbyrename(2)once the new build is complete on disk, keeping the lastKEEP(default 14) builds..woodpecker/nightly.yml— thepublish-sitestep that runs the two. It runs last, after the forge publish, and withstatus: [success, failure], so a forge outage — the night the second target matters most — does not take the independent target down with it.
Everything else here is unbuilt: the server, the bare git repo, rngit,
issues/, AGENTS.md, stagit, the issue renderer and the sync job.
The five open questions at the end are all still open, and the three ssh
values are not configured, so the step that exists publishes nothing.
Motivation
Today the project is single-homed: code, release artifacts and the issue tracker all live with one hosting provider. Any single provider is a single point of failure — an outage, an account dispute, or a change in a service's terms can remove all three at once. The response is to stop being single-homed: stand up an independent home we fully control, keep both in sync indefinitely, and treat the loss of either as a non-event because the other is already live and complete. As a bonus, a self-hosted node is reachable over Reticulum, which clearweb forges are not — a natural fit for a Reticulum stack.
Principles
- Simplest thing that works. No CI framework, no forge platform. A
bare git repo, a static web server, a shell script, and
rngit. - Both networks, one source. Everything a user needs — code, issues, releases — is reachable over clearweb and Reticulum. Wherever possible this is achieved by the data being in the git repo, which is clonable over both, not by running a service twice.
- No single point of failure at the code. The authoritative copy is
an ordinary bare git repo.
rngit— young, "not tested extensively in the wild" per the RNS manual — sits beside it as a Reticulum remote, never underneath it. - AI-friendly by construction. An agent with no prior knowledge
reads one
AGENTS.mdand knows how to clone, file an issue, and cut a release, because the patterns are plain files and standard git, not a bespoke API. - Reticulum-compatibility is Priority 1. The Reticulum side uses
Mark's own
rngit, so we interoperate with the ecosystem by default.
Architecture
┌──────────────────────────────────────────┐
│ workhorse.de (leviculum.network) │
│ │
clearweb ─────┤ nginx ──► /srv/git/leviculum.git (bare) │ ◄─── SSH push
git clone │ └─► /srv/www (static: issues, │ (maintainers)
downloads │ release artifacts, git browse)│
│ │
Reticulum ────┤ rngit ──► same bare repo as rns:// remote │
rns:// clone │ ├─► rngit release (signed nightly) │
NomadNet page │ └─► serve_nomadnet = yes (browse) │
lncp/rncp │ │
│ nightly.sh (systemd timer): build+test, │
│ on green → publish to both fronts │
│ sync.sh (timer): mirror to/from Codeberg│
└──────────────────────────────────────────┘
The bare repo is the single source of truth for code. Issues live
inside it as plain-text files, so they travel with every clone on
either network. Codeberg is kept as a synchronised parallel home.
Component 1 — Git upstream (dual-network, permanently mirrored)
The authoritative copy is /srv/git/leviculum.git, a bare repo.
- Maintainer push: over SSH (
git push origin master). SSH key auth, no web push. The only write path to the code. - Clearweb read/clone: nginx serving the bare repo, smart HTTP via
git-http-backend(a CGI shipped with git — no extra software), or dumb-HTTP (git update-server-infoin a post-receive hook + static serving) for zero CGI. Anonymous clone/pull only. - Reticulum read/clone:
rngitwith[repositories] public = /srv/gitexposes the same repo atrns://<hash>/public/leviculum. - Codeberg stays a mirror: every maintainer pushes to both remotes
(
git remote set-url --add --push origin), so Codeberg and workhorse hold identical history. Because git commits are content-addressed, the two are byte-identical when in sync; there is no divergence to reconcile for code.
Periculum is a second repo in the same group, treated identically.
Component 2 — Issues as plain text in the repo
Issues are Markdown files committed into the repo. They are part of every clone, on both networks, with no running issue service and no API.
issues/
open/ 0042-lxmf-msgpack-unbounded-recursion.md
closed/ 0038-lrproof-hop-asymmetry.md
Each file is front-matter plus a Markdown body:
---
id: 42
title: LXMF msgpack skip recurses without a depth budget
labels: [bug, compat, priority:low]
state: open
created: 2026-08-17
author: <name or Reticulum identity hash>
---
Body in Markdown, cite code as path:line. Comments are appended below a
`---` rule, each with an author + date line.
- State is the directory (
open/vsclosed/);ls issues/openis the backlog, a state change is onegit mv, and it diffs cleanly. - Search is
grep -r. No index, no database. - Numbering:
scripts/new-issue.sh "title"picks the next free number by scanningissues/. Collisions are rare (nearly all issues originate from our own instances) and resolve at merge; switch to short hash IDs only if distributed creation ever makes that a real problem. - Attribution, if wanted, comes from signed git commits, not a per-issue crypto layer.
rngit work was considered and rejected for the issue store: it keeps
work items as msgpack in a side directory, writable only with rngit
installed and not legible at rest. Plain files win on both-network reach,
AI-transparency, and zero-service — at the cost of built-in signing,
which git commit signing replaces.
Component 3 — Nightly builds (a script, not a framework)
nightly.sh, run by a systemd timer:
1. fetch + checkout the tip of master into a clean build worktree
2. cargo build --release (both stacks)
3. run the tests. If red -> stop, notify, publish NOTHING.
4. on green only:
- stamp nightly-YYYYMMDD-<shortsha>
- copy artifacts to /srv/www/releases/nightly/ (clearweb download)
- repoint /srv/www/releases/nightly/latest (symlink)
- keep the previous N builds = last-known-good (add-then-swap, never
destroy-then-upload)
- rngit release <repo> create nightly-YYYYMMDD-<sha>:./dist (signed)
5. regenerate the static clearweb views (Component 5)
This designs out three review findings at once: PUB-0016 (nightly
shipped from an untested/red tree — green is now a precondition),
PUB-0005 (destroy-before-upload with no failure detection — add-then-
swap keeps history), and PUB-0011 (no rollback — retained builds plus
the latest symlink are exactly that).
Where this stands. All three are delivered, but not by the shape
above. The script-and-timer was written as if there were no nightly; what
landed put the same three properties on the one this project already has,
.woodpecker/nightly.yml. Its first step is the test gate and a non-zero
exit there terminates the workflow, so build, package and publish are
never reached with a red commit — PUB-0016. PUB-0005 and PUB-0011
are the receiver's: packaging/site/lev-receive-nightly points latest
only after the new build is complete on disk, and keeps the previous
builds. A standalone nightly.sh on the server buys none of the three
any more. It would be needed only if the build itself ever has to leave
the forge's runners, which is a separate decision and is not made here.
Component 4 — Downloads over both networks
- Clearweb: nginx serves
/var/www/leviculum/releases— the defaultRELEASES_ROOTofpackaging/site/lev-receive-nightly— at stable URLs such ashttps://leviculum.network/releases/nightly/latest/leviculum-nightly-amd64.deb. The build side of this is done and runs every night; the nginx side is not. - NomadNet page:
rngit serve_nomadnet = yesalready exposes a release list, file browser, commit history and refs to any NomadNet client; its Micron templates live in~/.rngit/templates/. - Reticulum file pull:
rngit release <repo> fetchis primary (it verifies the Ed25519 manifest signature before writing any bytes);lncp/rncp --fetchremain available for raw file pulls.
Component 5 — Clearweb read-only views
Two static views, regenerated by the nightly run and by a post-receive hook, so nothing dynamic runs on the clearweb side:
- Code browsing:
stagitrenders static HTML into/srv/www/code/. No CGI, no daemon. (cgitonly if dynamic browsing is later wanted.) - Issue browsing: a small script renders
issues/**/*.mdto a static/srv/www/issues/index and per-issue pages. Read-only on clearweb; writing an issue is a git commit over either network.
Component 6 — Keeping Codeberg and workhorse in sync
Code sync is trivial and symmetric: maintainers push the same commits to both remotes, and content-addressing guarantees they match.
Issue sync is the one genuinely hard part, because Codeberg keeps issues in its own database (web UI, external contributors) while our issues are plain files in the repo. Rather than a true bidirectional merge (which invites conflicts), pick one side as the source of truth and mirror the other one-way; this decision is deferred but the two shapes are:
- Repo as source of truth (matches the long-term single-home design):
plain-text issues are authoritative, a
sync.shjob pushes creates/edits/closes to Codeberg via the Gitea API for visibility. External Codeberg-only comments are pulled back on a best-effort basis and appended to the file. Simplest to reason about; Codeberg becomes a read-mostly shopfront. - Codeberg as source of truth (matches keeping Codeberg the primary
day-to-day tracker): humans file and discuss on Codeberg's web UI, and
sync.shexports the issue set to plain files in the repo so Reticulum users can read (not write) them. Loses write-from-Reticulum until the repo is ever made authoritative.
The choice is a later detail decision. Whichever side is chosen, the sync is a one-way export/import script on a timer, not a live service, and the Gitea API access already used for issue curation is sufficient.
AI-friendliness — the AGENTS.md contract
A single file at the repo root tells any agent everything in one screen:
how to clone over each network, that issues are Markdown under
issues/open and issues/closed, how to file one (scripts/new-issue.sh,
edit, commit, push), how to close one (git mv to closed/ with a
note), and that releases are cut only by scripts/nightly.sh. Because
the substrate is plain files and standard git, an agent needs no
project-specific API knowledge — the patterns are the ones every agent
already knows.
Where this stands. Nothing in the paragraph above is in this tree yet:
no AGENTS.md, no issues/ directory, no scripts/new-issue.sh, and no
scripts/nightly.sh — that last one is the Component 3 design whose own
Where this stands note records that its three properties landed on
.woodpecker/nightly.yml instead. A contract written today would name that
pipeline as what cuts releases. Said here because the sentence reads as a
description of the repository and is a description of the workhorse.de
design, and because the citation guard cannot catch it: it names files in
prose rather than as path:line citations, so a reader is the only check
there is.
Rollout
- Stand up (no announcement). Provision workhorse.de: bare repo,
nginx,
rngit,AGENTS.md,stagit, the issue renderer. Maintainer instances add workhorse as a second push remote. The download target needs no new software on the server —lev-receive-nightly, a receiving user and a key with it as forced command — and then the three Woodpecker secrets, created in the UI before the pipeline names them, in the order.woodpecker/nightly.ymlspells out. - Parallel run, indefinitely. Both homes live; every push goes to
both. Import the current Codeberg issues into
issues/once, then runsync.shon a timer. Exerciserngitclone/pull/release over real Reticulum. This is the steady state, expected to continue indefinitely. - If either home ever becomes unavailable: the other already holds
everything, so nothing is lost. If the self-hosted side is to become
the sole home, point the README solely at
leviculum.networkand carry on — no scramble, because both have been live in parallel all along.
What this resolves
- Single-provider dependency for code, releases and issues — the reason for a second, self-controlled home.
- Untested nightly builds, destroy-before-upload publishing, and no
rollback — designed out by
nightly.sh(Component 3). - Build docs that point at a path the musl target never produces —
folded into the new README and
AGENTS.md.
Open questions
- Server baseline: current state of workhorse.de (OS, nginx, an existing Reticulum instance, public IP, TLS via certbot).
- Reticulum transport to the server: TCP interface over the
clearweb, or a real RF path — affects
rngitreachability and announce cadence. - Issue sync direction: which side is source of truth (Component 6), and the field mapping for the one-time Codeberg import.
rngitacceptance bar: what track record over Reticulum is required before it is relied on even as a secondary path.- Signing: SSH vs GPG commit signing for attribution; whether release signing uses a dedicated identity.
Licensing and third-party notices
Leviculum is licensed under the GNU Affero General Public License,
version 3 or later. That is the licence of the work, it is what LICENSE
at the root of the repository contains, and everything below is additive
to it rather than a qualification of it.
Why a second file exists
Our binaries are musl-static by default. A statically linked binary does not merely sit next to its dependencies, it contains them, and a large part of the Rust ecosystem we depend on is MIT or BSD-3-Clause. Both licence families require their copyright line and their permission text to accompany a binary distribution, not only a source one. Apache-2.0 §4 says the same thing in more words.
Until Codeberg #288 the only licence text in any published artifact was
our own AGPL. Every .deb, every userspace tarball and the lnflash
bundle were therefore short a notice for each such crate — unintentional,
and awkward to repair after publication because published artifacts stay
published.
THIRD-PARTY-NOTICES at the repository root is that notice. It ships
inside every artifact:
| artifact | path |
|---|---|
leviculum, lnomad, lblogd .deb | /usr/share/doc/<pkg>/THIRD-PARTY-NOTICES |
| userspace tarballs | doc/THIRD-PARTY-NOTICES |
| lnflash bundle | THIRD-PARTY-NOTICES beside the binary |
| source tarball | tracked file, included by git archive |
cargo-deb does write a copyright file of its own, but it is derived
from our Cargo.toml metadata and describes our code alone, so its
presence never closed this.
How it is produced
Generated, never hand-curated: a list maintained by hand drifts from
Cargo.lock the first time somebody runs cargo add.
scripts/gen-notices.py drives cargo-about, pinned at 0.9.2 by
scripts/install-ci.sh. cargo-about reads each crate's own LICENSE
files out of the cargo cache, so the copyright lines in the output are
the crates' real ones rather than a template with the names left blank.
The script decides layout and ordering, drops our own AGPL crates in
favour of the leading section, and refuses to emit a file if any crate
was classified without a licence text.
Two dependency graphs go in, because leviculum-nrf is a separate
workspace with its own lockfile and cannot be reached from the root
manifest:
- Part 1, host binaries —
lnsd,lnstest,lncp,lnstatus,lnomad,lblogd,lnflash, resolved for both musl triples. - Part 2, the LNode firmware image — the
t114build the lnflash bundle carries, resolved forthumbv7em-none-eabihf.
Everything runs --frozen. Offline is not only hygiene: with network
access cargo-about falls back to fetching licence files from a crate's
upstream repository, and output that depends on whether the machine had
connectivity cannot be diffed.
Two configuration decisions are worth naming.
Dual-licensed crates. The common MIT OR Apache-2.0 is taken under
Apache-2.0. Either satisfies us; Apache-2.0 is simply explicit about what
a binary distribution must carry, where MIT leaves it to be reconstructed.
The preference is the order of the accepted list in about.toml.
Nordic's SoftDevice bindings. nrf-softdevice-s140 declares a
license-file and no SPDX license field, so cargo-about dropped it
with a warning — and, worse, its file crawler had been reading Nordic's
five-clause text as plain BSD-3-Clause, which it is not: clauses 4 and 5
add restrictions BSD does not have. leviculum-nrf/about.toml now
clarifies it as LicenseRef-Nordic-5-Clause with the licence file's
checksum pinned, and the generator reproduces that text verbatim.
This is distinct from the SoftDevice blob, which travels through the
lnflash bundle as a separate .hex with its own licence agreement beside
it and is never linked into anything. The bindings are compiled into our
UF2; the blob is not.
How it is kept honest
just notices-guard regenerates the file and diffs it byte for byte
against the checked-in copy. It runs in just fast, so it is on the
pre-push path: adding a dependency without regenerating turns the push
red, in the same session that added it.
The file is checked in rather than generated at artifact-build time on
purpose. The .deb, tarball and bundle builds then copy a tracked file
and need neither cargo-about nor a network, and the freshness question
lives in one place instead of four build scripts.
To fix a red guard:
just notices # regenerate
git add THIRD-PARTY-NOTICES
Two further checks assert the file actually arrives:
scripts/verify-deb-packaging.sh looks for it in each .deb, and
scripts/lnflash-bundle.sh asserts it — and both of its section
headings — inside the finished tarball rather than in the staging
directory, because the tarball is what ships.
What is not covered
leviculum-ffi installs through its own make install and is not part
of the nightly publish set, so no notice file is installed alongside
libleviculum. A consumer linking the static archive takes on the same
obligation; when that library starts being published as an artifact it
needs the same treatment.
Installation
Requirements
- Rust stable toolchain
- Git
Optional, depending on what you want to test:
- Python 3 (for interop tests)
- Docker (for integration tests)
- 2-4 RNode modems via USB (for LoRa integration tests)
No system C libraries are required. All cryptography is compiled from Rust source.
Debian/Ubuntu setup
# Rust toolchain
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source $HOME/.cargo/env
# Interop tests
sudo apt install python3
# Integration tests
sudo apt install docker.io
sudo usermod -aG docker $USER
# LoRa tests and embedded firmware (USB serial access)
sudo usermod -aG dialout $USER
Build from source
git clone https://codeberg.org/Lew_Palm/leviculum.git
cd leviculum
cargo build --release --bin lnsd --bin lnstest --bin lncp
The workspace pins x86_64-unknown-linux-musl as its build target (see the
comments in .cargo/config.toml for why), so the binaries are in
target/x86_64-unknown-linux-musl/release/, not target/release/.
That pin names an architecture and cargo has no per-host conditional for
it, so on an arm64 host (a Raspberry Pi, for instance) the build succeeds
and produces x86_64 binaries that cannot run. Set your own target first —
the build then warns no more, and the binaries are in
target/aarch64-unknown-linux-musl/release/ instead:
rustup target add aarch64-unknown-linux-musl
export CARGO_BUILD_TARGET=aarch64-unknown-linux-musl
The .deb packages are published for arm64 as well and need none of this.
Verify the build; the output carries the version and the build commit:
./target/x86_64-unknown-linux-musl/release/lnsd --version
Running the daemon
./target/x86_64-unknown-linux-musl/release/lnsd
The daemon resolves its config directory in the same order as Python
Reticulum (configdir, RNS/Reticulum.py:231), implemented in
default_config_dir (leviculum-std/src/config.rs:1063):
/etc/reticulum— if/etc/reticulum/configexists. This is what the.debpackage sets up.~/.config/reticulum— if that directory'sconfigexists.~/.reticulum— the fallback, and where a source build with no prior config ends up.
Add -v (debug) or -vv (trace) for more verbose logging.
Development
Cargo aliases
Common workflows are available as cargo aliases (defined in .cargo/config.toml):
| Command | What it does |
|---|---|
cargo test-core | Run all leviculum-core unit tests |
cargo test-std | Run all leviculum-std unit tests |
cargo test-interop | Run interop tests against Python Reticulum |
cargo lint | Run clippy on all crates |
cargo fmt --all -- --check | Check formatting |
Test levels
Tests are organized by what they require:
Unit tests -- just Rust, no extra dependencies:
cargo test-core
cargo test-std
Interop tests -- require Python 3 and the vendored Reticulum:
git submodule update --init reference/Reticulum
cargo test-interop
Scenario tests -- multi-node scenarios live in the sibling
periculum checkout, expected at
../periculum. They require Docker and pre-built release binaries:
cargo build --release --bin lnsd --bin lnstest --bin lncp --bin lora-proxy
periculum run ../periculum/conformance ../periculum/regression
LoRa integration tests -- require physical RNode modems connected via USB:
LoRa scenarios live in periculum's hardware/ corpus. They exercise real
over-the-air transfers between RNode radios running Reticulum firmware, and
between those and LNodes running leviculum's own firmware. A scenario names
the set of boards it needs (profile = "rnode_pair", "rnode_quad",
"rnode_lnode_pair", ...), which is resolved against the bench description
in periculum's rig.toml. A scenario the bench cannot serve reports
SKIPPED_INFRA naming what was missing — never a failure. The per-profile
scenario counts are in periculum's hardware/README.md.
Hardware setup:
- Connect RNodes via USB. They appear as
/dev/ttyACM0,/dev/ttyACM1, etc. - Your user must be in the
dialoutgroup:sudo usermod -aG dialout $USER - Override device paths with environment variables if needed:
LEVICULUM_RNODE_0=/dev/ttyUSB0 LEVICULUM_RNODE_1=/dev/ttyUSB1
Running LoRa tests:
# See what the corpus holds, and what this bench can serve
periculum list ../periculum/hardware
periculum devices --probe
# Single scenario
periculum run ../periculum/hardware/lora_link_rust.toml
# The whole hardware corpus
periculum run ../periculum/hardware
# Override radio parameters (bandwidth in Hz)
LORA_BANDWIDTH=125000 periculum run ../periculum/hardware/lora_lncp_push.toml
Each LoRa test must pass on all three bandwidth profiles (62.5 kHz, 125 kHz,
250 kHz). The TOML files define 62.5 kHz; use LORA_BANDWIDTH to switch.
Some tests use the lora-proxy binary for fault injection (dropping frames
to test retransmit recovery). Build it before running proxy tests:
cargo build --release --bin lora-proxy
Embedded cross-compilation
Embedded targets are not downloaded automatically. Install them when needed:
rustup target add thumbv7em-none-eabihf # nRF52840
rustup target add thumbv6m-none-eabi # RP2040
cargo check-nrf52
cargo check-embedded
Before submitting changes
cargo fmt --all -- --check
cargo lint
cargo test-core
cargo test-interop
Configuration
lnsd reads the same INI-style configuration file as Python Reticulum
(rnsd). The format is a drop-in: a config that rnsd accepts, lnsd
accepts, and the two share the shared-instance IPC socket so client
tools (rnstatus, rncp, lnstest diag, Sideband, Nomadnet) attach to
either daemon without changes. Keys lnsd does not implement are
tolerated, not rejected — an unknown key never makes lnsd refuse a
config a current rnsd would load (ini_config.rs:435-440).
File location and lookup order
Pass an explicit config directory with --config DIR (lnsd.rs,
-c/--config). With no flag, lnsd resolves the directory using the
same order as Python Reticulum (config.rs:701-718):
/etc/reticulum— if/etc/reticulum/configexists$HOME/.config/reticulum— if that directory'sconfigexists$HOME/.reticulum— fallback, used even if absent
The config file is always named config inside that directory
(config.rs:720-722). The storage directory defaults to
<config_dir>/storage and can be overridden with --storage
(lnsd.rs, -s/--storage).
This order is why the Debian package can install a system-wide config
under /etc/reticulum and have Python clients connect to the live
daemon with no extra flags (config.rs:707-711).
INI vs TOML detection
lnsd accepts both the Python INI format and native TOML. Detection is
by content, not just extension (config.rs:662-689):
- An explicit
.tomlextension forces TOML. - A file containing
[[(the ConfigObj subsection marker Python uses for interfaces) is parsed as INI. - Otherwise TOML is tried first, then INI as a fallback.
In practice your config file uses the Python INI form shown
throughout this page. Boolean values accept Yes, yes, True,
true, 1, on (and their false counterparts); anything else is read
as false (ini_config.rs:929-940).
A file with a syntax error is refused
A line that is neither a section header nor a key = value pair is a parse
error, and lnsd exits non-zero without starting rather than reading the
rest of the file as nothing (ini_config.rs:207-209). The error names the
line number and the line, and where the format is ambiguous it reports both
the INI and the TOML verdict (config.rs:949-956):
$ lnsd --config /etc/reticulum
lnsd: configuration error: Failed to parse config /etc/reticulum/config: not
valid Reticulum INI (Invalid line 1 ('[reticulum'): matched as neither section
nor keyword) and not valid TOML (...)
This is the reference's behaviour, not a house rule: ConfigObj raises
Invalid line ('[reticulum') (matched as neither section nor keyword), and
rnsd logs Could not parse the configuration at <path> and exits 255
(RNS/Reticulum.py:330-333). Every bracket shape ConfigObj accepts still
loads here, including a header with a trailing comment ([reticulum] # note)
and [reticulum = x, which ConfigObj reads as a key rather than a header
(ini_config.rs:44-62).
What is NOT refused, because the reference does not refuse it either: an
interface whose type we do not implement is skipped with a warning and the
daemon runs with the rest (ini_config.rs:317-323), the same way rnsd logs
Could not locate external interface module and carries on. A key in a
section lnsd does not read is likewise kept out of the config and logged
(ini_config.rs:178-191).
The [reticulum] section
Core daemon settings. Every key below is parsed in
ini_config.rs:389-571; defaults come from config.rs:394-430.
| Key | Type | Default | Meaning |
|---|---|---|---|
enable_transport | bool | true | Route announces and serve paths for other peers. lnsd defaults this to true (it is a daemon); the Python library default is false. (config.rs:27-28, 202) |
use_implicit_proof | bool | true | Use implicit proof for link identification. (config.rs:30-31, 203; ini_config.rs:397-399) |
share_instance | bool | false | Listen on the abstract Unix socket \0rns/<instance_name> for local clients. Required for lnstest diag, rnstatus, Sideband etc. to attach. (config.rs:37-40, 205; key share_instance → shared_instance, ini_config.rs:394-396) |
instance_name | string | default | Names the shared-instance socket: \0rns/<instance_name>. Use a unique name to run two daemons side by side. (config.rs:41-44, 206; ini_config.rs:332-334) |
shared_instance_type | unix/tcp | unset | Parsed for rnsd compatibility. Only tcp/unix are stored; tcp clears shared_instance_socket (tcp disables AF_UNIX upstream). lnsd currently serves only the abstract AF_UNIX socket. (config.rs:45-52; ini_config.rs:400-409, 284-289) |
shared_instance_socket | path | unset | Explicit AF_UNIX socket path (RNS 1.3.x). Parsed for compatibility; cleared when shared_instance_type = tcp. (config.rs:53-58; ini_config.rs:410-412) |
respond_to_probes | bool | false | Answer rnprobe requests by signing a proof for each probe packet. (config.rs:54-60, 146; ini_config.rs:398-400) |
remote_management_enabled | bool | false | Enable remote management. (config.rs:61-63, 147; ini_config.rs:395-399) |
storage_path | path | unset | Where identity, known destinations and packet hashlist live. Relative values resolve against the config dir. (config.rs:97-98; storage_path (ini_config.rs:560)) |
flush_interval | u64 (sec) | 3600 | Seconds between periodic storage flushes. Crash protection only — normal shutdown always flushes. (config.rs:67-73, 149; ini_config.rs:407-411) |
control_channel_capacity | usize | 256 | Capacity of the lossless control-plane event channel (announces, paths, link/resource lifecycle). Raise on servers under heavy announce load. (config.rs:74-82, 150) |
data_channel_capacity | usize | 128 | Capacity of the droppable data-plane event channel; full means normal backpressure (silent drop). Reliable channel messages are exempt: the node stops proofing them to the sender instead of dropping them, so this value also bounds how far a slow reader lets a channel run ahead. (config.rs:83-90, 151) |
keepalive_interval | u64 (sec) | unset | Override link keepalive interval. When set, every link uses this interval and the stale-link timeout scales with it (stale after twice the keepalive). Local timing only, no wire change. Useful for slow links. (config.rs:91-98, 152; ini_config.rs:408-419) |
mgmt_announce_interval_secs | u64 (sec) | 7200 | Seconds between this node's management announces, its rnstransport.probe and remote-management destinations. Leviculum-only; for test rooms that need a probe destination to re-announce within minutes, a field node keeps the default. Below 5 the config is refused with the key named. (config.rs:187-195, 319-325; ini_config.rs:512-522) |
storage_profile | desktop/compact | desktop | Transport-table sizing profile (Codeberg #421). compact is sized to leave a Raspberry Pi Zero 2W (512 MB shared with the GPU, no swap) usable. An unrecognised value keeps desktop. (config.rs:197-205; ini_config.rs:465-476) |
path_table_cap | usize | profile | Maximum path_table entries, and with them path_states, path_requests and discovery_path_requests. Desktop 32768, compact 8192. The path table expires after seven days, so on a node up less than a week this is its only bound. (config.rs:206-213; ini_config.rs:477-479) |
reverse_table_cap | usize | profile | Maximum reverse_table entries. Desktop 200000, compact 16384. Entries expire after 8 minutes, so the working size is forwarding rate times that window; a field node measured 73 901. (config.rs:214-221; ini_config.rs:480-482) |
link_table_cap | usize | profile | Maximum link_table entries — links this node routes for, not its own (that is max_links). Desktop 8192, compact 1024. (config.rs:222-229; ini_config.rs:483-485) |
announce_table_cap | usize | profile | Maximum announce_table entries, the pending-rebroadcast queue. Each holds a full copy of an announce packet. Desktop 16384, compact 2048. (config.rs:230-236; ini_config.rs:486-488) |
destination_cap | usize | profile | Maximum entries in the destination-keyed tables: announce_cache, announce_rate_table, known_ratchets, known_dest_use. Desktop 50000, compact 4096. One key for four tables because they share one population. (config.rs:237-247; ini_config.rs:489-491) |
control_channel_capacity and data_channel_capacity are read from TOML
only; they have no INI key in apply_reticulum_key (ini_config.rs:389-571)
and are best set in a TOML config or left at their defaults.
storage_path is read from both formats and resolves in one order
everywhere: lnsd --storage, then the config key, then
<config_dir>/storage (Python's only choice, Reticulum.py:246). The
client tools resolve it the same way (resolve_storage_path,
config.rs:1130), so lnstatus, lncp, lnpath and lnprobe open the
same directory as the daemon and derive the same RPC authkey from its
transport_identity. Point the key at an external disk and nothing else
has to be told about it — but note that --storage moves the daemon
alone, and the clients then still follow the config.
flush_interval and keepalive_interval are Leviculum tuning
extensions — Python Reticulum ignores them. Battery-powered or SD-card
deployments may want a longer flush_interval; slow links benefit from
a fixed keepalive_interval:
[reticulum]
# Seconds between periodic storage flushes (crash protection only,
# normal shutdown always flushes). Default: 3600.
flush_interval = 3600
# Link keepalive interval in seconds. When set, every link uses this
# interval instead of the RTT-derived default. Default: unset.
keepalive_interval = 360
Table ceilings (Codeberg #421)
Every transport table has a maximum size, and a full table evicts rather than refuses: the new entry always lands, an old one goes. Without a ceiling each table's real bound was arrival rate times expiry window — a number the neighbours choose, not the operator.
The eviction order is oldest-first for most tables and argued per table in
TableCaps and the field docs of memory_storage.rs; three tables
deviate, because plain FIFO would do damage there: link_table drops an
unvalidated link request before a live link, announce_cache drops an
unretained destination before a pinned one, and receipts drops a
terminal receipt before a pending one — and a pending one it does have to
drop is still reported as a timeout, which is what Python does on the same
overflow (Transport.py:558-561).
The defaults are entries × modelled bytes from the same model the
lnstatus diagnostic dump prints, not round numbers: desktop totals about
339 MB across the tables these keys bound, compact about 22 MB. Pick the
profile first and override individual tables only where the deployment
differs:
[reticulum]
# A Pi Zero 2W with a busy uplink: compact everywhere, but a reverse
# table large enough for the traffic it actually forwards.
storage_profile = compact
reverse_table_cap = 40000
The [interfaces] section
Interfaces are ConfigObj subsections under [interfaces], each named in
double brackets [[Name]]. The name is free-form; the type key
selects the interface implementation. Twelve types pass the
supported-type filter (interface_type (ini_config.rs:301-326)):
TCPServerInterface, TCPClientInterface, UDPInterface,
AutoInterface, RNodeInterface, RNodeMultiInterface,
SerialInterface, PipeInterface, KISSInterface,
AX25KISSInterface, I2PInterface, BLEInterface.
BackboneInterface and BackboneClientInterface are accepted too:
they are wire-identical to TCP and are mapped onto the TCP interface at
parse time, as Python does (normalize_backbone_interface
(ini_config.rs:853-883)). An interface of any other type is skipped
with a log line (tracing::warn (ini_config.rs:317-322)), not an
error.
The per-type tables below cover the six types most deployments use. All
interface keys are parsed in apply_interface_key
(ini_config.rs:617-835); the struct they land in, with its defaults,
is InterfaceConfig (config.rs:242-520).
Keys common to every interface
| Key | Type | Default | Meaning |
|---|---|---|---|
type | string | (required) | Interface type, one of the eleven above. (type (ini_config.rs:619)) |
enabled | bool | true | Bring this interface up; the legacy spelling interface_enabled is honoured too. (enabled (ini_config.rs:623); InterfaceConfig::enabled (config.rs:401-403)) |
outgoing | bool | true | Allow sending outgoing packets. (outgoing (ini_config.rs:624); InterfaceConfig::outgoing (config.rs:454-456)) |
bitrate | u64 (bps) | per type | Override the interface's own bitrate figure, which feeds announce bandwidth capping and timing. Values below MINIMUM_BITRATE (constants.rs:386-389), 5 bps, are ignored. (bitrate (ini_config.rs:672-678); InterfaceConfig::bitrate (config.rs:457-462)) |
buffer_size | usize | per type | Channel buffer size. (buffer_size (ini_config.rs:725); InterfaceConfig::buffer_size (config.rs:595-597)) |
TCP server (TCPServerInterface)
| Key | Type | Default | Meaning |
|---|---|---|---|
listen_ip | string | unset | Address to bind. (listen_ip (ini_config.rs:641)) |
listen_port | u16 | unset | Port to listen on. (listen_port (ini_config.rs:632)) |
[interfaces]
[[Loopback TCP]]
type = TCPServerInterface
enabled = Yes
listen_ip = 127.0.0.1
listen_port = 45999
TCP client (TCPClientInterface)
| Key | Type | Default | Meaning |
|---|---|---|---|
target_host | string | unset | Remote host to connect to. (target_host (ini_config.rs:640)) |
target_port | u16 | unset | Remote port. (target_port (ini_config.rs:634)) |
reconnect_interval | u64 (sec) | 5 | Delay between reconnect attempts. (reconnect_interval (ini_config.rs:726); InterfaceConfig::reconnect_interval_secs (config.rs:597-598)) |
max_reconnect_tries | u64 | unlimited | Give up after this many attempts; unset means never. (max_reconnect_tries (ini_config.rs:727); InterfaceConfig::max_reconnect_tries (config.rs:599-600)) |
[interfaces]
[[RNS TCP Node Germany 002]]
type = TCPClientInterface
enabled = Yes
target_host = 193.26.158.230
target_port = 4965
UDP (UDPInterface)
| Key | Type | Default | Meaning |
|---|---|---|---|
listen_ip | string | 0.0.0.0 | Local bind address. (listen_ip (ini_config.rs:641)) |
listen_port | u16 | unset | Local bind port. (listen_port (ini_config.rs:632)) |
forward_ip | string | unset | Broadcast/forward address or hostname. Names are resolved at runtime and re-resolved periodically; a resolution failure is a logged interface error, not a config error. (forward_ip (ini_config.rs:642); InterfaceConfig::forward_ip (config.rs:534-544)) |
forward_port | u16 | unset | Broadcast/forward port. (forward_port (ini_config.rs:643)) |
port | u16 | unset | Fills both listen_port and forward_port; either explicit key wins over it. (port (ini_config.rs:523)) |
device | string | unset | Kernel interface name; its IPv4 broadcast address fills both listen_ip and forward_ip. Either explicit key wins over it. (device (ini_config.rs:629)) |
Bind and forward are independent, as in rnsd: an interface with only bind
parameters receives without transmitting, one with only forward parameters
transmits without listening, and only an interface that would do neither is a
configuration error.
AutoInterface (AutoInterface)
Discovers other Reticulum nodes on the same broadcast domain via multicast. No router or DHCP needed; the link must carry multicast.
| Key | Type | Default | Meaning |
|---|---|---|---|
group_id | string | unset | Multicast group identifier; isolate co-located meshes by setting different IDs. (group_id (ini_config.rs:739); InterfaceConfig::group_id (config.rs:603-605)) |
discovery_scope | string | unset | Multicast scope: link, admin, site, organisation, global. (discovery_scope (ini_config.rs:740); InterfaceConfig::discovery_scope (config.rs:605-606)) |
discovery_port | u16 | 29716 | Discovery (announce) port. (discovery_port (ini_config.rs:741); InterfaceConfig::discovery_port (config.rs:607-608)) |
data_port | u16 | 42671 | Data port. (data_port (ini_config.rs:742); InterfaceConfig::data_port (config.rs:609-610)) |
devices | string (CSV) | unset | Whitelist of NIC names to use. (devices (ini_config.rs:744); InterfaceConfig::devices (config.rs:611-612)) |
ignored_devices | string (CSV) | unset | Blacklist of NIC names to skip. (ignored_devices (ini_config.rs:744); InterfaceConfig::ignored_devices (config.rs:613-614)) |
multicast_loopback | bool | unset (inherits true) | Multicast loopback (IPV6_MULTICAST_LOOP), the carrier self-echo mechanism. Unset inherits the default true, matching Python-RNS; set no to opt out. (multicast_loopback (ini_config.rs:745); InterfaceConfig::multicast_loopback (config.rs:617-620)) |
multicast_address_type | string | unset (inherits temporary) | Multicast address type of the discovery group, temporary or permanent. It is part of the group address, so peers must agree on it: an lnsd node left on the default next to a permanent-type rnsd peer group discovers nobody, and nothing on either side says why. Unset inherits temporary, the group Python joins when the key is absent. A value that is neither spelling is refused at startup rather than resolved to temporary the way Python resolves it. (multicast_address_type (ini_config.rs:740-746); InterfaceConfig::multicast_address_type (config.rs:619-624); MulticastAddressType (interfaces/auto_interface/mod.rs:42-91)) |
BLE (BLEInterface)
Joins the Columba BLE mesh (the ble-reticulum protocol, v2.2 wire
format with the v0.3.0 capability record) as a dual-role BlueZ node: it
advertises and serves the Columba GATT layout like an LNode board does,
and it scans for and connects to nearby peers under the same
connection-direction rule the boards and phones apply. One section is
one Reticulum interface — a single broadcast domain across all live BLE
links. Requires a BlueZ (bluetoothd) host with a BLE-capable adapter;
if the adapter is missing or powered off at startup the interface keeps
retrying rather than failing the daemon.
The interface is off unless a [[BLE Interface]] section exists in the
config; the daemon never brings BLE up on its own. The advertised name
is derived from the daemon identity as LN-<hex8> exactly like the
firmware's, so scanner listings show lnsd and boards the same way.
Key names follow the reference ble-reticulum package where its options
map onto this implementation:
| Key | Type | Default | Meaning |
|---|---|---|---|
device | string | default adapter | BlueZ adapter to use, e.g. hci0. (device (ini_config.rs:629)) |
max_connections | usize | 4 | Simultaneous BLE link cap, both GATT roles counted together — and split by role inside it since #432: one slot is this node's own outgoing dial, the other max_connections - 1 are incoming links, the boards' shape (CENTRAL_LINKS / PERIPH_LINKS). Dialling is gated on the outgoing slot and advertising on the incoming ones, so a node that has dialled still advertises and a node with a full GATT server still dials. The default is the firmware's MAX_LINKS (4), not the reference's 7: 3-4 links is the protocol's reliable ceiling. (max_connections (ini_config.rs:811); InterfaceConfig::max_connections (config.rs:738-741)) |
min_rssi | i16 (dBm) | -85 | Sightings weaker than this are not dialled. (min_rssi (ini_config.rs:812); InterfaceConfig::min_rssi (config.rs:742-744)) |
discovery_interval | f64 (sec) | 5 | Pause between the 2-second BLE scan windows. (discovery_interval (ini_config.rs:813); InterfaceConfig::discovery_interval (config.rs:745-747)) |
enable_central | bool | true | Run the scanning + dialling central role. (enable_central (ini_config.rs:814); InterfaceConfig::enable_central (config.rs:748-750)) |
enable_peripheral | bool | true | Run the advertising + GATT-server peripheral role. Disabling both roles is a config error. (enable_peripheral (ini_config.rs:815); InterfaceConfig::enable_peripheral (config.rs:750-752)) |
initiate_only | string (CSV) | unset (every peer) | Peers this interface may DIAL: BLE addresses (AA:BB:CC:DD:EE:FF, - or no separator) or peer identities in hex, 8 digits (the four bytes an advertiser publishes as its hint) or all 32 (a board's [IDENTITY] line, truncated to those four). Unset or empty dials whoever the connection-direction rule picks, the behaviour that predates the key. It narrows dialling and nothing else: a peer left off the list that connects to US is admitted and served exactly as before, and nothing on the wire changes — it sees a node that has not dialled it yet. The digit count decides which is which — 12 is an address, 8 or 32 an identity — so both spellings can be copied out of a log line (BLE_SCAN_DECISION addr=, a board's BLE_CENTRAL_ADDR, its [IDENTITY]). A malformed entry is a startup error, not a dropped line. (initiate_only (ini_config.rs:811-813); InterfaceConfig::initiate_only (config.rs:754-763); PeerAllowlist (interfaces/ble/links.rs:1262)) |
accept_only | string (CSV) | unset (every peer) | Peers whose INCOMING link this interface SERVES — the symmetric counterpart of initiate_only, same vocabulary, same validation, same "unset or empty means everyone". A non-empty list narrows who we serve and nothing else: a peer left off it is still dialled if initiate_only allows it, and a peer left off initiate_only is still served if this list names it. An unlisted peer that connects is refused at the identity handshake — the first moment an inbound BLE connection has said who it is, since under RPA its address names nobody — so it never becomes a link, never enters the fan-out and is never reported to the core as a peer. Each refusal emits one BLE_LINK_NOT_ADMITTED peer=<hex8> identity=<hex32> addr=<a> role=peripheral listed=<n> action=disconnect line, so a run that turned strangers away is distinguishable from a run nobody tried. A malformed entry is a startup error, not a dropped line. (accept_only (ini_config.rs:825-827); InterfaceConfig::accept_only (config.rs:764-779); the refusal point (interfaces/ble/links.rs:740)) |
[interfaces]
[[BLE Interface]]
type = BLEInterface
enabled = yes
# device = hci0
# max_connections = 4
# min_rssi = -85
# Dial nothing but these two; still answer anyone who dials us.
# initiate_only = b2a8bea1, AA:BB:CC:DD:EE:FF
# Serve nothing but these two; a stranger's connection is refused.
# accept_only = b2a8bea1, AA:BB:CC:DD:EE:FF
A node that should be a leaf rather than a hub — one uplink out, still
reachable from anybody near it — is initiate_only naming that uplink.
A node that should be a leaf and invisible as well adds
enable_peripheral = no, which is the stronger statement: it stops
advertising, so no peer can dial it either.
The two keys answer independent questions and accept_only is the one
for a room the operator does not control. enable_peripheral = no
refuses every incoming link, including the ones the deployment wants;
accept_only names the set it wants and refuses the rest, which is what
a measurement in a flat needs — a Faraday cage stops LoRa, it does not
stop the phone in someone's pocket from dialling a Columba advertiser.
Setting both keys to the same list pins a closed mesh: we dial nobody
else and we serve nobody else. Setting neither is the default and is
what every existing deployment does.
RNode and Serial (RNodeInterface, SerialInterface)
RNodeInterface drives an RNode LoRa modem; SerialInterface is a raw
serial HDLC link. They share the serial-port and LoRa keys.
Divergence from Python: there, only RNodeInterface honours the LoRa
keys (frequency, bandwidth, spreadingfactor, codingrate,
txpower) and SerialInterface reads port settings only. Leviculum's
SerialInterface honours them too and configures the attached LNode's
radio over the serial port — the LNode frames HDLC, so it cannot be
driven by the KISS-framed RNodeInterface.
Serial keys:
| Key | Type | Default | Meaning |
|---|---|---|---|
port | string | unset | Serial device path, e.g. /dev/ttyACM0. (port (ini_config.rs:523); InterfaceConfig::port (config.rs:479-480)) |
speed / baudrate | u32 | unset | Serial baud rate (either spelling). (speed (ini_config.rs:645); InterfaceConfig::speed (config.rs:551-552)) |
databits | u8 | unset | Data bits. (databits (ini_config.rs:646); InterfaceConfig::databits (config.rs:553-554)) |
parity | string | unset | none, even, or odd. (parity (ini_config.rs:647); InterfaceConfig::parity (config.rs:555-556)) |
stopbits | u8 | unset | Stop bits. (stopbits (ini_config.rs:648); InterfaceConfig::stopbits (config.rs:557-558)) |
LoRa keys, derived from source — the meanings below describe the radio
parameters the interface configures; the fields sit together in the
RNode block of InterfaceConfig (InterfaceConfig::frequency
(config.rs:623-654)):
| Key | Type | Default | Meaning |
|---|---|---|---|
frequency | u64 (Hz) | unset | LoRa centre frequency. (frequency (ini_config.rs:728); InterfaceConfig::frequency (config.rs:658-660)) |
bandwidth | u32 (Hz) | unset | LoRa bandwidth. (bandwidth (ini_config.rs:729); InterfaceConfig::bandwidth (config.rs:637-638)) |
spreadingfactor / spreading_factor | u8 | unset | LoRa spreading factor (either spelling). (spreadingfactor (ini_config.rs:730); InterfaceConfig::spreading_factor (config.rs:679-680)) |
codingrate / coding_rate | u8 | unset | LoRa coding rate (either spelling). (codingrate (ini_config.rs:731); InterfaceConfig::coding_rate (config.rs:681-682)) |
txpower / tx_power | i8 (dBm) | unset (resolves to the board maximum, 22 dBm) | Transmit power (either spelling). Unset asks the board for its maximum — a board that can do less clamps and says so — rather than the 0 dBm (1 mW) Python-Reticulum resolves it to, which has no symptom at the node. An explicit txpower = 0 still means 0. Above roughly 7 dBi of antenna gain, 22 dBm conducted exceeds the EU 27 dBm ERP allowance and has to be set down. (txpower (ini_config.rs:732); InterfaceConfig::tx_power (config.rs:683-684); resolve_tx_power (rnode.rs:761); deviation) |
flow_control | bool | unset | Wait for the RNode's CMD_READY before the next TX. (flow_control (ini_config.rs:753); InterfaceConfig::flow_control (config.rs:701-702)) |
airtime_limit_short | f64 (%) | unset | Short-term airtime cap, percent (0.0–100.0). (airtime_limit_short (ini_config.rs:754); InterfaceConfig::airtime_limit_short (config.rs:703-704)) |
airtime_limit_long | f64 (%) | unset | Long-term airtime cap, percent (0.0–100.0). (airtime_limit_long (ini_config.rs:755); InterfaceConfig::airtime_limit_long (config.rs:705-706)) |
csma_enabled | bool | unset | Carried in the LNode radio-config frame and reported back, but current firmware no longer obeys it: LoRa channel access (pre-TX jitter plus CAD listen-before-talk) is always on, matching the RNode firmware, which offers no CSMA disable either. Only firmware older than the change still honours the flag. (csma_enabled (ini_config.rs:756); InterfaceConfig::csma_enabled (config.rs:707-708)) |
preamble_symbols | u16 (symbols) | unset (derived from the PHY) | LoRa preamble length pushed to LNode firmware in the radio-config frame (SerialInterface only). Unset derives what an RNode peer programs for the same PHY — 24 symbols at SF7/BW125, the 18-symbol floor from SF8 down — so a mixed pair agrees on the wire; set it only to pin a value against a non-conforming peer. A pin above roughly 20 symbols / 164 ms on air is warned about at startup and not refused: SX127x receivers (every RNode) were measured going deaf above that, losing every frame from the interface silently and one-way, while an SX126x peer copes (Codeberg #315). Not the same key as the KISS preamble (TX delay in ms), which never reaches a LoRa modem. (preamble_symbols (ini_config.rs:737); InterfaceConfig::preamble_symbols (config.rs:687-699); derive_preamble_symbols (rnode.rs:917); preamble_ceiling_warning (interfaces/serial.rs)) |
A board that does not take the config
A SerialInterface with a LoRa block pushes the radio config at the board up
to three times, waiting 2 s for the firmware's ACK each time. Then — ACK or no
ACK — it sends the radio query (TYPE_RADIO_QUERY,
Codeberg #349) and waits
a further 2 s for the board's report, which the firmware answers out of what
its LoRa task actually configured. The ACK alone does not settle it: the
legacy config frame's entire vocabulary is three bytes or silence, so a board
that only wrote the config to its flash page acks it exactly like a board that
keyed it (Codeberg #363).
- Report received, same profile — the board is running what was asked for. The interface comes up on it. (Without an ACK this is the same outcome: the config did arrive, only its receipt did not.)
- Report received, different profile — the interface comes up priced at
the profile the board reported, and logs
RADIO_BRINGUP iface=<name> outcome=running-differswith the requested and the running parameters side by side. - Query refused as busy or not-running — the board is talking and says no
radio is running this boot, which on an LNode means it booted
lora=off. The interface refuses to come up and logsRADIO_BRINGUP iface=<name> outcome=dead-radio lora=off. An ACK does not change this; it is the case that ACK is unable to distinguish. - Neither answered — the interface refuses to come up. It logs
RADIO_BRINGUP iface=<name> outcome=refusednaming both frames and both waits, reports Down tornstatus, and the daemon keeps running with its other interfaces. It does not retry on that port; fix the board and restart.
Both refusals report Down to rnstatus and leave the rest of the daemon
running.
A modem the host cannot price is a modem the host does not drive. Airtime accounting and transmit spacing are computed from the PHY, so an interface pricing SF7 in front of a board keying SF12 under-counts duty by an order of magnitude and hands the serial queue frames faster than the modem can key them; guessing the other way is a silent lie about airtime that nothing downstream can tell from a measurement.
Firmware older than #349 does not answer the radio query — it refuses it as a
frame it does not know, or says nothing. That is not a board stating its radio
is off, so such a board still comes up on its ACK, with a debug line saying the
running profile could not be read; an LNode running it that also misses the
config ACK is refused rather than driven blind.
(radio_bring_up, radio_pricing_phy (interfaces/serial.rs))
Test-only: test_drop_direct_ingress (bool, default off) emulates
out-of-range placement on a co-located rig: the interface drops every
received frame whose wire hops byte (raw[1]) is 0 — frames heard
directly from their originator — before they reach the transport, while
relayed copies (hops ≥ 1) pass. Two endpoints with this knob on one
bench are mutually deaf but both hear a relay, giving the A–B–C
repeater topology without attenuators; the consumer is the 3-node relay
hardware scenario (periculum hardware/lora_3node_relay.toml).
Python-RNS has no such option — frames are dropped locally on ingress
and nothing on the air changes, so this is a test-harness affordance,
not a wire or semantic deviation. Supported on RNodeInterface and
SerialInterface; incompatible with IFAC (which prepends material
before the flags byte), and that combination is refused at startup.
IFAC (Interface Access Codes)
IFAC keys apply to any interface and authenticate / isolate a virtual network on the link. They are common to all interface types:
| Key | Type | Default | Meaning |
|---|---|---|---|
networkname / network_name | string | unset | Network name for IFAC (either spelling). (networkname (ini_config.rs:790); InterfaceConfig::networkname (config.rs:667-669)) |
passphrase / pass_phrase | string | unset | IFAC passphrase (either spelling). (passphrase (ini_config.rs:705); InterfaceConfig::passphrase (config.rs:669-670)) |
ifac_size | usize (bits) | unset | IFAC size, specified in bits in the file and stored as bytes (bits / 8). Values below 8 bits are dropped, as Python drops them, so the interface falls back to its per-type default. (ifac_size (ini_config.rs:787-793); InterfaceConfig::ifac_size (config.rs:671-672)) |
networkname and passphrase are secrets: lnstest diag redacts them
before serialising a bundle (see the lnstest diag section).
Example configurations
Simple AutoInterface node
A node that talks to other Reticulum peers on the same LAN, no transport routing:
[reticulum]
enable_transport = No
share_instance = Yes
[interfaces]
[[Default Interface]]
type = AutoInterface
enabled = Yes
TCP-server transport node
A routing entrypoint that accepts inbound TCP peers and bridges them with the local LAN:
[reticulum]
enable_transport = Yes
share_instance = Yes
instance_name = entrypoint
[interfaces]
[[Public TCP]]
type = TCPServerInterface
enabled = Yes
listen_ip = 0.0.0.0
listen_port = 4965
[[Local LAN]]
type = AutoInterface
enabled = Yes
LoRa RNode node
A node on a LoRa RNode modem (radio values below are an EU 868 MHz example; set them for your region and hardware):
[reticulum]
enable_transport = Yes
share_instance = Yes
[interfaces]
[[LoRa RNode]]
type = RNodeInterface
enabled = Yes
port = /dev/ttyACM0
frequency = 867200000
bandwidth = 125000
spreadingfactor = 8
codingrate = 5
txpower = 14
See the upstream Reticulum Manual for the protocol-level meaning of the radio and IFAC parameters.
lnsd Quickstart for Beta Testers
This page gets you from "I have the .deb file" to "my node is on the
mesh and I know how to tell if it isn't", plus the one-liner you run when
something is off so the bug report has everything we need.
For the protocol itself, the upstream Reticulum Manual
is the reference. This page is about getting lnsd running on your
machine.
Prerequisites
- Linux, x86_64 or aarch64. (macOS and embedded targets exist but are
out of scope for the beta
.debpath.) - The nightly
.debfor your architecture. Download links are on the releases page. The binaries inside are statically linked against musl, so the package installs on Debian ≥ 9 and Ubuntu ≥ 16.04 regardless of host glibc. - A few free TCP/UDP ports on your machine for the configured interfaces (default ports below).
You do not need to install Rust, Python, or Docker for the beta flow.
Install
sudo apt install ./leviculum-nightly-amd64.deb # or -arm64
The package:
- Installs
lnsd,lnstest,lncp, andlnstatusunder/usr/bin/, with a man page for each. - Creates a system user
leviculumand a group of the same name. - Drops a default config file at
/etc/reticulum/config(mode 644) and creates the config directory/etc/reticulummode 2775 (group-writable + setgid, so files created inside it inherit theleviculumgroup). - Enables and starts the
lnsd.servicesystemd unit.
For the native tools (lnstest, lncp, lnstatus) and Python tools (rnstatus,
rnpath, rnprobe, Sideband, Nomadnet, …) to talk to the running
daemon, your user has to be in the leviculum group:
sudo usermod -aG leviculum "$USER"
# log out and back in, or `newgrp leviculum` for this shell only
Verify the installation:
lnsd --version # e.g. 0.7.0-nightly.20260419-5a5df20
lnstest --version
systemctl is-active lnsd
is-active should print active. If it prints failed, jump to
Troubleshooting.
Minimum-viable config
The default /etc/reticulum/config is conservative: it brings up a
single AutoInterface for local LAN peers, with transport routing
disabled. That's enough to talk to other Reticulum nodes on the same
LAN, but it does not connect you to the wider mesh.
A reasonable beta-tester config has two interfaces: one for the LAN, one
TCP uplink to a public entrypoint. Edit /etc/reticulum/config to:
[reticulum]
# Pass announces and serve paths for other peers. Leave off if your
# machine is mobile or sleeps a lot.
enable_transport = Yes
# Required for `lnstest diag`, `rnstatus`, Sideband etc. to attach to
# this daemon. The default config already sets this.
share_instance = Yes
[interfaces]
# 1. Local mesh: discovers and talks to every other Reticulum node
# on the same broadcast domain. No router/DHCP needed. Multicast
# has to reach the link (most home LANs do; corporate Wi-Fi often
# does not).
[[Default Interface]]
type = AutoInterface
enabled = Yes
# 2. TCP uplink to a public entrypoint. Pick a node from the
# community directory: https://directory.rns.recipes/ (entrypoints
# rotate; for redundancy add two or three, and see the Reticulum
# manual's "Bootstrapping Connectivity" section for the
# discover_interfaces auto-peering option). Example below: the
# RNS TCP Node Germany 002 entry.
[[RNS TCP Node Germany 002]]
type = TCPClientInterface
enabled = Yes
target_host = 193.26.158.230
target_port = 4965
Then restart the daemon so it picks up the new config:
sudo systemctl restart lnsd
lnstest diag (below) is the easiest way to confirm both interfaces came
up.
Start the daemon
The systemd unit handles this for you on install. The relevant commands:
sudo systemctl start lnsd # or restart
sudo systemctl stop lnsd
sudo systemctl status lnsd
journalctl -u lnsd -f # live log tail
journalctl -u lnsd --since '10 min ago'
Logs go to the journal. Increase verbosity by editing the unit's
ExecStart to add -v (debug) or -vv (trace), then
sudo systemctl daemon-reload && sudo systemctl restart lnsd. The
RUST_LOG environment variable also works (see lnsd --help).
To run lnsd by hand without systemd (useful for ad-hoc debugging):
sudo systemctl stop lnsd
sudo -u leviculum /usr/bin/lnsd -v --config /etc/reticulum
Check it's working
Three commands. Run them as a user that is in the leviculum group.
1. lnstest diag
This is the main health-check. It connects to the running daemon over the shared-instance socket and renders a single-file diagnostic bundle:
lnstest diag --config /etc/reticulum
A healthy bundle looks roughly like this (your transport id, paths,
and byte counters will differ):
===== Leviculum diagnostic bundle =====
----- Versions / build -----
lnstest version: 0.7.0
build profile: release
target: x86_64 / linux
daemon version: not exposed by the shared-instance RPC ...
----- Config -----
config dir: /etc/reticulum
config file: /etc/reticulum/config
config file: present, parsed OK
Effective config (TOML, secrets redacted; the raw file is NOT included
because it may contain secrets):
[reticulum]
enable_transport = true
shared_instance = true
instance_name = "default"
...
----- Daemon view (shared-instance RPC) -----
instance name: default
RPC socket: \0rns/default/rpc
authkey: derived from /etc/reticulum/storage/transport_identity (not shown)
## interface_stats
transport id: 0123456789abcdef0123456789abcdef
daemon uptime: 12m 34s (754s)
interfaces (2):
- AutoInterface[Default Interface] type=AutoInterface status=up rxb=482 txb=917 peers=2
- tcp_client_0 type=TCPClientInterface status=up rxb=14211 txb=8332
raw:
{ ... full JSON dump of the same data, one object per interface ... }
## path_table
known paths: 7
[ ... JSON array of {hash, interface, hops, expires, ...} ... ]
## link_count
relayed links: 0
## link_table
links (0):
raw:
[ ... JSON array of the links this node TERMINATES; note that
`link_count` above counts a different table, the links it RELAYS.
This section is a Leviculum extension and shows <unavailable>
against a Python rnsd ... ]
----- System -----
os: linux kernel: 6.12.73+deb13-amd64
distro: Debian GNU/Linux 13 (trixie)
lnsd pid: 12345
lnsd VmRSS: 18432 kB
lnsd open fds: 27
----- Recent events -----
No structured event-log file specified ...
===== end of diagnostic bundle =====
What to look at first:
status=upon every interface in theinterface_statssection. An interface that came up but lost its medium reportsstatus=down.- Non-zero
rxb/txbon the interfaces you expect traffic on (AutoInterfaceonce any other Reticulum node is on the same LAN, the TCP uplinktcp_client_Nas soon as it connects). peers=…on theAutoInterfaceline: how many other Reticulum nodes are visible on the LAN.known paths: Nwith N > 0 once announces have crossed the mesh. Brand-new daemons that haven't heard any announces yet showknown paths: 0for the first few seconds — that's normal.transport idis your node's identity (the public half). It is safe to share; the private half lives in/etc/reticulum/storage/transport_identityand is never included inlnstest diagoutput.
2. lnstest selftest --help
Sanity-checks that lnstest itself is installed and runnable:
lnstest selftest --help
The actual lnstest selftest exercise needs one or two relay nodes you
control. The full command and options are in lnstest selftest --help.
3. rnstatus (optional — Python tools)
The .deb does not install Python Reticulum. If you want rnstatus
/ rnpath / rnprobe / Sideband, install it in its own environment.
Debian 12+ and Ubuntu 24.04+ refuse pip install into the system
Python with an externally-managed-environment error (PEP 668), so use
pipx (or a venv):
sudo apt install pipx
pipx install rns
rnstatus
With a plain virtual environment instead:
sudo apt install python3-venv
python3 -m venv ~/.rns-venv
~/.rns-venv/bin/pip install rns
~/.rns-venv/bin/rnstatus
Python tools auto-detect /etc/reticulum/config and connect to the
running lnsd through the same shared-instance socket. No extra flags
are needed. (The native lnstatus from the .deb covers the same
ground as rnstatus; the Python install is only needed for rnpath,
rnprobe, Sideband, and friends.)
Connect to the wider mesh
With the config above, two things happen as soon as lnsd starts:
- Announcing. Your node sends an announce for its probe destination on every enabled interface. Other transport-enabled nodes pass that announce on, so within seconds your node is visible to peers on the LAN and within a few minutes to peers reachable through the TCP uplink.
- Learning paths. When other nodes announce, your daemon stores
a path to each announced destination (destination hash, the
interface it was heard on, hop count, expiry).
lnstest diag'sknown paths: Nis that table's size.
When you want to talk to a specific destination (e.g. send a file with
lncp), the daemon either has a path already (immediate) or requests
one (a path-request packet, a few seconds, then immediate). You don't
have to do anything to make path discovery happen — it runs whenever
the daemon is up.
For the protocol-level picture, read Bootstrapping Connectivity in the upstream Reticulum manual.
Troubleshooting
lnsd will not start
systemctl status lnsd
journalctl -u lnsd --since '10 min ago' | tail -50
Common causes:
- Config not parsed. Look for a "Failed to parse config" line in
the journal.
lnstest diag --no-rpcshows the parse status without needing the daemon up:config file: present but FAILED to parse: <details> - Abstract socket already in use. Another
lnsdorrnsdis running under the sameinstance_name. Stop it (sudo systemctl stop lnsdthenpkill -f rnsdif applicable), or set a uniqueinstance_namein your config. - Permission on the storage directory. The
leviculumuser has to be able to write/etc/reticulum/storage/. The.debsets the permissions correctly on install; a manualchowntoroot:rootbreaks the daemon. Fix:sudo chown -R leviculum:leviculum /etc/reticulum sudo chmod 2775 /etc/reticulum
No peers found / known paths: 0
Check lnstest diag's interface_stats section:
AutoInterfaceshowspeers: 0andrxb: 0— multicast isn't reaching the link. Likely causes: corporate Wi-Fi (multicast blocked); a Linux bridge or container network without multicast forwarding; no other Reticulum node on the segment.TCPClientInterfaceshowsstatus=upbutrxb: 0— TCP connected but the remote isn't sending anything, which usually means the remote is up but has no transport peers itself, or the entrypoint has been retired. Try a different entrypoint, or rely on AutoInterface + a TCP uplink to a known-good node you control.TCPClientInterfacenot listed at all — the daemon hasn't connected yet (look forEstablishing TCP connectionlines injournalctl -u lnsd) or DNS for the target host doesn't resolve.
Give it ~30 seconds after starting lnsd before concluding there's a
problem — the first round of announces and the initial TCP connect
take a moment.
Native and Python tools cannot reach lnsd
Symptom: in the daemon-view section, lnstest diag shows either
cannot derive RPC authkey: …/storage/transport_identity: Permission denied followed by (daemon queries skipped) (your user cannot read
the identity file, almost always a missing group membership), or
<unavailable: …> on the individual queries (the daemon is down or
you targeted the wrong instance). rnstatus errors with "Reticulum is
not running".
- Confirm your user is in the
leviculumgroup:
If not,id | tr , '\n' | grep leviculumsudo usermod -aG leviculum "$USER"and log out / back in. - Confirm the daemon really is up and has
share_instance = Yes:systemctl is-active lnsd grep -i share_instance /etc/reticulum/config - Confirm both client and daemon are using the same config directory.
The client defaults to
/etc/reticulumif it exists, then~/.config/reticulum, then~/.reticulum— the full resolution order is in Installation.lnstest diag --config /etc/reticulumis explicit.
Submitting a bug report
Run lnstest diag and attach its output to your report:
lnstest diag --config /etc/reticulum --output /tmp/lnstest-diag.txt
The bundle is plain UTF-8 text, designed to be safe to attach: IFAC
passphrase and networkname are redacted before serialisation; the
node identity private key is never read into the bundle (only its
SHA-256 is used, internally, to derive the shared-instance RPC
authkey). The bundle does contain your node's hostnames, configured
TCP targets, byte counters, and known-destinations table — review it
once before posting to a public tracker if your topology is sensitive.
If lnsd is in a structured event-log run
(LEVICULUM_EVENT_LOG=/var/log/lnsd-events.log in the service unit's
Environment=), include the tail of that file too:
lnstest diag --event-log /var/log/lnsd-events.log \
--output /tmp/lnstest-diag.txt
Otherwise the bundle already points the reviewer at journalctl -u lnsd, which is enough.
See also
lnsd --help,lnstest --help,lncp --help,lnstatus --helpfor the full command and option reference; the.debalso installs a man page for each (man lnsd, …).- Configuration for the format reference.
- Installation for the source-build path.
- The upstream Reticulum Manual for the protocol itself.
lnstest
lnstest is the Reticulum test and diagnostics tool. It manages
identities, runs an interactive session against a daemon, exercises the
stack with a self-test, and collects diagnostic bundles. It works
against either lnsd or Python rnsd through the shared-instance
interface. For file transfer use the standalone lncp tool.
Reticulum test and diagnostics tool
Usage: lnstest [OPTIONS] <COMMAND>
Commands:
identity Identity management
selftest Run integration self-test through relay node(s)
connect Interactive session: connect to rnsd and enter command loop
diag Collect a diagnostic bundle from a running lnsd (or rnsd) for bug reports
help Print this message or the help of the given subcommand(s)
Options:
-c, --config <CONFIG> Config directory of the daemon to ask
-v, --verbose Enable verbose logging
--corrupt-every <CORRUPT_EVERY> Corrupt ~1 byte per N bytes on TCP write (fault injection)
-h, --help Print help
-V, --version Print version
-c/--config, -v/--verbose, and --corrupt-every are global flags
available on every subcommand. --config names a daemon's config
directory — the one lnsd --config was given — not a file: diag and
selftest both read the config and the identity under it to reach that
daemon's shared instance. --corrupt-every is a fault-injection tool for
testing and should be left off in normal use.
identity
Manage Reticulum identities. An identity file holds the 64-byte private key (X25519 + Ed25519); the public half and the 16-byte hash are derived from it.
Usage: lnstest identity <COMMAND>
Commands:
generate Generate a new identity
show Show identity information
identity generate
Creates a fresh identity. With -o/--output FILE the private key is
written to the file and the hash is printed; without it, the hash and
public key are printed and nothing is saved (lnstest.rs:1033-1052).
lnstest identity generate -o my-identity.bin
Generated new identity
Hash: 0123456789abcdef0123456789abcdef
Saved to: my-identity.bin
Without -o, the public key is printed instead of being saved
(lnstest.rs:1047-1051):
lnstest identity generate
Generated new identity
Hash: 0123456789abcdef0123456789abcdef
Public key: <64 hex bytes>
identity show
Loads a saved identity file and prints its path, hash, and public key
(lnstest.rs:1054-1066):
lnstest identity show my-identity.bin
Identity: my-identity.bin
Hash: 0123456789abcdef0123456789abcdef
Public key: <64 hex bytes>
File transfer
lnstest has no file-copy subcommand. Use the standalone
lncp tool, the drop-in for Python's rncp, which attaches
to a running lnsd (or rnsd) through the shared instance.
connect
Open an interactive session against a running daemon (lnsd or rnsd)
and enter a command loop. The address is the daemon's TCP interface
(host:port); lnstest verifies TCP connectivity before building the node
(lnstest.rs:486-488).
Usage: lnstest connect [OPTIONS] <ADDR>
Arguments:
<ADDR> Address of the rnsd to connect to (host:port)
Options:
-c, --config <CONFIG> Config directory of the daemon to ask
--identity <IDENTITY> Path to identity file (default: generate ephemeral)
With no --identity, an ephemeral identity is generated for the session
(lnstest.rs:497-501). On connect, the session announces itself and prints
its identity and destination hashes (lnstest.rs:537-547):
Identity: <hash>
Destination: <hash>
Announced as lnstest-cli
Type /help for commands.
>
Interactive commands
The command loop accepts these (lnstest.rs:592-820):
| Command | Action |
|---|---|
/peers | List discovered destinations |
/link <hash> | Initiate a link to a destination (32-char hex) |
/target <hash> | Set a single-packet destination (32-char hex) |
/untarget | Clear the single-packet target |
/send <msg> | Send data on the active link or to the target |
/close | Close the active link |
/announce | Re-announce this destination |
/quiet | Hide announce/path messages |
/verbose | Show announce/path messages |
/status | Show node status (identity, destination, paths, peers) |
/help | Show this help |
/quit | Exit |
<bare text> | Send as data on the active link or to the target |
selftest
Run an end-to-end self-test through one or two relay nodes you control.
The relay addresses are given as host:port.
Usage: lnstest selftest [OPTIONS] [TARGETS]...
Arguments:
[TARGETS]... Address(es) of relay node(s) (host:port). One or two addresses
Options:
-c, --config <CONFIG>
Config directory of the daemon to ask
--duration <DURATION>
Test duration in seconds [default: 180]
--rate <RATE>
Messages per second per direction [default: 1]
--mode <MODE>
Which test phases to run [default: all]
--messages <MESSAGES>
Messages per direction in the ratchet exchange [default: 10]
--discovery-timeout <DISCOVERY_TIMEOUT>
Discovery timeout in seconds (Phase 2: mutual path discovery) [default: 60]
The --mode flag selects which phases run; the values are all,
link, packet, ratchet-basic, ratchet-enforced,
bulk-transfer, and ratchet-rotation (default all). --duration
defaults to 180 seconds, --rate to 1 message per second per
direction, and --discovery-timeout to 60 seconds.
--messages sizes the ratchet exchange: how many messages each
direction sends under --mode ratchet-basic and --mode ratchet-enforced, 10 by default. --duration and --rate do not
reach that count — they size Phase 5 (sustained link exchange) and
Phase 8 (single-packet exchange), and no ratchet mode runs either — so
--messages is the only way to widen a ratchet measurement. The other
two ratchet modes keep counts that are part of what they test:
bulk-transfer sends 100 each direction, and ratchet-rotation sends
5 before and 5 after the key change so the two halves can be compared.
Given the flag, they print a line saying it sized nothing.
A wider exchange is worth asking for when a percentage has to decide something. A pass/fail bar sits inside the confidence interval 20 packets support, which is about 25 points wide, so a run of 10 each way can distinguish full delivery from collapse and little in between; 40 each way narrows it enough for a bar in the eighties to be decided either way.
# 40 each direction, for a delivery floor that has to be decidable
lnstest selftest 127.0.0.1:4242 127.0.0.1:4242 --mode ratchet-basic --messages 40
# Full self-test through one relay
lnstest selftest 192.0.2.10:4965
# Just the link phase, two relays, shorter run
lnstest selftest --mode link --duration 60 192.0.2.10:4965 192.0.2.11:4965
Sizing the drain window on a slow link
Every single-packet phase waits for what is still in flight before it reads the receive counter. On a radio link that wait has to be sized from the link: ten 147-byte frames over a 2734 bps LoRa link need five seconds of air, and a fixed sleep shorter than that counts the frames still on the air as lost.
The tool has no radio of its own — it is a TCP client, usually two hops from one — so it asks the daemon that owns the radio. Give it that daemon's config directory:
lnstest -c /root/.reticulum selftest 127.0.0.1:4242 peer:4242 --mode ratchet-basic
It resolves the shared instance from that directory, reads
interface_stats, and sizes each phase's window from the reported
on-air bitrate and pre-TX jitter ceiling of the most constraining radio
interface. The run says which state it is in, on its own line:
[selftest] Link sizing: the daemon's `RNodeInterface[/dev/ttyUSB0]` (2734 bps, jitter ceiling 2926 ms)
[selftest] Phase 6: drain budget from the daemon's `RNodeInterface[…]` (2734 bps, …): \
10 frames x 147B at 2734 bps = 6.1s air (payload 5.1s +20% preamble/header/medium access) \
+ 2.9s handover (interface pre-TX jitter ceiling) = 9.1s
Without -c, against a daemon that reports no radio, or against one
that does not report a jitter ceiling (a Python rnsd, or an older
lnsd), it degrades rather than guessing: the reason is printed on the
same line, and the phase falls back to the fixed wait, or to the
airtime term alone when only the ceiling is missing.
diag
Collect a self-contained diagnostic bundle from a running daemon for bug
reports. diag queries the shared-instance RPC for the daemon's live
view (interface stats, path table, link count), bundles it with the
secret-redacted config, version and build info, and system info, and
prints to stdout (or to a file with --output).
Usage: lnstest diag [OPTIONS]
Options:
-c, --config <CONFIG>
Config directory of the daemon to ask
--output <OUTPUT>
Write the bundle to this path instead of stdout
--instance-name <INSTANCE_NAME>
Shared-instance name to query (default: from config, else "default")
--event-log <EVENT_LOG>
Tail this structured event-log file into the bundle
--no-rpc
Skip the daemon RPC queries; emit only config / versions / system
--no-rpc (lnstest.rs:419-421) skips the daemon queries — useful for
checking config parse status when the daemon is down.
A bundle is assembled from these sections in order (diag.rs:63-190):
- Versions / build —
lnstestversion, build profile, target. (The daemon version is not exposed by the RPC; check the daemon's startup log if needed.) - Config — config dir and file, parse status, then the effective config rendered as TOML with secrets redacted. The raw file is never included.
- Interfaces (configured) — each configured interface from the parsed config.
- Daemon view (shared-instance RPC) — instance name, RPC socket
path
\0rns/<name>/rpc, and liveinterface_stats,path_table, andlink_countqueries. - System — OS, kernel, distro, and the daemon's pid / RSS / open fds.
- Recent events — tail of the structured event log if
--event-logis given.
A trimmed bundle (your transport id, paths, and counters will differ):
===== Leviculum diagnostic bundle =====
----- Versions / build -----
lnstest version: 0.7.0
build profile: release
target: x86_64 / linux
daemon version: not exposed by the shared-instance RPC ...
----- Config -----
config dir: /etc/reticulum
config file: /etc/reticulum/config
config file: present, parsed OK
Effective config (TOML, secrets redacted; the raw file is NOT included
because it may contain secrets):
[reticulum]
enable_transport = true
shared_instance = true
instance_name = "default"
...
----- Daemon view (shared-instance RPC) -----
instance name: default
RPC socket: \0rns/default/rpc
authkey: derived from /etc/reticulum/storage/transport_identity (not shown)
## interface_stats
transport id: 0123456789abcdef0123456789abcdef
daemon uptime: 12m 34s (754s)
interfaces (3):
- Shared Instance[rns/default] type=LocalServerInterface status=up rxb=0 txb=0 clients=1
- AutoInterface[Default Interface/eth0/aabbccdd] type=AutoInterface status=up rxb=482 txb=917 peers=2
- TCPInterface[RNS TCP Node Germany 002/193.26.158.230:4965] type=TCPClientInterface status=up rxb=14211 txb=8332
## path_table
known paths: 7
[ ... JSON array of {hash, via, hops, expires, ...} ... ]
## link_count
relayed links: 0
----- System -----
os: linux kernel: 6.12.73+deb13-amd64
distro: Debian GNU/Linux 13 (trixie)
lnsd pid: 12345
lnsd VmRSS: 18432 kB
lnsd open fds: 27
----- Recent events -----
No structured event-log file specified ...
===== end of diagnostic bundle =====
Secret redaction
The bundle is designed to be safe to attach to a public tracker. IFAC
passphrase and networkname are redacted before the config is
serialised, and the node's private key is never read into the bundle —
only its hash is used internally to derive the RPC authkey
(diag.rs:119-128, 162-168). The bundle still contains your
hostnames, configured TCP targets, byte counters, and known-paths table,
so review it once before posting if your topology is sensitive.
lnstest diag --config /etc/reticulum --output /tmp/lnstest-diag.txt
See the lnsd Quickstart for how to read the bundle as a health check.
Network status, paths, and probing
lnstest deliberately does not reimplement Python's rnstatus,
rnpath, or rnprobe. For a running daemon's status, paths, and
interfaces, read the interface_stats and path_table sections of
lnstest diag, or point the Python rnstatus / rnpath / rnprobe
tools at the same shared instance — they attach to lnsd transparently.
lncp
lncp is the Reticulum file-transfer tool. It sends a file to a
destination or listens for incoming transfers, and is wire-compatible
with Python's rncp — an lncp listener accepts an rncp sender and
vice versa.
Reticulum File Transfer Utility
Usage: lncp [OPTIONS] [FILE] [DESTINATION]
Arguments:
[FILE] File to send (send mode)
[DESTINATION] Destination hash, 32 hex characters (send mode)
Options:
--config <CONFIG> Path to alternative Reticulum config directory
-v, --verbose... Increase verbosity
-q, --quiet... Decrease verbosity
-l, --listen Listen for incoming transfer requests
-w <TIMEOUT> Fetch / transfer phase timeout in seconds
-s, --save <SAVE> Save received files in specified path
-O, --overwrite Allow overwriting received files
-n, --no-auth Accept requests from anyone
-b <ANNOUNCE_INTERVAL> Announce interval (-1=none, 0=once at startup, N=every N sec) [default: 0]
-p, --print-identity Print identity and destination info and exit
-i <IDENTITY> Path to identity file to use
-S, --silent Fully silent: no progress output and no log output at all (equivalent to -qq)
-C, --no-compress Disable automatic compression
-f, --fetch Fetch file from remote listener
-F, --allow-fetch Allow authenticated clients to fetch files
-j, --jail <JAIL> Restrict fetch requests to specified path
-P, --phy-rates Display physical layer transfer rates
-a <ALLOWED> Allow identity hash (can be specified multiple times)
-h, --help Print help
-V, --version Print version
Modes
Send
Give a FILE and a 32-hex-character DESTINATION hash. lncp
establishes a link to the destination and transfers the file:
lncp report.pdf 0123456789abcdef0123456789abcdef
The file is compressed automatically unless you pass -C/--no-compress.
Listen
-l/--listen waits for incoming transfer requests. The listener prints
its own destination hash so the sender knows where to aim:
lncp -l -s ~/incoming
Fetch
-f/--fetch pulls a file from a remote listener instead of pushing to
it (the listener must allow this with -F/--allow-fetch).
Options
| Option | Meaning |
|---|---|
--config <DIR> | Use an alternative Reticulum config directory instead of the default lookup. |
-v/--verbose, -q/--quiet | Raise / lower log verbosity (stackable). |
-l/--listen | Listen for incoming transfer requests. |
-w <TIMEOUT> | Fetch/transfer phase timeout in seconds, counted after the link is established. Default: no timeout — the transfer runs to completion or until interrupted. Slow transports (LoRa) need no artificial cap; set this only for a hard wall-clock bound. |
-s/--save <PATH> | Save received files in this directory (listen mode). |
-O/--overwrite | Allow overwriting existing received files. |
-n/--no-auth | Accept requests from anyone (overrides -a). |
-b <INTERVAL> | Announce interval: -1 never, 0 once at startup, N every N seconds. Default: 0. |
-p/--print-identity | Print the destination hash and identity hash, then exit. |
-i <IDENTITY> | Use this identity file instead of the default. |
-S/--silent | Fully silent: no progress and no log output (equivalent to -qq). |
-C/--no-compress | Disable automatic compression. |
-f/--fetch | Fetch a file from a remote listener (instead of sending). |
-F/--allow-fetch | Allow authenticated clients to fetch files (listen mode). |
-j/--jail <PATH> | Restrict fetch requests to this directory (use with -F). |
-P/--phy-rates | Display physical-layer transfer rates. |
-a <HASH> | Allow a specific identity hash; repeatable to allow several. |
A few options only make sense in listen mode and warn otherwise:
-F/--allow-fetch warns when no -l is given, and -j/--jail warns
without -F (lncp.rs:199-204). -n/--no-auth overrides any -a
allow-list (lncp.rs:213-214).
Identity and authorisation
-p/--print-identity loads (or generates) the identity and prints the
rncp receive destination hash followed by the identity hash, then
exits (lncp.rs:334-352):
lncp -p
0123456789abcdef0123456789abcdef
Identity : fedcba9876543210fedcba9876543210
By default a listener only accepts senders whose identity hash you have
allowed with -a (repeatable). -n/--no-auth drops that check and
accepts anyone. An -a hash must be 32 hex characters / 16 bytes
(lncp.rs:314-332).
Examples
Send a file
On the receiver, start a listener and note its destination hash:
lncp -l -s ~/incoming -O
On the sender, transfer the file to that hash:
lncp ./report.pdf 0123456789abcdef0123456789abcdef
Listen and receive with authorisation
Allow only one known sender, saving into ~/incoming and showing
physical-layer rates:
lncp -l -s ~/incoming -a fedcba9876543210fedcba9876543210 -P
The sender finds its own identity hash with lncp -p.
lncp needs a running daemon (lnsd or rnsd) on the same shared
instance to reach the mesh — see the
lnsd Quickstart.
lnstatus
lnstatus is the Reticulum status tool. It shows the interfaces of a
running daemon and their traffic, announce, path-request, and link
statistics, and is output-compatible with Python's rnstatus — fed the
same interface_stats from the same daemon, lnstatus and rnstatus
render byte-identical output, so lnstatus | diff rnstatus passes.
Reticulum Network Stack Status
Usage: lnstatus [OPTIONS] [FILTER]
Arguments:
[FILTER] only display interfaces with names including filter
Options:
--config <CONFIG>
path to alternative Reticulum config directory
-a, --all
show all interfaces
-A, --announce-stats
show announce stats
-P, --pr-stats
show path request stats
-l, --link-stats
show link stats
-B, --burst
only show interfaces with active bursts
-t, --totals
display traffic totals
-s, --sort <SORT>
sort interfaces by [rate, traffic, rx, tx, rxs, txs, announces, arx, atx, prx, ptx, held]
-r, --reverse
reverse sorting
-j, --json
output in JSON format
--tables
add the transport's internal tables and collection sizes to the JSON output (requires -j)
-R <REMOTE>
transport identity hash of remote instance to get status from
-i <IDENTITY>
path to identity used for remote management
-w <TIMEOUT>
timeout before giving up on remote queries
-d, --discovered
list discovered interfaces
-D
show details and config entries for discovered interfaces
-m, --monitor
continuously monitor status
-I, --monitor-interval <MONITOR_INTERVAL>
refresh interval for monitor mode (default: 1) [default: 1]
-v, --verbose...
verbose logging (repeatable)
--instance-name <INSTANCE_NAME>
shared-instance name to query (default: from config, else "default")
-h, --help
Print help (see more with '--help')
-V, --version
Print version
lnstatus needs a running daemon (lnsd or rnsd) on the same shared
instance to query — see the
lnsd Quickstart.
Running it against a daemon
With lnsd (or Python rnsd) running, lnstatus with no arguments
prints every up interface and its counters:
lnstatus
It resolves the daemon exactly like the other tools: the config
directory (default lookup, or --config <DIR>) gives the shared-instance
name and the RPC authkey. If no shared instance is reachable, it reports
No shared RNS instance available to get status from and exits non-zero.
Give a FILTER to restrict the output to interfaces whose name contains
it:
lnstatus eth
Common flags
Extra statistics
-A/--announce-stats and -P/--pr-stats add announce and path-request
columns; -l/--link-stats adds link counts (queried separately from the
daemon); -t/--totals appends traffic totals:
lnstatus -A -P -l -t
-a/--all also shows interfaces that are currently down, and
-B/--burst restricts the output to interfaces with active bursts.
Sorting
-s/--sort <KEY> orders the interfaces by one of rate, traffic,
rx, tx, rxs, txs, announces, arx, atx, prx, ptx, or
held; -r/--reverse flips the order:
lnstatus -s traffic -r
Monitor mode
-m/--monitor clears the screen and re-renders on each interval;
-I/--monitor-interval <SECONDS> sets the refresh period (default 1):
lnstatus -m -I 2
JSON output
-j/--json emits the status as JSON instead of the rendered table, for
scripting:
lnstatus -j
The transport's tables
--tables adds one key, transport_tables, to that JSON object. It answers
how big every table the transport maintains is — path_table,
reverse_table, link_table (relayed links), announce_table,
announce_cache, tunnels, and local_links (links this node terminates) —
as a table_sizes list of {name, entries}:
lnstatus -j --tables
The rows themselves are a separate ask, --table-rows, which names the
tables you want them from (all for every one):
lnstatus -j --tables --table-rows path_table
lnstatus -j --tables --table-rows all
The split is about what the query costs the daemon. Answering a size is a
len(); answering with rows makes the daemon build one dictionary per row
before it can send anything, which on a node with 11 000 paths and 43 000
reverse entries was measured at 83 MB of daemon memory for a single call. A
status poll that only wanted to know how full the tables were was moving the
daemon's resident set by tens of megabytes, repeatedly. Asking for sizes now
costs about 35 KB regardless of how large the tables are; the rows cost what
they cost, to whoever actually wants them.
A table you did not ask rows for is absent from the response, not present
as an empty list, because an empty list is how this key says "the table is
empty". rows_for names the tables whose rows the response does carry, so a
reader never has to infer it. --table-rows requires --tables.
Beside the tables it carries collections: one row per collection the
daemon's storage holds, with name, entries and capacity (null where
the collection has no configured ceiling). It covers all of them, not just
the seven dumped above — including the packet dedup cache, which is the
largest structure in the daemon and appears as its two generations
packet_cache and packet_cache_prev rather than as a sum, because a
rotation frees one generation whole and a sum does not move when it happens.
That is how a resident set that steps up and falls back gets attributed to a
structure instead of guessed at.
rnstatus has no counterpart, so both flags require -j and never change
what a reference flag prints. lnstatus -j on its own is exactly what it was.
Two timestamps in there answer different questions. timestamp is our
clock — when this node learned the row. announce_emitted is the announcing
node's clock — the second it stamped into its announce, which is what peers
order competing announces by.
A daemon that does not implement the query — a Python rnsd, or an lnsd
older than this flag — makes lnstatus omit the key, print why on stderr, and
exit 0; the status you asked for is still printed. So an absent
transport_tables key means "this daemon cannot answer", while a present
key means it can — and inside it, a table named in rows_for whose list is
empty really is empty. Check for the key before reading it, and do not treat
its absence as an empty table.
Full field lists are in lnstatus(1).
Remote status (-R/-i/-w)
-R <hash> queries a remote transport instance's status over a link,
the way rnstatus -R does, and feeds the result to the same renderer,
so remote and local output match (run_remote (lnstatus.rs:453)).
<hash> is the remote instance's transport identity hash (32 hex
characters). -i <file> names the management identity and is
mandatory; it is proven to the remote over the link, so the remote
daemon only answers if it has remote management enabled and lists that
identity as allowed. -w <seconds> bounds the query (default 15,
matching Python's path-request timeout). -m re-queries the remote on
the monitor interval, and -j renders the remote status as JSON:
lnstatus -R 76fe5751a56067d1e84eef3e88eab85b -i ~/.reticulum/identities/mgmt -w 30
Discovered interfaces (-d/-D)
-d lists the interfaces this daemon has discovered on the network, in
the rnstatus discovered layout; -D renders the detailed layout with
ready-to-paste config entries (run_discovered (lnstatus.rs:362)).
Both read the local daemon's discovered-interface registry over the
shared-instance RPC and honour FILTER and -j:
lnstatus -D rnode
Examples
Full picture of a local daemon
lnstatus -a -A -P -l -t
Watch one interface live
lnstatus -m -I 2 rnode
lnstatus needs a running daemon (lnsd or rnsd) on the same shared
instance to reach the mesh — see the
lnsd Quickstart.
Field testing
A field walk leaves its only evidence on the boards' debug ports. A board that misbehaves in the field and is questioned afterwards has nothing to say: its debug log is not stored anywhere. What the log said during the walk is the measurement, so reading it is part of the walk, not an afterthought.
lnflash --watch is the shipped reader. It opens a board's debug CDC
(if00) with DTR and RTS raised (the firmware only transmits with both
set), prefixes every line with a wall-clock ISO-8601 timestamp with
milliseconds, appends to a file flushed per line, and survives the
board resetting, being reflashed or losing USB mid-walk: the gap is
logged as its own [WATCH] line and the port is reopened with a
bounded backoff. It never exits on EOF.
The sequence
-
Start the watch before the walk. With one board attached:
lnflash --watch --out walk-$(date +%F).logWith several, name the board's USB serial (or bus port):
lnflash --watch 183004F712B4A7FE --out walk-$(date +%F).log -
Note the file. The watch file is the walk's evidence; a walk whose log file cannot be named afterwards did not happen. Leave the watch running for the whole walk — the reconnect handling exists so a board reset in the field does not end the recording.
-
Summarize after:
lnflash --summarize walk-2026-09-04.logprints, per hour, how many LoRa receptions of each class the file holds (announce, data, path request), the last line seen per class, and how many reconnect gaps the watch bridged. Six announces and zero data lines in fifteen minutes of walking is a finding (Codeberg #365); the summary makes it visible in one read.
The watch never filters — classification lives entirely in
--summarize, so a wrong classifier can be fixed and re-run over the
same evidence.
A watch file line looks like this, wall clock first, the board's own
t= untouched:
2026-09-04T15:22:27.101+02:00 [LORA] RX 183 bytes rssi=-69 snr=5
lnflash --watch is not a daemon and has no background mode. For a
watch that outlives the terminal, run it under nohup or in a tmux
session yourself.
lnsd(1)
NAME
lnsd -- Reticulum network daemon
SYNOPSIS
lnsd [-c dir] [--storage dir] [-s] [--exampleconfig] [-v...] [-q...]
DESCRIPTION
lnsd runs the Reticulum network stack as a long-lived daemon process. It is a drop-in replacement for Python's rnsd. Other programs connect to it via shared instance IPC (Unix abstract socket).
On startup, lnsd reads config from the configuration directory, opens all configured interfaces, and begins routing packets. It keeps running until it receives SIGINT or SIGTERM.
Sending SIGUSR1 prints a diagnostic dump of internal state to stderr.
OPTIONS
-c, --config dir
: Path to the Reticulum configuration directory, the way rnsd --config takes one. The config file is <dir>/config. Without this option the default lookup order applies; see FILES.
--storage dir
: Storage directory path. Overrides the config file's storage_path, which in turn overrides the default <config_dir>/storage. It moves the daemon alone — the client tools read storage_path from the config — so a storage directory shared with lnstatus, lncp, lnpath or lnprobe belongs in the config file. Long-only on purpose: in rnsd, -s means --service, so the short letter stays reserved for that and the storage override is a Leviculum extension.
-s, --service
: Declare that the daemon is running as a service, accepted for compatibility with rnsd -s. lnsd keeps logging to standard output, which journald captures; it does not redirect to a log file.
--exampleconfig
: Print an example configuration to standard output and exit, like rnsd --exampleconfig. The output loads through lnsd's own config loader.
-v, --verbose : Increase log verbosity. Once for debug, twice for trace.
-q, --quiet : Decrease log verbosity. Once for warnings only, twice for errors only.
ENVIRONMENT
RUST_LOG
: Overrides the verbosity flags. See the tracing-subscriber documentation for filter syntax.
FILES
/etc/reticulum/config : System-wide configuration, used when it exists. The Debian package installs one here, which is also what lets Python clients find the running daemon without extra flags.
~/.config/reticulum/config : Per-user configuration, used when the system-wide file is absent.
~/.reticulum/config : Final fallback (INI format, same as Python Reticulum).
<config_dir>/storage/ : Storage directory for identities, known destinations, and cached path state.
The three configuration directories are tried in that order, matching Python-Reticulum's own lookup.
SIGNALS
SIGINT, SIGTERM : Graceful shutdown.
SIGUSR1 : Dump diagnostic state to stderr; see DIAGNOSTIC DUMP.
DIAGNOSTIC DUMP
The SIGUSR1 dump lists the estimated size of every tracked table, the === Total estimated === line, the process RSS, and three sections that say where the rest of the memory sits.
=== Allocator === prints glibc's own balance sheet (mallinfo2: arena, ordblks, hblks, hblkhd, uordblks, fordblks, keepcost) and the thread count. The shipped musl-static builds run musl's mallocng, which keeps no such totals, so there the section says it is unavailable and only the thread count is measured.
=== Mappings === buckets the process's anonymous writable mappings by size (under 64 KiB, 64 KiB to 1 MiB, 1 to 8 MiB, 8 to 64 MiB, above) with count and bytes per bucket, then names the ten largest by address. Thread stacks show up as regions of about 2 MiB each, so a growing count in one bucket points at an owner class without a debugger.
The gap: line starts from RSS plus swap and subtracts the tracked total, the allocator's free bytes (fordblks + keepcost) and its mmapped bytes (hblkhd) in turn, printing what is left after each. A term the libc cannot report is printed as n/a and subtracts nothing; hblkhd counts live large allocations, so it can overlap the tracked total.
EXAMPLES
Start with default config and verbose logging:
lnsd -v
Start with a custom config directory:
lnsd --config /etc/reticulum
SEE ALSO
lnstest(1), lncp(1), lnstatus(1)
lnstest(1)
NAME
lnstest -- Reticulum test and diagnostics tool
SYNOPSIS
lnstest [-c dir] [-v] command [args...]
DESCRIPTION
lnstest is the test and diagnostics tool for the Leviculum Reticulum stack. It drives integration self-tests, collects diagnostic bundles from a running daemon, manages identities, and opens interactive sessions. It works against either lnsd or Python rnsd through the shared-instance interface. For file transfer use lncp(1).
GLOBAL OPTIONS
-c, --config dir : Path to the Reticulum configuration directory.
-v, --verbose : Enable verbose logging.
--corrupt-every n : Fault injection: corrupt roughly one byte per n bytes written to TCP. For selftest, corruption is held back until mutual discovery has completed, so that announces cross a clean stream.
COMMANDS
lnstest identity generate [-o file]
Generate a new Reticulum identity and write it to file.
lnstest identity show file
Show the hash and public keys of the identity in file.
lnstest connect addr
Open an interactive session to a Reticulum daemon at addr (host:port). Supports link establishment, message exchange, and announce discovery. Type /help in the session for available commands.
lnstest selftest target [target]
Run integration self-tests through one or two relay nodes. Tests link establishment, channel data, ratchet operation, and bulk transfer.
Options:
--duration seconds : Test duration (default: 180).
--rate n : Messages per second per direction (default: 1).
--mode mode : Which phases to run: all, link, packet, ratchet-basic, ratchet-enforced, bulk-transfer, ratchet-rotation (default: all).
--messages n : Messages per direction in the ratchet exchange (default: 10). Sizes --mode ratchet-basic and ratchet-enforced; --duration and --rate do not reach that count, because they size the link and single-packet phases, which no ratchet mode runs. The other two ratchet modes keep their own counts — bulk-transfer sends 100 each direction, ratchet-rotation 5 before and 5 after the key change — and print a line saying so when the flag is given.
Every single-packet phase waits for what is still in flight before it reads the receive counter. On a radio link that wait must be sized from the link, and the tool has no radio of its own, so give it the global -c/--config pointing at the config directory of the daemon that owns the radio: it reads that daemon's interface_stats and sizes each window from the reported on-air bitrate and pre-TX jitter ceiling. Without it — or against a daemon reporting no radio — the phases keep a fixed wait, and the run prints which state it is in.
lnstest diag
Collect a self-contained diagnostic bundle from a running lnsd (or rnsd) for attaching to bug reports: versions/build, the secret-redacted config and configured interfaces, the daemon's live view via the shared-instance RPC (interface stats, path table, link count), best-effort system info, and an event-log pointer. Printed to stdout by default. Use the global -c/--config to point at the daemon's config directory.
Secrets are redacted — IFAC passphrase and networkname never appear, and the node identity private key (storage/transport_identity) is never read into the bundle (it is used only to derive the RPC authkey). Queries the daemon doesn't support (e.g. when run against Python rnsd) are reported as unavailable rather than failing.
Options:
--output path : Write the bundle to path instead of stdout (a one-line confirmation is printed to stderr).
--instance-name name
: Shared-instance name to query (default: from the config, else default).
--event-log path
: Tail this structured event-log file into the bundle (when lnsd was started with LEVICULUM_EVENT_LOG set).
--no-rpc : Skip the daemon RPC queries; emit only the config, versions, and system sections.
EXAMPLES
Generate a new identity:
lnstest identity generate -o my_identity
Run a self-test through a relay node:
lnstest selftest 192.0.2.10:4965 --mode link --duration 60
Collect a diagnostic bundle from a running daemon:
lnstest diag -c /etc/reticulum --output /tmp/lnstest-diag.txt
SEE ALSO
lnsd(1), lncp(1), lnstatus(1)
lncp(1)
NAME
lncp -- Reticulum file transfer utility
SYNOPSIS
lncp [options] file destination lncp [options] -l [-s dir]
DESCRIPTION
lncp transfers files over the Reticulum network. It is compatible with Python's rncp. It connects to a running daemon (lnsd or rnsd) via shared instance IPC.
In send mode, lncp sends file to the node identified by destination (a 32-character hex hash). In listen mode (-l), it waits for incoming file transfer requests.
OPTIONS
file : File to send (send mode).
destination : Destination hash, 32 hex characters (send mode).
--config dir : Path to alternative Reticulum configuration directory.
-v, --verbose : Increase verbosity. Repeat for more detail.
-q, --quiet : Decrease verbosity.
-l, --listen : Listen for incoming transfer requests.
-w seconds : Time out the fetch or transfer phase after this many seconds, counted from the moment the link is established (link establishment has its own internal budget). There is no timeout by default: slow transports such as LoRa need none, so set this only when you want a hard wall-clock bound.
-s, --save dir : Save received files in the specified directory.
-O, --overwrite : Allow overwriting existing files when receiving.
-n, --no-auth : Accept requests from anyone (no authentication).
-a hash : Allow a specific identity hash. May be given more than once.
-b interval : Announce interval in seconds. -1 = never, 0 = once at startup, N = every N seconds (default: 0).
-p, --print-identity : Print identity and destination info and exit.
-i file : Path to identity file to use.
-S, --silent : Fully silent: no progress output and no log output at all, equivalent to -qq.
-C, --no-compress : Disable automatic compression.
-f, --fetch : Fetch file from remote listener instead of pushing.
-F, --allow-fetch : Allow authenticated clients to fetch files.
-j path : Restrict fetch requests to the specified path.
-P, --phy-rates : Display physical layer transfer rates.
EXAMPLES
Send a file:
lncp myfile.tar.gz a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
Listen for incoming files and save to a directory:
lncp -l -s ~/received/
Listen with verbose logging, accepting from anyone:
lncp -l -n -v
SEE ALSO
lnsd(1), lnstest(1), lnstatus(1)
lnstatus(1)
NAME
lnstatus -- Reticulum network stack status
SYNOPSIS
lnstatus [options] [filter]
DESCRIPTION
lnstatus displays the status of the interfaces on a running Reticulum daemon. It is compatible with Python's rnstatus and produces the same per-interface layout. It connects to a running daemon (lnsd or rnsd) via shared instance IPC, querying interface_stats (and link_count for -l), so lnstatus | diff rnstatus against the same daemon passes.
Without a filter, all up interfaces are shown. Give a filter string to only display interfaces whose name contains it.
With -R it queries a remote transport instance over a link, the way rnstatus -R does, and feeds the result to the same renderer, so remote and local output match. With -d/-D it reads the local discovered-interface registry over the RPC and renders the rnstatus discovered layout.
OPTIONS
filter : Only display interfaces whose name contains this string.
--config dir : Path to alternative Reticulum configuration directory.
--instance-name name
: Shared-instance name to query. Defaults to the value from the configuration file, otherwise default.
-a, --all : Show all interfaces, including those that are down.
-A, --announce-stats : Show announce statistics.
-P, --pr-stats : Show path request statistics.
-B, --burst : Only show interfaces with active bursts.
-l, --link-stats
: Show link statistics: the number of entries in the daemon's transport link table, i.e. the links it relays (queries link_count from the daemon, the same value rnstatus -l reads).
-t, --totals : Display traffic totals.
-s, --sort key
: Sort interfaces by key: rate, traffic, rx, tx, rxs, txs, announces, arx, atx, prx, ptx, or held.
-r, --reverse : Reverse the sort order.
-j, --json : Output in JSON format.
--tables
: Add the size of every table the transport maintains, and the entry count
of every collection its storage holds, to the JSON output as a
transport_tables object. Sizes only; the rows are asked for with
--table-rows. Requires -j; not available with -R or
-d/-D. Leviculum extension — rnstatus has no counterpart, and a
daemon that does not implement it (a Python rnsd, or an older lnsd)
causes the key to be omitted, with a note on stderr and exit status 0.
See TRANSPORT TABLES below.
--table-rows TABLE[,TABLE...]
: Also include the ROWS of the named tables: path_table, reverse_table,
link_table, announce_table, announce_cache, tunnels,
local_links, or all for every one. Repeatable, and accepts a
comma-separated list. Requires --tables. A table not named here is
absent from the response rather than present and empty; its size is in
table_sizes either way. Rows are the expensive half of this query — see
WHAT THE QUERY COSTS below.
-N, --identities
: List every identity the daemon has learned from announces, one row per
announced destination, plus a derived: line with the lxmf.delivery
and rnstransport.probe destinations computed from the identity hash.
Leviculum extension — rnstatus has no counterpart, and a daemon that
does not implement the query (a Python rnsd, or an older lnsd)
causes an error and exit status 2. With -j the raw response is
printed instead. Not available with -R, -d/-D, --tables
or -m. See IDENTITY LISTING below.
-m, --monitor : Continuously monitor status, clearing and redrawing on each interval.
-I, --monitor-interval seconds : Refresh interval for monitor mode (default: 1).
-v, --verbose : Increase verbosity. Repeat for more detail.
--version : Print version and exit.
-R hash : Transport identity hash of a remote instance to query instead of the local one.
-i file : Identity used for remote management.
-w seconds : Timeout before giving up on remote queries.
-d, --discovered : List interfaces discovered on the network.
-D : Show details and config entries for discovered interfaces.
EXIT STATUS
0 : Success.
1 : No shared RNS instance available to get status from (could not derive the RPC authkey).
2 : The status query failed.
20 : Remote management (-R) was requested but the management identity is missing or unusable.
EXAMPLES
Show all interfaces:
lnstatus
Show announce and path request statistics for interfaces named like eth:
lnstatus -A -P eth
Sort interfaces by traffic, most first:
lnstatus -s traffic -r
Continuously monitor, refreshing every two seconds:
lnstatus -m -I 2
Emit machine-readable JSON:
lnstatus -j
Emit JSON with the size of every transport table and storage collection:
lnstatus -j --tables
Emit JSON with the path table's rows as well:
lnstatus -j --tables --table-rows path_table
List the identities heard from announces, with derived destinations ready to paste into lnprobe:
lnstatus -N
IDENTITY LISTING
rnpath -t shows destination hashes only, but probing a remote transport node
needs its rnstransport.probe destination, which is derived from its identity
hash — and the daemon knows that identity, because the announce carried the
public key. -N exposes it: one row per announced destination the daemon
still holds, with the identity hash, the announced destination hash, the name
(only when it matches an aspect the daemon registered itself; ? otherwise —
never guessed), and the live path toward it (hops, via as
interface/next-hop, last seen). Columns without a live path show -.
Under each row a derived: line prints the destinations computed from the
identity hash as sha256(sha256(name)[:10] + identity_hash)[:16] for the two
names that matter in practice: lxmf.delivery and rnstransport.probe.
Hashes are printed in full so they can be pasted into lnprobe or rnprobe.
The inventory is the daemon's announce cache (Python's equivalent store is
Identity.known_destinations), which is bounded by the announce-cache
cleaning the daemon already performs; an identity whose cached announce has
been evicted no longer appears.
TRANSPORT TABLES
With --tables, the -j object gains one additional key,
transport_tables. Nothing else about the output changes, so anything that
parses lnstatus -j today keeps working.
The object always holds table_sizes, rows_for and collections, plus one
list of rows per table named in --table-rows:
table_sizes
: How many rows each of the seven tables below has, as one {name, entries}
row per table, whether or not this response carries that table's rows.
Every entry is a len().
rows_for
: The tables whose rows this response carries — what --table-rows asked
for. Empty by default. A table not listed here has no key in the object;
a table listed here with an empty list is empty.
collections
: How large every collection the daemon's storage holds currently is — not
only the tables dumped below. One row per collection: name (the field
name in the storage, so a row can be read against the source), entries
(live count) and capacity (the ceiling the daemon enforces, or null
where it enforces none). The packet dedup cache appears as its two
generations, packet_cache and packet_cache_prev, never as a sum: a
rotation frees one generation whole, and a sum is flat across exactly
that event. A null capacity is an answer, not a gap — that collection
is bounded by expiry alone.
path_table
: Destinations this node knows a route to. Keys hash, timestamp, via,
hops, expires, interface are the same keys, with the same units, that
Python's own path_table RPC returns (Reticulum.get_path_table). Added:
announce_emitted.
reverse_table
: Where to send the reply to a packet this node forwarded: hash,
receiving_interface, outbound_interface, timestamp.
link_table
: Links this node relays: link_id, timestamp,
next_hop_interface, remaining_hops, receiving_interface, hops,
destination_hash, validated, proof_timeout.
announce_table
: Announces held for deferred rebroadcast: hash, timestamp,
retransmit_timeout, retries, receiving_interface, hops,
packet_length, local_rebroadcasts, block_rebroadcasts,
attached_interface.
announce_cache
: Known destinations whose last announce is still held: hash,
packet_length, retained, last_used.
tunnels
: Reconnectable peers and the paths held against them: tunnel_id,
interface, expires, and paths, each path carrying hash, hops,
via, expires, timestamp, announce_emitted.
local_links
: Links this node is an endpoint of — not the same table as link_table
above: link_id, state, destination_hash, age, interface.
Which clock a timestamp is
Two questions that are easy to confuse, and are answered by two different keys:
timestamp
: Our clock. When this node learned or last refreshed the row, in Unix
seconds. In path_table it is recovered from expires minus the lifetime
that path was granted, so it is exact for a path still on the interface it
was learned on.
announce_emitted
: The announcing node's clock. The whole-second emission stamp that node
wrote into its announce, as this node received it. Peers order competing
announces for one destination by this value, so it is a claim about a
remote machine's time, never about ours. 0 means no announce blob is
stored for the row.
What a count is for
entries without capacity does not say whether a node is near its limit,
which is the question an operator has, so the two travel together. Both are
reported, never enforced here: reading this changes no ceiling and adds none.
Use it to attribute memory. Multiply a count by what one entry of that collection costs and the products either account for the daemon's resident set or they do not — the difference is what separates a design that costs too much from a leak. Before this existed, seven tables of the twenty collections were visible, and the largest structure in the daemon, the dedup cache, was not among them.
What the query costs
A size is a len(). A row is not: the daemon builds one dictionary per row,
with a string key per field, before it can serialise anything, and the whole
structure is live at once. Measured on a node with 11 000 paths and 43 000
reverse entries, asking for those two tables' rows peaks at 83 MB of
daemon memory for one call; asking for sizes alone peaks at 35 KB, and
that figure does not move with the size of the tables.
This is why the rows are named rather than included. An operator polling a
node for how full its tables are — the common case, and the one that gets
polled in a loop — was moving the daemon's resident set by tens of megabytes
per call. If you want rows, ask for the table you want and not for all.
Absent is not empty
A daemon that implements the query answers with transport_tables present and
its tables possibly empty. A daemon that does not implement it causes the key
to be omitted entirely. Test the key's presence to tell the two apart; do
not read an absent key as an empty table.
SEE ALSO
lnsd(1), lnstest(1), lncp(1)
lnprobe(1)
NAME
lnprobe -- Reticulum probe utility
SYNOPSIS
lnprobe [options] full_name destination_hash
DESCRIPTION
lnprobe measures the reachability of a Reticulum destination. It is compatible with Python's rnprobe: the same command line, the same output, the same exit codes. It connects to a running daemon (lnsd or rnsd) via shared instance IPC, requests a path to the destination if none is known, then sends probe packets and reports the round-trip time and hop count taken from the delivery proof the probed destination signs for each probe.
full_name is the destination's full dotted name (for probing a transport node's probe responder: rnstransport.probe); destination_hash is its 32-character hexadecimal hash. The probed node answers only if it runs a probe responder (respond_to_probes in its configuration, or an LNode's built-in responder — the hash is on the board's [IDENTITY] boot line and in the lnflash --set-name read-back).
OPTIONS
--config dir : Path to alternative Reticulum configuration directory.
-s, --size bytes : Size of the probe packet payload in bytes. Default 16.
-n, --probes count : Number of probes to send. Default 1.
-t, --timeout seconds : Timeout before giving up, per probe and for the initial path request. Default 12 seconds plus the daemon's first-hop timeout for the destination, which scales with the next-hop interface's bitrate — a probe over a slow LoRa hop waits longer by default.
-w, --wait seconds : Time to wait between probes. Default 0.
-v, --verbose : Show the next hop and interface for each probe; repeat to raise log verbosity.
EXIT STATUS
0 when every probe was answered. 1 when no path to the destination could be found. 2 when at least one probe went unanswered (the summary line reports the loss). 3 when the requested probe size does not fit the Reticulum MTU.
EXAMPLES
Probe a transport node's probe responder:
lnprobe rnstransport.probe 6a1ab9ea64747f298c1f205dfcf0f5a3
Send 10 probes of 100 bytes, one second apart:
lnprobe -n 10 -s 100 -w 1 rnstransport.probe 6a1ab9ea64747f298c1f205dfcf0f5a3
SEE ALSO
lnsd(1), lnstest(1), lnstatus(1), lncp(1), lnpath(1)
lnpath(1)
NAME
lnpath -- Reticulum path query utility
SYNOPSIS
lnpath [options] destination_hash
DESCRIPTION
lnpath queries the path to a Reticulum destination, waits for it to arrive, and drops one a running daemon holds. It connects to a running daemon (lnsd or rnsd) via shared instance IPC. Without -d it requests a path if none is known, waits out the -w window, and reports the hop count, the next hop and the interface the traffic leaves on. With -d it removes the path from the daemon's table, which is the table that routes; the client's own copy dies with the process and dropping it would change nothing.
The tool implements the path-query verb of Python's rnpath -- query, wait, drop -- with the reference tool's arguments, output and exit codes for those three. It deliberately offers no flag for the reference tool's other roles: the path and rate views (-t, -r, -m), the blackhole administration verbs (-b, -B, -U, -p), and remote management of another instance (-R, -i, -W). Use rnpath for those, or lnstatus --tables, which prints the same path table rnpath -t shows plus the tables the reference tool cannot reach.
OPTIONS
--config dir : Path to alternative Reticulum configuration directory.
-d, --drop : Remove the path to the destination from the daemon's path table.
-w seconds
: Timeout before giving up on the path request. Default 15 seconds, the reference stack's PATH_REQUEST_TIMEOUT. This argument governs the whole wait: nothing is added to it and no floor is applied, so a caller that asks for two seconds waits two seconds and one that asks for sixty waits sixty.
-v, --verbose : Raise log verbosity. Diagnostics go to standard error; standard output carries only the verdict line.
EXIT STATUS
0 when a path was found, or when one was dropped. 1 when no path was found within the window, when the destination argument is not a 32-character hexadecimal hash, when there was no path to drop, or when the daemon could not be reached.
EXAMPLES
Query a path, waiting up to the default 15 seconds:
lnpath 6a1ab9ea64747f298c1f205dfcf0f5a3
Wait a minute for a path over a slow LoRa hop:
lnpath -w 60 6a1ab9ea64747f298c1f205dfcf0f5a3
Drop a path so the next request rediscovers it:
lnpath -d 6a1ab9ea64747f298c1f205dfcf0f5a3
SEE ALSO
lnsd(1), lnstest(1), lnstatus(1), lncp(1), lnprobe(1)
lnomad(1)
NAME
lnomad -- terminal browser for NomadNet micron pages
SYNOPSIS
lnomad [options] [url]
DESCRIPTION
lnomad fetches and renders NomadNet micron pages over Reticulum, either interactively in a terminal UI or once to standard output with --print. It connects to a running daemon (lnsd or rnsd) through the shared instance, so one of them must be running.
Node discovery runs continuously while the UI is up: nomadnetwork.node announces are folded into the places panel (d) as they arrive. Started without a url on a terminal, lnomad opens its start screen with that panel showing; without a terminal a url is required.
A /file/ URL downloads instead of rendering; see --output for where the file lands.
Pictures are drawn in the page. Micron has no image construct, so a page can only link to a picture in its node's file area; lnomad recognises such a link by its target — a /file/ path named .png, .jpg, .jpeg or .gif — and, when the link stands alone on its line, fetches it and draws it there. The terminal's own graphics protocol is used when it has one (Kitty, iTerm2 or Sixel); otherwise the picture is drawn with Unicode half-blocks, and where neither is possible a line naming the file, its format and its size takes its place. A page's pictures are fetched after it is displayed, one at a time, so nothing waits for them, Esc cancels the rest, and no more than eight are fetched from any one page. Fetched pictures are kept in memory (see --image-cache), so stepping back to a page shows them again without spending the airtime twice. Use --images off to fetch none at all.
With a picture focused (Tab), Enter saves it to the download directory and o opens it in the system viewer. Neither transfers it a second time.
OPTIONS
url
: Page to open, as <address>[:/page/x.mu[`f=v|...]], where address is the node's 32-hex-character destination address and everything after the : is the request path. A bare address opens the node's default page. Leaving the address out but keeping the : (:/page/x.mu) makes it a URL to a local page — local to the page in view — so it is only meaningful inside the UI, never on the command line. A bare request path with no : names no node and is rejected, as it is by NomadNet. Optional on a terminal, where omitting it opens the start screen.
--config dir : Reticulum configuration directory. Defaults to the platform default, the same one lncp(1) uses.
--instance name : Shared-instance name to connect to, overriding the configuration file.
--print : Fetch, render and print the page once, then exit. Non-interactive.
--output path
: Where a /file/ download is saved. An existing directory, or a path ending in /, receives the file under the name the server sent, falling back to the URL basename; any other path names the exact file to write. By default the file lands in the current directory, and an existing file is preserved by appending (1), (2) and so on.
--width columns : Render width. Defaults to the detected terminal width, otherwise 80.
--timeout seconds : Per-request fetch timeout (default: 30).
--images mode
: Inline images. auto draws a page's pictures where they are linked, using the terminal's graphics protocol or Unicode half-blocks; off leaves every image link an ordinary link and fetches nothing. Ignored with --print and when output is not a terminal, neither of which draws pictures (default: auto).
--image-cache megabytes
: How much fetched image data to keep in memory for revisits (default: 10). A page visited again then costs no airtime for pictures it has already shown. When the budget is exceeded, the pictures untouched longest are dropped until it fits again; a single picture larger than the whole budget is never kept. 0 disables the cache. Memory only — nothing is written to disk.
--no-color : Disable ANSI colour in the rendered output. Also suppresses inline images: a picture painted out of coloured blocks is not what a reader asking for no colour had in mind.
--theme theme
: Colour theme for the interactive UI: auto detects the terminal background, light and dark force a theme. Ignored with --print and when output is not a terminal (default: auto).
--color depth
: Terminal colour depth. auto picks true colour when COLORTERM is truecolor or 24bit and otherwise falls back to the xterm-256 palette; truecolor and 256 force the depth. --no-color still overrides this (default: auto).
ENVIRONMENT
COLORTERM
: Consulted by --color auto to choose between true colour and the xterm-256 palette.
XDG_CONFIG_HOME
: Base for lnomad's own directory (lnomad/), holding the bookmarks (bookmarks.toml), the per-node identify decisions (identify.toml) and the identity lnomad reveals when identifying (identity). Defaults to ~/.config.
XDG_DOWNLOAD_DIR
: Where a /file/ link followed inside the UI is saved, inline pictures included. Defaults to $HOME/Downloads. The --output option covers downloads started from the command line instead.
EXAMPLES
Open a node's default page interactively:
lnomad a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
Print a specific page without entering the UI:
lnomad --print a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4:/page/index.mu
Open the start screen and pick a node from the places panel:
lnomad
Download a file to a chosen directory:
lnomad --output ~/downloads/ a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4:/file/manual.pdf
SEE ALSO
lnsd(1), lblogd(1), lncp(1)
lblogd(1)
NAME
lblogd -- dev blog server, on the web and on NomadNet
SYNOPSIS
lblogd --config file lblogd --config file --print-hash
DESCRIPTION
lblogd serves a directory of Markdown posts on two sides at once: as a NomadNet page node over Reticulum, and as a web server on the clearnet. Posts are plain Markdown files with an optional TOML frontmatter block; adding a file and reloading the service publishes it to both sides.
The NomadNet side is a shared-instance client, so a Reticulum daemon must be running — either lnsd(1) or Python's rnsd — under the instance name named in the configuration file. It need not be running yet: at startup lblogd waits up to a minute for the daemon's IPC socket, which is what gets it through a boot where both start together and the daemon binds its socket a few seconds later. After that minute it exits, and the packaged service restarts it until a daemon answers.
A post may use the whole standard Markdown feature set, basic and extended: tables, footnotes, strikethrough, task lists, definition lists, heading identifiers, highlights, sub- and superscript, and bare URLs and e-mail addresses, which become links without being written as such. The web side renders all of it as HTML. Micron, the NomadNet page format, has fewer constructs than HTML, so a few degrade on the mesh: struck text is dimmed, a highlight becomes a background colour, and H~2~O and X^2^ use the Unicode subscripts and superscripts where they exist and keep their markers where they do not. Footnotes lose only the jump, not their content: the reference stays as [1] and the definitions are collected behind a divider at the end of the page. A heading identifier has nothing to attach itself to on the mesh and is dropped there. Emoji shortcodes such as :joy: are not translated on either side and stay as written.
Images travel as files. Micron, the NomadNet page format, has no image construct at all, so a picture referenced from a post is published as a file and linked from the page: NomadNet saves it to the reader's download directory, and lnomad(1) draws it inline. On the web the same reference becomes an ordinary <img>. Write  in a post and put mast.jpg in the file area; the same name is then served as /files/mast.jpg over HTTP and /file/mast.jpg over Reticulum. ./mast.jpg, files/mast.jpg and /files/mast.jpg all name that file. A reference with a scheme is left alone: nothing on the mesh can fetch an https:// image, so it degrades to its alt text there and stays an external image on the web.
The area is flat, and a requested name can carry no path separator, so no request can reach outside it. max_file_bytes bounds a single file, 10 MiB by default; anything larger is skipped with a line on standard error rather than served, because over a LoRa interface an unbounded transfer denies service to every other reader of the node for as long as it runs.
A domain usually says more than "here are my posts". pages_dir holds Markdown pages that are not entries — a landing page, a page about the code — parsed in the post format and rendered like the about page: no date, no byline, never in the post index, never in the feed. Each is served at /<name> on the web and /page/<name>.mu on the mesh. The name index is special: with a pages_dir/index.md the site gets a landing page at / and /page/index.mu, and the post index moves down to /blog and /page/blog.mu. Without one nothing moves, so a configuration that names no pages_dir serves exactly what it served before any of this existed. Posts keep /posts/<slug> and the feed keeps /feed.xml on purpose: a feed entry is identified by its URL, so moving posts would show every entry again as new in every reader.
The [links] section is the other half — names that point at something off this server:
[links]
code = "https://codeberg.org/Lew_Palm/leviculum"
issues = "https://codeberg.org/Lew_Palm/leviculum/issues"
Each becomes /<name> on the web and /page/<name>.mu on the mesh, and the two answer differently because they must. The web answers 302, not 301: a forge changes host, and a browser that cached a 301 would keep going to the old one long after the configuration said otherwise. The mesh answers with a short page naming the URL as text, because a NomadNet client cannot follow a web link at all. A name that is both a link and a page in pages_dir shows that page's text above the URL on the mesh, while the web still redirects. A link target must be an absolute http:// or https:// URL.
Every page, on both sides, carries a small nav line: the landing page, the blog, then each page by name and each link in the order the configuration lists them. With neither pages nor links there is nothing to put in it and none is emitted.
The web's top level and the mesh's /page/<name>.mu are shared namespaces, so page and link names are checked against one list: the routes blog, posts, files and feed.xml; about, when an about page is configured; every post's slug, which owns /page/<slug>.mu; and index for a link, since as a page name that is the landing page. A collision is a startup error naming both the offender and what already answers there; on a reload it is refused and the previous content keeps serving. Names follow the same slug rules as a post: plain lowercase ASCII letters, digits and hyphens. Unlike the file area, pages_dir is not optional by existence — the operator named it, so a directory that is not there is a startup error rather than a site that quietly lost its landing page.
Every page, on both sides, also carries the source offer AGPL section 13 requires: the running version, the licence, and where the source is — a link on the web, the bare URL as text on the mesh. It needs no configuration and cannot be switched off, because a reader of a served page is a user interacting with AGPL software over a network and is owed the Corresponding Source. The [source] url key exists for the one case the compiled-in default gets wrong: an operator running a modified lblogd owes their readers their own tree, not this project's repository. An empty value is a startup error rather than a footer that offers nothing.
The web side either obtains its own certificate from Let's Encrypt, or runs plain behind a reverse proxy that terminates TLS. Note that the canonical page URL and the Atom feed are derived from the configured domains list even when certificate handling is switched off, so a deployment behind a proxy still has to set that list.
COUNTING
lblogd appends one record per day to a counts file — counts.log under data_dir unless [counter] path says otherwise, and not at all if [counter] enabled is false. There is no UI and no HTTP endpoint; the file is the interface:
DAY date=2026-08-07 tz=UTC mesh_requests=41 mesh_sessions=12 mesh_identified_requests=0 web_requests=308 web_not_found=57 clock_behind=0 written=2026-08-07T23:59:12Z
The counts are requests and links, and are named as such. They are not visitors. On the mesh a request arrives on a Reticulum link, and a link is a session, not a person: one reader browsing five pages over one link is one session and five requests, and the same reader tomorrow is a different session. Reticulum discloses who a peer is only when the peer chooses to identify, which fetching a public page never asks for — mesh_identified_requests exists so that its zero is measured rather than assumed. On the web side there is a peer address, and lblogd never reads it, on disk or in memory: an address would buy a "unique visitors" figure that CGNAT, rotating IPv6 privacy addresses and crawlers make wrong anyway, at the price of making a blog server hold personal data. web_not_found is separate so that scans for pages that do not exist can be subtracted instead of inflating the total.
Dates are UTC calendar days, the same midnight a post's mtime fallback uses, and each record says tz=UTC so a bare date is still readable a year later.
The file is append-only and each record carries that day's whole running total, so the last record for a date wins and a kill -9 mid-write can lose at most the line being written — never an earlier day. A clock that steps backwards never reopens a day already written; its counts land on the open day and clock_behind records that it happened. The open day is written every five minutes, at every rollover, and once more on SIGTERM, and a restart resumes the day from the file rather than starting it again at zero. Each start compacts the file to one record per date.
awk '$1=="DAY" {print $2, $4}' /var/lib/lblogd/counts.log is the intended reader.
OPTIONS
--config file : Path to the TOML configuration file. Required.
--print-hash : Resolve the node's destination hash and the request paths it would serve — the pages first, including the static pages and the links, then the files — print them, and exit without starting any server. Needs no running daemon, so it doubles as a dry run for publishing: the posts and the file area are read exactly as serve mode reads them, with the same errors.
FILES
/etc/lblogd/config.toml : Configuration file installed by the Debian package. Registered as a conffile, so local edits survive upgrades.
/var/lib/lblogd/posts/ : Where the packaged service reads posts from: one Markdown file per post.
pages_dir
: Static pages: one Markdown file per page, named by its file stem. Not configured by default. index.md in it becomes the landing page and moves the post index to /blog. Reloaded with the posts.
/var/lib/lblogd/files/
: The file area: pictures and other files a post references. Set with files_dir, which defaults to a files directory beside posts_dir. The directory need not exist; without it the blog simply serves no files. Reloaded with the posts.
/var/lib/lblogd/counts.log
: The per-day request counts. See COUNTING. Moved with [counter] path, suppressed entirely with [counter] enabled = false.
/var/lib/lblogd/ : Node identity and, when certificate handling is enabled, the ACME certificate cache.
SIGNALS
SIGHUP
: Re-read the posts directory, the static pages and the file area. The packaged service maps systemctl reload lblogd onto this.
SIGTERM, SIGINT
: Stop. The open day's counts are written out first, so systemctl stop and systemctl restart do not lose the day so far.
EXIT STATUS
lblogd exits non-zero when the configuration cannot be loaded, when a post cannot be parsed at startup, and when no Reticulum daemon becomes reachable on the configured shared instance within the startup wait. Once running it is more forgiving: a reload that fails leaves the previous content serving.
EXAMPLES
Print the mesh address the blog would announce, without starting it:
lblogd --config /etc/lblogd/config.toml --print-hash
Publish a post to the packaged service:
sudo cp my-post.md /var/lib/lblogd/posts/
sudo systemctl reload lblogd
Publish a picture the post refers to as :
sudo install -o lblogd -g lblogd -m 640 mast.jpg /var/lib/lblogd/files/
sudo systemctl reload lblogd
SEE ALSO
lnsd(1), lnomad(1)
lnpnd(1)
NAME
lnpnd -- LXMF propagation node daemon for Reticulum
SYNOPSIS
lnpnd [--config dir] [--rnsconfig dir] [overrides] lnpnd --status [--peers] [-r hash] [--identity path] [--timeout secs] lnpnd --sync peer [-r hash] lnpnd --break peer [-r hash] lnpnd --exampleconfig
DESCRIPTION
lnpnd runs an LXMF propagation node: a store-and-forward mailbox on a Reticulum mesh. Clients (Sideband, MeshChat, lxmd-based tools) configure the destination hash it prints at startup as their propagation node; messages for offline recipients wait in its store until the recipient drains its mailbox, and the node peers with other propagation nodes -- lnpnd, lxmd or board-hosted -- and syncs stored messages both ways.
The daemon attaches to a Reticulum shared instance that is already running -- lnsd(1), or Python's rnsd -- the way lnstatus(1) and lncp(1) do; it does not start a Reticulum stack of its own. The instance need not be running yet: at startup lnpnd waits up to a minute for the daemon's IPC socket, which is what gets it through a boot where both start together and the daemon binds its socket a few seconds later. After that minute it exits, and the packaged service restarts it until a daemon answers.
lnpnd is a drop-in counterpart to Python's lxmd. It reads lxmd's configuration format and key names from an lxmd-shaped config directory, carries the same remote-management interface (an lxmf.propagation.control destination answering /pn/get/stats, /pn/peer/sync and /pn/peer/unpeer), and its query verbs print output in lxmd's shape -- lxmd --status --remote <hash> works against an lnpnd node, lnpnd --status --remote <hash> against a stock lxmd node, and scripts written for one read the other.
Like lxmd, the daemon also keeps a mailbox of its own: an LXMF delivery destination on the node identity, announced with the configured display name and stamp cost. Each received message is written to the messages directory in the reference's packed-container file format, and the configured on_inbound program runs with the file's path as its argument.
OPTIONS
--config dir
: Path to the lnpnd config directory (see FILES). Default: /etc/lnpnd if it holds a config, else ~/.config/lnpnd if it does, else ~/.lnpnd.
--rnsconfig dir
: Path to the Reticulum config directory whose instance_name decides which shared instance to join. Default: the platform's Reticulum config directory.
--instance name : Join this shared-instance name directly, overriding the Reticulum config file's.
--data-dir dir
: Where the message store, peer table and client state live. Default: <configdir>/storage.
-i, --on-inbound path
: Executable to run for each message received in the daemon's own mailbox; overrides the config file's on_inbound. The program receives the full path to the written message file as its argument.
-s, --service
: Log to <configdir>/logfile instead of the terminal. The packaged systemd unit does not use this -- under systemd the journal captures stderr.
-p, --propagation-node : Accepted for lxmd command-line compatibility; lnpnd always runs the propagation node role.
-v, -q
: Raise / lower the log level from the config's loglevel.
--exampleconfig : Print a verbose configuration example to stdout and exit.
Daemon settings can also be given as flags (--stamp-cost, --peering-cost, --max-peers, --static-peers, --autopeer, --autopeer-maxdepth, --remote-peering-cost-max, --max-inbound-syncs, --from-static-only, --transfer-limit-kb, --sync-limit-kb, --store-limit-kb, --announce-interval-secs, --name); a flag overrides the config file's value for the same setting.
REMOTE MANAGEMENT
The query verbs run as a client against a node's control destination, identifying with an identity the node has allowed. A node always allows its own identity, so the local verbs work on a fresh installation with nothing configured; control_allowed in the config adds other people's identity hashes. Without --remote they query the local daemon using the config directory's identity; with --remote hash (a propagation destination hash) they query that node, with --identity path naming the identity file to identify with. The counterpart tool's verbs are interchangeable with these.
A node that refuses a query answers the refusal, and the verb exits 204. Python's lxmd registers its control paths behind ALLOW_LIST, which sends nothing at all to an identity it refuses; against such a node a refusal is indistinguishable from silence and exits 200, so lnpnd names the possibility on stderr rather than leaving "timed out" as the only account.
--status : Print the node's status: store utilisation, costs, peer counts, traffic counters.
--peers : Print the peered nodes with their state, costs, sync keys and traffic.
--sync peer : Ask the node to sync with peer (a destination hash) now.
-b, --break peer : Break the node's peering with peer.
--timeout secs : Timeout for query operations (default 5, sync/break 10).
Exit codes follow lxmd's: 200 timeout, 203 no identity / bad hash, 204 access denied, 205 invalid data, 206 peer not found, 207 empty response.
CONFIGURATION
The config file is lxmd's format and keys (lxmd --exampleconfig and lnpnd --exampleconfig describe the same file). All keys of the reference's [propagation], [lxmf] and [logging] sections are accepted. Most are honoured identically: announce intervals and costs, autopeering and its depth, static peers, max_peers, from_static_only, max_inbound_syncs, auth_required (with the allowed file), control_allowed, storage and transfer limits, display_name, stamp_cost, delivery_transfer_max_accepted_size, on_inbound, loglevel, and the ignored file.
Keys accepted but not acted on, so a config file shared with lxmd parses cleanly (each is warned about at startup):
announce_at_start (in [propagation])
: lnpnd always announces the propagation node shortly after start, so yes is already the case and no has nothing to switch off. The [lxmf] key of the same name is honoured: it governs the daemon's own delivery destination.
prioritise_destinations
: lnpnd's store eviction is size- and age-driven only. The reference uses this list to keep favoured destinations when the store overflows; lnpnd's store design (shared with the board-hosted node, where the list would not fit) does not carry per-destination priority.
sequential_pn_stamp_validation, static_peers_bypass_sequential
: Stamp validation in lnpnd always runs sequentially on one worker thread, in arrival order, for every peer. The reference's toggles choose between parallel and sequential validation; lnpnd's single worker is the sequential behaviour, so yes is already the case and no has nothing to switch on.
Two defaults differ deliberately from the reference and are wire-legal: propagation_stamp_cost_target defaults to 0 (accept uploads without proof-of-work; lxmd never announces below 13) and peering_cost defaults to 0. Note that a stock lxmd peer can never sync toward a node announcing peering cost 0 -- its peering_key_ready short-circuits on a falsy cost -- so nodes that want stock peers to push messages to them should announce at least 1.
enable_node = no is refused: the propagation node is what lnpnd is. For a mailbox-only daemon use lxmd; for a client, lnmsg.
FILES
The config directory keeps lxmd's layout:
config : The configuration file. Created from the example on first daemon start if missing.
identity
: The node's identity -- its mesh address. Never replaced automatically; the same file format Python's RNS uses, so an identity can move between daemons. A start that finds no identity here creates one and says so before it joins the shared instance, logging a PN_IDENTITY_CREATED line with the new propagation destination hash. That line is the only notice a new address exists, so a node that is meant to continue an existing one must have that node's identity file copied here before its first start.
allowed
: With auth_required = yes: identity hashes (one hex hash per line) allowed to drain mailboxes from this node.
ignored : Destination hashes (one per line) whose messages the daemon's own mailbox drops.
storage/
: The data directory (unless --data-dir moves it): messagestore/ (the propagation store), peers/ (the peer table), messages/ (the daemon's own received mail, one packed-container file per message, readable by the reference's LXMessage.unpack_from_file), and the node's client state.
The Debian package installs the config directory at /etc/lnpnd and the data directory at /var/lib/lnpnd, both owned by the lnpnd service user. It enables the unit but does not start it: at install time neither an identity nor a configuration is in place, and a daemon started then would mint an address of its own and announce it. Place the identity, then systemctl start lnpnd. An upgrade over a running node restarts it and mints nothing; an upgrade over a stopped one leaves it stopped.
EVENTS
With LEVICULUM_EVENT_LOG=<path> set, the daemon appends one structured line per accepted upload (PN_ACCEPT), rejected upload (PN_REJECT, with a fixed reason word), mailbox request (PN_GET), eviction (PN_EVICT), peer-table change (PN_PEER), offer round (PN_OFFER), sync round (PN_SYNC) and own-mailbox delivery (PN_MAILBOX), plus a PN_STORE line every 8 minutes carrying store size against the limit. PN_STORE fires whether or not traffic arrived, so it is also the log's liveness heartbeat: a quiet node and a dead one are otherwise the same absence of lines.
scripts/analyze-lnpnd.py summarises such a log in one streaming pass with constant memory, which matters because a public node's event log grows to tens of gigabytes. packaging/logrotate/ caps it: a logrotate config, a oneshot and a 15-minute timer, with packaging/logrotate/README.md on why the rotation has to be copytruncate and what the ceiling comes to.
SEE ALSO
lnsd(1), lnstatus(1), lncp(1)
The Python counterpart: lxmd from the LXMF distribution.
lnflash(1)
NAME
lnflash -- flash, configure and watch LNode boards
SYNOPSIS
lnflash [options]
lnflash --set-time | --set-telemetry | --set-media [spec] | --set-name [name] | --set-position lat,lon[,alt] | --set-tx-power dbm | --set-tx-spacing ms | --set-ble-tx-gap ms | --store-storm count[,bytes] | --announce
lnflash --watch [serial-or-port] [--out file]
lnflash --summarize file
DESCRIPTION
lnflash brings an LNode board up on the firmware bundle beside the binary: it finds attached boards, brings each into its bootloader, confirms from the bootloader what the board is, checks the SoftDevice precondition and writes. After a flash it offers radio settings and a telemetry target. It needs no network and no external programs; writing needs root because the bootloader drive is a root:disk block device.
The configure sessions (--set- something) talk to boards that are already running and never flash: activation is configuration, not firmware. Each session ends the run, so only one of them can be given at a time.
--watch and --summarize are the field-testing pair: the first records a board's debug log with wall-clock timestamps, the second reads such a recording back. See FIELD WATCH below.
FLASHING OPTIONS
--bundle path : Where the bundle is. Otherwise $LNFLASH_BUNDLE, then next to the binary, then /usr/share/lnflash.
--board name : Only flash this board; refuse anything else.
--dry-run : Report what is attached and what would happen; change nothing.
--yes : Answer yes to every confirmation. Fails rather than waits when a board needs a physical double-tap.
--check-bundle : Verify the bundle's manifest and payload checksums, then exit.
--radio-preset name, --radio-freq hz, --radio-bw hz, --radio-sf n, --radio-cr n, --radio-txpower dbm, --no-radio : The radio settings written after a flash. A preset (eu868, us915, au915) and explicit values are two ways to state one configuration; pick one. --no-radio leaves the board's stored settings alone.
--telemetry address, --telemetry-profile name, --telemetry-key hex, --no-telemetry : The telemetry target offered after a flash, or the answer for --set-telemetry.
--quiet : Print less. With --watch, print nothing (an --out file is then required).
CONFIGURE SESSIONS
Each finds every running LNode on the bus, talks to it over the control envelope on the transport CDC (if02), reports what each board answered and exits. --set-time teaches the boards the host clock; --set-tx-spacing sets the on-air transmit spacing (not persisted); --set-tx-power sets the transmit power (persisted); --set-position/--clear-position pin or release a fixed position; --set-media reads or sets which carriers a board meshes over; --set-name/--clear-name read or set what the board is called; --set-telemetry configures the telemetry target; --set-ble-tx-gap sets the gap the board leaves between packets on one Bluetooth connection, 0 to 5000 ms (not persisted; 0 imposes nothing); --announce makes each board make every announce it makes on its own cadence, immediately and on all interfaces: its LXMF delivery destination exactly as its telemetry path does, and on a board running the propagation node role its lxmf.propagation destination too — a board without a calendar clock withholds the delivery announce and says so (run --set-time first), while the role's is not clock-gated; --store-storm count[,bytes] asks each board to append that many synthetic records to its message store (bytes defaults to 304, the measured field median; at most 1000 records of at most 1024 bytes). A flag given with no value, where allowed, only reads the boards back.
--store-storm is a bench instrument for the board's message store and the only thing that writes to it today: nothing on a board stores messages yet. It provokes the flash page erases a filling store causes, so their cost to Bluetooth throughput and LoRa airtime can be measured without waiting for a mesh to fill the region. The board acks the request and appends on its own task, so watch its debug port (--watch) for the STORE storm line that reports what landed, and STORE op_fail for every flash operation the SoftDevice refused. Nothing is persisted as configuration and the records are tagged as synthetic; a board that is still running a previous storm refuses a new one.
FIELD WATCH
--watch [serial-or-port]
: Open a running board's debug CDC (if00) with DTR and RTS raised — the firmware transmits only with both set — and keep reading. Every line is prefixed with an ISO-8601 wall-clock timestamp with milliseconds (the board's own t= stays in the line). If the port vanishes (reset, reflash, unplug) the gap is logged as its own [WATCH] line and the port is reopened with a bounded backoff; the watch never exits on EOF. With no value and exactly one running board, that board is watched; with several, name a board's USB serial or bus port (e.g. 3-2.4). A value containing a slash is opened directly as a serial port path. The watch runs until interrupted. It is not a daemon and has no background mode; run it in a terminal, or under nohup(1) yourself.
--out file : Append every watched line to file as well as stdout, flushed per line so a crash loses nothing. The file is the evidence a field walk leaves.
--summarize file
: Read a watch file and print, per hour, how many LoRa receptions of each class it holds — announce, data, path request — plus the last line seen per class and the number of reconnect gaps. Classification uses the flags= byte in a [LORA] RX line when present (the low two bits are the Reticulum packet type; a data packet to the well-known path request destination is a path request) and the class word otherwise; a bare RX n bytes line counts as unclassified. The watch itself never filters: the file is the evidence, the summary is a view of it.
EXAMPLES
Watch the only attached board, keeping the log:
lnflash --watch --out walk-$(date +%F).log
Watch one of several boards by serial, silently:
lnflash --watch 183004F712B4A7FE --out walk.log --quiet
Summarize the walk afterwards:
lnflash --summarize walk.log
EXIT STATUS
0 when every addressed board did what was asked (for --watch: never reached; the watch runs until killed).
On a flash run the three outcomes are kept apart, because they need different things done about them. Each board's line states which build it is running, which mechanism read it — the [FW_BUILD] banner the board emitted after the reset lnflash triggered — which port that line was read on, which device node the open was proved against, and how long after the port was flushed the line arrived. A line that arrived in the first milliseconds was already in flight and is worth doubting; a board that says nothing is reported as unknown and no sha is named for it.
0 : Every board was written and named the build in this bundle on its debug port.
1 : The flash failed: a board never came back, or came back naming a different build, or nothing was written to it.
2 : Every board took the write and none contradicted it, and at least one could not be read back. The firmware is on the board as far as anything here knows; which build it is running is unknown. Read it back again (--watch, or re-run the flash) rather than assuming the write failed.
Every other session exits 0 when every addressed board did what was asked and 1 otherwise.
SEE ALSO
lnsd(1), lnstatus(1), lnprobe(1)
The flashing design and its evidence: docs/src/concepts/lnode-flashing.md. Field testing: docs/src/guide/field-testing.md.
LNode Firmware: Supported Boards
The LNode firmware turns an nRF52840-based board into a standalone
Reticulum transport node. It runs the same leviculum-core transport
engine that powers the Linux daemon, cross-compiled for Cortex-M4F, and
routes packets between three interfaces: USB serial (HDLC framing to a
host), the SX1262 LoRa radio, and BLE. There is no PC in the data path;
the device is a router in its own right.
The transport engine is the same
leviculum-corelibrary that powers the Linux daemon, compiled for Cortex-M4F. (leviculum-nrf/README.md:4)
On the wire the firmware speaks the RNode LoRa framing protocol, so an
LNode and an RNode interoperate on the same LoRa network. On the host
side it connects to lnsd or rnsd over USB serial with HDLC framing.
On the BLE side it implements the Columba v2.2 protocol for the Columba
Android app. (leviculum-nrf/README.md:6,
leviculum-nrf/src/bin/t114.rs:3-8)
What the firmware does
Each firmware binary registers exactly three Reticulum interfaces and runs an event-driven main loop that dispatches packets between them:
| Interface | ID | Medium | HW MTU |
|---|---|---|---|
serial_usb | 0 | USB CDC-ACM, HDLC framing to host | 564 |
lora_sx1262 | 1 | SX1262 LoRa radio | 255 |
ble | 2 | BLE peripheral, Columba v2.2 | 564 |
(Interface registration and MTUs:
set_interface_name (leviculum-nrf/src/bin/t114.rs:265-282) and
leviculum-nrf/src/bin/rak4631.rs:325-366. The main loop selecting over
the three RX sources plus a timer deadline begins at
leviculum-nrf/src/bin/t114.rs:605.)
Transport routing is enabled in the node builder, so an LNode forwards
packets and serves paths for other peers, exactly like a
transport-enabled lnsd.
(enable_transport (leviculum-nrf/src/bin/t114.rs:203),
leviculum-nrf/src/bin/rak4631.rs:238)
Hardware coverage
We build one firmware per pinout family, not per product. A family is a set of boards whose SX1262 wiring is identical, which happens whenever the radio ships together with the MCU as one module: every carrier board built around that module then inherits the same wiring. One build therefore covers many products, and we only add a build when a board's radio wiring genuinely differs.
The same principle applies inside a family. Peripherals that a carrier board adds are detected at run time or degrade to nothing, so a single image serves the bare module and the fully populated product alike.
The policy behind this page, including when a specialised build is justified, is How far one firmware build reaches; how a board is identified before anything is written is Flashing an LNode.
How to read the tables
| Level | Means |
|---|---|
| Verified | We own this board and run it. Failures here are bugs we must fix. |
| Expected | Radio wiring checked against the vendor reference and identical to a verified board of the same family. Never run by us. Report results. |
| Not covered | Different radio wiring. Our image will not drive the radio; do not flash it. |
Expected is not a support promise. It means the one thing that decides whether the radio comes up at all, the SX1262 wiring, matches. Everything a carrier adds beyond that, displays, GNSS, Ethernet, accelerometers, e-paper, is not driven by our firmware on these boards even where the vendor firmware drives it. The node routes packets; the extra hardware stays dark.
Family A: RAK4630 module
rak4631 binary. Radio pins are internal to the RAK4630 module and
therefore identical across every carrier: NSS P1.10, SCK P1.11,
MOSI P1.12, MISO P1.13, BUSY P1.14, DIO1 P1.15, NRESET P1.06,
power enable P1.05. DIO2 drives the antenna switch; there is no
external TX or RX enable line on any of them.
| Product | Level | Note |
|---|---|---|
| RAK WisMesh Pocket V2 (RAK19026 carrier) | Verified | Display, GNSS and battery supported |
| RAK4631 bare module | Verified | Same image, peripherals absent |
| RAK WisMesh Pocket Mini (RAK19003) | Expected | |
| RAK WisMesh Repeater / Hub (RAK2560) | Expected | Solar, IP67 |
| RAK WisMesh Tap | Expected | TFT not driven |
| RAK WisMesh Tag | Expected | |
| WisBlock with RAK13800 Ethernet | Expected | Ethernet not driven |
| WisBlock with RAK14000 e-paper | Expected | E-paper not driven |
| NomadStar Meteor Pro | Expected | |
| MonteOps HW1 | Expected | |
| GAT562 Mesh Trial Tracker | Expected | |
| MeshTiny | Expected | |
| muzi R1 Neo | Expected |
Verified against Meshtastic's own variant definitions under
variants/nrf52840/, where all twelve carriers repeat the same seven pin
numbers, and against the RAK4630 datasheet quoted in
variants/nrf52840/rak4631/variant.h.
The non-radio pins this build drives were checked the same way, and they
hold up: the two LEDs on P1.03 and P1.04 are the module's own, and
P0.13 / P0.14 carry I2C and P0.15 / P0.16 the first serial port on
every carrier that defines them at all. That is not luck. The RAK4630
brings these signals out on fixed module pins and the WisBlock carriers
follow that convention, so a module-defined family stays coherent beyond
the radio. No carrier was found driving an output into a pin this build
also drives.
The one carrier worth naming is the e-paper-on-RX/TX variant, which puts the display's SPI where the others put I2C and the serial port. Nothing there fights our outputs, but the pins carry traffic that means nothing to that hardware.
Do not confuse RAK4631 with RAK3401. Meshtastic's
rak3401_1wattvariant declares the same PlatformIO board name, but it is a different radio module with different pins, its own SPI bus and a 1 W power amplifier. Our image would drive the wrong pins on it. This is why identification uses the bootloader'sBoard-ID, never a board name that vendor trees reuse.
Family B: Heltec T114
t114 binary. NSS P0.24, SCK P0.19, MOSI P0.22, MISO P0.23,
BUSY P0.17, DIO1 P0.20, NRESET P0.25, TCXO at 1.8 V via DIO3, DIO2
as antenna switch.
| Product | Level | Note |
|---|---|---|
| Heltec Mesh Node T114 | Verified | Status display and GNSS (L76K) supported |
| Heltec MeshSolar | Blocked | Radio matches, but our status LED sits on the battery controller's emergency-shutdown pin |
| LILYGO T-Echo | Do not flash | Two pin conflicts, see below |
| LILYGO T-Echo Plus | Do not flash | Same as T-Echo |
Matching radio pins are not sufficient, and this family is where that becomes concrete. The
bsp-t114build drives an ST7789 panel blind, because the panel cannot be detected, plus an LED and a GPS UART. Those pins are as much part of the image as the radio pins, and on a related board they land on whatever that board put there.On LILYGO T-Echo the collisions are severe. Our TFT power-enable output
P0.03meetsPIN_EINK_BUSY, which is an output of the e-paper controller, and our TFT clockP1.08meetsGPS_TX_PIN, an output of the GPS receiver. Both are two drivers on one line. Our TFT data lineP0.12meetsPIN_POWER_EN, so the display driver would switch the board's peripheral power on and off as a side effect of drawing. This is a hardware hazard, not a board that merely fails to transmit.On Heltec MeshSolar the radio wiring, the LoRa SPI bus and even the GPS UART line up exactly, and none of our TFT pins is occupied. One pin spoils it: our status LED
P1.03is that board'sBQ4050_EMERGENCY_SHUTDOWN_PIN. Blinking a heartbeat onto the battery controller's shutdown input is not acceptable, so this stays blocked until the LED becomes a board fact that can be left unset.
Method note. Membership in a pinout family is a necessary condition, never a sufficient one. Before any board moves to Expected, every pin the image drives has to be checked against that board's own definition, not only the seven radio pins. The three entries above passed the radio check and failed this one.
Unlike family A, this family is not one module: these are separate boards that happen to share a wiring convention, so a new Heltec or LILYGO model is not covered by default. Heltec Mesh Pocket, Heltec T1, Heltec T096 and LILYGO T-Echo Lite each wire the radio differently and are not covered either.
So today this build serves exactly one product, the T114. Sharing a radio pinout turned out to be the easy half.
The
Board-IDdoes not separate this family from its neighbours, and that is a hazard rather than an inconvenience. Meshtastic records the same bootloader product stringHT-n5262for the T114, for MeshSolar and for the Heltec Mesh Pocket, whose radio is wired differently and which is not covered here. Both our tools match that string exactly (board_for_id(lnflash/src/manifest.rs:495),leviculum-nrf/tools/uf2-runner.sh:108), so if theINFO_UF2.TXTBoard-IDis identical too, neither can tell a Mesh Pocket from a T114. We cannot check that without the hardware. Until someone does, treat aHT-n5262match as a family hint and confirm the model by other means before writing.This is the general rule behind both this warning and the XIAO case below: a
Board-IDis only a safe key when it is bound to the same unit as the radio wiring. On the RAK4630 both belong to the module, so the key is exact. Heltec binds the identifier to a bootloader shared across models while the wiring belongs to the model, and Seeed binds it to the MCU module while the radio sits outside it. Both of those decouple, and a decoupled key cannot carry a write decision alone.
LILYGO T-Echo and T-Echo Plus are the harmless side of the same coin:
they report a different Board-ID (TTGO_eink by Meshtastic's record),
so our tools decline them today. The firmware would run; the tooling
needs the identifier before it can.
Elecrow ThinkNode M1 is not covered, although its seven radio pins match. It runs its TCXO at 3.3 V where this family uses 1.8 V, and our build compiles 1.8 V in. Supporting it needs that value to become a board fact rather than a family fact.
Family C: XIAO nRF52840 + Wio-SX1262
solarnode binary. NSS P0.04, SCK P1.13, MOSI P1.15, MISO P1.14,
BUSY P0.29, DIO1 P0.03, NRESET P0.28, TCXO at 1.8 V via DIO3, DIO2
as antenna switch and an external RX enable on P0.05. That last pin
is what this family adds to the shared code: DIO2 steers only the
transmit side of the Wio-SX1262's switch, so the receive side is a host
GPIO the driver asserts for a listening window and releases before every
key-up, through a state machine that refuses both paths at once
(FrontEnd, leviculum-nrf/rx-arming/src/front_end.rs:97;
LoRaRxEnable, leviculum-nrf/src/boards/solarnode.rs:70).
| Product | Level | Note |
|---|---|---|
| Seeed SenseCAP Solar Node P1-Pro | Bring-up | On the rig since 2026-09-15; radio and GNSS not yet confirmed on the bench |
| Seeed XIAO nRF52840 + Wio-SX1262 kit | Expected | Same two modules, same seven pins plus RXEN |
| Wio Tracker L1 / L1 e-ink | Not covered | Different carrier, LEDs and battery sense not checked |
The image drives, besides the radio, one LED on P0.19, the battery
divider on P0.31/P0.14 and the XIAO L76K GNSS on P1.11/P1.12.
There is no display. That the kit above can be Expected at all is
largely because the modules decide the radio and the carrier decides the
rest: a carrier that omits the L76K reports no-hardware at run time
rather than needing an image of its own.
The battery sampler is in (Codeberg #233), and it reads this board's
divider: the XIAO module's own 1 MΩ over 510 kΩ, which puts the whole
measurable range at 10 660 mV against the T114's 17 698. Two things about
it are this board's alone. Its enable pin P0.14 is active low — it
sinks the low side of the divider rather than switching a load — so the
polarity travels with the pin (battery::DividerEnable) instead of being
a shared constant. And its 338 kΩ of source resistance is past the
100 kΩ the nRF52840 specifies the 10 µs acquisition window for, so the
board states its divider as the two resistors rather than as their ratio
and BatteryScale::for_divider derives a 20 µs window from them; the
[BAT] init line carries the result as acq_us=. The other two boards'
dividers are inside the default and their sampling is unchanged.
What the divider does not settle is the pack's cell topology, and
nothing in this firmware guesses it: the cell count is classified from the
board's own first reading, never read out of a constant. What the divider
does settle is a bound: at 10 660 mV of full scale this board cannot see a
series pack above two cells, since 3S sits above the range and would put
more than the ADC's 3.6 V on the pin. A reading above 9 V is rejected as
implausible and the board says [WARN] [BAT] implausible first reading
rather than publishing a percentage.
The topology itself is settled, and it is 1S4P. It is a hardware fact
rather than a firmware one, so it lives in the board file's
BATTERY_DIVIDER doc with its sources, and it is worth stating here
because a reader who only knows "four 18650s" will guess 2S2P — the two
give the same ~49 Wh, so the energy figure cannot decide it. The charger
can. The divider hangs on the XIAO's VBAT net, whose charger on the
module's own schematic is U2 BQ25100, a linear charger for one cell at a
fixed 4.2 V; the carrier charges the same net through a second single-cell
part, the CN3165 Seeed names as this product's charging management chip,
also fixed at 4.2 V, from a 5 V Type-C or 5 V solar input with nothing on
the board that steps up. So the four cells are in parallel: 13.4 Ah at one
cell's voltage. The rig unit reads pack_mv=4123..4154 with cells=1S
across eleven captures, which is where a 1S pack held by a 4.2 V charger
sits. A plausible reading on this board is therefore
pack_band_mv(1) = 2500..4330 mV, and a test that expects 6000..8660 from
it is asserting a pack this product does not have.
The bootloader cannot tell these apart. nRF52840-SeeedXiao-v1 names the
MCU module, and a DIY XIAO with an entirely different radio wired to the
same pads reports exactly the same string, so lnflash has no flashing
entry for this board and must not be given one on that evidence
(Codeberg #233). It does have a control-only catalogue entry, keyed on
the USB ID our own firmware publishes (1209:0003), so --watch,
--announce, --set-time and the --radio-* flags reach this board
like any other; a bundle carrying an image named after it is refused when
it loads. Writing firmware goes through just flash-solarnode, which is
told the Board-ID explicitly by a person who can see which board is on
the bench, and for the first flash the image goes onto the mass-storage
volume by hand.
ESP32 class: Heltec WiFi LoRa 32 V4
A different crate and a different stage. Everything above is
leviculum-nrf on nRF52840 and describes firmware that routes packets.
The ESP32 class is leviculum-esp, and as of step 1 it is a skeleton:
one binary, heltec_v4, built for xtensa-esp32s3-none-elf. The row
below is in this table so the board is not invisible, not because it is
comparable to the families above.
| Product | Level | Note |
|---|---|---|
| Heltec WiFi LoRa 32 V4 | Skeleton | Boots and identifies itself; no radio traffic, no interfaces, no transport |
What it can do after this step. Bring the SoC up, open the USB
Serial/JTAG port, and emit [FW_BUILD] git_sha=<short> dirty=<true|false> t=<ms> at boot and every five seconds after it, plus a [BOARD] line
naming the board file the image was compiled against. Blink the status
LED. Take the seven radio pins and hold the SX1262's SPI port open, with
leviculum_core::sx126x's CommandBus and RegisterBus implemented over
esp-hal SPI.
What it cannot do. Anything on the air. No opcode is issued to the radio, the front-end amplifier is never enabled, there is no interface, no transport, no identity, no persistence and no BLE. It does not interoperate with anything.
Radio pins, read off the manufacturer's schematic (revisions 4.2 and 4.3,
which agree on all seven): NSS GPIO8, SCK GPIO9, MOSI GPIO10,
MISO GPIO11, NRESET GPIO12, BUSY GPIO13, DIO1 GPIO14. The TCXO is
supplied from the SX1262's DIO3 through a ferrite bead; the voltage it
needs is not established — the schematic names the part only as
"32MHz" — and the constant is deliberately absent rather than guessed
(leviculum-esp/src/boards/heltec_v4.rs).
This board has a front end, and the two published schematics disagree about how it is steered. The V4 is the high-power variant: the SX1262 reaches the antenna through a KCT8103L PA/LNA whose CTX and CPS control inputs are wired to different sources in revision 4.2 and revision 4.3 — in 4.2 the SX1262's DIO2 drives CTX and
GPIO46drives CPS, in 4.3 DIO2 drives CPS andGPIO5drives CTX. Only CSD (GPIO2) agrees. Which is right decides the transmit path, so the board revision has to be read off the physical board before anything keys up. Step 1 does not need the answer and does not pretend to have it.
There is no lnflash entry and no UF2: the ESP32-S3 has no mass-storage
bootloader. The image is written with espflash over the same USB port
the banner comes out of, which is the SoC's own USB peripheral — there is
no USB-to-UART bridge on this board.
Not covered today
Each of these is a separate pinout family around the SX1262 this firmware already drives, reachable by adding one board file rather than by changing shared code:
| Family | Products |
|---|---|
| ThinkNode M6 | Elecrow ThinkNode M6, muzi BASE |
| ProMicro + E22 | nRF52 ProMicro DIY, DLS Minimesh Lite |
| Individual wirings | Heltec Mesh Pocket, B&Q Nano G2 Ultra, LILYGO T-Echo Lite, Canary One, MS24SF1, MeshLink, TWC Mesh v4 |
A different radio family is not on that list, and the T1000-E is the board that makes the distinction worth drawing. The Seeed SenseCAP Card Tracker T1000-E carries the nRF52840 every nRF family above runs on, so its MCU, its bootloader and its USB path are all familiar; its LoRa transceiver is a Semtech LR1110, a different part with a different command set. No board file reaches that. The SX126x driver every build on this page shares does not carry over at all, which makes this device dearer to support than either of the other two boards waiting for attention: the Solar Node P1-Pro above is the same radio die on a resolved pin map, and the Heltec V4 brings a new MCU family and a new toolchain but reuses the radio driver unchanged. What the port would actually cost, and which of the radio-adjacent crates survive it untouched, is in How far one firmware build reaches under "The axis the policy does not have"; Codeberg #406 is the record.
The one thing it has that no board above has is a 3-axis accelerometer,
which is the movement signal the announce cadence currently has to infer
from a position delta (MovementDetector,
leviculum-nrf/announce-policy/src/cadence.rs:226). If that work ever
needs a hardware answer instead, this is the device that can give one.
The XIAO family, now family C above, is the one case where the
bootloader cannot answer which board it is: the MCU module is a XIAO and
the radio is a separate part, so a SenseCAP Solar Node and a DIY XIAO
with different radio wiring both report nRF52840-SeeedXiao-v1. Having
a build for it does not change that. Boards like that need a second
discriminator before anything may be written.
Known open question
Our RAK build sets the SX1262 TCXO to 3.3 V, following the RNode
firmware, which selects MODE_TCXO_3_3V_6X for this board
(leviculum-nrf/src/boards/rak4631.rs:39-43). Meshtastic and MeshCore
both run the same module at 1.8 V. The value lives in the module, so it
applies to every carrier in family A equally. Our Pocket V2 works with
3.3 V, but the divergence against two references is unresolved and should
be settled before the family is presented as broadly supported.
Cargo features and binaries
Three firmware binaries are defined, one per board family:
[[bin]]
name = "t114"
path = "src/bin/t114.rs"
[[bin]]
name = "rak4631"
path = "src/bin/rak4631.rs"
[[bin]]
name = "solarnode"
path = "src/bin/solarnode.rs"
(leviculum-nrf/Cargo.toml:412-422)
The board-support-package (BSP) features select the runtime for a given
board. Exactly one BSP feature must be enabled per build; a
compile_error! in lib.rs enforces the mutual exclusion.
(leviculum-nrf/src/lib.rs:34-43)
| Feature | Effect | Cite |
|---|---|---|
bsp-t114 | T114 BSP (+ SoftDevice BLE + status display + GNSS + battery) | leviculum-nrf/Cargo.toml:351 |
bsp-rak4631 | RAK4631 BSP (+ SoftDevice BLE) | leviculum-nrf/Cargo.toml:335 |
bsp-solarnode | SenseCAP Solar Node P1-Pro BSP (+ SoftDevice BLE + battery + GNSS). No display | leviculum-nrf/Cargo.toml:372 |
display | SSD1306 OLED, probed at run time | leviculum-nrf/Cargo.toml:374 |
gnss | NMEA0183 GNSS (ZOE-M8Q on the V2 baseboard, L76K on the T114 and the Solar Node) | leviculum-nrf/Cargo.toml:375 |
battery | pack-voltage monitor: the BATTERY log line, the panel's voltage and, on the V2, the telemetry field. Unconditional under bsp-t114 (the divider is on every T114) and under bsp-solarnode (it is on the XIAO module), opt-in on the V2 via rak-baseboard | leviculum-nrf/Cargo.toml:381 |
rak-baseboard | aggregate of display + gnss + battery | leviculum-nrf/Cargo.toml:382 |
Note on BLE: Both firmware entry points register a BLE interface and call
leviculum_nrf::ble::init(leviculum-nrf/src/bin/t114.rs:372,leviculum-nrf/src/bin/rak4631.rs:429). The Cargosoftdevicefeature, and therefore the BLE stack, is pulled in by both BSP features (leviculum-nrf/Cargo.toml:335,leviculum-nrf/Cargo.toml:188).
The baseboard peripherals are each gated behind their own Cargo feature
(leviculum-nrf/Cargo.toml:374-382) and spawned only when that feature
is on (leviculum-nrf/src/bin/rak4631.rs:365-392). Because each of them
either probes for its hardware or degrades to nothing when it is absent,
the aggregate build is what we ship for the whole family rather than a
Pocket-V2-only image.
The mapping from board to binary and features used by the flash recipes:
| Board | Binary | Features |
|---|---|---|
| Heltec Mesh Node T114 | t114 | bsp-t114 |
| RAK4631 (bare module) | rak4631 | bsp-rak4631 |
| WisMesh Pocket V2 (full baseboard) | rak4631 | bsp-rak4631,rak-baseboard |
(Feature sets as invoked in the just flash, just flash-rak4631, and
just flash-rak4631-pocket recipes: Justfile:1887, Justfile:1915,
Justfile:1929.)
What the lnflash bundle carries
The distributable bundle carries an image for both families: bsp-t114
for the T114 and bsp-rak4631,rak-baseboard for the RAK4630 module
(Codeberg #261). The RAK row it ships is the last one in the table above,
not the middle one — the bare module runs the baseboard image, and the
paragraph above is why. There is deliberately no way for a user to choose
between them: the manifest cannot express two images for one Board-ID,
because a question nobody can answer from looking at their board is not a
question worth asking.
Which boards the bundle knows at all is lnflash/catalogue.toml, and it
is a shorter list than the tables above on purpose. A row here says our
image would drive that board's radio; a catalogue entry with a flashing
section says the bootloader can be told apart from every other board's,
which is the stricter of the two claims and the only one a write may rest
on. A catalogue entry without one — the Solar Node's, Codeberg #233 —
makes the control commands reach the board and nothing else; a bundle
naming such a board fails to load. See
Building and flashing, "Which boards the bundle carries".
The two are held together mechanically rather than by care, because the
same board facts now sit in three files and Codeberg #262 records what
that costs here: eleven Justfile citations in flashing.md had drifted
by roughly 250 lines before anyone noticed. Every board the catalogue
knows has to be named on this page, and every identifier a session rests
on — the Board-ID a write matches, the USB IDs a bootloader and a
running application answer on, the drive label a user is told to look
for — has to appear somewhere in this book
(every_board_the_catalogue_knows_is_named_on_the_coverage_page,
lnflash/tests/doc_board_catalogue.rs:171;
every_identifier_a_session_rests_on_is_written_down_in_the_book,
lnflash/tests/doc_board_catalogue.rs:196). The check runs in the
direction a board change travels: the catalogue leads and the prose
follows, so adding a board to lnflash without writing it down here is
red. It does not claim the sentence around an identifier is right — the
book quotes identifiers on purpose that are not ours and must never be
catalogue keys, Meshtastic's 2886:0059 and LILYGO's TTGO_eink among
them (Codeberg #262).
Build target
All firmware builds target the hard-float Cortex-M4 triple:
thumbv7em-none-eabihf
(leviculum-nrf/README.md:15. Add it with rustup target add thumbv7em-none-eabihf.)
Default radio profile
The radio parameters are compiled into the firmware and must match the RNode configuration on the same LoRa network.
| Parameter | Value |
|---|---|
| Frequency | 869.463 MHz (ReticulumNet consensus, EU ISM band) |
| Spreading factor | SF8 |
| Bandwidth | 125 kHz |
| Coding rate | CR4/5 |
| TX power | 22 dBm |
(leviculum-nrf/README.md:8. The profile the firmware loads at boot,
eu_medium (leviculum-nrf/src/lora.rs:432-461), applied at
leviculum-nrf/src/bin/t114.rs:305 and
leviculum-nrf/src/bin/rak4631.rs:421.)
See Flashing for how to build and write these binaries to a board, and Recovery for the bootloader-entry details.
LNode Firmware: Building and Flashing
There are two ways to put our firmware on a board, and they exist for different people.
lnflash | just flash* | |
|---|---|---|
| for | anyone with a board | developers and CI |
| needs | the bundle, and root | this checkout and the embedded toolchain |
| builds firmware | no, it carries it | yes, from the working tree |
| identifies the board | from its bootloader | from the USB id you configure |
| boards today | T114, RAK4631 | T114, RAK4631 |
If you just want our firmware on a board, use lnflash. If you are
changing the firmware and want your build on a board, use just flash.
Physical-device steps. The author of this page cannot flash a board, so any step that writes to or resets real hardware is marked derived from source — requires the physical device. The commands themselves are quoted verbatim from the
Justfileandleviculum-nrf/README.md; only the outcome on hardware is un-verified here.
lnflash, the distributable flasher
lnflash is a single static binary with the firmware beside it. It
needs no toolchain, no Python, no network, and nothing installed: the
point of the bundle is that a stranger can unpack it and run it.
wget https://codeberg.org/Lew_Palm/leviculum/releases/download/nightly/lnflash-nightly-amd64.tar.gz
tar xzf lnflash-nightly-amd64.tar.gz
cd lnflash-*
sudo ./lnflash
(Justfile:51-52)
That URL is the whole answer to "how do I get your firmware onto my
board" and it is the one this page previously left out: it described the
bundle without saying where it comes from, so the only path a reader
could follow was a build from source (Codeberg #295). The rolling nightly
carries one image per board in the list at scripts/lnflash-bundle.sh,
and just check-firmware-images keeps that list, the README's board
table and the release body from disagreeing about it.
It works out what the board is, rather than being told. That matters
because a board arrives carrying whatever its last owner put on it:
stock firmware, Meshtastic, MeshCore, RNode firmware, ours, or a build
that crashes before it reaches USB. Each of those picks its own USB
identity, so the running firmware cannot be trusted to say what the
hardware is. lnflash therefore finds candidates on the USB bus, brings
each into its bootloader, and only there asks what the board actually
is, from the bootloader's own INFO_UF2.TXT. The identity that a write
rests on can only come from that reading, which is enforced in the type
system rather than by convention (lnflash/src/lib.rs:15-21). Then it
checks the SoftDevice precondition, installs a matching SoftDevice first
if needed, writes the firmware, and reads the board's debug port back to
confirm what is now running.
Nothing is written before all of that has been shown and confirmed.
Root is required. The bootloader's drive is a root:disk block
device, and lnflash mounts it itself rather than assuming a desktop
automounter that a headless host does not have. Without root it will
identify the attached boards and then stop.
(lnflash/src/main.rs:37-38)
One key press is sometimes unavoidable. Getting into the bootloader
by software has to be implemented by whatever firmware is currently
running. Ours implements it, so every re-flash is touch-free. Stock
Meshtastic does not, so a first flash away from it needs a physical
double-tap of RESET, the second press within about half a second of the
first. lnflash detects that case and asks for it in plain words.
There is no universal software trigger, and a tool that claimed
otherwise would be lying.
Options
--dry-run reports what is attached and what would happen, changing
nothing at all, not even rebooting a board into its bootloader.
--check-bundle verifies the bundle's own checksums and exits.
--board NAME refuses to write if what is attached is a different
board. --yes skips confirmation for automation and fails rather than
waits when a board needs the manual double-tap. Radio settings can be
given at flash time with --radio-preset (eu868, us915, au915) or
the individual --radio-freq, --radio-bw, --radio-sf, --radio-cr
and --radio-txpower flags; --no-radio leaves the board's stored
configuration alone. (lnflash/src/main.rs:42-402. The board keeps what
it is given across resets and across the next flash, so this is part of
the flash rather than a later configuration step.)
The bundle is looked for in this order: --bundle PATH, then
$LNFLASH_BUNDLE, then the directory holding the binary, then
/usr/share/lnflash. (lnflash/src/main.rs:42-45)
The full user-facing text ships inside the bundle as its README
(lnflash/payload/README-bundle.md), including what the alarming but
harmless "the drive went away mid-flush" message means.
Building a bundle
just lnflash-bundle
Cross-compiles the firmware, converts it to UF2, builds the musl-static
binary, stages Nordic's SoftDevice next to Nordic's own licence file,
generates a manifest with checksums, and verifies the result. Output
lands under target/lnflash/. The first run takes minutes because of
the firmware build; SKIP_FIRMWARE=1 reuses an existing ELF while
iterating on the bundle itself. (Justfile:53-62)
Everything in the bundle comes from this checkout. A bundle built out of
a foreign tree would be exactly the hidden dependency our
clone-and-deploy policy forbids. (Justfile:55-57)
Which boards the bundle carries
Today: the T114 and the RAK4631 (WisMesh Pocket V2 and every other carrier built around the RAK4630 module). Boards are data rather than code, so a new board is a catalogue entry plus a firmware build, not a new binary — and an entry without a firmware build is an empty promise, so the shipped bundle carries what we actually build.
The RAK4631 image is the bsp-rak4631,rak-baseboard build, the same one
just flash-rak4631-pocket produces. Not because it is the richer build,
but because How far one firmware build
reaches already decided it: one
build serves a pinout family, and everything the Pocket V2 baseboard adds
degrades harmlessly on a bare module — the display is found by an I2C
probe and its task exits when nothing answers, the button is Pull::Up
so an absent one reads as not pressed, the GNSS task parks on a silent
UART, and the battery task publishes to a subscriber that is not running.
The bundle therefore does not ask which RAK you have, and the manifest
has no way to express two images for one Board-ID.
scripts/lnflash-bundle.sh walks a board list rather than naming boards
in its steps, so a third board is one more line in that list: the
firmware build, the UF2 conversion, the staging, the manifest sections
and the licence assertions against the finished tarball all derive from
it.
The SenseCAP Solar Node is known but not flashed here (Codeberg
#233). lnflash talks to it like any other board — --watch,
--announce, --set-time, --set-name, the --radio-* flags — because
those reach a board that is up and identifying itself. Writing firmware
to it is a different question and the answer is no: the Board-ID its
bootloader publishes, nRF52840-SeeedXiao-v1, belongs to the XIAO module
rather than to this product, and a DIY XIAO with the radio wired
elsewhere reports the same string. So the bundle carries no image for it,
--board solarnode is refused, and a flash session that finds it on the
bus names it, says why, and leaves it alone. It is flashed from this
checkout with just flash-solarnode, by a person who can see which board
is on the bench.
The SoftDevice carve-out. The T114 entry ships Nordic's S140 7.3.0
beside its licence, so a factory board carrying 6.1.1 is repaired and
then flashed. The RAK4631 entry ships no SoftDevice. It states the same
>=7.0.1, <8.0.0 constraint, but whether a factory Pocket V2 carries
something that constraint refuses is unmeasured — our only RAK has
run 7.3.0 since we first flashed it. A board that violates the constraint
with no remedy in the bundle is refused with Nothing was written rather
than written blind. The full reasoning, and what one reading of a stock
board would take to close it, is under "The SoftDevice carve-out" in
Flashing an LNode.
First flash on a Pocket V2 needs the pinhole. That board has no
externally accessible RESET, so when the 1200-baud touch does not take —
which is every board still running stock Meshtastic — lnflash asks for
a needle double-tap in the hidden pinhole beside the USB socket by name,
and points at Recovery. The bundle does not depend on the
meshtastic CLI for this; just dfu-rak4631 below stays available in
this checkout, but a stranger with the tarball needs only a needle.
The design behind all of this, including why the bootloader rather than the application is the board's identity, is in Flashing an LNode.
The developer path: building from this checkout
The rest of this page covers building the firmware here and flashing it
with the just flash* recipes.
Prerequisites
Install the Rust embedded toolchain, the ARM cross-compiler (needed by
nrf-sdc for C-header bindgen), flip-link, and add your user to the
dialout group for serial-port access. Log out and back in after the
usermod so the new group membership takes effect.
rustup target add thumbv7em-none-eabihf
rustup component add llvm-tools
cargo install flip-link
sudo apt install gcc-arm-none-eabi
sudo usermod -aG dialout $USER
flip-link is the firmware linker. It relocates the stack to the bottom of RAM so a stack overflow faults cleanly against the RAM floor instead of silently corrupting memory. It is link-time only, with zero runtime cost.
(leviculum-nrf/README.md:12-19)
--release is mandatory
Always build and flash with --release. The debug profile does not fit
the nRF52840 flash — the image overflows FLASH by several hundred KB at
link time.
The debug profile does not fit the nRF52840 flash (the image overflows FLASH by several hundred KB at link time) — always build and flash with
--release; alljust flash-*recipes already do. (leviculum-nrf/README.md:65-67)
Every just flash* recipe already passes --release, so following the
recipes below keeps you safe. The release profile is size-optimized
(opt-level = "z", lto = true, codegen-units = 1); DWARF debug info
is kept in the .elf (strip = "none", debug = true) for HardFault
post-mortem analysis, but the UF2 only carries loadable sections, so the
debug info does not bloat what lands on the device.
(leviculum-nrf/Cargo.toml:399-409)
The build/flash workflow
The firmware crate leviculum-nrf is its own Cargo workspace, separate
from the repo-root workspace, and is cross-compiled. The flash recipes
therefore cd leviculum-nrf before invoking cargo. (Justfile:1884-1885)
A plain build (no flash) is:
cargo build --release
(leviculum-nrf/README.md:23)
Flashing wraps cargo run: the runner builds the release binary, then
copies the resulting UF2 onto each board's UF2 bootloader drive. The
UF2 conversion and copy happen inside the cargo run step — a bare
cargo build produces only the ELF.
Build the firmware with
cargo build --release. Flash withjust flash(from the repo root), which wrapscargo run --release --bin t114. (leviculum-nrf/README.md:23)
Touch-free vs. manual double-tap
For the T114, flashing is touch-free in the common case: the host
opens the board's transport CDC port at 1200 baud, the firmware
intercepts the line-coding change, writes a retained-register magic, and
soft-resets into the Adafruit UF2 bootloader. No button press.
(leviculum-nrf/README.md:27)
A physical double-tap of RESET is still needed when the firmware on a
specific T114 has crashed or never reached USB init (panic before the
handler is installed, stack overflow, hardware fault). The runner detects
this per device via a UF2-drive-polling timeout and prompts for that
specific board only; the rest of the batch keeps flashing touch-free.
(leviculum-nrf/README.md:38)
The WisMesh Pocket V2 (RAK4631) running stock Meshtastic has no
1200-baud-touch handler and no externally accessible RESET pin, so its
first flash needs either just dfu-rak4631 (a Meshtastic admin
command, below) or the manual needle double-tap in the hidden pinhole.
Once our firmware is on the board, subsequent flashes use the touch path
automatically. (Justfile:1911-1913, Justfile:1963-1972. See
Recovery for the pinhole detail.)
The flash recipes
Each recipe below is quoted from the Justfile. The cargo invocation is
derived from source — requires the physical device to actually write
firmware (it builds the same on any host, but only does something useful
with a board attached).
just flash — every T114
Flashes every attached T114 sequentially. Flashing all of them is deliberate: if only one were flashed, a later multi-node test could run against mixed firmware versions. Use this as your default for T114s.
cd leviculum-nrf && cargo run --release --bin t114 --features bsp-t114
(Justfile:1886-1888; rationale leviculum-nrf/README.md:25)
just flash-one PORT — a single T114
Flashes one T114 by port path or udev symlink. Use it for A/B firmware testing (one board on a new build, one on the old).
just flash-one /dev/leviculum-transport
just flash-one /dev/ttyACM3
Expands to:
cd leviculum-nrf && LEVICULUM_FLASH_ONLY=<PORT> cargo run --release --bin t114 --features bsp-t114
(Justfile:1895-1900; usage forms leviculum-nrf/README.md:31-36)
just flash-rak4631 — every RAK4631 (bare module)
Flashes every attached RAK4631 / WisMesh Pocket V2 with the bare-module build (no baseboard peripherals).
cd leviculum-nrf && LEVICULUM_USB_PID=0002 LEVICULUM_BOARD_NAME=RAK4631 \
LEVICULUM_UF2_BOARD_ID=WisBlock-RAK4631-Board \
cargo run --release --bin rak4631 --features bsp-rak4631
(Justfile:1914-1916)
just flash-rak4631-one PORT — a single RAK4631
Flashes one RAK4631 by port path or udev symlink.
just flash-rak4631-one /dev/ttyACM0
just flash-rak4631-one /dev/leviculum-rak-transport
Expands to:
cd leviculum-nrf && LEVICULUM_FLASH_ONLY=<PORT> LEVICULUM_USB_PID=0002 \
LEVICULUM_BOARD_NAME=RAK4631 LEVICULUM_UF2_BOARD_ID=WisBlock-RAK4631-Board \
cargo run --release --bin rak4631 --features bsp-rak4631
(Justfile:1918-1922)
just flash-rak4631-pocket — WisMesh Pocket V2, full baseboard
Flashes with all RAK19026 baseboard peripherals enabled (display, GNSS,
battery). --features rak-baseboard aggregates the three baseboard
features. Use this for a complete WisMesh Pocket V2.
cd leviculum-nrf && LEVICULUM_USB_PID=0002 LEVICULUM_BOARD_NAME=RAK4631 \
LEVICULUM_UF2_BOARD_ID=WisBlock-RAK4631-Board \
cargo run --release --bin rak4631 --features bsp-rak4631,rak-baseboard
(Justfile:1924-1930; rak-baseboard aggregate
leviculum-nrf/Cargo.toml:382)
just dfu-rak4631 PORT — DFU entry for stock Meshtastic
Triggers the Adafruit UF2 bootloader on a stock-Meshtastic WisMesh
Pocket V2 in software. Stock Meshtastic has no 1200-bps-touch handler and
the device has no externally accessible RESET pin, so this firmware-side
admin command is the only software-only DFU entry. Needed only for
the first flash from Meshtastic; after our firmware lands,
just flash-rak4631 uses the touch path and this recipe is no longer
needed. Requires the meshtastic CLI on PATH (pip install meshtastic).
just dfu-rak4631 /dev/ttyACM0
Runs:
meshtastic --port /dev/ttyACM0 --enter-dfu
(Justfile:1963-1972)
A note on disconnecting consumers
Flashing a board takes over its transport serial port. Any running
consumer of that port (for example an active lnsd pointed at it) loses
its connection when the board is flashed. The flash action is explicit
and active; no persistence is promised across it.
(leviculum-nrf/README.md:40)
The device keeps its Reticulum identity in internal flash and preserves
it across firmware updates, so re-flashing does not change the node's
address. (leviculum-nrf/README.md:42. More in Recovery.)
Verifying the build before you flash
cargo build --release (above) confirms the image links and fits flash.
If you want to lint the firmware as CI does:
just lint-nrf
(Builds both BSP feature sets under clippy with -D warnings:
Justfile:75-77.)
Next: Serial ports for wiring the flashed board into
lnsd.
LNode Firmware: USB Serial Ports
A flashed LNode presents two USB CDC-ACM serial ports to the host. Knowing which is which is the difference between reading a debug log and talking the Reticulum transport protocol.
The two ports
The firmware exposes two CDC-ACM serial ports. The lower-numbered
port is the debug log output; the higher-numbered port is the
Reticulum transport interface that carries HDLC frames. The actual
/dev/ttyACM* numbers depend on what else is plugged into USB.
The firmware exposes two USB CDC-ACM serial ports. The lower-numbered port is the debug log output. The higher-numbered port is the Reticulum transport interface that carries HDLC frames. The actual
/dev/ttyACM*numbers depend on other connected USB devices. (leviculum-nrf/README.md:44-46)
Each CDC-ACM class occupies two USB interfaces (a Communication interface plus a Data interface), so the two ports map onto four USB interface numbers:
| Port | USB interface nums | Carries |
|---|---|---|
| Debug | 00 (comm) + 01 (data) | human-readable log lines |
| Transport | 02 (comm) + 03 (data) | Reticulum HDLC frames |
(leviculum-nrf/udev/99-leviculum.rules, header comment.)
Stable device paths via udev
Because the /dev/ttyACM* enumeration order is not stable, install the
shipped udev rules to get fixed symlinks:
sudo cp udev/99-leviculum.rules /etc/udev/rules.d/
sudo udevadm control --reload-rules
(leviculum-nrf/README.md:50-53)
After the next plug-in, the symlinks point at the correct ports regardless of enumeration order. The names are board-family specific, keyed off the per-board USB PID:
| Board | USB VID:PID | Debug symlink | Transport symlink |
|---|---|---|---|
| T114 | 1209:0001 | /dev/leviculum-debug | /dev/leviculum-transport |
| RAK4631 / Pocket V2 | 1209:0002 | /dev/leviculum-rak-debug | /dev/leviculum-rak-transport |
(Symlink names and PIDs: leviculum-nrf/udev/99-leviculum.rules. The
firmware-side USB VID/PID constants:
leviculum-nrf/src/boards/t114.rs:172-173 for 1209:0001,
leviculum-nrf/src/boards/rak4631.rs:188-189 for 1209:0002.)
Multiple boards of the same kind. The short symlinks (
/dev/leviculum-transport) land on whichever device udev sees first. The rules also emit per-serial-number symlinks (/dev/leviculum-transport-<SERIAL>); use those when more than one board of the same family is attached. (leviculum-nrf/udev/99-leviculum.rules, header comment andSYMLINK+="leviculum-transport-%s{serial}"lines.)
Without the rules installed there is still a stable path: systemd's own
/dev/serial/by-id/ entries carry the firmware's USB strings and the
board serial, and the CDC interface number distinguishes the two ports
the same way (-if00 debug, -if02 transport):
/dev/serial/by-id/usb-leviculum_leviculum_T114_<SERIAL>-if00 debug
/dev/serial/by-id/usb-leviculum_leviculum_T114_<SERIAL>-if02 transport
Reading the debug port
The debug port is plain text at 115200 baud:
picocom /dev/leviculum-debug -b 115200
(leviculum-nrf/README.md:59-60)
On the debug port you will see the boot banner, the firmware git SHA and
the periodic diagnostics the firmware emits: the [FW_BUILD] banner
every 5 s, the [STACK] watermark lines, and the LoRa TX/RX events.
(fw_build_banner, leviculum-nrf/src/bin/t114.rs:1342-1352, for the
banner task.) Do
not point lnsd at the debug port; it carries log text, not HDLC
frames.
The hashes a prober needs are among them, on one line:
[IDENTITY] identity=<32 hex> probe=<32 hex> lxmf=<32 hex> lxmf_propagation=<32 hex>
probe= is the rnstransport.probe destination — the address
rnprobe wants — and a
destination this boot did not register reads none rather than a
string of zeroes. The line is emitted once the boot has registered its
destinations and then again in the 5 s banner
(leviculum-nrf/src/identity.rs, log_banner), so attaching late
costs at most one banner period.
It has to be on the critical log path, and it is. Until 2026-09-17 the three older lines —
LNode started -- identity: …,[IDENTITY] t114_node=…,[IDENTITY] t114_probe=…— went throughlog_fmt, which is runtime-gated: with no reader attached yet it counts the line and returns before the ring buffer and before the reset-surviving tail (leviculum-nrf/src/log.rs,log_fmt). The gate opens on the first DTR-assert or after 30 s, both later than the lines were written, so a reader saw only the gate's own summary,[LOG_GATE] opened, dropped N runtime lines pre-attach, and nothing re-emitted them (Codeberg #234). The per-board duplicates are gone; the banner above says the same values onlog_critical!and repeats them.
Querying panic evidence over the debug port
The debug port is not entirely write-only: it accepts one command. A
single p byte makes the firmware replay its persistent panic
evidence — the [PANIC_COUNT] total=N line and, if a post-mortem
record is stored, the full [HARDFAULT_PMRT] / [PANIC_PMRT] block,
bracketed by [PM_QUERY] begin / [PM_QUERY] done markers
(leviculum-nrf/src/usb.rs, debug_reader_task;
leviculum-nrf/src/lib.rs, postmortem_query).
This exists because the boot-time replay of the same block is emitted
exactly once into the 8 KiB log ring: after the 30 s headless fallback
opens the runtime-drain gate, runtime output laps the ring, so a host
that attaches later never sees it. The post-mortem records in retained
RAM (the .retained region in leviculum-nrf/memory.x, placed where
the Adafruit bootloader provably never writes)
survive the boot read (it marks them seen rather than erasing them)
and soft resets — power loss wipes them, and a reflash must be assumed
to — so the query can retrieve the evidence any time after the crash,
as long as the board stays powered.
That "retained RAM" is younger than the feature it carries. Until
9b4d82a (2026-08-31) the same five cross-boot records lived in
.uninit, which flip-link packs against the top of RAM — and the top
of RAM is where the Adafruit bootloader starts its stack
(__StackTop = 0x20040000, nrf_common.ld, confirmed by the initial
SP in the shipped bootloader's vector table). Every reset runs that
bootloader before our reset handler, so the hardfault post-mortem (top
36 B) and the boot trace (top 48 B) were overwritten on every boot
and could never have been read back; the panic post-mortem, the panic
counter (Codeberg #65) and the persistent log tail sat lower in the
same 3.1 KiB and survived only as far as the bootloader's stack
happened not to reach on a given boot. The evidence was a live positive
control on the rig: three consecutive commanded resets out of a running
system, with the reset cause latched as sreq=1, still read
prev_magic=absent prev_boot=0. The fix was placement, not logic — a
dedicated RETAINED region below the bootloader's stack floor and
outside every region it declares, held there by two link-time
ASSERTs. Read a [PANIC_COUNT] or [PM_QUERY] result from firmware
older than 9b4d82a as unreliable rather than as a zero.
The committed helper drives the whole exchange:
scripts/lnode-panic-query.sh /dev/leviculum-rak-debug
It asserts DTR+RTS (the debug port transmits only with DTR raised),
sends p, and prints the tagged response lines. Exit 0 means a
complete response was captured; on older firmware without the query
command it times out with exit 1. Do not power-cycle a board whose
evidence you still need — the retained region lives in RAM, and power
loss is the one thing that wipes it.
Pointing a daemon at the transport port
The transport port carries HDLC-framed Reticulum packets. It is not
an RNode: a standalone LNode runs a complete stack in its own firmware
and is the daemon's neighbour node, not its radio. The firmware
implements no RNode KISS command set — there is no CMD_DETECT,
CMD_FW_VERSION or CMD_PLATFORM responder anywhere in
leviculum-nrf/ — so RNodeInterface cannot drive it, and neither can
rnodeconf. The interface type is SerialInterface.
SerialInterfaceis a raw serial HDLC link […] Leviculum'sSerialInterfacehonours [the LoRa keys] too and configures the attached LNode's radio over the serial port — the LNode frames HDLC, so it cannot be driven by the KISS-framedRNodeInterface. (docs/src/guide/configuration.md:327-336)
[interfaces]
[[LNode T114]]
type = SerialInterface
enabled = yes
port = /dev/leviculum-transport
speed = 115200
databits = 8
parity = none
stopbits = 1
frequency = 869463000
bandwidth = 125000
txpower = 22
spreadingfactor = 8
codingrate = 5
For a RAK4631 / WisMesh Pocket V2 the only change is the port
(/dev/leviculum-rak-transport).
Who applies the LoRa keys. Under lnsd the five LoRa keys are sent
to the board as a radio-config frame at interface startup
(leviculum-std/src/interfaces/serial.rs:719), so the config decides
the channel. Under Python-RNS rnsd they are inert: its
SerialInterface reads port settings only and pushes nothing to the
board, which then keeps whatever profile is in its flash — the compiled
eu_medium default (869.463 MHz, BW 125 kHz, SF8, CR4/5, 22 dBm;
leviculum-nrf/src/lora.rs:461-490, RadioConfig::eu_medium) or the
preset chosen at flash time. The values above are that default written
out, so a Python-driven LNode and an lnsd-driven one land on the same
channel. Changing the channel of a Python-driven board is a reflash
(lnflash --radio-preset), not a config edit.
After editing /etc/reticulum/config, restart the daemon so it picks up
the new interface:
sudo systemctl restart lnsd
(Same restart flow as any config change; see the lnsd Quickstart.)
Confirming the link came up
Run the standard health-check and look for the new interface in the
interface_stats section with status=up and non-zero counters once
LoRa traffic flows:
lnstest diag --config /etc/reticulum
(lnstest diag usage and the interface_stats reading are described in the
lnsd Quickstart.)
Finding the node's destination hash
A standalone LNode answers probes on one destination,
rnstransport.probe, and announces it 15 s after boot and then every
2 hours (schedule_initial_mgmt_announce, leviculum-core/src/node/mod.rs:748-749;
MGMT_ANNOUNCE_INTERVAL_MS, leviculum-core/src/constants.rs:219).
The hash is carried in the announce itself, but it is also printed on
the debug port — the probe= field of the [IDENTITY] banner, repeated
every 5 s (see Reading the debug port). That
is the quicker route when the board is cabled. Receiving an announce is
the route that needs no cable, and the one below.
With the interface configured and the daemon running, press the board's reset button and wait about 20 s. The daemon reopens the port by itself after the board re-enumerates, then records the announce:
rnpath -t
<6a1ab9ea64747f298c1f205dfcf0f5a3> is 1 hop away via <6a1ab9ea64747f298c1f205dfcf0f5a3> on SerialInterface[LNode T114]
The entry on the LNode's own interface is the board. The leading hash
is the destination; the via hash is the node's transport ID, which is
the same value here because a directly attached neighbour announces at
hop 0. Probing it takes the aspect name as well, since the name cannot
be recovered from the hash:
rnprobe rnstransport.probe 6a1ab9ea64747f298c1f205dfcf0f5a3
Valid reply from <6a1ab9ea64747f298c1f205dfcf0f5a3>
Round-trip time is 126.497 milliseconds over 1 hop
Miss the 15 s window and the next announce is 2 hours out; resetting the board again is quicker.
The probe destination is the only addressed service the firmware
offers. Remote management is not enabled on the standalone binary
(leviculum-nrf/src/bin/t114.rs:183 sets respond_to_probes and
nothing else), so rnstatus -R and rnpath -R have no responder;
rncp, rnsh and rnx have no counterpart either. What the board
does beyond that — forwarding announces, answering path requests,
relaying packets — needs no hash from the operator and shows up as
paths via the LNode in rnpath -t.
For the full key-by-key reference of the serial and LoRa keys, and the
meaning of the optional ones (flow_control, airtime_limit_*,
preamble_symbols), see the RNode and Serial section of the
Configuration
chapter.
The USB control envelope
The LNode's transport CDC carries HDLC-framed Reticulum packets, plus a
small out-of-band control plane between an attached host (lnflash,
lnsd) and the firmware. Until Codeberg #238 that control plane was one
hand-cut magic per feature — a radio-config frame and a reset frame, each
recognised by shape. Three pending features each wanted a third magic,
which is how a channel becomes unextendable. This page documents the one
envelope every control frame rides in now, and how the two legacy magics
retire.
Wire truth lives in leviculum-core/src/envelope.rs; this page explains
it. If they disagree, the code and its tests win.
Frame layout
One envelope per HDLC frame:
[0xA4, 0xA5] [type: u8] [len: u16 BE] [payload: len bytes]
The length is strict: a frame whose payload is shorter or longer than
len is malformed. A reader that knows the envelope but not the type
answers a named refusal and stays in sync — the HDLC delimiter bounds the
frame, the header names what was skipped. Nothing envelope-shaped is ever
answered with silence; the legacy magics predate that rule and keep their
old manners (below).
Frame types
Commands (host → board):
| type | name | payload |
|---|---|---|
| 0x01 | RADIO_CONFIG | the legacy frame's parameter block (13–19 B), no magic |
| 0x02 | RESET | empty |
| 0x03 | WALL_TIME | unix seconds, u64 BE (8 B) |
| 0x04 | CAPABILITIES | empty (a query) |
| 0x05 | TELEMETRY_TARGET | see below — set or clear the telemetry target |
| 0x06 | TX_SPACING | on-air transmit spacing in ms, u16 BE (2 B) |
| 0x07 | RADIO_QUERY | empty (a query, #349) — answered with RADIO_REPORT |
| 0x08 | FIXED_POSITION | see below — set or clear the user-set position |
| 0x09 | MEDIA_PROFILE | one flag byte (bit0 lora, bit1 ble) — answered with MEDIA_REPORT |
| 0x0A | MEDIA_QUERY | empty (a query) — answered with MEDIA_REPORT |
| 0x0B | POSITION_SOURCE_QUERY | empty (a query) — answered with POSITION_SOURCE_REPORT |
| 0x0C | NODE_NAME | see below — set or clear the operator-chosen name |
| 0x0D | NODE_NAME_QUERY | empty (a query) — answered with NODE_NAME_REPORT |
| 0x0E | IDENTITY_QUERY | empty (a query) — answered with IDENTITY_REPORT |
| 0x0F | ANNOUNCE | empty — announce now: the LXMF delivery destination, and the propagation destination where that role runs (#376, #384) |
| 0x10 | BLE_TX_GAP | BLE inter-packet gap in ms, u16 BE (2 B), 0..=5000 (#376) |
| 0x11 | STORE_STORM | record count and body size, two u16 BE (4 B), 1..=1000 and 0..=1024 (#384) |
| 0x12 | PN_CONFIG | announced stamp cost and required peering cost, one byte each (2 B); 0xFF in a field keeps the persisted value (#384) |
| 0x13 | MGMT_ALLOW | see below — set or clear the remote-management allow-list (#235) |
| 0x14 | MGMT_ALLOW_QUERY | empty (a query) — answered with MGMT_ALLOW_REPORT |
Responses (board → host):
| type | name | payload |
|---|---|---|
| 0x81 | ACK | [acked_type] |
| 0x82 | REFUSAL | [refused_type, reason] |
| 0x83 | CAPABILITY_REPORT | [version, accepted types...] |
| 0x84 | RADIO_REPORT | the RADIO_CONFIG parameter block the radio is running (#349) |
| 0x85 | MEDIA_REPORT | [running_flags, configured_flags] in the MEDIA_PROFILE flag encoding |
| 0x86 | POSITION_SOURCE_REPORT | one flag byte (bit0 fixed position set, bit1 GNSS built in and active) |
| 0x87 | NODE_NAME_REPORT | [flags, mesh_len, mesh…, ble_len, ble…] — see below |
| 0x88 | IDENTITY_REPORT | [flags, identity(16), probe(16), lxmf(16)], 49 B fixed |
| 0x89 | MGMT_ALLOW_REPORT | [flags, count, count × identity(16)] — see below |
| 0x8A | DUTY_HOLD | [state, lt_ms: u32 BE], 5 B; state 0x00 lifted, 0x01 held. Unsolicited on each hold edge, and behind every MEDIA_REPORT — see below |
Refusal reasons: 0x01 unknown type, 0x02 malformed, 0x03 value
refused, 0x04 busy, 0x05 unsupported (the envelope layer knows the
type but this binary carries no consumer for it — retrying or rebooting
cannot help, only different firmware can), 0x06 not persisted (see
below), 0x07 no calendar clock (the command needs one and the board has
none yet — seed it with a GNSS fix or --set-time and retry), 0x08 not
running (the value is durably on the flash page and applies at the next
reset, but the carrier it configures did not come up this boot — see
below). The version in the capability report (1) names the envelope framing itself;
new frame types extend the accepted list without bumping it.
DUTY_HOLD (0x8A) — the board says when its LoRa queue is held
[state: u8] [lt_ms: u32 BE] state 0x00 lifted, 0x01 held
The one board → host frame that is not only an answer (leviculum#501).
When the regulatory airtime budget runs out, the board's LoRa TX gate
holds its queue (leviculum#493), and a board driven by lnsd over this
port has to say so, or the daemon in front of it keeps relaying announces
into a queue that cannot send them, advertising routes through a relay
that cannot carry the link setups they invite. The firmware sends the
frame unsolicited on each edge of the hold, and once more right behind
every answer to MEDIA_QUERY, so a host that attached mid-hold learns the
state without waiting for the next edge. lt_ms is the keyed airtime of
the rolling hour as the board's ledger last stood.
A second frame after the MEDIA_REPORT rather than a longer report: every host already in the field decodes the report strictly at two bytes, and skips a frame type it does not know. A state byte other than the two is refused by the decoder, never read as lifted; a payload longer than five bytes is read by its first five.
lnsd asks with one MEDIA_QUERY each time its serial interface attaches,
after the radio bring-up, mirrors the state into the interface's
duty_hold flag, logs DUTY_HOLD iface=… state=held|lifted lt_ms=… on
each change (the line its RNode path emits), and clears the flag with
cause=iface_down when the port goes away. This is the USB protocol
between our board and our host, not the radio: nothing of it goes on
the air, and no Python-RNS peer ever sees it.
MGMT_ALLOW (0x13) and MGMT_ALLOW_QUERY (0x14) — who may read the board
A board with a remote-management allow-list serves
rnstransport.remote.management with the /status handler, so rnstatus -R <board> and lnstatus -R <board> read it exactly as they read a
daemon (#235, #86). The list is identity hashes — the querying instance's
own identity, the one it signs the link with, not a destination.
MGMT_ALLOW's payload is [count, count × 16 B], at most eight
identities (leviculum_core::mgmt_allow_store::MGMT_ALLOW_MAX_IDENTITIES,
which argues the bound against the flash slot and the #388 heap budget).
count == 0 is the explicit clear. A count above the bound is refused
0x03 (value) — never truncated, because a permission set that arrives
different from the one that was sent is one nobody authorised. A payload
whose length disagrees with its own count is refused 0x02 (malformed).
Both frames are answered with MGMT_ALLOW_REPORT, not an ack, because
two states differ and only the board knows both: bit0 of its flags says a
record is on the page, bit1 says this boot registered the management
destination. The destination is created while the node is built, from the
record read at boot, so a list set now is served after the next reset —
and an identity revoked now is still being served until then. The report
also echoes what the board stored (duplicates dropped), so a host prints
the list the board acknowledged rather than a repeat of its own argv. A
record that did not reach flash is refused 0x06 (not persisted): the
whole point of a set is the next boot.
This frame is the only way the list can be written, and it is
USB-only. That is a property of where the parser sits, not of a check
inside it: classify_control_frame has exactly one caller in the
firmware, the transport CDC read path (leviculum-nrf/src/usb.rs), and
bytes arriving from the LoRa or BLE interface go to the node core as
Reticulum packets, where the same frame is dropped — its first byte
0xA4 has the IFAC bit set and no radio carrier on a board runs IFAC.
Driven on a real NodeCore in
leviculum-core/src/node/mvr_mgmt_allow_is_usb_only.rs.
An absent or empty list registers nothing at all on a board: no
destination, no handler, no announce. That is a deliberate deviation from
the daemon, which registers the handler and consults an empty list per
request the way Python does (leviculum-std/src/config.rs). A daemon
sits on a machine with an operator and a login; a board is left on a mast,
and an unattended node announcing a management destination with nobody on
the list is advertising a door.
Host side: lnflash --management-identity <hex> (repeatable) and
lnflash --clear-management, plus the flash-time question beside the
radio one.
ANNOUNCE (0x0F) and BLE_TX_GAP (0x10) — the #376 bench instruments
ANNOUNCE makes the board make every announce it makes on its own
cadence, immediately and on all interfaces. That is TWO announces on a
board running the propagation role, and one on a board without it:
- the LXMF delivery destination, exactly as the telemetry path
announces it before a report — same destination, same app data, and the
same clock gate: without a calendar clock the board withholds it, logs
[ANNOUNCE] withheld reason=no-clockon the debug port and, if it has nothing else to announce, refuses with reason0x07. (The gate is not cosmetic: the emission timestamp inside the announce is what peers rank paths by — seedocs/src/protocol-notes/announce-dedup-and-path-replacement.md— so an uptime-stamped announce would poison the path under measurement.) On success:[ANNOUNCE] sent dst=<hex8> reason=host. - the lxmf.propagation destination, where the role runs (#384),
exactly as the role announces it on its 300 s interval — and NOT
clock-gated, because a clockless board still announces the role with its
uptime timebase (#384 item 6) and the contact that invites is what
delivers a clock seed. On success:
[ANNOUNCE] sent dst=<hex8> reason=pn-host. Withheld only when the record store did not mount (PN announce withheld reason=store-unmounted): a role that cannot prove an upload must not invite one.
Both, and not just the first, because Reticulum's identity cache is keyed
by destination hash: a client that heard the delivery announce still
cannot address the board's mailbox, so set_outbound_propagation_node
fails with identity for <hash> not known until the role's announce
arrives too.
The ack therefore grades the COMMAND: OK once at least one announce left
the board, reason 0x07 (no clock) when the only announce this board has
was withheld by the clock gate, and reason 0x03 (unsupported) when it
has none to make. Which ones went out is on the [ANNOUNCE] sent ... reason= lines, beside the usual BLE_TX_PKT lines. One-shot; nothing is
persisted. Host side: lnflash --announce, and periculum's
announce_board step, which relays the same frame through the board's
owning daemon (the daemon holds the data port TIOCEXCL).
BLE_TX_GAP sets the gap the BLE drain leaves between the last fragment
of one packet and the first fragment of the next packet on the same
connection handle. With no value set the pumps serve the compiled
default of 100 ms (#376, the measured desk value —
leviculum-ble-tx's DEFAULT_TX_GAP_MS); any set value overrides it,
0 disables the gap entirely, and values above 5000 ms are refused
with reason 0x03. Interface-layer only, per connection — the fan-out
and the core never learn of it — and volatile like TX_SPACING: a reset
restores the default. The board logs [BLE ] tx_gap_ms=<n> when the
value takes effect and BLE_TX_GAP conn=<h> waited_ms=<n> once per
deferred packet. Host side: lnflash --set-ble-tx-gap <ms>.
STORE_STORM (0x11) — the #384 bench instrument
Appends records synthetic records of size body bytes each to the
board's message store (leviculum-nrf/src/record_store.rs, the record log
on the 64 KiB region memory.x reserves behind the image). Nothing else
writes to that store yet: there is no LXMF propagation node, and this frame
exists so the one cost the store imposes on the rest of the board can be
measured before anything depends on it. That cost is erases — a 4 KiB page
erase holds the flash for ~85 ms (nRF52840 PS, NVMC) and the SoftDevice
has to fit it between radio events — so the question "what does a filling
store do to BLE throughput and LoRa airtime" needs a way to provoke the
erases without waiting for a mesh to fill 16 pages.
Bounds are in classify_control_frame, so every binary refuses the same
values: records must be 1..=1000 and size 0..=1024, and anything else
is refused with reason 0x03. A board whose store did not mount, or which
is still running the previous storm, refuses with 0x04 (busy) — those two
conditions are the firmware's to see, not the classifier's.
The ack means the request was accepted, not that the records are on the page: the store task appends them on its own time, which is the point (the measurement runs while it writes). What reports the result is the board's debug port:
STORE mount state=<ours|formatted> pages=<n> live=<n> free_bytes=<n> t=<ms>
STORE storm records=<n> size=<n> appended=<n> failed=<n> seq=<n> ms=<n>
STORE op_fail op=<erase|write> attempt=<n> t=<ms>
STORE stats appends=<n> fails=<n> sealed_pages=<n> t=<ms>
Nothing is persisted as configuration, and the records carry a synthetic
tag so a later purge can find them. Host side:
lnflash --store-storm <count>[,<bytes>].
NODE_NAME (0x0C) and NODE_NAME_REPORT (0x87)
The name an operator chooses for a board, replacing both derived
defaults at once — the LXMF announce's display name (LNode-<hex8>, what
Columba lists) and the BLE device name (LN-<hex8>, what a phone shows in
its Bluetooth settings). A board answering to two different names in two
places would be worse than the hex it replaced. The name is display only:
it never touches the identity, so two boards may carry the same name and
stay distinguishable everywhere it matters.
Set payload, the FIXED_POSITION set/clear shape on a variable-length value:
[set: u8] ([name: 1..=32 bytes of UTF-8])
set is 0x00 (clear, back to the derived defaults; 1-byte payload) or
0x01. No length byte — the envelope header already carries the frame
length. The 32-byte bound is airtime policy, not a wire limit: the
name rides in every announce, so leviculum_core::node_name derives it
from the announce's on-air cost and leviculum-lxmf/tests/ announce_name_airtime.rs pins every number in that derivation. Invalid
UTF-8, control characters, surrounding whitespace and an over-long name
are all refused as malformed rather than silently shortened: a name that
arrives different from the one that was typed is worse than an error.
The report answers both frames:
[flags: u8] [mesh_len: u8] [mesh…] [ble_len: u8] [ble…]
flags bit0 is "a name is stored" (as opposed to both names being
derived) and bit1 is "the BLE surfaces are one reset behind". Unknown bits
are kept, not refused.
The two names are the effective ones, not the stored record, because a
host cannot derive either: the two defaults are different strings built
from an identity hash the host never sees, and the BLE name is
additionally shortened to leviculum_ble_tx::DEVICE_NAME_LEN (11 bytes)
on a codepoint boundary. They also adopt the name at different moments —
the mesh name is in force for the next announce, while the advertisement
was built once at boot and cannot be rebuilt under a live SoftDevice — and
bit1 is the board saying so. That is the MEDIA_REPORT
running-versus-configured argument on a second feature.
A board that has not yet published its identity hash (USB comes up several
statements into the firmware's main, the node only after the LoRa
bring-up's awaited SPI transactions) answers busy and applies nothing,
so the host's retry is a real retry. unsupported is reserved for a
binary that carries no name gate at all.
What an answer on the persist path means (#358)
Four frames write a flash record: TELEMETRY_TARGET (0x05), FIXED_POSITION (0x08), MEDIA_PROFILE (0x09) and NODE_NAME (0x0C). For those four the answer carries a durability promise:
When the client's call returns, a reset cannot lose the setting.
The board therefore does not answer them until its store task confirms
the record is on the page. An ACK — or, for the media profile and the
node name, their report — means written, not merely applied. A write the store task
gave up on comes back as a refusal with reason 0x06: the board is
running the value, and cannot promise it survives a reboot. That is a
different sentence from busy (retry) and from value refused (the
value was fine), so a client can tell it apart and say so.
The wait is bounded at 2.5 s, inside the 3.5 s window lnflash gives one
control conversation; a store task that never confirms is reported as
0x06 rather than left holding the port. Until #358 the answer went out
between the RAM apply and the page write, so a scripted set followed by
a reset — periculum's per-scenario media application, lnflash, any
automation — could reboot the board inside the window and lose the
setting. A sleep in front of the reset does not close it: the store task
may be working an earlier queued write, and a constant cannot bound a
queue.
What the RADIO_CONFIG answer means
The same promise-shape on the radio path: a RADIO_CONFIG ack means the
radio is running this configuration — the serial task waits until the
LoRa task confirms the apply and the running config matches what was
delivered, bounded at 1.2 s on top of the 500 ms delivery grace, inside
the tightest host window (lnsd's legacy sender waits 2 s per attempt).
A config delivered but not yet confirmed — a retune deferring to a frame
mid-air on a slow profile, or a reconfig that failed on the SPI bus —
answers busy, and a retry after the apply lands is acked immediately
because the running config already matches. Until this wait existed the
ack went out on channel delivery, measurably 1.1–18.9 s before the
apply while the LoRa loop parked in single-mode RX, and even when the
reconfig then failed.
That retry is answered for the config that is queued, not for the
channel. The config channel holds one slot, and a host re-sending the
config it was just told was busy finds that slot still holding its own
first copy. That is not a delivery that failed: the board answers busy
again and acks as soon as the apply lands, rather than spending the
host's attempt on undeliverable. Only a slot held by a different
config is refused, and only after the 500 ms grace. Before 2026-09-23 a
repeat was refused: lora_path_discovery_wide_mixed had its config
delivered on the first attempt (from a site=yield RX window the loop
did not wake from, so the apply missed the 1.2 s wait), then attempts two
and three were refused as undeliverable and the cell was skipped
no_ack_after_3. Every RX window the LoRa loop can park in now wakes on
a queued config, which is the other half of the same fix.
One boot state changes the promise: a board whose boot did not bring the
LoRa carrier up (a lora=off media profile in flash) has no LoRa task,
so no config can be delivered or applied before the next reset. The
config goes to the flash store, and the answer waits for the confirmed
page write (#358) — a failed write refuses with 0x06 (persist), a
confirmed one refuses with 0x08 (not running). A refusal rather
than an ack because the only true claim here is a reboot comes back on
this configuration, which is not the claim an ack makes: until #363 both
states sent the same three bytes, and a host that reads that ack as "the
board is on this PHY" prices every frame at a modulation nothing is
keying. 0x08 is the mirror of 0x06: persist is
applied-but-not-durable, not-running is durable-but-not-applied, and
neither is a rejection — the value was taken both times.
The legacy magic frame keeps its ACK in this state, because ack-or-silence
is its whole vocabulary: there is no room in three bytes for a carrier
flag, and silence reads as "the frame never landed" to a sender whose next
act is the reset that applies the page. That is the contract the test
harness relies on when it pushes the scenario channel one reboot early and
resets afterwards. A legacy host that needs to know whether the radio is
running the configuration asks RADIO_QUERY, which a board with no LoRa
task refuses as busy rather than answering out of the flash page. Both
dialects decide this in one place
(leviculum_core::envelope::radio_config_answer and
legacy_radio_config_acked) so the pair cannot drift.
Before 2026-09-22 such a boot fed the config to the taskless channel
instead: the first one wedged its single slot for the rest of the boot,
every later one was refused as busy, and one BLE-profiled boot cost a
corpus run all 26 of its LNode cells
(SKIPPED_INFRA reason=lnode_radio_config_failed result=no_ack_after_3).
The media-profile frames
[flags: u8] bit0 = lora, bit1 = ble; set means the carrier is enabled
Both media frames are answered with a MEDIA_REPORT rather than an ACK,
because the two profiles it carries can honestly differ. running is
what the board is carrying traffic on right now; configured is what a
reset would come up with. They part exactly when a carrier that did not
come up at boot is switched on: the board has no driver task to start,
and an ack would claim it did. A flag byte with a bit outside the two
known carriers is malformed, never masked down to "that carrier is
off" — the firmware does not get to invent a reading of a carrier it
does not know.
The default, for a board with no stored profile, is both carriers on:
absence of a record must change nothing about a fielded board. Concept
and semantics: docs/src/concepts/media-profiles.md.
The wall-time frame calls the calendar seam
(set_wall_time_unix_secs(.., TimeSource::Host)); the seam's sanity
window decides between the ack and a value refused refusal, and an
accepted seed logs [TIME_SEED] source=host and flips the banner's
[TIME_SOURCE] to host — the exact mirror of the GNSS path.
The transmit-spacing frame (#345)
[spacing_ms: u16 BE]
The gap the board's LoRa interface leaves between the end of one packet's
airtime and the key-up of the next. It is applied inside
transmit_all_frames, the last thing before the radio is keyed, so it is a
gap between two packets on the air rather than between two hand-overs, and
whatever the transmit path already spent since the previous packet ended
(the CAD, the SPI traffic, the log lines) is counted against the requested
gap rather than added to it. The split frames of one packet are unaffected:
they still go out back-to-back, because the receiver's reassembler requires
that.
Every u16 value is legal, 0 included — 0 is the compiled default and
imposes nothing, so the only malformed frame is one of the wrong length.
The value is not persisted: it is a measurement instrument (the sweep of
the telemetry announce/report spacing, #345), and a reset returns the board
to the default. The board logs [LORA_TX_SPACING] intended_ms=… waited_ms=… gap_ms=… at every key-up; gap_ms is the gap that was measured, and -1
is the first packet since boot, which has no previous airtime edge to be
measured from.
lnflash --set-tx-spacing <MS> is the host side.
The telemetry-target frame (#236)
[profile: u8] [dest_hash: 16] [key_present: u8] ([public_key: 64])
key_present is 0x00 or 0x01, never inferred from the length: per
the #236 UX decisions (2026-08-22) the public key is optional and
hash-only is the common case — the user knows the LXMF address, the node
resolves the key over the air.
Profile ids:
| id | name | meaning |
|---|---|---|
| 0x00 | OFF | clear the target — telemetry off |
| 0x01 | TRACKER | movement-driven cadence |
| 0x02 | STATION | slow heartbeat only; the default profile |
0x00 is the clear encoding. It rides in the profile slot rather than
in a magic destination hash because that slot's whole job is to say
which cadence applies, and "none" belongs in its vocabulary; the rest of
the payload is still parsed and must still be well formed, so a clear
frame is not a licence to send a short one. The destination hash and key
of a clear frame are ignored, and encode_telemetry_clear zeroes them
rather than echoing a target back for no reason.
An id the firmware does not know is not a refusal: the destination
is kept and the default profile's cadence runs, because a newer host's
cadence preference is not worth losing a configured target over. Which
profile is actually running is in the board's [TELEMETRY] banner.
Firmware from before #236 answers this type with an unknown type
refusal and leaves it out of its capability report, which is precisely
how a #236-aware host detects a pre-#236 board.
How lnflash drives it
Telemetry is configuration, not firmware, so the same frame is reachable from the flash flow and without flashing anything:
| flag | effect |
|---|---|
| (none) | after the radio step: Send telemetry? [y/N], default no |
--telemetry <ADDRESS> | implies yes; 32 hex chars, spaces/colons/case tolerated |
--telemetry-profile <tracker|station> | which cadence; default station |
--telemetry-key <128 hex> | the key-present form; absent = hash-only, the common case |
--no-telemetry | send profile 0x00 — clear whatever the board had stored |
--set-telemetry | the same configuration on running boards, no flash |
Answering no at the prompt sends nothing; --no-telemetry sends a
clear frame. The difference matters on a board that already has a target:
silence leaves it, the clear frame removes it.
A yes needs exactly one input — the LXMF address — because that is what
users have. Nothing detects a terminal: Ui::ask answers "no answer" for
--yes and for a piped or closed stdin alike, and every prompt treats
that as its stated default, so a scripted run cannot block.
What the host reports back is the ack. The node's own
[TELEMETRY] target=… state=off|no-position-source|awaiting-key|ready
line goes to the debug CDC (if00), which lnflash holds open only for
the post-flash boot check — so it is named as the place to read the rest
rather than read back over a second connection.
The consequence sentence. A target alone does not make a board
report: sending the position is the switch for sending everything
(docs/src/concepts/telemetry.md), so a board with neither a fixed
position nor a GNSS receiver stores the target and stays silent. After
an ack, --set-telemetry therefore asks the board itself
(POSITION_SOURCE_QUERY, on the same open port) and, when the answer is
"neither", says so:
3-2.4: target stored; nothing will be sent until a position source
exists — set one with --set-position.
Honest, not a refusal: the target is valid configuration and it is stored. A board that answers with a source is told nothing of the kind, and a board that does not answer the query at all — firmware without it, or a binary with no reporter, which refuses it by name — is told nothing either. Guessing here would put a false warning in front of an operator whose board is fine.
The fixed-position frame
[set: u8] ([latitude_e6: i32 BE] [longitude_e6: i32 BE]
[alt_present: u8] ([altitude_e2: i32 BE]))
A user-set position as the telemetry source. set is 0x00 (clear, the
1-byte payload is the whole command) or 0x01; alt_present follows the
telemetry target's key-present rule — an explicit flag byte, never
inferred from the length. Units are the telemetry wire's own scaled
integers: degrees × 1e6, metres × 1e2, so the coordinates the user typed
are the coordinates that go on the air. A latitude beyond ±90° or a
longitude beyond ±180° is refused as malformed.
Semantics (decided 2026-08-30): while set, the fixed position replaces
the position sensor entirely, in every profile — no blending, no
fallback surprises — and the explicit clear returns the node to sensor
reporting, which for a GNSS-less binary means no position. The board
persists it beside the telemetry target (same flash page, so it survives
resets and UF2 updates), marks the source in its report line as
possrc=fixed|gnss, and puts it on the wire in Sideband's own
fixed-location shape: accuracy 0.01 m, speed and bearing 0, altitude 0
when unset (Location.update_data, synthesized branch, Sideband
2000d81).
The ack is capability-gated exactly like the telemetry target's: only
the reporter reads the position, so a binary without one answers the
unsupported refusal rather than acking a pin nothing will ever report.
How lnflash drives it
| flag | effect |
|---|---|
--set-position LAT,LON[,ALT] | set it on every running board, then exit; no flash |
--clear-position | back to sensor reporting |
The value is decimal degrees, comma or space separated, sign or
hemisphere letter (52.52,13.405,34, "52.52N 13.405E", 36.85S,73.04W
all parse; a letter and a sign together do not). The optional third value
is the altitude in metres. Degrees/minutes/seconds notation is refused by
name rather than misparsed.
Why an envelope frame can never be a packet
The channel's other occupant is HDLC-framed Reticulum traffic, so every control frame must be unmistakable. Three facts hold it:
- The first magic byte
0xA4has the IFAC bit set, and this channel runs without IFAC — no peer on it emits a packet whose first byte matches, and firmware from before the envelope drops a received envelope frame in packet parsing for the same reason. - Every frame a host may send before it knows the peer speaks the envelope — the capability probe, wall time, reset — is shorter than the 19-byte minimum Reticulum wire packet, so it cannot be packet-shaped at all.
- Frames at that size or beyond (radio config at 24 B, telemetry target at up to 87 B, a set fixed position at exactly 19 B) are only sent after a capability report proved the peer is envelope-speaking firmware. This ordering is load-bearing: an envelope speaker must probe before it sends any envelope frame of 19 bytes or more.
Compatibility window, and how it retires
The two legacy magics stay accepted, with their legacy answers, so both field directions keep working:
- Old host tool → new firmware: the legacy 21-byte config magic and
the 4-byte reset magic are classified ahead of the envelope
(
classify_control_frame) and answered with the legacy two-byte-style acks (RADIO_CONFIG_ACK,RADIO_RESET_ACK). An invalid legacy config keeps its historical silence; audible refusals begin with the envelope. - New host tool → old firmware:
lnflashopens every control conversation with a capability probe. Firmware that answers gets envelope frames; firmware that stays silent (pre-envelope) gets the legacy config magic as a fallback, and--set-timereports "this firmware predates the control envelope" by name instead of guessing.
lnsd still speaks the legacy config magic on every connect; it migrates
to the envelope in its own batch.
Retirement happens in that order: first lnsd and every shipped host
tool speak the envelope (probing, with fallback), then — after a release
cycle in which lnflash bundles only envelope-speaking firmware, so any
field board a current tool meets accepts it — the firmware drops the two
legacy classifier arms and the host tools drop the fallback. Each step is
observable: a host that still needs the fallback logs it, and a board
that still receives legacy magics is running firmware older than the
bundle that introduced the envelope.
Adding a fourth frame type
The definition of done for #238: allocate the next type constant in
leviculum-core/src/envelope.rs, give it a payload codec with tests, add
a ControlAction variant and its executor arm in
leviculum-nrf/src/usb.rs, and append the type to
ACCEPTED_CONTROL_TYPES so the capability report advertises it. The
framing, the refusal path, the probe, and both host speakers stay
untouched.
LNode Firmware: Bootloader Entry and Recovery
The nRF52840 boards use the Adafruit UF2 bootloader: it appears as a
mass-storage drive, and writing a .uf2 file to that drive flashes the
device. This page covers how to enter that bootloader (touch-free and
manual), the board-specific caveats, what survives a re-flash, and what
to do when USB stays dark.
Physical-device steps. The author of this page cannot operate a board. Every step that presses a button, taps a pinhole, or observes a drive appearing is derived from source — requires the physical device. The commands and mechanisms are quoted from
leviculum-nrf/README.mdand theJustfile; only the hardware outcome is un-verified here.
Entering the UF2 bootloader
Touch-free (1200-baud), the common case for T114
When the LNode firmware is already running, the host can drop it into the bootloader without any physical interaction: it opens the board's transport CDC port at 1200 baud, the firmware intercepts the line-coding change, writes a retained-register magic value, and soft-resets into the Adafruit UF2 bootloader.
The host opens each T114's transport CDC port at 1200 baud, the firmware intercepts the line-coding change, writes a retained-register magic, and soft-resets into the Adafruit UF2 bootloader. No physical button press required. (
leviculum-nrf/README.md:27)
All just flash* recipes use this path automatically when the device is
running our firmware. (derived from source — requires the physical
device.)
Manual RESET double-tap, the fallback
A physical double-tap of the RESET button forces the UF2 bootloader
regardless of firmware state. You need it when the firmware on a specific
board has crashed or never reached USB init — a panic before the
1200-baud handler is installed, a stack overflow, or a hardware fault. In
a just flash batch the runner detects this per device via the
UF2-drive-polling timeout and prompts for that specific board only; the
rest of the batch keeps flashing touch-free.
the firmware on a specific T114 has crashed or never reached USB init (panic before the handler is installed, stack overflow, hardware fault). The runner detects this per device via the UF2-drive-polling timeout and prompts for that specific T114 only. (
leviculum-nrf/README.md:38)
(derived from source — requires the physical device.)
WisMesh Pocket V2 (RAK4631): the hidden-pinhole caveat
The RAK WisMesh Pocket V2 has no externally accessible RESET pin, so the ordinary double-tap-the-button trick does not apply. On this board:
-
First flash from stock Meshtastic. Stock Meshtastic has no 1200-baud-touch handler, so the touch-free path does not work yet. Use the software DFU command instead:
just dfu-rak4631 /dev/ttyACM0which runs
meshtastic --port /dev/ttyACM0 --enter-dfu. This firmware-side admin command is the only software-only DFU entry on a board with no accessible RESET pin. Requires themeshtasticCLI (pip install meshtastic). (Justfile:1963-1972) -
Manual fallback. Where the software command is unavailable, the bootloader is reached by a needle double-tap in the hidden pinhole — there is no visible reset button; the reset contact is reachable only through a small pinhole, double-tapped with a needle. (This pinhole detail comes from project field notes, not from the firmware source; the source confirms only that the device "has no externally accessible RESET pin",
Justfile:1964-1965.) -
The same pinhole gives a plain reset with a single tap, which is what When USB stays dark asks for first: one tap restarts the board and keeps RAM, two taps enter the bootloader. The pinhole is the Pocket V2's only reset, so on this board the evidence-preserving recovery and the flashing dance go through the same hole and differ only in the number of taps.
-
After our firmware lands, subsequent flashes use the touch handler in
src/usb.rsand the DFU recipe is no longer needed. (Justfile:1966-1967)
Do not flash foreign nRF52 firmware onto the Pocket V2 without a recovery plan. Project field experience is that prebuilt third-party nRF52 firmware may not boot on this RAK board (USB stays dark). Because the only software DFU entry is firmware-side, a board that boots into a non-responsive image and exposes no RESET pin can be hard to recover. (This caveat is project knowledge; it is not stated in the firmware source, which documents only the missing RESET pin and the firmware-side DFU command.)
All steps in this section are derived from source / project notes — requires the physical device.
Identity persistence across updates
A re-flash does not change the node's Reticulum address. The device stores its Reticulum identity in internal flash and preserves it across firmware updates.
The device stores its Reticulum identity in internal flash and preserves it across firmware updates. (
leviculum-nrf/README.md:42)
Mechanically, the firmware loads the identity from a dedicated flash page at boot and only generates (and saves) a new one when none is present:
if id_store.load() => Some(identity) -> "Identity loaded from flash"
else -> generate new, then save
(leviculum-nrf/src/bin/t114.rs:181-285,
leviculum-nrf/src/bin/rak4631.rs:216-320. The identity lives on the
board's identity_flash_page, e.g. 0xEC000 on the T114,
leviculum-nrf/src/boards/t114.rs:177.) Flashing new firmware rewrites
the program region but leaves that page intact, so the node keeps its
address. You can confirm the loaded identity on the debug port: the boot
log prints Identity loaded from flash
(leviculum-nrf/src/bin/t114.rs:241) and an [IDENTITY] line with the
full hash (leviculum-nrf/src/bin/t114.rs:616, and again on the 5 s
banner). Both are on the boot-critical log path, so attaching after the
board has come up still shows them (Codeberg #234).
When USB stays dark
A board that enumerates nothing is also a board that cannot say why, and
the only witness is in RAM: a breadcrumb record in the RETAINED region
carries how far the last boot got, plus that boot's POWER.RESETREAS.
It rides through a reset and dies with the power
(leviculum-nrf/memory.x, the RETAINED comment; capture,
leviculum-nrf/src/boot_trace.rs:57). So the order of the recovery
steps decides whether a dark board is diagnosable or only a tally mark.
Codeberg #359 has paid that price once already: the recurrence of
2026-09-02 was recovered with a power cycle and answered
prev_magic=absent reset_reason=0x00000000 on the next boot, which is
the instrument being honest, not the instrument failing.
If the board enumerates nothing on USB after a flash or a bad image:
-
Do not remove power, and on a board with a battery do not pull the cell. Power loss is the one thing that wipes the retained region. On a battery-backed board a host-side VBUS cycle is not a recovery anyway: the crashed image keeps running off the cell, so the cycle costs nothing and buys nothing.
-
Single-tap RESET. One tap is a pin reset: the core restarts and RAM is left alone. On a T114 that is the button; on a Pocket V2 it is one needle tap in the hidden pinhole (see above), not two. A double tap is the bootloader, not a reset, and the bootloader prints no trace: it is the app that reads the record and logs it. It can also destroy it — in OTA-DFU mode the bootloader enables the SoftDevice itself, whose RAM then reaches up over the retained band (
leviculum-nrf/memory.x, theRETAINEDcomment). Keep the double tap for step 4, once the trace has been read. -
Read the debug port at 115200 baud. The first line of the boot banner is the trace:
picocom /dev/leviculum-debug -b 115200BOOT_TRACE prev_magic=ok prev_phase=usb-up prev_boot=17 reset_reason=0x00000004prev_phaseis the last milestone the DEAD boot completed, so it names where that boot stopped: anything beforemain-loopsays it hung right after the named milestone, andmain-loopsays no boot after that one ever reachedmainat all, which puts the hang in the bootloader or in startup rather than in the firmware. The milestone names and thereset_reasondecode are in Structured event logs. The same port replays the previous boot's HardFault/panic post-mortem and the persistent log: look for[HARDFAULT_PMRT],[PANIC_PMRT], and[PERSISTENT_LOG](leviculum-nrf/src/bin/t114.rs:97-158;leviculum-nrf/README.md:59-60). -
Only now force the bootloader manually. On a T114, double-tap RESET to get the UF2 drive regardless of the running image (
leviculum-nrf/README.md:38). On a Pocket V2, use the hidden-pinhole needle double-tap (see above) — the board has no accessible RESET pin (Justfile:1964-1965). -
Re-flash the known-good LNode firmware once the UF2 drive appears:
just flash(T114) orjust flash-rak4631/just flash-rak4631-pocket(RAK4631). See Flashing.
If the board comes back on the single tap, steps 4 and 5 are not needed and the trace is the report. If it stays dark through the pin reset, the trace is gone either way and the bootloader is the next move.
(All hardware steps: derived from source / project notes — requires the physical device.)
ESP32 RNodes vs. nRF52 LNodes. The bricking risk above is specific to the nRF52 LNodes. The ESP32-based RNodes (LilyGO T-Beam) have a mask-ROM download bootloader and cannot be bricked: a failed flash is always recoverable by re-running the flash recipe. The nRF52 LNodes (T114, RAK4631) are different — a bad external image can leave the device USB-dark, which is why a recovery plan matters here. (
Justfile:1974-1978)
Debugging with the Debug Probe (SWD)
A reliable, bootloader-free workflow for debugging the Leviculum nRF52840 LNode firmware (RAK4631 / T114) over SWD, using a Raspberry Pi Debug Probe and probe-rs. It replaces the UF2-bootloader / 1200-baud-touch / pinhole flashing dance and adds a USB-independent log (RTT) plus full register and memory access.
What it is for
Use SWD when you need to look inside the firmware, not just talk to it:
- Firmware crashes, hard faults and SoftDevice asserts.
- Reboots under load (the USB-CDC log port drops exactly when the board re-enumerates or reboots, so USB logging loses the interesting moment).
- Register and memory inspection (RESETREAS, heap, the reset-cause markers).
- Reliable flashing every time, with no bootloader and no double-tap.
RTT streams the firmware log straight through a reboot, because SWD is a separate physical bus from USB.
How: the commands
Everything runs through scripts/probe-debug.sh <cmd> [board], wrapped by the
just probe recipe. Default board is rak4631; pass t114 for the other.
| Command | What it does |
|---|---|
just probe info | chip and debug-port info; confirms wiring and APPROTECT open |
just probe reset | reset the target over SWD |
just probe flash rak4631 | build and flash via SWD (no bootloader) then reset |
just probe rtt rak4631 | stream the live RTT log (firmware built with rtt) |
just probe gdb rak4631 | start a probe-rs GDB server on :1337 |
just probe read <hex-addr> <n> | read n bytes of target memory |
For the reboot-cause repro under LoRa load, use scripts/catch-reboot.sh [board], described in the worked example below.
The probe binary and its behaviour are configurable by env var: LEVICULUM_PROBE
(default 2e8a:000c, the RPi Debug Probe CMSIS-DAP), PROBE_RS (default
~/.cargo/bin/probe-rs), and LEVICULUM_RTT=1 to build the RTT debug firmware.
One-time setup
- Install probe-rs:
cargo install probe-rs-tools(needs >= 0.31). Optionallygdb-multiarch,binutils-arm-none-eabi,picotool,tiofor GDB and probe maintenance. - Install a udev rule so the probe is reachable without root, for example
/etc/udev/rules.d/69-probe-rs.rulesfrom the probe-rs docs. sudo works as a fallback if you skip this. - The probe's OWN firmware must be >= 2.2.0 (CMSIS-DAP v2). Update it if probe-rs
complains: see
scripts/probe-debug.sh fw-update.
Wiring
Probe D connector (Debug / SWD) to the board SWD pads:
Probe D cable | Board pad |
|---|---|
| orange | SWCLK |
| yellow | SWDIO |
| black | GND |
Do NOT connect 3V3/VTref or RST (the board is self-powered; RST is not needed).
- RAK4631 (inside the WisMesh Pocket V2): the RAK4631 module's 5-pin SWD port,
pads labelled
SWDIO SWCLK RST 3V3 GND. Pinch the V2 case open to reach it. - T114: header
P1, SWCLK =P1.13, SWDIO =P1.15, GND = any GND pin.
Our lab rig (schneckenschreck via VFIO)
On our rig the probe plugs into a USB controller that is VFIO-passed-through to the
VM schneckenschreck, so it appears there as 2e8a:000c Raspberry Pi Debug Probe (CMSIS-DAP). This is one example topology, not a requirement: a probe on a plain
host USB port works the same way. The VFIO path has its own failure modes, noted
under Known issues below.
RTT (live log over SWD)
Build the firmware with the rtt feature, then it mirrors every log line to an RTT
channel that probe-rs streams, unaffected by a USB reboot:
cd leviculum-nrf
cargo build --release --bin rak4631 --features bsp-rak4631,rak-baseboard,rtt
just probe flash rak4631 # flashes the rtt build over SWD
just probe rtt rak4631 # live log, including straight through a reboot
Production builds omit rtt and are byte-identical to before.
GDB (for faults and live inspection)
just probe gdb rak4631 # starts the server on :1337
gdb-multiarch leviculum-nrf/target/thumbv7em-none-eabihf/release/rak4631 \
-ex 'target extended-remote :1337'
CAUTION on the RAK: it runs the SoftDevice (BLE). Halting the core (a GDB breakpoint) for more than a few ms can make the SoftDevice assert and reset the chip. RTT (non-halting background memory access) is the safe default; use GDB breakpoints only briefly and expect the SoftDevice may not tolerate a long halt.
Worked example: catching the #50 reboot
scripts/catch-reboot.sh rak4631 drives the airtime-max repro (SF10) while
continuously capturing the board debug port (USB if00, reopen-on-EOF so it survives
the reboot). After the first reboot it reports the cause from the boot banner:
[RESET_SITE] name=<touch|panic|hardfault|none(external)>: which sys_reset fired.none(external)means none of our code, so a dependency or SoftDevice reset.[RESETREAS],[PANIC_COUNT],[SD_FAULT]from the boot banner.- the ~25 log lines before the reboot (the context).
It reads the marker from USB (the boot banner), NOT RTT: probe-rs attach halts
the core, which perturbs the SoftDevice, whereas USB if00 capture is non-invasive.
Tune the load with the RUNS, LORA_SF, LORA_CR, LORA_BANDWIDTH env vars.
Known issues and lessons (VFIO-passed-through probe)
These are lessons from our lab rig where the probe is VFIO-passed-through. On a plain host USB port the wedge modes below are unlikely, but the RTT/SoftDevice caveat still applies.
probe-rs attach(RTT) HALTS the core to set up RTT. On the RAK (SoftDevice) this freezes the firmware while attached and can leave it halted on exit. For reading the reboot CAUSE usecatch-reboot.sh(USB if00, non-invasive); use rtt only for short live inspection, andjust probe resetafterwards.- A long-running
probe-rs gdbserver can WEDGE in D-state (uninterruptible) on the VFIO USB and destabilise the whole rig USB (probe-rs ops and even USB serial reads start to hang;pkill -9and/procreads block). Keep gdb sessions SHORT and bounded. To recover:scripts/probe-debug.sh recover(re-enumerates the probe USB); if that is not enough, physically replug the Debug Probe (and the RAK if its USB is wedged), or reboot the VM. Thenjust probe reset. - For catching a reset CAUSE without gdb, the robust method is the reset-site marker
plus
catch-reboot.sh(USB), not a live gdb breakpoint.
Troubleshooting
probe-rslists two probes (CMSIS-DAP + ESP JTAG): the scripts always select--probe 2e8a:000c, so this is handled.- "Failed to open the debug probe" or udev warnings: the udev rule is missing or the user lacks access; re-run the one-time setup. (sudo works as a fallback.)
- "firmware ... outdated ... minimum 2.2.0": update the probe firmware (fw-update).
- "could not select JTAG": harmless; the scripts force
--protocol swd. - Probe present but no SWD: check the three wires (orange=SWCLK, yellow=SWDIO, black=GND) and that the board is powered.
- probe-rs or USB serial reads hang for minutes: a wedged
probe-rs gdb(see above); runrecoveror replug the probe.
Building on Leviculum in Rust: Choosing a Layer
Leviculum is a Rust workspace, not a single crate. The Reticulum stack is split into layers so that the same protocol engine can run on a tokio server, a bare-metal nRF52 radio, or behind a C ABI. As an application developer your first decision is which layer you build against. This chapter explains the four crates, the dependency direction between them, and gives a decision table.
The companion chapters are the Rust API tutorial (a
hands-on leviculum-std walkthrough), the Rust API reference
(verified signatures of the key types), and Embedded development
(building on leviculum-core directly). If you are writing C rather than Rust,
the C API overview and How-To are
your counterparts to those chapters.
The four layers
leviculum-ffi (C ABI) leviculum-nrf (nRF52 firmware)
│ │
▼ │
leviculum-std (std, tokio) │
│ │
▼ ▼
leviculum-core (no_std, sans-IO)
The dependency direction is strict and one-way. leviculum-std builds on
leviculum-core; leviculum-ffi wraps leviculum-std; leviculum-nrf wraps
leviculum-core directly (it never pulls in std or tokio). Nothing depends on
a layer above it.
leviculum-core — the no_std, sans-IO engine
leviculum-core is the protocol. It is no_std (it pulls in alloc, but not
the standard library), performs no I/O of its own, and owns no runtime. It is
sans-IO: you feed it received bytes, it returns a
TickOutput describing the
packets to send and the events that occurred, and you dispatch those yourself.
Time, persistence, and the network are abstracted behind three traits —
Clock, Storage,
and Interface — that you implement for your
platform.
Build against leviculum-core when you have your own runtime or event loop and
do not want tokio: embedded firmware, an integration into a different async
executor, a simulator, or a host program that wants byte-level control. See
Embedded development.
leviculum-std — the full std/tokio application layer
leviculum-std is what most Rust applications use. It supplies the platform
pieces leviculum-core abstracts: a SystemClock, file-backed storage with
Python-compatible on-disk formats, and concrete interfaces (TCP client and
server, UDP, AutoInterface for LAN discovery, RNode/LoRa, raw serial). On top of
those it runs the sans-IO core inside a tokio event loop and exposes an async,
handle-based API: build a node with ReticulumNodeBuilder,
start() it, take an EventReceiver,
and use LinkHandle / PacketSender
to send.
Build against leviculum-std when you are writing a normal Rust program on
Linux/macOS that talks to a Reticulum mesh. This is the path the
tutorial and the examples under
leviculum-std/examples/ take.
leviculum-ffi — the C ABI wrapper
leviculum-ffi exposes leviculum-std through a C-compatible ABI: opaque
handles, integer error codes, a pollable event fd. It is the layer behind
leviculum.h and libleviculum.so. If you are writing Rust you do not use it —
you use leviculum-std directly, which is what leviculum-ffi itself does
internally. It exists so that non-Rust programs (C, and anything that can call a
C library) get the same engine.
If your application is in C, stop here and read the C API overview and How-To instead; they are the C counterpart to this Rust documentation.
leviculum-nrf — the reference firmware
leviculum-nrf is standalone firmware for nRF52 boards (the T114 and RAK4631
LoRa nodes), built with the Embassy async embedded
framework. It targets thumbv7em-none-eabihf and depends on
leviculum-core directly with default-features = false — no std, no tokio.
It is both a usable firmware and the worked reference for how to drive the
sans-IO core on bare metal; the embedded chapter walks through its
main loop.
You do not "build on" leviculum-nrf the way you build on a library; you fork it
or read it as the canonical example of a leviculum-core integration on a real
device.
Decision table
| You are building… | Use | Why |
|---|---|---|
| A Linux/macOS app or daemon talking to a mesh | leviculum-std | Async handle API, real interfaces, file storage, tokio loop already wired |
A drop-in tool reusing a running lnsd/rnsd | leviculum-std | connect_to_shared_instance over the shared-instance IPC |
| A relay / transport node | leviculum-std | enable_transport(true), see relay_daemon.rs |
| A C program (any non-Rust language with C FFI) | leviculum-ffi | Stable C ABI, opaque handles, pollable fd — see the C API chapters |
| Firmware on an nRF52 LoRa board | leviculum-nrf | Reference firmware; fork or adapt it |
| Firmware on a different MCU / a custom async runtime | leviculum-core | Implement Clock/Storage/Interface, drive the sans-IO loop yourself |
| A simulator or byte-level test harness with no I/O | leviculum-core | Feed bytes, inspect TickOutput, no runtime imposed |
Adding the dependency
None of these crates are published on crates.io. Depend on them by path (in a
workspace checkout) or by git. For a leviculum-std application:
# By path — adjust to wherever, and under whatever name, you cloned the
# repository; this example assumes a sibling directory named `leviculum`
[dependencies]
leviculum-std = { path = "../leviculum/leviculum-std" }
tokio = { version = "1", features = ["full"] }
# Or by git
# leviculum-std = { git = "https://codeberg.org/Lew_Palm/leviculum" }
For embedded work depend on leviculum-core instead, with default features off:
[dependencies]
leviculum-core = { path = "../leviculum/leviculum-core", default-features = false }
The workspace is edition 2021 and licensed AGPL-3.0-or-later. Version numbers
are the crate manifests' to state, not this page's — read them from
Cargo.toml (leviculum-nrf versions independently of the workspace).
Rust API Tutorial: Building on leviculum-std
This chapter builds a small application on leviculum-std, the std/tokio layer.
By the end you will have created a node, attached an interface, registered a
destination, sent both a single packet and link data, and consumed
NodeEvents. Every snippet is
adapted from a real example under leviculum-std/examples/; each step names the
file it comes from so you can read the full program. For exact signatures of
everything used here, see the Rust API reference.
If you have not yet decided that leviculum-std is the right layer, read
Choosing a layer first.
Setup
Add the dependency and tokio. The crates are not on crates.io, so use a path (workspace checkout) or git:
[dependencies]
# Adjust the path to wherever, and under whatever name, you cloned the repository
leviculum-std = { path = "../leviculum/leviculum-std" }
tokio = { version = "1", features = ["full"] }
tracing-subscriber = "0.3"
The examples all assume a running Reticulum daemon to attach to. Start a Python
rnsd (or a Leviculum lnsd) listening on 127.0.0.1:4242, then run an example
with, for instance, cargo run --example simple_send.
Step 1: build and start a node
The entry point is ReticulumNodeBuilder.
You add interfaces on the builder, call build().await, then start().await.
This is the opening of every example; here it is from simple_send.rs:
use leviculum_std::driver::ReticulumNodeBuilder; #[tokio::main] async fn main() -> Result<(), Box<dyn std::error::Error>> { tracing_subscriber::fmt::init(); // Build a node with a TCP interface to a local daemon. let mut node = ReticulumNodeBuilder::new() .add_tcp_client("127.0.0.1:4242".parse()?) .build() .await?; node.start().await?; // ... use the node ... node.stop().await?; Ok(()) }
build() loads or generates the node's transport identity (persisted under the
storage path) and prepares interfaces, but does not run anything. start()
spawns the tokio event loop and brings the interfaces online. stop() flushes
state and tears the loop down. If you are constructing a node outside an async
context, build_sync() is the non-async equivalent of build().
Other interfaces are added the same way:
add_tcp_server(addr), add_udp_interface(listen, forward),
add_auto_interface() (IPv6 multicast LAN discovery), and
add_rnode_interface(...) for LoRa. A relay node adds enable_transport(true),
as in relay_daemon.rs:
#![allow(unused)] fn main() { // Adapted from relay_daemon.rs let mut node = ReticulumNodeBuilder::new() .enable_transport(true) .add_tcp_client(peer) .build() .await?; }
Step 2: take the event receiver and consume events
Everything inbound — announces, paths, link lifecycle, link data — reaches you as
NodeEvent values on an
EventReceiver. Take it once with
take_event_receiver() and call recv().await in a loop. From simple_send.rs:
#![allow(unused)] fn main() { let mut events = node .take_event_receiver() .ok_or("Failed to get event receiver")?; while let Some(event) = events.recv().await { println!("Received event: {:?}", event); } }
recv() behaves like a tokio::sync::mpsc::Receiver::recv (it is cancel-safe in
tokio::select!) and returns None only once the node has shut down. The
echo_server.rs example shows the real shape: match on the variants you care
about and ignore the rest.
#![allow(unused)] fn main() { use leviculum_std::NodeEvent; // Adapted from echo_server.rs loop { tokio::select! { Some(event) = events.recv() => match event { NodeEvent::LinkEstablished { link_id, is_initiator } => { println!("link up: {:02x?} (we initiated: {})", &link_id.as_bytes()[..4], is_initiator); } NodeEvent::LinkDataReceived { link_id, data } => { println!("{} bytes on {:02x?}: {:?}", data.len(), &link_id.as_bytes()[..4], String::from_utf8_lossy(&data)); } NodeEvent::MessageReceived { link_id, msgtype, sequence, data } => { println!("msg type 0x{:04x} seq {} on {:02x?}", msgtype, sequence, &link_id.as_bytes()[..4]); } NodeEvent::AnnounceReceived { announce, interface_index } => { println!("announce from {:02x?} on iface {}", &announce.destination_hash().as_bytes()[..4], interface_index); } other => println!("other: {:?}", other), }, _ = tokio::signal::ctrl_c() => break, } } }
Note the two receive variants. MessageReceived is the channel-multiplexed path
(sequenced, retransmitted) most link applications use; LinkDataReceived is the
lower-level raw-link-packet path (for example a Python peer calling
RNS.Packet(link, data).send()). The chat.rs example handles both.
Step 3: register and announce a destination
To be reachable you register a local destination and announce it. A destination
is built from your identity, a direction, a type, an app name, and aspect
strings. This is from the api module's own test, which is the most compact
worked registration in the tree:
#![allow(unused)] fn main() { use leviculum_std::{Destination, Direction, DestinationType, generate_identity}; let id = generate_identity(); let dest = Destination::new( Some(id), Direction::In, DestinationType::Single, "leviculum-test", &["api"], )?; let dh = *dest.hash(); // 16-byte DestinationHash, read before moving dest node.register_destination(dest); // consumes dest // Announce it; the optional payload rides along in the announce. node.announce_destination(&dh, Some(b"hi")).await?; }
Read dest.hash() before calling register_destination, which takes the
Destination by value. Incoming (Direction::In) destinations are auto-accepted
for links by the core (Python-RNS parity): when a peer opens a link to one, the
stack accepts and proves it automatically and you see a LinkEstablished event —
there is no separate accept call.
Step 4: send a single packet
For fire-and-forget delivery use a PacketSender,
the single-packet handle. A path to the destination must already be known (learn
it from an announce, or call request_path). Adapted from the PacketSender
doctest in driver/sender.rs:
#![allow(unused)] fn main() { let endpoint = node.packet_sender(&dest_hash); let _packet_hash = endpoint.send(b"Hello!").await?; }
send returns the truncated packet hash, which you can match against a later
PacketDeliveryConfirmed event if the destination proves delivery.
Step 5: open a link and send on it
A link is an encrypted session. Open one with connect, passing the destination
hash and its 32-byte Ed25519 signing key (the signing half of the peer's
identity, learned from its announce). You get back a
LinkHandle. Adapted from the LinkHandle
doctest in driver/stream.rs:
#![allow(unused)] fn main() { let handle = node.connect(&dest_hash, &signing_key).await?; // The handle is usable immediately, but the link is not yet established. // Watch for NodeEvent::LinkEstablished on the event receiver before relying // on delivery, then send: handle.send(b"Hello!").await?; }
connect returns as soon as the link request is dispatched; the link is pending
until a LinkEstablished event fires for its link_id. send absorbs pacing and
busy conditions by retrying internally; try_send is the non-blocking variant
that surfaces backpressure instead. Responses arrive as MessageReceived /
LinkDataReceived events on the receiver you took in step 2. Close with
handle.close().await when done.
On the responder side you do not call connect. Once a LinkEstablished event
fires for a link you did not initiate (is_initiator == false), the link is
already live; mint a writable handle for it with node.link_handle(&link_id) and
send on that.
Where to go next
simple_send.rsandecho_server.rs— the minimal node + event loop.chat.rs— both receive variants, node status (active_link_count,pending_link_count).relay_daemon.rs— a transport node andtransport_stats().link_test.rs/link_integration_test.rs— these drop down toleviculum-core'sLinkdirectly against a Pythonrnsd, useful if you want to see the wire-level handshake rather than the high-level handle API.
The full method list of every type is in the generated rustdoc. Build it with:
cargo doc --no-deps --open -p leviculum-std
For verified signatures of the types used above, continue to the Rust API reference.
Rust API Reference
This chapter is a reference for the key entry points and core value types of the
Leviculum Rust API, organized by type. Each signature carries a file:line
citation to the source as of this writing. It is deliberately not exhaustive:
the complete per-type method list is generated rustdoc (see
Full rustdoc at the end). Use this chapter to orient, then
rustdoc for the long tail.
The hands-on introduction is the tutorial; the layer overview is Choosing a layer.
All leviculum-std types are re-exported from the crate root
(leviculum-std/src/lib.rs:67-94), so use leviculum_std::{NodeEvent, LinkHandle, …} works without naming submodules.
leviculum-std (std / tokio)
Reticulum
The configuration-driven entry point, wrapping a ReticulumNode. Defined at
leviculum-std/src/reticulum.rs:13. Use this when your node is described by a
Config (an INI file or a programmatic Config); use ReticulumNodeBuilder
when you assemble interfaces in code.
| Signature | Purpose |
|---|---|
fn new() -> Result<Self> — reticulum.rs:22 | Build from the default config path, or defaults if absent |
fn with_config(config: Config) -> Result<Self> — reticulum.rs:37 | Build from an explicit Config |
fn with_config_daemon(config: Config) -> Result<Self> — reticulum.rs:60 | Like with_config but with no application event channel (daemon mode); take_event_receiver() then returns None |
async fn start(&mut self) -> Result<()> — reticulum.rs:76 | Spawn the event loop |
async fn stop(&mut self) -> Result<()> — reticulum.rs:82 | Stop and persist |
fn is_running(&self) -> bool — reticulum.rs:89 | Whether the loop is running |
fn config(&self) -> &Config — reticulum.rs:94 | Borrow the active config |
fn take_event_receiver(&mut self) -> Option<EventReceiver> — reticulum.rs:154 | Take the event stream, once |
ReticulumNodeBuilder
The programmatic builder. Defined at leviculum-std/src/driver/builder.rs:39;
re-exported as leviculum_std::ReticulumNodeBuilder. Each setter consumes and
returns self.
| Signature | Purpose |
|---|---|
fn new() -> Self — builder.rs:96 | Builder with defaults |
fn identity(self, identity: Identity) -> Self — builder.rs:219 | Pin an explicit identity (else one is generated/persisted) |
fn add_tcp_client(self, addr: SocketAddr) -> Self — builder.rs:270 | Connect outward to a Reticulum node |
fn add_tcp_server(self, addr: SocketAddr) -> Self — builder.rs:323 | Listen for inbound connections |
fn add_udp_interface(self, listen: SocketAddr, forward: SocketAddr) -> Self — builder.rs:384 | One datagram per packet |
fn add_rnode_interface(self, port: String, frequency: u64, bandwidth: u32, spreading_factor: u8, coding_rate: u8, tx_power: i8) -> Self — builder.rs:424 | LoRa interface; required radio settings |
fn add_serial_interface(self, port: String, speed: u32, databits: u8, parity: String, stopbits: u8) -> Self — builder.rs:483 | KISS over raw serial |
fn add_auto_interface(self) -> Self — builder.rs:610 | IPv6 multicast LAN discovery |
fn enable_transport(self, enabled: bool) -> Self — builder.rs:664 | Act as a relay/forwarder |
fn config(self, config: Config) -> Self — builder.rs:244 | Use a pre-loaded Config |
fn config_file(self, path: PathBuf) -> Self — builder.rs:254 | Load an INI config file |
fn storage_path(self, path: PathBuf) -> Self — builder.rs:262 | Identity / known-destinations / ratchet store dir |
fn connect_to_shared_instance(self, name: impl Into<String>) -> Self — builder.rs:716 | Attach to a running lnsd/rnsd instead of bringing up own interfaces |
fn without_events(self) -> Self — builder.rs:211 | Daemon mode: no application event channel |
async fn build(self) -> Result<ReticulumNode, Error> — builder.rs:1039 | Build the node (not yet running) |
fn build_sync(self) -> Result<ReticulumNode, Error> — builder.rs:803 | Same as build, outside an async context |
ReticulumNode
The running node. Defined at leviculum-std/src/driver/mod.rs:1319; re-exported
as leviculum_std::ReticulumNode. Selected methods:
| Signature | Purpose |
|---|---|
async fn start(&mut self) -> Result<(), Error> — driver/mod.rs:1558 | Spawn the event loop, bring interfaces up |
async fn stop(&mut self) -> Result<(), Error> — driver/mod.rs:2084 | Stop and flush |
fn is_running(&self) -> bool — driver/mod.rs:2416 | Loop state |
fn register_destination(&self, destination: Destination) — driver/mod.rs:2424 | Make a local destination reachable (consumes it) |
async fn announce_destination(&self, dest_hash: &DestinationHash, app_data: Option<&[u8]>) -> … — driver/mod.rs:3404 | Announce a registered destination |
async fn connect(&self, dest_hash: &DestinationHash, dest_signing_key: &[u8; 32]) -> Result<LinkHandle, Error> — driver/mod.rs:2591 | Open a link; returns a pending handle |
fn link_handle(&self, link_id: &LinkId) -> LinkHandle — driver/mod.rs:2862 | Writable handle for an already-established inbound link |
fn packet_sender(&self, dest_hash: &DestinationHash) -> PacketSender — driver/mod.rs:3764 | Single-packet send handle |
async fn send_single_packet(&self, …) -> … — driver/mod.rs:3712 | Send one unreliable datagram |
fn take_event_receiver(&mut self) -> Option<EventReceiver> — driver/mod.rs:2878 | Take the event stream, once |
fn identity_hash(&self) -> [u8; 16] — driver/mod.rs:2712 | The node's own identity hash |
fn has_path(&self, dest_hash: &DestinationHash) -> bool — driver/mod.rs:3085 | Whether a path is known |
fn hops_to(&self, dest_hash: &DestinationHash) -> Option<u8> — driver/mod.rs:3196 | Hop count to a destination |
async fn request_path(&self, dest_hash: &DestinationHash) -> Result<(), Error> — driver/mod.rs:3108 | Send a PATH_REQUEST; result arrives as PathFound |
fn get_identity(&self, dest_hash: &DestinationHash) -> Option<Identity> — driver/mod.rs:3093 | Identity learned from an announce (its signing key feeds connect) |
fn transport_stats(&self) -> TransportStats — driver/mod.rs:3311 | rnstatus-style counters |
fn is_transport_enabled(&self) -> bool — driver/mod.rs:3779 | Relay mode flag |
The stable, curated facade leviculum_std::api — NodeBuilder (leviculum-std/src/api/mod.rs:60),
Node (leviculum-std/src/api/mod.rs:238) — re-projects this surface with core internals
hidden; it is what leviculum-ffi wraps. Notable facade-only helpers:
api::generate_identity() (api/mod.rs:35), api::version() (api/mod.rs:42),
api::version_string() (api/mod.rs:51), and Node::connect_with_key
(api/mod.rs:456) / Node::accept_link (api/mod.rs:472).
LinkHandle
Send-only async handle for a link. Defined at leviculum-std/src/driver/stream.rs:47;
re-exported as leviculum_std::LinkHandle. Incoming data is delivered via
NodeEvent, not on the handle.
| Signature | Purpose |
|---|---|
fn link_id(&self) -> &LinkId — stream.rs:74 | The link's id |
fn is_closed(&self) -> bool — stream.rs:79 | Handle state |
async fn try_send(&self, data: &[u8]) -> Result<(), Error> — stream.rs:88 | Non-blocking send; surfaces Busy / PacingDelay |
async fn send(&self, data: &[u8]) -> Result<(), Error> — stream.rs:110 | Send, retrying pacing/busy internally |
async fn close(&mut self) -> Result<(), Error> — stream.rs:147 | Graceful close (sends LINKCLOSE) |
PacketSender
Send-only async handle for single packets, the single-packet analog of
LinkHandle. Defined at leviculum-std/src/driver/sender.rs:44; re-exported as
leviculum_std::PacketSender.
| Signature | Purpose |
|---|---|
fn dest_hash(&self) -> &DestinationHash — sender.rs:70 | The target destination |
async fn send(&self, data: &[u8]) -> Result<[u8; TRUNCATED_HASHBYTES], Error> — sender.rs:92 | Send one unreliable packet; returns the truncated packet hash. A path must already be known |
EventReceiver and NodeEvent
EventReceiver is the merged event stream, defined at
leviculum-std/src/driver/mod.rs:379. It internally fronts a lossless control
plane and a droppable data plane (Codeberg #71), draining control first.
| Signature | Purpose |
|---|---|
async fn recv(&mut self) -> Option<NodeEvent> — driver/mod.rs:436 | Next event, control plane prioritized; None once shut down. Cancel-safe |
fn try_recv(&mut self) -> Result<NodeEvent, TryRecvError> — driver/mod.rs:476 | Non-blocking receive |
NodeEvent is the event enum, defined in core at
leviculum-core/src/node/event.rs:45 and re-exported as
leviculum_std::NodeEvent. It is #[non_exhaustive], so always include a
catch-all arm. The variants most applications match (field names verbatim from
source):
| Variant | Fields | Source |
|---|---|---|
AnnounceReceived | announce: ReceivedAnnounce, interface_index: usize | event.rs:24 |
PathFound | destination_hash: DestinationHash, hops: u8, interface_index: usize | event.rs:32 |
PacketReceived | destination: DestinationHash, data: Vec<u8>, interface_index: usize | event.rs:59 |
PacketDeliveryConfirmed | packet_hash: [u8; TRUNCATED_HASHBYTES] | event.rs:69 |
LinkEstablished | link_id: LinkId, is_initiator: bool | event.rs:84 |
MessageReceived | link_id: LinkId, msgtype: u16, sequence: u16, data: Vec<u8> | event.rs:95 |
LinkDataReceived | link_id: LinkId, data: Vec<u8> | event.rs:111 |
LinkClosed | (see source) | event.rs:155 |
MessageReceived is the channel-multiplexed (sequenced) receive path;
LinkDataReceived is the raw-link-packet path. The full variant list (resources,
requests/responses, identify, stale/recovered, control-plane overflow) is in
event.rs and in rustdoc.
Config
Configuration, defined at leviculum-std/src/config.rs:12; re-exported as
leviculum_std::Config. pub reticulum: ReticulumConfig (config.rs:15) and
pub interfaces: HashMap<String, InterfaceConfig> (config.rs:17).
| Signature | Purpose |
|---|---|
fn load<P: AsRef<Path>>(path: P) -> Result<Self> — config.rs:901 | Load an INI config (the rnsd/lnsd format) |
fn default_config_dir() -> PathBuf — config.rs:1063 | Default config directory |
fn default_config_path() -> PathBuf — config.rs:1072 | Default config file path |
leviculum-core (no_std, sans-IO)
The core is the no_std engine the std layer drives. You use these types directly
only when building on leviculum-core — see Embedded development.
All are re-exported from leviculum-core/src/lib.rs:161-186.
NodeCore<R, C, S>
The sans-IO protocol engine, generic over an RNG R: CryptoRngCore, a clock
C: Clock, and storage S: Storage. Defined at
leviculum-core/src/node/mod.rs:409. It never performs I/O; every method that
can produce output returns a TickOutput the
caller must dispatch.
| Signature | Purpose |
|---|---|
fn new(identity: Identity, config: TransportConfig, proof_strategy: ProofStrategy, max_incoming_resource_size: usize, rng: R, clock: C, storage: S) -> Self — node/mod.rs:568 | Construct directly |
fn register_destination(&mut self, dest: Destination) — node/mod.rs:629 | Register a local destination |
fn announce_destination(&mut self, dest_hash: &DestinationHash, app_data: Option<&[u8]>) -> Result<TickOutput, AnnounceError> — node/mod.rs:876 | Build and queue an announce |
fn send_single_packet(&mut self, dest_hash: &DestinationHash, data: &[u8]) -> Result<([u8; TRUNCATED_HASHBYTES], TickOutput), SendError> — node/mod.rs:1078 | Build an unreliable data packet |
fn connect(&mut self, dest_hash: DestinationHash, dest_signing_key: &[u8; 32]) -> (LinkId, bool, TickOutput) — node/link_management.rs:252 | Build a link request |
fn send_on_link(&mut self, link_id: &LinkId, data: &[u8]) -> Result<TickOutput, SendError> — node/link_management.rs:693 | Send on an established link |
fn close_link(&mut self, link_id: &LinkId) -> TickOutput — node/link_management.rs:600 | Close a link |
fn handle_packet(&mut self, iface: InterfaceId, data: &[u8]) -> TickOutput — node/mod.rs:2280 | Feed received bytes from an interface |
fn handle_timeout(&mut self) -> TickOutput — node/mod.rs:2508 | Run periodic maintenance (call at the next deadline) |
fn next_deadline(&self) -> Option<u64> — node/mod.rs:2539 | Earliest timer deadline (ms); when to call handle_timeout |
A node is more often built with NodeCoreBuilder (node/builder.rs:40), whose
fn build<R, Clk, S>(self, rng: R, clock: Clk, storage: S) -> NodeCore<R, Clk, S>
(node/builder.rs:247) supplies the platform triple. Setters include
identity, proof_strategy, and enable_transport.
Core TickOutput and Action
TickOutput is what every core method returns. Defined at
leviculum-core/src/transport.rs:148. It is #[must_use] — dropping it silently
loses outbound packets and events.
| Field | Type | Source |
|---|---|---|
actions | Vec<Action> — I/O for the driver to execute | transport.rs:150 |
events | Vec<NodeEvent> — application-visible events | transport.rs:152 |
next_deadline_ms | Option<u64> — when to next call handle_timeout | transport.rs:155 |
Action is the I/O the driver performs, defined at leviculum-core/src/transport.rs:123:
| Variant | Fields | Source |
|---|---|---|
SendPacket | iface: InterfaceId, data: Vec<u8>, peer: Option<[u8; 16]> | transport.rs:125 |
Broadcast | data: Vec<u8>, exclude_iface: Option<InterfaceId> | transport.rs:132 |
The helper dispatch_actions(interfaces: &mut [&mut dyn Interface], actions: Vec<Action>, ifac_configs: &BTreeMap<usize, IfacConfig>) -> DispatchResult
(transport.rs:221) routes Actions to interfaces with broadcast-exclusion and
IFAC wrapping handled in core, so every driver gets it for free.
Value types
Identity — a key pair or public-only identity. Defined at
leviculum-core/src/identity.rs. Re-exported as leviculum_std::Identity.
| Signature | Purpose |
|---|---|
fn generate<R: CryptoRngCore>(rng: &mut R) -> Self — identity.rs:81 | New random identity |
fn from_public_key_bytes(bytes: &[u8]) -> Result<Self, IdentityError> — identity.rs:125 | Public-only identity |
fn from_private_key_bytes(bytes: &[u8]) -> Result<Self, IdentityError> — identity.rs:139 | From the raw 64-byte private key (Python-compatible) |
fn hash(&self) -> &[u8; IDENTITY_HASHBYTES] — identity.rs:167 | The 16-byte identity hash |
fn public_key_bytes(&self) -> [u8; IDENTITY_KEY_SIZE] — identity.rs:172 | 64 bytes: X25519 [0..32], Ed25519 [32..64] |
fn has_private_keys(&self) -> bool — identity.rs:197 | Whether it can sign/decrypt |
fn sign(&self, message: &[u8]) -> Result<…, IdentityError> — identity.rs:202 | Ed25519 sign |
fn verify(&self, message: &[u8], signature: &[u8]) -> Result<bool, IdentityError> — identity.rs:214 | Ed25519 verify |
Destination — a local or remote destination. Defined at
leviculum-core/src/destination.rs. Re-exported as leviculum_std::Destination.
| Signature | Purpose |
|---|---|
fn new(identity: Option<Identity>, direction: Direction, dest_type: DestinationType, app_name: &str, aspects: &[&str]) -> Result<Self, DestinationError> — destination.rs:349 | Construct a destination |
fn hash(&self) -> &DestinationHash — destination.rs:476 | Its 16-byte hash |
fn direction(&self) -> Direction — destination.rs:497 | In / Out |
DestinationHash — a 16-byte address (newtype, destination.rs:158):
fn new(bytes: [u8; TRUNCATED_HASHBYTES]) -> Self (destination.rs:162),
fn as_bytes(&self) -> &[u8; TRUNCATED_HASHBYTES] (destination.rs:167),
fn into_bytes(self) -> [u8; TRUNCATED_HASHBYTES] (destination.rs:172).
Direction (destination.rs:160) and DestinationType (destination.rs:132)
are the small enums passed to Destination::new. Packets are constructed
internally (leviculum_core::packet::Packet); applications work with
destinations and links, not raw packets.
Platform traits
The three abstractions you implement to run the core on a platform. Defined in
leviculum-core/src/traits.rs and re-exported from lib.rs:141.
| Trait | Required methods (selected) | Source |
|---|---|---|
Clock | fn now_ms(&self) -> u64 | traits.rs:421 |
Storage | key-value persistence: has_packet_hash, get_path/set_path, link/announce tables, identities, ratchets (large trait) | traits.rs:196 |
Interface | id, name, mtu, is_online, fn try_send(&mut self, data: &[u8]) -> Result<(), InterfaceError> | traits.rs:280 |
Provided Storage implementations: NoStorage (traits.rs:920, zero-sized
no-op for stubs and stateless devices), MemoryStorage
(leviculum-core/src/memory_storage.rs, BTreeMap-backed with caps), and
EmbeddedStorage (leviculum-core/src/embedded_storage.rs:91, heapless-backed
for flash-constrained targets; fn new() -> Self at embedded_storage.rs:344).
leviculum-std adds a file-backed Storage with Python-compatible on-disk
formats.
Full rustdoc
This chapter covers the load-bearing surface; the exhaustive method list is the generated rustdoc. Build and open it with:
cargo doc --no-deps --open -p leviculum-std # std/tokio layer
cargo doc --no-deps --open -p leviculum-core # no_std core
Embedded Development: Building on leviculum-core
This chapter is for building on leviculum-core directly: embedded firmware, a
custom async runtime, a simulator, or any host program that wants byte-level
control without tokio. The core is no_std (it uses alloc, but not the
standard library) and sans-IO — it performs no I/O and owns no runtime. You feed
it bytes, it hands back a TickOutput,
and you do the I/O. The worked reference is the nRF52 firmware in leviculum-nrf,
cited throughout.
If you can use std and tokio, prefer leviculum-std and read the
tutorial instead — leviculum-std is itself a driver
for this same core. See Choosing a layer for the trade-off.
The dependency
Depend on leviculum-core with default features off. It is not on crates.io, so
use a path or git:
[dependencies]
# Adjust the path to wherever, and under whatever name, you cloned the repository
leviculum-core = { path = "../leviculum/leviculum-core", default-features = false }
No std, no tokio. You bring your own executor (Embassy, RTIC, a bare loop) and
your own allocator. The reference firmware leviculum-nrf targets
thumbv7em-none-eabihf and uses Embassy.
The sans-IO contract
The core is a state machine with exactly three ways in, and one way out. The way
out is always a TickOutput (leviculum-core/src/transport.rs:531), carrying
actions to perform, events that occurred, and next_deadline_ms, the time at
which you must next tick the timer. It is #[must_use]: dropping it loses
outbound packets and events.
received bytes ─► handle_packet(iface, data) ─┐
timer expired ─► handle_timeout() ├─► TickOutput { actions, events, next_deadline_ms }
│
└─► you: dispatch actions, react to events,
schedule the next timeout
The three entry points (signatures in the reference):
handle_packet(iface, data)—leviculum-core/src/node/mod.rs:1034. Feed one received frame, tagged with theInterfaceIdit arrived on.handle_timeout()—leviculum-core/src/node/mod.rs:1362. Run periodic maintenance (path expiry, announce rebroadcasts, keepalives, retransmissions). Call it at or beforenext_deadline.next_deadline()(leviculum-core/src/node/mod.rs:2539). The earliest timer deadline in milliseconds, orNoneif no timer is pending. Sleep until this, or until a packet arrives, whichever comes first.
App-initiated operations (register_destination, announce_destination,
connect, send_on_link, send_single_packet) likewise return a TickOutput
you must dispatch.
The driver loop
The shape is: compute the next deadline, wait for whichever of "a packet on any
interface" or "the deadline" happens first, call the matching entry point,
dispatch the resulting actions. This is exactly the leviculum-nrf T114 main
loop: the deadline comes from next_deadline
(leviculum-nrf/src/bin/t114.rs:763-768) and the wait from Embassy's select4
(leviculum-nrf/src/bin/t114.rs:821-840). The board selects over nine event
sources; the loop below narrows that to the three interfaces (serial, LoRa,
BLE) and the timer:
#![allow(unused)] fn main() { // Adapted from the T114 main loop cited above. loop { let deadline = node .next_deadline() .map(Instant::from_millis) .unwrap_or(Instant::MAX); match select4( serial.incoming_rx.receive(), lora_channels.incoming_rx.receive(), ble_channels.incoming_rx.receive(), Timer::at(deadline), ) .await { Either4::First(data) => { let output = node.handle_packet(InterfaceId(0), &data); let mut ifaces: [&mut dyn Interface; 3] = [&mut serial_iface, &mut lora_iface, &mut ble_iface]; dispatch_actions(&mut ifaces, output.actions, &ifac_configs); } Either4::Second(data) => { let output = node.handle_packet(InterfaceId(1), &data); let mut ifaces: [&mut dyn Interface; 3] = [&mut serial_iface, &mut lora_iface, &mut ble_iface]; dispatch_actions(&mut ifaces, output.actions, &ifac_configs); } Either4::Third(data) => { let output = node.handle_packet(InterfaceId(2), &data); let mut ifaces: [&mut dyn Interface; 3] = [&mut serial_iface, &mut lora_iface, &mut ble_iface]; dispatch_actions(&mut ifaces, output.actions, &ifac_configs); } Either4::Fourth(()) => { let output = node.handle_timeout(); let mut ifaces: [&mut dyn Interface; 3] = [&mut serial_iface, &mut lora_iface, &mut ble_iface]; dispatch_actions(&mut ifaces, output.actions, &ifac_configs); } } } }
Three things to notice:
next_deadline()drives the timer. MapNoneto "wait forever" (Instant::MAX) so you wake only when something actually needs doing — there is no fixed tick rate.InterfaceId(n)tags the source. The index you pass tohandle_packetmust match the interface's ownid(), so the core's routing tables and broadcast-exclusion stay consistent.dispatch_actionsdoes the routing. Rather than matching on eachActionyourself, hand the wholeactionsvec plus your&mut dyn Interfaceslice todispatch_actions(leviculum-core/src/transport.rs:695). Broadcast exclusion, interface selection, and IFAC wrapping live in core, so every driver gets them for free. Bind what it returns: theDispatchResultis#[must_use]because dropping it discards the retries the core asked for, the interface errors it saw, and any action it could not route.
This loop ignores output.events because a leaf firmware node has no application
logic to react to them; a richer firmware would drain output.events here the
way the std event loop
drains the EventReceiver.
Building the node
NodeCoreBuilder (leviculum-core/src/node/builder.rs:40) takes the platform
triple — RNG, Clock, and
Storage — in its build call. The T114
firmware builds its node the same way (NodeCoreBuilder,
leviculum-nrf/src/bin/t114.rs:202-220):
#![allow(unused)] fn main() { // Adapted from the T114 builder cited above. let mut builder = NodeCoreBuilder::new() .enable_transport(true) .max_incoming_resource_size(8 * 1024) .respond_to_probes(true); if let Ok(Some(identity)) = id_store.load() { builder = builder.identity(identity); } let mut node = builder.build_boxed(rng, EmbassyClock, EmbeddedStorage::new()); }
build consumes the builder and the platform triple and returns the
NodeCore<R, C, S>. A driver with an inline-storage S must not call it:
use build_boxed (leviculum-core/src/node/builder.rs:329), which allocates
first and configures through the box. Box::new(builder.build(..)) holds a
full-size NodeCore as a by-value local on the way into the box, and with
EmbeddedStorage that is upwards of 40 KB twice over — it gave the T114 a
94 KB main frame on a 128 KB stack and corrupted SoftDevice RAM on the
deeper paths.
Implementing the platform traits
Three traits decouple the core from your hardware. Their signatures are in the reference; here is what to supply.
Clock
A monotonic millisecond clock. The whole trait is one required method. The
nRF52 implementation wraps Embassy's timer
(leviculum-nrf/src/clock.rs):
#![allow(unused)] fn main() { use leviculum_core::traits::Clock; pub struct EmbassyClock; impl Clock for EmbassyClock { fn now_ms(&self) -> u64 { embassy_time::Instant::now().as_millis() } } }
now_secs, has_elapsed, and deadline have default implementations
(leviculum-core/src/traits.rs:471-483); you only provide now_ms. It must be
monotonic.
Interface
The send side of an interface — id, name, mtu, is_online, and the
non-blocking try_send (leviculum-core/src/traits.rs:280). The receive side is
deliberately not in the trait: receiving is platform-specific (an interrupt, a
DMA buffer, an Embassy channel), and you feed received bytes into the core via
handle_packet yourself. try_send returns InterfaceError::BufferFull
(non-fatal, packet dropped — Reticulum is best-effort) or
InterfaceError::Disconnected. A minimal always-ready interface looks like the
test impl in traits.rs:
#![allow(unused)] fn main() { use leviculum_core::traits::{Interface, InterfaceError}; use leviculum_core::transport::InterfaceId; struct MyRadio { /* hardware handle */ } impl Interface for MyRadio { fn id(&self) -> InterfaceId { InterfaceId(1) } fn name(&self) -> &str { "my-radio" } fn mtu(&self) -> usize { 500 } fn is_online(&self) -> bool { true } fn try_send(&mut self, data: &[u8]) -> Result<(), InterfaceError> { // hand `data` to the radio's TX queue, non-blocking Ok(()) } } }
A constrained medium (LoRa) overrides next_slot_ms
(leviculum-core/src/traits.rs:365) to report the next airtime-fit time, so the
core schedules retries against capacity without knowing any radio physics — the
interface-isolation rule. For a fast link the default
("always ready") is correct.
An interface that carries several point-to-point links behind one
InterfaceId (BLE) overrides try_send_to_peer
(leviculum-core/src/traits.rs:317): the core passes the 16-byte
identity of the peer it addressed the packet at — the same value the
interface reports on peer-up/peer-lost — or None for a broadcast.
The interface maps that identity to its own link(s); the core never
sees a link. The default implementation drops the hint, which is why
a single-peer interface implements nothing.
Storage
Key-value persistence for the path table, link table, announce caches,
identities, ratchets, and dedup hashes (leviculum-core/src/traits.rs:500). It
is a large trait; you do not write it from scratch:
NoStorage(leviculum-core/src/traits.rs:920) — zero-sized, every lookup returns nothing. Use it for a stateless node or a smoke test.EmbeddedStorage(leviculum-core/src/embedded_storage.rs:91,EmbeddedStorage::new()at:344) —heapless-backed, fixed-capacity, the production choice for flash-constrained devices. This is what the nRF52 firmware uses.MemoryStorage(leviculum-core/src/memory_storage.rs) — BTreeMap-backed with configurable caps, for hosts with more memory.
Implement Storage yourself only to add real persistence (e.g. to flash); the
file-backed implementation in leviculum-std is the worked example of wrapping
MemoryStorage with disk writes.
Summary
- Depend on
leviculum-corewithdefault-features = false. Nostd, no tokio,allocrequired. - Drive the loop:
next_deadline()→ wait for a packet or the deadline →handle_packet/handle_timeout→dispatch_actions(output.actions). - Implement
Clock(trivial),Interface(send side only — you feed RX in viahandle_packet), and pick aStorage(NoStorage/EmbeddedStorage/MemoryStorage, or your own). - The full worked driver is
leviculum-nrf/src/bin/t114.rs; the full method list iscargo doc --no-deps -p leviculum-core(see the reference).
C API: Overview and Concepts
Leviculum ships a C API so an application can use the Reticulum network stack
the way it uses any normal Unix C library: a clean header, opaque handle
types, integer error codes, and composition with the application's own event
loop. This chapter explains the model that the How-To and the
API Reference build on. For the design rationale behind these
choices, see the design-of-record at docs/leviculum-api-design.md.
Every symbol is prefixed lev_ (functions) or LEV_ (constants). The header
is leviculum.h, the library is libleviculum.so.
Installing and linking
Once the development package is installed, building against Leviculum is the usual two lines:
#include <leviculum.h>
cc app.c $(pkg-config --cflags --libs leviculum)
The pkg-config call expands to -lleviculum plus the include and library
paths. To build from source and install the header, the shared object (with its
SONAME and dev symlinks), the static archive, and the pkg-config file:
make -C leviculum-ffi install PREFIX=/usr/local # builds, then installs
To link Leviculum statically while glibc stays dynamic, pass --static so
pkg-config adds the archive's system dependencies, and force the archive:
cc app.c $(pkg-config --cflags leviculum) \
-l:libleviculum.a $(pkg-config --static --libs-only-l leviculum | sed 's/-lleviculum//')
See Installation for the full toolchain setup. The
install is verified end to end (dynamic and static, x86_64 and aarch64) by
scripts/verify-packaging.sh.
Opaque handles
Every complex object is an opaque pointer. The application never sees a struct
layout, so the ABI stays stable across versions. Each handle has a constructor
and a matching free function; _free(NULL) is always a no-op.
| Handle | Represents | Created by | Freed by |
|---|---|---|---|
leviculum_t | a node (runtime, engine, event bridge) | lev_builder_build | lev_free |
lev_builder_t | node configuration before build | lev_builder_new | lev_builder_free |
lev_identity_t | a key pair or public-only identity | lev_identity_generate, lev_identity_from_*, lev_identity_load_file, lev_link_remote_identity | lev_identity_free |
lev_destination_t | a local destination | lev_destination_new | lev_destination_free |
lev_link_t | one link to a peer | lev_connect, lev_connect_with_key, lev_accept_link | lev_link_free |
lev_event_t | one drained event | lev_next_event, lev_wait_event | lev_event_free |
Two builders are single-use: lev_builder_build and lev_register_destination
take the contents of their handle and leave an empty shell that the caller
still frees.
Addresses are not handles. A destination hash, a link id, and an identity hash
are each a fixed 16-byte value (LEV_ADDR_LEN); a resource hash is 32 bytes
(LEV_RESOURCE_HASH_LEN). They cross the boundary as plain uint8_t arrays.
Error handling
Functions that can fail return int: 0 (LEV_OK) on success, a negative
LEV_ERR_* code on failure. Constructors that return a handle return NULL
on failure. Two helpers turn a code into text:
lev_strerror(code)returns a static, never-freed string for the code.lev_last_error()returns a thread-local string with the specific detail of the most recent failing call on the calling thread (which argument, which address). It is owned by the library and must not be freed.
int rc = lev_start(node);
if (rc != LEV_OK) {
fprintf(stderr, "start failed: %s (%s)\n", lev_strerror(rc), lev_last_error());
}
The full code list is in the reference.
Buffers: the read(2) convention
Every function that returns bytes into a caller buffer uses the same shape,
modelled on read(2):
int lev_identity_hash(const lev_identity_t *id,
uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
- The caller owns
bufand passes its capacitycapplus anout_len. - On success the library writes the bytes and sets
*out_lento the count. - If
capis too small (orbufisNULL), nothing is written,*out_lenis set to the required size, and the call returnsLEV_ERR_BUFFER_TOO_SMALL. Passingbuf == NULLis therefore a valid size query.
uint8_t hash[LEV_ADDR_LEN];
uintptr_t len = sizeof(hash);
if (lev_identity_hash(id, hash, sizeof(hash), &len) == LEV_OK) {
/* `hash` holds `len` bytes */
}
The library never hands C a raw pointer to free: all freeing goes through a
typed lev_*_free, which removes C-free-versus-Rust-dealloc mistakes.
Out-parameters for returned values
Status and value are never multiplexed into one return. A call that both can
fail and produces a value returns the int status and writes the value
through an out-parameter:
uint8_t packet_hash[LEV_ADDR_LEN];
int rc = lev_send_datagram(node, dest, data, len, packet_hash, 3000);
lev_link_t *link = NULL;
int rc2 = lev_connect(node, dest, 5000, &link); /* link in *out */
Strings and bytes
- Opaque byte payloads (keys, hashes, datagram data, link data, resource data) are always a pointer plus a length, never NUL-terminated, and may contain zero bytes.
- Human-readable strings the library consumes (storage path, destination
app_name, requestpath) are NUL-terminated UTF-8 C strings. - A destination's aspects are passed as a
const char *const *array plus a count. - Library-returned static strings (
lev_strerror,lev_last_error,lev_version_string) are NUL-terminated and must not be freed.
The event model: a pollable fd
Everything inbound (received announces, link data, request and response
arrivals, resource progress and completion) reaches the application as events.
A node exposes a single readable file descriptor that the application adds to
its own poll/epoll/select loop:
struct pollfd p = { .fd = lev_event_fd(node), .events = POLLIN };
poll(&p, 1, -1);
lev_event_t *ev;
while (lev_next_event(node, &ev) == LEV_OK && ev) {
switch (lev_event_type(ev)) {
case LEV_EVENT_ANNOUNCE_RECEIVED: /* ... */ break;
case LEV_EVENT_LINK_DATA: /* ... */ break;
}
lev_event_free(ev);
}
The fd is level-triggered: it is readable exactly while the queue is
non-empty. After each wake, drain with lev_next_event until it yields NULL.
lev_wait_event(node, &ev, timeout_ms) is a convenience that blocks for the
next event without your own loop. The event side is single-consumer: do not
call the two drain functions concurrently for the same node.
The fd is owned by the library and closed by lev_free. The shutdown order is
mandatory: stop reacting to the fd, remove it from your loop, then call
lev_free. Polling the fd after lev_free is a use-after-close.
Event handles are fully self-owned (payloads are copied out at dequeue), so an
event stays valid until lev_event_free regardless of later calls. Read its
fields with the typed accessors (lev_event_link_id, lev_event_data,
lev_event_request_id, lev_event_resource_hash, and so on); an accessor that
does not apply to the event type returns LEV_ERR_INVALID_ARG.
Threading and blocking
The tokio runtime is created and owned inside the node and never exposed.
- A
leviculum_tis thread-safe: its methods may be called concurrently from multiple threads. - The event side is single-consumer (above).
- Every potentially-blocking call takes a
timeout_ms(negative means wait forever); on expiry it returnsLEV_ERR_TIMEOUT. The link data path istry_send-first:lev_link_try_sendnever blocks and returnsLEV_ERR_AGAINunder backpressure, whilelev_link_sendretries up to its deadline. lev_free,lev_stop, and the other blocking calls must run on a plain OS thread, never on a worker thread of another runtime (for example a host async runtime); doing so would panic the embeddedblock_on.- The log callback may fire on any internal worker thread and must not call
back into any
lev_*function.
No panic crosses the boundary
Every exported function wraps its body so that an internal Rust panic is caught
and converted to LEV_ERR_PANIC (or NULL for a constructor) instead of
unwinding into C, which would be undefined behaviour. After a caught panic the
affected node should be freed and not reused.
One-time setup and logging
lev_init() performs idempotent process setup (logging subscriber and panic
hook). It is optional, since other entry points run it lazily, but call it
explicitly to configure logging before the first node. Logging is silent by
default; raise it with lev_log_set_level(LEV_LOG_INFO) and route records with
lev_log_set_callback, or leave the default which writes to stderr.
With these conventions in hand, the Tutorial builds a complete,
useful program (levcat, a pipe over the mesh) step by step, the
How-To is the recipe book for every flow, and the
API Reference documents every function.
Tutorial: Build levcat, a Pipe over Reticulum
This tutorial builds one small, complete, genuinely useful program from scratch:
levcat, a bidirectional pipe over the mesh, the netcat of Reticulum. Run it
in two terminals and it is a chat. Feed it a file and it is a file transfer
(levcat connect ... < file). Drop it in a shell pipeline and it carries bytes
between machines, over TCP, over LoRa, over anything Reticulum reaches.
Along the way you learn the patterns every Leviculum C program needs: bring up a
node, announce and discover a destination, open a link, and — the heart of it —
run the node's event loop inside your own poll(2) loop, alongside your own
file descriptors. After this you can write your own Leviculum program.
This builds on the Overview (opaque handles, the read(2) buffer
convention, the pollable event fd); skim it first. The How-To is the
recipe companion, and the API Reference has every signature. The
finished program is leviculum-ffi/examples/c/levcat.c, compiled and tested in
the repo, so the code here is real, not pseudo-code.
What we build
Two roles share one transport and one steady-state loop:
levcat listen <storage> <bind host:port> # the listening end
levcat connect <storage> <peer host:port> <dest-hex> # the dialing end
The listener registers a destination, announces it, and prints its address. The connector is handed that address, finds a path to it, and opens a link. Once linked, both ends pump stdin to the link and link data to stdout.
1. Skeleton
Start with argument parsing, one-time init, and a signal flag so Ctrl-C exits
cleanly. lev_init() is optional (other calls run it lazily) but it is the
place to set up logging before anything else.
#include <errno.h>
#include <poll.h>
#include <signal.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include "leviculum.h"
static volatile sig_atomic_t stop = 0;
static void on_signal(int s) { (void)s; stop = 1; }
int main(int argc, char **argv) {
signal(SIGINT, on_signal);
signal(SIGTERM, on_signal);
lev_init();
if (argc == 4 && strcmp(argv[1], "listen") == 0)
return run_listen(argv[2], argv[3]);
if (argc == 5 && strcmp(argv[1], "connect") == 0)
return run_connect(argv[2], argv[3], argv[4]);
fprintf(stderr, "usage:\n %s listen <storage> <bind host:port>\n"
" %s connect <storage> <peer host:port> <dest-hex>\n",
argv[0], argv[0]);
return 2;
}
2. Bring up a node
A node is built, then started. The builder is an opaque handle you configure and then consume; see reference: node lifecycle and builder. Both roles share this helper, differing only in the interface they add — a TCP server for the listener, a TCP client for the connector.
static leviculum_t *build_start(const char *storage,
void (*configure)(lev_builder_t *, const char *),
const char *arg, lev_identity_t *id) {
lev_builder_t *b = lev_builder_new();
if (!b) return NULL;
if (lev_builder_storage_path(b, storage) != LEV_OK) { lev_builder_free(b); return NULL; }
if (id) lev_builder_identity(b, id);
configure(b, arg); /* add the interface */
leviculum_t *node = lev_builder_build(b);
lev_builder_free(b); /* build empties the builder; still free it */
if (!node) return NULL;
if (lev_start(node) != LEV_OK) { lev_free(node); return NULL; }
return node;
}
static void cfg_server(lev_builder_t *b, const char *addr) { lev_builder_add_tcp_server(b, addr); }
static void cfg_client(lev_builder_t *b, const char *addr) { lev_builder_add_tcp_client(b, addr); }
3. The listening end
The listener owns a destination: an address other nodes can reach. We generate
an identity, register an incoming single destination under the app name
levcat with the aspect pipe, and read back its 16-byte hash. Then we
announce it so the network learns a path, and print the address — to stderr,
because stdout is the data pipe and must stay clean.
lev_identity_t *id = lev_identity_generate();
leviculum_t *node = build_start(storage, cfg_server, bind_addr, id);
const char *aspects[] = {"pipe"};
lev_destination_t *dest =
lev_destination_new(id, LEV_DIRECTION_IN, LEV_DEST_SINGLE, "levcat", aspects, 1);
uint8_t dh[LEV_ADDR_LEN];
size_t dhl = sizeof(dh);
lev_destination_hash(dest, dh, sizeof(dh), &dhl);
lev_register_destination(node, dest);
lev_destination_free(dest);
char hexhash[2 * LEV_ADDR_LEN + 1];
hex(dh, LEV_ADDR_LEN, hexhash); /* lev_hex_encode wrapper */
fprintf(stderr, "destination: %s\n", hexhash);
Now wait for someone to dial in. We re-announce in a loop (so a peer that starts
later still discovers us) and watch for a LEV_EVENT_LINK_REQUEST. When it
arrives we read the link id from the event and accept it. See
How-To: announcing and discovering.
lev_link_t *link = NULL;
while (!stop && !link) {
lev_announce(node, dh, NULL, 0, 2000);
for (int i = 0; i < 3 && !link; i++) {
lev_event_t *ev = NULL;
if (lev_wait_event(node, &ev, 200) != LEV_OK || !ev) continue;
if (lev_event_type(ev) == LEV_EVENT_LINK_REQUEST) {
uint8_t lid[LEV_ADDR_LEN];
size_t l = sizeof(lid);
lev_event_link_id(ev, lid, sizeof(lid), &l);
lev_accept_link(node, lid, 5000, &link);
}
lev_event_free(ev);
}
}
One subtlety: accepting a link does not make it immediately usable for
sending. The responder's link becomes active only after the initiator's RTT
exchange, signalled by the responder's own LEV_EVENT_LINK_ESTABLISHED. Sending
before that returns LEV_ERR_SEND ("link not active"). So we wait for it before
pumping, writing through any data that arrives meanwhile so none is lost:
int active = 0;
while (link && !stop && !active) {
lev_event_t *ev = NULL;
if (lev_wait_event(node, &ev, 200) != LEV_OK || !ev) continue;
int t = lev_event_type(ev);
if (t == LEV_EVENT_LINK_ESTABLISHED) active = 1;
else if (t == LEV_EVENT_LINK_MESSAGE) emit_message(ev); /* don't drop early data */
else if (t == LEV_EVENT_LINK_CLOSED) stop = 1;
lev_event_free(ev);
}
if (active) pump(node, link);
lev_wait_event is the blocking drain we use during setup; the steady-state
loop (pump, below) uses the pollable fd instead. Every event must be freed
with lev_event_free.
4. The dialing end
The connector is given the listener's address as hex. Decode it to 16 bytes,
bring up a node with a TCP client interface, and wait for a path: the listener's
announce arrives over the link and installs one. lev_request_path nudges it
along; lev_has_path reports when it is ready. See
reference: paths, connect, and links.
uint8_t dest[LEV_ADDR_LEN];
size_t dlen = sizeof(dest);
lev_hex_decode((const uint8_t *)dest_hex, strlen(dest_hex), dest, sizeof(dest), &dlen);
leviculum_t *node = build_start(storage, cfg_client, peer_addr, NULL);
lev_request_path(node, dest, 2000);
for (int i = 0; i < 300 && lev_has_path(node, dest) != 1; i++) {
lev_event_t *ev = NULL;
if (lev_wait_event(node, &ev, 200) == LEV_OK && ev) lev_event_free(ev);
}
With a path in hand, open the link. lev_connect returns as soon as the request
is sent — the link is usable only after the handshake, which the engine signals
with LEV_EVENT_LINK_ESTABLISHED. Wait for it, then start pumping.
lev_link_t *link = NULL;
lev_connect(node, dest, 8000, &link);
int established = 0;
for (int i = 0; i < 100 && !established; i++) {
lev_event_t *ev = NULL;
if (lev_wait_event(node, &ev, 200) == LEV_OK && ev) {
if (lev_event_type(ev) == LEV_EVENT_LINK_ESTABLISHED) established = 1;
lev_event_free(ev);
}
}
if (established) pump(node, link);
5. The pump loop — the heart of it
Both ends now have a link and run the same loop. This is the pattern that makes
Leviculum composable: the node exposes a single readable file descriptor
(lev_event_fd), so you put it in your own poll(2) set right next to your own
fds. Here that is stdin. One poll waits for either: local input to send, or
a network event to receive. See the event model.
static void pump(leviculum_t *node, lev_link_t *link) {
struct pollfd fds[2];
fds[0].fd = STDIN_FILENO; fds[0].events = POLLIN;
fds[1].fd = lev_event_fd(node); fds[1].events = POLLIN;
while (!stop) {
int r = poll(fds, 2, 1000);
if (r < 0) { if (errno == EINTR) continue; break; }
if (fds[0].revents & POLLIN) { /* local input -> link */
uint8_t buf[CHUNK];
ssize_t n = read(STDIN_FILENO, buf, sizeof(buf));
if (n <= 0) { /* EOF: flush, then close — see below */ return; }
if (lev_link_send(link, buf, (size_t)n, 5000) != LEV_OK) return;
}
if ((fds[1].revents & POLLIN) && drain_to_stdout(node)) return; /* link -> stdout */
}
}
Two details:
- Chunking. A link's reliable channel has a maximum message size, so we read
stdin in
#define CHUNK 256-byte pieces that fit on any interface.lev_link_sendis the reliable, sequenced send; it blocks up to its deadline, retrying backpressure internally. (The non-blocking sibling islev_link_try_send, which returnsLEV_ERR_AGAINinstead of waiting — see How-To: links and exchanging data.) - Receiving.
lev_link_sendon one side surfaces as aLEV_EVENT_LINK_MESSAGEon the other. We drain every pending event and copy each message's bytes to stdout, using the read(2)-style accessor (size query, then fill):
static void emit_message(lev_event_t *ev) {
size_t need = 0;
lev_event_data(ev, NULL, 0, &need); /* size query */
uint8_t *d = malloc(need ? need : 1);
size_t got = need;
lev_event_data(ev, d, need, &got); /* fill */
fwrite(d, 1, got, stdout);
fflush(stdout);
free(d);
}
static int drain_to_stdout(leviculum_t *node) {
int closed = 0;
lev_event_t *ev = NULL;
while (lev_next_event(node, &ev) == LEV_OK && ev) {
int t = lev_event_type(ev);
if (t == LEV_EVENT_LINK_MESSAGE) {
emit_message(ev);
} else if (t == LEV_EVENT_LINK_CLOSED) {
closed = 1;
}
lev_event_free(ev);
}
return closed;
}
The fd is level-triggered: it stays readable while the queue is non-empty,
so after each wake we drain with lev_next_event until it yields NULL. The
event side is single-consumer — never drain the same node from two threads.
6. Closing cleanly
When local input ends (Ctrl-D, or the end of a piped file), we are done sending.
Give the reliable channel a moment to deliver the last bytes — draining any
final inbound meanwhile — then close our end. The peer sees LEV_EVENT_LINK_CLOSED
and exits too, so a cat file | levcat connect ... terminates instead of
hanging. This is the if (n <= 0) branch of the pump:
for (int g = 0; g < 10 && !stop; g++) {
struct pollfd ef = {fds[1].fd, POLLIN, 0};
if (poll(&ef, 1, 100) > 0 && drain_to_stdout(node)) break;
}
lev_close_link(link, 2000);
return;
(A production tool would do a real half-close so the reverse direction can keep flowing; we keep it minimal.) Then tear the node down in order — the link first, then the node:
lev_link_free(link); /* NULL-safe; closes the link if still open */
lev_stop(node); /* persists state, stops the loop */
lev_free(node); /* releases the runtime and the event fd */
lev_identity_free(id); /* listener only */
The shutdown order is mandatory: stop reacting to the event fd before
lev_free, which closes it.
7. Build and run it
Compile against the installed library with pkg-config (installing and linking):
cc levcat.c $(pkg-config --cflags --libs leviculum) -o levcat
Open two terminals. In the first, listen:
$ ./levcat listen /tmp/levcat-a 127.0.0.1:4242
destination: a1b2c3d4e5f6... # printed on stderr
In the second, connect with that address, then type on either side:
$ ./levcat connect /tmp/levcat-b 127.0.0.1:4242 a1b2c3d4e5f6...
hello from the other terminal
That is a chat. It is also a pipe — send a file and end with Ctrl-D, or:
# receiver
./levcat listen /tmp/levcat-a 127.0.0.1:4242 > received.tar
# sender
tar c somedir | ./levcat connect /tmp/levcat-b 127.0.0.1:4242 <dest-hex>
Nothing here is TCP-specific. Swap lev_builder_add_tcp_* for
lev_builder_add_rnode (or load a config file) and the same program pipes bytes
across a LoRa mesh.
Where to go next
You have used the core of the API: node setup, announce and discovery, links, and the event loop. From here:
- Bulk data with progress, compression, and metadata:
How-To: resource transfer (and the full
leviculum-ffi/examples/c/lncp.cfile-copy tool). - Lightweight RPC: How-To: request and response.
- Best-effort single packets: How-To: datagrams.
- Inspecting the stack: How-To: diagnostics.
- Every signature and constant: the API Reference.
Full program: leviculum-ffi/examples/c/levcat.c.
C API: How-To, Building Applications
This chapter shows how the functions combine into working programs. It assumes
the model from the Overview: opaque handles, integer error
codes, read(2) buffers, and the pollable event fd. Each recipe gives the
functions involved and a focused snippet; the complete, compiling programs are
the acceptance tests under leviculum-ffi/examples/c/, named per recipe. For a
single program built end to end from these pieces, see the
Tutorial.
Error checks are abbreviated in the snippets for readability. In real code,
check every int return against LEV_OK and report lev_last_error() (see
Errors and logging).
A minimal node
Build a node, attach an interface, start it, and shut it down. The builder is
single-use: lev_builder_build consumes its configuration and you still free
the empty handle.
#include <leviculum.h>
#include <stdio.h>
int main(void) {
lev_init();
printf("leviculum %s\n", lev_version_string());
lev_builder_t *b = lev_builder_new();
lev_builder_storage_path(b, "/var/lib/myapp/reticulum");
lev_builder_add_tcp_client(b, "127.0.0.1:4242"); /* a Reticulum hub */
leviculum_t *node = lev_builder_build(b);
lev_builder_free(b); /* build emptied it */
if (!node) {
fprintf(stderr, "build failed: %s\n", lev_last_error());
return 1;
}
if (lev_start(node) != LEV_OK) {
fprintf(stderr, "start failed: %s\n", lev_last_error());
lev_free(node);
return 1;
}
/* ... run the application ... */
lev_stop(node);
lev_free(node); /* lev_free also stops a still-running node */
return 0;
}
Interfaces are added on the builder: lev_builder_add_tcp_client,
lev_builder_add_tcp_server, lev_builder_add_udp,
lev_builder_add_auto_interface. Use lev_builder_identity to pin a specific
identity (otherwise one is generated), and lev_builder_enable_transport(b, 1)
to act as a relay.
Full program: leviculum-ffi/examples/c/phase_a.c.
Running the event loop
Everything inbound arrives as events. Add lev_event_fd(node) to your loop,
and on each wake drain with lev_next_event until it yields NULL.
#include <poll.h>
int fd = lev_event_fd(node);
for (;;) {
struct pollfd p = { .fd = fd, .events = POLLIN };
poll(&p, 1, -1);
lev_event_t *ev;
while (lev_next_event(node, &ev) == LEV_OK && ev) {
switch (lev_event_type(ev)) {
case LEV_EVENT_ANNOUNCE_RECEIVED: on_announce(ev); break;
case LEV_EVENT_LINK_REQUEST: on_link_request(ev); break;
case LEV_EVENT_LINK_MESSAGE: on_link_message(ev); break;
/* ... */
}
lev_event_free(ev);
}
}
If you do not want to own a loop, block for one event at a time:
lev_event_t *ev = NULL;
if (lev_wait_event(node, &ev, 1000) == LEV_OK && ev) { /* up to 1s */
/* handle ev */
lev_event_free(ev);
}
Rules: the fd is level-triggered (readable while the queue is non-empty); the
two drain functions are single-consumer (one thread at a time); and the
shutdown order is stop reacting to the fd, then lev_free. Reading an event's
fields uses the typed accessors shown in the recipes below.
Running as or with a daemon
A node need not bring up its own interfaces in code. Three builder calls cover the daemon use cases.
Load an RNS-style config (the same INI rnsd/lnsd read), so interfaces,
transport, and the shared instance come from a file an operator edits. This is
also how a C node reaches LoRa without programmatic radio setup, the config
names an RNodeInterface or SerialInterface and the stack brings it up.
lev_builder_t *b = lev_builder_new();
lev_builder_config_file(b, "/etc/leviculum/config");
leviculum_t *node = lev_builder_build(b);
lev_builder_free(b);
lev_start(node); /* now a daemon: run the event loop until signalled */
Offer a shared instance, so other local programs and the Reticulum tools
(rnstatus, rnpath, rnprobe) attach to this one stack instead of each
opening the radio:
lev_builder_share_instance(b, "leviculum"); /* opens the IPC + RPC endpoint */
Or attach to a running daemon as a client, the way rncp/rnx do, instead of
bringing up interfaces of your own:
lev_builder_connect_shared_instance(b, "leviculum");
A NULL path or name returns LEV_ERR_INVALID_ARG. The daemon.c example is
a worked acceptance program for all three calls. The lncp.c file-copy tool
has both styles: its recv/send modes bring up their own interface, while its
recv-shared/send-shared modes attach to a running lnsd by instance name,
so several tools share one daemon's radio.
Radio interfaces (LoRa and serial)
For off-grid mesh, add an RNode (LoRa) or a raw serial interface programmatically, no config file needed:
lev_builder_t *b = lev_builder_new();
/* RNode: device, frequency Hz, bandwidth Hz, spreading factor, coding rate,
* tx power dBm. */
lev_builder_add_rnode(b, "/dev/ttyUSB0", 867200000, 125000, 8, 5, 0);
/* Serial: device, speed, data bits, parity ("N"/"E"/"O"), stop bits. */
lev_builder_add_serial(b, "/dev/ttyACM0", 115200, 8, "N", 1);
The device is opened at lev_start, so a wrong path surfaces there, not at the
setter (which only rejects a NULL path with LEV_ERR_INVALID_ARG). A serial
port is raw KISS with no handshake; an RNode performs the RNode detect and
config handshake on start. For the optional RNode knobs (airtime limits, flow
control, buffer size), load a config file instead. The radio.c example brings
a node up over a serial interface.
Identities
An identity is a key pair. Generate one, persist it, and reload it next run. The on-disk format is the raw 64-byte private key, compatible with Python Reticulum.
lev_identity_t *id;
id = lev_identity_load_file("/var/lib/myapp/identity");
if (!id) { /* first run: make one */
id = lev_identity_generate();
lev_identity_save_file(id, "/var/lib/myapp/identity");
}
uint8_t hash[LEV_ADDR_LEN];
uintptr_t len = sizeof(hash);
lev_identity_hash(id, hash, sizeof(hash), &len); /* the 16-byte address */
A combined key is 64 bytes (LEV_IDENTITY_KEY_LEN): the X25519 encryption key
in bytes 0..32 and the Ed25519 signing key in bytes 32..64. Applications
rarely split it by hand, because lev_connect resolves the signing key for
you (see below). Use lev_builder_identity(b, id) to give a node a fixed
identity, and lev_identity_free(id) when done.
An identity also signs, verifies, encrypts, and decrypts directly, for crypto tooling and signed application data, interoperable with Python peers (Ed25519 for signatures, X25519+AES for encryption):
uint8_t sig[64];
uintptr_t n = sizeof(sig);
lev_identity_sign(id, msg, msg_len, sig, sizeof(sig), &n);
int ok = lev_identity_verify(id, msg, msg_len, sig, n); /* 1 valid, 0 not */
/* Encrypt to a peer's public-only identity; only its private key recovers it. */
uint8_t ct[512];
uintptr_t ctl = sizeof(ct);
lev_identity_encrypt(peer, msg, msg_len, ct, sizeof(ct), &ctl);
Sign, encrypt, and decrypt write read(2) style (a NULL buffer queries the
length); signing and decryption need the private key and return
LEV_ERR_CRYPTO on a public-only identity, while verify needs only the public
key.
Full programs: leviculum-ffi/examples/c/phase_a.c and crypto.c.
Announcing and discovering
To be reachable, a node registers an incoming destination and announces it.
Other nodes learn the destination (its address, identity, and a path) from the
announce, which arrives as LEV_EVENT_ANNOUNCE_RECEIVED.
Announcing side:
const char *aspects[] = { "inbox" };
lev_destination_t *dest = lev_destination_new(
id, LEV_DIRECTION_IN, LEV_DEST_SINGLE, "myapp", aspects, 1);
uint8_t dh[LEV_ADDR_LEN];
uintptr_t dhl = sizeof(dh);
lev_destination_hash(dest, dh, sizeof(dh), &dhl); /* read before registering */
lev_register_destination(node, dest); /* consumes dest */
lev_destination_free(dest); /* free the empty shell */
lev_announce(node, dh, NULL, 0, 2000); /* optional app_data, here none */
For forward secrecy, call lev_destination_enable_ratchets(dest, now_ms) on an
inbound destination before registering it (now_ms is the current time in
milliseconds); peers, including Python ones, then encrypt to a rotating ratchet
key. lev_destination_ratchet_public(node, dh, ...) reads the current key. See
leviculum-ffi/examples/c/ratchet.c.
For delivery proofs, call lev_destination_set_proof_strategy(dest, strategy)
before registering. LEV_PROOF_ALL auto-proves every received packet (Python's
PROVE_ALL). LEV_PROOF_APP raises a LEV_EVENT_PACKET_PROOF_REQUESTED event
whose data is the 32-byte packet hash; the app decides and calls
lev_send_proof(node, dest_hash, packet_hash, timeout_ms). See
leviculum-ffi/examples/c/proof.c.
Receiving side, in the event loop:
case LEV_EVENT_ANNOUNCE_RECEIVED: {
uint8_t peer[LEV_ADDR_LEN];
uintptr_t n = sizeof(peer);
lev_event_dest_hash(ev, peer, sizeof(peer), &n); /* who announced */
/* optional payload via lev_event_data(ev, ...) */
break;
}
After processing the announce, the receiver has a path and the announcer's
cached identity, so lev_has_path(node, peer) returns 1 and lev_connect will
work.
Full program: leviculum-ffi/examples/c/phase_b.c.
Links and exchanging data
A link is an encrypted session to a destination. lev_connect resolves the
peer's signing key from the identity cached by an announce, so you pass only
the destination hash:
lev_link_t *link = NULL;
int rc = lev_connect(node, peer, 5000, &link);
if (rc == LEV_ERR_UNKNOWN_DEST) { /* no announce seen yet */ }
else if (rc == LEV_ERR_NO_PATH) { lev_request_path(node, peer, 3000); }
else if (rc == LEV_OK) { /* link is pending; wait for established */ }
The connecting node watches for LEV_EVENT_LINK_ESTABLISHED; the destination
node watches for LEV_EVENT_LINK_REQUEST and accepts it:
case LEV_EVENT_LINK_REQUEST: {
uint8_t lid[LEV_ADDR_LEN];
uintptr_t n = sizeof(lid);
lev_event_link_id(ev, lid, sizeof(lid), &n);
lev_link_t *accepted = NULL;
lev_accept_link(node, lid, 5000, &accepted);
/* keep `accepted` to send on this link */
break;
}
Send and receive link data. lev_link_send blocks up to its deadline,
retrying backpressure; lev_link_try_send returns LEV_ERR_AGAIN instead of
blocking. It sends over the link's reliable channel (sequenced and
retransmitted, the same RawBytesMessage Python peers use), so the peer sees a
LEV_EVENT_LINK_MESSAGE, with a message type and a sequence number:
lev_link_send(link, (const uint8_t *)"hello", 5, 5000);
case LEV_EVENT_LINK_MESSAGE: {
uint8_t buf[512];
uintptr_t n = sizeof(buf);
uint16_t msgtype = 0, sequence = 0;
if (lev_event_data(ev, buf, sizeof(buf), &n) == LEV_OK) {
lev_event_msgtype(ev, &msgtype); /* 0 for raw bytes */
lev_event_sequence(ev, &sequence); /* per-channel send order */
/* `n` bytes received */
}
break;
}
A peer that sends a raw, unsequenced link packet instead of using the channel
(for example Python's RNS.Packet(link, data).send()) arrives as the
lower-level LEV_EVENT_LINK_DATA, which carries only link_id and data.
Close with lev_close_link(link, 2000) and release with lev_link_free(link)
(which also closes an open link). A LEV_EVENT_LINK_CLOSED event reports a
link that drops for any reason.
Full program: leviculum-ffi/examples/c/phase_c.c.
Proving identity on a link
By default a link is anonymous. Either side can prove an identity to the peer;
the peer is notified with LEV_EVENT_LINK_IDENTIFIED and can read it back.
/* prover */
lev_link_identify(node, my_link_id, my_identity, 3000);
/* peer, in the event loop */
case LEV_EVENT_LINK_IDENTIFIED: {
lev_identity_t *who = lev_link_remote_identity(node, my_link_id);
if (who) {
uint8_t h[LEV_ADDR_LEN];
uintptr_t n = sizeof(h);
lev_identity_hash(who, h, sizeof(h), &n); /* the peer's address */
lev_identity_free(who);
}
break;
}
The 16-byte identity hash is also the payload of the
LEV_EVENT_LINK_IDENTIFIED event (lev_event_data).
Full program: leviculum-ffi/examples/c/phase_c.c.
Request and response
For a request/response service, the responder registers a handler for a path on its destination; the requester sends a request over a link. Request and response payloads are msgpack-encoded values.
Responder:
lev_register_request_handler(node, dh, "/echo",
LEV_REQUEST_POLICY_ALLOW_ALL, NULL, 0);
case LEV_EVENT_REQUEST_RECEIVED: {
uint8_t link_id[LEV_ADDR_LEN], req_id[LEV_ADDR_LEN], data[512];
uintptr_t a = sizeof(link_id), b = sizeof(req_id), c = sizeof(data);
lev_event_link_id(ev, link_id, sizeof(link_id), &a);
lev_event_request_id(ev, req_id, sizeof(req_id), &b);
lev_event_data(ev, data, sizeof(data), &c); /* the request body */
/* path is available via lev_event_path(ev, ...) */
lev_send_response(node, link_id, req_id, data, c, 3000); /* echo it */
break;
}
Requester (over an established link, whose id comes from lev_link_id):
uint8_t req[] = { 0xA4, 'p','i','n','g' }; /* msgpack "ping" */
uint8_t request_id[LEV_ADDR_LEN];
lev_send_request(node, link_id, "/echo", req, sizeof(req), 5000, request_id);
case LEV_EVENT_RESPONSE_RECEIVED: {
uint8_t rid[LEV_ADDR_LEN], body[512];
uintptr_t a = sizeof(rid), b = sizeof(body);
lev_event_request_id(ev, rid, sizeof(rid), &a); /* match request_id */
lev_event_data(ev, body, sizeof(body), &b);
break;
}
A request that gets no reply within its deadline surfaces as
LEV_EVENT_REQUEST_TIMEOUT. To restrict callers, use
LEV_REQUEST_POLICY_ALLOW_LIST with an array of n_ids 16-byte identity
hashes.
Full program: leviculum-ffi/examples/c/phase_d.c.
Datagrams
A datagram is a single, unreliable packet to a destination. A path must
already be known. Delivery is best-effort: a LEV_EVENT_PACKET_RECEIVED on the
other side, and a delivery confirmation only if the destination returns a
proof.
uint8_t packet_hash[LEV_ADDR_LEN];
int rc = lev_send_datagram(node, dest_hash, (const uint8_t *)"hi", 2,
packet_hash, 3000);
if (rc == LEV_ERR_NO_PATH) { lev_request_path(node, dest_hash, 3000); }
/* receiver */
case LEV_EVENT_PACKET_RECEIVED: {
uint8_t buf[256];
uintptr_t n = sizeof(buf);
lev_event_data(ev, buf, sizeof(buf), &n);
break;
}
Full program: leviculum-ffi/examples/c/phase_d.c.
Resource transfer
A resource carries bulk data (a file) over a link, in segments, with optional compression and msgpack metadata. The receiver chooses a strategy: accept all, reject all, or be asked per transfer.
Receiver sets a strategy on the link, then accepts when advertised:
lev_set_resource_strategy(node, link_id, LEV_RESOURCE_ACCEPT_APP);
case LEV_EVENT_RESOURCE_ADVERTISED:
lev_accept_resource(node, link_id, 3000); /* or lev_reject_resource */
break;
case LEV_EVENT_RESOURCE_COMPLETED: {
uint8_t buf[65536];
uintptr_t n = sizeof(buf);
lev_event_data(ev, buf, sizeof(buf), &n); /* the assembled data */
/* metadata via lev_event_metadata(ev, ...) if present */
break;
}
Sender initiates the transfer and tracks progress:
uint8_t resource_hash[LEV_RESOURCE_HASH_LEN];
lev_send_resource(node, link_id, file_data, file_len,
NULL, 0, /* optional msgpack metadata */
1, /* auto-compress */
resource_hash, 5000);
case LEV_EVENT_RESOURCE_PROGRESS: {
double frac;
lev_event_progress(ev, &frac); /* 0.0 .. 1.0 */
break;
}
LEV_EVENT_RESOURCE_COMPLETED carries the data only on the receiver;
LEV_EVENT_RESOURCE_FAILED reports a transfer that did not finish.
lev_send_resource returns once the transfer is initiated: the receiver then
pulls the parts part by part. A sending program must keep its node alive and
running the event loop until the transfer is done, the receiver must keep the
link it accepted open (freeing a link closes it), and the receiver applies its
resource strategy on the link before the resource arrives. Exiting the sender
right after the call returns aborts an in-flight transfer.
Full programs: leviculum-ffi/examples/c/phase_e.c, and
leviculum-ffi/examples/c/lncp.c, a complete two-process file-copy tool
(lncp send / lncp recv) that exercises the whole stack end to end.
Errors and logging
Every fallible call returns int. Pair the code with the thread-local detail:
int rc = lev_connect(node, peer, 5000, &link);
if (rc != LEV_OK) {
fprintf(stderr, "connect: %s (%s)\n", lev_strerror(rc), lev_last_error());
}
LEV_ERR_AGAIN (from lev_link_try_send) and LEV_ERR_TIMEOUT are normal,
retryable conditions, not hard failures. Logging from the stack itself is off
by default; turn it on and route it to your own sink:
static void log_sink(int level, const char *msg, void *user) {
(void)user;
fprintf(stderr, "[lev %d] %s\n", level, msg);
}
lev_init();
lev_log_set_callback(log_sink, NULL);
lev_log_set_level(LEV_LOG_INFO);
The callback may run on an internal thread and must not call back into any
lev_* function. For hex display of an address, use lev_hex_encode and
lev_hex_decode.
Diagnostics
For an rnstatus-style view, lev_transport_stats reads the transport
counters and the path-table size:
uint64_t sent, received, dropped, paths;
lev_transport_stats(node, &sent, &received, NULL, NULL, &dropped, &paths);
Any out-pointer may be NULL to skip it.
For an rnpath-style listing, take a frozen snapshot of the path table, read
its entries by index, and free it:
lev_path_table_t *table = lev_path_table_snapshot(node);
for (int i = 0; i < lev_path_table_count(table); i++) {
uint8_t dest[LEV_ADDR_LEN];
uint8_t hops;
lev_path_table_entry(table, i, dest, &hops, NULL, NULL, NULL, NULL);
/* `dest` reachable in `hops` hops */
}
lev_path_table_free(table);
The snapshot is a point-in-time copy, so reads never race a changing table.
Interface stats work the same way (lev_interface_stats_snapshot /
_count / _name / _entry / _free), giving each interface's name, online
status, and byte counters for an rnstatus-style interface listing. See
leviculum-ffi/examples/c/stats.c.
Putting it together
A typical application wires these into one loop: it loads or generates an
identity, builds and starts a node with an interface, registers and announces a
destination, then runs the event loop, reacting to announces by connecting,
to link requests by accepting, and to data, request, and resource events by
serving the application. The phase_b.c through phase_e.c programs are
complete two-node demonstrations of exactly these flows, runnable via
cargo test-ffi.
C API: Reference
Every public function and constant of leviculum.h, grouped by area. The
header leviculum-ffi/leviculum.h is generated from the Rust source and is the
canonical statement of the exact prototypes; this reference is kept in sync
with it and adds semantics. For the model behind these signatures (handles,
the read(2) buffer convention, out-parameters, the event fd, threading), read
the Overview first.
Conventions used below:
- Functions return
int:LEV_OK(0) on success, a negativeLEV_ERR_*on failure; constructors returnNULLon failure. After any failure,lev_last_error()holds a detail string. - A
(uint8_t *buf, uintptr_t cap, uintptr_t *out_len)triple is read(2) style:buf == NULLor too-smallcapreturnsLEV_ERR_BUFFER_TOO_SMALLwith*out_lenset to the required size. timeout_msis a deadline in milliseconds; negative means wait forever; on expiry the call returnsLEV_ERR_TIMEOUT.
Opaque types
| Type | Created by | Freed by | Thread-safety |
|---|---|---|---|
leviculum_t | lev_builder_build | lev_free | thread-safe; events single-consumer |
lev_builder_t | lev_builder_new | lev_builder_free | one thread |
lev_identity_t | lev_identity_generate, lev_identity_from_private_key, lev_identity_from_public_key, lev_identity_load_file, lev_link_remote_identity | lev_identity_free | one thread |
lev_destination_t | lev_destination_new | lev_destination_free | one thread |
lev_link_t | lev_connect, lev_connect_with_key, lev_accept_link | lev_link_free | sends thread-safe; do not close/free concurrently with other calls on the same link |
lev_event_t | lev_next_event, lev_wait_event | lev_event_free | one thread |
lev_path_table_t | lev_path_table_snapshot | lev_path_table_free | one thread |
lev_interface_stats_t | lev_interface_stats_snapshot | lev_interface_stats_free | one thread |
Initialisation and logging
int lev_init(void);
int lev_log_set_level(int level);
int lev_log_set_callback(lev_log_callback cb, void *user);
typedef void (*lev_log_callback)(int level, const char *message, void *user);
lev_initruns one-time process setup (logging subscriber, panic hook) once, through an internalOnce. Idempotent and thread-safe. Optional: other entry points run it lazily.lev_log_set_levelsets the global verbosity to one of theLEV_LOG_*constants. ReturnsLEV_ERR_INVALID_ARGfor an out-of-range level.lev_log_set_callbackroutes log records tocb(withuserpassed back unchanged), or restores the stderr default whencbisNULL. The callback may run on any internal worker thread, receives a NUL-terminated message valid only for the call, and must not call back into anylev_*function.
| Constant | Value | Meaning |
|---|---|---|
LEV_LOG_OFF | 0 | no logging (default) |
LEV_LOG_ERROR | 1 | errors only |
LEV_LOG_WARN | 2 | warnings and above |
LEV_LOG_INFO | 3 | info and above |
LEV_LOG_DEBUG | 4 | debug and above |
LEV_LOG_TRACE | 5 | everything |
Versioning
const char *lev_version_string(void);
uint32_t lev_version_number(void);
lev_version_stringreturns the workspace version fromCargo.tomlas a static, never-freed string.lev_version_numberpacks it as(major << 16) | (minor << 8) | patch, a host-byte-order integer for in-process comparison only.
Errors
const char *lev_strerror(int code);
const char *lev_last_error(void);
lev_strerrorreturns a static message for aLEV_ERR_*code; safe any time, never freed.lev_last_errorreturns the thread-local detail string for the most recent failing call on the calling thread, orNULLif there is none. Owned by the library, valid until the next failing call on the same thread, never freed.
Error codes
| Constant | Value | Meaning |
|---|---|---|
LEV_OK | 0 | success |
LEV_ERR_NULL_PTR | -1 | a required pointer argument was NULL |
LEV_ERR_INVALID_ARG | -2 | malformed argument (bad length, unparseable string) |
LEV_ERR_BUFFER_TOO_SMALL | -3 | caller buffer too small; *out_len holds the needed size |
LEV_ERR_NOT_RUNNING | -4 | the node event loop is not running |
LEV_ERR_IO | -5 | an I/O or storage error |
LEV_ERR_CONFIG | -6 | a configuration error |
LEV_ERR_CRYPTO | -7 | a cryptographic operation failed |
LEV_ERR_NO_PATH | -8 | no path to the destination is known |
LEV_ERR_LINK | -9 | a link operation failed (closed, inactive, handshake) |
LEV_ERR_SEND | -10 | a send failed (no route, payload too large) |
LEV_ERR_RESOURCE | -11 | a resource transfer operation failed |
LEV_ERR_REQUEST | -12 | a request or response operation failed |
LEV_ERR_TIMEOUT | -13 | the operation timed out |
LEV_ERR_AGAIN | -14 | non-fatal backpressure; retry later |
LEV_ERR_UNKNOWN_DEST | -15 | no cached identity for the destination |
LEV_ERR_NO_HANDLER | -16 | nothing was registered under the name the call asked to remove |
LEV_ERR_PANIC | -127 | a panic was caught at the FFI boundary |
Identity
struct lev_identity_t *lev_identity_generate(void);
struct lev_identity_t *lev_identity_from_private_key(const uint8_t *key, uintptr_t len);
struct lev_identity_t *lev_identity_from_public_key(const uint8_t *key, uintptr_t len);
struct lev_identity_t *lev_identity_load_file(const char *path);
int lev_identity_save_file(const struct lev_identity_t *id, const char *path);
void lev_identity_free(struct lev_identity_t *id);
int lev_identity_hash(const struct lev_identity_t *id, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_identity_public_key(const struct lev_identity_t *id, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_identity_private_key(const struct lev_identity_t *id, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_identity_has_private_keys(const struct lev_identity_t *id);
int lev_identity_sign(const struct lev_identity_t *id, const uint8_t *msg, uintptr_t msg_len, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_identity_verify(const struct lev_identity_t *id, const uint8_t *msg, uintptr_t msg_len, const uint8_t *sig, uintptr_t sig_len);
int lev_identity_encrypt(const struct lev_identity_t *id, const uint8_t *plaintext, uintptr_t plaintext_len, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_identity_decrypt(const struct lev_identity_t *id, const uint8_t *ciphertext, uintptr_t ciphertext_len, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
lev_identity_generatemakes a new random full identity;NULLon failure.lev_identity_from_private_key/lev_identity_from_public_keybuild an identity from a 64-byte combined key (lenmust equalLEV_IDENTITY_KEY_LEN); the public-key variant yields a public-only identity.NULLon failure.lev_identity_load_filereads the raw 64-byte private key file (the Python-Reticulum format);NULLif missing, wrong size, or invalid.lev_identity_save_filewrites the private key topathatomically;LEV_ERR_CRYPTOif the identity is public-only.lev_identity_hashwrites the 16-byte identity hash;_public_keyand_private_keywrite the 64-byte combined keys (_private_keyreturnsLEV_ERR_CRYPTOfor a public-only identity). All read(2) style.lev_identity_has_private_keysreturns 1 for a full identity, 0 otherwise (and 0 onNULL).lev_identity_signwrites the 64-byte Ed25519 signature ofmsgread(2) style;LEV_ERR_CRYPTOif the identity is public-only.lev_identity_verifyreturns 1 if the signature is valid, 0 if not (including a wrong-length signature), and a negativeLEV_ERR_*on a NULL argument; it needs only the public key.lev_identity_encryptencryptsplaintextto the identity's public key (the Reticulum X25519+AES scheme) and writes the ciphertext read(2) style; encryption is randomised, so a length query and the real call differ in bytes but not length.lev_identity_decryptreverses it with the private key and returnsLEV_ERR_CRYPTOfor a public-only identity or a ciphertext that fails to authenticate.
Every returned lev_identity_t is owned by the caller and freed with
lev_identity_free.
| Constant | Value | Meaning |
|---|---|---|
LEV_ADDR_LEN | 16 | destination, link, and identity hash length |
LEV_IDENTITY_KEY_LEN | 64 | combined key length (public or private) |
LEV_X25519_KEY_LEN | 32 | encryption half, bytes 0..32 |
LEV_SIGNING_KEY_LEN | 32 | Ed25519 signing half, bytes 32..64 |
Node lifecycle and builder
struct lev_builder_t *lev_builder_new(void);
void lev_builder_free(struct lev_builder_t *b);
int lev_builder_identity(struct lev_builder_t *b, const struct lev_identity_t *id);
int lev_builder_storage_path(struct lev_builder_t *b, const char *path);
int lev_builder_add_tcp_client(struct lev_builder_t *b, const char *addr);
int lev_builder_add_tcp_server(struct lev_builder_t *b, const char *addr);
int lev_builder_add_udp(struct lev_builder_t *b, const char *listen_addr, const char *forward_addr);
int lev_builder_add_auto_interface(struct lev_builder_t *b);
int lev_builder_add_rnode(struct lev_builder_t *b, const char *port, uint64_t frequency, uint32_t bandwidth, uint8_t spreading_factor, uint8_t coding_rate, int8_t tx_power);
int lev_builder_add_serial(struct lev_builder_t *b, const char *port, uint32_t speed, uint8_t databits, const char *parity, uint8_t stopbits);
int lev_builder_enable_transport(struct lev_builder_t *b, int enabled);
int lev_builder_event_capacity(struct lev_builder_t *b, uintptr_t control_cap, uintptr_t data_cap);
int lev_builder_link_keepalive(struct lev_builder_t *b, uint64_t secs);
int lev_builder_config_file(struct lev_builder_t *b, const char *path);
int lev_builder_share_instance(struct lev_builder_t *b, const char *name);
int lev_builder_connect_shared_instance(struct lev_builder_t *b, const char *name);
struct leviculum_t *lev_builder_build(struct lev_builder_t *b);
int lev_start(struct leviculum_t *node);
int lev_stop(struct leviculum_t *node);
int lev_is_running(const struct leviculum_t *node);
int lev_identity_hash_self(const struct leviculum_t *node, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
void lev_free(struct leviculum_t *node);
lev_builder_newallocates a builder;lev_builder_freereleases it (lev_builder_free(NULL)is a no-op).- The setters configure the node:
_identitypins a (cloned) identity,_storage_pathsets the state directory, the_add_*calls add interfaces (TCP addresses arehost:port),_enable_transporttoggles relay mode, and_event_capacitysets the event-queue sizes (control and data planes; a 0 keeps the current default). Each setter returnsLEV_ERR_INVALID_ARGif the builder was already consumed. lev_builder_link_keepaliveoverrides the link keepalive interval, in seconds, for every link the node creates; the stale-link timeout scales with it (a link goes stale after twice the keepalive). The value is clamped to the protocol minimum. The default (no call) derives the interval from the link RTT. Useful for slow links, and for makingLEV_EVENT_LINK_STALEobservable quickly. The same knob is thekeepalive_intervalkey in a config file.lev_builder_add_rnodeadds a LoRa interface over an RNode:portis the serial device, then the required radio settings (frequencyandbandwidthin Hz,spreading_factor,coding_ratedenominator,tx_powerin dBm).lev_builder_add_serialadds a raw KISS serial interface:port,speed,databits,parity("N","E", or"O"),stopbits. Both returnLEV_ERR_INVALID_ARGon a NULL device path (or NULL parity). For the optional RNode tuning (airtime limits, flow control, buffer size) use a config file. The device is opened atlev_start, not when the setter runs.lev_builder_config_fileloads an RNS-style INI config (the same formatrnsd/lnsdread) frompath; its[reticulum]and[interfaces]sections add to whatever the builder set programmatically. Loading a config brings up every interface type it names, including RNode and Serial, so a C node reaches LoRa through a config file.lev_builder_share_instancemakes the node offer a shared instance undername: it opens a local IPC endpoint and thernstatus/rnpath/rnprobeRPC server, so other local programs (and tools) attach to this one stack.lev_builder_connect_shared_instancemakes the node a client of a shared instance namednameinstead of bringing up its own interfaces, the wayrncp/rnxattach to a running daemon. ANULLpath or name returnsLEV_ERR_INVALID_ARG.lev_builder_buildproduces aleviculum_tand empties the builder; you still calllev_builder_freeon the empty handle.NULLon failure.lev_startspawns the event loop and brings up interfaces;lev_stoppersists state and tears it down; a stopped node can be started again.lev_starton a running node returnsLEV_ERR_CONFIG, as does starting a node configured to serve a shared instance whose name another daemon on the host already serves —lev_last_errorthen carries the instance name. A node meant to use the running daemon should connect as a client instead, not serve.lev_is_runningreturns 1 while the loop runs (0 onNULL).lev_identity_hash_selfwrites the node's own 16-byte identity hash.lev_freestops a running node and releases it (lev_free(NULL)is a no-op). Call it, and the other blocking calls, from a plain OS thread.
The event-side functions on a node (lev_event_fd, lev_next_event,
lev_wait_event) are documented under Events.
Destinations and announce
struct lev_destination_t *lev_destination_new(const struct lev_identity_t *identity,
int direction, int dest_type,
const char *app_name,
const char *const *aspects, uintptr_t n_aspects);
void lev_destination_free(struct lev_destination_t *dest);
int lev_destination_hash(const struct lev_destination_t *dest, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_destination_enable_ratchets(struct lev_destination_t *dest, uint64_t now_ms);
int lev_destination_ratchet_public(const struct leviculum_t *node, const uint8_t *dest_hash, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_destination_set_proof_strategy(struct lev_destination_t *dest, int strategy);
int lev_send_proof(const struct leviculum_t *node, const uint8_t *dest_hash, const uint8_t *packet_hash, int timeout_ms);
int lev_register_destination(const struct leviculum_t *node, struct lev_destination_t *dest);
int lev_announce(const struct leviculum_t *node, const uint8_t *dest_hash,
const uint8_t *app_data, uintptr_t app_data_len, int timeout_ms);
int lev_send_datagram(const struct leviculum_t *node, const uint8_t *dest_hash,
const uint8_t *data, uintptr_t data_len, uint8_t *out_hash, int timeout_ms);
lev_destination_newbuilds a destination from an identity (may beNULL; required for some types, forbidden forLEV_DEST_PLAIN), a direction, a type, anapp_name, and an array ofn_aspectsNUL-terminated aspect strings.NULLon failure.lev_destination_hashwrites the 16-byte hash; read it before registering. ReturnsLEV_ERR_INVALID_ARGonce the destination has been consumed.lev_destination_enable_ratchetsturns on forward secrecy for an inbound destination before it is registered;now_msis the current time in milliseconds, seeding ratchet rotation.LEV_ERR_INVALID_ARGfor an outbound destination or one already registered.lev_destination_ratchet_publicreads the current 32-byte ratchet public key of a registered destination (read(2) style), orLEV_ERR_INVALID_ARGif it has no ratchets. Ratcheted destinations interoperate with Python peers.lev_destination_set_proof_strategysets, before registration, how a destination proves delivery of received packets:LEV_PROOF_NONE(default, never),LEV_PROOF_APP(emitLEV_EVENT_PACKET_PROOF_REQUESTEDso the app decides, then callslev_send_proof), orLEV_PROOF_ALL(auto-prove every packet, Python's PROVE_ALL).lev_send_proofsends a delivery proof for thepacket_hashfrom a proof-requested event;LEV_ERR_SENDif no return path exists.lev_register_destinationregisters the destination on the node so it can be announced and accept links and packets. It consumes the destination (the handle is emptied; still free it).LEV_ERR_INVALID_ARGif already registered.lev_announcebroadcasts a registered destination (by 16-byte hash) on all interfaces, with optionalapp_data.lev_send_datagramsends one unreliable packet to a destination and writes the 16-byte packet hash intoout_hash. A path must be known (LEV_ERR_NO_PATHotherwise).
| Constant | Value | Meaning |
|---|---|---|
LEV_DIRECTION_IN | 0 | incoming: receives announces, links, packets |
LEV_DIRECTION_OUT | 1 | outgoing: a source address for sending |
LEV_DEST_SINGLE | 0 | point-to-point, ephemeral encryption |
LEV_DEST_GROUP | 1 | shared-key broadcast |
LEV_DEST_PLAIN | 2 | unencrypted |
Paths, connect, and links
int lev_has_path(const struct leviculum_t *node, const uint8_t *dest_hash);
int lev_hops_to(const struct leviculum_t *node, const uint8_t *dest_hash, uint8_t *out);
int lev_request_path(const struct leviculum_t *node, const uint8_t *dest_hash, int timeout_ms);
int lev_connect(const struct leviculum_t *node, const uint8_t *dest_hash,
int timeout_ms, struct lev_link_t **out);
int lev_connect_with_key(const struct leviculum_t *node, const uint8_t *dest_hash,
const uint8_t *signing_key, int timeout_ms, struct lev_link_t **out);
int lev_accept_link(const struct leviculum_t *node, const uint8_t *link_id,
int timeout_ms, struct lev_link_t **out);
int lev_link_send(const struct lev_link_t *link, const uint8_t *data, uintptr_t len, int timeout_ms);
int lev_link_try_send(const struct lev_link_t *link, const uint8_t *data, uintptr_t len);
int lev_link_id(const struct lev_link_t *link, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_link_is_closed(const struct lev_link_t *link);
int lev_link_identify(const struct leviculum_t *node, const uint8_t *link_id,
const struct lev_identity_t *identity, int timeout_ms);
struct lev_identity_t *lev_link_remote_identity(const struct leviculum_t *node, const uint8_t *link_id);
int lev_close_link(struct lev_link_t *link, int timeout_ms);
void lev_link_free(struct lev_link_t *link);
lev_has_pathreturns 1 if a path to the destination is known, else 0 (negative on a NULL argument).lev_hops_towrites the hop count into*outor returnsLEV_ERR_NO_PATH.lev_request_pathasks the network for a path; the result arrives as an event andlev_has_paththen returns 1.lev_connectopens a link by destination hash, resolving the peer's signing key from the identity cached by an announce;*outreceives the link. ReturnsLEV_ERR_UNKNOWN_DESTif no identity is cached andLEV_ERR_NO_PATHif no path is known (it does not auto-request one).lev_connect_with_keyis the same with an explicit 32-byte Ed25519 signing key, for out-of-band peers.lev_accept_linkaccepts an incoming link request (16-byte link id from aLEV_EVENT_LINK_REQUESTevent);*outreceives the link.lev_link_sendsends data, retrying backpressure up to the deadline (thenLEV_ERR_TIMEOUT);lev_link_try_sendnever blocks and returnsLEV_ERR_AGAINunder backpressure. Inbound data arrives asLEV_EVENT_LINK_DATA.lev_link_idwrites the 16-byte link id.lev_link_is_closedreturns 1 if closed (0 onNULL).lev_link_identifyproves an identity to the peer (who seesLEV_EVENT_LINK_IDENTIFIED);lev_link_remote_identityreturns the peer's identity as a new handle the caller frees, orNULLif the peer has not identified.lev_close_linkcloses gracefully (idempotent);lev_link_freereleases the handle, closing an open link first.
Request and response
int lev_register_request_handler(const struct leviculum_t *node, const uint8_t *dest_hash,
const char *path, int policy,
const uint8_t *allow_identity_hashes, uintptr_t n_ids);
int lev_send_request(const struct leviculum_t *node, const uint8_t *link_id, const char *path,
const uint8_t *data, uintptr_t data_len,
int response_timeout_ms, uint8_t *out_request_id);
int lev_send_response(const struct leviculum_t *node, const uint8_t *link_id,
const uint8_t *request_id, const uint8_t *data, uintptr_t data_len, int timeout_ms);
int lev_send_response_resource(const struct leviculum_t *node, const uint8_t *link_id,
const uint8_t *request_id, const uint8_t *data, uintptr_t data_len,
int timeout_ms);
int lev_send_file_response(const struct leviculum_t *node, const uint8_t *link_id,
const uint8_t *request_id, const uint8_t *data, uintptr_t data_len,
const uint8_t *metadata, uintptr_t metadata_len, int timeout_ms);
int lev_deregister_request_handler(const struct leviculum_t *node, const uint8_t *dest_hash,
const char *path);
lev_register_request_handlerregisters a handler forpathon a local destination. ForLEV_REQUEST_POLICY_ALLOW_LIST,allow_identity_hashesisn_ids * 16bytes of identity hashes; otherwise passNULL, 0. Registering overwrites a previous handler for the same destination and path;lev_deregister_request_handlerretires one.lev_send_requestsends a request on an established link topathand writes the 16-byte request id intoout_request_id.datais the msgpack-encoded payload (NULL, 0for none);response_timeout_msis the request-response deadline. The response (LEV_EVENT_RESPONSE_RECEIVED) or a timeout (LEV_EVENT_REQUEST_TIMEOUT) arrives as an event.lev_send_responsereplies to a received request (link id and request id from theLEV_EVENT_REQUEST_RECEIVEDevent);datamust be one valid msgpack-encoded value.lev_send_response_resourcereplies to the same request when the answer does not fit in one packet:lev_send_responseis bounded by the link MDU and returnsLEV_ERR_REQUESTabove it, and this call sends the answer as a resource instead. Same arguments, samedatacontract — one valid msgpack-encoded value, with no[request_id, response]wrapper of your own, because the library prepends the request id itself. Use this and notlev_send_resourcefor an over-MDU answer: a plain resource carries no request id, so the requester never correlates it and waits out its deadline.lev_send_file_responsereplies with a file rather than a value, the wire form a NomadNet/file/download has. This is the one of the three response calls whose name will not tell you it is different:datais sent as RAW bytes with no[request_id, response]wrapper, andmetadata(mandatory, one valid msgpack value, typically{"name": <basename>}) travels beside it. Its presence on the wire is what marks the response raw rather than wrapped, so it must not beNULL. The requester reads the bytes withlev_event_dataand the metadata withlev_event_metadataoff the oneLEV_EVENT_RESPONSE_RECEIVEDevent.lev_deregister_request_handlerretires the handler forpath:LEV_OKwhen one was registered and is now gone,LEV_ERR_NO_HANDLERwhen there was none, so a caller can tell "retired" from "never registered" without keeping its own book. Requests to a path with no handler are dropped without an answer, so a requester sees its deadline expire asLEV_EVENT_REQUEST_TIMEOUT— the same way any unserved path fails.
| Constant | Value | Meaning |
|---|---|---|
LEV_REQUEST_POLICY_ALLOW_NONE | 0 | drop all requests |
LEV_REQUEST_POLICY_ALLOW_ALL | 1 | allow any identity |
LEV_REQUEST_POLICY_ALLOW_LIST | 2 | allow only listed identity hashes |
Resource transfer
int lev_send_resource(const struct leviculum_t *node, const uint8_t *link_id,
const uint8_t *data, uintptr_t data_len,
const uint8_t *metadata, uintptr_t metadata_len,
int auto_compress, uint8_t *out_hash, int timeout_ms);
int lev_set_resource_strategy(const struct leviculum_t *node, const uint8_t *link_id, int strategy);
int lev_accept_resource(const struct leviculum_t *node, const uint8_t *link_id, int timeout_ms);
int lev_reject_resource(const struct leviculum_t *node, const uint8_t *link_id, int timeout_ms);
lev_send_resourcesends bulk data over a link and writes the 32-byte resource hash intoout_hash.metadata, if present, must be msgpack-encoded;auto_compressis 0 or 1. The call blocks only for the initial dispatch; progress and completion arrive as events.lev_set_resource_strategysets how incoming resources on a link are handled (one of theLEV_RESOURCE_*constants).lev_accept_resource/lev_reject_resourceanswer aLEV_EVENT_RESOURCE_ADVERTISEDevent under the AcceptApp strategy.
| Constant | Value | Meaning |
|---|---|---|
LEV_RESOURCE_ACCEPT_NONE | 0 | reject all incoming resources |
LEV_RESOURCE_ACCEPT_ALL | 1 | accept all automatically |
LEV_RESOURCE_ACCEPT_APP | 2 | advertise to the app to accept or reject |
LEV_RESOURCE_HASH_LEN | 32 | resource hash length |
Events
int lev_event_fd(const struct leviculum_t *node);
int lev_next_event(struct leviculum_t *node, struct lev_event_t **out);
int lev_wait_event(struct leviculum_t *node, struct lev_event_t **out, int timeout_ms);
void lev_event_free(struct lev_event_t *ev);
int lev_event_type(const struct lev_event_t *ev);
int lev_event_link_id(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_dest_hash(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_request_id(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_resource_hash(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_path(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_data(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_metadata(const struct lev_event_t *ev, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_event_progress(const struct lev_event_t *ev, double *out);
int lev_event_dropped_count(const struct lev_event_t *ev, uint64_t *out);
int lev_event_msgtype(const struct lev_event_t *ev, uint16_t *out);
int lev_event_sequence(const struct lev_event_t *ev, uint16_t *out);
int lev_event_is_sender(const struct lev_event_t *ev);
int lev_event_interface_id(const struct lev_event_t *ev, uint64_t *out);
int lev_event_close_reason(const struct lev_event_t *ev, int *out);
int lev_event_delivery_error(const struct lev_event_t *ev, int *out);
int lev_event_transfer_size(const struct lev_event_t *ev, uint64_t *out);
int lev_event_data_size(const struct lev_event_t *ev, uint64_t *out);
int lev_event_segment_index(const struct lev_event_t *ev, uint32_t *out);
int lev_event_total_segments(const struct lev_event_t *ev, uint32_t *out);
lev_event_fdreturns the readable fd to add to apoll/epoll/selectloop. The library owns it and closes it inlev_free; never close it.lev_next_eventdequeues without blocking: on success*outis an event handle, orNULLwhen the queue is empty.lev_wait_eventblocks up totimeout_ms(negative forever);*outisNULLif the timeout elapses. It wakes promptly when an event arrives (the event fd is the real wake source) and otherwise rechecks at most every 250 ms, so an infinite wait still returns soon after an event lands. Both are single-consumer for a node. Free each event withlev_event_free.lev_event_typereturns the event'sLEV_EVENT_*type (0 onNULL).- The accessors read a field of the event, read(2) style for the byte fields:
_link_id,_dest_hash,_request_id(16 bytes each),_resource_hash(32 bytes),_path(UTF-8 bytes, not NUL-terminated),_data(the primary payload, possibly empty),_metadata(msgpack bytes)._progresswrites adoublein0.0..1.0for resource-progress events;_dropped_countwrites the count of aLEV_EVENT_CONTROL_OVERFLOWevent;_msgtypeand_sequencewrite the message type and sequence of aLEV_EVENT_LINK_MESSAGEevent. An accessor that does not apply to the event type returnsLEV_ERR_INVALID_ARG. lev_event_is_senderreturns 1 on the sender side of a resource event (_PROGRESS/_COMPLETED/_FAILED), 0 on the receiver side. A sender'sLEV_EVENT_RESOURCE_COMPLETED(empty data) signals that an outgoing transfer finished; a receiver's carries the data. A node that both sends and receives resources uses this to tell the two apart.LEV_EVENT_RESOURCE_STARTEDis not in that list because the engine emits it on the receiver only, so 0 there is the truth rather than a gap. One non-resource event also sets the flag:LEV_EVENT_LINK_ESTABLISHEDreturns 1 for a link this node initiated and 0 for an inbound one. Everything else returns 0.lev_event_interface_idwrites the node-assigned id of the interface an_ANNOUNCE_RECEIVED,_PATH_FOUNDor_PACKET_RECEIVEDevent arrived on. It is an id, not a position: resolve it by walkinglev_interface_stats_snapshotand comparinglev_interface_stats_id. It is the same numberinglev_path_table_entryreports asinterface_index.lev_event_close_reasonwrites one of theLEV_CLOSE_*constants for aLEV_EVENT_LINK_CLOSEDevent. The reason decides how to reconnect:LEV_CLOSE_BLACKHOLEDmust not be retried at all,LEV_CLOSE_TIMEOUTwants the path re-resolved first.lev_event_delivery_errorwrites one of theLEV_DELIVERY_*constants for aLEV_EVENT_DELIVERY_FAILEDevent._TIMEOUTand_LINK_FAILEDmean re-send;_INVALID_PROOFmeans the peer answered with a proof that did not verify, so a re-send produces the same result and the destination's identity needs re-resolving instead.lev_event_transfer_sizeandlev_event_data_sizewrite the encrypted transfer size and the uncompressed payload size of a resource, onLEV_EVENT_RESOURCE_ADVERTISED(what the accept-or-reject decision turns on) and onLEV_EVENT_RESOURCE_PROGRESS(the only place an auto-accepting receiver sees them, since no advertisement is surfaced underACCEPT_ALL).lev_event_segment_indexandlev_event_total_segmentswrite the 1-based position of aLEV_EVENT_RESOURCE_COMPLETEDsegment and how many segments the transfer has. A multi-segment resource fires one completion per segment under the sameresource_hash, so these are how a receiver orders the chunks and recognises the last one (segment_index == total_segments); metadata is present on segment 1 only.
| Constant | Value | Link close reason |
|---|---|---|
LEV_CLOSE_NORMAL | 0 | closed deliberately; reconnect freely |
LEV_CLOSE_TIMEOUT | 1 | handshake did not complete; re-resolve the path |
LEV_CLOSE_INVALID_PROOF | 2 | a proof did not verify |
LEV_CLOSE_PEER_CLOSED | 3 | the peer closed it |
LEV_CLOSE_STALE | 4 | inactive past the keepalive deadline |
LEV_CLOSE_CHANNEL_EXHAUSTED | 5 | a channel message ran out of retries |
LEV_CLOSE_BLACKHOLED | 6 | peer is blackholed; do not retry |
LEV_CLOSE_OTHER | 255 | a reason this ABI version does not name |
| Constant | Value | Delivery failure |
|---|---|---|
LEV_DELIVERY_TIMEOUT | 0 | no proof before the receipt expired; re-send |
LEV_DELIVERY_LINK_FAILED | 1 | the carrying link failed; re-send |
LEV_DELIVERY_INVALID_PROOF | 2 | proof did not verify; re-resolve the identity |
LEV_DELIVERY_OTHER | 255 | a failure this ABI version does not name |
Event types
| Constant | Value | Fields available |
|---|---|---|
LEV_EVENT_OTHER | 0 | catch-all for events without a typed projection |
LEV_EVENT_ANNOUNCE_RECEIVED | 1 | dest_hash, data (app_data), interface_id |
LEV_EVENT_PATH_FOUND | 2 | dest_hash, interface_id |
LEV_EVENT_LINK_REQUEST | 3 | link_id, dest_hash |
LEV_EVENT_LINK_ESTABLISHED | 4 | link_id, dest_hash, is_sender |
LEV_EVENT_LINK_CLOSED | 5 | link_id, dest_hash, close_reason |
LEV_EVENT_LINK_DATA | 6 | link_id, data |
LEV_EVENT_PACKET_RECEIVED | 7 | dest_hash, data, interface_id |
LEV_EVENT_CONTROL_OVERFLOW | 8 | dropped_count |
LEV_EVENT_REQUEST_RECEIVED | 9 | link_id, dest_hash, request_id, path, data |
LEV_EVENT_RESPONSE_RECEIVED | 10 | link_id, request_id, data, metadata |
LEV_EVENT_REQUEST_TIMEOUT | 11 | link_id, request_id |
LEV_EVENT_RESOURCE_ADVERTISED | 12 | link_id, resource_hash, transfer_size, data_size |
LEV_EVENT_RESOURCE_STARTED | 13 | link_id, resource_hash |
LEV_EVENT_RESOURCE_PROGRESS | 14 | link_id, resource_hash, progress, transfer_size, data_size, is_sender |
LEV_EVENT_RESOURCE_COMPLETED | 15 | link_id, resource_hash, data, metadata, segment_index, total_segments, is_sender |
LEV_EVENT_RESOURCE_FAILED | 16 | link_id, resource_hash, is_sender |
LEV_EVENT_LINK_IDENTIFIED | 17 | link_id, data (16-byte identity hash) |
LEV_EVENT_LINK_MESSAGE | 18 | link_id, data, msgtype, sequence (reliable channel) |
LEV_EVENT_PACKET_PROOF_REQUESTED | 19 | dest_hash, data (32-byte packet hash), interface_id |
LEV_EVENT_LINK_PROOF_REQUESTED | 20 | link_id, data (32-byte packet hash) |
LEV_EVENT_LINK_DELIVERY_CONFIRMED | 21 | link_id, data (32-byte packet hash) |
LEV_EVENT_LINK_STALE | 22 | link_id (link inactive past keepalive) |
LEV_EVENT_LINK_RECOVERED | 23 | link_id (stale link resumed) |
LEV_EVENT_PATH_LOST | 24 | dest_hash (path expired) |
LEV_EVENT_PACKET_DELIVERY_CONFIRMED | 25 | data (16-byte packet hash) |
LEV_EVENT_DELIVERY_FAILED | 26 | data (16-byte packet hash), delivery_error |
LEV_EVENT_LINK_DELIVERY_FAILED | 27 | link_id, data (32-byte packet hash) |
Diagnostics
int lev_transport_stats(const struct leviculum_t *node,
uint64_t *out_packets_sent, uint64_t *out_packets_received,
uint64_t *out_packets_forwarded, uint64_t *out_announces_processed,
uint64_t *out_packets_dropped, uint64_t *out_path_count);
struct lev_path_table_t *lev_path_table_snapshot(const struct leviculum_t *node);
int lev_path_table_count(const struct lev_path_table_t *table);
int lev_path_table_entry(const struct lev_path_table_t *table, uintptr_t index,
uint8_t *dest_hash, uint8_t *hops, uint8_t *next_hop,
int *has_next_hop, uint64_t *interface_index, uint64_t *expires_ms);
void lev_path_table_free(struct lev_path_table_t *table);
struct lev_interface_stats_t *lev_interface_stats_snapshot(const struct leviculum_t *node);
int lev_interface_stats_count(const struct lev_interface_stats_t *table);
int lev_interface_stats_name(const struct lev_interface_stats_t *table, uintptr_t index,
uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_interface_stats_entry(const struct lev_interface_stats_t *table, uintptr_t index,
int *online, int *is_local_client,
uint64_t *rx_bytes, uint64_t *tx_bytes);
int lev_interface_stats_id(const struct lev_interface_stats_t *table, uintptr_t index,
uint64_t *out_id);
void lev_interface_stats_free(struct lev_interface_stats_t *table);
int lev_tcp_listen_addr(const struct leviculum_t *node, uintptr_t index,
uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
lev_transport_statsreads the node's transport counters and the current path-table size into the out-parameters, the basis for anrnstatus-style view. Any out-pointer may beNULLto skip that counter; aNULLnode returnsLEV_ERR_NULL_PTR. The values are a point-in-time snapshot.lev_path_table_snapshotreturns an owned, frozen copy of the path table for anrnpath-style view, orNULLon aNULLnode; free it withlev_path_table_free(NULLis a no-op). Because it is frozen, reads never race a changing table.lev_path_table_countgives the number of entries.lev_path_table_entryreads one entry by index into the out-parameters:dest_hashandnext_hop(each at leastLEV_ADDR_LENbytes when non-NULL),hops,has_next_hop(1 for a relayed path, 0 for a direct one),interface_index, andexpires_ms. Any out-pointer may beNULL;LEV_ERR_INVALID_ARGifindexis out of range.lev_interface_stats_snapshotreturns an owned, frozen copy of the interface list for anrnstatus-style interface view, freed withlev_interface_stats_free.lev_interface_stats_countgives the number of interfaces.lev_interface_stats_namereads the interface name read(2) style (variable length), andlev_interface_stats_entryreads the scalar fields (online,is_local_client,rx_bytes,tx_bytes) into out-parameters. Both returnLEV_ERR_INVALID_ARGfor an out-of-range index.lev_interface_stats_idreads the node-assigned id of the interface at a position. The position is not the identity: the node numbers interfaces as they are registered and never renumbers, so a removed interface leaves a gap. This is the accessor that resolves the ids the rest of the API hands out —lev_path_table_entry'sinterface_indexandlev_event_interface_id— to an entry whose name and counters can be read.lev_tcp_listen_addrwrites the boundip:portof the node's TCP server listener numberindex(in start order), read(2) style. A server added with port 0 reports the kernel-assigned port here afterlev_start— bind:0and read the port back instead of probing a free port up front and racing other processes for the re-bind.LEV_ERR_INVALID_ARGifindexis out of range, including before start.
Helpers
int lev_hex_encode(const uint8_t *data, uintptr_t len, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
int lev_hex_decode(const uint8_t *hex, uintptr_t hex_len, uint8_t *buf, uintptr_t cap, uintptr_t *out_len);
lev_hex_encodewrites2 * lenlowercase hex bytes (not NUL-terminated), read(2) style.lev_hex_decodewriteshex_len / 2bytes;LEV_ERR_INVALID_ARGon an odd length or a non-hex digit.
RNode Interface Protocol Research
Research based on Python RNS v1.1.3, source files:
RNS/Interfaces/RNodeInterface.py(1558 lines)RNS/Interfaces/RNodeMultiInterface.py(1149 lines)RNS/Interfaces/Interface.py(302 lines, base class)RNS/Interfaces/KISSInterface.py(standard KISS, for comparison)
1. Serial Protocol
1.1 Framing
The RNode serial protocol uses KISS framing, not HDLC. This is a critical distinction from the TCP/Serial framing used elsewhere in Reticulum.
| Constant | Value | Purpose |
|---|---|---|
FEND | 0xC0 | Frame delimiter (start and end) |
FESC | 0xDB | Escape byte |
TFEND | 0xDC | Escaped FEND (after FESC) |
TFESC | 0xDD | Escaped FESC (after FESC) |
Frame format:
[FEND 0xC0] [CMD byte] [escaped payload...] [FEND 0xC0]
Escaping (KISS standard):
When the payload contains 0xC0 (FEND) or 0xDB (FESC), they are replaced:
0xDB->0xDB 0xDD(FESC TFESC) -- escape is applied FIRST0xC0->0xDB 0xDC(FESC TFEND)
Note the escape ordering: FESC bytes are escaped first, then FEND bytes.
This matches Python's data.replace(bytes([0xdb]), bytes([0xdb, 0xdd])).replace(bytes([0xc0]), bytes([0xdb, 0xdc])).
Comparison with HDLC framing (used for TCP):
| Property | KISS (RNode) | HDLC (TCP) |
|---|---|---|
| Delimiter | 0xC0 | 0x7E |
| Escape byte | 0xDB | 0x7D |
| Escape method | Substitution (0xDC/0xDD) | XOR with 0x20 |
| CRC | None | None (Reticulum simplified HDLC) |
| First byte after delimiter | Command byte | Payload starts immediately |
Key difference: KISS frames carry a command byte after the opening FEND. Standard HDLC frames do not. The RNode protocol is a KISS superset with RNode-specific command extensions.
1.2 Command Set (complete table)
Configuration Commands (Host -> Device, Device -> Host as confirmation)
| Command | Byte | Direction | Payload | Description |
|---|---|---|---|---|
CMD_DATA | 0x00 | Both | Raw packet bytes (KISS-escaped) | Reticulum packet data |
CMD_FREQUENCY | 0x01 | Both | 4 bytes, big-endian, Hz | Set/report operating frequency |
CMD_BANDWIDTH | 0x02 | Both | 4 bytes, big-endian, Hz | Set/report channel bandwidth |
CMD_TXPOWER | 0x03 | Both | 1 byte, dBm | Set/report TX power |
CMD_SF | 0x04 | Both | 1 byte (5-12) | Set/report spreading factor |
CMD_CR | 0x05 | Both | 1 byte (5-8) | Set/report coding rate (4/5 through 4/8) |
CMD_RADIO_STATE | 0x06 | Both | 1 byte: 0x00=off, 0x01=on, 0xFF=ask | Set/report radio on/off state |
CMD_RADIO_LOCK | 0x07 | Device->Host | 1 byte | Report radio lock state |
CMD_DETECT | 0x08 | Both | Host sends 0x73, device responds 0x46 | Device presence detection handshake |
CMD_LEAVE | 0x0A | Host->Device | 0xFF | Host is disconnecting (shutdown notification) |
CMD_ST_ALOCK | 0x0B | Both | 2 bytes, big-endian, value/100 = percent | Short-term airtime limit |
CMD_LT_ALOCK | 0x0C | Both | 2 bytes, big-endian, value/100 = percent | Long-term airtime limit |
CMD_READY | 0x0F | Device->Host | (none meaningful) | Device ready for next TX packet |
Statistics Commands (Device -> Host, unsolicited)
| Command | Byte | Direction | Payload | Description |
|---|---|---|---|---|
CMD_STAT_RX | 0x21 | Device->Host | 4 bytes, big-endian | Total RX packet count |
CMD_STAT_TX | 0x22 | Device->Host | 4 bytes, big-endian | Total TX packet count |
CMD_STAT_RSSI | 0x23 | Device->Host | 1 byte (unsigned + 157 offset) | Last packet RSSI |
CMD_STAT_SNR | 0x24 | Device->Host | 1 byte (signed * 0.25 dB) | Last packet SNR |
CMD_STAT_CHTM | 0x25 | Device->Host | 11 bytes (see below) | Channel time/utilization stats |
CMD_STAT_PHYPRM | 0x26 | Device->Host | 12 bytes (see below) | Physical layer parameters |
CMD_STAT_BAT | 0x27 | Device->Host | 2 bytes: [state, percent] | Battery status |
CMD_STAT_CSMA | 0x28 | Device->Host | 3 bytes: [band, min, max] | CSMA contention window params |
CMD_STAT_TEMP | 0x29 | Device->Host | 1 byte (value - 120 = Celsius) | CPU temperature |
System Commands
| Command | Byte | Direction | Payload | Description |
|---|---|---|---|---|
CMD_BLINK | 0x30 | Host->Device | (unknown) | Blink LED for identification |
CMD_RANDOM | 0x40 | Device->Host | 1 byte | Hardware random byte |
CMD_FB_EXT | 0x41 | Host->Device | 1 byte: 0x00=disable, 0x01=enable | External framebuffer control |
CMD_FB_READ | 0x42 | Both | Host sends 0x01; Device responds with 512 bytes | Read framebuffer |
CMD_FB_WRITE | 0x43 | Host->Device | [line_byte] + [8 bytes line data] | Write framebuffer line |
CMD_BT_CTRL | 0x46 | Host->Device | (unknown) | Bluetooth control |
CMD_PLATFORM | 0x48 | Both | Host sends 0x00; Device responds with platform byte | Query/report platform |
CMD_MCU | 0x49 | Both | Host sends 0x00; Device responds with MCU byte | Query/report MCU type |
CMD_FW_VERSION | 0x50 | Both | Host sends 0x00; Device responds with 2 bytes [major, minor] | Query/report firmware version |
CMD_ROM_READ | 0x51 | Host->Device | (unknown) | Read ROM data |
CMD_RESET | 0x55 | Both | Host sends 0xF8; Device sends 0xF8 on reset | Hard reset / reset notification |
CMD_DISP_READ | 0x66 | Both | Host sends 0x01; Device responds with 1024 bytes | Read display buffer |
Multi-Interface Commands (RNodeMultiInterface only)
| Command | Byte | Direction | Payload | Description |
|---|---|---|---|---|
CMD_INTERFACES | 0x71 | Both | Host queries; Device responds with 2 bytes per interface [vport, type] | List available radio interfaces |
CMD_SEL_INT | 0x1F | Host->Device | 1 byte: interface index | Select subinterface for next command |
CMD_INT0_DATA | 0x00 | Device->Host | Packet data | Data received on interface 0 |
CMD_INT1_DATA | 0x10 | Device->Host | Packet data | Data received on interface 1 |
CMD_INT2_DATA | 0x20 | Device->Host | Packet data | Data received on interface 2 |
CMD_INT3_DATA | 0x70 | Device->Host | Packet data | Data received on interface 3 |
CMD_INT4_DATA | 0x75 | Device->Host | Packet data | Data received on interface 4 |
CMD_INT5_DATA | 0x90 | Device->Host | Packet data | Data received on interface 5 |
CMD_INT6_DATA | 0xA0 | Device->Host | Packet data | Data received on interface 6 |
CMD_INT7_DATA | 0xB0 | Device->Host | Packet data | Data received on interface 7 |
CMD_INT8_DATA | 0xC0 | Device->Host | Packet data | Data received on interface 8 |
CMD_INT9_DATA | 0xD0 | Device->Host | Packet data | Data received on interface 9 |
CMD_INT10_DATA | 0xE0 | Device->Host | Packet data | Data received on interface 10 |
CMD_INT11_DATA | 0xF0 | Device->Host | Packet data | Data received on interface 11 |
Note: CMD_INT8_DATA (0xC0) collides with FEND. This appears to be
an oversight or intentional oddity in the multi-interface protocol. It
means interface 8 data cannot actually be distinguished from a frame
delimiter. In practice, multi-interface devices may not populate all 12
slots.
Error Codes (in CMD_ERROR payload)
| Error | Byte | Description |
|---|---|---|
ERROR_INITRADIO | 0x01 | Radio initialization failed |
ERROR_TXFAILED | 0x02 | Transmission failed |
ERROR_EEPROM_LOCKED | 0x03 | EEPROM is locked |
ERROR_QUEUE_FULL | 0x04 | TX queue full (single-interface only) |
ERROR_MEMORY_LOW | 0x05 | Memory exhausted (single-interface only) |
ERROR_MODEM_TIMEOUT | 0x06 | Modem communication timeout (single-interface only) |
Platform Constants
| Platform | Byte | Description |
|---|---|---|
PLATFORM_AVR | 0x90 | AVR-based RNode |
PLATFORM_ESP32 | 0x80 | ESP32-based RNode |
PLATFORM_NRF52 | 0x70 | nRF52-based RNode |
Radio Chip Types (Multi-Interface only)
| Chip | Byte | Frequency Range |
|---|---|---|
SX127X | 0x00 | Sub-GHz (137 MHz - 1 GHz) |
SX1276 | 0x01 | Sub-GHz |
SX1278 | 0x02 | Sub-GHz |
SX126X | 0x10 | Sub-GHz |
SX1262 | 0x11 | Sub-GHz |
SX128X | 0x20 | 2.4 GHz (2.2 GHz - 2.6 GHz) |
SX1280 | 0x21 | 2.4 GHz |
1.3 Initialization Sequence
The host performs this exact sequence after opening the serial port:
Step 1: Open serial port
Baud: 115200
Data bits: 8
Stop bits: 1
Parity: None
Flow control: None (xonxoff=False, rtscts=False, dsrdtr=False)
Timeout: 0 (non-blocking reads)
Step 2: Wait 2.0 seconds
This is a hard-coded sleep to let the device settle after USB enumeration or power-on. Critical for reliability.
Step 3: Start read loop thread
A background thread begins reading bytes from the serial port and parsing KISS frames.
Step 4: Send detect + query commands (single frame sequence)
C0 08 73 C0 50 00 C0 48 00 C0 49 00 C0
This decodes as four back-to-back KISS frames:
FEND CMD_DETECT DETECT_REQ(0x73) FEND-- "Are you an RNode?"CMD_FW_VERSION 0x00 FEND-- "What firmware version?"CMD_PLATFORM 0x00 FEND-- "What platform?"CMD_MCU 0x00 FEND-- "What MCU?"
Note: Frames 2-4 rely on FEND at end of previous frame serving as start of next frame (KISS allows this).
For RNodeMultiInterface, an additional query is appended:
5. CMD_INTERFACES 0x00 FEND -- "List your radio interfaces"
Step 5: Wait for detect response (200ms for serial, 5s for TCP/BLE)
The read loop parses incoming bytes. When it sees CMD_DETECT with payload
0x46 (DETECT_RESP), it sets self.detected = True. The FW_VERSION,
PLATFORM, and MCU responses are also parsed and stored.
Step 6: Validate firmware version
Required minimum: major >= 1, minor >= 52 (for single-interface). Required minimum: major >= 1, minor >= 74 (for multi-interface).
Step 7: Configure radio parameters
Sends these commands in sequence:
CMD_FREQUENCYwith 4-byte big-endian frequency in HzCMD_BANDWIDTHwith 4-byte big-endian bandwidth in HzCMD_TXPOWERwith 1-byte TX power in dBmCMD_SFwith 1-byte spreading factor (5-12)CMD_CRwith 1-byte coding rate (5-8)CMD_ST_ALOCKwith 2-byte short-term airtime limit (if configured)CMD_LT_ALOCKwith 2-byte long-term airtime limit (if configured)CMD_RADIO_STATEwith0x01(RADIO_STATE_ON)
For multi-interface: each command is preceded by CMD_SEL_INT with the
subinterface index, and configurations are sent per-subinterface.
Step 8: Validate radio state
Wait 250ms (serial) / 1.0s (BLE) / 1.5s (TCP), then compare the
device-reported values (r_frequency, r_bandwidth, etc.) against the
configured values. Frequency must match within 100 Hz.
Step 9: Mark interface online
Wait 300ms, then set self.online = True.
1.4 Data Transfer
Outgoing (Host -> Device)
A Reticulum packet is wrapped as:
[FEND 0xC0] [CMD_DATA 0x00] [KISS-escaped packet bytes] [FEND 0xC0]
For multi-interface, data is sent as:
[FEND] [CMD_SEL_INT 0x1F] [interface_index] [FEND] [FEND] [CMD_DATA 0x00] [KISS-escaped packet bytes] [FEND]
The packet bytes are the raw Reticulum packet (header + payload), with NO additional metadata. RSSI/SNR are not included in outgoing packets.
Incoming (Device -> Host)
The device sends received packets as:
[FEND 0xC0] [CMD_DATA 0x00] [KISS-escaped packet bytes] [FEND 0xC0]
For multi-interface, the device uses interface-specific data commands:
[FEND 0xC0] [CMD_INTn_DATA] [KISS-escaped packet bytes] [FEND 0xC0]
where CMD_INTn_DATA indicates which radio interface received the packet.
Accompanying metadata (separate KISS frames, sent before the data frame):
The device sends RSSI and SNR as separate KISS frames before or after the data frame:
CMD_STAT_RSSI(0x23): 1 byte, unsigned. Actual RSSI = value - 157 dBmCMD_STAT_SNR(0x24): 1 byte, signed. Actual SNR = value * 0.25 dB
These are stored on the interface object and cleared after process_incoming
delivers the packet:
self.r_stat_rssi = None
self.r_stat_snr = None
Periodically reported statistics (unsolicited, from device)
The device periodically sends these frames without host request:
CMD_STAT_CHTM (0x25) -- Channel Time, 11 bytes:
Bytes 0-1: airtime_short (BE u16, /100 = percent)
Bytes 2-3: airtime_long (BE u16, /100 = percent)
Bytes 4-5: channel_load_short (BE u16, /100 = percent)
Bytes 6-7: channel_load_long (BE u16, /100 = percent)
Byte 8: current_rssi (unsigned, -157 offset)
Byte 9: noise_floor (unsigned, -157 offset)
Byte 10: interference (unsigned, -157 offset; 0xFF = no interference)
When the device sends it, and why that dates the keying: the firmware
calls kiss_indicate_channel_stats() as the last statement of
update_airtime() (RNode_Firmware.ino:712), and update_airtime() is the
last statement of both flush_queue() (:606) and pop_queue() (:644).
So a CHTM frame follows every keyed burst -- after LoRa->endPacket()
returned and add_airtime() (:751) folded that burst's airtime cost into
the bins -- on top of the idle cadence of roughly one per second. A rise
in airtime_short is therefore the modem's own receipt that it keyed:
airtime_bins is written by add_airtime() alone, and airtime is their
two-bin ratio (:698) over 15000 ms (Config.h:183), scaled by 100*100, so
one raw unit is 1.5 ms of airtime.
The arrival alone is not the receipt. transmit() with radio_online
false answers CMD_ERROR TXFAILED and keys nothing, yet its caller still
emits a CHTM, with airtime_short unmoved; a failed endPacket() (:744)
answers MODEM_TIMEOUT + TXFAILED and hard-resets, so no CHTM follows at
all. And the frame dates a burst, not a frame: below
LORA_GUARD_THRESHOLD_BPS = 14 kbps (Config.h:89), which is every LoRa
PHY we run, flush_queue() drains the whole queue before the single CHTM.
AVR RNodes compile the frame out entirely.
The interface logs each decoded CHTM as LORA_CHTM iface=<n> airtime_short=<pct> airtime_long=<pct> channel_load_short=<pct> channel_load_long=<pct>, on the same target and level as LORA_TX, so a
capture that holds the handovers holds the keying account beside them.
The frame a modem consumes without transmitting
Measured twice on the rig at firmware 1.85, on two different boards
(t-beam-1 and t-beam-2, 2026-09-24): a frame accepted over serial --
LORA_TX written, the KISS frame complete -- never keyed. No listener
decode, no decode at the far node, no airtime_short step in the next
CHTM while the medium was free, and the frame was not in the modem's
queue afterwards either: the next frame aired alone.
The firmware answers nothing on six of its ten paths from an accepted
CMD_DATA frame to no transmission. The two that produce exactly this
shape -- written, never aired, not queued afterwards -- are the length
guards in the two queue drains, flush_queue()
(RNode_Firmware.ino:586-589) and pop_queue()
(RNode_Firmware.ino:626-628): both pop the packet's start and length
off their FIFOs before testing them, so a rejected packet is already
gone. What feeds those guards a bad length is a queue-accounting desync
that nobody has placed yet; 1.85 also carries the pop_queue
accounting bug that 8382d4a fixes upstream after the tag, which
leaves queue_height permanently high once a packet is discarded.
stat_tx is never incremented in 1.85, and CMD_READY answers a
single queue-full bit that no realistic backlog reaches, so there is no
counter to read.
What there is, is the receipt above. The interface now consumes it. For
every frame it hands over it records the ledger as it stood, what the
frame should cost (its airtime at the running PHY plus the one header
byte transmit() prepends, RNode_Firmware.ino:720-724, in raw CHTM
units), and the window the answer has to arrive in: the frame's own
post-TX hold plus one firmware stat cadence. When a CHTM lands in that
window with the ledger unmoved, the interface says so:
LORA_TX_UNACCOUNTED iface=<n> len=<B> handover_t=<unix ms> chtm_t=<unix ms> expected_delta=<pct>
WARN level, on the same target as LORA_TX and LORA_CHTM, and
counted as tx_unaccounted in the interface stats. Every ambiguity
resolves away from the accusation, because a false one would poison the
instrument (judge_airtime,
leviculum-std/src/interfaces/rnode.rs:1316): any rise at all counts as
keyed, a falling ledger is read as a bin ageing out of the two-bin
window rather than as a swallowed frame, a rise in airtime_long
absolves even when the short window has not moved, a DCD busy fraction
that could cover the whole hold is read as a frame still legitimately
queued, and a configured airtime lock at or below the reported airtime
silences it. expected_delta is reported, never thresholded.
Behind the observable sits a workaround: the interface keeps its copy
of an unaccounted frame and re-hands it ONCE, naming the retry
(LORA_TX_REHAND iface=<n> len=<B> handover_t=<unix ms>). Once,
because a modem that swallows the retry too is not going to send this
frame, and repeating into it only buys latency for everything behind
it; the second verdict counts the frame as lost instead
(RNODE_TX_QUEUE_DROP ... reason=unaccounted_twice). What a re-hand
costs the far end is one duplicate, and the dedup cache drops it on
arrival (has_packet_hash,
leviculum-core/src/transport.rs:3956, which emits DEDUP_DROP). The
classes that cache exempts -- announces, and link requests and proofs
addressed to us -- are exactly the classes whose retries the stack
already expects, so a second copy there is processed as a retry rather
than as corruption.
No upstream report (project policy); this page and the code are the record.
Note: For multi-interface (RNodeMultiInterface), CMD_STAT_CHTM is only 8 bytes (no RSSI/noise_floor/interference fields):
Bytes 0-1: airtime_short (BE u16, /100 = percent)
Bytes 2-3: airtime_long (BE u16, /100 = percent)
Bytes 4-5: channel_load_short (BE u16, /100 = percent)
Bytes 6-7: channel_load_long (BE u16, /100 = percent)
CMD_STAT_PHYPRM (0x26) -- Physical Parameters, 12 bytes (single) / 10 bytes (multi):
Bytes 0-1: symbol_time (BE u16, /1000 = milliseconds)
Bytes 2-3: symbol_rate (BE u16, baud)
Bytes 4-5: preamble_symbols (BE u16)
Bytes 6-7: preamble_time (BE u16, milliseconds)
Bytes 8-9: csma_slot_time (BE u16, milliseconds)
Bytes 10-11: difs_time (BE u16, milliseconds) -- ONLY in single-interface
CMD_STAT_CSMA (0x28) -- CSMA Parameters, 3 bytes (single-interface only):
Byte 0: contention_window_band
Byte 1: contention_window_min
Byte 2: contention_window_max
CMD_STAT_BAT (0x27) -- Battery Status, 2 bytes:
Byte 0: battery_state (0x00=unknown, 0x01=discharging, 0x02=charging, 0x03=charged)
Byte 1: battery_percent (0-100, clamped)
CMD_STAT_TEMP (0x29) -- CPU Temperature, 1 byte:
Byte 0: temperature + 120 (actual temp = value - 120 Celsius)
Valid range: -30 to +90 Celsius
1.5 Flow Control
The RNode protocol implements software flow control via the CMD_READY mechanism:
- When
flow_control=Trueis configured, the host setsinterface_ready = Falseafter sending each packet. - The device sends a
CMD_READYframe when it has finished transmitting and is ready for the next packet. - Upon receiving
CMD_READY, the host callsprocess_queue():- If packets are queued, pops the first one and sends it.
- If no packets are queued, sets
interface_ready = True.
- If
interface_readyis False whenprocess_outgoing()is called, the packet is appended topacket_queueinstead of being sent immediately.
When flow_control=False (the default), interface_ready starts True and
is never set to False by the host. The CMD_READY frames from the device
still trigger process_queue(), which is a no-op if the queue is empty.
The packet queue is a simple FIFO list with no maximum size and no priority. Overflow is not explicitly handled.
2. Radio Configuration
2.1 Parameters and Encoding
Frequency (CMD_FREQUENCY, 0x01)
- 4 bytes, big-endian unsigned integer
- Unit: Hertz
- Example: 868.0 MHz =
0x33B13B40 - Encoding:
[freq >> 24, (freq >> 16) & 0xFF, (freq >> 8) & 0xFF, freq & 0xFF] - KISS-escaped after encoding
Bandwidth (CMD_BANDWIDTH, 0x02)
- 4 bytes, big-endian unsigned integer
- Unit: Hertz
- Encoding identical to frequency
- Valid range: 7,800 Hz to 1,625,000 Hz
TX Power (CMD_TXPOWER, 0x03)
- 1 byte, unsigned for single-interface (0-37 dBm)
- 1 byte, signed for multi-interface (-9 to +37 dBm). Encoded as
txpower.to_bytes(1, signed=True)on send; decoded asbyte - 256 if byte > 127 else byteon receive. - Unit: dBm
Spreading Factor (CMD_SF, 0x04)
- 1 byte, unsigned
- Valid range: 5-12
Coding Rate (CMD_CR, 0x05)
- 1 byte, unsigned
- Valid range: 5-8
- Represents 4/5 through 4/8
Short-term Airtime Limit (CMD_ST_ALOCK, 0x0B)
- 2 bytes, big-endian unsigned integer
- Encoding:
int(percent * 100)-- so 50.0% becomes 5000 - KISS-escaped after encoding
Long-term Airtime Limit (CMD_LT_ALOCK, 0x0C)
- Same encoding as ST_ALOCK
Radio State (CMD_RADIO_STATE, 0x06)
- 1 byte
0x00= OFF,0x01= ON,0xFF= ASK (query current state)
2.2 Hardware Variants
The Python code identifies devices by platform and MCU:
Platforms:
| Platform | Value | Has Display | Notes |
|---|---|---|---|
| AVR | 0x90 | No | Original Arduino-based RNode |
| ESP32 | 0x80 | Yes | Most common modern RNode |
| NRF52 | 0x70 | Yes | Nordic-based RNode |
Radio chips (Multi-Interface only):
| Chip Family | Frequency Range | Notes |
|---|---|---|
| SX127X (SX1276/SX1278) | 137 MHz - 1 GHz | Sub-GHz LoRa |
| SX126X (SX1262) | 137 MHz - 1 GHz | Sub-GHz LoRa, newer |
| SX128X (SX1280) | 2.2 GHz - 2.6 GHz | 2.4 GHz LoRa |
Frequency validation:
- Single-interface: 137 MHz to 3 GHz (broad range, device validates further)
- Multi-interface SX127X/SX126X: 137 MHz to 1 GHz
- Multi-interface SX128X: 2.2 GHz to 2.6 GHz
Hardware MTU: Fixed at 508 bytes for all RNode variants.
2.3 Firmware Detection
The host queries firmware version as part of the detect sequence:
FEND CMD_FW_VERSION 0x00 FEND
Response is 2 KISS-escaped bytes: [major, minor].
Required minimum versions:
- RNodeInterface (single radio): 1.52
- RNodeMultiInterface (multi radio): 1.74
If firmware is below minimum, Python calls RNS.panic() with instructions
to update via rnodeconf.
Validation logic:
if maj_version > REQUIRED_MAJ:
firmware_ok = True
elif maj_version >= REQUIRED_MAJ and min_version >= REQUIRED_MIN:
firmware_ok = True
2.4 Transport Variants: USB, TCP, BLE
The RNode can be accessed over three transport types. The serial protocol is identical over all three; only the physical transport differs.
USB Serial (default)
- Baud: 115200, 8N1
- Uses pyserial
Serialobject directly - Read timeout: 100ms
- Detect wait: 200ms
TCP (port specified as tcp://hostname)
- Port: 7633 (TCPConnection.TARGET_PORT)
- Uses raw TCP socket with
TCP_NODELAY - Has keepalive mechanism: sends detect command every 3.5s (ACTIVITY_KEEPALIVE = ACTIVITY_TIMEOUT - 2.5 = 3.5s)
- Read timeout: 1500ms
- Detect wait: 5.0s
- TCP keepalive probes: every 2s, after 5s idle, 12 probes, 24s user timeout
BLE (port specified as ble:// or ble://name or ble://AA:BB:CC:DD:EE:FF)
- Uses Nordic UART Service (NUS) over BLE GATT:
- Service:
6E400001-B5A3-F393-E0A9-E50E24DCCA9E - RX Char:
6E400002-B5A3-F393-E0A9-E50E24DCCA9E(host writes to device) - TX Char:
6E400003-B5A3-F393-E0A9-E50E24DCCA9E(device notifies host)
- Service:
- Requires device to be bonded (paired)
- Read timeout: 1250ms
- Detect wait: 5.0s
- Uses
bleakPython library for BLE access - Write chunk size limited to
max_write_without_response_size
3. Medium Access
3.1 CSMA/CA
CSMA is handled entirely by the RNode firmware, not the host.
The host has no CSMA logic. It simply sends packets to the device. The device reports its CSMA parameters back to the host for informational/ monitoring purposes via:
CMD_STAT_PHYPRM(0x26): symbol time, symbol rate, preamble symbols, preamble time, CSMA slot time, DIFS timeCMD_STAT_CSMA(0x28): contention window band, min, max
The host stores these values but does not use them for transmission decisions. All carrier sensing, backoff, and collision avoidance is performed by the RNode firmware.
3.2 Airtime Calculation
On-air bitrate calculation (host-side, for capacity planning):
bitrate = sf * (4.0 / cr) / (2**sf / (bandwidth/1000)) * 1000
Where:
sf= spreading factor (5-12)cr= coding rate (5-8, representing 4/5 through 4/8)bandwidth= channel bandwidth in Hz
This gives the effective data rate in bits per second.
Channel utilization (device-reported):
The device periodically sends CMD_STAT_CHTM with:
r_airtime_short: Short-term airtime percentage (own TX)r_airtime_long: Long-term airtime percentage (own TX)r_channel_load_short: Short-term channel load (all observed activity)r_channel_load_long: Long-term channel load (all observed activity)
The host does NOT compute channel utilization itself. It relies entirely on the device's reporting.
Airtime limiting:
If configured, the host sends airtime limits to the device:
CMD_ST_ALOCK: Short-term airtime limit percentageCMD_LT_ALOCK: Long-term airtime limit percentage
The device enforces these limits in firmware.
3.3 Timing and Jitter
The RNode interface applies NO send-side jitter or timing delays.
Unlike Transport's PATHFINDER_RW (0.5s random window for announce
rebroadcasts), the RNode interface has no equivalent jitter mechanism.
All timing is either:
- Device-side CSMA (carrier sensing in firmware)
- Transport-layer announce scheduling (handled by Transport, not the interface)
- Announce rate cap (in base Interface class, based on bitrate):
Defaulttx_time = (len(packet) * 8) / self.bitrate wait_time = tx_time / announce_capannounce_cap= 2% of interface bandwidth.
The 80ms sleep in the read loop idle path is purely to prevent busy-waiting when no data is available, not a timing mechanism.
Callsign beaconing:
If id_interval and id_callsign are configured, the interface
periodically transmits the callsign as raw packet data (not a Reticulum
packet). The timer resets on each TX. The first_tx timestamp records
when the first actual (non-callsign) packet was transmitted.
4. Queue Management
4.1 TX Pipeline
process_outgoing(data)
|
|-- Is interface online?
| No -> drop silently
|
|-- Is interface_ready?
| No -> queue(data) -> append to self.packet_queue
| Yes:
| |-- If flow_control: set interface_ready = False
| |-- KISS-escape the data
| |-- Build frame: FEND + CMD_DATA(0x00) + escaped_data + FEND
| |-- serial.write(frame)
| |-- Increment txb counter
Key observations:
packet_queueis an unbounded Python list (no max size)- FIFO ordering, no priority
interface_readystarts as False, set to True only after successful device configuration (in configure_device)- With flow_control=False (default), interface_ready is always True once online, so the queue is never used
- With flow_control=True, the queue drains one packet at a time via CMD_READY callbacks
4.2 RX Pipeline
readLoop() [background thread]
|
|-- Read 1 byte from serial
|-- Parse KISS frame state machine:
| |-- FEND: start new frame, reset command
| |-- First byte after FEND: set as command byte
| |-- CMD_DATA (0x00): accumulate into data_buffer with KISS unescaping
| |-- CMD_* (config): accumulate into command_buffer, parse when complete
| |-- FEND while in CMD_DATA frame: frame complete
|
|-- On complete CMD_DATA frame:
| process_incoming(data_buffer)
| |-- Increment rxb counter
| |-- self.owner.inbound(data, self) [delivers to Transport]
| |-- Clear r_stat_rssi and r_stat_snr
|
|-- On complete CMD_* frame:
| Parse and store in corresponding r_* fields
|
|-- Timeout handling:
| If partial frame and no data for > self.timeout ms:
| Clear buffer, reset state machine
Buffer size limit: self.HW_MTU (508 bytes). If data_buffer reaches
this size, additional bytes are silently dropped until the next FEND.
4.3 Threading Model
Single-interface (RNodeInterface):
Main thread: Read loop thread:
| |
configure_device() --> readLoop() [daemon]
| |
process_outgoing() ---- | <-- serial.read(1)
setFrequency() ---- | --> parse KISS frames
setBandwidth() ---- | --> update r_* fields
... ---- | --> process_incoming() -> owner.inbound()
| --> process_queue() [on CMD_READY]
|
| [80ms sleep when no data]
Both threads access self.serial (the pyserial object). There is NO
explicit locking between the write path (main thread) and the read path
(readLoop thread). pyserial's internal buffering provides some safety,
but this is technically a race condition in the Python implementation.
For BLE and TCP transports, separate TX/RX queues with locks are used:
ble_rx_lock/ble_tx_locktcp_rx_lock/tcp_tx_lock
Multi-interface (RNodeMultiInterface):
Same model, but the read loop dispatches to the correct sub-interface
based on CMD_INTn_DATA command bytes. The CMD_SEL_INT command in the
read loop updates self.selected_index, which determines which
sub-interface receives configuration confirmations.
5. Interface Lifecycle
5.1 INI Configuration
The [[RNode Interface]] section accepts these config keys:
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
name | string | Yes | -- | Interface name |
port | string | Yes | -- | Serial port path, tcp://host, or ble://... |
frequency | int | Yes | 0 | Operating frequency in Hz |
bandwidth | int | Yes | 0 | Channel bandwidth in Hz |
txpower | int | No | 22 (board maximum) | TX power in dBm. Absent asks for the board maximum, not the reference's 0 — see the pinned deviation. An explicit 0 still means 0. |
spreadingfactor | int | Yes | 0 | Spreading factor (5-12) |
codingrate | int | Yes | 0 | Coding rate (5-8) |
flow_control | bool | No | False | Enable TX flow control |
id_interval | int | No | None | Callsign beacon interval in seconds |
id_callsign | string | No | None | Callsign for beaconing (max 32 bytes UTF-8) |
airtime_limit_short | float | No | None | Short-term TX airtime limit (0-100%) |
airtime_limit_long | float | No | None | Long-term TX airtime limit (0-100%) |
RNodeMultiInterface adds:
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
port | string | Yes | -- | Serial port path |
Each sub-interface is defined as a nested section with:
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
interface_enabled | bool | No | (inherits parent enabled) | Enable this sub-interface |
vport | int | Yes | -- | Virtual port index on device |
frequency | int | Yes | -- | Frequency in Hz |
bandwidth | int | Yes | -- | Bandwidth in Hz |
txpower | int | No | 22 (board maximum) | TX power in dBm, resolved per subinterface the same way as above |
spreadingfactor | int | Yes | -- | Spreading factor |
codingrate | int | Yes | -- | Coding rate |
flow_control | bool | No | False | TX flow control |
airtime_limit_short | float | No | None | Short-term airtime limit |
airtime_limit_long | float | No | None | Long-term airtime limit |
outgoing | bool | No | True | Whether TX is allowed |
5.2 Connection Management
Startup:
- Validate configuration parameters
- Open serial port
- If open succeeds: configure_device (detect, init radio, validate)
- If open fails: start reconnect_port thread
Reconnection:
reconnect_port()runs in a loop:- Sleep 5 seconds (
RECONNECT_WAIT) - Try to open port and configure device
- Repeat until
onlineordetached
- Sleep 5 seconds (
- The readLoop also triggers reconnection when it catches an exception (serial port error, device reset, etc.)
- ESP32 devices send
CMD_RESET 0xF8when they reset while online, which the host treats as an error triggering reconnection.
Shutdown (detach):
- Set
self.detached = True - Disable external framebuffer
- Set radio state to OFF
- Send CMD_LEAVE
- Close BLE/TCP connections if applicable
Ingress limiting:
RNodeInterface overrides should_ingress_limit() to always return
False. This means RNode interfaces never throttle incoming announces
at the interface level (Transport still applies its own limiting).
5.3 Statistics
The interface tracks and exposes:
Counters (host-maintained):
rxb: Total bytes received (incremented in process_incoming)txb: Total bytes transmitted (incremented in process_outgoing)
Device-reported:
r_stat_rx: Total device RX packet count (4-byte)r_stat_tx: Total device TX packet count (4-byte)r_stat_rssi: Last packet RSSI in dBm (byte - 157)r_stat_snr: Last packet SNR in dB (signed_byte * 0.25)r_stat_q: Signal quality percentage (computed from SNR and SF)r_airtime_short/r_airtime_long: TX airtime percentagesr_channel_load_short/r_channel_load_long: Channel load percentagesr_battery_state/r_battery_percent: Battery infor_temperature/cpu_temp: CPU temperature in Celsius
Signal quality calculation:
q_snr_min = Q_SNR_MIN_BASE - (sf - 7) * Q_SNR_STEP # where BASE=-9, STEP=2
q_snr_max = Q_SNR_MAX # 6
q_snr_span = q_snr_max - q_snr_min
quality = clamp(((snr - q_snr_min) / q_snr_span) * 100, 0, 100)
RSSI decoding:
All RSSI values use the same offset: actual_dBm = raw_byte - 157.
6. Physical Device Info
Connected Device
Device: /dev/ttyACM0
USB Vendor: 1a86 (QinHeng Electronics)
USB Model: USB Single Serial (55d4)
USB Serial: 5896004228
Driver: cdc_acm
Symlinks: /dev/serial/by-id/usb-1a86_USB_Single_Serial_5896004228-if00
Probe Results
Detection: Successful (DETECT_RESP = 0x46)
Firmware: 1.85
Platform: ESP32 (0x80)
MCU: 0x81
Battery report: Received (CMD_STAT_BAT 0x27, state=0x00 unknown, percent=0%)
The device also sent an unsolicited CMD_STAT_BAT frame during the detect
sequence, which is expected -- the device reports battery status
periodically.
The QinHeng Electronics CH340/CH9102 USB-serial chip (VID 1a86, PID 55d4) is commonly used on ESP32 development boards, specifically the Heltec and LilyGO T-Beam variants commonly used for RNode.
7. Implementation Notes for Rust
7.1 What Maps to Our Interface Trait (Send Side)
The outgoing path is straightforward: process_outgoing(data) takes a raw
Reticulum packet and wraps it in a KISS frame. This maps to our Interface
trait's send method. The KISS framing (FEND + CMD_DATA + escape + FEND) is
a simple transformation.
The flow control queue (interface_ready / packet_queue) is host-side state that should live on the interface struct. When flow_control is enabled, the interface buffers packets until the device signals CMD_READY.
7.2 What Needs Its Own Async Task (Receive Side, Serial I/O)
The Python implementation uses a daemon thread for readLoop(). In our
async Rust architecture, this maps to an async task that:
- Reads bytes from the serial port (async serial I/O)
- Parses the KISS frame state machine
- Dispatches complete frames:
- CMD_DATA -> feed to NodeCore via
handle_packet() - CMD_STAT_* -> update interface metadata
- CMD_READY -> trigger queue drain
- CMD_ERROR -> handle errors
- CMD_DATA -> feed to NodeCore via
The serial port read should use tokio-serial or similar async serial
crate. The KISS deframer runs in the same task (no separate thread needed).
The read loop is the only path that needs to be truly async. All writes (config commands, data packets) can be synchronous or fire-and-forget since there's no write-side acknowledgment protocol.
7.3 Where Send-Side Jitter Fits
There is no send-side jitter in the RNode interface itself. All timing is handled by:
- RNode firmware: CSMA/CA with carrier sensing
- Transport layer: Announce rebroadcast random window (
PATHFINDER_RW = 0.5s) - Interface base class: Announce rate cap (
announce_cap = 2%)
The announce rate cap and queue management from the base Interface class should be implemented in the transport/driver layer, not in the RNode interface itself. The interface is a dumb pipe -- it takes packets from the send queue and KISS-frames them to the serial port.
7.4 State the Interface Needs to Maintain
Configuration (set once):
- frequency, bandwidth, txpower, sf, cr
- st_alock, lt_alock
- flow_control flag
- id_callsign, id_interval
- port path, transport type (USB/TCP/BLE)
Device-reported (updated from read loop):
- r_frequency, r_bandwidth, r_txpower, r_sf, r_cr, r_state, r_lock
- r_stat_rssi, r_stat_snr (per-packet, cleared after delivery)
- r_airtime_short, r_airtime_long, r_channel_load_short, r_channel_load_long
- r_symbol_time_ms, r_symbol_rate, r_preamble_symbols, r_preamble_time_ms
- r_csma_slot_time_ms, r_csma_difs_ms
- r_csma_cw_band, r_csma_cw_min, r_csma_cw_max
- r_battery_state, r_battery_percent, r_temperature
- detected, firmware_ok, maj_version, min_version
- platform, mcu, display
Runtime (host-managed):
- online flag
- interface_ready flag (for flow control)
- packet_queue (if flow_control enabled)
- rxb, txb counters
- first_tx timestamp (for callsign beaconing)
7.5 Relationship to Existing KISS Framing in leviculum-core
The existing framing code in leviculum-core/src/framing/hdlc.rs is
HDLC framing, NOT KISS framing. They are different protocols.
Key differences:
| Property | HDLC (existing) | KISS (needed for RNode) |
|---|---|---|
| Flag byte | 0x7E | 0xC0 (FEND) |
| Escape byte | 0x7D | 0xDB (FESC) |
| Escape method | XOR with 0x20 | Substitution: 0xDC (TFEND) or 0xDD (TFESC) |
| Command byte | None | First byte after FEND is command |
| Used for | TCP interfaces | Serial RNode interface |
We need a new framing/kiss.rs module alongside the existing hdlc.rs.
The module structure should be:
framing/
mod.rs -- re-exports both
hdlc.rs -- existing, for TCP
kiss.rs -- new, for RNode serial
The KISS module needs:
- Constants: FEND, FESC, TFEND, TFESC
fn kiss_escape(data: &[u8]) -> Vec<u8>fn kiss_frame(cmd: u8, data: &[u8]) -> Vec<u8>struct KissDeframerwith state machine for parsing incoming bytes (tracking command byte, escape state, buffer)
The KissDeframer should yield (command: u8, data: Vec<u8>) tuples,
not raw byte buffers like the HDLC Deframer.
7.6 Is This Standard KISS or a Superset?
It is a KISS superset. Standard KISS TNC protocol (as used in amateur radio) defines:
| Standard KISS | RNode Extension |
|---|---|
CMD_DATA (0x00) | Same |
CMD_TXDELAY (0x01) | Repurposed as CMD_FREQUENCY |
CMD_P (0x02) | Repurposed as CMD_BANDWIDTH |
CMD_SLOTTIME (0x03) | Repurposed as CMD_TXPOWER |
CMD_TXTAIL (0x04) | Repurposed as CMD_SF |
CMD_FULLDUPLEX (0x05) | Repurposed as CMD_CR |
CMD_SETHARDWARE (0x06) | Repurposed as CMD_RADIO_STATE |
CMD_RETURN (0xFF) | Not used by RNode |
| (none) | 0x07-0x0F: RNode-specific config commands |
| (none) | 0x21-0x29: RNode statistics |
| (none) | 0x30-0x55: RNode system commands |
| (none) | 0x66: Display read |
| (none) | 0x71, 0x1F: Multi-interface commands |
| (none) | 0x90: Error reporting |
The framing layer (FEND/FESC/TFEND/TFESC) is identical to standard KISS. The command bytes 0x01-0x06 overlap with standard KISS but have completely different semantics (frequency vs. txdelay, etc.).
This means our KISS framing module should implement the framing layer
generically, and the RNode command interpretation should be in a separate
module (e.g., interfaces/rnode.rs in leviculum-std).
7.7 Architecture Mapping
leviculum-core/src/framing/kiss.rs -- KISS framing (FEND/FESC escaping)
Layer 0, no_std compatible
Pure data transformation, no I/O
leviculum-std/src/interfaces/rnode.rs -- RNode interface implementation
Owns serial port (async I/O)
KISS command interpretation
Radio configuration state machine
Flow control queue management
leviculum-std/src/driver/ -- Existing driver integrates RNode
interface alongside TCP
The KISS framing module belongs in leviculum-core because it's a pure data transformation (like HDLC). The RNode interface logic belongs in leviculum-std because it performs I/O (serial port access).
7.8 Implementation Priority
For a minimal working RNode interface:
- KISS framing module (framing/kiss.rs) -- escape/unescape, frame/deframe
- RNode command parser -- interpret command bytes and payloads
- Initialization sequence -- detect, query, configure, validate
- Data path -- TX: KISS-frame packets; RX: deframe and deliver
- Statistics -- parse RSSI/SNR/channel stats from device
- Flow control -- CMD_READY queue management
- Reconnection -- handle disconnect/reconnect
- Multi-interface support -- CMD_SEL_INT, CMD_INTn_DATA
BLE and TCP transport support for RNode can be deferred; USB serial is the primary use case.
SX1261/2 Datasheet Reference (Driver Development Extract)
Source: Semtech SX1261/2 Data Sheet, Rev 2.2, DS.SX1261-2.W.APP, December 2024.
This document extracts the sections relevant for SX1262 LoRa driver development. For the complete datasheet, see semtech.com.
8. Digital Interface and Control
The SX1261/2 is controlled via a serial SPI interface and a set of general purpose input/output (DIOs). At least one DIO must be used for IRQ and the BUSY line is mandatory. BUSY indicates that the chip is ready for new command only if this signal is low.
8.1 Reset
A complete "factory reset" can be issued by toggling pin NRESET. It is automatically followed by the standard calibration procedure and any previous context is lost. The pin should be held low for typically 100us for the Reset to happen.
8.2 SPI Interface
The SPI interface uses a synchronous full-duplex protocol: CPOL = 0, CPHA = 0 (Mode 0). Only the slave side is implemented.
- MOSI is generated by the master on the falling edge of SCK and is sampled by the slave on the rising edge of SCK.
- MISO is generated by the slave on the falling edge of SCK.
- A transfer is always started by the NSS pin going low. MISO is high impedance when NSS is high.
- SPI runs on the external SCK clock to allow high speed up to 16 MHz.
SPI Timing Requirements (Table 8-1)
| Symbol | Description | Min | Max | Unit |
|---|---|---|---|---|
| t1 | NSS falling to SCK setup time | 32 | - | ns |
| t2 | SCK period | 62.5 | - | ns |
| t6 | NSS falling to MISO delay | 0 | 15 | ns |
| t7 | SCK falling to MISO delay | 0 | 15 | ns |
| t8 | SCK to NSS rising hold time | 31.25 | - | ns |
| t9 | NSS high time | 125 | - | ns |
| t10 | NSS falling to SCK setup when switching from SLEEP to STDBY_RC | 100 | - | us |
| t11 | NSS falling to MISO delay when switching from SLEEP to STDBY_RC | 0 | 150 | us |
8.2.2 SPI Timing When Leaving Sleep Mode
One way for the chip to leave Sleep mode is to wait for a falling edge of NSS. The delay between the falling edge of NSS and the first rising edge of SCK must take into account the wake-up sequence and the chip initialization. During Sleep mode and the initialization phase, BUSY is set high. Once the chip is in STDBY_RC mode, BUSY goes low and the host can start sending a command. This is also true for startup at battery insertion or after a hard reset.
8.3.1 BUSY Control Line
The BUSY control line indicates the status of the internal state machine. When BUSY is held low, the internal state machine is in idle mode and the radio is ready for a command.
For all "write" commands, BUSY is asserted high after time T_SW. T_SW from NSS rising edge to BUSY rising edge is max 600 ns in all cases.
"Read" commands are handled directly without the internal state machine and BUSY remains low after a read command.
Switching Times (Table 8-2)
| Transition | T_SW_Mode Typical (us) |
|---|---|
| SLEEP to STBY_RC cold start | 3500 |
| SLEEP to STBY_RC warm start | 340 |
| STBY_RC to STBY_XOSC | 31 |
| STBY_RC to FS | 50 |
| STBY_RC to RX | 83 |
| STBY_RC to TX | 126 |
| STBY_XOSC to TX | 105 |
8.4 Digital Interface Status versus Chip Modes (Table 8-3)
| Mode | DIO3 | DIO2 | DIO1 | BUSY | MISO | MOSI | SCK | NSS |
|---|---|---|---|---|---|---|---|---|
| Reset | PD | PD | PD | PU | HIZ | HIZ | HIZ | IN |
| Start-up | PD | PD | PD | PU | HIZ | HIZ | HIZ | IN |
| Sleep | PD | PD | PD | PU | HIZ | HIZ | HIZ | IN |
| STBY_RC | OUT | OUT | OUT | OUT | OUT | IN | IN | IN |
| STBY_XOSC | OUT | OUT | OUT | OUT | OUT | IN | IN | IN |
| FS / RX / TX | OUT | OUT | OUT | OUT | OUT | IN | IN | IN |
PU = pull up 50kOhm, PD = pull down 50kOhm, HIZ = high impedance, OUT = output, IN = input.
During Reset, Start-up, and Sleep: MISO is High-Impedance. Any SPI read during these states returns undefined data.
8.5 IRQ Handling (Table 8-4)
| Bit | IRQ | Description | Modulation |
|---|---|---|---|
| 0 | TxDone | Packet transmission completed | All |
| 1 | RxDone | Packet received | All |
| 2 | PreambleDetected | Preamble detected | All |
| 3 | SyncWordValid | Valid Sync Word detected | FSK |
| 4 | HeaderValid | Valid LoRa Header received | LoRa |
| 5 | HeaderErr | LoRa Header CRC error | LoRa |
| 6 | CrcErr | Wrong CRC received | All |
| 7 | CadDone | Channel activity detection finished | LoRa |
| 8 | CadDetected | Channel activity detected | LoRa |
| 9 | Timeout | Rx or Tx Timeout | All |
Note: If DIO2 or DIO3 are used to control the RF Switch or the TCXO, the IRQ is not generated even if it is mapped to the pins.
9. Operational Modes
Operating Modes (Table 9-1)
| Mode | Enabled Blocks |
|---|---|
| SLEEP | Optional registers, backup regulator, RC64k oscillator, data RAM |
| STDBY_RC | Top regulator (LDO), RC13M oscillator |
| STDBY_XOSC | Top regulator (DC-DC or LDO), XOSC |
| FS | All of the above + Frequency synthesizer at Tx frequency |
| TX | Frequency synthesizer and transmitter, Modem |
| RX | Frequency synthesizer and receiver, Modem |
9.1 Startup
At power-up or after a reset, the chip goes into STARTUP state. The BUSY pin is set to high. When the digital voltage and RC clock become available, the chip can boot up and the CPU takes control. At this stage the BUSY line goes down and the device is ready to accept commands.
9.2 Calibration
The calibration procedure is automatically called in case of POR. Blocks calibrated: RC64k, RC13M, PLL, RX ADC, Image. Once calibration is finished, the chip enters STDBY_RC mode.
9.2.1 Image Calibration for Specific Frequency Bands (Table 9-2)
| Frequency Band [MHz] | Freq1 | Freq2 |
|---|---|---|
| 430 - 440 | 0x6B | 0x6F |
| 470 - 510 | 0x75 | 0x81 |
| 779 - 787 | 0xC1 | 0xC5 |
| 863 - 870 | 0xD7 | 0xDB |
| 902 - 928 | 0xE1 (default) | 0xE9 (default) |
By default, the image calibration is made in the 902-928 MHz band. When using a TCXO, the calibration fails and the user should request a complete calibration after calling SetDIO3AsTcxoCtrl(...).
9.7 Transmit (TX) Mode
In TX mode, after enabling and ramping-up the Power Amplifier (PA), the contents of the data buffer are transmitted. The timeout can be used as a security to ensure that if the TxDone IRQ is never triggered, the TxTimeout prevents waiting indefinitely. In TX mode, BUSY goes low as soon as the PA has ramped-up and transmission of preamble starts.
10. Host Controller Interface
10.1 Command Structure (Table 10-1)
| Byte | 0 | [1:n] |
|---|---|---|
| Data from host (MOSI) | Opcode | Parameters |
| Data to host (MISO) | RFU | Status |
During byte 0 (the opcode byte), MISO returns RFU (Reserved for Future Use) -- NOT the status byte. The status byte appears starting at byte 1.
10.2 Transaction Termination
The host terminates an SPI transaction with the rising NSS signal. The host must not raise NSS within the bytes of a transaction. All parameters must be sent before raising NSS.
11. List of Commands
11.1 Operational Mode Commands (Table 11-1)
| Command | Opcode | Parameters | Description |
|---|---|---|---|
| SetSleep | 0x84 | sleepConfig | Set Chip in SLEEP mode |
| SetStandby | 0x80 | standbyConfig | Set Chip in STDBY_RC or STDBY_XOSC mode |
| SetFs | 0xC1 | - | Set Chip in Frequency Synthesis mode |
| SetTx | 0x83 | timeout[23:0] | Set Chip in Tx mode |
| SetRx | 0x82 | timeout[23:0] | Set Chip in Rx mode |
| SetCad | 0xC5 | - | Set chip in RX mode with CAD parameters |
| SetTxContinuousWave | 0xD1 | - | Test command: CW at selected frequency |
| SetRegulatorMode | 0x96 | regModeParam | Select LDO or DC_DC+LDO |
| Calibrate | 0x89 | calibParam | Calibrate RC13, RC64, ADC, PLL, Image |
| CalibrateImage | 0x98 | freq1, freq2 | Image calibration at given frequencies |
| SetPaConfig | 0x95 | paDutyCycle, HpMax, deviceSel, paLUT | Configure PA |
| SetRxTxFallbackMode | 0x93 | fallbackMode | Mode after TX/RX done |
11.2 Register and Buffer Access Commands (Table 11-2)
| Command | Opcode | Parameters |
|---|---|---|
| WriteRegister | 0x0D | address[15:0], data[0:n] |
| ReadRegister | 0x1D | address[15:0] |
| WriteBuffer | 0x0E | offset, data[0:n] |
| ReadBuffer | 0x1E | offset |
11.3 DIO and IRQ Control (Table 11-3)
| Command | Opcode | Parameters |
|---|---|---|
| SetDioIrqParams | 0x08 | IrqMask[15:0], Dio1Mask[15:0], Dio2Mask[15:0], Dio3Mask[15:0] |
| GetIrqStatus | 0x12 | - |
| ClearIrqStatus | 0x02 | ClearIrqParam[15:0] |
| SetDIO2AsRfSwitchCtrl | 0x9D | enable |
| SetDIO3AsTcxoCtrl | 0x97 | tcxoVoltage, timeout[23:0] |
11.4 RF, Modulation and Packet Commands (Table 11-4)
| Command | Opcode | Parameters |
|---|---|---|
| SetRfFrequency | 0x86 | rfFreq[31:0] |
| SetPacketType | 0x8A | protocol |
| SetTxParams | 0x8E | power, rampTime |
| SetModulationParams | 0x8B | ModParam1..8 |
| SetPacketParams | 0x8C | (preamble, header, payload, crc, iq) |
| SetBufferBaseAddress | 0x8F | TX base address, RX base address |
| SetLoRaSymbNumTimeout | 0xA0 | SymbNum |
13. Command Details (Selected)
13.1.2 SetStandby
| Byte | 0 | 1 |
|---|---|---|
| Data from host | 0x80 | StdbyConfig |
StdbyConfig: 0 = STDBY_RC, 1 = STDBY_XOSC.
13.1.4 SetTx
| Byte | 0 | 1-3 |
|---|---|---|
| Data from host | 0x83 | timeout[23:0] |
- Starting from STDBY_RC mode, the oscillator is switched ON followed by PLL, then the PA ramps up.
- When the last bit has been sent, an IRQ TX_DONE is generated, the PA ramps down, and the chip goes back to STDBY_RC mode.
- A TIMEOUT IRQ is triggered if TX_DONE is not generated within the timeout period.
- Timeout duration = Timeout * 15.625us
- Timeout = 0x000000: No timeout, device stays in TX until packet is transmitted and returns to STBY_RC.
13.1.5 SetRx
| Byte | 0 | 1-3 |
|---|---|---|
| Data from host | 0x82 | timeout[23:0] |
| Timeout | Duration |
|---|---|
| 0x000000 | Single mode: stays in RX until reception, then returns to STBY_RC |
| 0xFFFFFF | Continuous mode: remains in RX until host sends a mode change command |
| Others | Timeout active: returns to STBY_RC on timeout or reception. Max timeout is 262s. |
13.1.11 SetRegulatorMode
| Byte | 0 | 1 |
|---|---|---|
| Data from host | 0x96 | regModeParam |
regModeParam: 0 = Only LDO, 1 = DC_DC+LDO (used for STBY_XOSC, FS, RX and TX modes).
13.1.12 Calibrate Function
| Byte | 0 | 1 |
|---|---|---|
| Data from host | 0x89 | calibParam |
calibParam is a bitmask: Bit 0=RC64k, 1=RC13M, 2=PLL, 3=ADC pulse, 4=ADC bulk N, 5=ADC bulk P, 6=Image. 0x7F = calibrate all. Total calibration time ~3.5ms. BUSY is high during calibration.
13.1.14 SetPaConfig
| Byte | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| Data from host | 0x95 | paDutyCycle | hpMax | deviceSel | paLut |
- deviceSel: 0 = SX1262, 1 = SX1261
- paLut: reserved, always 0x01
- For SX1262, paDutyCycle should not be higher than 0x04.
PA Optimal Settings (Table 13-21)
| Output Power | paDutyCycle | hpMax | deviceSel | paLut | SetTxParams power |
|---|---|---|---|---|---|
| +22dBm | 0x04 | 0x07 | 0x00 | 0x01 | +22dBm |
| +20dBm | 0x03 | 0x05 | 0x00 | 0x01 | +22dBm |
| +17dBm | 0x02 | 0x03 | 0x00 | 0x01 | +22dBm |
| +14dBm | 0x02 | 0x02 | 0x00 | 0x01 | +22dBm |
We do not drive the PA from this table. These four rows are the most
efficient pairing for four particular outputs, not the way an output is
selected: every row leaves SetTxParams at +22 and lets the PA row set the
power, which reaches exactly these four points and nothing between or below
them. leviculum_core::sx126x::program_tx_power instead writes the +22 row
once (PA_CONFIG_HIGH_POWER) and passes the configured power to
SetTxParams, clamped to -9..=22 — the whole range, and the same shape the
RNode firmware uses (reference/RNode_Firmware/sx126x.cpp:714-735), so an
LNode and an RNode configured to the same number radiate the same. The cost
is PA efficiency, i.e. supply current, at the three lower points; the benefit
is that the other 28 points exist at all (Codeberg #349).
13.1.15 SetRxTxFallbackMode
| Byte | 0 | 1 |
|---|---|---|
| Data from host | 0x93 | fallbackMode |
| Fallback Mode | Value | Description |
|---|---|---|
| FS | 0x40 | Go to FS mode after TX/RX |
| STDBY_XOSC | 0x30 | Go to STDBY_XOSC after TX/RX |
| STDBY_RC | 0x20 | Go to STDBY_RC after TX/RX (default) |
13.2.1 WriteRegister
| Byte | 0 | 1 | 2 | 3 | ... | n |
|---|---|---|---|---|---|---|
| MOSI | 0x0D | addr[15:8] | addr[7:0] | data@addr | ... | data@addr+(n-3) |
| MISO | RFU | Status | Status | Status | ... | Status |
13.2.2 ReadRegister
| Byte | 0 | 1 | 2 | 3 | 4 | ... |
|---|---|---|---|---|---|---|
| MOSI | 0x1D | addr[15:8] | addr[7:0] | NOP | NOP | ... |
| MISO | RFU | Status | Status | Status | data@addr | ... |
Note: The host must send an NOP after the 2 bytes of address to start receiving data bytes on the next NOP sent.
13.2.3 WriteBuffer
| Byte | 0 | 1 | 2 | ... | n |
|---|---|---|---|---|---|
| MOSI | 0x0E | offset | data@offset | ... | data@offset+(n-2) |
| MISO | RFU | Status | Status | ... | Status |
13.2.4 ReadBuffer
| Byte | 0 | 1 | 2 | 3 | ... |
|---|---|---|---|---|---|
| MOSI | 0x1E | offset | NOP | NOP | ... |
| MISO | RFU | Status | Status | data@offset | ... |
Note: An NOP must be sent after sending the offset.
13.3.1 SetDioIrqParams
| Byte | 0 | 1-2 | 3-4 | 5-6 | 7-8 |
|---|---|---|---|---|---|
| Data from host | 0x08 | IrqMask[15:0] | DIO1Mask[15:0] | DIO2Mask[15:0] | DIO3Mask[15:0] |
The interrupt causes a DIO to be set if the corresponding bit in DioxMask AND IrqMask are both set. For example, to route TxDone to DIO1: set bit 0 of both IrqMask and DIO1Mask.
13.3.3 GetIrqStatus
| Byte | 0 | 1 | 2-3 |
|---|---|---|---|
| MOSI | 0x12 | NOP | NOP |
| MISO | RFU | Status | IrqStatus[15:0] |
13.3.4 ClearIrqStatus
| Byte | 0 | 1-2 |
|---|---|---|
| MOSI | 0x02 | ClearIrqParam[15:0] |
13.3.5 SetDIO2AsRfSwitchCtrl
| Byte | 0 | 1 |
|---|---|---|
| MOSI | 0x9D | enable |
enable=1: DIO2 controls RF switch. DIO2=1 during TX, DIO2=0 otherwise.
13.3.6 SetDIO3AsTcxoCtrl
| Byte | 0 | 1 | 2-4 |
|---|---|---|---|
| MOSI | 0x97 | tcxoVoltage | delay[23:0] |
tcxoVoltage (Table 13-35)
| Value | Output Voltage |
|---|---|
| 0x00 | 1.6V |
| 0x01 | 1.7V |
| 0x02 | 1.8V |
| 0x03 | 2.2V |
| 0x06 | 3.0V |
| 0x07 | 3.3V |
Delay duration = delay[23:0] * 15.625us
The XOSC_START_ERR flag is raised at POR or wake-up from Sleep in cold-start condition when TCXO is used. This is expected and should be cleared with ClearDeviceErrors.
Note: The user should take the delay period into account when going into Tx or Rx mode from STDBY_RC mode, since the time needed to switch modes increases with the duration of delay.
13.4.1 SetRfFrequency
| Byte | 0 | 1-4 |
|---|---|---|
| MOSI | 0x86 | RfFreq[31:0] |
RF_frequency = RF_Freq * F_XTAL / 2^25, where F_XTAL = 32 MHz.
To compute RF_Freq from Hz: RF_Freq = freq_hz * 2^25 / 32_000_000
13.4.4 SetTxParams
| Byte | 0 | 1 | 2 |
|---|---|---|---|
| MOSI | 0x8E | power | RampTime |
power: -9 to +22 dBm (encoded as 0xF7 to 0x16) for high power PA (SX1262).
| RampTime | Value | Time (us) |
|---|---|---|
| SET_RAMP_10U | 0x00 | 10 |
| SET_RAMP_20U | 0x01 | 20 |
| SET_RAMP_40U | 0x02 | 40 |
| SET_RAMP_80U | 0x03 | 80 |
| SET_RAMP_200U | 0x04 | 200 |
| SET_RAMP_800U | 0x05 | 800 |
13.4.5 SetModulationParams (LoRa)
| Byte | 0 | 1 | 2 | 3 | 4 | 5-8 |
|---|---|---|---|---|---|---|
| MOSI | 0x8B | SF | BW | CR | LdOpt | unused (0x00) |
- ModParam1 = SF (Spreading Factor)
- ModParam2 = BW (Bandwidth)
- ModParam3 = CR (Coding Rate)
- ModParam4 = LdOpt (Low Data Rate Optimization)
13.5.1 GetStatus
| Byte | 0 | 1 |
|---|---|---|
| MOSI | 0xC0 | NOP |
| MISO | RFU | Status |
Status Byte Format (Table 13-76)
| Bit 7 | Bits 6:4 | Bits 3:1 | Bit 0 |
|---|---|---|---|
| Reserved | Chip mode | Command status | Reserved |
Chip mode:
| Value | Mode |
|---|---|
| 0x0 | Unused |
| 0x2 | STBY_RC |
| 0x3 | STBY_XOSC |
| 0x4 | FS |
| 0x5 | RX |
| 0x6 | TX |
Command status:
| Value | Meaning |
|---|---|
| 0x0 | Reserved |
| 0x2 | Data is available to host |
| 0x3 | Command timeout |
| 0x4 | Command processing error |
| 0x5 | Failure to execute command |
| 0x6 | Command TX done |
13.5.2 GetRxBufferStatus
| Byte | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| MOSI | 0x13 | NOP | NOP | NOP |
| MISO | RFU | Status | PayloadLengthRx | RxStartBufferPointer |
13.5.3 GetPacketStatus (LoRa)
| Byte | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| MOSI | 0x14 | NOP | NOP | NOP | NOP |
| MISO | RFU | Status | RssiPkt | SnrPkt | SignalRssiPkt |
- Actual signal power = -RssiPkt/2 (dBm)
- Actual SNR = SnrPkt/4 (dB)
15. Known Limitations
15.1 Modulation Quality with 500kHz LoRa Bandwidth
Before any packet transmission, bit #2 at register address 0x0889 shall be set to:
- 0 if the LoRa BW = 500kHz
- 1 for any other LoRa BW or (G)FSK configuration
Must be applied before each packet transmission.
15.2 Better Resistance to Antenna Mismatch (TX PA Clamp)
During chip initialization on the SX1262, the register TxClampConfig at address 0x08D8 should be modified. Bits 4-1 must be set to "1111" (default value "0100").
value = ReadRegister(0x08D8)
value = value | 0x1E
WriteRegister(value, 0x08D8)
Must be done after POR or wake-up from cold start.
15.3 Implicit Header Mode Timeout Behavior
After ANY Rx with Timeout active sequence, stop the RTC and clear the timeout event:
WriteRegister(0x00, 0x0902)
value = ReadRegister(0x0944)
value = value | 0x02
WriteRegister(value, 0x0944)
15.4 Optimizing the Inverted IQ Operation
Bit 2 at address 0x0736 must be set to:
- "0" when using inverted IQ polarity
- "1" when using standard IQ polarity
Key Register Addresses
| Address | Name | Description |
|---|---|---|
| 0x0740 | LoRaSyncword | LoRa sync word (2 bytes, MSB first). 0x1424=private, 0x3444=public. |
| 0x0889 | TxModulation | BW500 workaround (bit 2) |
| 0x08AC | RxGain | 0x94=power saving (default), 0x96=boosted gain |
| 0x08D8 | TxClampConfig | PA clamp workaround (bits 4:1 = 0xF) |
| 0x0736 | IqPolarity | Inverted IQ workaround (bit 2) |
| 0x0902 | RtcControl | RTC stop (write 0x00 after Rx with timeout) |
| 0x0944 | EventMask | Clear timeout event (bit 1) |
| 0x029F | RetentionList count | Number of retention registers |
| 0x02A0-0x02A1 | RetentionList[0] | First retention register address (0x08AC) |
Init Sequence Summary (from Semtech reference driver + RNode)
- Hardware reset (NRESET LOW 100us, HIGH, wait BUSY LOW)
- SetStandby(STBY_RC)
[0x80, 0x00] - SetRegulatorMode(DC_DC)
[0x96, 0x01] - SetDIO2AsRfSwitchCtrl(enable)
[0x9D, 0x01] - ClearDeviceErrors
[0x07, 0x00, 0x00] - SetDIO3AsTcxoCtrl(1.8V, timeout)
[0x97, 0x02, t2, t1, t0] - Calibrate(all)
[0x89, 0x7F]— wait BUSY LOW (~3.5ms) - CalibrateImage(863-870MHz)
[0x98, 0xD7, 0xDB] - SetPacketType(LoRa)
[0x8A, 0x01] - SetRfFrequency(freq)
[0x86, f3, f2, f1, f0] - SetPaConfig(0x04, 0x07, 0x00, 0x01) for +22dBm SX1262
- SetTxParams(power, ramp)
[0x8E, pwr, ramp] - SetBufferBaseAddress(0, 0)
[0x8F, 0x00, 0x00] - SetModulationParams(SF, BW, CR, LDRO)
[0x8B, sf, bw, cr, ldro, 0,0,0,0] - SetPacketParams(preamble, header, len, crc, iq)
[0x8C, ...] - Write LoRa sync word to register 0x0740
- Apply workarounds: TxClamp (0x08D8), BW500 (0x0889), IQ (0x0736)
- Set RxGain to 0x96 (boosted) at register 0x08AC
TX Sequence
- SetDioIrqParams(TxDone|Timeout on DIO1)
[0x08, mask_hi, mask_lo, dio1_hi, dio1_lo, 0,0, 0,0] - ClearIrqStatus(all)
[0x02, 0xFF, 0xFF] - WriteBuffer(0, payload)
[0x0E, 0x00, data...] - SetTx(timeout)
[0x83, t2, t1, t0]— timeout=0 for no timeout - Wait for DIO1 HIGH (TxDone IRQ)
- ClearIrqStatus(TxDone)
[0x02, 0x00, 0x01]
RX Sequence
- SetDioIrqParams(RxDone|Timeout|CrcErr on DIO1)
- ClearIrqStatus(all)
[0x02, 0xFF, 0xFF] - SetRx(timeout)
[0x82, t2, t1, t0] - Wait for DIO1 HIGH
- GetIrqStatus — check RxDone vs Timeout vs CrcErr
- GetRxBufferStatus
[0x13, ...]— get length + start pointer - ReadBuffer(start, length)
[0x1E, start, NOP, data...] - GetPacketStatus
[0x14, ...]— get RSSI, SNR - ClearIrqStatus
- Apply workaround 15.3 (stop RTC after Rx with timeout)
Structured event logs
Test-harness scaffolding (Codeberg #39 piece 1, Stage 6) for capturing mesh-protocol events as parseable lines so multi-node failures can be diagnosed from a single merged log instead of N hand-correlated process traces.
Format
Each emitted event renders to a single line:
EVENT_NAME node=<n> key1=val1 key2=val2 ... t=<rel-ms>
Rules:
EVENT_NAMEfirst. Comes from the literal string passed as theeventfield in atracing::debug!call.node=second. Value comes fromLEVICULUM_EVENT_NODEenvironment variable; defaults tolocal.- All other keys appear alphabetically sorted between
node=andt=. t=last. Millisecond offset from layer registration time.
Records that don't carry an event = "..." field are silently
dropped, so the legacy printf-style tracing::debug!("[FOO] ...")
sites stay valid alongside the converted ones.
Per-packet journey contract
The packet-level events PKT_TX, PKT_RX, PKT_FORWARD, PKT_DROP
and DEDUP_DROP form the journey contract an external collector uses
to stitch one packet's path across nodes:
-
They are emitted on the dedicated tracing target
leviculum_core::pkt(DEBUG), so a collector can enable exactly this stream viaRUST_LOG=leviculum_core::pkt=debugwithout the rest of the transport noise. The event-log layer sees every record regardless of target. -
Each carries
ph, the first 16 hex chars of the dedup packet hash (SHA-256 over the hashable part, which stripshopsandtransport_id).phis therefore stable across hops and across Type1/Type2 header conversion: the same value appears in the sender'sPKT_TX, every relay'sPKT_RX/PKT_FORWARDand the receiver'sPKT_RX, or in thePKT_DROP/DEDUP_DROPwhere the packet died. -
PKT_DROPrenders itsreasonas the kebab-caseDropReason(no-path,plain-group-multihop,forward-max-hops, ...).unknown-contextis the one reason that says nothing about the packet's validity: it means the packet was addressed to US and carries a context byte this build assigns no meaning to, so nothing above transport could interpret it. The same packet addressed to someone else is relayed normally and never reaches this counter — the context byte is semantic, not routing information. A risingunknown-contexton a node that is also an endpoint means a peer speaks a dialect (newer RNS, third implementation) we do not. -
no-such-interfaceis the onePKT_DROPon the OUTBOUND path, and the one with a different field set: the action was routed to an interface the driver's dispatch slice does not contain, so it never became a received packet and carriesiface_outandleninstead ofdst,typeandiface_in. Non-zero means the driver and the core disagree about interface numbering — a configuration fault in the driver, not a mesh condition, which is why it does not share a counter withno-path. -
A relay whose outbound path points back out of the arrival interface forwards there — same-interface relay on a shared medium is a normal hop, not a drop (see Python-RNS Compatibility). Its
PKT_FORWARDcarriesiface_outequal toiface_in. -
PKT_TXon aBroadcastaction reportsiface=bcast: the sans-I/O core does not know the concrete interface set the driver expands the broadcast to; journeys stitch byph. -
PKT_TXandPKT_RXboth carryhops, but they are counted at different points and the difference is the contract:PKT_TX hopsis the hop count of the packet as transmitted — the byte that goes on the wire, read from the packed buffer being handed to the driver.PKT_RX hopsis the hop count after receipt, i.e. after the receiver's increment (Transport::incoming_hop_count, mirroring PythonTransport.py:1457).
So for one
phcrossing one radio hop,rx hops = tx hops + 1. A collector reads a journey's DIRECTION from exactly that relation: the node observing the packet at the lower hop count transmitted it, the node at one more heard that transmission. WithouthopsonPKT_TX, a node that ORIGINATES a packet (a firmware node, or a daemon's own announces, path requests and link proofs) contributes no hop count at all and drops out of that relation.The relation is deliberately +1 only for a real medium crossing. Over the local-IPC hop — a
LocalClientinterface, or the uplink to a shared instance —incoming_hop_countundoes its own increment, so thererx hops = tx hops. That is Python's behaviour (Transport.py:1481-1484) and it is correct: the IPC hop is not a network hop. A collector pairing on+1therefore ignores IPC hops, which is what it should do. -
The hash is never computed twice for one packet: emission sites reuse the dedup/cache hash where it exists and otherwise hash only while the
leviculum_core::pkttarget is enabled. With the target disabled the whole contract is zero-cost. -
Deliberate exclusions: the high-volume overheard drop (
overheard-transport-id) and IFAC drops stay counter-only (PKT_DROP_SUMMARY); announce-pipeline drops (replay, rate-limit, ingress-burst, over-max-hops, blackhole) are covered by the announce event family and the summary counters.
Architecture
All test threads, including tokio multi-thread workers, route through the same global subscriber registered once via Layer composition.
Specifically: tracing_setup::init_tracing_with_event_log() builds a
Registry::default().with(fmt_layer).with(event_log_layer) chain
and installs it via set_global_default once per process (Once-
guarded). Every thread, every spawned future, every tokio worker
inherits this global subscriber. This is the load-bearing
architectural choice that lets a #[tokio::test(multi_thread)] mvr
see events emitted from worker threads.
Per-test buffer isolation is built on top: init_event_log()
returns an EventLogHandle whose Arc<Mutex<Vec<String>>> buffer
is registered in the layer's active-handles list. The layer's
on_event iterates the active list and pushes the formatted line
to every active buffer. When the test's handle drops, it removes
itself from the list.
Concurrency consequence: every active buffer receives every event
the layer sees, regardless of which test emitted it. Tests that
assert on buffer contents must filter by event name to avoid
cross-test pollution. Use disjoint event names per test
(EV_BASIC, EV_VIOLATION, …); mvr tests already enforce
--test-threads=1 so this only affects unit tests.
How to wire a test
Inside any test, before the test body runs:
#![allow(unused)] fn main() { let _evlog = leviculum_std::test_support::event_log::init_event_log(); }
The handle is RAII: when the binding goes out of scope it removes
itself from the layer's active-handles list. If the test thread
is panicking at drop-time, the buffer dumps to stderr with a
=== EVENT LOG DUMP … banner that cargo test surfaces in the
failure listing.
Use init_event_log_to_file(path) instead of init_event_log()
when the test wants to assert on the dumped content directly
(std::fs::read_to_string(path)).
To make the test fail loud on undocumented schema gaps, end the test body with:
#![allow(unused)] fn main() { leviculum_std::assert_no_schema_violations!(_evlog); }
It panics if any EVENT_SCHEMA_VIOLATION line appears in the
buffer.
How to add an event
Two steps, both in the same commit:
-
Convert the call site. Replace the printf-style
tracing::debug!with structured fields:#![allow(unused)] fn main() { tracing::debug!( event = "FOO", iface = %iface_name, dst = %HexShort(&dst_hash), hops = packet.hops, len = bytes.len(), ); }%for Display,?for Debug. Values must be ASCII without whitespace,=, or non-printable characters — otherwise the field-value validator fires (see below). For Rust keywords liketype, use the raw identifierr#type.No trailing message.
tracing::debug!(event = "FOO", a = 1, "some prose")renders the prose under amessagefield whose spaces split the line for every token-based parser. What the sentence would have said belongs in a structured field or in a comment at the call site. -
Add a catalogue entry in
leviculum-std/src/event_log.rs'sEVENT_CATALOG:#![allow(unused)] fn main() { EventSchema { name: "FOO", required_keys: &["iface", "dst", "hops", "len"], }, }required_keysis the INTERSECTION of the keys the name's call sites set — the contract is "present on every emission". Where sites differ in a way worth checking, the name gets several entries and a record passes if any one shape is fully present (Codeberg #320); note that a shorter shape dominates a longer one that merely extends it, so a second entry only earns its place when neither shape contains the other. The subscriber checks that every catalogued event's emission satisfies some declared shape; a record that satisfies none produces aEVENT_SCHEMA_VIOLATIONline in the dumped buffer alongside the original event.
Step 2 is not on the honour system. leviculum-std's
#[cfg(test)] mod event_catalog_completeness walks the
tracing::*!(event = "...") sites of every workspace member's src/
and fails on a name EVENT_CATALOG is missing, and on a site that
passes a message argument. It runs under cargo test --workspace --lib, which is what just fast and the forge gate run.
It exists because the honour system had already failed: the miauhaus
soak of 2026-09-18 (397 023 881 events) found eight emitted names
undeclared, LINK_ENTRY_SET among them at 612 639 emissions. An
undeclared name is not cosmetic — the layer validates a shape only for
a name it finds in the catalogue, so an undeclared event can lose a
required field forever and nothing says a word. What caught
LINK_ENTRY_SET's broken next_hop, 252 669 times, was the
field-VALUE check, which runs regardless of the catalogue.
Catalogue entries without a live emitting site are explicitly
discouraged: the runtime-validation layer can't detect them, so
they silently rot. Only add entries you have a corresponding
emit for. (That direction is still unchecked, and the catalogue
carries entries whose emitter lives outside this workspace:
SILENCE_LNODE_ENTER/SILENCE_LNODE_EXIT are periculum's.)
Firmware-side events
The nRF firmware emits the same line grammar, but not through this
machinery. leviculum-nrf is no_std and cannot depend on
leviculum-std, so there is no tracing subscriber, no node=
field, and no runtime schema validation; the line goes into the
debug-CDC ring buffer behind the module prefix that the rest of the
firmware log uses:
[BLE ] BLE_TX_DROP kind=packet len=312 frag=1 of=2 sent=1 reason=stalled code=0 waits=0 dropped=3 t=48210
The prefix does not disturb the grammar — the event name is still
one whitespace-delimited token and grep BLE_TX_DROP over a
captured debug-port log still works — but the event deliberately
does not appear in EVENT_CATALOG. A catalogue entry the
subscriber can never see emitted is exactly the silent rot the rule
above forbids. Firmware events are documented here and at their
call site instead.
Who writes t= on a firmware line
Nobody at the call site. Since #344 the firmware's log formatter
(leviculum-nrf/src/log.rs, shape in leviculum-log-line) appends
t=<uptime-ms> to every line it emits — event lines, plain
[LORA]/[INFO] lines, boot banners, tracing records alike. A
call site that also writes its own t= renders the field twice;
BLE_TX_DROP did, and stopped.
The stamp is taken when the line is formatted, not when it is drained: the ring is emptied in 64-byte USB packets on a 100 ms loop, so host arrival times measure that loop and nothing else.
A line may still legitimately carry two t= fields. The boot
replay of the persistent tail wraps a line from the previous boot,
its stamp included, inside a line of this boot:
[INFO!] [PERSISTENT_LOG] [LORA] RX 41 bytes t=91422 t=137
Both are true. A line's own stamp is always its last t= —
the same rule merge_event_logs already applies, so the two agree.
Epochs do not. A firmware t= is milliseconds of board uptime;
a host-side t= is milliseconds since that subscriber's init.
Merging a debug-port capture into a host event log with
merge_event_logs therefore orders each stream correctly within
itself and says nothing across the two.
Current firmware events:
| Event | Emitted by | Meaning |
|---|---|---|
BLE_TX_DROP | leviculum-nrf/src/ble/notify.rs | A BLE packet or keepalive was abandoned part-way through its fragments. frag= is the fragment that failed, sent= how many did go out, reason= one of stalled (no HVN-TX-COMPLETE within the bound), disconnected, budget, sd_error (with the raw code in code=), internal. conn= is the SoftDevice connection handle the TX targeted — the fan-out sends one copy per live link, and without this key the desk log of 2026-09-08 could not say WHICH link ate an sd_error (#365). dropped= is the cumulative counter, so one line states both the incident and the running total. |
BLE_TX_PKT | leviculum-nrf/src/ble/notify.rs (notify pump) and leviculum-nrf/src/ble/columba.rs (central write loop); line rendered by leviculum-ble-tx's TxPktLine, verbatim-pinned by its host test | One line per multi-fragment packet handed to a link, whatever became of it (#373). conn= the SoftDevice connection handle, len= the whole packet, frags= how many fragments it split into, sent= how many the stack accepted. A healthy hand-over reads sent==frags; sent<frags is a loss on this node and is always accompanied by a BLE_TX_DROP naming the reason. The desk log that motivated it had two of five relayed two-fragment packets vanish on the BLE hop with no line anywhere — BLE_TX_DROP only fires on an abandoned packet, so a fully-accepted packet that still never arrived left nothing to grep. Single-fragment packets and keepalives stay unlogged (they are the bulk of the traffic and the failure mode needs ≥2 fragments). |
BLE_RX_ABANDON | leviculum-nrf/src/ble/columba.rs (both GATT roles) | A link's reassembly was discarded before completion (#373): a new START over an unfinished head, a total contradicting the reassembly in progress, or the hard reset after a garbage frame. Each is one or more whole Reticulum packets this receiver lost — previously a silent state reset. slot= the link's drain slot, lost= what this frame cost, total= the link's running count (BleDefragmenter::abandoned_count). lnsd's Columba interface emits the same event name identity-keyed (see the host-side section below), so a merged bench timeline carries both receivers. |
BLE_TX_RESYNC | leviculum-nrf/src/ble/notify.rs | The connection was dropped deliberately after a torn fragment stream, because the wire protocol has no abort marker and a reconnect is the only in-band reset of the peer's reassembler (#255). |
BLE_TX_GAP | leviculum-nrf/src/ble/columba.rs (both pumps) | The per-link inter-packet gap deferred a packet (#376): conn= the connection handle, waited_ms= how long the packet's first fragment was held back after the previous packet's last. The gap defaults to 100 ms (leviculum-ble-tx's DEFAULT_TX_GAP_MS, the measured desk value); lnflash --set-ble-tx-gap overrides it for measurement, 0 disables. On a paced link with real traffic this line is the expected signature; its absence under back-to-back traffic means the knob was set to 0. lnsd emits the same event name on its notify pipe (link=notify) and central links (addr=). |
BLE_TX_HELD | leviculum-nrf/src/ble/columba.rs (peripheral pump) | The drain held a packet because the peer cannot receive yet (#376): conn=, reason=not-subscribed. Once per connection, however many packets wait. Background: a notify before the central writes the TX CCCD fails with sd_error code=13313 (BLE_ERROR_GATTS_SYS_ATTR_MISSING) and the packet dies — the field T114 lost the first packet of a fresh connection this way at 11:31:18 on 2026-09-09. Held packets wait in the link's queue and drain after the subscription (or the handshake, whichever is later); policy host-tested in leviculum-ble-tx's hold module. |
[ANNOUNCE] | leviculum-nrf/src/announce.rs, and the TYPE_ANNOUNCE arm of each binary | Frozen shape — the desk recipe reads this. When the board announced one of its destinations, and why (#376, #384). [ANNOUNCE] sent dst=<hex8> reason=<r>; dst says WHICH destination and reason says what occasioned it. For lxmf.delivery: reason=peer-up peer=<hex8> (a BLE peer finished its identity handshake; the announce goes on that peer's link alone and appears as BLE_TX_ROUTE beside it, never as BLE_TX_FLOOD), reason=periodic (the timer, every 30 minutes, on every interface, so it is a BLE_TX_FLOOD) or reason=host (lnflash --announce, periculum's announce_board). For lxmf.propagation on a board running the role: reason=pn-periodic (the role's own 300 s interval) or reason=pn-host (the same host command — one TYPE_ANNOUNCE frame produces both lines, because a client needs the propagation destination to address the mailbox and the delivery announce does not carry it). The telemetry path's own pre-report announce is not named here; it is the broadcast that precedes a [TELEMETRY] send line. [ANNOUNCE] withheld reason=no-clock is the clock gate — an announce stamped from uptime can never replace a path at the receiver, so a board without a plausible wall clock says why it is silent instead of poisoning path tables; written once per change, not once per retry. reason=rate-limited peer=<hex8> under the [BLE ] prefix is the per-identity 15-minute limit, which is why a phone rotating its BLE address every minute does not buy an announce every minute. lnsd emits the same two occasions as the trace events ANNOUNCE_TX reason=peer-up peer= iface= count= and ANNOUNCE_WITHHELD reason= peer= iface=. |
BLE_TX_ROUTE / BLE_TX_FLOOD / BLE_TX_ROUTE_MISS | leviculum-nrf/src/ble/mod.rs (the fan-out task) | What the core's per-packet delivery hint made of one outbound packet (#376). Exactly one of the three per packet, so a capture accounts for everything the interface was handed. BLE_TX_ROUTE peer=<hex8> conn=<h> slot=<n> len=<n> — the packet was addressed at one peer and went on that peer's link alone. The core takes the peer from the path it routed over (via_peer, Codeberg #365), or, for a proof, from the arrival it is answering: the ingress peer travels with the deferred ProofRequested event, so a probe's proof leaves as BLE_TX_ROUTE and not as a flood. This is what stops a report for the phone from also travelling to the board beside it, which used to forward it back and give the phone two copies. BLE_TX_FLOOD links=<n> len=<n> — no hint, so a broadcast: announces and path requests still reach every live link. BLE_TX_ROUTE_MISS peer=<hex8> len=<n> — the addressed peer holds no live link here and the packet is dropped, NOT flooded (flooding would spend the other links' airtime on a packet they cannot deliver and rebuild the relayed duplicate); its running total is route_miss= on BLE_COUNTERS, and a rising value means the path table outlived a link the #365 cull should have taken. lnsd emits the same three names with iface= and, on BLE_TX_ROUTE, conn=central|peripheral instead of a SoftDevice handle. |
BLE: RX | leviculum-nrf/src/ble/columba.rs (both GATT roles) | One line per reassembled inbound packet: BLE: RX <n>B conn=<h> frags=<k> — conn= tells the phone's link from the neighbour board's, frags= how the PEER fragmented the packet (BleDefragmenter::last_completed_fragments), which is the only place a peer's real fragment size is visible (#376: Columba as peripheral claimed "MTU 20 bytes" while our central negotiated a large ATT MTU). The binaries' former BLE RX <n> bytes line was dropped for it: one reception, one line. |
BLE_CONN_PARAMS | leviculum-nrf/src/ble/columba.rs (both roles); line rendered by leviculum-ble-tx's ConnParamsLine, byte-pinned by its host tests | What a link actually runs at (#385). One line per link, at the connection event, beside BLE: connected on the peripheral side and beside the successful dial on the central one: BLE_CONN_PARAMS conn=<h> role=central|peripheral interval_ms=<n.nn> latency=<n> timeout_ms=<n>. Neither stack requests connection parameters, so these are the central's choice inherited whole — and the supervision timeout among them is exactly how long a radio disturbance may last before the link dies (the #385 bench: a board and lnsd one metre apart lost their link 36 times in 16.2 hours, every connection at a 45 ms interval with a 420 ms supervision timeout, every death HCI reason 0x08). Both time fields are printed in milliseconds, converted from the two different raw scales the controller reports (interval in 1.25 ms steps, hence the two decimals; supervision timeout in 10 ms steps), so no reader of a field log has to remember which scale belongs to which field. latency= is a count of skippable connection events, not a time, and is printed unconverted. A peripheral link emits the line a second time when it ends, marked when=close and carrying the same conn= (#385): the SoftDevice updates the connection's stored parameters on BLE_GAP_EVT_CONN_PARAM_UPDATE without handing the event to application code (nrf-softdevice's gap.rs), so a renegotiation cannot be logged as it happens — but the stored copy survives the disconnect, so re-reading it at teardown states the values the link actually ended on. That pair is the only evidence of what a central did with the board's update request (below): opened 420 ms, closed 4000 ms means honoured; closed 420 ms means not. A central link emits the open line only. lnsd cannot emit this event at all; see the host-side section below. |
BLE_CONN_PARAMS_REQ | leviculum-nrf/src/ble/columba.rs (peripheral role only); line rendered by leviculum-ble-tx's ConnParamsReqLine, decision in its judge_supervision_timeout, both byte- and boundary-pinned by host tests | Whether the board asked its central for a supervision timeout it can survive, and what came of asking (#385). One line per peripheral link, right after the link's BLE_CONN_PARAMS: BLE_CONN_PARAMS_REQ conn=<h> timeout_ms=<n> result=sent|refused|skipped. The board is the peripheral in exactly the two cases that are bad or unknown — an lnsd central gives it BlueZ's 420 ms, a phone gives it something unmeasured — and the host cannot set these values through the API lnsd uses, so the peripheral asking is the only lever. The rule is conditional: below a 2000 ms floor the link asks for 4000 ms (the same value ConnectConfig::default asks for in the central role, so the two roles agree), at or above the floor it asks for nothing, because an update request that fights an already-good value is a regression. timeout_ms= is the value asked for on sent/refused and the value that passed the floor on skipped. result=sent says only that the request left the board — on a peripheral it is an L2CAP connection parameter update request, which a central may honour, ignore, or answer with something else entirely, and only the when=close line says which. refused is a local refusal by the SoftDevice and is not retried. Board-to-board links never emit anything but skipped: their central already asks for 4 s. No lnsd counterpart, for the same BlueZ reason as BLE_CONN_PARAMS. |
BLE_DRAIN_TABLE_FULL | leviculum-nrf/src/ble/columba.rs | A connection could not claim a per-connection HVN drain slot; slots= is the table size. Expected never: it means more live connections than ble::MAX_LINKS. |
BLE_GATT_WRITE_OVERSIZE / BLE_GATT_NOTIFY_OVERSIZE | leviculum-nrf/src/ble/columba.rs (peripheral write path / central notification path); line rendered by leviculum-ble-tx's OversizeLine, verbatim-pinned by its host test | An inbound GATT value exceeded the characteristic bound and was dropped whole (#387): conn= the SoftDevice connection handle, len= the peer's wire length, max= the bound (GATT_VALUE_MAX, 253 = our ATT MTU grant of 256 − 3). Truncating instead of dropping would hand the defragmenter a cut Columba fragment that completes a reassembly with garbage. Before #387 this input was not a line but a board panic — the vendored GattValue conversion unwraps on length (nrf-softdevice gatt_traits.rs:97), and the field T114 died on it twice on 2026-09-12; the same event also disproved the assumption that the SoftDevice rejects over-max_len writes before an event exists. Running total is oversize= on BLE_COUNTERS. Expected never from a healthy peer: a full-MTU Columba fragment is exactly 253 bytes and is accepted (the old width of 251, the link-layer DLE payload, was 2 bytes short of that, which is how a healthy phone panicked the board). |
[MEDIA] | leviculum-nrf/src/media.rs | Frozen shape — assertions read this. Which carriers this node meshes over: lora=on|off ble=on|off src=default|flash. The two carrier fields are what the board is running (a carrier configured on but not started this boot reads off here, which is the honest answer), and src= says whether the profile came off the flash page or from the both-on default. Emitted once per boot after both spawn decisions, then re-emitted with [FW_BUILD] every 5 s so a capture attached after the boot window still reads the carriers off the board. A carrier held down by the profile also emits carrier=<lora|ble> state=down reason=profile under the same prefix at the point its bring-up would have been, and packets dropped because their medium is off emit MEDIA_TX_DROP iface=<name> packets=<n> bytes=<n> reason=carrier-off — not one line per packet but on the first drop of a run and at each decade after it (the 1st, 10th, 100th …), because every log line also writes the 2 KiB post-crash tail and a carrier that is off drops one packet per announce; MEDIA_TX_RESUMED iface=<name> packets=<n> bytes=<n> closes the run with its exact totals when the carrier takes a packet again (host tests in leviculum-nrf/media-state). See docs/src/concepts/media-profiles.md. |
BATTERY | leviculum-nrf/src/battery.rs on all three boards; line rendered by leviculum-battery-scale's BatteryLine, byte-pinned by its host tests | What the pack is doing (#380). BATTERY mv=<n> min_mv=<n> max_mv=<n> pct=<n>|none cells=<n>S, once at boot and then every 30 s, under the [BAT] prefix. mv= is the filtered pack voltage — the same number the status panel shows — while min_mv/max_mv are the lowest and highest of the raw 1 Hz samples in the period, which is where a sag under transmit load appears instead of being averaged into the mean. pct= comes from the per-cell LiPo OCV curve, cells= from the classification made on the boot's first reading — and pct=none is the case where the pack voltage is outside the band that classification implies, see BATTERY_PCT below. Before this the module fed the display and said nothing: a Pocket V2 that restarted twice on a 90 minute field walk produced zero battery lines, so how close the pack had been to the edge could not be asked afterwards. It is a margin instrument and not a brownout detector — the sampler is a second apart, a brownout is microseconds wide, and the reset takes the log with it (that walk's every boot came up reset_reason=0x00000000, BOOT_TRACE prev_magic=absent). A first reading outside anything a LiPo pack can be (below 2.5 V or above 9 V — on the T114 the divider reaches 17.7 V and on the Solar Node 10.7 V, so a floating input lands there) is NOT classified as 2S: it falls back to 1S and says so as [WARN] [BAT] implausible first reading. The [BAT] init line beside it carries full_scale_mv=, the board's whole measurable range, which is 6228 on the Pocket V2, 17698 on the T114 and 10660 on the Solar Node, and acq_us=, the SAADC acquisition window the board's divider needs — 10 on the first two and 20 on the Solar Node, whose 1 M∕510 k divider presents 338 kΩ of source where the part specifies 10 µs for 100 kΩ (#233). The two boot lines bypass the RUNTIME_DRAIN_OPEN gate and the periodic ones do not: the gate stays shut until DTR-assert or 30 s of uptime, and a board on battery in a field has no host to assert DTR — which is precisely the run whose first statement is worth keeping. |
BATTERY_PCT | leviculum-nrf/src/battery.rs on all three boards; line rendered by leviculum-battery-scale's BatteryPercentLine, band and boundaries pinned by its host tests | Whether the charge percentage can be believed (#380). BATTERY_PCT reportable=0|1 pack_mv=<n> cells=<n>S band_lo_mv=<n> band_hi_mv=<n>, under the [BAT] prefix, said only when that answer CHANGES — plus once at boot if the answer is already no. The cell count is decided from one reading at boot and held for the boot, and every per-cell voltage after it is the pack voltage divided by it; a percentage carries neither a unit nor the count it was divided by, so a wrong one is indistinguishable from a right one at the far end of a mesh. While the pack voltage stays inside the band its classification implies, pct= carries a number; outside it the board publishes no percentage (pct=none, and no battery sensor in the telemetry report at all) and emits this line once. The band's ceiling is the OCV curve's own 100 % point carried one step of its top segment further (4.33 V per cell), so the guard and the percentage cannot disagree about what a cell is. Its floor is deliberately BELOW the curve's floor, at the 2.5 V per cell protection cut-off: between 2.5 and 3.0 V a pack is nearly empty, which is a real state that must report 0 % rather than go quiet exactly when the battery is about to give out. The voltage is never withheld and the classification is never revised — this declines to build on a boot-time decision, it does not re-take it. Critical rather than drain-gated, like the [BAT] init line: one line per transition is rare, and a field board on battery has no host to open the gate on the run where the gap appears. |
[NAME ] | leviculum-nrf/src/name.rs | What this board is called on each of its two display surfaces: mesh=<name> ble=<name> src=derived|flash. mesh= is the display name the LXMF announces carry (what Columba lists), ble= the GAP/advertised name a scanner sees, and src= whether both come off the flash page or from the names derived from the identity (LNode-<hex8> and LN-<hex8>, which are different strings, not a truncation of one another). The two differ when an operator's name is longer than the BLE bound of 11 bytes — visibly, which is the point. Emitted once per boot beside the [MEDIA] banner, then re-emitted with [FW_BUILD], so a capture says under which names the board is visible without an operator having to remember what they set. Set over the control envelope with lnflash --set-name (docs/src/firmware/usb-control-envelope.md, NODE_NAME 0x0C). |
SD_RAM_FLOOR | leviculum-nrf/src/ble/mod.rs | One line per boot, before Softdevice::enable. wanted= is the app RAM base the S140 says this BLE configuration needs, floor= is ORIGIN(RETAINED) from memory.x (the SoftDevice's ceiling — the retained cross-boot records and then the flip-link stack sit above it), margin= their signed difference. fits=0 never appears — the boot panics instead. |
ADV | leviculum-nrf/src/ble/columba.rs | One line per boot when the advertising payloads are built. adv_bytes=/scan_bytes= are the built PDU sizes against cap=31, peripheral_only= is the v0.3.0 capability bit, periph_links= the number of incoming link slots the payloads serve (#372), and free_slots= how many of them the boot advertisement offers — the live count the record carries in capability bits 1-3 (#375), which is why the advertisement is rebuilt at every advertising start rather than once here. Emitted on the critical log path (like SD_RAM_FLOOR): it fires before the host's DTR-assert opens the runtime drain, and the gated path would silently drop it. |
BLE_SCAN_DECISION | leviculum-nrf/src/ble/columba.rs | The scanner saw a Columba peer and applied the v2.2 address sort with the v0.3.0 capability override. addr= is the peer's current address as 12 hex digits, caps_record= whether a readable capability record was present (caps= is meaningless when 0), free_slots= the DECODED free incoming-slot count the peer advertised or the token unknown when it advertised none (#375 — unknown is not 0: a peer that says nothing is preferred as if all its slots were free, so reading it as zero inverts the conclusion), rule= the decision rule that fired, initiate= whether this side dials. Emitted once per (address, decision) change, not per PDU. The count orders the collection window (most free slots first, then lowest address) and nothing else — no admission or refusal reads it. |
BLE_CENTRAL_*, BLE_LINK_SELF, BLE_LINK_DUP | leviculum-nrf/src/ble/columba.rs | Central-role connection lifecycle: BLE_CENTRAL_ADDR/CONNECT/FAIL/UP/DOWN, plus the two admission decisions. BLE_LINK_SELF addr=<a> action=disconnect — the peer presented our own identity. BLE_LINK_EXPIRE role=peripheral|central slot=<n> conn=<h> silence_ms=<n> — the board's link expiry (#382): this link delivered nothing at all, payload AND keepalives, for LINK_TIMEOUT_MS (45 s), and the session disconnected it. It is the board's counterpart of lnsd's BLE_LINK_DOWN … reason=timeout and the one mechanism that clears a link nobody dials; silence_ms says how far past the bound it ran. BLE_LINK_DUP peer=<hex8> addr=<a> action=refuse rule=<r> origin=incoming|outgoing old_conn=<h> new_conn=<h> old_mtu=<n> new_mtu=<n> old_silence_ms=<n> old_data_silence_ms=<n|never> — that identity already holds a link the decision kept, so the newcomer is dropped (#360 round 2). BLE_LINK_REPLACED peer=<hex8> addr=<a> rule=<r> origin=incoming|outgoing old_conn=<h> new_conn=<h> old_mtu=<n> new_mtu=<n> old_silence_ms=<n> old_data_silence_ms=<n|never> moved=<n> dropped=<n> — the NEWER connection took the peer over and the old link was disconnected by us at the decision; moved=/dropped= are the packets carried over from the old link's queue. One rule behind both lines and rule= says which branch decided: abandoned — the old link delivered nothing at all, keepalives included, for LINK_ABANDONED_MS (30 s, two keepalive intervals), so its peer has walked away from it; same_role — both connections carry the same role (a rotated address dialled twice in one direction), so the peer has no role preference to copy and we decide alone, keeping our own old dial or the peer's new one; columba_mtu / columba_identity — preferred_ble_role, the port of Columba's preferredBleRole, evaluated from the PEER's perspective, deciding on the usable-MTU comparison or on its identity tie-break. old_mtu=/new_mtu= are the usable MTUs as compared, in the peer's ledger (a connection the peer has not bookkept reads MIN_USABLE_MTU, 20). old_silence_ms= (any frame, keepalives included) is the abandonment test's input; old_data_silence_ms was round 1's input and is now reported only — never for a link that carried no payload. A link that has really stopped answering entirely is cleared by BLE_LINK_EXPIRE, dial or no dial. A refusal also emits BLE_DIAL_DEAD_END addr=<a> reason=dup_refused ttl_s=<n>: the address is backed off, or the scanner re-offers it within seconds. Running totals are refused= and displaced= on BLE_COUNTERS; beside a rotating Columba phone displaced= climbs about once per ~90 s rotation and refused= stays near zero. |
[LORA] RX | leviculum-nrf/src/lora.rs; fields read by leviculum_core::packet::peek_wire_class, offsets host-tested in packet.rs | One line per reassembled reception: RX <n> bytes rssi=<dBm> snr=<dB> flags=0x<hh> dst=<hex8> ctx=0x<hh>. flags= is the packet's Reticulum header flags byte (its low two bits are the packet type) and dst= the first 8 hex digits of the destination hash — the two keys lnflash --summarize classifies announce, data, path request and proof from. ctx= is the context byte, and it is the only field that separates the two announce shapes a relay hands down: a relayed announce reads flags=0x51 ctx=0x00, a path response carrying the same announce reads flags=0x51 ctx=0x0b (PacketContext::PathResponse) — identical flags=, identical dst=, and the second is the one Transport::handle_announce keeps out of the announce table by design. Without ctx= a capture cannot tell "the board should have passed this up" from "by design it did not". A packet too short to carry the header its flags claim (flags, hops, the destination its header-type bit implies, and the context byte after it) keeps the bare RX <n> bytes rssi= snr= shape. lnsd's RNode interface emits the same rssi=/snr= keys on its LORA_RX iface=… len=… trace event and its RX … bytes from radio line, paired to the data frame the way the reference interface pairs them: the firmware indicates CMD_STAT_RSSI/CMD_STAT_SNR immediately before each data frame, so the last-seen stat values are that frame's own report (keys absent until the first stat frame). |
ANNOUNCE_LEARNED_NOT_RELAYED | leviculum-nrf/src/events.rs, from the core's NodeEvent::AnnounceLearnedNotRelayed; line rendered by leviculum-log-line's AnnounceLearnedNotRelayedBody, byte-pinned by its host tests | Not a loss — read the name literally. ANNOUNCE_LEARNED_NOT_RELAYED closed=<r> discovery=none|open|expired dest=<hex8>. An announce reached this node, was validated, and updated its path table; the node then had no route to pass it on. Nothing was discarded, so no drop bucket moves and nothing in PKT_DROP_SUMMARY will ever account for it — which is exactly why the line exists, because a route that was never taken otherwise leaves no trace at all. It matters most on a board: a board registers no shared-instance local client, so handle_announce's local-client forward cannot run there and the announce table is the only general route to its serial host (leviculum-core/src/node/mvr_board_announce_uplink.rs). closed= says which term of the announce-table gate closed — path_response (the announce carried PacketContext::PATH_RESPONSE, which the table excludes by design, Python Transport.py:1886), rate_blocked, or rate_limited. The event fires only on a node that relays announces at all (enable_transport, or an announce from a local client): on a node with transport off, "learned and not relayed" is the configured steady state of every announce it hears, and a line per reception would say nothing. discovery= reports the discovery TABLE at the moment the announce arrived, not the history of who asked: a pending request is reaped on the first tick past DISCOVERY_TIMEOUT_MS, so a genuinely late path answer reads none, and none does not distinguish "nobody asked" from "somebody asked and the record is gone". expired is the narrow case of an announce that beat the reaper. A duplicate announce inside the rate window never reaches this line at all — it leaves handle_announce at the earlier rate-limited return, which is a counted drop. Off the boards the same event is a crate::tracing::debug! with event="ANNOUNCE_LEARNED_NOT_RELAYED". |
PKT_RELAY | leviculum-nrf/src/events.rs, from the core's NodeEvent::RelayDecided; line rendered by leviculum-log-line's RelayDecidedBody, byte-pinned by its host tests | What this board did with ONE packet a neighbour addressed to it for relay (#346). PKT_RELAY outcome=forwarded|no-path|duplicate|forward-max-hops ph=<hex16> dst=<hex8> hops=<n> iface_out=<n>|none. Off the boards the same decision is already readable as the journey events PKT_FORWARD, PKT_DROP and DEDUP_DROP; on a board none of them exists, because the firmware builds leviculum-core without tracing and every debug!/trace! in the core is a no-op there (leviculum-core/src/lib.rs:83-100). Until this line the only transport-layer account a board could give was the periodic [TRANSPORT] counter line, which says how many packets were forwarded or dropped in the last 30 s and never which — chasing #344 cost a 5.5 h receiver log, a three-port millisecond capture and packet-length arithmetic to establish something the board could have said in one line. ph= is the FULL 8-byte journey correlator, byte-identical to the ph= on a peer's lnsd journey events, so the two logs stitch on one id; dst= is the first 4 bytes of the destination hash, a prefix of the host events' 16-byte dst=. Scope is the ADDRESSED relay path only, and the absence of a line is itself a reading. A packet whose transport header names another node produces nothing: on a shared medium a relay hears every packet routed via its neighbours, and an event per reception is the 99 %-noise problem the counter-only overheard path exists to avoid (drops_overheard_transport_id still accounts for it, and [TRANSPORT] overheard= still prints it). Relayed announces and broadcasts are addressed to nobody and are likewise not reported here. So "a peer says it sent this packet and no PKT_RELAY names it" means the board either never heard it or was never named as its next hop — not that the relay swallowed it. |
[DROP] | leviculum-nrf/src/events.rs, from the core's NodeEvent::PacketDropped; line rendered by leviculum-log-line's PacketDroppedBody, byte-pinned by its host tests; rate limit in leviculum-nrf/drop-budget | One packet this board heard and threw away, with the taxonomy's own reason (#346) and — since #421 — the packet's name. [DROP] reason=<kebab> ph=<hex16>|none dst=<hex8> iface=<n>. The complement of PKT_RELAY, never a second copy: the core's two event sites are disjoint, so one dropped packet is one line, on one of the two. ph= is the FULL 8-byte journey correlator, byte-identical to the ph= on a peer's lnsd journey events and on PKT_RELAY, carried wherever the deciding site already held a hash — the link-request, link-echo, addressed-data and node-layer (decrypt-miss / dead-link, #421's new coverage) drops. ph=none names the two sites that decide without ever hashing: the HEADER_2 overheard gate (whose 99 %-volume path deliberately computes no per-packet SHA-256) and the path-request refusals; there dst= plus t= remain the correlator. So “did the relay forward my report, and if not, why” is one grep for the 16-hex ph value over a board capture and the sender's log: it returns the sender's PKT_TX, the relay's PKT_RELAY or [DROP], and nothing at all means the board never heard it — the answer #344 needed a 5.5 h receiver log, a three-port capture and packet-length arithmetic to reconstruct. (That grep is also what a periculum cell would assert on; the parser is deliberately not part of #421.) Rate-limited to LINES_PER_WINDOW per second (leviculum-nrf/drop-budget, derivation pinned by the_ring_derivation_still_holds); a clipped window closes with [DROP] suppressed=<n> window_ms=<w>, so a storm and a trickle never read the same. |
ANN_TX | leviculum-nrf/src/events.rs, from the core's NodeEvent::AnnounceTransmitted; line rendered by leviculum-log-line's AnnounceTransmittedBody, byte-pinned by its host tests | Why this board just talked (#405). ANN_TX occasion=transit|local|uncapped|path-response|reoffer dst=<hex8> hops=<n> iface=<n>|all, one line per announce transmission. Since #402 the board registers an airtime cap on its LoRa interface, whose holdoff allows two or three transit announces in a window where a capture may show fifteen announce-sized transmissions — and every kind of announce transmission used to leave the same trace, so none of them could be attributed. occasion=transit is a relayed announce the cap passed (immediately, or out of the cap's queue once the holdoff expired); local one this board originated, which bypasses the cap by design and whose rate is the announce policy's, never the cap's (hops=0 on the same line says so); uncapped a relayed announce on an interface carrying no cap at all — the board's BLE and serial interfaces beside a capped LoRa one, and the reading to expect if a PHY was never registered; path-response an answer somebody asked for, which is requested rather than propagated and does not mean this board relays announces at all. reoffer (#383) is a stored announce handed to one peer whose first link on a multi-peer interface just came up — the announce a relay ladder that retired before the link existed would otherwise never deliver; it goes on that peer's link alone and is paced by the #402 cap where one is registered. Transmissions only. An announce the cap held back or dropped emits nothing here — that is ANN_TX_SUPPRESSED, which exists only in the core's tracing, and the boards have no tracing; silence on this line is not evidence a board stayed quiet. The board's own [ANNOUNCE] sent reason= line is the other half and stays as it is: it names the occasion of the announces this board ORIGINATES (peer-up, periodic, host), at the moment the policy decides, where this line is one per transmission at the moment it leaves. Off the boards the same statement is the host event ANN_TX, which carries the same occasion= beside dst=<hex32> hops= iface=<name>. |
[TELEMETRY] send | leviculum-nrf/src/telemetry.rs | Frozen shape — the #365 proofs read this. One line per telemetry report attempt, emitted at the moment of the routing decision: send dst=<hex8> via=<iface-name> next_hop=<hex8|direct> online=<y|n>. "Sent to a live carrier" is this line (online=y) followed by the report target=… line once the dispatch settles; "not sent" is report withheld target=<hex8> reason=<r> instead, where reason=no-path means no path entry at all and reason=iface-offline means an entry exists but its interface is offline (carrier off, or a peerless BLE domain) — both are followed by a path request on the carriers still online. lnsd states the same decision for every originated packet as the host events OUTBOUND_ROUTE dst= iface= next_hop= online= and OUTBOUND_WITHHELD dst= iface= next_hop= reason= (schema-validated via EVENT_CATALOG, asserted in leviculum-std/tests/obs_outbound_route_events.rs). |
[TELEMETRY] proof / path dropped, asking again / retry / gave up | leviculum-nrf/src/telemetry.rs | Frozen shape — the #373 proofs read these. The proof wait of a sent report (state machine host-tested in leviculum-nrf/telemetry-policy): the transport tracks a receipt per single packet, and a report that gets no proof within the receipt timeout (leviculum-core/src/transport.rs compute_receipt_timeout) is retransmitted ONCE with the identical payload — same position, same time — then given up. [TELEMETRY] proof pkt=<hex8> after=<ms> on a verified proof (after= measured from the first send, whether the first send or the retransmission landed); [TELEMETRY] path dropped, asking again dst=<hex8> pkt=<hex8> when the proof failed to come while a path was held: the path is dropped and asked for once before the retry (#344, the reference's drop_path then request_path, reference/LXMF/LXMF/LXMRouter.py:2743-2751), and the retry waits for the answer, so a capture tells a rediscovered retry from a plain one; [TELEMETRY] retry pkt=<hex8> reason=no-proof at the moment the retransmission goes out; [TELEMETRY] gave up pkt=<hex8> after=<ms> after the second loss, and the next scheduled report carries on. pkt= is always the FIRST send's packet hash, so the two or three lines of one report correlate on one id. The retry goes through the ordinary send path (an offline path is re-resolved, the interface applies its airtime rules) and moves no cadence: min_interval_ms for the next report still counts from the original attempt. Firmware only — lnsd has no telemetry sender. |
[TELEMETRY] discarded | leviculum-nrf/src/telemetry.rs | What arrived at the board's lxmf.delivery destination and could not be kept. A board announces that destination because a receiver verifies its reports against the key the announce carries, but it has no inbox, no message store and no links: the reporter keeps a Sideband telemetry request and discards everything else. [TELEMETRY] discarded from=<hex8> len=<n> reason=no-inbox is a real LXMF message thrown away — somebody wrote to this board and nobody will ever read it; [TELEMETRY] discarded len=<n> reason=not-a-message is bytes encrypted to that hash that are no LXMF message at all. The destination proves NOTHING (ProofStrategy::None, via LxmfNode::delivery_destination_without_inbox), so neither case is confirmed to its sender; the line exists because the version that dropped these in silence let a board mark a peer's message DELIVERED and then throw it away. Firmware only. |
BOOT_TRACE | leviculum-nrf/src/boot_trace.rs (shape host-tested in leviculum-nrf/boot-trace) | One line per boot, first thing in the boot banner: what the PREVIOUS boot's breadcrumb record in retained RAM (.retained, a memory.x region the Adafruit bootloader provably never touches — its original .uninit home sat under the bootloader's stack and was wiped on every reset) says. BOOT_TRACE prev_magic=ok|absent prev_phase=<milestone> prev_boot=<n> reset_reason=<hex>. prev_phase is the last boot milestone the previous boot COMPLETED (enter-main, persist-read, usb-up, lora-task/lora-skipped, sd-enabled, ble-task, main-loop) — a healthy reboot reads main-loop; anything earlier says the previous boot HUNG right after that milestone, which is the whole point: a boot that dies before USB enumerates is otherwise invisible (ledger local-pocket-dark-d66209e). prev_magic=absent (with prev_phase=absent prev_boot=0) is the honest first-boot-after-power-loss answer, never a fabricated phase; prev_phase=unknown-0x<byte> marks a record left by an image with a different phase table. reset_reason is raw POWER.RESETREAS at entry to main (cleared after the read so each boot reports only its own cause); the decoded per-bit view follows on the [RESET_REASON] line, now on both boards. |
LORA_TX_MUTED / LORA_TX_UNMUTED | leviculum-nrf/src/lora.rs (admit_for_transmit); lines rendered by leviculum-log-line's lora_tx_muted / lora_tx_unmuted, byte-pinned by its host tests | Whether this board's transmitter is switched off (#410). The host's radio_silent flag drops every outgoing LoRa frame at the driver boundary while the receiver keeps running, so a muted board has a live receive loop, [MEDIA] lora=on, a climbing transport counter and no [LORA] TX — and used to have no line saying why either: the frame is swallowed before the acquisition jitter, the CAD, the CSMA verdict and the airtime lock, each of which logs. LORA_TX_MUTED packets=<n> bytes=<n> reason=host-silent on the first frame of a run and at each decade after it (the 1st, 10th, 100th …), for MEDIA_TX_DROP's reason — every line also writes the 2 KiB post-crash tail and a muted board drops one frame per announce; LORA_TX_UNMUTED packets=<n> bytes=<n> closes the run with its exact totals when a frame reaches the radio again. The companion fact is silent= on the [LORA] active config: line, which states the flag at the moment a host sets or clears it. radio_silent is deliberately not persisted (leviculum_core::radio_config_store), so a reset always ends a mute — which is why the mute is not visible in a capture taken after one. |
LORA_QUEUE_DROP | leviculum-nrf/src/lora.rs (the hold branch of lora_task); rule and its derivation in leviculum-nrf/queue-budget (HOLD_MAX_AGE_MS, hold_verdict), lines rendered by leviculum-log-line's lora_queue_drop_stale / lora_queue_drop_suppressed, both byte-pinned by their host tests | A frame the regulatory duty lock aged out instead of keying (#433). Read the companion fact first: a duty-cycle hold does not thin traffic, it ages it — FIFO under a hold means each dip of the ledger below the cap admits one frame and re-pins the lock, so under sustained load the queue drains at the cap's rate with its order intact and everything reaching the air is as old as the standing backlog (measured at 145.6 s on the Pocket V2, 2026-09-27). LORA_QUEUE_DROP reason=stale age_ms=<n> bytes=<n> total=<n> says the interface threw such a frame away where it waited: age_ms= is the wait since LoRaInterface::try_send accepted it, total= the count since boot. Only while [LORA_AIRTIME_LOCK] … holding is the reason for the wait — a frame held by CSMA, by an acquisition-jitter draw or behind a burst gap is keyed whatever its age. Type-blind: an age and a length, never a packet kind, because the interface reads neither. Rate-limited by leviculum-drop-budget exactly like the core's [DROP] lines, so a purge cannot evict the [STACK]/[TRANSPORT]/panic lines a capture was taken for; LORA_QUEUE_DROP suppressed=<n> window_ms=<n> names what a clipped window held back. The running total is also the lora_stale= field on the periodic [TRANSPORT] line — the one field there that is not a core counter, because an interface-level loss has none by design. Why the rule exists and where its 18 s comes from: docs/src/concepts/regulatory-airtime.md. |
LORA_TX_STALE | leviculum-std/src/interfaces/rnode.rs (count_stale_drop, reached from pop_live_frame and pop_live_vport_frame); rule in leviculum-nrf/queue-budget (HOLD_MAX_AGE_MS, DutyHolds) | The lnsd twin of LORA_QUEUE_DROP reason=stale (#433). LORA_TX_STALE iface=<name> age_ms=<n> len=<n> held_ms=<n> says the RNode interface's host-side send queue threw a frame away at dequeue instead of handing it to the modem: age_ms= is the wait since the interface accepted it, len= the payload length, held_ms= how much of that wait overlapped a duty hold. The host cannot read the modem's airtime lock, so a duty hold here is what the lock looks like from the serial line: the CMD_READY gate held shut with frames waiting for longer than a full frame's airtime plus one re-query interval explains. Only a frame past HOLD_MAX_AGE_MS (18 s) whose wait overlapped such a hold is dropped; a frame old for any other reason is handed over. Type-blind: an age and a length, never a packet kind. Counted as tx_stale_drops on the interface's interface_stats row and in tx_queue_drops (it is a loss); lnstatus renders it as TX stale. WARN level, one line per drop. |
BOOT_COUNT | leviculum-nrf/src/boot_count.rs (layout, wrap and line host-tested in leviculum-nrf/boot-count) | How often this board restarted while nobody was watching (#380). One line per boot, beside BOOT_TRACE and [RESET_REASON]: BOOT_COUNT n=<n> reset_reason=0x<hex> retained=0|1 since_erase=<n>. n= is the boot number, and it is the one number here that survives a POWER loss — BOOT_TRACE's prev_boot= comes out of retained RAM and restarts at 1 whenever the supply goes away, which is exactly the case a field board produces (the 90 minute walk that motivated this reported reset_reason=0x00000000 prev_magic=absent on every one of its boots). retained= is that same RAM's verdict carried onto this line, so the pair separates a power loss (retained=0 with reset_reason=0x00000000) from a watchdog or a commanded reset. The count lives in one page of internal flash, 0xD9000 on all three boards — memory.x's BOOT region, directly below the record store — one 16-byte record appended per boot; since_erase= is how many of the page's 256 slots are spent, so it counts up to 256 and the boot that finds the page full erases it and carries n= forward — a restart of since_erase= with n= unbroken is that erase and not a lost count. Ungated like [MEDIA] and [QSPI]: a board on battery in a field has no host to open the drain gate. BOOT_COUNT_WRITE_FAILED n=<n> follows the line when the flash refused the write — the boot number is still stated, it simply will not be there next time. |
Host-side BLE events
lnsd's Columba BLE interface (leviculum-std/src/interfaces/ble/)
emits BLE_SCAN_DECISION with the same fields as the firmware's
line of the same name — addr=, caps_record=, caps=,
free_slots=, rule=, initiate= — so a merged rig timeline shows both sides of one
mutual sighting deciding, and the two rule= values must be
complementary (one initiate…, one wait…). Its link lifecycle is
BLE_LINK_UP / BLE_LINK_DOWN (with role=central|peripheral and
peer=<hex8>, the same hex the firmware logs and the LN-<hex8>
name carries), the admission decisions are the firmware's
BLE_LINK_SELF / BLE_LINK_DUP / BLE_LINK_REPLACED names — lnsd's
lines carry rule=, origin=incoming|outgoing, old_mtu=/new_mtu=,
old_silence_ms= and the reported old_data_silence_ms= exactly as
the firmware's do (#360 round 2), plus old_role= on a replacement —
and one lnsd has no firmware counterpart for:
BLE_LINK_NOT_ADMITTED peer=<hex8> identity=<hex32> addr=<a> role=peripheral listed=<n> action=disconnect — the accept_only
config key refused this incoming link at its identity handshake, so
the peer never became a link. identity= carries the whole hash
beside peer='s four bytes because both are spellings the key
accepts, and listed= is how many peers the list names, so a capture
shows that a list is in force without the config beside it. The line
exists so a measuring host that turned strangers away does not read
like a host nobody tried; its dialling-side sibling is
BLE_DIAL_NOT_ALLOWED, which fires when initiate_only is what
stopped a dial. The churn penalty (#417) is the firmware's
BLE_CHURN_LEDGER with the bare token spelled state=:
BLE_CHURN_LEDGER state=armed identity=<hex8> run=<n> reason=expiry_central|expiry_peripheral|replaced window_ms=<n> cooldown_ms=<n> when a rotating-addressed identity's run of
departures (a BLE_LINK_DOWN reason=timeout or a BLE_LINK_REPLACED)
closes the rotating fallback class, and BLE_CHURN_LEDGER state=held run=<n> until_ms=<n> now_ms=<n> once per cooldown when that closed a
candidate out of the collection window. Same ledger, same constants
(ChurnPolicy::MEASURED), no A/B switch on lnsd. A fan-out drop on a congested
link is BLE_TX_FANOUT_DROP. The delivery-hint decision
is the firmware's BLE_TX_ROUTE / BLE_TX_FLOOD /
BLE_TX_ROUTE_MISS (#376), with conn=central|peripheral naming the
role of the link rather than a SoftDevice handle; routing at a
peripheral-role peer still reaches the other subscribed centrals,
because BlueZ fans one notification out to every subscriber and the
Columba service has a single notify characteristic. A reassembly discarded
before completion is BLE_RX_ABANDON (#373) with iface=,
peer=<hex8>, lost= (packets this frame cost) and total= (the
link's running count) — the firmware's line of the same name is
slot-keyed instead of identity-keyed, everything else matches.
There is no lnsd counterpart of the firmware's BLE_CONN_PARAMS
(#385), and the gap is in BlueZ rather than in this interface: the
D-Bus Device1 interface publishes no connection interval, slave
latency or supervision timeout — bluer's device properties (bluer
0.17.4, src/device.rs) run from Name and Rssi to ServicesResolved
and BatteryPercentage and contain none of the three — and neither does
sysfs, whose per-connection directory (/sys/class/bluetooth/hci<n>:<handle>)
carries only uevent. The values exist one layer down, in the LE
Connection Complete and LE Connection Update Complete events on the HCI
transport, reachable only through an HCI monitor socket (what btmon
reads) with the privileges that implies. So on a bench with an lnsd
central the parameters are read from the capture or the central's kernel,
as #385 did; on a link to a phone the board's line is the only source,
which is why the firmware half exists. The same blindness applies to
the answer lnsd gives: when a board asks its lnsd central for a longer
supervision timeout (BLE_CONN_PARAMS_REQ), whether the kernel granted
it is readable from the board's when=close line or from an HCI
capture, and from nothing lnsd itself logs.
These are host events, so
unlike the firmware's they do appear in EVENT_CATALOG and are
schema-validated.
Reading "was the radio listening at instant X"
[SX_RX_ARM] is emitted at the SetRx that arms the SX1262, once
per receive window:
[SX_RX_ARM] site=idle timeout_ms=0 dark_ms=3 t=123456
t=is the instant the receiver went live. It is not derivable from the window's completion line:[T114_LORA_LOOP] op=rx_* duration_ms=brackets the wholereceive()call, IRQ setup and buffer readout included, sot - duration_mslands before the arming, not on it.dark_msis the gap back to the previous window's end, computed on the board. The boot arm has no previous window and saysdark_ms=firstrather than a digit.siteis which of the loop's six listening windows this is —idle,ack,csma,jitter,hold,yield— because their timeouts overlap and the length alone does not identify them.
The two together close the span: window n was listening from its
own t= until t(n+1) - dark_ms(n+1). Everything outside those
spans is standby, including CAD and TX. So the last window in a
capture has no closing instant — its successor is what supplies it.
The line is written through log_fmt, so a board with nothing
attached to the debug CDC (RUNTIME_DRAIN_OPEN == false) never
formats it.
Reading a LINK_DIED age
LINK_DIED comes from the core (link_management.rs, target
leviculum_core::link) and carries its age in milliseconds, to be
read against the threshold_ms on the same line:
elapsed_since_activity_msis the time since the link's last inbound packet, and readsnonefor a link that never had one — every handshake that did not complete. There the expected number does not exist, and the raw subtraction would print the process uptime in its place: the field base of 2026-09-27 reported 91 responder culls aselapsed_since_activity_ms=9581430besidethreshold_ms=85824, for culls that were 72 to 90 s old and therefore on time (#354).since_request_msappears ondetail=handshake_timeoutand is the age of the handshake itself, measured from the request — the initiator's send, the responder's proof. That is the clockestablishment_timeout_ms()and hencethreshold_msis measured on, so those two are the pair to compare: a handshake culled on time shows a difference of one tick.
Validation behaviour
Two violation classes, both non-blocking — the original event line is never suppressed.
Schema violation (per-handle)
EVENT_SCHEMA_VIOLATION event=<NAME> missing=[a,b] caller=file:line t=<ms>
Emitted when a catalogued event misses required keys at emission.
Each active handle's catalogue lookup chains the production
EVENT_CATALOG with the handle's own extra_schemas, so
test-only schemas don't pollute the production catalogue.
Field-value violation (per-event)
EVENT_FIELD_VIOLATION event=<NAME> field=<key> value_problem=<kind> caller=file:line t=<ms>
Emitted when a field's stringified value contains ASCII
whitespace, =, or non-printable characters. Such values break
the whitespace-tokenised parser used by Stage-7's
jl --filter <key>=<value> filter. <kind> is one of
whitespace, equals, non_printable.
The fix at the call site is to pick a value form that doesn't
need escaping — substitute _ for spaces, drop = from value
strings, etc. The original event line is still emitted; the
tester sees the violation alongside, treats it as a bug.
User-named values: render them at the emission site
A value that can carry text a user chose — an interface name from
the config file ([[TCP Uplink]]) or from discovery
(autoconnect/Dark Doodad 23), a filesystem path, an instance
name — is not a source bug when it contains a space: it is
legitimate input that the EMISSION SITE has to render as a single
token. Wrap it in leviculum_std::event_log::Scalar (or, inside
leviculum-core, event_scalar::Scalar; the interface-name
formatters IfaceName / IfaceNameOpt already do it for every
iface = %… field):
#![allow(unused)] fn main() { tracing::debug!(event = "BLE_LINK_UP", iface = %Scalar(&self.name), …); }
The sink still rescues an unwrapped value (sanitize_scalar), and
for the handful of fields it can recognise as names by key
(iface, iface_in, iface_out, in_iface, out_iface) it does
so without raising a violation. That list cannot be completed
from the sink side — next_hop carries an interface name at one
site and a hash at another — so a name wrapped at the site is the
only form that is correct at every field. Substitution, not
quoting: jl, jldiff and every awk/grep one-liner split on
whitespace, so a quoted value with a space is still several tokens
to all of them.
Multi-process workflow
Spawned subprocesses (e.g. an lnsd child of an integration
test) emit to a per-process file when given two env vars:
LEVICULUM_EVENT_LOG=/tmp/leviculum-events-<pid>.log \
LEVICULUM_EVENT_NODE=node-a \
./lnsd ...
LEVICULUM_EVENT_LOG=<path>— child appends each event line (and any field-violations) to<path>as it emits. When unset, the subscriber writes only to the in-memory buffer used for panic-dump.LEVICULUM_EVENT_NODE=<name>— supplies thenode=value.
After the children exit, the parent merges all per-process files:
#![allow(unused)] fn main() { use leviculum_std::test_support::event_log::merge_event_logs; let merged: Vec<String> = merge_event_logs(&[ PathBuf::from("/tmp/leviculum-events-12345.log"), PathBuf::from("/tmp/leviculum-events-12346.log"), ]); }
merge_event_logs reads every input, parses the trailing t=<n>
token of each line, and returns the union sorted by t (stable
on tie). Lines without parseable t= sort to the end with
their relative order preserved.
Per-process clock note: t= values are millisecond offsets from
each subscriber's local init time, not a shared wall clock.
Merged ordering is monotone across the union but doesn't directly
say which real-world event came first across hosts. For
wall-clock correlation add a per-emission timestamp field to the
catalogue (ts=<unix-ms>) and sort on that instead.
Production-daemon integration
lnsd honours both env vars at startup via
leviculum_std::test_support::event_log::install_global_subscriber().
When LEVICULUM_EVENT_LOG is unset, the install path is
functionally equivalent to the previous
tracing_subscriber::fmt().init() call — no event-log layer is
built, so runtime overhead is whatever the fmt layer would
otherwise impose.
rnsd is Python-side (reference/Reticulum); structured event
capture for it is out of scope for the current Rust-side work.
See also
- Codeberg #39 piece 1 (this document's spec).
leviculum-std/src/test_support/event_log.rs(implementation + catalogue).leviculum-std/src/test_support/tracing_setup.rs(Registry composition + Once-guard).leviculum-std/tests/event_log_subscriber.rs(unit tests).leviculum-std/tests/event_log_multiprocess.rs(multi-process merge integration test).- Stage 7:
jl/jldifffilter tools that consume this format.
jl and jldiff — filtering and comparing structured event logs
Stage 7 / Codeberg #39 piece 4. Two CLI tools that consume the
structured event-log format established by
structured-event-logs.md:
jl— filter and slice an event log. Reads from stdin or one or more files, applies AND-combined filters, emits matching lines unchanged.jldiff— compare two event logs by an alignment-key tuple. Partitions events into LEFT_ONLY / RIGHT_ONLY / MATCHED_DIFFER / MATCHED_IDENTICAL buckets.
Examples below are verified by
tests/jl_jldiff_docs.rs. If you change a worked example here, update the test; if a worked example breaks, the doc is wrong, not the test.
Test infrastructure
The tools are exercised by six test files in leviculum-std/tests/:
jl_filter.rs— Phase A unit/integration tests forjlin isolation.jldiff_compare.rs— Phase B unit/integration tests forjldiffin isolation.jl_jldiff_workflow.rs— end-to-end Subscriber → binary tests.jl_jldiff_fixtures.rs— checked-in real-shape log files driving expected outputs.jl_jldiff_edge_cases.rs— boundary and adversarial inputs.jl_jldiff_docs.rs— every example below is mirrored here as a test.
When a Stage-6 format change drifts these tools, all six of those files are likely to fail at once; that is the intended signal.
Format recap
EVENT_NAME node=<name> k1=v1 k2=v2 ... kN=vN t=<rel-ms>
EVENT_NAMEfirst.node=<name>always second when present.- Other fields alphabetically sorted.
t=<ms>always last; integer ms relative to subscriber init.
Lines that do not fit the structured shape (banners, free text,
cargo-test output) pass through jl unchanged. See
structured-event-logs.md for the
full format spec, the runtime catalogue, and the violation-line
synthesis rules.
jl — filter binary
jl [--filter <expr>]... [--node <name>] [--since-event <NAME>] [--until-event <NAME>] [INPUT...]
| Flag | Effect |
|---|---|
--filter <expr> | Filter expression. Repeatable; AND-combined. |
--node <name> | Shorthand for --filter node=<name>. |
--since-event <NAME> | Drop everything before the first event whose EVENT_NAME is <NAME>. The matching event is included. At most one. |
--until-event <NAME> | Drop everything at and after the first event whose EVENT_NAME is <NAME>. The matching event is excluded. At most one. |
INPUT... | Optional file paths. Without any, reads stdin. Multiple files are read in order; output preserves order. |
Filter expression forms:
| Form | Meaning |
|---|---|
key=value | exact match |
key=* | event has that key (any value) |
key=prefix* | value starts with prefix |
t<N, t>N, t<=N, t>=N | numeric t comparison |
The event key is special-cased: event=PKT_RX matches BOTH a
real PKT_RX ... line (where the EVENT_NAME first token is
PKT_RX) AND a synthetic violation line whose explicit event=
field is PKT_RX. This makes the filter consistent across real
events and the EVENT_SCHEMA_VIOLATION / EVENT_FIELD_VIOLATION
lines that reference them.
Example 1: filter to one event-name
Input:
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
ANN_RX node=alpha dst=abc1 hops=0 iface=lora0 path_response=false t=20
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=21
PKT_RX node=alpha dst=abc2 hops=1 iface=lora0 len=64 type=Data t=80
Command:
jl --filter event=PKT_RX
Output:
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
PKT_RX node=alpha dst=abc2 hops=1 iface=lora0 len=64 type=Data t=80
Example 2: slice between two markers
Input:
PKT_LOCAL node=alpha dst=abc1 iface=lora0 matched=true t=10
PKT_LOCAL node=alpha dst=abc1 iface=lora0 matched=true t=20
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=30
PKT_RX node=alpha dst=abc2 hops=0 iface=lora0 len=64 type=Data t=40
PKT_RX node=alpha dst=abc3 hops=0 iface=lora0 len=64 type=Data t=50
PKT_DROP node=alpha dst=abc4 hops=3 iface_in=lora0 reason=ttl_expired type=Data t=60
PKT_RX node=alpha dst=abc5 hops=0 iface=lora0 len=64 type=Data t=70
Command:
jl --since-event PATH_ADD --until-event PKT_DROP
Output:
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=30
PKT_RX node=alpha dst=abc2 hops=0 iface=lora0 len=64 type=Data t=40
PKT_RX node=alpha dst=abc3 hops=0 iface=lora0 len=64 type=Data t=50
Example 3: time window
Input (same as Example 1), with one extra later event:
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
ANN_RX node=alpha dst=abc1 hops=0 iface=lora0 path_response=false t=20
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=21
PKT_RX node=alpha dst=abc2 hops=1 iface=lora0 len=64 type=Data t=80
PKT_RX node=alpha dst=abc3 hops=2 iface=lora0 len=64 type=Data t=200
Command:
jl --filter t>=20 --filter t<100
Output:
ANN_RX node=alpha dst=abc1 hops=0 iface=lora0 path_response=false t=20
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=21
PKT_RX node=alpha dst=abc2 hops=1 iface=lora0 len=64 type=Data t=80
jldiff — compare binary
jldiff --align-on <key>[,<key>...] LEFT_FILE RIGHT_FILE
The alignment-key tuple groups events on each side. Each group's
events are paired by file order (1st left ↔ 1st right, …); surplus
events on either side go to LEFT_ONLY / RIGHT_ONLY. Events
missing one of the align-keys are unalignable and surface in the
appropriate _ONLY bucket with an [unalignable: missing key X]
annotation.
Output format:
=== LEFT_ONLY (N events) ===
<event line>
...
=== RIGHT_ONLY (N events) ===
<event line>
...
=== MATCHED_DIFFER (N pairs) ===
L: <left event line>
R: <right event line>
DIFF: key=lvalue|rvalue [key=lvalue|rvalue ...]
=== MATCHED_IDENTICAL (N pairs) ===
MATCHED_IDENTICAL is count-only — events with no field
differences are not re-listed. The t= field is reported in DIFF
lines when it differs (which is normal — alignment keys are how
you say "same logical event"; t shifts naturally between runs).
Example 4: compare two mvr-test runs
a.log:
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=11
ANN_RX node=alpha dst=abc1 hops=0 iface=lora0 path_response=false t=20
b.log:
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=15
PATH_ADD node=alpha dst=abc1 hops=2 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=16
Command:
jldiff --align-on event,dst a.log b.log
Output:
=== LEFT_ONLY (1) ===
ANN_RX node=alpha dst=abc1 hops=0 iface=lora0 path_response=false t=20
=== RIGHT_ONLY (0) ===
=== MATCHED_DIFFER (2 pairs) ===
L: PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
R: PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=15
DIFF: t=10|15
L: PATH_ADD node=alpha dst=abc1 hops=0 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=11
R: PATH_ADD node=alpha dst=abc1 hops=2 iface=lora0 next_hop=alpha ok=true source=announce table_len=1 t=16
DIFF: hops=0|2 t=11|16
=== MATCHED_IDENTICAL (0 pairs) ===
Example 5: multi-key alignment (lnsd vs Python-RNS)
a.log (lnsd):
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
PKT_RX node=alpha dst=abc1 hops=0 iface=tcp1 len=64 type=Data t=20
b.log (Python-RNS):
PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=12
PKT_RX node=alpha dst=abc1 hops=0 iface=tcp1 len=64 type=Data t=22
Command:
jldiff --align-on event,dst,iface a.log b.log
Output:
=== LEFT_ONLY (0) ===
=== RIGHT_ONLY (0) ===
=== MATCHED_DIFFER (2 pairs) ===
L: PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=10
R: PKT_RX node=alpha dst=abc1 hops=0 iface=lora0 len=64 type=Data t=12
DIFF: t=10|12
L: PKT_RX node=alpha dst=abc1 hops=0 iface=tcp1 len=64 type=Data t=20
R: PKT_RX node=alpha dst=abc1 hops=0 iface=tcp1 len=64 type=Data t=22
DIFF: t=20|22
=== MATCHED_IDENTICAL (0 pairs) ===
The multi-key tuple (event, dst, iface) keeps the two interfaces
separate even though both have the same event and dst —
without iface in the key, jldiff would multi-occurrence-pair
them in file order, which is fine but obscures the per-interface
view.
Workflow notes
When an mvr-test fails, the dump goes to stderr framed by
=== EVENT LOG DUMP ... === banners. Pipe it through jl to
narrow:
just mvr 2>&1 | jl --filter event=PATH_ADD
The banners and any free-text lines around the dump pass through unchanged; only the structured events filter.
For an A/B comparison between two runs, capture each run's output
to a file and run jldiff:
# Run a baseline; capture only the structured events.
just mvr 2>&1 | jl > baseline.log
# Run again after a change.
just mvr 2>&1 | jl > candidate.log
# Diff aligned on the event identity.
jldiff --align-on event,dst,iface baseline.log candidate.log
For multi-process logs, the Stage-6
merge_event_logs
helper produces a t-ordered union; jl and jldiff then operate
on the merged file as if it came from a single subscriber.
See also
structured-event-logs.md— Stage-6 format spec, subscriber architecture, runtime catalogue.- Codeberg #39 — the test framework epic this batch closes.
Storage Trait Split Analysis
Deep analysis of every Storage trait method: callers, frequency, embedded relevance, and proposed sub-trait groupings.
Method Inventory: 71 methods across 15 data categories
Group 1: Packet Dedup (3 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 1 | has_packet_hash | Transport (process_incoming) | hot -- every inbound packet | ESSENTIAL |
| 2 | add_packet_hash | Transport (9 sites: send/receive/proof/data) | hot -- every packet | ESSENTIAL |
| 3 | remove_packet_hash | DEAD CODE -- 0 production calls | never | dead |
Embedded impl: Fixed-size ring buffer, e.g. [[u8; 32]; 512] with
write cursor. Has-check is linear scan (512 x 32B = 16KB). Cannot be
no-op -- without dedup, packets loop forever.
Group 2: Path Table (12 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 4 | get_path | Transport (19 sites) | hot | ESSENTIAL |
| 5 | set_path | Transport (4 sites) | frequent | ESSENTIAL |
| 6 | remove_path | NodeCore (1), Transport (3), RPC (1) | sometimes | ESSENTIAL |
| 7 | path_count | NodeCore (1), Transport (1), Driver (1) | rarely | nice-to-have |
| 8 | expire_paths | Transport (clean_path_states) | periodic | ESSENTIAL |
| 9 | earliest_path_expiry | Transport (next_deadline) | periodic | ESSENTIAL |
| 10 | has_path | NodeCore (3), Transport (4), Driver (1) | frequent | ESSENTIAL |
| 11 | path_entries | Transport (2: path_table_entries, drop_all_paths_via) | rarely | nice-to-have |
| 12 | get_path_state | Transport (1: path_is_unresponsive) | sometimes | nice-to-have |
| 13 | set_path_state | Transport (3: mark_path_unresponsive/responsive) | sometimes | nice-to-have |
| 14 | clean_stale_path_metadata | Transport (clean_path_states) | periodic | nice-to-have |
| 15 | remove_paths_for_interface | NodeCore (1), Transport (1) | rarely | ESSENTIAL |
Embedded impl: Fixed-size array, e.g. [Option<PathEntry>; 32]
with LRU eviction on set. PathEntry is ~50 bytes, total ~1.6KB. Cannot
be no-op -- node can't route without paths.
Path state (methods 12-14) is separable -- unresponsive tracking is a quality-of-life feature. An embedded node could skip it and just remove stale paths via expiry.
Group 3: Announce Processing (11 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 16 | get_announce | Transport (5 sites) | sometimes | ESSENTIAL |
| 17 | get_announce_mut | NodeCore (1), Transport (2) | sometimes | ESSENTIAL |
| 18 | set_announce | Transport (4 sites) | sometimes | ESSENTIAL |
| 19 | remove_announce | Transport (2: check_announce_rebroadcasts) | sometimes | ESSENTIAL |
| 20 | announce_keys | Transport (2: next_deadline, check_announce_rebroadcasts) | periodic | ESSENTIAL |
| 21 | get_announce_cache | NodeCore (1), Transport (3) | sometimes | ESSENTIAL |
| 22 | set_announce_cache | NodeCore (3), Transport (1) | sometimes | ESSENTIAL |
| 23 | clean_announce_cache | Transport (1: clean_path_states) | periodic | nice-to-have |
| 24 | get_announce_rate | Transport (1: check_announce_rate) | sometimes | OPTIONAL |
| 25 | set_announce_rate | Transport (3: check_announce_rate) | sometimes | OPTIONAL |
| 26 | announce_rate_entries | Transport (1: rate_table_entries) | rarely | OPTIONAL |
Embedded impl: AnnounceEntry array, e.g. [Option<AnnounceEntry>; 16] (~2KB). Announce cache stores raw bytes -- variable size, harder
for embedded (up to ~500 bytes each, so 16 x 500 = ~8KB). Cannot be
no-op for core announces (16-22). Rate limiting (24-26) CAN be no-op --
node just doesn't rate-limit.
Group 4: Path Requests (3 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 27 | get_path_request_time | Transport (2: send_path_request rate limiting) | sometimes | ESSENTIAL |
| 28 | set_path_request_time | Transport (1) | sometimes | ESSENTIAL |
| 29 | check_path_request_tag | Transport (1: handle_path_request dedup) | sometimes | ESSENTIAL |
Embedded impl: Small fixed array, e.g. [([u8; 16], u64); 16] for
request times (~384B), ring buffer for tags. Cannot be no-op -- without
request dedup, path request storms occur.
Group 5: Receipts (5 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 30 | get_receipt | Transport (4: get_receipt, mark_delivered, proof handling) | sometimes | OPTIONAL |
| 31 | set_receipt | Transport (3: create_receipt, create_receipt_with_timeout, mark_delivered) | sometimes | OPTIONAL |
| 32 | remove_receipt | 0 production calls (only via expire_receipts) | never | dead (direct) |
| 33 | expire_receipts | Transport (1: check_receipt_timeouts) | periodic | OPTIONAL |
| 34 | earliest_receipt_deadline | Transport (1: next_deadline) | periodic | OPTIONAL |
Embedded impl: Fixed array, e.g. [Option<PacketReceipt>; 8]
(~1KB). CAN be no-op -- node works without delivery proofs. Links still
establish; resources still transfer. You just don't get explicit delivery
confirmation for single packets.
Group 6: Known Identities (2 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 35 | get_identity | NodeCore (1: send_single_packet), Driver (1) | sometimes | ESSENTIAL |
| 36 | set_identity | NodeCore (1: remember_identity) | sometimes | ESSENTIAL |
Embedded impl: Fixed array, e.g. [([u8; 16], Identity); 16].
Identity is ~128 bytes, total ~2.3KB. Borderline essential -- without it,
node can't encrypt to a destination whose announce it already saw but
isn't currently cached in the announce table. Could be no-op if the node
only talks to destinations it just heard announce.
Group 7: Transport Relay (14 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 37 | get_link_entry | Transport (3: forward_link_routed, process_data, is_for_local_client_link) | sometimes | RELAY ONLY |
| 38 | get_link_entry_mut | Transport (1: mark link validated on proof) | rarely | RELAY ONLY |
| 39 | set_link_entry | Transport (1: handle_link_request -- insert bidirectional route) | rarely | RELAY ONLY |
| 40 | remove_link_entry | 0 production calls (only via expire/cleanup) | never | dead (direct) |
| 41 | has_link_entry | Transport (3: dedup exemptions, is_link_routed check) | sometimes | RELAY ONLY |
| 42 | expire_link_entries | Transport (1: clean_link_table) | periodic | RELAY ONLY |
| 43 | earliest_link_deadline | Transport (1: next_deadline) | periodic | RELAY ONLY |
| 44 | remove_link_entries_for_interface | NodeCore (1), Transport (1) | rarely | RELAY ONLY |
| 45 | get_reverse | Transport (1: proof routing) | sometimes | RELAY ONLY |
| 46 | set_reverse | Transport (3: forward_packet, link-routed data, proof handling) | sometimes | RELAY ONLY |
| 47 | remove_reverse | Transport (1: proof routing) | sometimes | RELAY ONLY |
| 48 | has_reverse | 0 production calls (default impl, test only) | never | dead |
| 49 | expire_reverses | Transport (1: clean_reverse_table) | periodic | RELAY ONLY |
| 50 | remove_reverse_entries_for_interface | NodeCore (1), Transport (1) | rarely | RELAY ONLY |
Embedded impl: CAN be full no-op for leaf nodes
(enable_transport=false). A leaf node never relays, never builds
link/reverse tables. If an embedded node IS a relay, needs fixed arrays:
[Option<LinkEntry>; 16] (~1KB), [Option<ReverseEntry>; 32] (~2KB).
Group 8: Discovery Path Requests (5 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 51 | set_discovery_path_request | Transport (1: handle_path_request) | sometimes | RELAY ONLY |
| 52 | get_discovery_path_request | Transport (3: handle/retry/send_discovery) | sometimes | RELAY ONLY |
| 53 | remove_discovery_path_request | Transport (2: send_discovery_path_response) | sometimes | RELAY ONLY |
| 54 | expire_discovery_path_requests | Transport (1: clean_path_states) | periodic | RELAY ONLY |
| 55 | discovery_path_request_dest_hashes | Transport (2: next_deadline, retry) | periodic | RELAY ONLY |
Embedded impl: CAN be full no-op for leaf nodes. Only transport
nodes forward path requests on behalf of others. Leaf nodes send their
own path requests via send_path_request (Group 4), not this mechanism.
Group 9: Ratchets (7 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 56 | get_known_ratchet | Transport (1: set_local_client ratchet replay) | rarely | OPTIONAL |
| 57 | remember_known_ratchet | Transport (1: handle_announce), NodeCore (2: announce_destination, check_mgmt_announces) | rarely | OPTIONAL |
| 58 | has_known_ratchet | 0 production calls | never | dead |
| 59 | known_ratchet_count | 0 production calls (test only) | never | dead |
| 60 | expire_known_ratchets | Transport (1: clean_path_states) | periodic | OPTIONAL |
| 61 | store_dest_ratchet_keys | NodeCore (2: announce_destination, check_mgmt_announces) | rarely | OPTIONAL |
| 62 | load_dest_ratchet_keys | NodeCore (1: register_destination) | rarely | OPTIONAL |
Embedded impl: CAN be full no-op. Node works without forward secrecy -- announces are still validated, links still established, data still encrypted. Ratchets add key rotation for post-compromise security. On a RAM-constrained device, this is the first thing to skip.
Group 10: Shared Instance (7 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 63 | add_local_client_dest | Transport (1: handle_announce for local client) | rarely | SHARED ONLY |
| 64 | remove_local_client_dests | Transport (1: set_local_client cleanup) | rarely | SHARED ONLY |
| 65 | has_local_client_dest | 0 production calls (test only) | never | dead |
| 66 | set_local_client_known_dest | Transport (1: handle_announce) | rarely | SHARED ONLY |
| 67 | has_local_client_known_dest | 0 production calls (test only) | never | dead |
| 68 | local_client_known_dest_hashes | Transport (2: set_local_client, clean_path_states) | rarely | SHARED ONLY |
| 69 | expire_local_client_known_dests | Transport (1: clean_path_states) | periodic | SHARED ONLY |
Embedded impl: Full no-op. Embedded nodes don't share instances. Zero correctness impact. Shared instance is a desktop/server feature (multiple programs sharing one daemon via Unix sockets). An embedded node IS the daemon.
Group 11: Persistence & Diagnostics (2 methods)
| # | Method | Caller | Frequency | Embedded? |
|---|---|---|---|---|
| 70 | flush | Driver (2: save_persistent_state, auto_interface) | rarely | OPTIONAL |
| 71 | diagnostic_dump | Driver (1) | rarely | OPTIONAL |
Already have empty default implementations. No action needed.
Dead Code Summary
7 methods with zero production callers:
| Method | Notes |
|---|---|
remove_packet_hash | Defined but never called anywhere |
has_reverse | Only the default impl delegates to get_reverse; no external callers |
remove_link_entry | Only called indirectly via expire_link_entries |
remove_receipt | Only called indirectly via expire_receipts |
has_known_ratchet | Test-only |
known_ratchet_count | Test-only |
has_local_client_known_dest | Test-only (one NodeCore test) |
has_local_client_dest is also test-only in Transport tests.
Proposed Sub-Trait Split
Tier 1: CoreStorage -- 28 methods
Every node needs these. Without them the protocol doesn't function.
Packet dedup: has_packet_hash, add_packet_hash (2)
Path table: get_path, set_path, remove_path,
path_count, expire_paths,
earliest_path_expiry, has_path,
path_entries, remove_paths_for_interface (9)
Path state: get_path_state, set_path_state,
clean_stale_path_metadata (3)
Announces: get_announce, get_announce_mut,
set_announce, remove_announce,
announce_keys, get_announce_cache,
set_announce_cache, clean_announce_cache (8)
Path requests: get_path_request_time,
set_path_request_time,
check_path_request_tag (3)
Identities: get_identity, set_identity (2)
Persistence: flush (1)
Embedded minimum: ~30KB RAM total
| Collection | Layout | Size |
|---|---|---|
| Packet ring | [[u8; 32]; 512] | 16KB |
| Path table | [Option<PathEntry>; 32] | ~2KB |
| Announce table | [Option<AnnounceEntry>; 16] | ~1KB |
| Announce cache | [Option<([u8;16], Vec<u8>)>; 16] | ~8KB (variable, biggest concern) |
| Path requests | [([u8;16], u64); 16] | ~0.5KB |
| Path states | [Option<([u8;16], PathState)>; 32] | ~1KB |
| Identities | [Option<([u8;16], Identity)>; 16] | ~2KB |
Cannot be no-op.
Tier 2: ReceiptStorage -- 5 methods
get_receipt, set_receipt, remove_receipt,
expire_receipts, earliest_receipt_deadline
Embedded: [Option<([u8;16], PacketReceipt)>; 8] ~1KB.
Can be no-op. Node works; single-packet delivery proofs are lost.
Links, resources, and channels all function -- they have their own proof
mechanisms.
Why separate from CoreStorage: An nRF52840 sensor that only sends data and doesn't care about delivery confirmation saves 1KB RAM and 5 method implementations.
Tier 3: TransportRelayStorage -- 19 methods
Link table: get_link_entry, get_link_entry_mut,
set_link_entry, remove_link_entry,
has_link_entry, expire_link_entries,
earliest_link_deadline,
remove_link_entries_for_interface (8)
Reverse table: get_reverse, set_reverse, remove_reverse,
has_reverse, expire_reverses,
remove_reverse_entries_for_interface (6)
Discovery: set_discovery_path_request,
get_discovery_path_request,
remove_discovery_path_request,
expire_discovery_path_requests,
discovery_path_request_dest_hashes (5)
Embedded: Full no-op for leaf nodes. Only matters if
enable_transport=true.
Can be no-op. Leaf node can't relay, but communicates fine as an
endpoint.
Why one trait instead of three: Link table, reverse table, and discovery requests are always used together -- they're all transport-mode infrastructure. A node is either a relay or it isn't. There's no use case for "relay with link table but no reverse table."
Tier 4: RatchetStorage -- 7 methods
get_known_ratchet, remember_known_ratchet,
has_known_ratchet, known_ratchet_count,
expire_known_ratchets,
store_dest_ratchet_keys, load_dest_ratchet_keys
Embedded: [Option<([u8;16], [u8;32], u64)>; 8] ~0.5KB if
implemented.
Can be full no-op. Forward secrecy is a security enhancement.
Without ratchets, announce encryption still works via the destination's
static keys. An embedded sensor node may not need post-compromise key
rotation.
Why separate: Security feature with storage cost. Embedded devices with extreme RAM constraints can skip it. Also the only group that spans both Transport and NodeCore callers in a way that's cleanly separable.
Tier 5: SharedInstanceStorage -- 7 methods
add_local_client_dest, remove_local_client_dests,
has_local_client_dest,
set_local_client_known_dest, has_local_client_known_dest,
local_client_known_dest_hashes,
expire_local_client_known_dests
Embedded: Full no-op. Zero correctness impact. Shared instance is a desktop/server feature (multiple programs sharing one daemon via Unix sockets). An embedded node IS the daemon.
Why separate: Entirely irrelevant to embedded. Also the most likely candidate for removal from the trait hierarchy entirely -- it could be a compile-time feature flag instead.
Tier 6: AnnounceRateStorage -- 3 methods
get_announce_rate, set_announce_rate, announce_rate_entries
Embedded: [Option<([u8;16], AnnounceRateEntry)>; 16] ~0.5KB if
implemented.
Can be no-op. Without rate limiting, node processes all announces.
On a small network (typical for embedded LoRa), announce volume is low
enough that rate limiting is unnecessary.
Why separate from CoreStorage: Rate limiting is operator policy, not protocol correctness. A network of 5 LoRa nodes doesn't need it.
Composition Design
#![allow(unused)] fn main() { // Tier 1 -- every node trait CoreStorage { /* 28 methods */ } // Tier 2-6 -- optional capabilities trait ReceiptStorage { /* 5 methods */ } trait TransportRelayStorage { /* 19 methods */ } trait RatchetStorage { /* 7 methods */ } trait SharedInstanceStorage { /* 7 methods */ } trait AnnounceRateStorage { /* 3 methods */ } // Backward-compatible supertrait -- existing code unchanged trait Storage: CoreStorage + ReceiptStorage + TransportRelayStorage + RatchetStorage + SharedInstanceStorage + AnnounceRateStorage {} // Blanket impl impl<T> Storage for T where T: CoreStorage + ReceiptStorage + TransportRelayStorage + RatchetStorage + SharedInstanceStorage + AnnounceRateStorage {} }
Transport<C, S> and NodeCore<R, C, S> keep S: Storage -- zero
changes to existing code. MemoryStorage and FileStorage implement all
sub-traits and get Storage for free.
For embedded:
#![allow(unused)] fn main() { struct EmbeddedStorage { // Only CoreStorage collections // ~30KB RAM } impl CoreStorage for EmbeddedStorage { /* real impls */ } impl ReceiptStorage for EmbeddedStorage { /* no-ops */ } impl TransportRelayStorage for EmbeddedStorage { /* no-ops */ } impl RatchetStorage for EmbeddedStorage { /* no-ops */ } impl SharedInstanceStorage for EmbeddedStorage { /* no-ops */ } impl AnnounceRateStorage for EmbeddedStorage { /* no-ops */ } // Gets Storage automatically via blanket impl }
Trade-offs & Uncertainties
Confident assessments
- Groups 5 (SharedInstance) and 3 (TransportRelay) are cleanly separable -- no leaf-node code path touches them in production.
- Group 4 (Ratchets) is cleanly separable -- all call sites have graceful None/no-op fallback.
- The 7 dead-code methods should be removed regardless of whether the trait is split.
Uncertainties
1. CoreStorage is still 28 methods. That's a lot for a "minimal"
trait. I considered splitting Path and Announce into separate sub-traits,
but they're called from the same Transport methods (handle_announce
touches both path table and announce table in the same function).
Splitting would require where S: PathStorage + AnnounceStorage bounds
scattered across Transport methods -- high friction for zero embedded
benefit since both are essential.
2. Announce cache is the RAM wildcard. Each cached announce is up to ~500 bytes of raw wire data. 16 entries = 8KB. On an nRF52840 with 256KB RAM this is manageable, but on a smaller MCU it could dominate. The cache is needed for path responses and link requests from remote nodes. An endpoint that only initiates (never responds to path requests) could skip it -- but that's a very narrow use case.
3. Whether the split is worth the complexity. Right now, NoStorage
already serves as the "skip everything" option, and MemoryStorage with
capacity limits would cover the "real embedded" case. The sub-trait
split adds type-system guarantees but also adds 6 trait definitions, 6
impl blocks per storage type, and ongoing maintenance burden. If there's
only one embedded target (nRF52840), a capacity-limited MemoryStorage
might be strictly better.
4. Conditional compilation is the pragmatic alternative. Instead of
sub-traits, use #[cfg(feature = "transport")] to gate
TransportRelay collections in MemoryStorage. Simpler, less
generic-parameter noise, but loses per-instance flexibility (can't mix
endpoint and transport nodes in the same binary).
Recommendation
Remove the 7 dead methods now. Defer the sub-trait split until the first
embedded target actually needs it. The current MemoryStorage with
configurable capacity limits (already has packet_hash_cap,
identity_cap) extended to all collections covers the nRF52840 case
without any trait refactoring. The sub-trait design above is the right
split IF the refactor becomes necessary -- but it's a premature
abstraction today.
Broadcast behaviour: Python-RNS parity reference
This document is the source-of-truth reference that our Rust
leviculum-core broadcast code must match. It records what
Python-Reticulum does for every broadcast-related mechanism,
citing reference/Reticulum/RNS/Transport.py (and neighbouring files)
by line. The companion mapping table at the end records the
Rust-side implementation or intentional divergence for each item.
The rule (Lew, 2026-04-15): Leviculum matches Python-RNS exactly for on-wire packet counts, packet types, and protocol semantics. Timing may diverge — jitter-window shape and interface pacing are free — as long as the counts and types stay identical.
State of the citations (2026-09-23). Every Rust-side citation on
this page was re-resolved in that pass and is current. The Python side
was only spot-checked, and the vendored RNS/ tree has moved under it
since the page was written: the checks that were made are corrected in
place below, but the line numbers into Transport.py, Packet.py,
Destination.py, Interface.py and Reticulum.py that are not
mentioned here have not been re-established and must be treated as
stale until someone walks them. Confirmed still correct: the constants
at Transport.py:68, :69, :70 and :83, the dedup storage at
:106, :107 and :175, the retry loop at :576-591,
Transport.request_path at :2771, and the management-announce
citations :193, :194, :283 and :963. Corrected below: the
dedup check site, packet_hashlist_prev, the PATHFINDER_RW
constant, the announce-table insert and its local-client special case,
and the outbound/inbound entry points. Everything else on the
Python side is unverified. Section 14 says why this is its own task.
1. Overview: what can appear on the wire
Python-Reticulum emits five distinct packet classes that can be broadcast or unicast:
| Class | Packet type | Scope | Who originates |
|---|---|---|---|
| Self-announce | Packet.ANNOUNCE | Broadcast | Destination.announce() |
| Forwarded announce | Packet.ANNOUNCE | Broadcast | Transport relay on received announce |
| Path-request | Packet.DATA with transport_type = BROADCAST | Broadcast | Transport.request_path() or client call |
| Path-response | Packet.ANNOUNCE with context = PATH_RESPONSE | Targeted | Transport answering a path-request |
| Link-request | Packet.LINKREQUEST | Unicast | Link.__init__ on initiator |
Everything below walks each class.
2. Self-originated announce
Trigger
Destination.announce(app_data, path_response=False, ...) at
reference/Reticulum/RNS/Destination.py:243. Builds an announce
packet, calls announce_packet.send() once at line 322.
On-wire behaviour
Packet.send() at reference/Reticulum/RNS/Packet.py:273-299 calls
Transport.outbound(self) exactly once and returns a receipt (or
False). There is no retry loop on the send path. A second
call on the same packet raises IOError (Packet.py guard).
Fan-out across interfaces
Inside Transport.outbound() at
reference/Reticulum/RNS/Transport.py:1092 (the interior line numbers
in this section are unverified, see the banner above): for broadcast
packets (the "else" branch after the targeted-path and
transport-id branches), the code iterates Transport.interfaces
(line 1027) and transmits on each. There is no
if interface != packet.receiving_interface filter in the
announce path. For self-originated announces receiving_interface
is None anyway (the packet was created locally) so the question
is moot, but the point is relevant when we contrast with the
forwarded-announce path below.
Mode-based filtering is applied in this loop at lines 1040-1084
for MODE_ACCESS_POINT, MODE_ROAMING, MODE_BOUNDARY. These
modes suppress the rebroadcast on specific interfaces depending
on where the destination sits in the mesh. Bandwidth-cap logic
(lines 1089-1162) defers transmissions when the interface is
saturated.
Summary: Python self-announce = exactly 1 on-wire broadcast per
call, via one-shot Packet.send(). Count = 1.
3. Received-for-forwarding announce
Reception
Transport.inbound(raw, interface) at Transport.py:1389 is
the entry point for everything received on an interface. The
packet-hash dedup check at line 1376 is:
if not packet.packet_hash in Transport.packet_hashlist and
not packet.packet_hash in Transport.packet_hashlist_prev:
return True
Transport.packet_hashlist at Transport.py:106 is set().
Transport.packet_hashlist_prev at line 107 is the rolling
previous window used to keep the dedup memory constant-bounded.
A duplicate return here bails out of inbound() before any
announce-specific handling. This is the only mechanism that
prevents the same packet from being processed twice — critical
for the broadcast-back-to-source echo pattern that B1 relies on.
Insertion into announce_table
For announces (packet.packet_type == ANNOUNCE) that pass dedup,
the code path at Transport.py:1866-1908 initialises an
announce_table entry:
retries = 0 # line 1866
local_rebroadcasts = 0 # line 1868
block_rebroadcasts = False # line 1869
attached_interface = None # line 1870
retransmit_timeout = now + (RNS.rand() * PATHFINDER_RW) # line 1872
PATHFINDER_RW = 0.5 (seconds) at line 70, so the first
retransmission is scheduled within 0–500 ms of receipt.
Line 1891-1895 is the special case for announces that arrived from a local client (shared-instance peer over the local socket):
if Transport.from_local_client(packet):
retransmit_timeout = now
retries = Transport.PATHFINDER_R
This sets retries = 1 right away. Combined with the retry-loop
guard below, this makes local-client-sourced announces fire
only 1 time from the scheduler, not 2.
Retry loop
The periodic job at Transport.py:576-591 walks announce_table:
for destination_hash in Transport.announce_table:
announce_entry = Transport.announce_table[destination_hash]
if announce_entry[IDX_AT_RETRIES] > 0 and
announce_entry[IDX_AT_RETRIES] >= Transport.LOCAL_REBROADCASTS_MAX:
# "local rebroadcast limit reached"
completed_announces.append(destination_hash)
elif announce_entry[IDX_AT_RETRIES] > Transport.PATHFINDER_R:
# "retry limit reached"
completed_announces.append(destination_hash)
else:
if time.time() > announce_entry[IDX_AT_RTRNS_TMO]:
announce_entry[IDX_AT_RTRNS_TMO] =
time.time() + Transport.PATHFINDER_G + Transport.PATHFINDER_RW
announce_entry[IDX_AT_RETRIES] += 1
# ... build rebroadcast packet and send
With the constants:
| Constant | Value | Citation |
|---|---|---|
PATHFINDER_R | 1 | Transport.py:68 |
PATHFINDER_G | 5 s | Transport.py:69 |
PATHFINDER_RW | 0.5 s | Transport.py:70 |
LOCAL_REBROADCASTS_MAX | 2 | Transport.py:77 |
Deterministic walk — non-local-client source
Entry inserted with retries = 0, retransmit_at = now + rand*0.5s.
| Tick | retries in | Guard A | Guard B | Action | retries out |
|---|---|---|---|---|---|
| 1 | 0 | 0 > 0 && 0 >= 2 = false | 0 > 1 = false | fire, schedule next | 1 |
| 2 | 1 | 1 > 0 && 1 >= 2 = false | 1 > 1 = false | fire, schedule next | 2 |
| 3 | 2 | 2 > 0 && 2 >= 2 = true | — | remove | — |
Count = 2 rebroadcasts per received non-local-client announce.
Deterministic walk — local-client source
Entry inserted with retries = 1, retransmit_at = now.
| Tick | retries in | Guard A | Guard B | Action | retries out |
|---|---|---|---|---|---|
| 1 | 1 | 1 > 0 && 1 >= 2 = false | 1 > 1 = false | fire, schedule next | 2 |
| 2 | 2 | 2 > 0 && 2 >= 2 = true | — | remove | — |
Count = 1 rebroadcast per received local-client-sourced announce.
Immediate local-client forward
Lines 1788-1833: after the table insertion, Python also emits the announce immediately to every local-client interface that is not the receiving interface:
for local_interface in Transport.local_client_interfaces:
if packet.receiving_interface != local_interface:
new_announce = RNS.Packet(...)
new_announce.send()
This is the only place in the announce path where receiving_interface
filtering happens. It only applies to local-client interfaces — the
fanout onto LoRa, TCP, UDP interfaces is unfiltered. This confirms
that for the mixed LoRa-Serial + LoRa-RF topology our tests care
about, Python does not skip the received interface when
rebroadcasting.
Fan-out per rebroadcast fire
Each fire builds a new announce packet (lines 540-561), calls
send() → Transport.outbound(), which applies the mode
filtering and bandwidth-cap logic. The receiving interface is
implicitly included in the for interface in Transport.interfaces
loop (no exclusion check). Echoes are absorbed by the
packet_hashlist check at line 1376 when they arrive back.
Block-rebroadcasts path
announce_entry[IDX_AT_BLCK_RBRD] set to True (indices at
line 557 of the retry loop) reroutes the rebroadcast as a
PATH_RESPONSE packet (announce_context = PATH_RESPONSE,
line 537). This is how path-responses ride the same scheduler.
4. Path-request
Trigger
Transport.request_path(destination_hash, ...) at
Transport.py:2771 is the main producer. Clients call into it
via Destination.request_path() or explicit transport calls.
On-wire behaviour
At line 2561-2587: builds a Packet with
packet_type = Packet.DATA and
transport_type = Transport.BROADCAST, then calls
packet.send() once. Same one-shot pattern as self-announce.
Fan-out goes through the same Transport.outbound() broadcast
loop at Transport.py:1180-1317.
Count = 1 on-wire broadcast per path-request call. No retries in the scheduler for path-requests.
Rate-limiting
Path-requests are subject to PATH_REQUEST_MI = 20 seconds
minimum interval per destination (Transport.py:83) — clients
requesting the same path more often are throttled upstream of
Transport.outbound().
5. Path-response
Trigger
Two paths produce a PATH_RESPONSE:
- Active answer: Transport receives a path-request, has the
path, calls
Destination.announce(path_response=True, tag=...)with the matching identity. This produces aPacket.ANNOUNCEwithcontext = PATH_RESPONSE(Destination.py:309-310, 319-322) and sends it once. - Rebroadcast with block_rebroadcasts: the retry loop at
Transport.py:576-604emits path-responses whenannounce_entry[IDX_AT_BLCK_RBRD]is set. Same 2-fire count as a regular received-announce rebroadcast.
On-wire semantics
Path-responses are a packet-type subset of announces. The
fan-out logic is the same as announces. Consumers distinguish
by packet.context == PATH_RESPONSE.
Special routing
In Transport.outbound() at lines 1167+ (targeted-transport
branch), a packet with transport_id set AND a known next-hop
in path_table is routed to a single specific interface via
SendPacket, not broadcast. This is what happens when a
path-response is specifically addressed to the path-requester
rather than broadcast. In our Rust code the answering site stamps
the requesting interface onto the announce-table entry it inserts —
target_interface (transport.rs:11313) — and the retry scheduler
hands an entry carrying one to that interface alone instead of
broadcasting it: target_iface (transport.rs:12019-12043).
(Re-read 2026-09-23. The citation this paragraph carried was
written backwards, 4336 down to 4286, and pointed at neither site:
4286 sits inside announce_table_entries (transport.rs:5299), an
RPC export. The only non-test target_interface: Some(...) in the
file is the one cited above.)
Held announces during path-response scheduling
Inserting the path-response entry into announce_table would
overwrite an announce for the same destination still waiting in
its rebroadcast grace. Python holds any such entry in
Transport.held_announces before the insertion
(Transport.py:2991-2999) and reinserts it when the response
entry fires in the retry loop (Transport.py:630-633), so the
targeted response goes out first and the network-wide
rebroadcast afterwards. The Rust counterpart is
hold_displaced_announce plus the reinsertion sites in
check_announce_rebroadcasts (Codeberg #170), with one
declared deviation: a held network-wide rebroadcast is never
displaced by a later path-response entry, where Python
overwrites the held slot on every request and can lose the
rebroadcast to back-to-back requests.
6. Link-request
Trigger
Link.__init__(destination=...) on the initiator. Internally
calls Packet(destination, link_data, Packet.LINKREQUEST, ...)
and sends it.
On-wire behaviour
Packet.LINKREQUEST (Packet.py:62) is unicast, not
broadcast. At Transport.py:2091: local-destination link
requests are dispatched to the destination's attached interface
directly. Non-local paths route through next-hop. There is no
broadcast fanout.
Count = 1 unicast packet per link initiation. Not relevant to broadcast parity directly, but enumerated here for completeness.
7. Dedup (packet_hashlist)
| Item | Value | Citation |
|---|---|---|
| Storage | set() | Transport.py:106 |
| Previous-window storage | set() | Transport.py:107 |
| Max size | 1 000 000 entries | Transport.py:175 |
| Check site | line 1376 | Transport.py |
| Rotation | half-cleared when reaches hashlist_maxsize/2 | approximate, see cull job |
The dedup check is the only mechanism that prevents the self-heard echo when we (Rust) stop excluding the receiving interface from the rebroadcast fanout. Verifying the check fires reliably is a hard requirement for B1.
8. ANNOUNCE_CAP — per-interface rate limiter
Constants
| Constant | Value | Citation |
|---|---|---|
Reticulum.ANNOUNCE_CAP | 2 (percent of bandwidth) | Reticulum.py:114 |
Interface instances set
interface.announce_cap = Reticulum.ANNOUNCE_CAP/100.0 = 0.02
at Reticulum.py:819. Each interface also has
interface.bitrate (bps).
Logic
The rate limiter is consulted only for forwarded announces
(packet.hops > 0). Self-originated announces bypass it because
they only fire once and are not worth deferring.
At Transport.py:1252-1311:
if (packet.hops > 0):
if not hasattr(interface, "announce_cap"): ...
if not hasattr(interface, "announce_allowed_at"):
interface.announce_allowed_at = 0
if time.time() >= interface.announce_allowed_at and interface.bitrate:
tx_time = len(packet.raw) * 8 / interface.bitrate
wait_time = tx_time / interface.announce_cap
interface.announce_allowed_at = time.time() + wait_time
# proceed with immediate TX
else:
# queue for later
if not len(interface.announce_queue) >= Reticulum.MAX_QUEUED_ANNOUNCES:
interface.announce_queue.append(packet)
wait_time = tx_time / 0.02 = 50 × tx_time: each forwarded
announce "books" 50× its own airtime on the interface before the
next forwarded announce is allowed immediate TX.
Queue drain
When announce_allowed_at rolls past and there are queued
announces, the interface's process_announce_queue() pops the
next one and emits it. This is a per-interface deferred-send
mechanism, not a transport-wide one.
9. LOCAL_REBROADCASTS_MAX
Covered in section 3 (retry loop). The enforcement sites are:
Transport.py:582: retry-loop guard A. Prevents emission whenretries >= LOCAL_REBROADCASTS_MAX.Transport.py:1728: secondary site that removes an entry fromannounce_tablewhen a duplicate announce arrives and the local rebroadcast counter has saturated. This is the "I'm hearing too many copies of this announce from others, stop my own rebroadcast too" path.
10. Management announce keepalive
Constants
| Constant | Value | Citation |
|---|---|---|
mgmt_announce_interval | 7 200 s (2 h) | Transport.py:194 |
| Initial-fire trick | last_mgmt_announce = now - interval + 15 | Transport.py:283 |
Behaviour
Transport.py:283 runs at startup and sets last_mgmt_announce
to 15 seconds ago minus the full interval, so the next check at
Transport.py:963 fires ~15 s after startup. Each fire walks
Transport.mgmt_destinations (a list of transport-control
destinations like probe responders and blackhole destinations,
populated at lines 220-241, 367 during Transport.start()) and
announces each.
After each successful batch the code updates
Transport.last_mgmt_announce = time.time().
Purpose
Without this keepalive, a node that loses its initial one-shot
Destination.announce() is unreachable until the next manual
announce. The 2-h re-announce gives the mesh a periodic refresh
without flooding the network with announce traffic.
11. Interface modes
Python-Reticulum distinguishes five interface modes
(Interfaces/Interface.py:45-50):
| Mode | Constant | Intent |
|---|---|---|
MODE_FULL | 0x01 | Default. Fully participating transport node. |
MODE_POINT_TO_POINT | 0x02 | Directed link, no announce flooding. |
MODE_ACCESS_POINT | 0x03 | Gateway to clients. Special path expiry. |
MODE_ROAMING | 0x04 | Mobile node. Selective rebroadcast. |
MODE_BOUNDARY | 0x05 | Edge between mesh segments. Selective rebroadcast. |
MODE_GATEWAY | 0x06 | Inter-mesh gateway. |
These are consulted in Transport.outbound() at lines 1040-1084
to suppress rebroadcast on specific interfaces.
block_rebroadcasts at the announce-table entry level is a
related per-entry flag.
Leviculum does not implement interface modes. All interfaces
behave as MODE_FULL. This is a documented divergence that Phase
A audit records; if a future scenario surfaces that requires
mode behaviour, a separate task lands them. Until then, our
fanout is "unfiltered over the broadcast-capable interface set",
which is behaviourally equivalent to Python with all interfaces
in MODE_FULL.
12. Rust ↔ Python parity matrix
Legend: ✓ matches, ≈ matches in count/semantics with timing or structural divergence, ⚠ gap not yet addressed, ✗ does not match.
| Mechanism | Python reference | Rust today | Status | Notes |
|---|---|---|---|---|
| Self-announce one-shot | Destination.py:322, Packet.py:294 | one emission per call: send_on_all_interfaces (transport.rs:4292) | ✓ | History: the 3 extra retries this row recorded were removed by B3, 08128e97 (2026-04-15) |
| Self-announce on-wire count | 1 | 1 for an ordinary destination; 2 for a management destination | ≈ | B3 brought it to 1; 2d9234da (2026-09-01) gave management destinations the reference's local-client second emission — declared deviation, see section 13 |
| Self-announce fanout | all interfaces (MODE_FULL assumed) | no exclusion argument | ✓ | send_on_all_interfaces (transport.rs:4292) |
| Received-announce rebroadcast count | 2 (non-local-client), 1 (local-client) | 2 and 1: the insert picks the start value from the source, retries (transport.rs:6826-6830) | ✓ | History: B2, 79ac5204 (2026-04-15), dropped PATHFINDER_RETRIES to 1 and reordered the init |
| Received-announce fanout | all interfaces; echo dedup'd on RX | send_on_all_interfaces (no exclude) | ✓ | Matches Python. B1 verified by test_announces_forwarded_through_transport. |
| Packet-hash dedup on RX | Transport.py:1227 | has_packet_hash (transport.rs:3956) | ✓ | Identical semantics, rolling window |
PATHFINDER_G grace | 5 s | 5 000 ms | ✓ | PATHFINDER_G_MS (constants.rs:186) |
PATHFINDER_RW jitter | 0.5 s | 500 ms (+ optional airtime factor) | ≈ | Option α permitted timing divergence |
LOCAL_REBROADCASTS_MAX | 2 | 2 | ✓ | LOCAL_REBROADCASTS_MAX (constants.rs:153); enforced in the retry loop, local_rebroadcasts (transport.rs:11874), and on a duplicate arrival, local_rebroadcasts (transport.rs:6400) |
ANNOUNCE_CAP | 2 % | 2 % | ✓ | DEFAULT_ANNOUNCE_CAP_PERCENT (constants.rs:384); state in InterfaceAnnounceCap (transport.rs:827-834), holdoff at allowed_at_ms (transport.rs:12209-12221) |
announce_queue / deferred-send | interface.announce_queue | InterfaceAnnounceCap.queue | ✓ | Same intent, Rust-side uses Vec |
mgmt_announce_interval | 7 200 s | 7 200 000 ms | ✓ | MGMT_ANNOUNCE_INTERVAL_MS (constants.rs:219); check_mgmt_announces (node/mod.rs:2360-2452) |
| mgmt-announce initial 15 s trick | Transport.py:283 | schedule_initial_mgmt_announce (node/mod.rs:2399-2405) with MGMT_ANNOUNCE_INITIAL_DELAY_MS (node/mod.rs:241) | ≈ | Verified by B4 audit; Rust adds a per-node draw on top, MGMT_ANNOUNCE_INITIAL_JITTER_MS (node/mod.rs:260) |
| mgmt-announce iterates all dests | Python walks mgmt_destinations | check_mgmt_announces walks mgmt_destinations | ✓ | Verified by B4 audit |
| Path-request one-shot broadcast | Transport.py:2771-2809 | transport.rs (to verify in B7) | ≈ | B7 audit |
| Path-response targeted | targeted-transport branch, section 5 | target_iface (transport.rs:12019-12043) | ✓ | Preserved |
| Interface modes (FULL/ROAMING/…) | 5 modes | none (all = FULL) | ⚠ | Documented gap; separate task |
block_rebroadcasts | per-entry flag | AnnounceEntry.block_rebroadcasts | ✓ | Verified by B7 audit |
13. Phase A resolutions of semantic ambiguities
B2 retry-count alignment
Question: PATHFINDER_R = 1 — does this mean 1 retry after
the initial or 1 TX total?
Resolution (walking the Python loop, section 3): Python fires
2 times per received non-local-client announce, bounded by
LOCAL_REBROADCASTS_MAX = 2 not by PATHFINDER_R. The
PATHFINDER_R guard (retries > PATHFINDER_R) would fire at
retries = 2 but LOCAL_REBROADCASTS_MAX fires first at
retries >= 2. In other words, for the default constants the
PATHFINDER_R guard is redundant with LOCAL_REBROADCASTS_MAX
in the non-local-client path.
Rust equivalent target: 2 fires per received non-local-client announce. Achievable in two ways:
- A. Set
PATHFINDER_RETRIES = 1and change the entry-insert so a received non-local-client announce starts atretries = 0rather thanretries = 1. The retry-loop guards already readretries > PATHFINDER_RETRIESandlocal_rebroadcasts >= LOCAL_REBROADCASTS_MAX; both fire at the right count. - B. Set
PATHFINDER_RETRIES = 2and leave the insert atretries = 1. Same on-wire count.
B2 committed path A — it more closely mirrors Python's constants and counter semantics, so future upstream-audit readers see 1:1 constants.
Path A is history, not an outstanding step (re-read
2026-09-23). B2 landed as 79ac5204 (2026-04-15); retries: 1 no
longer occurs anywhere in transport.rs, so the edit this step
describes cannot be made against the tree and the line it used to
cite now holds unrelated code. Where the mechanism lives today:
- the insert picks the start value from the source of the
announce —
PATHFINDER_RETRIESfor a local client,0otherwise — atretries(transport.rs:6826-6830); - the two guards are
PATHFINDER_RETRIES(transport.rs:11873) andlocal_rebroadcasts(transport.rs:11874), incheck_announce_rebroadcasts; PATHFINDER_RETRIES(constants.rs:157) is 1.
B1 fanout alignment
Question: if we remove exclude_iface, can dedup reliably
catch the self-echo, and does it play well with Python peers?
Resolution: yes. Outgoing broadcasts go through
send_on_all_interfaces (transport.rs:4292), which calls
self.storage.add_packet_hash() before emitting the
Action::Broadcast. The check that reads that set on arrival is
has_packet_hash (transport.rs:3956), in process_incoming. The only edge case is the
dedup window rollover at HASHLIST_MAXSIZE = 1 000 000 entries —
a packet that is ~1M packets old could theoretically come back.
Not a concern in practice for single-day bench runs.
Python interop subtlety (discovered 2026-04-15 when the B1
change was first landed, caused a 3-node TCP relay test to fail,
then resolved by spacing out the test's announce emissions): the
Python reference has a per-interface ingress control at
reference/Reticulum/RNS/Interfaces/Interface.py:117-138. When two
announces arrive on the same interface faster than
IC_BURST_FREQ_NEW = 3.5/s (≈ 285 ms apart), Python activates
burst mode for at least IC_BURST_HOLD = 60 s then penalises for
IC_BURST_PENALTY = 300 s. Held announces are released by
process_held_announces every interface_jobs_interval = 5 s,
but only once the cooldown expires.
In a LoRa topology the multi-second airtime per transmit naturally spaces announces below this threshold, so ingress control never activates. In a TCP relay topology a Rust node that receives announces from both peers in rapid succession — and with B1 fans them both out on every interface, with only the retry scheduler's 0-500 ms jitter spacing them — can trip Python's ingress control on the receiving side.
This is not a Rust bug; it is Python's intended rate-limit
behaviour that naive TCP-only tests can expose. The regression
guard test test_announces_forwarded_through_transport spaces
its two announce_destination calls by two seconds to keep
the spawned-peer interface's ia_freq below 3.5 /s. Production
scenarios where two daemons announce in tight succession through
a Rust relay remain subject to Python's ingress limits — exactly
as they would be through a Python relay.
Management announces get a second emission (2026-09-01)
Decision: a management destination's announce is emitted once
by send_on_all_interfaces (transport.rs:4292) and then once more
by schedule_own_announce_retry (transport.rs:4291), which inserts
an announce-table entry at retries = PATHFINDER_RETRIES so the
scheduler fires it exactly once and retires it. Landed as 2d9234da.
Why it is not a parity break: the reference gives the same two
emissions to any announce reaching rnsd from a shared-instance
client — inserted at retries = PATHFINDER_R and fired once more by
the job loop. Only an announce originated inside the transport
process misses out, and our management destinations live inside the
daemon. The deviation rule holds on all three clauses: the wire bytes
are the same announce re-emitted, a peer's packet-hash dedup already
absorbs the duplicate, and the P1 gain was measured on the residual
ble_lora_transport reds of 2026-09-01. The retry is cancelled as
soon as a neighbour is heard passing the announce on, so a healthy
mesh pays nothing. Full argument and citations in the doc comment on
schedule_own_announce_retry (transport.rs:4291).
Scope: management destinations only. An ordinary
Destination.announce() is still one-shot, matching Python exactly.
Mode-less Rust
Decision: Leviculum continues without interface modes.
Documented as a deliberate scope reduction. Our scenarios and the
Python peer we interop against all use MODE_FULL implicitly.
A future Bug that requires MODE_ROAMING or similar gets its
own task; this parity doc predates and outscopes that work.
14. Usage
This document is the audit target for both sides:
- When we upgrade the vendored
RNS/tree to a new upstream release, the Python line numbers here are the first thing to re-verify. A changed line number is a hint the behaviour may have shifted; a changed mechanism is a new parity task. - When we add a new broadcast code path to
leviculum-core, we extend the parity matrix (section 12) and add a test underleviculum-std/tests/rnsd_interop/that verifies the new path matches what a live Python peer sees.
The parity matrix is the contract. Everything else in this document is the reading behind the entries.
Hop counting
Why this document exists
The hop counter is one unsigned byte in the packet header. It is also load bearing. It decides which header form a packet takes, when a circulating packet is killed, which path replaces which, and whether a link proof is accepted. Two stacks that disagree about it cannot establish links with each other.
This page records the rules as the reference implements them, and where leviculum diverges. Every
claim cites a line in reference/Reticulum/RNS/Transport.py (or Packet.py / Link.py) so it can be
checked rather than believed.
The invariant
packet.hops counts the links a packet has traversed. Each node that receives it adds one,
including the receiving node itself. The IPC connection between a shared instance and one of its
local clients is not a link on the mesh and is never counted.
Life of a hop counter
1. Birth
Packet.py:135 sets self.hops = 0. It travels as header byte 1 (Packet.py:181 on pack,
Packet.py:245 on unpack). It is outside the signature, so a relay may legally change it.
2. Receipt
Transport.py:1457: packet.hops += 1, unconditionally, for every inbound packet.
3. The two IPC exceptions
Transport.py:1478-1484:
if len(Transport.local_client_interfaces) > 0:
if Transport.is_local_client_interface(interface): packet.hops -= 1
elif Transport.interface_to_shared_instance(interface): packet.hops -= 1
Read the structure carefully. The elif belongs to the OUTER if. A node that has local clients
(it IS a shared instance) subtracts only for packets arriving from a client. A node with no local
clients (it IS a client of some instance) subtracts for packets arriving from that instance. The two
branches are mutually exclusive. The net effect is that an IPC hop is free in both directions.
After this step the counter has a meaning that the rest of the stack relies on:
hops == 0the packet came from a local clienthops == 1the packet came from a direct neighbour
4. Announce rebroadcast
Transport.py:2009: new_announce.hops = packet.hops. The already incremented value goes back on
the wire. Each relay therefore contributes exactly one, never two.
5. Path table
Transport.py:1868: announce_hops = packet.hops, written to IDX_PT_HOPS at
Transport.py:2014. A path entry records the length of the route the ANNOUNCE travelled to reach
us. This is not necessarily the length of the route a packet to that destination will take. See
"What remaining_hops actually means" below.
6. Path acceptance
Transport.py:1765: if packet.hops <= Transport.path_table[dst][IDX_PT_HOPS]: and
Transport.py:2371: if announce_hops <= old_hops or time.time() > old_expires:.
A path is replaced only by an equal or shorter one, or once the old one has expired. This rule is what drives every node toward the same shortest tree, and it is why in a homogeneous mesh a stored hop count and a live route length agree.
7. Cache re-emission and path responses
Transport.py:326 and :379 increment a cached announce on reload, with the comment "reading a
packet from cache is equivalent to receiving it again over an interface".
Transport.py:2956: packet.hops = Transport.path_table[destination_hash][IDX_PT_HOPS] when
answering a path request, and Transport.py:618: new_packet.hops = announce_entry[4].
A path learned from a path response therefore inherits the responder's STORED count, not a freshly measured one. Staleness propagates through this channel.
leviculum matches this as of 2026-07-10 (D3, fixed on branch path-response-hops). When a transport
node answers a path request from a network peer (handle_path_request case 2b, transport.rs:6995)
it now emits self.storage.get_path(&requested_hash).map(|p| p.hops), the receipt-incremented stored
count, exactly as :2956 does. It previously emitted cached_packet.hops, the AS-RECEIVED wire byte
(set_announce_cache stores the raw pre-increment buffer; the receipt increment at transport.rs:2682
touches only the in-memory packet). That value is stored - 1, so every peer learning through our
transport path response was one hop short, and the deficit COMPOUNDED on each re-learn through a
leviculum transport. Case 1 (local dest) and case 2a (local-client answer, explicit +1) were already
correct.
8. Link table
Built at Transport.py:1615-1625, keyed by the link id:
| Index | Contents | Source |
|---|---|---|
3 IDX_LT_REM_HOPS | remaining hops | path_table[dst][IDX_PT_HOPS] (:1563) |
5 IDX_LT_HOPS | taken hops | packet.hops of the LinkRequest |
6 IDX_LT_DSTHASH | original destination hash | packet.destination_hash |
Note the trap: for a link packet, packet.destination_hash IS the link id (Transport.py:1498
looks the link table up with it). The address of the actual destination survives only at index 6,
and the healing loop below depends on it.
9. Link proof validation, the strict check
Transport.py:1656:
if packet.hops == link_entry[IDX_LT_REM_HOPS] or packet.hops == link_entry[IDX_LT_HOPS]:
and :1664 / :1668 use WHICH of the two matched to choose the forwarding direction. On the
local client link path, Transport.py:2176 applies the single == IDX_LT_REM_HOPS check. A proof
matching neither frozen value is dropped.
10. Endpoint check
Link.py:282 sets expected_hops = Transport.hops_to(destination), and Transport.py:2228 checks
packet.hops == link.expected_hops or link.expected_hops == PATHFINDER_M. The establishment timeout
also scales with hops (Link.py:207).
11. Loop bound
PATHFINDER_M = 128 (Transport.py:63). Transport.py:1750 requires
packet.hops < PATHFINDER_M + 1. The counter is the only thing that terminates a circulating
packet. Lowering it hands the packet extra life.
12. Header form for local clients
Transport.py:1356, :1367, :1565-1577. hops == 0 means the destination is directly
reachable, send Header1. hops == 1 means it needs transport, convert to Header2 and attach a
transport id. A counter that is off by one changes the packet form.
What remaining_hops actually means
It is the hop count of the route the ANNOUNCE took to reach this relay. It is frozen into the link
table when the LinkRequest is forwarded. The route the link then uses is chosen hop by hop by the
next_hop entry of every relay along the way. The two coincide only while all those relays agree
on the same tree. Rule 6 is what makes them agree in a homogeneous mesh.
Therefore a mismatch between packet.hops of a returning proof and the frozen remaining_hops is
not an arithmetic error. It is a statement that this relay's view of the topology disagrees with
the topology the packet actually traversed.
The control loop that makes strictness safe
The strict check of step 9 is not a bare guard. It is the SENSOR of a healing loop:
-
A proof whose hop count matches neither frozen value is dropped.
-
The link is therefore never validated, and expires (
Transport.py:693,LINK_TIMEOUT). -
clean_link_tablerequests a fresh path for the ORIGINAL destination (index 6), throttled byPATH_REQUEST_MI = 20seconds (Transport.py:83), under four conditions::710no path is known:717the failed link was initiated by a LOCAL CLIENT (lr_taken_hops == 0):726the destination was previously direct (hops_to(dst) == 1):748the initiator was direct (lr_taken_hops == 1)
and marks the path unresponsive (
Transport.py:2721) when transport is enabled. -
The path is relearned. The next attempt agrees, and the link establishes.
A stack that suppresses the drop also suppresses the LINK-FAILURE healing path. A relay that
rewrites a mismatching hop count so the proof is accepted makes the link succeed once and blocks
clean_link_table from ever re-requesting the path for that entry. It does NOT guarantee the entry
is never corrected at all: a fresh equal-or-shorter announce still replaces it via rule 6,
independent of the link-failure loop. So recurrence is a FIELD property (observed: a five-minute
heartbeat on hamster, 2026-07-10) — evidence that no corrective announce arrived, not a guarantee
the code makes it inevitable.
Where leviculum diverges
Recorded 2026-07-10 against reference/Reticulum as vendored.
| Rule | Reference | leviculum | Verdict |
|---|---|---|---|
| Receipt increment | :1498 | transport.rs:2498 | matches |
| IPC exception, instance side | :1523 | transport.rs:1750 | matches |
| IPC exception, client side | :1525 | transport.rs:1750 (else-arm of the has_local_clients gate) | matches — fixed 2026-07-10 (D2, commit 06aadaff); was absent |
| Announce rebroadcast | :2050 | transport.rs:7324 | matches |
| Path table store | :1909, :2055 | transport.rs:4360 | matches |
| Path acceptance | :1806, :2412 | transport.rs:4456 (should_update) | matches |
| Path-response hop emission | :2997 (packet.hops = path_table[dst][IDX_PT_HOPS]), :618 | transport.rs:6995 (case 2b emits the stored path-table count) | matches — fixed 2026-07-10 (D3, commit path-response-hops); previously emitted cached_packet.hops = the pre-increment wire byte (stored - 1) |
| Link entry fields | :1615-1625 | storage_types.rs:60 (destination_hash at :76) | matches, including the destination hash |
| LRPROOF relay check | :2215-2206 (single == remaining_hops, drop else; the :1697 disjunction is gated OUT for LRPROOF at :1687) | transport.rs:5223; rewritten by default, DROPPED behind lrproof_rewrite_on_asymmetry=false | deliberate deviation (default); the flagged strict branch drops like the reference, but see the mapping caveat below |
| Healing, no path | :737 | transport.rs:8536 | matches |
Healing, local client link (taken_hops == 0) | :744 | transport.rs:8827 | matches — fixed 2026-07-10 (D1, commit 74ac655); was absent |
| Healing, destination direct | :753 | transport.rs:8544 | matches |
Healing, initiator direct (taken_hops == 1) | :775 | transport.rs:7729 | matches |
The deliberate deviation, and its cost
On a mismatch we log a warning and REWRITE the forwarded proof's hop count to the frozen value, so
that a strict Python client accepts it (transport.rs:5247, commit 5d0833d7). It buys
interoperability today: without it, NomadNet cannot establish a link through our relay.
It also costs three things:
- It suppresses the sensor. The link validates,
clean_link_tableskips it (if entry.validated { continue; }), no path request is issued, and the wrong path survives. Measured in the field: the same warning recurs on an exact five minute heartbeat, indefinitely. Since #330 the second half of that sentence no longer holds: the wrong path does not survive, because a signature-validated proof now re-balances the path entry and the link entry in place (see "What we do since #330" below). The sensor is still suppressed and the sweep still asks for nothing — it no longer has anything to ask for. - It sometimes LOWERS the counter. Measured on miauhaus 2026-07-10:
packet_hops=7rewritten to3. That is four hops of extra life handed to a packet thatmax_hopswas meant to kill. - It overwrites a measurement with an assertion. Downstream consumers of
hopsreceive what this relay believes rather than what the packet did.
The rewrite must stay until the cause is fixed and the warning is shown to fall silent. What is MEASURED is that the warning recurs every ~300 s with the rewrite ON. That removing it would break NomadNet is an INFERENCE (drop -> strict client rejects the proof -> link fails), not yet a measurement: no flag-off live run has been done. Do not deploy the flag off without one.
The strict behaviour now exists behind a flag
The reference-exact strict check is implemented behind TransportConfig.lrproof_rewrite_on_asymmetry
(transport.rs), default true. The default keeps the rewrite above unchanged, so this is a no-op
in the field. Set to false, the forward site DROPS a proof whose hop count matches neither frozen
operand rather than rewriting it:
- the
next_hopdirection (destination -> initiator) drops unlesspacket.hops == remaining_hops. For a proof this maps toTransport.py:2176— the SINGLE== IDX_LT_REM_HOPScheck whose only else (:2206) drops. This is the arm the field case takes.
MAPPING CAVEAT (found by adversarial review 2026-07-10): the :1656 disjunction does NOT apply to
LRPROOF at all — its transit block is gated at Transport.py:1646 with packet.context != RNS.Packet.LRPROOF. So for proofs the reference has exactly ONE relay path (:2174-2206, single
check, drop else) and NO initiator-side LRPROOF forwarding. Our received-direction arm therefore
has no LRPROOF counterpart in the reference; it is practically moot (proofs flow
destination -> initiator), but it is a leviculum choice, not reference parity. Earlier drafts of
this page and a code comment mis-cited the :1656/:1664/:1668 arms for proofs — corrected.
The drop is the healing SENSOR. Whether the loop actually CLOSES is NOT yet established. The mvr
(mvr_hop_asymmetry.rs, flag off) shows the sensor fires — the proof is dropped, the link stays
unvalidated, and clean_link_table issues a path request — but its convergence step is CIRCULAR and
must not be read as proof of healing: the path request is discarded (handle_timeout() result
dropped, no node answers it), and the short arm is relearned only because the test HAND-FEEDS a
fresh announce. That same injected announce would heal the rewrite-ON world identically (rule 6),
so the mvr does not isolate the flag as the cause of convergence. In the field, a path RESPONSE
inherits the responder's STORED count (rule 7, "staleness propagates"), so a re-request can relearn
the SAME stale count and loop "fail, request, fail". Convergence is guaranteed only for one-level
divergence answered by the correct next hop. This is the open risk the interop A/B and a live
flag-off run must settle before the default can change.
The flag stays false-capable but true-default until an interop A/B and a live NomadNet-retry
check confirm the strict drop heals on the air as it does in the mvr; only then can false become
the default.
Upstream changed its mind: 1.5.x re-balances instead of dropping (Codeberg #330)
Everything above this line describes the reference as of 1.3.5, which is what
reference/Reticulum is pinned to and therefore what every interop test in this tree
measures against. RNS 1.5.x replaced the strict drop with a re-balance. The source facts,
read against 1.5.2 (ea98db4f, 2026-08-29) — every Python line number in this section is
1.5.2's and does NOT resolve inside the pinned reference/Reticulum, so the citation guard
cannot check it; re-read them against a 1.5.x checkout, never the submodule:
Transport.py:153—ALLOW_LINK_PATH_REBALANCE = True, a class constant, no config surface.- Relay site,
Transport.py:2614-2634. Whenpacket.hops != link_entry[IDX_LT_REM_HOPS]and the proof arrived onIDX_LT_NH_IF, the signature is validated first; if it is valid and the entry is not yetIDX_LT_VALIDATED, the relay ADOPTS the measurement —link_entry[IDX_LT_REM_HOPS] = packet.hops(:2632) andpath_entry[IDX_PT_HOPS] = packet.hopsfor the link's destination (:2634). Control then falls into the unchangedpacket.hops == IDX_LT_REM_HOPSforward arm, which now matches, so the proof is forwarded carrying its own true hop count. The re-balance happens at most once per link entry: the forward arm setsIDX_LT_VALIDATED, and the re-balance is gated on that flag being unset. That gate, not a hop equality, is 1.5.x's loop breaker here. - Terminus site,
Transport.py:2680-2707. For a pending link withpacket.hops != link.expected_hopsandstatus == PENDING, the signature is validated againstlink_id + peer_pub + peer_sig_pub + signalling_bytes; if valid andlink.rebalancedis unset,link.expected_hops = packet.hops(:2704) and the path entry's hops follow (:2707). The unchanged== expected_hopscheck then matches andvalidate_proofruns.Link.py:267-268adds the two fields;Link.py:525re-adoptsexpected_hopsfrom the RTT packet once the link is active. - No third site. The general link-table repeat arm is untouched, and LRPROOF is still excluded from it. The MAPPING CAVEAT above still holds in 1.5.x: the reference has exactly one relay path for proofs and no initiator-side LRPROOF forwarding.
What that means for the three sites we have:
- Cross-interface relay arm. We already deliver — that is the #38 rewrite. The difference
is not delivery, it is bookkeeping: 1.5.x heals
remaining_hopsand the path entry and then tells the truth on the wire; we heal neither and rewrite the wire instead. 1.5.x's re-balance is the healing loop this page says the rewrite suppresses, reached without the link having to fail first. - Terminus. We have no hop gate at all.
handle_link_proof(node/link_management.rs) checks phase, state and signature, never a hop count, andexpected_hopsdoes not exist anywhere inleviculum-core. So the half of #330 that reads "links over asymmetric paths form on Python but not on us" does not describe our initiator: ours accepts any hop count and always has. What ours does not do is 1.5.x's table healing. - Shared-medium arm. This is the one place we drop where 1.5.x forwards.
NH_IFandRCVD_IFare the same interface there, so 1.5.x'sreceiving_interface == IDX_LT_NH_IFtest passes and the re-balance arm fires (source read, not measured). We drop the proof as an echo, because on one medium the strict hop match is our only loop breaker — thelora_3node_relaystorm of 2026-08-12, pinned bymvr_lrproof_echo_storm.rs. Adopting 1.5.x here swaps that loop breaker for theIDX_LT_VALIDATEDgate. That is a rig question, not a desk one.
Why this is not a port. Forwarding the proof with its true hop count is exactly what a
1.3.5 initiator rejects: Transport.py:2228 in the pinned reference gates on
packet.hops == link.expected_hops. Adopting the relay site verbatim therefore re-opens #38
against every 1.3.5 peer in the mesh, and lrproof_hop_undercount_interop_tests.rs — which
drives a real Python initiator out of reference/Reticulum — passes today only because the
relay rewrites the count down to the frozen value. Upstream can do this because it fixed both
ends in the same release; we cannot assume both ends.
So #330 is a choice between three behaviours, and rule 5 below no longer decides it on its own now that "the reference" names two generations that disagree:
- a) keep the rewrite — links form for 1.3.5 and 1.5.x initiators alike, tables stay stale, we keep lying about the count on the wire;
- b) adopt 1.5.x verbatim — tables heal, the wire is honest, links through us stop forming for 1.3.5 initiators over asymmetric paths;
- c) heal the tables, keep the rewrite — correct the path entry from the proof's
measurement while still forwarding the frozen count, so the next link over that destination
freezes the right
remaining_hopsand the asymmetry drains within one link lifetime. Neither reference does this, so it is a deviation-rule argument and needs the deviation-rule evidence.
Deciding between them is a measurement, not a reading: the interop A/B this page already
demands for the strict flag, run against both a 1.3.5 and a 1.5.x peer. The fixture for the
relay half already exists — mvr_hop_asymmetry.rs builds the honest asymmetric topology and
asserts both arms of lrproof_rewrite_on_asymmetry — so a fix pass starts from a working
reproduction, not from scratch.
Decided 2026-09-26: (c). The next section records what was implemented, which half of the A/B was measured, and which half is still owed.
What we do since #330: option (c), measured on the 1.3.5 half
Implemented 2026-09-26, leviculum-core/src/transport.rs (relay arm) and
leviculum-core/src/node/link_management.rs (terminus arm). Read against the 1.5.0 tag
(e32d4df7), whose line numbers differ from the 1.5.2 ones quoted above:
Transport.py:150—ALLOW_LINK_PATH_REBALANCE = True.- Relay,
Transport.py:2540(if packet.hops != link_entry[IDX_LT_REM_HOPS] and Transport.ALLOW_LINK_PATH_REBALANCE:) with the adoption at:2555-2560(if peer_identity.validate(signature, signed_data) and not link_entry[IDX_LT_VALIDATED]:thenlink_entry[IDX_LT_REM_HOPS] = packet.hopsandpath_entry[IDX_PT_HOPS] = packet.hops). The 1.3.5 line it replaced isTransport.py:2176, the bareif packet.hops == link_entry[IDX_LT_REM_HOPS]:whose only else drops. - Terminus,
Transport.py:2608with the adoption at:2627-2637(link.rebalanced = time.time(),link.expected_hops = packet.hops,path_entry[IDX_PT_HOPS] = packet.hops). The 1.3.5 line it replaced isTransport.py:2228,if packet.hops == link.expected_hops or link.expected_hops == RNS.Transport.PATHFINDER_M:, which matched no pending link otherwise and letcreate_linktime out.
What we adopted, and what we did not:
- Adopted, both arms. On a hop mismatch whose Ed25519 signature holds, the proof's count
replaces the frozen one in the link entry (relay) or on the
Link(terminus), and the path entry for the link's DESTINATION follows. Onlyhopsmoves — not the interface, not the next hop, not the expiry, notlink_entry.hops. Preconditions are the reference's: the relay arm requires!validated(so a returning echo cannot move the count a second time) and a recalled peer signing key (Python reaches its rebalance throughIdentity.recall; without an identity it raises and adopts nothing). At the terminus the once-only property is structural: the link leavesPendingOutgoingon the same proof and the phase gate refuses every later one, which is what Python'slink.rebalancedflag buys. - Not adopted: the honest wire. The forwarded copy still carries the PRE-adoption frozen
count, the #38 rewrite. This is option (c) above and it is a deviation from 1.5.0, which
forwards
packet.hopsunchanged. The reason is measured, not inferred:lrproof_hop_undercount_interop_tests.rsdrives a real Python 1.3.5 initiator out ofreference/Reticulumbehind our relay over the asymmetric topology, and both of its cells pass with the adoption in place (2026-09-26). Forwarding the true count instead would hand that initiator a proof itsTransport.py:2228gate rejects — the initiator froze its expectation from the announce WE rebroadcast, i.e. from the stale count. A 1.5.0 initiator accepts the frozen count too: it equals what its own path table says, so its re-balance arm simply does not fire. - Not adopted: the shared-medium arm. Unchanged, still an echo drop. 1.5.0 would re-balance
there (its
receiving_interface == IDX_LT_NH_IFtest passes when the two interfaces are one) and bound the loop withIDX_LT_VALIDATEDinstead of the hop equality. Swapping our loop breaker for that one is a rig question — thelora_3node_relaystorm of 2026-08-12, pinned bymvr_lrproof_echo_storm.rs— and no desk argument settles it.
Deviation rule, clause by clause: the wire format is untouched (a hop byte, as before); semantic compatibility improves, because the set of initiators that establish through us over an asymmetric path is unchanged for 1.3.5 and unchanged for 1.5.0, while our own tables stop being wrong; and priority 1 gains the drain — the next link to that destination freezes the re-balanced count, so the asymmetry does not recur for the life of the path entry.
What this does NOT settle, and is still owed:
- The 1.5.x half of the interop A/B. Nothing in this tree runs a 1.5.x daemon
(
reference/Reticulumis pinned at 1.3.5 and every interop cell drives that), so "a 1.5.0 initiator accepts the frozen count" is a source reading, not a measurement. - The stale-downstream window. Once our path entry is re-balanced, the mismatch stops firing, so the rewrite stops firing with it — and a downstream 1.3.5 initiator whose own expectation is still the stale count now disagrees with what we forward. It re-agrees when the next announce from that destination reaches it through us. Between the re-balance and that announce, a link attempt from such a peer can fail where the pre-#330 rewrite would have papered over it. Python 1.5.0 has the same window and pays it in full (it never rewrites); we pay it only after the first successful link. Measuring it needs the 1.5.x A/B fixture above plus a second link attempt inside the window.
The fixtures are in mvr_hop_asymmetry.rs: the relay shape
(relay_adopts_validated_proof_hop_count_into_link_and_path), the terminus shape
(initiator_adopts_validated_proof_hop_count_into_link_and_path), and one negative control per
arm pinning that a forged signature adopts nothing.
The guard #330 needed: a proof that took the short way back does not move the route (#332)
#330 landed the adoption and asked for an mvr "before any change" that shows what the
adoption costs when the proof's route is not a shortening of the path entry's route but a
DIFFERENT route. Periculum pass 327 (2026-09-27,
the periculum tree's own report 2026-09-10-pathchoice-sweep, section 10) measured it in the
emulated pathchoice cells: twelve arms under measure, rnsd (1.3.5 in the containers)
carrying 8/8 transfers on every relayed arm, lnsd reading 7/8, 3/8, 4/8, 8/8, 8/8 at
L = 0.3/0.5/0.7/0.9/1.0. Ten of ten failed lnsd attempts had sent their link request over
the direct lossy pair; all 21 relayed link requests in the run belonged to attempts that
succeeded. The route moved without an announce:
PATH_ADD hops=2 next_hop=<bravo> reason="new_destination"
LINK_ENTRY_SET remaining_hops=2 <- attempt 1, relayed, ok
LRPROOF arrived dest=… iface=serial_0 hops=1
WARN LRPROOF hop asymmetry: rewriting forwarded hops to the frozen count … packet_hops=1 remaining_hops=2
event="PATH_REBALANCE" dst=… from=2 to=1
LINK_ENTRY_SET remaining_hops=1 <- every later attempt, direct
The mechanism is one line of arithmetic. PathEntry::needs_relay() is
hops > 1 && next_hop.is_some() (storage_types.rs:60), and it is the sole switch that puts
a transport header on an originated packet (transport.rs::send_to_destination,
route_via_transport; connect reads it too). Writing hops = 1 into an entry whose
next_hop still names the relay therefore does not shorten a route, it DELETES one: the relay
is still recorded, still the only way to the destination, and no longer addressed by anything
we send. Every later attempt is a coin toss on the pair that lost the first one. At L >= 0.9 no
proof crosses the pair at all, nothing rebalances, and the arm reads 8/8 — the damage is done
by the ONE frame that gets through.
The rule, as of #332: a rebalance may not adopt a hop count that turns needs_relay()
false while next_hop still names a transport peer. rebalance_path_hops
(leviculum-core/src/transport.rs) refuses such a count, leaves the entry untouched, reports
PathRebalance::HeldForNextHop to its caller and emits
PATH_REBALANCE_HELD dst= from= refused= next_hop= iface=. Both adoption sites go through that
one function, so the rule holds at the relay arm and at the terminus alike. Nothing else
changes: link_entry.remaining_hops and link.hops() still adopt, and the forwarded copy still
carries the frozen count (#38's rewrite).
What the entry should do with the information instead: nothing. The proof proves that one
frame crossed a route of that length, not that the route is ours to use — the path entry has no
interface and no next hop for it, and a rebalance has no authority to invent either, because a
route arrives by announce. Keeping hops = 2 keeps the entry internally consistent and keeps
the relay that has been delivering. The direct sighting is genuinely worth keeping, but where
route CHOICE can weigh it (#230's second-best route), not in the field that decides whether a
header is written; #332 deliberately does not build that.
Is this a mis-port or a deviation? A deviation — Python 1.5.2 has the same hole. Read
against /home/lew/coding/Reticulum at ea98db4f (1.5.2; these numbers do not resolve inside
the pinned 1.3.5 submodule):
- Relay site,
Transport.py:2632-2634:link_entry[IDX_LT_REM_HOPS] = packet.hops path_entry = Transport.path_table.get(link_destination) if path_entry: path_entry[IDX_PT_HOPS] = packet.hops - Terminus site,
Transport.py:2704-2707:link.expected_hops = packet.hops path_entry = Transport.path_table.get(link.destination.hash) if path_entry: path_entry[IDX_PT_HOPS] = packet.hops
Neither reads IDX_PT_NEXT_HOP, and neither clears it. And Python routes on the same predicate
we do. In its outbound path, Transport.py:1396-1429 (1.5.2), the line
if path_entry[IDX_PT_HOPS] > 1: inserts the transport header with
new_raw += path_entry[IDX_PT_NEXT_HOP], and the else that closes the chain
"know[s] the destination is directly reachable" and transmits HEADER_1. So a 1.5.2 node whose
2-hop entry is rebalanced to 1 stops addressing its relay for exactly the same reason ours did.
One difference is worth recording because it narrows Python's exposure without closing it: the
relay site is additionally gated on packet.receiving_interface == link_entry[IDX_LT_NH_IF]
(Transport.py:2615), so a proof that comes back on another interface rebalances nothing there.
On one shared carrier — the pathchoice cells, and any LoRa mesh — that test passes and the hole
is open. The terminus site has no interface test at all.
Deviation rule, clause by clause: wire unchanged (the proof is still accepted and still
forwarded with the frozen count, #38's rule stays; only the path table, which is ours alone,
declines a write); semantics unchanged for peers (no peer can observe a path entry; what a
peer observes is a relay that keeps being addressed, i.e. what it observed before #330);
priority 1 measurably served — the baseline is pass 327's K/8 transfer column above, and
the prediction for the reviewer's rerun is 8/8 on all five relayed lnsd arms.
The fixtures are in leviculum-core/src/node/mvr_link_proof_rebalance_next_hop.rs: the
single-carrier reproduction in the emulated cell's own shape (all three nodes on one interface,
bravo measured to forward the first request and to forward the second one too), the same
mechanism with the proof arriving on a second carrier (where the stranded request left on the
RELAY's carrier, not the direct one — the rebalance never moves interface_index either), and
the guard in isolation with its positive controls: a 3 -> 2 adoption that keeps the relay still
happens, and an entry naming no next hop still moves freely.
The ceiling, and what 1.5.x does at it
PATHFINDER_M is 128. In 1.3.5 that is a reachability limit and nothing else: a hop byte of 128 or
more is parsed, delivered and forwarded, and only the announce gate at Transport.py:1750 cares.
In 1.5.x it is also a parse limit — Packet.py:248 raises on a received hop byte of 128 or above
and Transport.py:1356 refuses to emit one — so the same byte that merely travels too far on a
1.3.5 peer is unreadable to a 1.5.x one. Our receipt increment can reach it from a legal wire
value, which is why the emit gates ask Transport::hop_ceiling() rather than config.max_hops.
Receipt stays liberal. The walk is in
Four things RNS 1.5.x changed, which also records what
local_hops_delta does to the meaning of hops == 0.
Rules to obey
- Never doctor a hop count to make a check pass. The check exists to expose a disagreement, and something downstream is listening for that disagreement.
remaining_hopsis not the length of the route a packet will take.- For a link packet,
packet.destination_hashis the link id. The original destination is a separate field. Do not use one where the other belongs. This mistake produced a silently useless diagnostic on 2026-07-10. - Any change to hop counting is checked against the reference first, and lands behind a test that fails before the change and passes after it.
- When the reference and leviculum disagree about a compatibility relevant mechanism, the reference is right. Where 1.3.5 and 1.5.x disagree with EACH OTHER, this rule names no winner: see the 1.5.x re-balance section above before invoking it.
Field evidence, 2026-07-10
Two relays, both running the same build, both logging both frozen counts and the interface branch.
hamster packet_hops=4 hops=0 remaining_hops=5 dir=next_hop (five times, one every 300 s)
hamster packet_hops=4 hops=0 remaining_hops=3 dir=next_hop
miauhaus packet_hops=7 hops=1 remaining_hops=3 dir=next_hop
miauhaus packet_hops=4 hops=1 remaining_hops=3 dir=next_hop
hops == 0 identifies a link initiated by a local client. Both signs of the mismatch occur, and the
magnitude reaches four. No constant per relay counting error can produce that, and the counting was
shown above to match the reference. What remains is the meaning of remaining_hops.
Announce dedup and path replacement (Python-RNS reference facts)
The reference facts behind the #376 desk measurement (a board's announce
reaches Columba only through the other board, never on the direct BLE
link). Four questions, answered strictly from the vendored
reference/Reticulum tree (Python-RNS 1.3.5), with citations, plus a
fifth section (added for Codeberg #231) on what 1.5.2 — the version a
periculum rnsd arm pins — does differently in the same arm. This page
states what the reference does; it decides nothing.
Sibling pages: Hop counting, Broadcast Python-RNS parity.
1. Direct vs. forwarded copy: same packet hash
Question. An announce received directly (header type 1, wire hops 0) versus the same announce forwarded by a transport node (header type 2, wire hops 1, transport id inserted): same packet hash?
Answer: yes, the hash is identical. The hashable part is built by
get_hashable_part (reference/Reticulum/RNS/Packet.py:355):
def get_hashable_part(self):
hashable_part = bytes([self.raw[0] & 0b00001111])
if self.header_type == Packet.HEADER_2:
hashable_part += self.raw[(RNS.Identity.TRUNCATED_HASHLENGTH//8)+2:]
else:
hashable_part += self.raw[2:]
return hashable_part
Three exclusions make the two copies hash the same:
- The hop count is excluded entirely. It is header byte 1 —
hops(reference/Reticulum/RNS/Packet.py:245) — and both branches above start at byte 2 or later. - The transport id is excluded. For header type 2 the slice starts
after the 16-byte transport id (
TRUNCATED_HASHLENGTH//8 + 2= 18). - The header-type and transport-type bits are masked off. Byte 0 is
packed as
packed_flags(reference/Reticulum/RNS/Packet.py:171) —header_type << 6 | context_flag << 5 | transport_type << 4 | destination.type << 2 | packet_type— and the& 0b00001111mask keeps only destination type and packet type. Header type (bit 6), transport type (bit 4) and the on-air IFAC flag (bit 7) all vanish.
So a relay changing hops, inserting its transport id and flipping the
header to type 2 does not change the packet hash: the two copies are the
same announce to every dedup structure keyed on packet_hash.
2. The second copy in Transport.inbound: order, and the announce exemption
Order. The dedup check runs first, the path-table update later, and both copies pass the dedup:
packet.hops(reference/Reticulum/RNS/Transport.py:1457) is incremented for every inbound packet, right after unpack.- The filter runs —
packet_filter(reference/Reticulum/RNS/Transport.py:1486) gates all further processing. - On acceptance the hash is remembered at once —
add_packet_hash(reference/Reticulum/RNS/Transport.py:1506) — before any announce processing (deferred only for link-table traffic and LR proofs). - Announce processing, including the path-table update, comes much
later in the same call —
validate_announce(reference/Reticulum/RNS/Transport.py:1691).
The exemption. Inside packet_filter
(reference/Reticulum/RNS/Transport.py:1336), a packet whose hash is
already in packet_hashlist
(reference/Reticulum/RNS/Transport.py:1376) is dropped — except an
announce for a SINGLE destination, which is accepted anyway:
if not packet.packet_hash in Transport.packet_hashlist and ...: return True
else:
if packet.packet_type == RNS.Packet.ANNOUNCE:
if packet.destination_type == RNS.Destination.SINGLE:
return True
So the second copy of the same announce is not discarded by the hashlist. It runs the full announce path again, and what it may change is decided there, by the random-blob replay check and the hops comparison of §3 — not by dedup. For #376 this matters in both directions: hashlist dedup cannot explain a missing direct announce, and hearing the relayed copy first does not inoculate the node against the direct copy.
3. The path replacement rule
All of this sits under the hop cap and non-local condition —
PATHFINDER_M (reference/Reticulum/RNS/Transport.py:1750) — with
announce_emitted (reference/Reticulum/RNS/Transport.py:1753) the
emission timestamp read out of the announce's random blob (bytes 5..10;
announce_emitted, reference/Reticulum/RNS/Transport.py:3191), and
the table side aggregated as the maximum over the recorded blobs
(timebase_from_random_blobs,
reference/Reticulum/RNS/Transport.py:3182).
Unknown destination: added unconditionally — should_add
(reference/Reticulum/RNS/Transport.py:1831).
Fewer or equal hops than the table entry — path_table
(reference/Reticulum/RNS/Transport.py:1765):
if packet.hops <= Transport.path_table[packet.destination_hash][IDX_PT_HOPS]:
path_timebase = Transport.timebase_from_random_blobs(random_blobs)
if not random_blob in random_blobs and announce_emitted > path_timebase:
should_add = True
(path_timebase, reference/Reticulum/RNS/Transport.py:1772.) Two
conditions, both required:
- the random blob must be new — the same announce heard again, e.g. the direct copy after the relayed copy, has the same blob and does not replace the path, however many hops it saves;
- the emission timestamp must be strictly newer than the newest one recorded for the destination. A later announce whose clock is behind the recorded one loses even at fewer hops — which is why the #376 announce instrument keeps the telemetry path's clock gate.
More hops than the table entry: ignored, unless one of three escapes fires, in order —
- the path has expired:
path_expires(reference/Reticulum/RNS/Transport.py:1793), still requiring an unseen blob; - the emission is strictly newer:
path_announce_emitted(reference/Reticulum/RNS/Transport.py:1809), still requiring an unseen blob; - same emission, but the recorded path has been marked unresponsive:
path_is_unresponsive(reference/Reticulum/RNS/Transport.py:1822).
4. No hops-0 drop rule
Question. Does Python drop or downgrade announces received with hops 0 on any interface class — anything that could make Columba ignore a direct board announce?
Answer: no such rule; searched Transport.inbound. There is no
condition anywhere in Transport.inbound that keys on hops == 0 for
an announce or treats a directly received announce worse than a relayed
one. (packet.hops, reference/Reticulum/RNS/Transport.py:1457,
increments every inbound packet before any decision, so a direct
announce is processed at hops == 1; the only decrements are the two
shared-instance IPC cases — is_local_client_interface,
reference/Reticulum/RNS/Transport.py:1482 — which are not radio
interfaces.)
The gates that do exist on the way to the path table, for any hops value, are:
- Signature validation —
validate_announce(reference/Reticulum/RNS/Identity.py:532): fails, and the announce is silently dropped. - Ingress limiting —
should_ingress_limit(reference/Reticulum/RNS/Transport.py:1705): on an interface withingress_controlenabled (should_ingress_limit,reference/Reticulum/RNS/Interfaces/Interface.py:145), an announce for an unknown destination can be held —hold_announce(reference/Reticulum/RNS/Transport.py:1706) — rather than processed; a pending path request bypasses the hold. - PLAIN/GROUP announces are always invalid
(
packet_filter,reference/Reticulum/RNS/Transport.py:1336);lxmf.deliveryannounces are SINGLE and unaffected. - The §3 replacement conditions, in particular the strict emission-timestamp comparison.
Sections 1 to 4 recorded 2026-09-09 against reference/Reticulum as
vendored (1.3.5).
5. 1.5.2's same-emission arm: gravity, not hops
§3 answers for 1.3.5, the vendored tree. It is not the whole answer for
the Python a comparison run actually faces: periculum's rnsd arm pins
1.5.2 (rns_pin in every pathchoice_*_rnsd.json), and 1.5.2 has an
acceptance path on EQUAL emission that 1.3.5 does not. Codeberg #231's
readings 308 and 309 disagreed because each had read one of the two
versions, so both are recorded here.
The line references in this section are deliberately NOT written in
the path:line citation form. 1.5.2 is not vendored — the tree read
was ~/coding/Reticulum at _version.py 1.5.2 — so the citation guard
would resolve a Transport.py line citation written here against the
1.3.5 copy and pass it on existence alone, a green that means nothing
(docs/src/concepts/checks-and-citations.md §"could not be checked").
Prose line numbers are checkable by a reader and by nothing else, which
is the truth about them.
The outer gate is unchanged — packet.hops <= the known hop count — and
the first test inside it is the 1.3.5 one, verbatim: 1.5.2
Transport.py lines 2237-2238, against
reference/Reticulum/RNS/Transport.py:1771-1772 here. What 1.5.2 adds
is the else, at 1.5.2 Transport.py lines 2245-2252:
# If the same announce is received later on an interface
# with higher gravity, allow updating the path table to
# use this interface instead.
if announce_emitted != path_timebase: should_add = False
elif announce_gravity == None or current_gravity == None: should_add = False
else:
if announce_gravity <= current_gravity: should_add = False
else:
... should_add = True
Three facts follow:
- Hop count is not in the acceptance test, in either version. It
appears only in the
packet.hops <=gate above. A second copy of an emission already installed does not replace the path for having taken fewer hops in 1.3.5 or in 1.5.2. - What 1.5.2 does accept on equal emission is a higher-
gravityinterface.gravityis an operator-configured per-interface preference — 1.5.2Reticulum.pylines 798-799 read it out of the interface stanza, and_add_interface(lines 1134-1136) fills thedefault_gravitywhen the stanza names none. It is a configured ranking, not a measurement. - On a default config 1.5.2 behaves exactly like 1.3.5 here.
1.5.2
Interfaces/Interface.pyline 75 setsDEFAULT_GRAVITY = 0, soannounce_gravity <= current_gravityholds and the arm rejects. Anrnsdarm that sets nogravityanywhere is measuring the strict 1.3.5 rule — and that is everypathchoice_*cell: nogravitykey appears anywhere under periculum'semulated/,conformance/orpericulum/adapters/, and each node in those cells has exactly one interface to begin with.
Recorded 2026-09-27 against ~/coding/Reticulum at 1.5.2. The nearest
thing we have to this arm is our own same-emission clause
(transport.rs, SameEmissionRule, Codeberg #231), which ranks by hop
count rather than by a configured preference.
6. The measured same-emission arm: path_choice = hops_and_loss
Since 2026-10-05 (Codeberg #230) the same-emission clause has a second,
opt-in rule beside the default hop ranking: path_choice = hops_and_loss (the [reticulum] key; TransportConfig::path_choice,
PathChoice in transport.rs). Where §5's 1.5.2 gravity re-ranks
equal-emission copies by an operator-CONFIGURED preference, this rule
re-ranks them by a MEASURED one: the direct link is priced at its
bidirectional ETX, estimated for free from the announce re-hear ratio —
of the emissions this transport provably learned about, the fraction
that also arrived as a direct copy (DirectLinkEvidence,
lower_route_cost). A clean two-hop route then beats a direct link
losing more than ~29 % of announces. Three properties bound it:
- Scope. Only the same-emission comparison changes. A newer emission still installs unconditionally in both hop arms (§3's rules, which is also the liveness guarantee when the relay dies), and the random-blob replay gate lets a same-blob copy through exactly when the cost arm will accept it. Path responses and local-client copies feed no evidence: a response is served once over one route, so the absence of a direct copy of it proves nothing.
- Degradation. Below
DIRECT_EVIDENCE_MIN_KNOWNknown emissions — a fresh destination, or a mesh whose announce cadence is too sparse to score — the rule decides exactly likehops. The default IShops; the knob is a deliberate opt-in and rnsd ignores the key in a shared config. - Evidence. The pinned replay
mvr_pathchoice_rule_table(leviculum-core, node/) drives both knob settings through the real acceptance gate over thepathchoice_loss*room shape at five loss levels and eight seeds: identical at 0 % loss, more delivered and fewer flaps at every measured loss ≥ 30 %, and the Python reference row (§3/§5 semantics as a pure fold) matches the shipped rule exactly, which pins the replay to the gate. The companionmvr_pathchoice_direct_vs_routeholds the norelay control: with no competing route the penalty can never unroute the only path.
When a transport node repeats a data packet
Established from reference/Reticulum while fixing #383, where a node that
was neither sender nor destination repeated all 30 probes it overheard on a
shared medium and buried 24 of the 30 answers.
1. Which data packets the reference forwards at all
Transport.inbound reaches its path-table forwarding branch only through this
condition (RNS/Transport.py:1559-1560):
if packet.transport_id != None and packet.packet_type != RNS.Packet.ANNOUNCE:
if packet.transport_id == Transport.identity.hash:
if packet.destination_hash in Transport.path_table:
Three gates in order: the packet must carry a transport id, that id must be
this node's own identity hash, and a path to the final destination must be
known. The complementary case is handled earlier, in
Transport.packet_filter (Transport.py:1341-1344): a non-announce packet
whose transport id names a different instance is rejected before any of this
runs. So a transport node acts on exactly one class of overheard traffic, the
class addressed to it by name.
Once inside, the hop count only selects the header rewrite
(Transport.py:1567-1581): remaining_hops > 1 keeps HEADER_2 and swaps in
the next hop, remaining_hops == 1 strips back to HEADER_1, remaining_hops == 0 just bumps the count. All three transmit.
A packet addressed directly to a destination one hop away is never handed to
a transport node. The sender decides this, in Transport.outbound
(Transport.py:1134-1166): a path table entry with hops > 1, or hops == 1
while the sender is behind a shared instance, gets the HEADER_2 transport
header with the next hop written into it. Anything else falls through to
# If none of the above applies, we know the destination is
# directly reachable, and also on which interface, so we
# simply transmit the packet directly on that one.
which puts HEADER_1 on the air with no transport id. That packet fails the
very first gate at Transport.py:1559 on every node that hears it, so a
neighbour holding its own path to the destination repeats nothing. Holding a
path is not an invitation to forward; being named is.
The one exception, and it is not really an exception: if the destination sits
behind a local client of a shared instance, the previous hop stripped the
transport id (clients are made to look directly reachable), so the instance
synthesizes it back before the gate (Transport.py:1543-1548):
if packet.transport_id == None and for_local_client:
packet.transport_id = Transport.identity.hash
for_local_client is a path table entry at hops == 0 (Transport.py:1513).
2. What the reference does repeat without being named
Two mechanisms, and a fix to the path-table gate must leave both alone.
Link table (Transport.py:1648-1686). Packets addressed to an established
link's id carry no transport id at all, and the relay repeats them purely off
its link_table entry. This is not overhearing: the entry exists only because
this node forwarded the LINKREQUEST earlier, as its designated next hop
(Transport.py:1625). The same-interface case is explicit about repeating
back onto the medium the packet arrived on:
# If receiving and outbound interface is
# the same for this link, direction doesn't
# matter, and we simply repeat the packet.
gated on the taken hop count matching one of the two frozen counts, which is what stops the relay-to-relay echo on one channel.
The hop counts are the whole gate. IDX_LT_VALIDATED is not consulted here
-- expiry is its only reader (Transport.py:687) -- so a relay that forwarded
the LINKREQUEST but lost the returning LRPROOF still carries the link's data.
Gating the repeat on it instead would strand a link its two endpoints consider
established, on one relay's RF luck (Codeberg #228).
Announces. Rebroadcast is transport-id-independent by design; the
packet_filter exemption at Transport.py:1342 exists for it.
3. Shared versus point-to-point
Nothing in any of this inspects whether the interface is a shared medium. The gate is a property of the packet, not of the carrier, and it produces the right behaviour on both: on a point-to-point link the only node that hears the packet is the one it was sent to, so the gate never fires; on a shared medium it is the only thing standing between one probe and N repeats.
The one interface comparison in the area, link_entry[IDX_LT_NH_IF] == link_entry[IDX_LT_RCVD_IF], compares two stored interface indices of one link
entry. It asks whether a relayed link happens to enter and leave by the same
interface, not whether that interface is shared.
Consequences for us
leviculum-core/src/transport.rs, handle_data, now gates the path-table
forward on transport_id == Some(own hash), with the for_local synthesis
arm. The LINKREQUEST path already carried the same gate (designated_hop),
added for the LRPROOF echo storm; data never got it.
Path::needs_relay() (storage_types.rs:60, hops > 1 && next_hop.is_some())
is not the predicate for this and never was. It answers how to rewrite the
header of a packet already accepted for forwarding, mirroring the
remaining_hops split above. Gating acceptance on it would kill the last hop
of every chain: in A-B-C where A cannot hear C, B is correctly the designated
hop and B's path onward to C is exactly one hop, so needs_relay() is false
for the very packet B must repeat.
Pinned by leviculum-core/src/node/mvr_overheard_direct_data.rs, whose
control 2 is that chain packet.
Four things RNS 1.5.x changed, and what each one costs us
Why this document exists
Codeberg #331 listed four differences read out of the 1.5.0 source diff. None of them was an observed break; each was an assumption about our side that nobody had checked. This page records the check. One of the four was a real defect and is fixed; two were already covered and now have tests saying so; one is latent and this page is the whole of the answer.
The reading was done against /home/lew/coding/Reticulum at 1.5.2 — reference/Reticulum in this
tree is pinned at 1.3.5 and does not contain any of the four changes. Line numbers below name the
generation they belong to, because on all four points the two generations disagree with each
other.
1. Keepalives on last_outbound — not our gap
What changed. 1.5.2 Link.py:749 widened the watchdog gate:
if now >= last_inbound + self.keepalive or now >= self.last_outbound + self.keepalive:
1.3.5 Link.py:792 has only the first half.
What it is actually for. The stale test inside that branch is now >= last_inbound + self.stale_time in both generations, and it is last_inbound-only in both. So the wider gate does
not make a 1.5.x destination tear a link down sooner — it cannot; entering the branch earlier with
fresh inbound just sends a keepalive and sleeps. What it fixes is the opposite end. A 1.3.5
INITIATOR that receives continuously and transmits nothing keeps its own last_inbound fresh,
never enters the branch, and never sends a keepalive; meanwhile the destination's last_inbound
ages out and the destination — of either generation — declares the link stale. 1.5.x closed that by
making the initiator's own silence a trigger.
Where we stand. Link::should_send_keepalive (leviculum-core/src/link/mod.rs) gates on
last_keepalive alone: an active initiator emits one every interval regardless of traffic in
either direction. That is a superset of both Python gates, so neither the 1.3.5 bug nor the 1.5.x
change can reach us. It is a superset by accident of a simpler rule, though, so
silent_initiator_keeps_sending_keepalives_while_inbound_is_fresh now pins it: a future "only
keepalive when idle" optimisation turns that test red before it reaches a peer.
2. Stream chunks six bytes larger — already our size
What changed. 1.3.5 Buffer.py:229 sizes a chunk channel.mdu - StreamDataMessage.OVERHEAD
with OVERHEAD = 2 + 6 (:56). The six is the channel envelope header, which channel.mdu
(Channel.py:652) has already subtracted — so 1.3.5 charged it twice. 1.5.x subtracts the
two-byte stream header only, and its chunks are six bytes larger, landing on the link MDU exactly.
Where we stand. max_data_len (leviculum-core/src/link/channel/buffer.rs:63) is
channel_mdu - STREAM_DATA_HEADER_SIZE. We already write 1.5.x-sized chunks and always have; the
larger chunk is not new traffic to us, only newly observable. On the read side
RawChannelReader::receive has no length gate at all — it appends whatever arrives — so the only
place a size could be refused is the len > mdu guard in Channel::send_raw, which is inclusive
at the boundary. a_1_5_x_sized_stream_chunk_fits_the_channel_and_reads_back_whole drives the
largest chunk 1.5.x can build through send, receive, unpack and reassembly and compares the bytes.
3. hops >= 128 rejected as malformed — we could emit it
What changed. 1.5.2 rejects the value at both ends. Packet.py:248:
if self.hops >= RNS.Transport.PATHFINDER_M:
raise ValueError(f"Invalid hop count {self.hops}")
That runs in unpack, before the header type is read, so the packet is not over-ranged — it is
unreadable, and nothing downstream sees it. Transport.py:1356 declines to emit one:
if packet.hops > Transport.PATHFINDER_M-1: return False. Neither check exists in 1.3.5, which
both accepts and forwards such a packet.
Where we stood. Our emit gates compared packet.hops > self.config.max_hops with max_hops = PATHFINDER_MAX_HOPS = 128, and packet.hops is the receipt-incremented count
(incoming_hop_count, transport.rs). A relayed packet arriving with a wire byte of 127 therefore
became hops = 128, passed a > 128 gate, and went back out stamped 128 — the first value a
1.5.x neighbour refuses to parse. Six emit sites were reachable this way: the three forward arms,
the capped announce broadcast, the targeted path response, and the shared-instance hand-off to
local clients (where a 1.5.x client raises on exactly the same byte).
What changed here. Transport::hop_ceiling() returns
config.max_hops.min(PATHFINDER_MAX_WIRE_HOPS) — the operator's reachability limit and the wire
ceiling, whichever binds first — and all six sites ask it.
forward_never_stamps_a_hop_byte_a_1_5_x_peer_cannot_parse is the reproduction: pre-fix it
forwards with fwd[1] == 128, post-fix the packet is dropped and accounted as forward-max-hops.
forward_still_relays_one_below_the_wire_hop_ceiling holds the other side, so the fix cannot
degrade into "stop forwarding".
Receipt is deliberately untouched. We still accept and deliver a hop byte 1.5.x would refuse.
Being liberal in what we accept costs a 1.3.5 peer nothing; being liberal in what we emit costs a
1.5.x peer the packet. Against the deviation rule: wire compatibility improves, semantic
compatibility is unaffected (no peer can expect delivery along a 128-hop path — every peer drops
announces above PATHFINDER_M), and the packets shed are unroutable or looping traffic.
4. local_hops_delta and the ephemeral transport identity — latent, and what to watch
Two 1.5.x options with no 1.3.5 counterpart.
local_hops_delta. Off by default (Reticulum.py:257); when the config option at
Reticulum.py:515 is set, Transport.py:337 draws (urandom(1) % 6) + 2 once per boot and stamps
it in place of the true hop byte on packets the node originates (Transport.py:1401, :1421,
:1435, :1596, gated by should_apply_delta at :1609, which requires packet.hops == 0 and
no shared instance). So on a mesh with it enabled, hops == 0 no longer means "origin" and a
delta-enabled neighbour reads as two to seven hops further away than it is.
What that touches here: nothing that breaks, and nothing that is right either. Our hops == 0
tests are all path-table entries meaning "local client behind the shared instance"
(transport.rs:10080, :10539) or our own locally-created announces (:10210) — neither is a
remote node's claim about itself, so neither can be lied to. The cost is metric only: a
delta-enabled peer loses every path race against an honest one, and its
ESTABLISHMENT_TIMEOUT_PER_HOP scaling is drawn from a fiction. There is no fix to make, only a
thing not to assume: a mixed-mesh hop-count measurement is meaningless unless
local_hops_delta is known to be off on every Python node in it, and it is not observable from
outside.
Ephemeral transport identity. 1.5.2 Transport.py:332-335: a node with transport disabled and
static_transport_identity unset — the default for every non-transport 1.5.x node — replaces its
stored Transport.identity with a fresh RNS.Identity() at every start. Its transport hash is
therefore per-boot.
What that touches here: we hold no remote transport id anywhere that survives our own restart.
There is no destination_table on disk — the path table is in memory and refreshed from announces
— and no code path compares a received transport_id against a remembered one; the only equality
test on a carried transport id is against our OWN hash (transport.rs:3737), to decide whether a
transport-routed packet is addressed to us. A rebooted peer's stale via entries are the ordinary
stale-path case, which announce refresh and PATHFINDER_EXPIRY_SECS already cover.
The thing that does break is a test fixture. Any interop or rig scenario that records a Python peer's transport hash in one run and expects it in the next is now reading a per-boot value; the symptom is an unplaceable identifier in a log, which is exactly the shape the null-hypothesis rule in CLAUDE.md was written for. Check the hash against the run's own nodes before treating it as a stranger.
What is NOT settled here
The fifth 1.5.x difference — the LRPROOF hop re-balance — is a design decision, not an audit result, and lives in Hop counting under #330. This page does not reopen it.
Testing — Developer Quick Reference
One-page orientation. See CI Pipeline for the automation details.
TL;DR
- Writing code: run
cargo test -p <crate-you-touched>as you go. - Committing: nothing happens. No hook tests a commit.
- Per batch of work: run
just standard(Tier 1) yourself. Nothing starts it for you. - Pushing: Tier 0 runs automatically and blocks on fail.
- Daily: Tier 3 (02:00) runs via systemd. Tier 2 is on demand — nothing starts it for you.
- Suite overview:
just status.
Prerequisites (system tools)
The interop and integration tests shell out to real external tools, so these system packages must be installed (all apt-installable on Debian):
- docker — the Tier 2 scenario corpora (run by periculum).
- python3 + Python RNS (
rns) — thernsd_interoptests drive a real Pythonrnsd/rnstatusas the compatibility reference. - socat — bridges a virtual serial pty pair so the serial-family interfaces (KISS, AX.25, Pipe) can be interop-tested against a real Python peer.
- nomadnet (optional) — the real NomadNet node used by the on-demand lnomad
acceptance (
scripts/lnomad_nomadnet_acceptance.sh, see The lnomad acceptance). Not part of any tier; install withpip install nomadnet. - i2pd (optional) — provides the SAM bridge on 127.0.0.1:7656 that the
I2PInterfacelive tests use. The default suite coversI2PInterfacewith an in-process mock SAM bridge, so i2pd is not needed to go green; it only gates the#[ignore]d live tests (cargo test -p leviculum-std i2pd_live -- --ignored). Enable the bridge withsam.enabled = truein/etc/i2pd/i2pd.confand start thei2pdservice. - cargo-fuzz + nightly (optional) — drive the wire-format parser fuzz
harness under
leviculum-core/fuzzandleviculum-std/fuzz(see Fuzzing the wire parsers). Not part of any tier; install withcargo install cargo-fuzz && rustup toolchain install nightly. - just, cargo, flock, notify-send — build/CI plumbing.
scripts/install-ci.sh checks for these at setup and prints the sudo apt install hint for any that are missing. Whenever a test starts depending on a
new tool, add it BOTH here and to that check list so a fresh machine can be set
up from scratch.
The four tiers
| Tier | When | Command | Time | Scope |
|---|---|---|---|---|
| 0 | on git push (hook) | just fast | ~3 min | fmt + clippy + workspace lib tests |
| 1 | on demand, once per batch1 | just standard | ~15 min (40 min cold2) | Tier 0 + core/tests + ffi + proxy + rnsd_interop |
| 2 | on demand: systemctl --user start leviculum-ci-tier2.service3 | just extensive | 30–90 min | Tier 1 + periculum conformance/ + regression/ |
| 3 | 02:00 daily (systemd timer) | just nightly | 2–6 h | Tier 2 + LNode flash-from-HEAD + periculum hardware/ |
Each tier includes every lower tier, so a green nightly proves the whole stack.
just guards is not a fifth tier: it is the coder pass's standing gate —
run beside cargo fmt, cargo clippy -D warnings and
cargo test --workspace after every batch — and it is the subset of Tier 0's
guards, censuses and selftests that costs under ten seconds each, about 20 s
warm in total. fast and standard stay the landing gate's, unchanged: a
green guards is an early verdict on part of Tier 0, never a substitute for
it. It exists because the recipes in it are the ones a coder never ran and
the landing gate did, which cost two landings an hour each on 2026-09-26
(check-supervised-spawns, check-source-invariant-census).
The fuzz crates' committed Cargo.lock files are a guards precondition
too (check-fuzz-lock, 0.1 s, Codeberg #442), because the land gate up to
bc07cee3 went red in fuzz-regress on a lock that 8042e34a's new path
dependency had made stale, with every fuzz target green.
That the list really is a subset is checked and not promised:
check-guards-subset (in guards itself) reads just --dump and refuses a
guards member fast never reaches, a member the two lists run in different
orders, and a member of fast that is in neither guards nor the ledger of
not-in-guards: reasons in the Justfile.
Results go to ~/.local/state/leviculum-ci/last-results.txt:
GREEN = passed, RED = failed, SKIPPED = deferred because another
test held the lock (see "Concurrent runs" below). No tier raises a
notify-send alarm today — read the ledger, or just status. See
CI Pipeline.
While writing code
Fast feedback. Run only what you changed:
cargo test -p leviculum-core --lib # touched core lib code
cargo test -p leviculum-std # touched std
cargo clippy -p leviculum-core # clippy for one crate
cargo fmt # apply formatter (not --check)
End of a batch of work:
just standard # Tier 1 (~15 min)
This is the "15-minute-budget" check that CLAUDE.md expects after every task, and typing it is the only thing that runs it.
Never run a full scenario corpus casually:
# DON'T do this without a reason — the containers and the USB handles
# collide with anything else running scenarios on the box.
periculum run ../periculum/conformance
If you must, use just extensive, which builds the binaries the nodes
mount first and is lock-protected.
Before pushing
Nothing to type. git push triggers .githooks/pre-push, which lints
the Woodpecker pipelines (.githooks/pre-push:204) and then runs
just fast (Tier 0, .githooks/pre-push:207). A red Tier 0 aborts the
push — fix, commit, and push again.
Three cheap guards run before those minutes are spent. Two are about
what reaches a public forge: only master and tags go to Codeberg
(.githooks/pre-push:31), and no commit carrying CLAUDE.md,
.mcp.json or .claude/ goes there at all. The third is about the
gates themselves — they test the working tree, so the working tree has
to be what is being pushed. A tree with uncommitted tracked changes
(.githooks/pre-push:144), or a push that would move master to
anything but HEAD (.githooks/pre-push:162), is refused before the
first gate starts: otherwise the verdict describes code that is not
being pushed, in either direction. Untracked files are exempt; they are
in no commit.
To push a sha this tree is not standing on, run
scripts/push-clean.sh <sha> [<remote>]. It keeps a clone outside the
tree ($LEV_PUSH_TREE, default ~/.cache/leviculum/push-tree), checks
the sha out there detached, initialises the submodules, points that
clone's remote at the URL this repository uses for it, and — the part a
hand-written git clone silently omits — sets core.hooksPath, so the
gates actually run on what is being pushed. The clone keeps its build
cache between runs.
just fast runs the guards and that script against scratch repositories
(just prepush-guard, ~0.3 s), so a guard broken by an edit is caught by
the next push instead of by the push it wrongly refuses. One of its cases
pushes from a clone with core.hooksPath unset and requires the proof to
fail there: a push that arrives is no evidence a hook ran.
The rest of the hook: it used to also block on Tier 2 staleness, at
5 commits/8 h (warn) and 10 commits/24 h (block); the block was
unsatisfiable and was removed on 2026-08-07, along with the
git push --no-verify habit it taught. See
CI Pipeline.
After committing
Nothing. There is no post-commit hook — deliberately, since
2026-08-07 (footnote 1 above). Tier 1 is just standard, typed once
per batch.
Checking state
just status # last result per tier
just logs # tail most recent Tier 1 log
cat ~/.local/state/leviculum-ci/last-results.txt # full history
LoRa hardware tests
Tier 3 only. Requires two Heltec T114 boards + two RNode radios connected via USB. Manual runs:
just flash # flashes ALL attached T114s; touch-free
# since the Bug #13 firmware change.
# Double-tap RESET only if the runner
# prompts you (crashed-firmware fallback).
just flash-one /dev/ttyACM3 # flash one specific T114 (A/B testing)
just nightly # full Tier 3 run
A single LoRa scenario in isolation:
periculum run ../periculum/hardware/<name>.toml
Hardware scenarios are not gated behind a flag: they live in
hardware/, and periculum decides from the scenario itself whether
this bench can serve it. One that binds a board the bench does not
hold reports SKIPPED_INFRA naming what was missing, never RED.
Radio duty-cycle lock is OFF by default in tests
The harness writes airtime_limit_long = 0 into every generated
radio interface (single RNode, multi-vport RNode, serial LNode), so
the firmware duty-cycle airtime lock never engages mid-run. Without
this, the driver's lawful-by-default ETSI cap (#55) silently stops a
saturating sender once its rolling-hour airtime hits 10 %: the modem
stops radiating while still accepting frames, which reads from above
as an intermittent resource stall. A test that itself exercises the
duty-cycle lock opts back in explicitly:
[radio]
frequency = 869463000
airtime_limit_long = 10 # percent; arms the 10 % ETSI cap
or per subinterface under [[nodes.x.rnode_interfaces]], or for a
one-off run with LORA_AIRTIME_LIMIT_LONG=10.
Concurrent runs
Only one scenario run can be in flight at a time — Docker names and USB handles would otherwise collide. A second invocation exits in under a second:
[leviculum] Another integration test is already running.
[leviculum] Current holder:
[leviculum] pid=12345
[leviculum] started=2026-04-14T02:01:33
[leviculum] pkg=periculum
[leviculum] binary=periculum
[leviculum] cwd=/path/to/leviculum
[leviculum] Wait for it to finish or stop that process, then retry.
A scheduled Tier 2 / Tier 3 that hits this case logs SKIPPED,
not RED, and sends a normal (not critical) notification. No
action needed — the next scheduled slot runs normally. In
practice this means: if you're doing late-night hardware work and
the 02:00 nightly fires, it silently defers. You don't need to
stop it.
Unit tests in the leviculum crates (leviculum-core,
leviculum-std, leviculum-ffi, leviculum-proxy,
leviculum-cli) run in parallel with a held scenario lock — they
never touch containers or boards.
Installing / updating the CI
just install-ci
Idempotent. Installs git hooks, systemd user units, state dirs,
separate cargo target dir, and the pinned tools that live in venvs the
runner owns — esptool, and the Python Reticulum periculum's host
type = "python" nodes run
(~/.local/state/leviculum-ci/rns-<pin>, built by
../periculum/scripts/install-python-runtime.sh; without it those
cells skip with reason=image_runtime_missing). Safe to re-run after
pulling.
The lnomad acceptance
lnomad is the terminal NomadNet browser. Its end-to-end acceptance drives a
real NomadNet node rather than a mock: NomadNet runs as the shared Reticulum
instance and as a node server hosting a known index.mu; lnomad --print
fetches and renders that page over the shared-instance path, and the script
asserts the rendered output contains the known content.
../periculum/periculum/assets/scripts/lnomad_nomadnet_acceptance.sh
It prints ACCEPT-PASS / LNOMAD-ACCEPT-COMPLETE and exits 0 on success, or
ACCEPT-FAIL: <reason> and non-zero otherwise. It creates an isolated RNS +
NomadNet config under a temp dir and always cleans up (kills nomadnet, removes
the temp dir) on exit.
This is an on-demand acceptance — it is NOT wired into any tier, because it
needs nomadnet (and its Python RNS) installed and takes ~30 s of real
announce/link setup. Requirements and overrides:
- nomadnet on PATH (or point
NOMADNETat the executable). - python3 + RNS on PATH (or point
PYat the interpreter). - The musl
lnomadrelease binary; the script builds it (cargo build --release -p lnomad) if it is missing. Override its location withLNOMAD, or the cargo target dir withCARGO_TARGET_DIR. - Tune timing with
NN_SETTLE(nomadnet startup, default 25 s) andLNOMAD_TIMEOUT(fetch timeout, default 40 s).
Fuzzing the wire parsers
The functions that parse UNTRUSTED bytes off the wire (packet, resource
advertisement, discovery announce app-data, announce field-slicer, IFAC,
HDLC/KISS deframers, I2P SAM reply lines) have a coverage-guided fuzz harness
(cargo-fuzz / libFuzzer). A parser that panics, overflows, hangs, or OOMs on
malformed input is a remote DoS, so each target asserts graceful Err/None.
There are two detached fuzz crates, one per library crate that owns a parser:
leviculum-core/fuzz—packet_unpack,resource_advertisement_unpack,discovery_announce,announce_from_packet,ifac_verify,hdlc_deframe,kiss_deframe.leviculum-std/fuzz—sam_parse(the I2P SAM reply-line parser and the base64/destination decoders it feeds).
announce_from_packet (the ReceivedAnnounce::from_packet field-slicer) and
sam_parse were added in Codeberg #108 alongside the #23 targets. Both reach
crate-internal parsers through a #[cfg(fuzzing)]-gated fuzz module in each
crate, so no fuzz-only surface leaks into the normal public API.
Running them
just fuzz runs every target in both crates; scripts/run-fuzz.sh is what it
calls. Until Codeberg #290 there was no recipe, no CI step and no schedule, so
the targets were run by nobody — and three defects found by hand in September
2026 sit exactly on top of three of them: unbounded msgpack recursion (#263)
and a wrapping bin32 length (#267) in resource_advertisement_unpack, an
uncapped HDLC accumulator (#271) in hdlc_deframe. A length field, a nesting
depth and an unbounded accumulator are what a fuzzer finds in minutes.
just fuzz # every target, 60 s each
just fuzz --seconds 900 # the budget a scheduled run wants
just fuzz hdlc_deframe # one target by name
just fuzz --list # what would run, without building
Exit codes separate the three outcomes, because a run that could not happen
must never look like a clean one: 0 every target ran its budget and found
nothing, 1 a crash (the input is kept, with its hash and a hexdump), 2
it could not run — missing nightly toolchain, missing cargo-fuzz, or a
fuzz_targets/*.rs that the crate manifest does not register as a [[bin]]
and that therefore no run reaches.
Still NOT part of any tier: it needs nightly + cargo-fuzz and the glibc host target (the workspace defaults to musl, which ASan does not want), and even a short run costs minutes — 345 s for all eight targets at 30 s each (measured 2026-09-19, warm registry; 95 s of that is the leviculum-std ASan build).
The corpus persists outside the checkout
The working corpus and any crash input live under
~/.local/state/leviculum-fuzz (LEVICULUM_FUZZ_STATE), not in the fuzz
crates:
~/.local/state/leviculum-fuzz/corpus/<crate>/<target>/ inputs libFuzzer kept
~/.local/state/leviculum-fuzz/findings/<crate>/<target>/ crash inputs
~/.local/state/leviculum-fuzz/logs/<timestamp>/ full per-target output
This is the difference between fuzzing and re-fuzzing. A corpus inside the tree would be thrown away by construction: the nightly runs on a fresh clone it deletes when green, so every scheduled run would restart from the checked-in seeds and re-explore the same shallow paths, and its value would plateau on night two. The first run above left 1837 inputs across the eight targets; the next run starts from them.
The runner is a one-line hook for a scheduled job — bash scripts/run-fuzz.sh --seconds <budget>, exit 1 on a crash — and its FUZZ_SUMMARY line is
deliberately not spelled SUMMARY, so a nightly that tallies scenario results
out of ^SUMMARY lines cannot silently fold fuzz counts into them.
Keeping the runner honest
just fuzz-selftest (on the Tier 0 push path, ~15 s,
scripts/test-run-fuzz.sh) injects the failures into a throwaway fuzz crate
and asserts what the runner concluded: a target that crashes, a target that
does not, a corpus that has to survive between runs, an unregistered target
file, and a host without cargo-fuzz. The last one is the case that must never
look green. It skips with a named reason where nightly or cargo-fuzz is absent,
so the push path does not inherit the toolchain requirement.
Any crash the fuzzer finds is fixed at the root AND pinned by a deterministic
regression unit test in the normal suite, so it stays fixed without the fuzzer.
See leviculum-core/fuzz/README.md and leviculum-std/fuzz/README.md for the
target lists and exposure ranking.
Golden rules
- Tests are never flaky. A failure is a real bug — diagnose and fix at the root, don't retry until green.
- Don't commit while tests are red.
#[ignore]is only for hardware-dependent tests. For CPU-expensive non-hardware tests, use a Cargo feature flag.
The ignored-test census
An #[ignore]d test is run by nothing. Codeberg #189 found one such
suite that had been broken for a month with no gate anywhere to say so,
so the size of that bucket is pinned per test unit in
scripts/ignored-counts.txt and checked by
scripts/check-ignored-counts.py at the end of just standard. The
census is exhaustive: every test executable in the workspace plus every
package's doc-tests, with units absent from the pin file expected to
have zero, so a new test binary is covered without an entry.
A new #[ignore] therefore fails Tier 1 until you either route the test
into a tier or raise its number in the pin file — a one-line diff, on
purpose, with the reason belonging in the commit message. python3 scripts/check-ignored-counts.py --print dumps the current census in
pin-file format.
Routing an ignored test by name (rather than lifting the ignore) is what
scripts/run-status-parity.sh does for the three status_parity tests,
which need to run serially. A test filter that matches nothing exits 0,
so any such script must also assert how many tests actually ran.
-
A
post-commithook detachedscripts/run-tier1.shafter every commit until 2026-08-07. A commit is not a unit anybody wants tested — a WIP commit, an amend and a commit mid-refactor each started the same forty-minute docker run — and there is nothing an author can do about a red gate that lands twenty minutes after the commit it judges. Removed; see CI Pipeline and the hook rule in Checks That Are Actually Checks. ↩ ↩2 -
just standardtyped by hand builds in the repo's owntarget/. The separateCARGO_TARGET_DIRat~/.cache/leviculum-ci-target— so IDE builds and CI builds don't fight over the same incremental cache — belongs toscripts/run-tier1.sh, which nothing has started since the post-commit hook went. Either way the first run against a cold target dir compiles the workspace from scratch (~40 min); subsequent runs are incremental (~15 min). ↩ -
Tier 2 had a 12:30/18:30 timer until 2026-06-12, when it was retired in favour of on-demand runs (
scripts/install-ci.shstep 9, which also deletes any timer a previous install left behind). This page went on advertising the timer, and.githooks/pre-pushwent on telling people to "wait for the next scheduled run" until 2026-08-07 — see CI Pipeline.scripts/ci-status.shprints how long it has been since a Tier 2 run was recorded. ↩
CI Pipeline
Two things run tests here, and they are not the same thing. Four
local tiers with different time budgets and triggers automate the
test discipline mandated by CLAUDE.md, on the developer's machine —
no GitHub Actions. On the forge, Woodpecker runs a smaller set on
every push, because the local tiers are hooks a fresh clone does not
have and --no-verify switches off. The tiers come first; the forge
pipelines are below them.
Tiers
| Tier | Name | Trigger | Budget | Test scope |
|---|---|---|---|---|
| 0 | fast | pre-push hook | ~3 min | fmt + clippy (host + nrf firmware workspace, both BSPs) + the firmware stack-frame gate + rustdoc gate + the third-party notice guard + the pre-push guard selftest + workspace lib tests |
| 1 | standard | on demand: just standard, once per batch | ~15 min (first run: 20-40 min cold compile) | Tier 0 + every integration-test target in the tree (the ones needing a flag or their own feature set by name, the rest computed by scripts/standard-integ.sh) + TCP-hub endurance smoke soak (see Soak and endurance) + the status_parity two-daemon suite + the ignored-test census |
| 2 | extensive | on demand: systemctl --user start leviculum-ci-tier2.service | ~30-90 min | Tier 1 + the periculum conformance/ and regression/ corpora (docker) |
| 3 | nightly | systemd timer 02:00 daily | ~2-6h | Tier 2 + LNode flash-from-HEAD + the periculum hardware/ corpus |
Each tier runs everything from the lower tiers as well, so a green nightly proves the entire stack.
just guards is not in this table because it is not a tier: it is the
coder-side subset of Tier 0's sub-ten-second guards and censuses (~20 s warm),
run beside fmt, clippy and the workspace tests after a batch so that a guard
red is found by its author rather than by the landing gate. fast and
standard remain what the push path and the landing gate run. See
Testing.
What the forge runs
Three Woodpecker workflows on ci.codeberg.org, all in
.woodpecker/:
| File | Fires on | Runs |
|---|---|---|
ci.yml | every push, every pull request, manual | just ci-gate — fmt, clippy over all targets, and every workspace test except the submodule-bound suites and tests named in scripts/ci-gate-integ.sh (Codeberg #312) |
commit-trailers.yml | every push, every pull request, manual | scripts/check-commit-trailers.sh over the pushed range |
nightly.yml | cron, plus pushes touching the packaging paths | the same gate, then the .deb + tarball build; the cron run also publishes |
ci.yml exists because until Codeberg #299 none of the others
covered an ordinary source commit: the nightly's push trigger is
filtered to the packaging paths, the trailer check reads messages
rather than code, and what stood between a Rust-only commit and the
public releases page was .githooks/pre-push — per-clone local
config, skipped by --no-verify, running the developer's toolchain
and not the pipeline's. So it carries no path: filter, and
scripts/check-ci-pipeline.sh (part of just fast) fails the push
path if any future edit gives it one.
just ci-gate is a Justfile recipe rather than commands spelled out
in YAML, and the container provisioning both pipelines need is
scripts/ci-gate.sh rather than two copies of the same apt lines:
one gate, one environment, no drift between the two files that run
it. The gate is deliberately not an alias for just fast — the
recipe's comment lists what a submodule-less host-target container
cannot prove (the firmware workspace, the cross-compiles, the
submodule pins), and those stay on the local push path, which has
the targets. Measured cost, cold, of the version that ran --lib
only: 3m23s including provisioning. With Codeberg #312's widening the
runner's own step timing is 965-1185 s (cron pipelines 476-486, push
488, provisioning included), before the release build of the
integration binaries that four mvr tests need was added to it.
What the gate runs, and what it leaves out. The forge container is a
stock rust:bookworm with musl-tools, socat and iproute2 added
by scripts/ci-gate.sh, and a clone without submodules (#300). In it
just ci-gate runs fmt, clippy over every target, the workspace lib
tests, just build-integ-bins (the release binaries the mvr tests
spawn), and then scripts/ci-gate-integ.sh: every integration-test
target the tree has, every bin unittest target and the doctests. Out
of that it leaves exactly what needs a reference/ submodule at run
time, because those fail rather than skip without it: three whole
targets (rnsd_interop, leviculum-lxmf's reference_lock,
lnmsg's python_interop) and seven tests by exact name inside mvr,
which includes the rnsd_interop harness by #[path] and so runs the
harness's own tests and the four tests that spawn a Python peer
through it. The script's header holds the
list, one written reason per entry, and checks every entry against the
tree so a stale one is a red gate. The citation guard runs in full but
skips the citations into the absent references and prints how many
(about 2600). Everything left out runs on the tier-2 nightly with the
submodules present, which is what the publish gate reads.
What may be published
The forge gate runs fmt, clippy and every test in the workspace
except three suites and seven submodule-bound tests inside mvr, and
rnsd_interop — the suite that measures
whether we still interoperate with a Python-RNS peer, which is half of
Priority 1 — is the largest of the three. It runs in neither forge
pipeline and cannot: it needs the reference/Reticulum submodule and a
python3, and both pipelines clone with submodules: false on purpose,
which is the property just check-plain-clone exists to hold (Codeberg
#300). Fetching the submodule into the release path would undo exactly
that. The other two are leviculum-lxmf's reference_lock and
lnmsg's python_interop, for the same reason;
scripts/ci-gate-integ.sh holds the list, one written reason per
entry, and computes everything else it runs from the tree.
So the interop verdict is imported rather than re-derived
(Codeberg #312). The tier-2 nightly already runs the whole workspace,
with submodules, over a fresh clone pinned to origin/master. When
that run is green it pushes a lightweight ref at the commit it tested:
refs/nightly/green/<YYYYMMDDTHHMMSSZ> -> <tested commit>
and scripts/publish-nightly.sh refuses, before it touches the forge,
any commit those refs do not cover. Three conditions, and a refusal
always names which one failed:
| Condition | Meaning |
|---|---|
NO-SIGNAL | there is no refs/nightly/green/* on the remote, or it could not be read |
NOT-COVERED | no green ref names this commit or a descendant of it — no nightly has seen this code |
TOO-OLD | the newest covering ref is older than the staleness bound (72 h) |
The 72 h bound is measured, not assumed: over 2026-08-22..09-22 the nightly timer produced 27 runs with a median gap of 24 h, every gap but one at or under 54.4 h, and one 96 h gap (2026-09-04 to 09-08) which is precisely the case the bound exists to stop. When the forge publishes on the fallback cron rather than on the trigger described below, the freshest ref it can read is normally the previous night's and is already ~24 h old. 72 h therefore accepts the ordinary day plus one missed night and refuses two.
What fires the publish
The publish is fired by the green ref, not by the clock. Since
2026-09-24 the reviewer host polls the forge every ten minutes and, as
soon as a new refs/nightly/green/* appears, fires the nightly cron
through the Woodpecker API (lev-nightly-publish-trigger, reviewer-host
tooling: it is not in this repo, and the API token stays on that host on
purpose). The scheduled cron, 0 4 * * * UTC, remains as a fallback
rather than as the normal path; both firing on the same day is harmless,
because the release is rolling and scripts/publish-nightly.sh writes
the same nightly tag with its assets overwritten. Ordering is the
whole point of the change: pipelines 455 (2026-09-23) and 458
(2026-09-24) both refused at the publish gate because the 04:00 cron ran
before the nightly host had pushed that day's green ref. The code was
fine and the night had been green; the signal simply was not there yet.
That refusal is also what an operator sees when the trigger did not
fire. The publish step refuses with NO-SIGNAL when no green ref can be
read from the remote at all, and with NOT-COVERED — printing "the
newest green ref is ..." and "is neither that commit nor an ancestor of
it" — when refs exist but none of them names this commit or a
descendant of it, which is what a night that has not pushed yet looks
like from the publish step. Neither is a build failure and neither
needs a code change:
bash scripts/check-nightly-green.sh --commit <sha> answers which ref
the forge can see, and the fix is to wait for the night's ref to land
and then start the nightly cron by hand from the Woodpecker UI
(Repo → Settings → Crons) if the trigger has not done it first.
Publishing anyway
A human can decide otherwise. Set, on the publish step:
LEVICULUM_PUBLISH_WITHOUT_NIGHTLY="why you are doing this"
It is deliberately not a boolean: the value is the reason, it must be at least 8 characters, and it is printed into the run's log where it stays with the build it excused. A value too short to be a reason is refused.
Two ways to set it, and neither needs a code change:
- Woodpecker — start the
nightly.ymlworkflow manually and add the variable in the run dialog. - By hand, from a checkout —
CI_REPO=Lew_Palm/leviculum CI_COMMIT_SHA=$(git rev-parse HEAD) CODEBERG_TOKEN=... LEVICULUM_PUBLISH_WITHOUT_NIGHTLY="..." bash scripts/publish-nightly.sh, withdist/already staged byscripts/collect-nightly-debs.sh.
bash scripts/check-nightly-green.sh --commit <sha> answers the
question on its own, without publishing anything, which is the first
thing to run when a nightly publish has gone red.
What holds the chain together
| Gate | Asserts |
|---|---|
just nightly-green-selftest | both scripts BEHAVE: each of the three refusals fires, the override works and needs a reason, and a red or absent rnsd_interop signs nothing |
just check-publish-nightly-gate | the mechanism is still CONNECTED: five links from the publish step to the manifest the signer reads, each broken on purpose in its own self-test |
Both are in just fast, so they run on the push path. They are
separate because the failure mode is available to both halves: a gate
wired into nothing, and a gate wired in that says yes to everything.
The signing half is scripts/nightly-green-ref.sh. It does not take
the night's verdict on trust for the one property this is about — it
reads the run's own manifest (scripts/run-with-manifest.py,
Guarantee B) and refuses to sign unless the rnsd_interop unit
executed and every test in it passed. "The nightly was green" must not
be able to mean "the suite never ran".
Installation
One command, idempotent:
just install-ci
It installs git hooks (via core.hooksPath = .githooks), runner
scripts, systemd user units, the separate cargo target dir, the
build-directory sweeper just sweep needs, the state dir, and the two
pinned tools that live in venvs the runner owns — esptool for RNode
flashing, and the Python Reticulum periculum's host type = "python"
nodes run (~/.local/state/leviculum-ci/rns-<pin>, see "Host python
cells need the image's Reticulum" below). Re-running is safe.
The installer detects the worktree it was run from and patches the
systemd-unit ExecStart paths to match — so a git worktree-based
second checkout (see "VM-mode install" below) installs its own units
that fire against itself.
VM-mode install (CI worktree on a long-running host)
For schneckenschreck or any other dedicated CI machine where the
nightly Tier-3 runs land, install with --vm-mode:
git worktree add ~/coding/libreticulum-ci master
cd ~/coding/libreticulum-ci
bash scripts/install-ci.sh --vm-mode
--vm-mode differs from the default install in two ways:
- The git-hook wiring (
core.hooksPath = .githooks) is skipped. The VM never commits or pushes; hooks would never fire. - A worktree-scoped marker file
(
.git/worktrees/<name>/leviculum-ci-vm-mode-marker) is created.run-tier2.shandrun-tier3-hw.shcheck this marker at the head of every run and, if present, invoke_repo-sync.shto dogit fetch + git checkout --force origin/master + git submodule update --recursive.
The marker is per-worktree, not per-user: a manual invocation of
run-tier2.sh from the developer's primary checkout will not
trigger the destructive --force checkout against the wrong tree.
The synced commit hash is appended to last-results.txt as
<timestamp> tier2 sync HEAD=<short-hash> (or tier3-hw for the
nightly), so you can correlate scheduled runs with the master commit
they tested.
The firmware stack-frame gate
just nrf-stack-frames builds both firmware binaries and reads the
frame-allocating sub sp immediates out of the linked ELF. Any frame
above 16 KB fails the gate.
The T114 stack is 128 008 B and grows down into the SoftDevice's RAM
floor. An overflow past _stack_end does not fault: it overwrites SD
state, and the board dies later in an SD internal assertion with a
useless PC. So a single oversized frame is both fatal and invisible,
which is why this is checked statically on every push rather than
observed at runtime.
The frame it was written for: Box::new(builder.build(..)) materialised
a by-value NodeCore — over 40 KB once EmbeddedStorage's inline
collections are counted — twice in main's poll frame. 94 720 B, 74 % of
the stack, ~13 KB of margin left for the whole call tree.
NodeCoreBuilder::build_boxed allocates first and configures through the
box, which drops that frame to 12 672 B.
The gate prefers arm-none-eabi-objdump and falls back to the rustup
llvm-tools llvm-objdump; install-ci.sh installs the latter.
Manual operation
just fast # Tier 0
just standard # Tier 1
just extensive # Tier 2
just nightly # Tier 3
just status # show recent runs across all tiers
First-run expectation
Tier 1 runs in a separate CARGO_TARGET_DIR (~/.cache/leviculum-ci- target/) so it doesn't fight your IDE's target/ for inkremental
caches. The first run after install-ci.sh compiles the whole
workspace and all test binaries from scratch — plan for 20-40
minutes. Subsequent runs are incremental, ~5-15 minutes.
Keeping the build directories bounded
Cargo adds; it never removes. Every changed input writes a new
hash-suffixed artefact next to the old one, so a target directory only
grows, and the growth rate is the point rather than any one build:
measured on the CI host on 2026-09-24, the deps directory under
target/x86_64-unknown-linux-musl/debug alone held 2705 files and
27 GB of that tree's 36 GB, and on 2026-09-09 one working day of gate
runs took the same tree to 137 GB and filled the root volume
(Codeberg #381). A full volume does not announce itself as a full
volume: hardware runs go red for want of space and look like the
stack.
just sweep # both workspaces, 30 GB and 4 GB caps
just sweep 20GB 2GB # tighter caps
Two directories, because this repository has two workspaces — the host
one at the root and the firmware one in leviculum-nrf — and sweeping
the root leaves the firmware's 6 GB untouched. Where they lie is asked
rather than assumed, so a tree that moved its artefacts with
CARGO_TARGET_DIR (Tier 1 and the nightly do) is swept where they
actually are.
cargo sweep --maxsize drops the oldest artefacts until the directory
fits the cap, which keeps exactly the ones the next build would reuse.
cargo clean is the blunt version of the same thing and costs a full
rebuild of everything.
What no cleanup may take is the compilation cache: with
RUSTC_WRAPPER=sccache set, that cache is what makes the rebuild after
a sweep cheap, and it bounds itself through SCCACHE_CACHE_SIZE.
Deleting it to free space buys one-off gigabytes and charges the next
build for them.
Notifications
Read this as history, not as behaviour. scripts/run-tier3.sh calls
notify-send on its verdict — -u critical for RED (sticky until
dismissed), -u normal for GREEN and for a lock-held skip, -u critical
again when the lock's holder is a suspected wedge. It is the only
tier runner that ever did. It is also not the script the nightly starts:
leviculum-ci-nightly.service runs scripts/run-tier3-hw.sh, which
writes the ledger and notifies nobody. So no tier notifies today. Results
are pull-only — just status, or
~/.local/state/leviculum-ci/last-results.txt.
Prerequisite: notify-send needs DBUS_SESSION_BUS_ADDRESS and
XDG_RUNTIME_DIR in the user systemd manager environment, which
exists only when you have a logged-in graphical session. On a
headless server, notifications are silently dropped — inspect
~/.local/state/leviculum-ci/last-results.txt instead.
Stale-block on push (removed 2026-08-07)
pre-push used to block the push when the last tier2 GREEN line in
last-results.txt was ≥ 10 commits or ≥ 24 hours old. It was removed,
not repaired. Only scripts/run-tier2.sh writes that line, nothing has
started it since the Tier 2 timer was retired on 2026-06-12, and the
remedy the block printed (just extensive) does not write it either —
so the block could not be cleared by doing what it said. It was
unsatisfiable for 46 days, and the 502 commits that landed in that
window all used git push --no-verify, which switches off the lint,
Tier 0, mvr and the trailer guard along with it.
scripts/ci-status.sh still reports how long it has been since a Tier 2
run was recorded. It states the age and blocks nothing.
Tier 1 after every commit (removed 2026-08-07)
.githooks/post-commit detached scripts/run-tier1.sh — just standard
under docker, 15 minutes warm and 20-40 cold — after every commit that was
not part of a rebase. It was removed, and the rule it failed is in
Checks That Are Actually Checks.
The short form: a commit is not a unit anybody wants tested. WIP commits,
amends and commits mid-refactor all started a forty-minute run, which is
why the runner needed a dirty-flag loop to coalesce them — it was
repairing a granularity that was wrong to begin with. It was also
invisible: batches were separately instructed to start just standard
under nohup, so the same tier ran twice per batch for a week before
anyone noticed the hook existed. And it ran docker in the background,
which tears down containers whatever else is on the box was using — the
standing rule against starting the full integ suite behind someone's back
exists for that collision, and this hook was doing it after every commit.
Tier 1 is now started explicitly, once per batch, by typing
just standard.
Logs
Location: ~/.local/state/leviculum-ci/
| File | Contents |
|---|---|
last-results.txt | one-line tally per run (<iso-timestamp> <tier> GREEN/RED <log-path>, or <tier> SKIPPED lock-held|lock-suspect <verdict fields> <log-path>) |
tier1-YYYYMMDD-HHMMSS-PID.log | full Tier 1 output (one file per run) |
tier2-YYYYMMDD-HHMMSS-PID.log | full Tier 2 output |
nightly-YYYYMMDD-HHMMSS-PID.log | full Tier 3 output |
tier1.lock | flock for Tier 1 concurrency control |
tier1.dirty | marker that Tier 1 needs to (re-)run |
Rotation: tier 1/2 logs are deleted after 14 days; nightly logs after 60 days. Done at the start of each runner script.
Each script run gets its own log file (timestamp + PID suffix). No
run ever overwrites another run's log — this is intentional so a
failure trace cannot vanish under a successful re-run. The path of
the specific log goes into last-results.txt so just status can
point at exactly the right file.
The scenario suites live in periculum
The multi-node scenarios that used to be reticulum-integ are now the
sibling periculum checkout,
which leviculum expects at ../periculum (override with
PERICULUM_ROOT, or the binary with PERICULUM_BIN). They are TOML
files, not #[test] functions, so the tier separation is a matter of
which directory a tier runs rather than of #[ignore]:
| Corpus | Binds hardware | Run by |
|---|---|---|
conformance/ | no | Tier 2 |
regression/ | no | Tier 2 |
hardware/ | yes | Tier 3 |
The split is machine-checked in periculum
(periculum/tests/corpus_admission.rs), so a scenario cannot drift into
the wrong tier by convention alone. A hardware/ scenario whose boards
this bench does not hold reports SKIPPED_INFRA naming what was
missing — never RED.
Run one scenario by hand:
periculum run ../periculum/hardware/lora_link_rust.toml
Host python cells need the image's Reticulum
A periculum node with type = "python" runs one of two Python
Reticulums. In a CONTAINER it runs the image's: assets/Dockerfile
installs the pinned rns==<pin> wheel last, where no resolver step can
move it. Run as a HOST process — every emulated/ cell, and a BLE node
— there is no image, so the runner resolves it onto a venv carrying the
same two packages the image carries, in the same order:
~/.local/state/leviculum-ci/rns-<pin> # $PERICULUM_PYTHON_RUNTIME_DIR overrides the parent
just install-ci provisions it, the way it provisions the other pinned
tool in a venv: one guarded call out to
../periculum/scripts/install-python-runtime.sh, which reads the pin
out of periculum's assets/Dockerfile rather than writing it down a
second time, installs the vendored LXMF first and the pinned wheel
last, and asserts the version before leaving the venv on disk. It is
idempotent — an already-provisioned venv costs one interpreter start —
and both the guard and a failure are warn-only: a host with no sibling
periculum checkout has no corpus to run and gets a note, not a failed
install.
Until it exists, a host python cell skips — SKIPPED_INFRA with
reason=image_runtime_missing, naming the script that builds it.
Nothing falls back to the vendored citation trees
(reference/Reticulum, 1.3.5), because a figure produced against a
stack nobody deploys is worse than a skip; a cell that wants the
sources our comments quote asks for them per node with
python_runtime = "reference". The pin is in the directory NAME for
the same reason: a pin bump is then a venv that does not exist yet, so
it reads as that named skip rather than as a run quietly served by the
version before the bump.
The BLE room needs a patched btvirt
The ble_room_* cells put N virtual LE controllers on one emulated air
so that N lnsd daemons can prove BLE mesh formation with no boards.
The emulator is bluez's btvirt, which Debian does not package, and
which cannot be a stock build: bluez 5.82 and older report the
central's connection handle to the peripheral and forward ACL data
under the sender's handle, so every room stalls from the second
concurrent link (N >= 3) and the emulator's bug is measured as a mesh
finding. Upstream commit 4ff7deaf8c fixes it; with it the ladder is
green to N = 16, which is the emulator's own MAX_BTDEV_ENTRIES
ceiling. just install-ci therefore builds btvirt itself
(scripts/install-btvirt.sh): it fetches the bluez source matching this
host, applies the vendored patch from scripts/patches/, and installs
the binary together with a provenance sidecar,
/usr/local/bin/btvirt.provenance, whose first non-comment line every
room prints as origin=:
BLE_ROOM_BTVIRT bin=/usr/local/bin/btvirt bytes=895400 origin="bluez 5.82-1.1 + upstream 4ff7deaf8c (emulator: use the handle of the receiving side); built 2026-09-27, md5 11f1de0c2b097787824cd6335ef676bc"
A room with no sidecar still runs and says origin=unrecorded, which
is a result nobody can attribute afterwards. Verify a bench without
installing anything with bash scripts/install-ci.sh --check. The build
is skipped on a host that already holds the patched binary, and it warns
rather than fails when a prerequisite is missing — only the bench that
runs the room needs it. The four prerequisites beyond the binary (the
hci_vhci module, a running system bluetoothd, no BLE MIDI GATT
service, cap_net_raw on btmon) are listed in
scripts/install-ci.sh at the btvirt step; nothing can provision them
for you.
The patch is applied by content, not by version number. Before touching
the source tree the script asks patch --dry-run which of three states
it is in: the fix reverse-applies cleanly (already there — leave it
alone), it applies forward (stock — apply it), or neither (refuse,
rather than build an emulator nobody can name). No released bluez known
here carries the fix yet, 5.82 being the newest and the commit dated
after it, so the first release that ships it will simply be recognised
and need no change to this script.
Both halves run with no root, no network and no compiler under just btvirt-selftest (in guards and fast, 0.12 s): the patch step
against a bluez source tree synthesised from the vendored patch's own
pre-image, and the provenance checker against each way the sidecar can
lie. What the synthesised tree cannot answer is whether the patch still
fits Debian's bluez — a patch and a fixture derived from it drift
together — and that is what patch --dry-run answers at provisioning
time.
Sixteen controllers is the ceiling. MAX_BTDEV_ENTRIES is 16 in
bluez's emulator/btdev.c, in the 5.82 this host builds from and in
upstream master (read 2026-09-27 at bluez HEAD 8b4a4176): btvirt -L -l16 runs, -l17 exits at once with "Failed to open Virtual HCI
device". Sixteen is therefore the largest room this bench can host, and
on the patched binary it is green. That is why periculum's
regression/ble_room_20.toml carries an [unsupported] section rather
than a red — a twenty-node room cannot be built here at all, and the
cell skips as infra after 21 s with the kernel showing no new
controllers. That skip is not a bug in lnsd.
Concurrent test protection
Two scenario runs on the same machine fight over Docker container names
and USB serial handles. To prevent that, periculum acquires a
process-wide file lock on ~/.local/state/leviculum-ci/test.lock before
bringing any node up.
Single invocation: transparent. No extra output.
Two simultaneous invocations: the second exits within a second with
a multi-line [leviculum] message naming the current holder —
pid, started time, cwd, optionally the test-name filter. Example:
[leviculum] Another integration test is already running.
[leviculum] Current holder:
[leviculum] pid=12345
[leviculum] started=2026-04-14T02:01:33
[leviculum] pkg=periculum
[leviculum] binary=periculum
[leviculum] cwd=/path/to/leviculum
[leviculum] Wait for it to finish or stop that process, then retry.
On-demand Tier 2 / scheduled Tier 3 runs that collide with a manual test
drop a marker file at ~/.local/state/leviculum-ci/lock-contention;
the runner scripts read the marker, classify the run as SKIPPED
(not RED), and delete it. No false-alarm pages.
Which kind of contention (Codeberg #309)
The marker is not a flag: it carries periculum's verdict on the process holding the lock, plus that process's identity.
| Field | Meaning |
|---|---|
verdict= | running, suspected_wedge, misrecorded, unattributable |
suspect= | true for every verdict but running |
holder_pid=, holder_age_secs= | who is holding it, and for how long |
detail= | periculum's sentence about the holder, with what to inspect |
A holder alive past 24 hours has by definition starved at least one nightly, and one whose recorded identity the kernel disagrees with is a bug shape nothing else can see. Both reach the ledger under their own token, so the case worth acting on is greppable:
<iso> tier3 SKIPPED lock-held verdict=running holder_pid=… holder_age_secs=… <log>
<iso> tier3 SKIPPED lock-suspect verdict=suspected_wedge holder_pid=… holder_age_secs=… <log>
Neither is RED. The verdict is a heuristic over metadata — a genuinely
enormous run looks like a wedge — and the contender never touches the
lock, so a false accusation would cost somebody killing a healthy
nightly. What changes is what the ledger says and, in
scripts/run-tier3.sh, whether the notification is normal or
critical.
The marker, not the exit code, is what the runners branch on.
periculum also carries the distinction in its exit status (2 for an
overlap, 4 for a suspect holder), but run-tier2.sh and run-tier3.sh
reach periculum through just extensive / just nightly, and the
nightly recipe rewrites its status to 1 whenever an LNode's firmware
could not be verified. The exit code is therefore a corroborating signal
there, and it is also the only channel that cannot say WHO.
scripts/run-tier3-hw.sh calls periculum directly and uses the codes for
one decision only: 2 and 4 both mean "look at the marker".
scripts/test-lock-contention.sh (just lock-contention-selftest) and
the contention cases in scripts/tier3-hw-selftest.sh hold this against
stubbed markers; neither needs a build, docker or the rig.
Inspecting the lock
cat ~/.local/state/leviculum-ci/test.lock # current (or last) holder
ls ~/.local/state/leviculum-ci/lock-contention # marker if present
Force-release
Not applicable. The kernel releases the flock the moment the holding
process closes its fd — on clean exit, panic, SIGINT, SIGKILL, and
even host reboot. There is no TTL, no heartbeat, no manual cleanup
path. A stale test.lock file on disk after a reboot is self-
healing: the next invocation opens it, flock succeeds immediately
(kernel state is empty post-reboot), and the stale content is
overwritten.
Scope
The lock protects only scenario runs. Unit tests in leviculum-core,
leviculum-std, leviculum-ffi, leviculum-proxy, and
leviculum-cli do not acquire it — they parallelise freely with an
in-progress scenario run. periculum validate and periculum list
do not acquire it either: they read scenario files and touch no node,
container or radio.
Filesystem requirement
Local filesystem only. flock semantics over NFS / sshfs are
implementation-defined. If your $HOME is on a network filesystem,
the lock behaviour is not guaranteed. This is a single-developer
dev-box tool; not an issue in practice.
Hardware test profiles (Tier 3)
Tier 3 runs the periculum hardware/ corpus over USB-attached
embedded devices. Different scenarios need different subsets of the
attached boards; the rest must not transmit, so their RF activity does
not contaminate the run.
No USB-hub power switching. Every board stays permanently powered
and passed through to the VM. RF isolation of non-participating
firmware nodes is done in software: the runner pushes radio_silent
over serial to every discovered board it did not bind. Per-port power
cycling correlated with hamster hardware-watchdog freezes (proven
2026-06-15) and was removed, together with the usbhub-helper and its
libvirt-passthrough caveats.
Which individual boards exist on this bench is site data and lives in
periculum's rig.toml (override with $PERICULUM_RIG). What kind of
board each is — how it is recognised over USB, which port carries which
role, what it can be asked to do — lives in periculum/devices/*.toml
and is the same everywhere. A scenario names the set of boards it needs:
profile = "rnode_lnode_pair"
which is resolved against the rig file. A scenario needing more boards
than the bench holds is SKIPPED_INFRA with a reason naming what was
missing — never RED. An absent board is not a protocol result.
Firmware identity
Before any hardware scenario runs, scripts/flash-lnodes-from-head.sh
flashes every attached LNode from the current commit and reads its
[FW_BUILD] banner back over the debug serial to confirm the board
really runs that commit. A board whose firmware cannot be confirmed
makes the tier RED and is named in the verdict
(firmware_unverified=<vid:pid>): a run against unknown firmware must
never be silently trusted. This step is leviculum's, not periculum's —
periculum tests whatever firmware it finds and leaves board preparation
out of scope on purpose.
Device-vanish watchdog
scripts/run-tier3-hw.sh polls lsusb once a second for the whole run,
cross-checks every sub-baseline reading against sysfs, and records one
journal line per event — every vanish and every return, not one latched
line per board. Under VFIO controller passthrough the host cannot inject
a phantom VM-side disconnect, so a board that leaves the bus really left
it; what that means, though, is decided afterwards. A disconnect
periculum's own BOARD_RESET lines say it commanded (it reboots every
bound board per scenario) is accounted and never RED. An unaccounted
one forces RED with the board named (board_vanish=<vid:pid> cause=<what the witness supports>), and every scenario verdict from the
vanish onwards is untrusted. The cause token is read off the board's
debug witness or reads cause=unknown; it is never asserted.
The journal also records what the board's own witness cannot see, because a board that loses power writes nothing:
- where each board sat — its USB bus path and the hub it hangs off,
snapshotted at baseline while the whole rig is still present, and
quoted back on the vanish line (
last_paths=,last_hubs=). The RED banner turns that into a per-hub count, and says so when more than one board was lost on a single hub: that is the shape a hub or power event has, and independent firmware failures do not have it. - what the kernel said — the
usb/hublines about those paths, taken at the moment of the vanish (kernel ... msg=).dmesgis a ring buffer that rolls over long before anyone reads a nightly, andUSB disconnectversusdisabled by hubor an over-current report is the whole difference between a board fault and a hub fault. An unreadable buffer is recorded asunavailable reason=..., never as silence.
Both were added after the 2026-08-12 run (Codeberg #251) lost two LNodes four minutes apart while a third board on another hub ran on, and left no artefact able to say whether one hub had dropped out or two firmwares had failed.
Troubleshooting
| Symptom | Action |
|---|---|
| Tier 1 never seems to run | It doesn't run itself. Type just standard. Nothing has started it automatically since the post-commit hook went (2026-08-07). |
| Notification never arrived | Expected: nothing that currently runs calls notify-send (see Notifications above). Check last-results.txt. |
| Tier 1 spuriously red | Check log; if Docker is involved, ensure no leftover containers (docker ps -a) |
| Timer didn't fire | systemctl --user list-timers, then journalctl --user -u leviculum-ci-nightly.timer. The nightly is the only timer this installer enables; Tier 2 has no timer. |
| Tier 2 looks like it never runs | It doesn't, unless started: systemctl --user start leviculum-ci-tier2.service. scripts/ci-status.sh prints how long it has been. |
| Disk filling up | Logs auto-rotate (14d/60d); the build directories do not — see Keeping the build directories bounded. just sweep caps both workspaces, just sweep 20GB 2GB harder. Tier 1's own directory is separate: cargo clean --target-dir ~/.cache/leviculum-ci-target. Never the sccache. |
Soak and endurance
Two independent lines of evidence back the claim that lnsd runs as a stable,
long-lived transport node: an in-repo soak test that runs as a CI gate, and a
permanent node in the public Reticulum mesh. A third measure, poison-tolerant
locking, keeps a task panic under a lock from crashing the whole daemon.
In-repo soak: the TCP-hub endurance test
leviculum-std/tests/rnsd_interop/loadtest_tcp_hub_tests.rs (Codeberg #101)
boots the real lnsd binary as an internet-facing transport hub, drives
sustained load plus connection churn against it, and samples the hub process's
/proc/<pid> RSS and open-fd count throughout. Because the hub is a separate
process, those samples are meaningful.
Topology
N raw TCP clients ─┐ ┌─ sink (Single dest, TCP client)
churn connections ─┼─▶ lnsd ──▶┘
┘ (transport hub)
A pool of steady TCP clients plus a set of churn workers (connections opened,
used, and closed in a tight loop) push sequence-numbered, encrypted single
packets through the hub to a sink daemon. The sink decrypts and folds each
(client, seq) into a per-source set, so delivery is verified exactly, not
sampled.
What it asserts
On every run report_and_assert enforces:
- 100% delivery. TCP is lossless, so every packet a client sends must arrive at the sink, contiguous and without duplicates. Any shortfall is a real hub bug, never noise. Connection-refused-under-load counts as zero-delivery, so backpressure failures cannot hide.
- RSS plateau (no per-connection leak). A steady population of connections legitimately costs memory, so growth from idle baseline to steady is expected. The leak signal is a continuous climb across the steady+churn phase, where thousands of connections are churned: the test compares the first vs second half of the steady-phase RSS samples and fails only when both a proportional and an absolute floor are exceeded, so a plateau with jitter never trips. A separate absolute ceiling over baseline is a runaway backstop.
- fd bounded under churn, released after teardown. Peak fd count must stay
under
baseline + steady_conns + churn_workers + margin, and after the clients close and drain, the count must fall back near baseline. A per-connection fd leak would blow past the ceiling and leave the end count elevated. - Clean hub log. The hub's log is scanned for fatal/bad lines; expected churn-teardown lines are allow-listed, anything else fails the run.
Two variants
| Test | Default load | Runtime | When |
|---|---|---|---|
loadtest_tcp_hub_smoke | 24 conns / 5 s | ~15 s | Tier 1 CI gate, every commit |
loadtest_tcp_hub_soak | 200 conns / 60 s | minutes | on demand / heavier validation |
Both are #[ignore]d because they spawn the lnsd binary, which the
leviculum-std test build does not itself produce — the binary is built first.
Running it
Use the entrypoint, which builds lnsd (release) so the test's locate_lnsd()
finds it, then runs the right variant:
bash scripts/run-soak.sh # smoke (~15 s + build)
bash scripts/run-soak.sh --full # heavy soak (minutes)
The script honours the ambient CARGO_TARGET_DIR so the binary lands where the
test looks, prints the effective parameters, and exits non-zero on failure. A
passing run ends with a PASS: block plus the rss plateau: and fds: lines.
Tuning
The soak reads these environment variables (defaults shown are the heavy-soak values; the smoke variant uses smaller ones):
| env | default | meaning |
|---|---|---|
LOADTEST_CONNS | 200 | steady concurrent TCP client connections |
LOADTEST_SECS | 60 | steady + churn duration (seconds) |
LOADTEST_PKT_MS | 50 | per-connection inter-packet interval (ms) |
LOADTEST_CHURN_WORKERS | 16 | connections repeatedly opened/closed |
LOADTEST_CHURN_PKTS | 4 | packets per churn connection before close |
LOADTEST_MAX_RSS_GROWTH_PCT | 40 | max steady-phase RSS growth |
LOADTEST_MAX_RSS_ABS_MIB | 300 | absolute RSS ceiling over baseline |
LOADTEST_DRAIN_SECS | 20 | post-load drain window for the fd check |
LOADTEST_SAMPLE_MS | 250 | RSS/fd/CPU sampler cadence |
LOADTEST_LNSD_BIN | auto | explicit path to the lnsd binary |
LEVICULUM_DELIVERY_LOG | unset | append one DELIVERY line per run to this file |
Sweeping delivery against load
The assertion is binary — 100 % or the run is red — which is right for a gate and useless for the question "at what load does the hub start to drop?". Codeberg #208 recorded one 99.5454 % run at 128 connections / 15 ms during the #198 A/B measurement, on a four-core host that was simultaneously running the measurement harness, and could attribute it to neither the hub nor the machine: the gate runs at 24 connections / 50 ms, and nobody had ever swept delivery against connection count and rate.
Every run therefore prints, and with LEVICULUM_DELIVERY_LOG=<file> also
appends, one line — before the assertions, so red runs contribute too:
DELIVERY test=lnsd_soak sent=377285 recv=375570 pct=99.5454 ci95=99.5234-99.5665 \
conns=128 pkt_ms=15 secs=20 churn_workers=16 churn_conns=42 cores=4 \
hub_cpu_pct=82.4 gen_cpu_pct=210.5 host_busy_pct=96.1
ci95 is the Wilson 95 % interval for the counts on the same line (the project
rule that a delivery ratio is never printed alone), at four decimals because a
hub run's denominator is in the hundreds of thousands, where two significant
digits would erase the very shortfall the line records. The cell coordinates are
on the line because two runs at different conns/pkt_ms offered different
volumes and cannot be pooled. The three CPU figures are what make a cell
interpretable, and the middle one — the load generator and the sampler, i.e. the
harness itself — is there because that is the cost #208 could not account for.
scripts/sweep-tcp-hub.sh drives the grid and reads the matrix back out of the
log:
bash scripts/sweep-tcp-hub.sh # 4x3 cells, 3 runs each
SWEEP_CONNS="128 192" SWEEP_PKT_MS="15 10" bash scripts/sweep-tcp-hub.sh
bash scripts/sweep-tcp-hub.sh --summarize <log> # re-read an earlier sweep
It refuses to start above SWEEP_MAX_LOAD1 (default 1.5), because the hub, the
sink and the generator all run on the sweep host and any other workload there is
indistinguishable from the hub being slow — the precise ambiguity #208 is about.
A red cell does not stop the sweep; the distribution is the point. What the
matrix is read for:
- a cell below 100 % while nothing is saturated is a defect in the hub;
- a cell below 100 % only at or past saturation is a load ceiling, and the finding is that the gate should name where the cliff is.
The sweep itself has not been run yet; #208 stays open until it has.
Where it runs regularly
The smoke soak is wired into the Tier 1 standard CI target (Justfile), which
is run once per batch of work, and again by just extensive and just nightly,
which depend on it — so a green soak is produced regularly and left on record.
Until 2026-08-07 a post-commit hook also ran it after every commit; that hook is
gone and Tier 1 is now started explicitly. See
CI Pipeline. The heavier --full soak is run on demand.
Real-world production soak: the miauhaus node
A permanent lnsd transport node (miauhaus) runs continuously in the public
Reticulum mesh, not in a lab harness. It has operated multi-day continuous as a
routing transport node, carrying real announce and path traffic, and has
survived a host reboot with no crash or self-reset observed. This is operational
endurance evidence alongside the synthetic soak: the daemon holds up under real,
unscripted mesh traffic over long uptimes.
Figures are kept deliberately conservative — multi-day continuous operation with no crash is what is directly observed and defensible; no precise uptime hours are claimed here.
Crash containment: poison-tolerant locking
Endurance is not only about not leaking; it is about a fault in one path not
taking down the whole process. The shared-state std::sync::Mutex sites in
leviculum-std were previously locked with .lock().unwrap(), so one task
panicking while holding a lock poisoned it and crashed every later locker —
turning an isolated task panic into a whole-daemon crash.
MutexRecover::lock_recover() (leviculum-std/src/sync_ext.rs) locks with
lock().unwrap_or_else(PoisonError::into_inner) and is applied uniformly to the
non-test std-mutex locks across the driver, interfaces, RPC, and event-log
paths. This is "continue degraded, do not crash," not "isolate one interface":
the dominant lock is the node-wide core mutex, held on every RPC, connect, and
dispatch, so recovery continues node-wide state that may be mid-mutation from the
task that panicked. The guarantee is that one task panic no longer cascades into
a whole-daemon crash that drops every peer the node routes for; the first such
recovery logs a tracing::warn so the degraded state is visible in the logs.
Reticulum Protocol Specification
This appendix is an exact, English specification of the Reticulum (RNS) protocol,
derived from and proven against the vendored Python reference
(reference/Reticulum, RNS 1.3.5, commit d5e62d4). It is the wire contract the
leviculum leviculum-core/leviculum-std crates interoperate against, and
the foundation the LXMF Protocol Specification builds
on.
Reticulum is a cryptography-based networking stack: self-sovereign identities, addressable destinations, encrypted packets, announces for discovery, links for sessions, resources for large transfers, and channels for ordered messaging, all medium-agnostic above a thin framing layer.
How to read this specification
- Normative statements use RFC 2119 keywords (MUST, SHOULD, MAY). Every
normative fact carries a
file:linecitation into the reference, a derivation with the arithmetic shown, or a labelled test vector[VEC-...]. - Informative sections describe internal reference behaviour (transport tables, retransmission timing, resource window adaptation, keepalive math) that an implementation MAY diverge from without breaking wire or semantic compatibility.
- Test vectors are the genuine byte output of the reference, regenerated by
vectors/gen_vectors.pyand pinned to the submodule commit. Ephemeral-key paths (encryption token, announces, link handshake) are frozen by injecting fixed randomness and time, and additionally carry a decrypt/verify/derive roundtrip so the semantic property is proven too.
Sections
- Introduction and scope
- Cryptographic primitives
- Identity
- Destination
- Packet format
- Announce
- Link
- Resource
- Channel and Buffer
- Transport
- Framing and IFAC
- Constants reference
- Coverage ledger
- Test vectors
- Symbol inventory (frozen)
Introduction and scope
Reference
This specification describes Reticulum as implemented by the pinned reference:
| Component | Version | Commit |
|---|---|---|
| Reticulum (RNS) | 1.3.5 | d5e62d4e15c5fe2e170f7bd9e120551671f21a27 |
Citations reference files under reference/Reticulum/RNS/ unless another path is
given.
Scope
Normative (specified exactly and proven):
- the cryptographic primitives as used (hashes, X25519, Ed25519, AES-CBC, HMAC, HKDF, the encryption token);
- identity key material, hashing, signing, and the ECIES encryption token;
- destination naming, hash derivation, and types;
- the packet header bitfield, both header types, addressing, context bytes, and sizes;
- announce layout, signing, and validation;
- link establishment (request, proof, link id, session key) and the link context-byte payloads;
- resource advertisement, parts, hashmap, requests, and proof;
- channel envelope and stream-data framing;
- the transport path request and path response packets;
- byte-stream framing (HDLC/KISS) and IFAC masking.
Informative (described, not byte-proven; an implementation MAY diverge):
- transport routing internals: path/announce/link/reverse/tunnel tables and their index shapes, announce retransmission timing and jitter, deduplication, table TTLs;
- resource flow-control window adaptation and timeout factors;
- keepalive interval math and link watchdog scheduling.
Out of scope: interface drivers beyond framing, the daemon/CLI tooling, and the shared-instance IPC.
The full enumeration and classification of reference symbols is the frozen Symbol inventory. The Coverage ledger maps every normative symbol to a section and a proof; a normative symbol with no mapping is a coverage gap.
Notational conventions
- RFC 2119 keywords (MUST, MUST NOT, SHOULD, MAY) mark normative requirements.
- Citations take the form
(Packet.py:177). - Byte layouts are shown as offset tables or annotated hex; concatenation is
a || b; a field width isname(16). - Test vectors are referenced by label
[VEC-...]and listed in full in Test vectors; machine-readable invectors/vectors.json. - Hashes are SHA-256 unless stated; the truncated hash is its leading 16
bytes (
Identity.py:383). Integers in masks/derivations are big-endian.
Vector kinds
- frozen — deterministic; the hex is the proof; reproduces byte for byte.
- frozen-injection — an ephemeral-key path (token, announce, link handshake,
path request) made reproducible by pinning
os.urandom,time, and X25519 ephemeral key generation in the harness, with a decrypt/verify/derive roundtrip recorded as the semantic proof. - computed — reconstructed per the source layout where building a live object needs a running link or transport (HEADER_2 packet, resource advertisement); the construction is byte-exact and cited.
Regenerating the vectors
From the repository root:
PYTHONPATH=reference/Reticulum \
python3 docs/src/appendix/reticulum/vectors/gen_vectors.py
The harness boots a headless RNS.Reticulum instance, fixes all key material,
runs the genuine reference code, asserts determinism for every frozen and
frozen-injection vector by rebuilding and comparing, and writes
vectors/vectors.json with the submodule commit in its meta block. Re-running
MUST reproduce the committed file byte for byte.
Cryptographic primitives
Reticulum builds every wire surface on a small set of primitives. An implementation MUST produce byte-identical results to these, because their outputs are hashed, signed, and exchanged on the wire.
Hashing
full_hash(x)is SHA-256 overx, 32 bytes (Identity.HASHLENGTH = 256bits,Identity.py:81,374).truncated_hash(x)is the leading 16 bytes offull_hash(x)(TRUNCATED_HASHLENGTH = 128bits,Identity.py:84,383).
Proven by [VEC-HASH]: full_hash("reticulum-spec") = 659fe249468c635cdfe90a12624abec49f0bd36ba66d467b4f7155c79e8addf2, truncated to
659fe249468c635cdfe90a12624abec4.
Key exchange and signing
- X25519 (
Cryptography/X25519.py): 32-byte keys;exchangeis constant-time ECDH yielding a 32-byte shared secret. Deterministic for a given key pair. - Ed25519 (
Cryptography/Ed25519.py): 32-byte seed, 32-byte public key, 64-byte signature; deterministic per RFC 8032.sign(m)andverify(sig, m).
Symmetric encryption
- AES-128/256-CBC (
Cryptography/AES.py) with a 16-byte IV and PKCS7 padding (Cryptography/PKCS7.py, block size 16).[VEC-AES]:AES-256-CBCof one block"0123456789abcdef"under key00010203…1fand IV000102…0fise23fc0b91c7bd64425c559736e9b0c58, and decrypts back.
HMAC and HKDF
- HMAC-SHA256 (
Cryptography/HMAC.py, RFC 2104), 32-byte digest. - HKDF-SHA256 (
Cryptography/HKDF.py:35),hkdf(length, derive_from, salt, context).[VEC-HKDF]:hkdf(32, derive_from=00..1f, salt=00..0f) = 2bc3faec9f360e81e77086b6e17a9ce8722a4cb3bc0ed90b4d78d37036e43a0f. The block counter is taken modulo 256 (Cryptography/HKDF.py:58), so unlike RFC 5869 there is no 255-block ceiling: outputs beyond 8160 bytes keep deriving, with the counter wrapping to 0 on block 256. IFAC needs this — its mask is as long as the packet.
Encryption token (modified Fernet)
The token (Cryptography/Token.py) is the AEAD-like envelope Reticulum uses for
SINGLE-destination and link encryption. Its layout is:
IV(16) || AES-CBC(plaintext, derived_key, IV) || HMAC-SHA256(...)(32)
with TOKEN_OVERHEAD = 48 (IV 16 + HMAC 32, Token.py:50). The HMAC
authenticates the IV and ciphertext; decrypt MUST verify it before decrypting
(Token.py:77,100). The full ECIES wrapper that prepends the ephemeral public
key is specified in Identity and proven by [VEC-ID-TOKEN].
Determinism
Hashing, HKDF, HMAC, Ed25519 signing, X25519 (given the keys), and AES-CBC (given key and IV) are deterministic and yield frozen vectors. Anything that generates an ephemeral key or IV (the encryption token, announces, link handshake) is non-deterministic in normal operation; this specification freezes those by injecting fixed randomness in the harness and additionally proves them by roundtrip (see Test vectors).
Identity
An identity is a pair of key pairs: X25519 for encryption and Ed25519 for
signing. This section is proven by [VEC-ID-HASH], [VEC-ID-SIGN], and
[VEC-ID-TOKEN].
Key material
The public key is the concatenation (Identity.py:811):
public_key = X25519_public(32) || Ed25519_public(32) # 64 bytes
and the private key is X25519_private(32) || Ed25519_seed(32)
(Identity.py:804,822-831). KEYSIZE = 512 bits (Identity.py:59),
SIGLENGTH = 512 bits (:81). [VEC-ID-HASH] records a 64-byte public key from
the fixed private material 00010203…3f.
Identity hash
identity_hash = truncated_hash(X25519_public || Ed25519_public) # 16 bytes
(Identity.py:859-864). [VEC-ID-HASH]: identity hash aca31af0441d81dbec71e82da0b4b5f5.
Name hash
NAME_HASH_LENGTH = 80 bits (Identity.py:84); the name hash is the leading 10
bytes of full_hash of the dotted destination name. Used in destination hashing
and announces; see Destination.
Signing
sign(m)is Ed25519 overm, 64 bytes, deterministic (Identity.py:931-941).validate(sig, m)verifies it (Identity.py:948-964).
[VEC-ID-SIGN]: signing the fixed message yields signature
bbfdcde5aa05197f… and validate returns true.
Encryption token (ECIES)
encrypt(plaintext) to a SINGLE destination produces (Identity.py:881-911):
ephemeral_X25519_public(32) || token(IV(16) || AES-CBC ciphertext || HMAC(32))
The derivation:
- generate an ephemeral X25519 key pair (
Identity.py:890); shared = ephemeral_private.exchange(target_X25519_public)(:844);derived_key = hkdf(length=64, derive_from=shared, salt=target_identity_hash, context=None)(:846-851);token = Token(derived_key).encrypt(plaintext)(:854).
DERIVED_KEY_LENGTH = 64 bytes (Identity.py:91): 32 for the AES-256 key and 32
for the HMAC key. decrypt recovers the ephemeral public key from the first 32
bytes, re-derives the key, and (when ratchets are present) tries each ratchet
before the base identity (Identity.py:872-928).
Because the ephemeral key and IV are random, the token is not reproducible in
normal operation. [VEC-ID-TOKEN] is a frozen-injection vector: under pinned
randomness the 112-byte token (32 ephemeral + 48 overhead + 32 ciphertext for a
27-byte plaintext) reproduces byte for byte, and decrypt(token) == plaintext
holds. An implementation MUST reproduce this construction so a Python peer can
decrypt its packets.
Ratchets
A ratchet is an ephemeral X25519 key offering forward secrecy. RATCHETSIZE = 256 bits (Identity.py:64); the ratchet id is the leading 10 bytes of
full_hash of the ratchet public key (_get_ratchet_id, Identity.py:417-418). A
destination MAY advertise its current ratchet in an announce (see
Announce); a sender then encrypts to the ratchet public key
instead of the identity's static key. Ratchet rotation and expiry
(RATCHET_EXPIRY, RATCHET_INTERVAL) are informative.
Destination
A destination is an addressable endpoint. This section is proven by
[VEC-DEST-HASH].
Naming and hashing
A destination name is the dotted concatenation of an app name and aspects, e.g.
test.vec (expand_name, Destination.py:96). The hashes are
(Destination.py:116-141):
name_hash = full_hash(app_name [. aspect ...])[:10] # 80 bits
destination_hash = truncated_hash(name_hash || identity_hash) # 16 bytes
[VEC-DEST-HASH] for app test, aspect vec, identity hash
069092a03c194639207219dd05f9c840: name hash 9da53eec82a28ce2f2e9 (10 bytes),
destination hash 07d4541d4fdc0abfacc9364fdf979ee1 (16 bytes). An implementation
MUST derive these identically, since the destination hash is the on-wire address
and is recomputed by every receiver of an announce.
Types
| Type | Value | Encryption | Citation |
|---|---|---|---|
SINGLE | 0x00 | per-recipient ECIES token to the identity (or ratchet) | Destination.py:63 |
GROUP | 0x01 | symmetric (shared key, not auto-distributed) | :64 |
PLAIN | 0x02 | none (cleartext) | :65 |
LINK | 0x03 | per-link session key | :66 |
The type occupies bits 3-2 of the packet header (see Packet format).
Direction and proof strategy
- Direction:
IN = 0x11,OUT = 0x12(Destination.py:79-80). OnlyINSINGLE destinations may be announced (Destination.py:251-255). - Proof strategy:
PROVE_NONE = 0x21,PROVE_APP = 0x22,PROVE_ALL = 0x23(Destination.py:69-71) controls whether the destination automatically returns delivery proofs (see Packet format proofs).
Encryption and decryption
Destination.encrypt/decrypt (Destination.py:585,611) delegate to the
identity's token for SINGLE destinations, applying the current ratchet when
enabled. This is the path LXMF and link setup use for SINGLE-addressed payloads.
Packet format
This section is normative and proven by [VEC-PKT-PLAIN], [VEC-PKT-ENC], and
[VEC-PKT-HEADER2].
Header bitfield
The first byte is a bitfield, packed by Packet.get_packed_flags
(Packet.py:169-175) and decoded in unpack (:247-251):
bit 7 IFAC flag (set by the interface, not by pack; see Framing/IFAC)
bit 6 header type 0 = HEADER_1, 1 = HEADER_2
bit 5 context flag context-specific; set when an announce carries a ratchet
bit 4 transport type 0 = BROADCAST, 1 = TRANSPORT
bits3-2 destination type 00 SINGLE, 01 GROUP, 10 PLAIN, 11 LINK
bits1-0 packet type 00 DATA, 01 ANNOUNCE, 10 LINKREQUEST, 11 PROOF
packed_flags = (header_type<<6) | (context_flag<<5) | (transport_type<<4) | (destination_type<<2) | packet_type.
HEADER_1 layout
flags(1) || hops(1) || destination_hash(16) || context(1) || data
offset 0 1 2 18 19
Packed by Packet.pack (Packet.py:177-239), read back by unpack
(:262-264). The fixed header is
HEADER_MINSIZE = 19 bytes (Reticulum.py:144).
[VEC-PKT-PLAIN] is a PLAIN HEADER_1 DATA packet,
0800fc0910664040482cd653166c8f225520006869:
08 flags: PLAIN(0x08) + DATA, HEADER_1, BROADCAST
00 hops
fc0910664040482cd653166c8f225520 destination_hash(16)
00 context (NONE)
6869 data = "hi"
The flags decode to ifac 0, header_type 0, context_flag 0, transport_type 0,
destination_type 2 (PLAIN), packet_type 0 (DATA). [VEC-PKT-ENC] is the SINGLE
encrypted equivalent (flags 00); under injection its encrypted body reproduces
byte for byte.
HEADER_2 layout
When the header type bit is set, a 16-byte transport id precedes the destination
hash (Packet.py:256-260):
flags(1) || hops(1) || transport_id(16) || destination_hash(16) || context(1) || data
offset 0 1 2 18 34 35
HEADER_MAXSIZE = 35 bytes (Reticulum.py:145). [VEC-PKT-HEADER2] shows the
constructed layout 4000 a0..af b0..bf 00 64617461: flags 40 (header_type 1),
transport id a0a1…af, destination hash b0b1…bf. HEADER_2 is emitted by
transport nodes forwarding toward a known next hop.
Sizes
MTU = 500 (Reticulum.py:93)
HEADER_MAXSIZE = 2 + (128/8)*2 + 1 = 35 (Reticulum.py:148)
MDU = 500 - 35 - 1 = 464 (Reticulum.py:152)
Packet.ENCRYPTED_MDU = 383 (Packet.py:106)
Packet.PLAIN_MDU = MDU = 464 (Packet.py:110)
Context bytes
The context byte (offset 18 for HEADER_1) selects the packet's role within its
type. Full set (Packet.py:72-92):
| Hex | Name | Used by |
|---|---|---|
| 0x00 | NONE | generic data |
| 0x01 | RESOURCE | resource part |
| 0x02 | RESOURCE_ADV | resource advertisement |
| 0x03 | RESOURCE_REQ | resource part request |
| 0x04 | RESOURCE_HMU | resource hashmap update |
| 0x05 | RESOURCE_PRF | resource proof |
| 0x06 | RESOURCE_ICL | resource initiator cancel |
| 0x07 | RESOURCE_RCL | resource receiver cancel |
| 0x08 | CACHE_REQUEST | cache request |
| 0x09 | REQUEST | link request (application) |
| 0x0A | RESPONSE | link response |
| 0x0B | PATH_RESPONSE | transport path response |
| 0x0C | COMMAND | command |
| 0x0D | COMMAND_STATUS | command status |
| 0x0E | CHANNEL | channel data |
| 0xFA | KEEPALIVE | link keepalive |
| 0xFB | LINKIDENTIFY | link identification |
| 0xFC | LINKCLOSE | link close |
| 0xFD | LINKPROOF | (deprecated) |
| 0xFE | LRRTT | link RTT |
| 0xFF | LRPROOF | link request proof |
Proofs
A delivery proof is a PROOF packet over the original packet hash
(Packet.get_hashable_part, :355; validate_proof, :498). Two forms exist:
explicit (packet_hash(32) || signature(64), EXPL_LENGTH = 96) and implicit
(signature(64), IMPL_LENGTH = 64). The reference currently emits explicit
proofs (Link.prove_packet, Link.py:383). Whether a destination proves is
governed by its proof strategy (see Destination).
Announce
An announce broadcasts a destination's keys so peers can address and route to it.
This section is proven by [VEC-ANN-NORATCHET] and [VEC-ANN-RATCHET].
Packet
An announce is an ANNOUNCE-type packet (packet_type = 0x01) to the announced
SINGLE destination. The context flag (header bit 5) is set when a ratchet is
included (Destination.py:310-311). [VEC-ANN-NORATCHET] flags byte 01
(ANNOUNCE, context flag 0); [VEC-ANN-RATCHET] flags byte with context flag 1.
Announce data layout
The packet data is (Destination.py:301):
public_key(64) || name_hash(10) || random_hash(10) || [ratchet(32)] || signature(64) || app_data
The ratchet field is present only when the context flag is set. Without it the
data is 64 + 10 + 10 + 64 = 148 bytes plus app_data; with it, 180 plus
app_data. [VEC-ANN-NORATCHET] carries 150 data bytes (148 + 2-byte app_data);
[VEC-ANN-RATCHET] carries 182 (180 + 2).
The random_hash is get_random_hash()[0:5] || int(time.time()).to_bytes(5, "big") (Destination.py:282): 5 random bytes and a
5-byte big-endian Unix timestamp. It makes each announce unique and lets
receivers reject replays.
Signed data
The signature covers (Destination.py:297-300):
signature = sign( destination_hash || public_key || name_hash || random_hash || [ratchet] || app_data )
The destination hash is signed but not transmitted in the announce data; the receiver recomputes it from the transmitted keys (below). This binds the keys to the address without spending 16 bytes on the wire.
Validation
A receiver validates an announce by (Identity.py:568-670):
- parsing
public_key = data[:64], thenname_hash,random_hash, optionalratchet(when the context flag is set),signature, andapp_dataat the offsets above (Identity.py:582-600); - reconstructing
signed_dataand verifying the signature against the transmitted public key (Identity.py:579); - recomputing
expected_destination_hash = truncated_hash(name_hash || identity_hash)and checking it matches (Identity.py:584-585); - remembering the public key, app_data, and (if present) the ratchet for future
encryption (
Identity.py:634,654-655).
[VEC-ANN-NORATCHET] and [VEC-ANN-RATCHET] are frozen-injection vectors:
under pinned randomness and time the announce reproduces byte for byte, and the
genuine validate_announce returns true. An implementation MUST reproduce the
signed-data order and the destination-hash recomputation, or a Python peer
rejects the announce.
Path response
A path response is an announce re-emitted with context PATH_RESPONSE (0x0B)
rather than NONE, in reply to a path request (see Transport).
The announce data is identical; only the packet context differs.
Link
A link is an ephemeral, forward-secret session between two destinations,
established by an ECDH handshake. This section is proven by [VEC-LINK].
Link request
The initiator sends a LINKREQUEST packet (packet_type = 0x02) whose data is
(Link.py:308-317):
ephemeral_X25519_public(32) || ephemeral_Ed25519_public(32) || signalling(3)
ECPUBSIZE = 64 (the two public keys). The 3-byte signalling field encodes the
proposed link MTU (21 bits) and mode (3 bits): (mtu & 0x1FFFFF) | ((mode<<5 & 0xE0)<<16), packed big-endian, low 3 bytes (Link.signalling_bytes,
Link.py:148-151). The only enabled mode is MODE_AES256_CBC = 0x01.
Link id
hashable = packet.get_hashable_part() # masked-flags byte || addressing || data
if len(packet.data) > ECPUBSIZE: # strip trailing signalling bytes
hashable = hashable[:-(len(packet.data) - ECPUBSIZE)]
link_id = truncated_hash(hashable) # 16 bytes
(Link.link_id_from_lr_packet, Link.py:340-347): the hashable part is trimmed by
the number of bytes the request data exceeds the 64-byte key block (i.e. the
signalling bytes) before hashing.
[VEC-LINK] link id 4725ac1375601d182afec3610f019b25. The link id replaces the
destination hash in the addressing of all subsequent link packets, and is also
the link's salt (below).
Proof and handshake
The responder replies with a PROOF packet, context LRPROOF (0xFF), data
(Link.py:371-377):
signature(64) || ephemeral_X25519_public(32) || signalling(3)
where signature = sign( link_id || responder_eph_X25519_pub || responder_eph_Ed25519_pub || signalling ) (Link.py:373). The initiator
validates it against the destination's known identity (Link.py:417-420).
Both sides then derive the session key (Link.handshake, Link.py:353-366):
shared = own_ephemeral_private.exchange(peer_ephemeral_public) # X25519 ECDH
session_key = hkdf(length=64, derive_from=shared, salt=link_id, context=None)
64 bytes for MODE_AES256_CBC (32 key + 32 HMAC). get_salt() returns the link
id and get_context() returns None (Link.py:643,646). [VEC-LINK] proves
the ECDH is symmetric (ecdh_agreement = true, both sides compute the same
shared secret) and records the resulting 64-byte session key
569ac51a07fb242f…. An implementation MUST derive the link id and session key
identically.
Link operations (context bytes)
Once active, link packets are DATA packets addressed by link id, encrypted with the session key, distinguished by context byte:
| Context | Hex | Encrypted | Payload | Citation |
|---|---|---|---|---|
LRPROOF | 0xFF | no | sig(64) + eph_pub(32) + signalling(3) | Link.py:371-377 |
LRRTT | 0xFE | yes | msgpack(float rtt) | Link.py:440 |
LINKIDENTIFY | 0xFB | yes | identity_public(32) + sign(link_id||public)(64) | Link.py:459-471 |
KEEPALIVE | 0xFA | no | single byte 0xFF | Link.py:848-851 |
LINKCLOSE | 0xFC | yes | link_id | — |
After proof, the initiator measures RTT and sends an LRRTT packet; either side
MAY identify (prove an identity over the link) by sending a LINKIDENTIFY
packet. Keepalive cadence, the stale/close watchdog, and MTU discovery are
informative.
Resource
A resource is a reliable, segmented transfer over a link for data larger than a
single packet. This section is proven by [VEC-RES-ADV] and [VEC-RES-PROOF].
Window adaptation and timeout scheduling are informative.
Advertisement
The sender advertises a resource with a RESOURCE_ADV packet (context 0x02)
carrying a msgpack dictionary (ResourceAdvertisement.pack, Resource.py:1330-1352):
| Key | Meaning |
|---|---|
t | transfer size (encrypted bytes) |
d | total uncompressed data size |
n | number of parts |
h | resource hash (32) |
r | random hash (4) |
o | original (first-segment) hash (32) |
i | segment index |
l | total segments |
q | associated request id, or nil |
f | flags byte |
m | hashmap segment (4-byte MAPHASH_LEN entries) |
The flags byte is (Resource.py:1304):
f = (has_metadata<<5) | (is_response<<4) | (is_request<<3) | (split<<2) | (compressed<<1) | encrypted
[VEC-RES-ADV] is a computed vector (a live advertisement needs a Resource
over a link): the dictionary with compressed and encrypted set yields flags
03 and packs to 146 bytes under the genuine msgpack. MAPHASH_LEN = 4 and
RANDOM_HASH_SIZE = 4.
Parts, requests, and hashmap
- A part is a RESOURCE packet (context 0x01) carrying up to one SDU of pre-encrypted data.
- The receiver requests parts with a RESOURCE_REQ packet (context 0x03) whose
first byte is the hashmap status (
HASHMAP_IS_EXHAUSTED = 0xFF/HASHMAP_IS_NOT_EXHAUSTED = 0x00), optionally followed by the last received map index and the requested 4-byte part hashes. - The sender extends the hashmap with a RESOURCE_HMU packet (context 0x04)
carrying
resource_hash || msgpack([segment_index, hashmap_bytes]).
The hashmap lets the receiver request parts by hash, and the sender stream
hashes as the window advances. The sliding window sizes (WINDOW, WINDOW_MIN,
WINDOW_MAX*) and the part/proof timeout factors are informative.
Proof and cancellation
On completion the receiver assembles the parts, verifies integrity, and the
sender sends a RESOURCE_PRF packet (context 0x05, unencrypted) carrying
(Resource.py:752-753):
proof = full_hash(data || resource_hash)
proof_data = resource_hash(32) || proof(32)
a single SHA-256 over the assembled data concatenated with the resource hash,
prefixed by the resource hash. validate_proof accepts when proof_data is 64
bytes and its second half matches the expected proof (Resource.py:779-783).
[VEC-RES-PROOF] records this construction. Either party may abort with
RESOURCE_ICL (0x06, initiator) or RESOURCE_RCL (0x07, receiver).
Compression and segmentation
A resource MAY be bz2-compressed (the compressed flag) and, when larger than a
single segment, split into total_segments segments chained by the original
hash. The maximum efficient single-segment size and metadata limits are
implementation guidance.
Channel and Buffer
A channel is an ordered, typed messaging layer over a link; a buffer is a
byte-stream abstraction built on a channel. This section is proven by
[VEC-CHAN-ENVELOPE] and [VEC-STREAM-HDR].
Channel envelope
Every channel message is wrapped in an envelope (Channel.Envelope.pack,
Channel.py:174-200):
msgtype(u16, big-endian) || sequence(u16, big-endian) || length(u16, big-endian) || payload
[VEC-CHAN-ENVELOPE]: msgtype 0xabcd, sequence 7, payload "channeldata" packs
to abcd0007000b6368616e6e656c64617461 — abcd type, 0007 sequence, 000b
length 11, then the payload. Message types 0xF000 and above are reserved for
system messages.
Stream data message
The buffer layer sends StreamDataMessages (type SMT_STREAM_DATA = 0xff00) over
a channel. Each carries a 2-byte header (Buffer.py:80-92):
header(u16, big-endian):
bits 0-13 stream_id (0..STREAM_ID_MAX = 0x3fff)
bit 14 compressed
bit 15 eof
then: data
header = (stream_id & 0x3fff) | (0x8000 if eof) | (0x4000 if compressed).
[VEC-STREAM-HDR]: stream id 0x0102, eof set, not compressed, payload
"streamdata" packs to 810273747265616d64617461 — 8102 header (eof bit set
over stream id 0x0102), then the data.
The combined overhead is 8 bytes (2-byte stream header + 6-byte channel envelope),
so MAX_DATA_LEN = Link.MDU - 8. Channel sequencing, acknowledgement, and
retransmission are informative.
Transport
Transport routes packets across the mesh: it discovers paths, forwards packets
toward known next hops, and propagates announces. The wire surfaces (path
request and path response packets) are normative and proven by
[VEC-PATH-REQUEST]. The routing internals (tables, retransmission timing,
deduplication, TTLs) are informative.
Path request
To discover a path, a node sends a path request: a DATA packet to the PLAIN
destination rnstransport.path.request (Transport.request_path,
Transport.py:2786-2787). The payload is (Transport.py:2849-2850):
target_destination_hash(16) || [transport_identity_hash(16)] || request_tag
The transport_identity_hash is included only when transport is enabled on the
requesting node; the request_tag is a random hash that deduplicates the request
(Transport.py:2780). [VEC-PATH-REQUEST] is a frozen-injection vector with
transport disabled: payload target(16) || tag(16), in a PLAIN HEADER_1 packet
(flags 08). An implementation MUST address the request to this PLAIN
destination and use this payload order so existing nodes answer it.
Path response
A node that holds a path answers by re-emitting the destination's cached announce
with the packet context set to PATH_RESPONSE (0x0B) instead of NONE
(Transport.py:595-596). The announce data is byte-identical to an ordinary
announce (see Announce); only the context differs, and the
rebroadcast is marked so it is not propagated further.
Routing internals (informative)
The following are reference behaviour an implementation MAY diverge from:
- Tables. Transport maintains path, announce, reverse, link, and tunnel
tables, each a list indexed by the
IDX_*constants (Transport.py:3547-3586). A path entry holds timestamp, next hop, hops, expiry, the announce random blobs, the receiving interface, and the cached packet hash. - Announce propagation. Announces are rebroadcast with a hop limit
PATHFINDER_M = 128, up toLOCAL_REBROADCASTS_MAX = 2local rebroadcasts, after a gracePATHFINDER_G = 5 splus random jitterPATHFINDER_RW = 0.5 s(PATHFINDER_M,Transport.py:63-77). - Path TTLs. Default
PATHFINDER_E = 7 days; access-point pathsAP_PATH_TIME = 1 day; roaming pathsROAMING_PATH_TIME = 6 hours(PATHFINDER_E,Transport.py:71-73). - Deduplication. A rolling table of recent packet hashes suppresses loops.
- Path request pacing.
PATH_REQUEST_TIMEOUT = 15 s,PATH_REQUEST_MI = 20 sminimum interval, and per-interface announce caps (PATH_REQUEST_TIMEOUT,Transport.py:79-83).
These values and structures are documented for fidelity; only the path request and path response packets above are normative.
Framing and IFAC
This section covers how packets are delimited on byte-stream interfaces (HDLC)
and how an interface authenticates and masks packets (IFAC). Both are normative
for wire interop on those media and proven by [VEC-HDLC] and [VEC-IFAC].
Medium specifics beyond framing (LoRa airtime, TCP particulars) are informative.
HDLC byte-stream framing
Byte-stream interfaces (TCP, serial, pipe) delimit packets with HDLC-style flags
and byte stuffing (Interfaces/TCPInterface.py:44-52,323):
FLAG = 0x7E
ESC = 0x7D
ESC_MASK = 0x20
frame = FLAG || escape(packet) || FLAG
escape: replace 0x7D -> 0x7D 0x5D, then 0x7E -> 0x7D 0x5E
A literal flag or escape byte inside the packet is replaced by the escape byte
followed by the original XORed with 0x20. [VEC-HDLC]: input 01 7E 02 7D 03
frames to 7e017d5e027d5d037e — leading flag, 01, escaped 7E→7D5E, 02,
escaped 7D→7D5D, 03, trailing flag. A receiver un-stuffs by reversing the
replacement between flags. (KISS interfaces use the analogous FEND/FESC framing.)
IFAC (interface access codes)
An interface configured with a passphrase derives a 64-byte IFAC identity and an
IFAC key, then authenticates and masks every packet (Transport.transmit,
Transport.py:1052-1087):
ifac = ifac_identity.sign(raw)[-ifac_size:]
mask = hkdf(length = len(raw) + ifac_size, derive_from = ifac, salt = ifac_key, context = None)
new_raw = (raw[0] | 0x80) || raw[1] || ifac || raw[2:]
masked[i] = new_raw[i] ^ mask[i] for i == 0 (then re-set bit 7), i == 1, and i > ifac_size+1
masked[i] = new_raw[i] for the ifac bytes (2 .. ifac_size+1, left unmasked)
The IFAC flag (header bit 7) is set, the ifac tag is inserted right after the
two header bytes, and everything except the tag itself is XOR-masked with the
HKDF stream. On receipt the interface reverses the mask, extracts the tag,
recomputes ifac_identity.sign(recovered_raw)[-ifac_size:], and drops the packet
on mismatch (Transport.inbound, Transport.py:1398-1435).
[VEC-IFAC] (ifac_size = 8) records the tag, the mask, the masked output
9f4a2cf485c1dfcea0…, the IFAC flag set in the masked header, and proves a full
mask/unmask roundtrip recovers the original packet and a matching tag
(unmask_roundtrip_ok = true). IFAC_MIN_SIZE = 1, and IFAC_SALT is a fixed
32-byte constant (Reticulum.py:146-147). An implementation sharing an interface
with Python peers MUST reproduce this masking exactly or its packets are dropped.
Constants reference
Grouped constants with values and citations. Derived sizes are captured in
vectors.json constants.
System (Reticulum.py)
| Constant | Value | Line |
|---|---|---|
MTU | 500 | 93 |
MDU | 464 | 152 |
TRUNCATED_HASHLENGTH | 128 bits | 145 |
HEADER_MINSIZE | 19 | 147 |
HEADER_MAXSIZE | 35 | 148 |
IFAC_MIN_SIZE | 1 | 149 |
Identity (Identity.py)
| Constant | Value | Line |
|---|---|---|
KEYSIZE | 512 bits | 59 |
RATCHETSIZE | 256 bits | 64 |
TOKEN_OVERHEAD | 48 | 77 |
HASHLENGTH | 256 bits | 80 |
SIGLENGTH | 512 bits | 81 |
NAME_HASH_LENGTH | 80 bits | 83 |
DERIVED_KEY_LENGTH | 64 | 90 |
Destination (Destination.py)
SINGLE 0x00, GROUP 0x01, PLAIN 0x02, LINK 0x03 (63-66); PROVE_NONE 0x21,
PROVE_APP 0x22, PROVE_ALL 0x23 (69-71); IN 0x11, OUT 0x12 (79-80).
Packet (Packet.py)
Packet types DATA 0x00, ANNOUNCE 0x01, LINKREQUEST 0x02, PROOF 0x03
(60-63). Header types HEADER_1 0x00, HEADER_2 0x01 (67-68). FLAG_SET 0x01,
FLAG_UNSET 0x00 (95-96). ENCRYPTED_MDU 383 (106), PLAIN_MDU 464 (110).
Context bytes 0x00-0xFF (72-92) — full table in Packet format.
Proof lengths EXPL_LENGTH 96, IMPL_LENGTH 64.
Link (Link.py)
ECPUBSIZE 64, KEYSIZE 32, LINK_MTU_SIZE 3, MTU_BYTEMASK 0x1FFFFF,
MODE_BYTEMASK 0xE0, MODE_AES256_CBC 0x01. States PENDING 0x00 ..
CLOSED 0x04 (informative). Context bytes KEEPALIVE 0xFA, LINKIDENTIFY 0xFB,
LINKCLOSE 0xFC, LINKPROOF 0xFD, LRRTT 0xFE, LRPROOF 0xFF.
Resource (Resource.py)
MAPHASH_LEN 4, RANDOM_HASH_SIZE 4, HASHMAP_IS_EXHAUSTED 0xFF,
HASHMAP_IS_NOT_EXHAUSTED 0x00, advertisement OVERHEAD 134 (1235). Context
bytes RESOURCE 0x01 .. RESOURCE_RCL 0x07 (Packet.py:73-79). Window sizes and
timeout factors are informative.
Channel / Buffer (Channel.py, Buffer.py)
SMT_STREAM_DATA 0xff00, STREAM_ID_MAX 0x3fff, combined OVERHEAD 8.
Transport (Transport.py) — informative
PATHFINDER_M 128, PATHFINDER_R 1, PATHFINDER_G 5 s, PATHFINDER_RW 0.5 s,
PATHFINDER_E 7 days, AP_PATH_TIME 1 day, ROAMING_PATH_TIME 6 h,
LOCAL_REBROADCASTS_MAX 2, PATH_REQUEST_TIMEOUT 15 s, PATH_REQUEST_MI 20 s
(50-83). Table index constants IDX_* (3547-3586).
Coverage ledger
Traceability matrix from the frozen Symbol inventory to the
specification. Every normative (N) symbol maps to a section and a proof; every
informative (I) and out-of-scope (X) symbol carries a reason. Proof: vector
(a [VEC-...]), computed (derivation shown), quoted (value cited verbatim),
n/a (informative/out-of-scope).
Reticulum.py / Cryptography
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| MTU, MDU, HEADER sizes, TRUNCATED_HASHLENGTH | Reticulum.py:93-152 | N | 04 | computed (vector constants) |
| IFAC_MIN_SIZE, IFAC_SALT | Reticulum.py:149-150 | N | 10 | quoted |
| full/truncated hash | Identity.py:373-390 | N | 01 | vector VEC-HASH |
| HKDF, HMAC, AES, PKCS7 | Cryptography/* | N | 01 | vector VEC-HKDF/HMAC/AES |
| Token (modified Fernet), TOKEN_OVERHEAD | Token.py | N | 01,02 | vector VEC-ID-TOKEN |
| X25519, Ed25519 | Cryptography/* | N | 01,02 | vector VEC-ID-SIGN/LINK |
Identity.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| KEYSIZE, HASHLENGTH, SIGLENGTH, NAME_HASH_LENGTH, DERIVED_KEY_LENGTH, RATCHETSIZE | 59-90 | N | 02 | quoted/computed |
| key material, identity hash, get_public_key | 750-810 | N | 02 | vector VEC-ID-HASH |
| sign / validate | 931-964 | N | 02 | vector VEC-ID-SIGN |
| encrypt / decrypt (token) | 827-928 | N | 02 | vector VEC-ID-TOKEN |
| validate_announce | 532-634 | N | 05 | vector VEC-ANN-* |
| ratchet id / generation | 417-425 | N | 02 | quoted |
| RATCHET_EXPIRY, recall/remember | 69,— | I | 02 | n/a (rotation/resolution) |
Destination.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| SINGLE/GROUP/PLAIN/LINK, PROVE_*, IN/OUT | 63-80 | N | 03,04 | quoted |
| name hash / destination hash | 116-141 | N | 03 | vector VEC-DEST-HASH |
| announce (data + signed data) | 243-317 | N | 05 | vector VEC-ANN-* |
| encrypt / decrypt | 585-611 | N | 03 | quoted (VEC-ID-TOKEN) |
| RATCHET_COUNT/INTERVAL, PR_TAG_WINDOW | 83-90 | I | 02 | n/a |
Packet.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| packet types, header types, context bytes, FLAG_* | 60-96 | N | 04 | quoted |
| ENCRYPTED_MDU / PLAIN_MDU | 106-110 | N | 04 | computed |
| pack / unpack | 177-272 | N | 04 | vector VEC-PKT-PLAIN/ENC/HEADER2 |
| get_hashable_part, validate_proof, EXPL/IMPL_LENGTH | 355,498 | N | 04 | quoted |
| PacketReceipt states | 408-415 | I | 04 | n/a (local) |
Link.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| ECPUBSIZE, KEYSIZE, LINK_MTU_SIZE, masks, MODE_AES256_CBC | — | N | 06 | quoted |
| link_id_from_lr_packet, set_link_id | 340-351 | N | 06 | vector VEC-LINK |
| handshake (session key), get_salt/get_context | 353-366,643 | N | 06 | vector VEC-LINK |
| prove, signalling_bytes | 371-377,148 | N | 06 | vector VEC-LINK |
| identify, send_keepalive, context payloads | 459,848 | N | 06 | quoted |
| states, KEEPALIVE/STALE timing, watchdog | — | I | 06 | n/a |
Resource.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| ResourceAdvertisement.pack, flags, keys | 1278-1355 | N | 07 | vector VEC-RES-ADV |
| MAPHASH_LEN, RANDOM_HASH_SIZE, HASHMAP_* | — | N | 07 | quoted |
| prove / validate_proof | 752-786 | N | 07 | vector VEC-RES-PROOF |
| context bytes RESOURCE..RESOURCE_RCL | Packet.py:73-79 | N | 04,07 | quoted |
| WINDOW*, timeout factors, advertise/assemble | — | I | 07 | n/a (flow control) |
| status enum | — | I | 07 | n/a |
Channel.py / Buffer.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| Envelope.pack/unpack | Channel.py:174-200 | N | 08 | vector VEC-CHAN-ENVELOPE |
| StreamDataMessage header, SMT_STREAM_DATA, STREAM_ID_MAX, OVERHEAD | Buffer.py:80-92 | N | 08 | vector VEC-STREAM-HDR |
| MessageState, CEType, sequencing | — | I | 08 | n/a |
Transport.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| transmit / inbound (IFAC) | 1051-1434 | N | 10 | vector VEC-IFAC |
| request_path (path request) | 2771-2787 | N | 09 | vector VEC-PATH-REQUEST |
| path response (PATH_RESPONSE rebroadcast) | 2943-2972 | N | 09 | quoted |
| PATHFINDER_*, path TTLs, pacing | 50-83 | I | 09 | n/a (routing) |
| path/announce/link/reverse/tunnel tables, IDX_* | 3547-3586 | I | 09 | n/a (internal) |
| dedup, jobs, table culling | 508,— | I | 09 | n/a |
Interfaces
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
| HDLC FLAG/ESC/ESC_MASK, escape | TCPInterface.py:44-52,323 | N | 10 | vector VEC-HDLC |
| KISS FEND/FESC framing | KISSInterface.py | N | 10 | quoted |
| interface drivers (TCP/LoRa/serial specifics) | Interfaces/* | X | — | n/a (medium drivers) |
Result
Every normative symbol maps to a section and a proof; no normative row has an
empty section or n/a proof. Informative and out-of-scope rows are reasoned.
Coverage is complete against the frozen inventory at commit d5e62d4. Re-auditing
after a reference bump: re-enumerate the source and diff the inventory; a new
symbol appears here unclassified.
Test vectors
The golden vectors that prove the binary claims in this specification. They are
the genuine output of the reference, generated by
vectors/gen_vectors.py and stored in
vectors/vectors.json.
Regenerate and verify (from the repository root):
PYTHONPATH=reference/Reticulum \
python3 docs/src/appendix/reticulum/vectors/gen_vectors.py
Pinned to RNS 1.3.5 @ d5e62d4. Fixed inputs: source identity private =
0001…3f, destination identity private = 4041…7f, time = 1700000000.0.
Vector kinds
- frozen — deterministic; the hex is the proof.
- frozen-injection — ephemeral-key path made reproducible by pinning
os.urandom/time/X25519 generation, with a decrypt/verify/derive roundtrip. - computed — reconstructed per source layout where a live link/transport is needed; byte-exact and cited.
Primitives
- VEC-HASH (frozen,
Identity.py:409-426):full_hash("reticulum-spec") = 659fe249468c635cdfe90a12624abec49f0bd36ba66d467b4f7155c79e8addf2; truncated659fe249468c635cdfe90a12624abec4. - VEC-HKDF (frozen,
HKDF.py:35):hkdf(32, 00..1f, salt=00..0f) = 2bc3faec9f360e81e77086b6e17a9ce8722a4cb3bc0ed90b4d78d37036e43a0f. - VEC-HMAC (frozen): HMAC-SHA256 over the fixed key/message.
- VEC-AES (frozen): AES-256-CBC one block →
e23fc0b91c7bd64425c559736e9b0c58, decrypts back.
Identity
- VEC-ID-HASH (frozen): 64-byte public key from
0001…3f; identity hashaca31af0441d81dbec71e82da0b4b5f5. - VEC-ID-SIGN (frozen): Ed25519 signature
bbfdcde5aa05197f…,validatetrue. - VEC-ID-TOKEN (frozen-injection,
encrypt,Identity.py:827-928): 112-byte tokenefd9ec3449e46df2…= ephemeral_pub(32) || IV(16) || ciphertext || HMAC(32);decrypt(token) == plaintext.
Destination
- VEC-DEST-HASH (frozen): app
test, aspectvec→ name hash9da53eec82a28ce2f2e9(10), destination hash07d4541d4fdc0abfacc9364fdf979ee1(16).
Packet
- VEC-PKT-PLAIN (frozen):
0800fc0910664040482cd653166c8f225520006869— flags08(PLAIN/DATA), hops00, dest(16), context00, data"hi". - VEC-PKT-ENC (frozen-injection): SINGLE encrypted HEADER_1 packet, flags
00. - VEC-PKT-HEADER2 (computed,
Packet.py:256-260):4000 a0..af b0..bf 00 64617461— flags40(HEADER_2), transport id(16), dest(16), context, data.
Announce
- VEC-ANN-NORATCHET (frozen-injection): flags
01(context flag 0); 150 data bytes79a631eede1bf9c9…= pubkey(64) || name(10) || random(10) || sig(64) || app_data(2);validate_announcetrue. - VEC-ANN-RATCHET (frozen-injection): context flag 1; 182 data bytes including
the 32-byte ratchet;
validate_announcetrue.
Link
- VEC-LINK (frozen-injection,
Link.py:308-366): request dataa4e09292…= eph_x25519(32) || eph_ed25519(32) || signalling(3); link id4725ac1375601d182afec3610f019b25; ECDH agreement true; 64-byte session key569ac51a07fb242f…=hkdf(64, ecdh, salt=link_id).
Resource
- VEC-RES-ADV (computed,
Resource.py:1275-1352): advertisement dict with flags03(compressed+encrypted), packs to 146 bytes. - VEC-RES-PROOF (frozen,
Resource.py:752-753):proof_data = resource_hash || full_hash(data || resource_hash)=d257b38bc5d6aa22…(64 bytes).
Channel and Buffer
- VEC-CHAN-ENVELOPE (frozen):
abcd0007000b6368616e6e656c64617461— typeabcd, sequence0007, length000b, payload"channeldata". - VEC-STREAM-HDR (frozen):
810273747265616d64617461— header8102(eof set over stream id 0x0102), data"streamdata".
Transport
- VEC-PATH-REQUEST (frozen-injection,
Transport.py:2780-2787): payload000102…0f(target 16) || tag(16), in a PLAIN packet (flags08).
Framing and IFAC
- VEC-HDLC (frozen,
TCPInterface.py:44-52):01 7E 02 7D 03frames to7e017d5e027d5d037e. - VEC-IFAC (frozen,
Transport.py:1112-1148): tag2cf485c1dfcea002, masked output9f4a2cf485c1dfcea0…, IFAC header flag set, mask/unmask roundtrip recovers the original packet and tag.
Reticulum reference symbol inventory (frozen)
Frozen ground-truth enumeration of the wire surfaces in the vendored Python RNS reference. The specification is written against it; the Coverage ledger maps every entry to a section and a proof. To refresh, re-enumerate the source at the pinned commit and diff; a new unclassified symbol is a coverage gap.
Pin
| Component | Version | Submodule commit |
|---|---|---|
Reticulum (reference/Reticulum) | RNS 1.3.5 | d5e62d4e15c5fe2e170f7bd9e120551671f21a27 |
Classification key
- N normative: crosses the wire or is observable by a peer; specified exactly and proven.
- I informative: internal behaviour an implementation may diverge on.
- X out of scope: interface drivers beyond framing, daemon/CLI, scaffolding.
Reticulum.py — system constants
| Symbol | Value | Line | Class |
|---|---|---|---|
MTU | 500 | 93 | N |
MDU | 464 | 152 | N |
TRUNCATED_HASHLENGTH | 128 | 145 | N |
HEADER_MINSIZE | 19 | 147 | N |
HEADER_MAXSIZE | 35 | 148 | N |
IFAC_MIN_SIZE | 1 | 149 | N |
IFAC_SALT | 32-byte hex | 150 | N |
Cryptography/*
| File | Symbol / method | Line | Class |
|---|---|---|---|
| Token.py | Token (modified Fernet), TOKEN_OVERHEAD=48, encrypt, decrypt, verify_hmac | 50,77,87,100 | N |
| X25519.py | X25519PrivateKey/PublicKey, generate, exchange | 126,139 | N |
| Ed25519.py | Ed25519PrivateKey/PublicKey, sign, verify | 53,69 | N |
| AES.py | AES_128_CBC, AES_256_CBC (16-byte IV, PKCS7) | — | N |
| HMAC.py | RFC 2104 HMAC-SHA256, new, digest | — | N |
| HKDF.py | hkdf(length, derive_from, salt, context) | 35 | N |
| PKCS7.py | block size 16 padding | — | N |
Identity.py (980 lines) — class Identity
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
KEYSIZE | 512 | 59 | N |
RATCHETSIZE | 256 | 64 | N |
RATCHET_EXPIRY | 2592000 | 69 | I |
TOKEN_OVERHEAD | 48 | 77 | N |
HASHLENGTH | 256 | 80 | N |
SIGLENGTH | 512 | 81 | N |
NAME_HASH_LENGTH | 80 | 83 | N |
TRUNCATED_HASHLENGTH | 128 | 84 | N |
DERIVED_KEY_LENGTH | 64 | 90 | N |
Methods
| Method | Line | Class |
|---|---|---|
full_hash / truncated_hash | 373/383 | N |
get_random_hash | 393 | N |
_get_ratchet_id / _ratchet_public_bytes / _generate_ratchet | 417-425 | N |
validate_announce | 532 | N |
encrypt / decrypt | 827/872 | N |
sign / validate | 931/948 | N |
update_hashes / get_public_key / get_private_key | 808/757/750 | N |
recall / remember | — | I (resolution) |
Destination.py (680 lines) — class Destination
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
SINGLE/GROUP/PLAIN/LINK | 0x00-0x03 | 63-66 | N |
PROVE_NONE/APP/ALL | 0x21-0x23 | 69-71 | N |
IN/OUT | 0x11/0x12 | 79-80 | N |
RATCHET_COUNT | 512 | 85 | I |
RATCHET_INTERVAL | 1800 | 90 | I |
PR_TAG_WINDOW | 30 | 83 | I |
Methods
| Method | Line | Class |
|---|---|---|
expand_name | 96 | N |
hash / hash_from_name_and_identity | 116/141 | N |
announce | 243 | N |
encrypt / decrypt | 585/611 | N |
Packet.py (603 lines) — class Packet, PacketReceipt
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
packet types DATA/ANNOUNCE/LINKREQUEST/PROOF | 0x00-0x03 | 60-63 | N |
header types HEADER_1/HEADER_2 | 0x00/0x01 | 67-68 | N |
context bytes NONE..LRPROOF | 0x00-0xFF | 72-92 | N |
FLAG_SET/FLAG_UNSET | 0x01/0x00 | 95-96 | N |
ENCRYPTED_MDU | 383 | 106 | N |
PLAIN_MDU | =MDU (464) | 110 | N |
PacketReceipt FAILED/SENT/DELIVERED/CULLED | 0,1,2,0xFF | 408-415 | I |
EXPL_LENGTH/IMPL_LENGTH | 96/64 | — | N (proofs) |
Methods
| Method | Line | Class |
|---|---|---|
get_packed_flags | 169 | N |
pack / unpack | 177/242 | N |
get_hashable_part | 355 | N |
validate_proof_packet / validate_proof | 443/498 | N |
Link.py (1538 lines) — class Link
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
ECPUBSIZE | 64 | — | N |
KEYSIZE | 32 | — | N |
LINK_MTU_SIZE | 3 | — | N |
MTU_BYTEMASK | 0x1FFFFF | — | N |
MODE_BYTEMASK | 0xE0 | — | N |
MODE_AES256_CBC | 0x01 | — | N |
states PENDING..CLOSED | 0x00-0x04 | — | I |
KEEPALIVE / STALE_TIME | 360/720 | — | I |
Methods
| Method | Line | Class |
|---|---|---|
link_id_from_lr_packet / set_link_id | 340/349 | N |
handshake | 353 | N |
prove / validate_proof | 371/396 | N |
signalling_bytes / mtu_from_lp_packet / mode_from_lp_packet | 148+ | N |
identify | 459 | N |
send_keepalive | 848 | N |
get_salt / get_context | 643/646 | N |
| watchdog, RTT scheduling, teardown | — | I |
Context-byte payloads (N)
LRPROOF 0xFF (sig(64)+eph_pub(32)+signalling(3)), LRRTT 0xFE (msgpack float),
LINKIDENTIFY 0xFB (pub(32)+sig(64)), LINKCLOSE 0xFC (link_id), KEEPALIVE
0xFA (single byte 0xFF), REQUEST/RESPONSE 0x09/0x0A.
Resource.py (1380 lines) — Resource, ResourceAdvertisement
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
MAPHASH_LEN | 4 | — | N |
RANDOM_HASH_SIZE | 4 | — | N |
HASHMAP_IS_EXHAUSTED/NOT | 0xFF/0x00 | — | N |
WINDOW*, *_TIMEOUT_FACTOR, MAX_RETRIES | various | — | I |
OVERHEAD (advertisement) | 134 | 1235 | N |
status NONE..CORRUPT | 0x00-0x08 | — | I |
Methods
| Method | Line | Class |
|---|---|---|
ResourceAdvertisement.__init__ / pack / unpack | 1278/1333 | N |
hashmap_update / request / receive_part | — | N (wire) / I (scheduling) |
prove / validate_proof | 752/782 | N |
advertise / assemble / window adaptation | 508/672 | I |
Context bytes RESOURCE 0x01, RESOURCE_ADV 0x02, RESOURCE_REQ 0x03,
RESOURCE_HMU 0x04, RESOURCE_PRF 0x05, RESOURCE_ICL 0x06, RESOURCE_RCL
0x07 (Packet.py:73-79) — N.
Channel.py (738 lines) / Buffer.py (371 lines)
| Symbol | Line | Class |
|---|---|---|
Envelope.pack/unpack (>HHH type,seq,len) | 192/179 | N |
StreamDataMessage header (14-bit id, compressed, eof) | Buffer 80-92 | N |
SMT_STREAM_DATA 0xff00, STREAM_ID_MAX 0x3fff, OVERHEAD 8 | — | N |
MessageState, CEType enums | — | I |
| channel sequencing/retransmission | — | I |
Transport.py (3585 lines) — class Transport
Normative surfaces
| Symbol | Line | Class |
|---|---|---|
transmit (IFAC masking) | 1051 | N |
inbound (IFAC unmasking) | 1398 | N |
request_path (path request packet) | 2771 | N |
path_request_handler / path response (PATH_RESPONSE rebroadcast) | 2866/2943 | N |
Informative
PATHFINDER_*, AP_PATH_TIME, ROAMING_PATH_TIME, LOCAL_REBROADCASTS_MAX,
PATH_REQUEST_* (50-83); the path/announce/link/reverse/tunnel tables and their
IDX_* shapes (3547-3586); announce retransmission timing and jitter; dedup;
table culling in jobs (508). All I.
Out of scope (X)
Interface drivers under RNS/Interfaces/* beyond the HDLC/KISS framing
documented in Framing and IFAC; the daemon/CLI tooling;
shared-instance IPC.
LXMF Protocol Specification
This appendix specifies the LXMF messaging protocol. Its canonical wire fixture
is generated from the Python LXMF 1.1.0 reference (reference/LXMF, commit 795fdaa)
running on Reticulum 1.3.5 (commit d5e62d4), and is the compatibility contract
for leviculum-lxmf. The original symbol inventory and much of the source-line
audit were captured against LXMF 0.9.6 (8499729); changed wire surfaces are
called out and tested against the active 1.1.0 lock.
LXMF (Lightweight Extensible Message Format) is the store-and-forward messaging layer of Reticulum. It defines how a message is structured, signed, encrypted, sized, and delivered (opportunistically as a single packet, directly over a link, via a propagation node, or on paper), plus the anti-spam stamp and ticket mechanisms. It moves opaque bytes over Reticulum primitives; it carries no media processing of its own.
The Rust implementation is intentionally client-only for propagation: it can
discover a propagation node, upload a recipient-encrypted message as a raw Link
Packet or Resource, and perform the /get list, download, and acknowledgement
exchange. Propagation-node hosting, transit storage, /offer, and peer
synchronisation are documented only as Python-reference context and are not
implemented by leviculum-lxmf. The current /get path supports canonical
single-segment Link request and response Resources. Packed request or response
values above Reticulum's 1,048,575-byte efficient Resource limit, and incoming
split request/response advertisements, are rejected until semantic segment
reassembly is implemented.
How to read this specification
- Normative statements use RFC 2119 keywords (MUST, SHOULD, MAY) and specify
behaviour an interoperable implementation has to reproduce. Every normative
fact carries either a
file:linecitation into the reference, a derivation with the arithmetic shown, or a labelled test vector[VEC-...]. - Informative sections describe internal reference behaviour (router scheduling, queues, persistence) that an implementation MAY diverge from without breaking wire or semantic compatibility.
- Test vectors are the genuine byte output of the reference, regenerated by
vectors/gen_vectors.pyand pinned to the submodule commits. A citation proves "the code says this"; a vector proves "these are the bytes".
Sections
- Introduction and scope
- Cryptographic primitives
- Identifiers and sizes
- Message binary format
- Fields
- Delivery methods and sizing
- On-air sequences
- Stamps and proof-of-work
- Tickets
- Announce application data
- Propagation
- Router internals (informative)
- Constants reference
- Coverage ledger
- Test vectors
- Python 1.1.0 and RNS 1.4.0 parity status
- Symbol inventory (frozen)
Introduction and scope
Reference
The active wire fixture and Rust compatibility tests use these pinned references:
| Component | Version | Commit |
|---|---|---|
| LXMF | 1.1.0 (_version.py:1) | 795fdaa2b0777c13033787d933d1afc94a2377cb |
| Reticulum (RNS) | 1.3.5 | d5e62d4e15c5fe2e170f7bd9e120551671f21a27 |
APP_NAME is "lxmf" (LXMF.py:1). Where the reference defers to a Reticulum
primitive (hashing, signing, encryption, MDU sizes), this document cites
reference/Reticulum and does not re-specify the primitive; its behaviour-as-used
is pinned by test vectors instead.
The frozen symbol inventory and original source-line audit were captured from
LXMF 0.9.6 (8499729). The canonical fixture is now generated from 1.1.0, and
sections changed since the audit carry current citations and vectors.
Scope
Normative (this document specifies exactly, and proves):
- the message binary layout, hashing input, signing input, and verification;
- the message payload msgpack structure and its type discipline;
- the fields dictionary and its identifiers;
- delivery method selection and the size thresholds that drive it;
- the on-air form of each delivery method (opportunistic, direct, propagated, paper);
- stamp construction, validity, value, and the ticket shortcut;
- announce application-data formats (delivery and propagation);
- the recipient-facing propagation
/getexchange and its error codes; - the Python reference's
/offerand propagation-ingest wire shapes, for protocol documentation only.
Informative (described, not byte-proven; an implementation MAY diverge):
- router job scheduling, outbound/inbound queue management, retry cadences and timeouts;
- on-disk persistence layout (message store, peers, tickets, costs, stats);
- propagation-node peer selection, rotation, and sync scheduling internals.
Out of scope: the lxmd daemon and CLI (Utilities/lxmd.py).
Rust implementation scope
leviculum-lxmf implements propagation only as a client: discovery, outgoing
link establishment, origin uploads as raw Link Packets or Resources, /get
list/download requests, and acknowledgement/purge. It does not implement
propagation-node hosting, transit storage, /offer, propagation peers, peer
rotation, or peer synchronisation. /get requests and responses are currently
bounded to one Reticulum Resource segment; Python-compatible splitting above
that boundary is deferred. References to the unimplemented mechanisms below
describe the Python protocol and do not imply a Rust server surface.
The historical enumeration of reference symbols and their normative / informative / out-of-scope classification is the frozen Symbol inventory. The Coverage ledger maps every normative symbol to a section and a proof; a normative symbol with no mapping is a coverage gap.
Notational conventions
- RFC 2119 keywords (MUST, MUST NOT, SHOULD, MAY) carry their usual meaning and mark normative requirements.
- Citations take the form
(LXMessage.py:364)and refer to the pinned reference file underreference/LXMF/LXMF/unless another path is given. - Byte layouts are shown as offset tables or annotated hex. Concatenation is
written
a || b. A field width in bytes is shown asname(16). - Test vectors are referenced by label, e.g.
[VEC-MSG-1], and are listed in full in Test vectors. They live in machine-readable form invectors/vectors.json. - Hashes are SHA-256 unless stated. Integers in stamp arithmetic are
big-endian (
LXStamper.py:66,76).
Regenerating the vectors
From the repository root:
PYTHONPATH=reference/Reticulum:reference/LXMF \
python3 docs/src/appendix/lxmf/vectors/gen_vectors.py
The harness fixes all identity key material and the message timestamp, runs the
genuine reference code, asserts determinism for every frozen vector, and writes
vectors/vectors.json with the submodule commits recorded in its meta block.
Re-running MUST reproduce the committed file byte for byte. Vectors whose output
depends on ephemeral encryption key material are marked roundtrip and are
proven by a decrypt round trip plus structural assertions rather than by frozen
ciphertext.
Cryptographic primitives
LXMF builds on Reticulum primitives for all hashing, signing, and encryption. An implementation MUST use primitives that produce byte-identical results to these, because their outputs are signed, hashed, and exchanged on the wire.
Hashing
full_hash(x)is SHA-256 overx, 32 bytes (RNS.Identity.HASHLENGTH= 256 bits). LXMF uses it for the message hash and message-id (LXMessage.py:368-369), the transient-id (LXMessage.py:434), and the stamp digest (LXStamper.py:65,75).truncated_hash(x)is the leading 16 bytes offull_hash(x)(RNS.Identity.TRUNCATED_HASHLENGTH= 128 bits). LXMF uses it for the ticket stamp shortcut (LXMessage.py:277,300).
Signing
Identity.sign(m)is Ed25519 overm, producing a 64-byte signature (RNS.Identity.SIGLENGTH= 512 bits). Ed25519 is deterministic (RFC 8032), so a given key and message always yield the same signature;[VEC-MSG-1]pins one.Identity.validate(sig, m)verifies an Ed25519 signature. The inbound path calls it assource.identity.validate(signature, signed_part)(LXMessage.py:809).
Encryption
Destination.encrypt(plaintext)encrypts to a SINGLE destination using Reticulum's ECDH scheme: a fresh ephemeral X25519 key per call, an HKDF-derived AES-128-CBC key, and an HMAC token, optionally keyed by the destination's current ratchet. Because the ephemeral key is fresh per call, the ciphertext is not reproducible across runs; LXMF vectors that involve encryption ([VEC-PROP-ENVELOPE],[VEC-PAPER-URI]) are proven by a decrypt round trip, not by frozen ciphertext.- LXMF calls
encryptfor the propagated and paper forms over the message tailpacked[16:](LXMessage.py:430,449), andDestination.decrypton receipt. - The encryption description strings
"AES-128"/"Curve25519"/"Unencrypted"(LXMessage.py:98-100) are local labels only, not on the wire.
Key derivation
Cryptography.hkdf(length, derive_from, salt, context)is HKDF-SHA-256. LXMF uses it only to build the stamp workblock (LXStamper.py:53-56); see Stamps and proof-of-work.
Identities
An Identity carries an X25519 key pair (encryption) and an Ed25519 key pair
(signing). The reference constructs deterministic identities for the vectors from
fixed 64-byte private material X25519(32) || Ed25519(32)
(gen_vectors.py, recorded in vectors.json meta.src_identity_prv_hex /
meta.dst_identity_prv_hex). An identity's 16-byte hash is truncated_hash of
its concatenated public keys; LXMF treats the hash as opaque and obtains it from
the Reticulum Destination.
Identifiers and sizes
All sizes below are the genuine class attributes of the reference, captured in
vectors.json constants.
Identifiers
| Identifier | Width | Definition | Citation |
|---|---|---|---|
| Destination hash | 16 | Reticulum SINGLE destination hash of lxmf/delivery | LXMessage.py:40 |
| Source hash | 16 | sender's lxmf/delivery destination hash | LXMessage.py:383-384 |
| Signature | 64 | Ed25519 over the signed part | LXMessage.py:41 |
| Message hash / message-id | 32 | `full_hash(dest | |
| Transient-id | 32 | full_hash(lxmf_data) for propagation | LXMessage.py:434 |
| Stamp | 32 | proof-of-work nonce | LXStamper.py:15 |
| Ticket | 16 | shared secret for stamp shortcut | LXMessage.py:42 |
The message hash and message-id are the same value (LXMessage.py:369); this
document uses "message-id". [VEC-MSG-1] shows message_id_hex == hash_hex.
Overhead and packet sizes
The fixed overhead and the three single-packet content limits are derived constants. The derivations (MUST evaluate to these values):
TIMESTAMP_SIZE = 8 (LXMessage.py:60)
STRUCT_OVERHEAD = 8 (LXMessage.py:61)
LXMF_OVERHEAD = 2*16 + 64 + 8 + 8 = 112 (LXMessage.py:62)
ENCRYPTED_PACKET_MDU = RNS.Packet.ENCRYPTED_MDU + 8 = 391 (LXMessage.py:67)
ENCRYPTED_PACKET_MAX_CONTENT = 391 - 112 + 16 = 295 (LXMessage.py:78)
LINK_PACKET_MDU = RNS.Link.MDU = 431 (LXMessage.py:83)
LINK_PACKET_MAX_CONTENT = 431 - 112 = 319 (LXMessage.py:89)
PLAIN_PACKET_MDU = RNS.Packet.PLAIN_MDU = 464 (LXMessage.py:93)
PLAIN_PACKET_MAX_CONTENT = 464 - 112 + 16 = 368 (LXMessage.py:94)
PAPER_MDU = ((2953 - (3 + 3)) * 6) // 8 = 2210 (LXMessage.py:105)
QR_MAX_STORAGE = 2953 (LXMessage.py:105); the + 16 terms restore the
destination hash that is excluded from LXMF_OVERHEAD accounting for the
encrypted single-packet forms. These are the thresholds the delivery-method
selector compares against; see Delivery methods and sizing.
Content size
The reference defines the content size of a packed message as
content_size = len(packed_payload) - TIMESTAMP_SIZE - STRUCT_OVERHEAD
= len(packed_payload) - 16
(LXMessage.py:388). This is the value compared against the limits above. Note
it is computed from the serialized payload length, not from the raw content
bytes, so msgpack framing and the title and fields count toward it.
Message binary format
This section is normative and is proven by [VEC-MSG-1], [VEC-MSG-2], and
[VEC-MSG-3].
Packed layout
A packed LXMF message is the concatenation (LXMessage.py:382-386):
destination_hash(16) || source_hash(16) || signature(64) || packed_payload
offset 0 16 32 96
An implementation MUST produce exactly this layout. packed_payload is the
msgpack serialization of the payload array (below). The total fixed prefix is 96
bytes.
Payload array
The payload is a msgpack array (LXMessage.py:362):
[ timestamp, title, content, fields ]
with an optional fifth element stamp appended when a stamp is generated
(LXMessage.py:371-373); see Stamps.
The msgpack type discipline is normative for a writer and is a common interop trap. It is deliberately not normative for a reader: see Reader tolerance below, which is the other half of the same trap.
| Element | msgpack type | Citation |
|---|---|---|
timestamp | float64 (f64), seconds since the Unix epoch | LXMessage.py:357,362 |
title | binary (bin), not string | LXMessage.py:196-197 |
content | binary (bin), not string | LXMessage.py:202-205 |
fields | map, integer keys (may be empty {}) | LXMessage.py:215-219 |
stamp (optional) | binary (bin), 32 bytes | LXMessage.py:373 |
title and content are stored and packed as bytes; the *_as_string
accessors only decode UTF-8 on demand (LXMessage.py:199,208). An implementation
MUST pack them as msgpack bin, never str. Mismatching this changes the
serialized bytes and therefore the message-id, so a Python peer rejects the
message.
Proof: annotated [VEC-MSG-1] payload
For timestamp = 1700000000.0, title = b"Hi", content = b"Hello",
fields = {}, the packed payload is 94cb41d954fc40000000c4024869c40548656c6c6f80:
94 fixarray, 4 elements
cb 41d954fc40000000 float64 = 1700000000.0 (timestamp)
c4 02 4869 bin8 len 2 = "Hi" (title)
c4 05 48656c6c6f bin8 len 5 = "Hello" (content)
80 fixmap, 0 entries = {} (fields)
The cb (float64), c4 (bin8), and 80 (fixmap) prefixes prove the type
discipline directly. [VEC-MSG-2] shows a non-empty fields map carrying an
integer key.
Reader tolerance
A reader MUST NOT require the writer's type discipline for timestamp. The
reference reads timestamp = unpacked_payload[0] (LXMessage.py:765) with no
type check, so any msgpack number — every unsigned and signed integer
width, float32, float64 — is accepted and reaches the application. Measured on
the pinned reference: uint32, uint64, positive and negative fixints,
int8..int64, float32, and the non-finite float64s all unpack with
signature_validated = true. So do nil, booleans and strings, which no
consumer can use as a time; a reader that carries the timestamp in a numeric
type MAY refuse those, and SHOULD refuse them by a distinguishable error.
LXMF's own writer never exercises this: time.time() is a Python float, which
umsgpack always packs as float64. The tolerance matters for the third-party
writers in a real mesh. [VEC-MSG-FOREIGN-UINT32],
[VEC-MSG-FOREIGN-FLOAT32] and [VEC-MSG-FOREIGN-NEGATIVE-FIXINT] record the
reference decoder's verdict on each form.
title and content are a different case: the reference does not type-check
them either, but set_title_from_bytes (LXMessage.py:196-197) stores
whatever it is handed, and a str-typed title produces a message whose bytes
no writer that follows this specification would have produced. A reader MAY
require bin for those.
Hashing input (message-id)
The message hash is (LXMessage.py:364-369):
hashed_part = destination_hash || source_hash || msgpack(payload_without_stamp)
message_id = full_hash(hashed_part)
The payload hashed here MUST NOT include the optional stamp element.
[VEC-MSG-1] records hashed_part_hex and the resulting message_id_hex.
On unpack the reference takes the hashed payload from one of two places, and
the difference is normative (LXMessage.py:751-762):
- No stamp (
len(unpacked_payload) == 4):packed_payloadis the slice taken from the wire, hashed verbatim. A reader MUST NOT re-encode it. Any encoding the writer chose therefore verifies, which is what makes the reader tolerance above usable rather than decorative. - Stamp present: the reference discards the received bytes and hashes
msgpack.packb(unpacked_payload[:4])(:758) — the decoded values packed again by the reader's own encoder. A writer whose encoding is not whatmsgpack.packbwould emit therefore fails verification on the stamped path even though it passes on the unstamped one. Measured: a uint32-encoded timestamp with a stamp verifies (Python re-packs it as uint32), the same value spelled as uint64 does not, and neither does float32.
The asymmetry is the reference's, not a specification choice, and an implementation MUST reproduce both branches or it will compute a different message-id than the sender for some legal inputs.
Signing input
The signature is (LXMessage.py:375-378):
signed_part = hashed_part || message_id (= dest || src || msgpack(payload) || message_id)
signature = source.sign(signed_part) (Ed25519, 64 bytes)
[VEC-MSG-1] records signed_part_hex, signature_hex, and
signature_valid = true (verified with source.identity.validate).
Unpack and verification
unpack_from_bytes (LXMessage.py:746-822) slices at the fixed offsets:
destination_hash = bytes[0:16], source_hash = bytes[16:32],
signature = bytes[32:96], packed_payload = bytes[96:]. It unpacks the
payload, and if the array has more than four elements treats element [4] as the
stamp and removes it before recomputing the hash (LXMessage.py:754-758); see
Hashing input for which bytes are hashed in each
case.
The reference also caches the received bytes (message.packed = lxmf_bytes,
LXMessage.py:799) and pack() returns them untouched when they are present
(if not self.packed, :355). An implementation that re-serialises an
unpacked message instead of returning what it received will emit bytes the
message's own signature does not cover.
Verification requires the source identity to be known (learned from its
announce). The outcome is one of (LXMessage.py:801-816):
- signature valid:
signature_validated = true; - signature present but invalid:
unverified_reason = SIGNATURE_INVALID (0x02); - source identity unknown:
unverified_reason = SOURCE_UNKNOWN (0x01).
[VEC-MSG-3] makes the source identity recallable, unpacks [VEC-MSG-1], and
records signature_validated = true, matches_source = true, and the recovered
title and content. An implementation MUST reproduce these offsets and the
stamp-stripping rule, or it will compute a different message-id than the sender.
Fields
The fields element of the payload (LXMessage.py:362) is a msgpack map with
integer keys. Keys are the FIELD_* identifiers; values are field-specific. An
empty map {} is valid and is the default. Field keys are packed as msgpack
integers, including the high-value debug keys which serialize as uint8.
Field identifiers (LXMF.py:8-51)
| Key | Name | Value convention |
|---|---|---|
| 0x01 | FIELD_EMBEDDED_LXMS | list of embedded LXM byte strings |
| 0x02 | FIELD_TELEMETRY | telemetry blob |
| 0x03 | FIELD_TELEMETRY_STREAM | telemetry stream blob |
| 0x04 | FIELD_ICON_APPEARANCE | appearance descriptor |
| 0x05 | FIELD_FILE_ATTACHMENTS | list of [name, bytes] |
| 0x06 | FIELD_IMAGE | [format, bytes] |
| 0x07 | FIELD_AUDIO | [audio_mode, bytes] (see audio modes) |
| 0x08 | FIELD_THREAD | thread reference |
| 0x09 | FIELD_COMMANDS | list of commands |
| 0x0A | FIELD_RESULTS | list of results |
| 0x0B | FIELD_GROUP | group metadata |
| 0x0C | FIELD_TICKET | [expires, ticket], see Tickets |
| 0x0D | FIELD_EVENT | event payload |
| 0x0E | FIELD_RNR_REFS | RNR references |
| 0x0F | FIELD_RENDERER | renderer hint (see renderers) |
| 0xFB | FIELD_CUSTOM_TYPE | custom type tag |
| 0xFC | FIELD_CUSTOM_DATA | custom data |
| 0xFD | FIELD_CUSTOM_META | custom metadata |
| 0xFE | FIELD_NON_SPECIFIC | unspecified |
| 0xFF | FIELD_DEBUG | debug payload |
An implementation MUST treat unknown field keys as opaque and preserve them
(the reference round-trips the whole fields map through msgpack). [VEC-MSG-2]
carries {0x0F: 0x02} (FIELD_RENDERER: RENDERER_MARKDOWN) and shows it packed
inside the payload.
Renderers (LXMF.py:99-102)
| Value | Name |
|---|---|
| 0x00 | RENDERER_PLAIN |
| 0x01 | RENDERER_MICRON |
| 0x02 | RENDERER_MARKDOWN |
| 0x03 | RENDERER_BBCODE |
Audio modes (LXMF.py:65-89)
Used as the first element of FIELD_AUDIO. Codec2 modes AM_CODEC2_450PWB
(0x01) through AM_CODEC2_3200 (0x09); Opus modes AM_OPUS_OGG (0x10) through
AM_OPUS_LOSSLESS (0x19); AM_CUSTOM (0xFF). These identify the audio codec and
profile of an attached clip; LXMF does not process the audio, it only carries the
mode byte and the encoded bytes.
Delivery methods and sizing
Method and representation constants
| Method | Value | Citation |
|---|---|---|
OPPORTUNISTIC | 0x01 | LXMessage.py:30 |
DIRECT | 0x02 | LXMessage.py:31 |
PROPAGATED | 0x03 | LXMessage.py:32 |
PAPER | 0x05 | LXMessage.py:33 |
| Representation | Value | Citation |
|---|---|---|
UNKNOWN | 0x00 | LXMessage.py:25 |
PACKET | 0x01 | LXMessage.py:26 |
RESOURCE | 0x02 | LXMessage.py:27 |
The representation records whether the message goes out as a single Reticulum Packet or as a Reticulum Resource (multi-packet transfer). PAPER uses neither.
Selection algorithm
pack() sets method and representation from desired_method and the content
size (LXMessage.py:390-458). The normative rules:
- If no method is desired, default to
DIRECT(LXMessage.py:392-393). - OPPORTUNISTIC is valid only for SINGLE or PLAIN destinations. If the
content size exceeds
ENCRYPTED_PACKET_MAX_CONTENT(295) for a SINGLE destination, the reference falls back toDIRECT(LXMessage.py:397-401). Otherwise representation isPACKET(LXMessage.py:404-415). For PLAIN destinations the limit isPLAIN_PACKET_MAX_CONTENT(368). - DIRECT: if content size
<= LINK_PACKET_MAX_CONTENT(319), representation isPACKET; otherwiseRESOURCE(LXMessage.py:417-424). - PROPAGATED: the message is wrapped into the propagation envelope (see
Propagation); if the envelope size
<= LINK_PACKET_MAX_CONTENTit is aPACKET, otherwise aRESOURCE(LXMessage.py:426-444). - PAPER: if the encrypted paper form
<= PAPER_MDU(2210) the representation is paper; otherwisepack()raises (LXMessage.py:446-458).
content_size is the serialized-payload measure from
Identifiers and sizes (LXMessage.py:388).
An implementation MUST apply the same thresholds so that a message a Python peer would send as a single packet is not sent as a resource (and vice versa), since the on-air framing differs.
On-air forms
| Method | Representation | On-air bytes | Citation |
|---|---|---|---|
| OPPORTUNISTIC | PACKET | packed[16:] (destination hash omitted) | LXMessage.py:634 |
| DIRECT | PACKET | full packed over a Link | LXMessage.py:636 |
| DIRECT | RESOURCE | full packed as a Resource over a Link | LXMessage.py:653-654 |
| PROPAGATED | PACKET/RESOURCE | propagation_packed to a propagation node | LXMessage.py:637-638,655-656 |
| PAPER | (paper) | lxm:// URI or QR | LXMessage.py:698-713 |
The opportunistic packet omits the leading 16-byte destination hash because the
Reticulum packet header already addresses the destination; the receiver
reconstructs the full message by prepending the known destination hash. This is
proven by [VEC-DLV-OPP] (on_air_hex == packed[16:]) and the direct full-bytes
form by [VEC-DLV-DIRECT].
On-air sequences
This section describes the on-air event sequence for each delivery method. The
bytes placed on the wire are normative (see
Delivery methods and sizing); the scheduling
(retry cadence, timeouts, path-request timing) is informative and lives in
Router internals. The relevant cadence constants are
MAX_DELIVERY_ATTEMPTS = 5 (LXMRouter.py:30), DELIVERY_RETRY_WAIT = 10 s
(LXMRouter.py:32), PATH_REQUEST_WAIT = 7 s (LXMRouter.py:33), and
MAX_PATHLESS_TRIES = 1 (LXMRouter.py:34).
Opportunistic
- If there is no path to the destination, request one and wait (informative
cadence). After
MAX_PATHLESS_TRIESthe message may be sent pathless. - Send a single Reticulum Packet whose payload is
packed[16:](the destination hash is omitted;LXMessage.py:634). - The message state becomes
SENT. Delivery is confirmed by a Reticulum proof; on timeout the router re-queues up toMAX_DELIVERY_ATTEMPTS.
No link is established. Suitable only for messages within the single-packet content limit.
Direct
- Ensure a path, then establish a Reticulum
Linkto the destination'slxmf/deliveryendpoint. - When the link is
ACTIVE(LXMessage.py:650):- if representation is
PACKET, send one Packet carrying the fullpackedbytes over the link (LXMessage.py:636); - if representation is
RESOURCE, transferpackedas a Reticulum Resource over the link (LXMessage.py:653-654), with compression negotiated per the peer's advertised support.
- if representation is
- On link failure before delivery, tear down and retry.
Propagated
- Establish a
Linkto the configured outbound propagation node. - Send
propagation_packed(the encrypted envelope, see Propagation) as a Packet or Resource depending on size (LXMessage.py:637-638,655-656). - Success marks the message
SENT(notDELIVERED): final delivery to the recipient happens asynchronously when the recipient syncs from the node.
Paper
No Reticulum transport. pack() produces the encrypted paper form; as_uri()
renders it as an lxm:// URI (LXMessage.py:698-713) or as_qr() as a QR code.
The recipient ingests the URI out of band.
State model (informative)
A message moves through the states GENERATING (0x00) -> OUTBOUND (0x01) -> SENDING (0x02) -> SENT (0x04) -> DELIVERED (0x08), with terminal REJECTED (0xFD), CANCELLED (0xFE), and FAILED (0xFF) (LXMessage.py:15-22). These
are local lifecycle states, not on-wire values, and an implementation MAY model
the lifecycle differently.
Stamps and proof-of-work
Stamps are an anti-spam proof-of-work bound to a message-id (delivery stamps) or
a transient-id (propagation stamps). This section is normative and is proven by
[VEC-STAMP-1] and [VEC-STAMP-PN]. An implementation MUST reproduce the
workblock, validity test, and value computation bit-for-bit, or its stamps will
not be accepted by a Python peer (and vice versa).
Workblock
stamp_workblock(material, expand_rounds):
workblock = b""
for n in range(expand_rounds):
workblock += hkdf(length=256,
derive_from=material,
salt=full_hash(material || msgpack(n)),
context=None)
return workblock
(LXStamper.py:49-60). Each round appends 256 bytes, so the workblock is
expand_rounds * 256 bytes. The salt for round n is
full_hash(material || msgpack(n)), where msgpack(n) is the msgpack encoding of
the integer n (LXStamper.py:55). The expand-round counts are:
| Context | Rounds | Workblock size | Citation |
|---|---|---|---|
| Delivery stamp | WORKBLOCK_EXPAND_ROUNDS = 3000 | 768 000 B | LXStamper.py:12 |
| Propagation stamp | WORKBLOCK_EXPAND_ROUNDS_PN = 1000 | 256 000 B | LXStamper.py:13 |
| Peering key | WORKBLOCK_EXPAND_ROUNDS_PEERING = 25 | 6 400 B | LXStamper.py:14 |
The Python reference holds the entire workblock in RAM. The Rust cooperative executor instead feeds one 256-byte HKDF block at a time into SHA-256, keeping constant workblock workspace while producing the same final digest.
Validity
stamp_valid(stamp, target_cost, workblock):
target = 1 << (256 - target_cost)
return int.from_bytes(full_hash(workblock || stamp), "big") <= target
(LXStamper.py:73-77). The digest is interpreted as a big-endian 256-bit
integer and compared against target. target_cost is the number of required
leading zero bits. The stamp itself is 32 random bytes (STAMP_SIZE,
LXStamper.py:15).
Value
stamp_value(workblock, stamp):
count leading zero bits of full_hash(workblock || stamp) # big-endian
(LXStamper.py:62-71). The value is the achieved number of leading zero bits.
Proof: [VEC-STAMP-1]
For a fixed 32-byte material, expand_rounds = 4, and target_cost = 8, the
harness builds the workblock (1024 bytes = 4 x 256), then deterministically
searches stamp = full_hash(material || counter_be8) over increasing counter
until stamp_valid holds. The vector records the winning counter, the stamp, the
digest, the target (0x0100…00, i.e. 1 << 248, one set bit then 248 zero
bits), valid = true, and stamp_value = 8. The reduced round count keeps the vector cheap to reproduce;
the algorithm it pins is identical to the production path, which differs only
in expand_rounds.
Proof: [VEC-STAMP-PN]
The propagation vector uses the complete 1000-round, 256000-byte logical workblock and pins its hash, one valid cost-8 stamp, and value. This guards the otherwise easy interop error of using the 3000-round delivery workblock for the outer propagation-node stamp.
Generation
generate_stamp(material, stamp_cost, expand_rounds) brute-forces random 32-byte
stamps until one is stamp_valid (generate_stamp, LXStamper.py:123-144). The
reference parallelizes this across processes on Linux and falls back to
single-process elsewhere
(LXStamper.py:178-376); the parallelism is informative, the resulting stamp is
not.
Rust execution model
CooperativeStamper::cooperative is the default: its future yields after a
bounded number of workblock rounds and candidate attempts. Router events carry
owned DeliveryStampRequest, InboundStampRequest, or
PropagationStampRequest values. Calling generate_with() or
validate_with() on one of those values borrows only the stamp executor, not
the router or NodeCore, so packet receive, Link, Resource, and router tasks
remain callable throughout calculation on a single-threaded executor. The
result is attached later with the matching router setter.
For outbound work, use set_outbound_stamp_result() or
set_outbound_propagation_stamp_result() with the same request after awaiting
the worker. These guarded setters reject a result if the advertised cost,
queued message, or encrypted transient changed while work was in flight.
Inbound messages remain queued until set_inbound_stamp_result() receives the
matching InboundStampRequest result. The lower-level stamp setters are kept
for trusted externally generated or restored stamps.
StampExecutor is the override boundary. Host applications MAY implement it
with Rayon, another worker pool, dedicated hardware, or a detached WASM worker;
the returned stamp is attached through the router's delivery-stamp or
propagation-stamp setter. This keeps leviculum-lxmf no_std + alloc and does
not make threads a protocol dependency.
Where stamps are required
- Delivery stamp: the recipient advertises a
stamp_costin its delivery announce (see Announce application data). The sender generates a stamp over the message-id and appends it as payload element[4](LXMessage.py:371-373,320). The recipient validates it withvalidate_stamp(LXMessage.py:273-294). - Propagation stamp: generated over the transient-id with
WORKBLOCK_EXPAND_ROUNDS_PNand the node's advertised cost (LXMessage.py:329-353). - Ticket shortcut: if a valid ticket is held, the stamp is
truncated_hash(ticket || message_id)and the value isCOST_TICKET = 256, bypassing proof-of-work (LXMessage.py:277-280,299-303). See Tickets.
Validation order
validate_stamp(target_cost, tickets) first tries each held inbound ticket: if
stamp == truncated_hash(ticket || message_id) the stamp is accepted with value
COST_TICKET (LXMessage.py:274-280). Otherwise it builds the workblock over the
message-id and runs stamp_valid (LXMessage.py:287-292). An implementation
MUST check tickets before proof-of-work to interoperate with ticketed senders.
Tickets
A ticket is a 16-byte shared secret (TICKET_LENGTH, LXMessage.py:42) that lets
a known correspondent skip proof-of-work. The recipient issues a ticket to a
sender; the sender then derives stamps from it cheaply.
Derivation
A ticketed stamp is (LXMessage.py:300, validated at :274):
stamp = truncated_hash(ticket || message_id)
with value COST_TICKET = 256 (LXMessage.py:53,301). On the receiving side,
validate_stamp accepts the message if stamp equals
truncated_hash(ticket || message_id) for any held inbound ticket
(LXMessage.py:274-280). An implementation MUST use truncated_hash (16 bytes),
matching the stamp width expectation of this path.
Issuing
generate_ticket(destination_hash, expiry) (LXMRouter.py:1073-1100) returns
[expires, ticket] where:
ticket = os.urandom(16)(LXMRouter.py:1096);expires = now + TICKET_EXPIRY(LXMRouter.py:1095).
An existing inbound ticket with more than TICKET_RENEW validity left is reused
rather than reissued (LXMRouter.py:1083-1089), and a new ticket is not issued to
a destination more often than TICKET_INTERVAL (LXMRouter.py:1076-1081).
Exchange
A ticket is delivered to a correspondent inside a message via FIELD_TICKET
(0x0C, LXMF.py:19), carrying the [expires, ticket] pair. The receiver remembers
it as an outbound ticket (remember_ticket, LXMRouter.py:1102-1105) and uses it
for subsequent stamps until it expires (get_outbound_ticket,
LXMRouter.py:1107-1113).
The Rust router follows the same default: enqueue() automatically derives and
attaches the 16-byte delivery stamp whenever TicketStore holds a valid
outbound ticket for the destination. This happens before propagated recipient
encryption, but the ticket stamp remains distinct from the required 32-byte
outer propagation-node stamp. To grant a reply ticket,
issue_ticket_field() returns the bounded/persisted FIELD_TICKET value that
must be passed to Message::create() so it is covered by the signature.
Timing constants
| Constant | Value | Seconds | Citation |
|---|---|---|---|
TICKET_EXPIRY | 21 days | 1 814 400 | LXMessage.py:49 |
TICKET_GRACE | 5 days | 432 000 | LXMessage.py:50 |
TICKET_RENEW | 14 days | 1 209 600 | LXMessage.py:51 |
TICKET_INTERVAL | 1 day | 86 400 | LXMessage.py:52 |
The validity windows are part of the interoperable behaviour: a ticket can
stamp messages until its encoded expiry, while the issuer retains its record
for the additional TICKET_GRACE cleanup window. The exact reuse and reissue
scheduling around them is informative.
Announce application data
LXMF carries application data in Reticulum announces. There are two formats: the
delivery announce (sent by a normal LXMF destination) and the propagation-node
announce. Both are normative and proven by [VEC-ANN-DELIVERY] and
[VEC-ANN-PROPAGATION].
Delivery announce
The delivery announce app_data is (LXMRouter.py:1034-1050):
msgpack([ display_name, stamp_cost, supported_functionality ])
display_name: the UTF-8 encoded display name asbin, orNone(LXMRouter.py:1038-1040).stamp_cost: an integer in(0, 255), orNone(LXMRouter.py:1042-1045).supported_functionality: a list of advertised feature codes. LXMF 1.0.1 emits[SF_COMPRESSION], whereSF_COMPRESSION = 0x00(LXMRouter.py:1047-1048;LXMF.py:140-142).
Format detection
The decoders distinguish this version-0.5.0+ format from the legacy format
(a bare UTF-8 display name) by sniffing the first byte: it is the new format iff
app_data[0] is in 0x90..0x9f (msgpack fixarray) or equals 0xdc (array16)
(LXMF.py:151-200). An implementation MUST emit a msgpack array so this sniff
succeeds; the current three-element array begins with 0x93.
For compatibility with earlier senders, the reference compression decoder
treats a missing or non-list third element as compression support. When the
third element is a list, compression is supported only if that list contains
SF_COMPRESSION (LXMF.py:187-200).
Proof: [VEC-ANN-DELIVERY]
msgpack([b"Alice", 8, [SF_COMPRESSION]]) produces
93c405416c696365089100, with first_byte = 0x93. The genuine decoders recover
display_name = "Alice", stamp_cost = 8, and compression support
(LXMF.py:151-200).
Propagation-node announce
The propagation announce app_data is a 7-element list
(LXMRouter.py:328-336):
msgpack([
legacy_flag, # 0: bool, legacy LXMF PN support
timebase, # 1: int, int(time.time())
propagation_enabled, # 2: bool
per_transfer_limit_kb, # 3: int
per_sync_limit_kb, # 4: int
[prop_cost, prop_flex, peering_cost], # 5: list of three ints
metadata, # 6: dict (PN_META_* keys)
])
Validity
pn_announce_data_is_valid (LXMF.py:224-250) requires: data decodes to a
list of length >= 7; data[1] (timebase), data[3], data[4] are integer-
coercible; data[2] is strictly True or False; data[5] is a list whose
first three elements are integer-coercible; and data[6] is a dict. An
implementation MUST satisfy all of these for its propagation announce to be
accepted.
Metadata map (LXMF.py:128-138)
Keys: PN_META_VERSION (0x00), PN_META_NAME (0x01), PN_META_SYNC_STRATUM
(0x02), PN_META_SYNC_THROTTLE (0x03), PN_META_AUTH_BAND (0x04),
PN_META_UTIL_PRESSURE (0x05), PN_META_CUSTOM (0xFF). The node name is
metadata[PN_META_NAME] as UTF-8 bytes (LXMRouter.py:321).
Proof: [VEC-ANN-PROPAGATION]
The vector builds the 7-element list with metadata = {PN_META_NAME: b"Node"},
stamp_costs = [16, 3, 18], and a fixed timebase, then proves with the genuine
helpers: pn_announce_data_is_valid = true, pn_name_from_app_data = "Node"
(LXMF.py:202-213), and pn_stamp_cost_from_app_data = 16
(LXMF.py:215-222).
Field 1 (timebase) is int(time.time()) in the real protocol and is pinned to a
constant in the vector.
leviculum-lxmf decodes this announce to discover a propagation node. Encoding
is documented here as a wire-format requirement; the Rust crate does not host a
propagation node or emit propagation-node announces.
Propagation
Propagation lets a sender deposit a message at an always-reachable node for an
offline recipient to collect later. The Rust implementation covers both client
directions: origin uploads and the recipient /get list, download, and
acknowledgement exchange. /offer, node ingest, the node-internal store, peer
selection, rotation, and sync scheduling are Python-reference documentation
only and are not implemented by leviculum-lxmf.
Propagation transfer envelope
A propagated message is wrapped as follows (LXMessage.py:426-436):
pn_encrypted_data = destination.encrypt(packed[16:])
lxmf_data = packed[:16] || pn_encrypted_data
transient_id = full_hash(lxmf_data)
if propagation_stamp: lxmf_data || = propagation_stamp # 32 bytes
propagation_packed = msgpack([ wall_clock_timestamp, [ lxmf_data ] ])
Normative points:
- The destination hash (
packed[:16]) stays in cleartext; the rest of the packed message (packed[16:], i.e. source hash, signature, payload) is encrypted to the recipient (LXMessage.py:430,433). transient_id = full_hash(lxmf_data)and is computed before any propagation stamp is appended (LXMessage.py:434-435).- The envelope is
msgpack([timestamp, [lxmf_data, ...]]): a timestamp followed by a list of one or morelxmf_datablobs (LXMessage.py:436). The peer-sync path reuses the same shape with many blobs (LXMPeer.py:466).
Origin upload (implemented)
An originating client sends a singleton envelope directly as Link data. If the
encoded envelope is at most 319 bytes it is a raw Link Packet; otherwise it is
a Resource (LXMessage.py:423-441,483-496,608-614). This is not a request and
does not use /offer.
Before upload, the client appends a separate 32-byte propagation-node stamp to
lxmf_data. Its work material is the pre-stamp transient_id, its target is
the full cost advertised by the selected node, and its workblock uses
WORKBLOCK_EXPAND_ROUNDS_PN = 1000. A 16-byte delivery-ticket stamp, when one
is available for the recipient, remains inside the signed clear message before
recipient encryption; it never replaces the outer propagation stamp.
The client serialises one upload per propagation Link. Packet proof or
sender-side Resource completion changes the message to SENT. Packet timeout,
Resource failure, or Link closure returns it to the bounded retry queue; Packet
and Resource failures tear down the Link before retry, matching Python. Link
establishment is charged as the logical attempt while submission on that Link
is not charged again. The node signal msgpack([0xF5]) rejects the message.
Ciphertext, transient ID, envelope timebase, advertised target cost, and
generated outer stamp are checkpointed so restoration and retries do not
re-encrypt or change the bytes. Process-local Link state and monotonic
deadlines are not reusable after a restart; restored entries become
immediately due while retaining their durable retry count and prepared bytes.
The router emits an owned PropagationStampRequest; its calculation borrows
neither the router nor NodeCore. The default PoW executor yields
cooperatively while expanding and searching the workblock, so Link and receive
events remain serviceable on a single-threaded runtime. Applications can
supply another StampExecutor (for example a Rayon-backed pool); builds
without the pow feature can attach a detached 32-byte stamp through the
router API.
Proof: [VEC-PROP-ENVELOPE]
Because destination.encrypt uses a fresh ephemeral key per call, the ciphertext
is not reproducible; the vector is a round-trip proof. It records deterministic
lengths and the cleartext dest_hash_prefix, checks the structure
destination_hash(16) || destination.encrypt(packed[16:]), checks that the
derived transient ID is full_hash(lxmf_data), and proves
destination.decrypt(pn_encrypted) == packed[16:]
(decrypt_recovers_inner_tail = true). The random ciphertext and transient-ID
bytes are deliberately omitted so the canonical fixture remains byte-stable.
An implementation MUST reproduce the framing and transient-ID derivation.
Propagation-node announce
See Announce application data for the 7-element node announce that advertises the node's limits and stamp costs.
/offer (Python propagation peers only; not implemented)
A Python propagation peer requests OFFER_REQUEST_PATH = "/offer"
(LXMPeer.py:14) over a
Link with the payload (LXMPeer.py:385,389):
offer = [ peering_key, [ transient_id, ... ] ]
where peering_key is the party's proof-of-work peering key (see below) and the
list is the transient-ids it offers. The node replies via offer_response
(LXMPeer.py:400); the reply is one of: False (node already has all),
True (node wants all), or a list (the subset the node wants). The wanted
messages are then pushed as one Resource carrying
msgpack([timestamp, [lxmf_data, ...]]) (LXMPeer.py:466-468).
leviculum-lxmf exposes no /offer request path, handler, peering-key engine,
or peer-sync state machine.
/get (collect from a node)
A recipient requests MESSAGE_GET_PATH = "/get" (LXMPeer.py:15) with the
payload (LXMRouter.py:1482-1504):
[ want, have ]
- if both
wantandhaveareNone, the node returns a list of the recipient's availabletransient_ids, sorted by size (LXMRouter.py:1491-1504); - otherwise
havelists transient-ids the client already holds (so the node can drop them) andwantlists the ones to send (LXMRouter.py:1506-1556).
The exact list, download, acknowledgement, list-response, and download-response
bytes are pinned by VEC-PROP-GET-LIST, VEC-PROP-GET-DOWNLOAD,
VEC-PROP-GET-ACK, VEC-PROP-LIST-RESPONSE, and
VEC-PROP-GET-RESPONSE in the test vectors.
Current Rust Resource boundary
Python passes the request timeout into an oversized /get request Resource,
then starts a fresh response deadline after the upload proof. The Rust client
matches those two timeout phases and the canonical request (q/u) and response
(q/p) advertisement flags.
The current Core correlation path represents a request or response Resource as
one semantic transfer. The fully packed Link request or response must currently
fit one Reticulum Resource segment: at most RESOURCE_MAX_EFFICIENT_SIZE =
1,048,575 bytes. An oversized outgoing request returns ResourceTooLarge
before LXMF stores its correlation, and an incoming split request/response
advertisement is ignored. Generic LXMF message Resources still use Core's
ordinary multi-segment transfer path; this limit applies specifically to Link
request/response Resources.
The router's default /get delivery transfer limit is 1000 KB (1,024,000
bytes), below the single-segment ceiling. Full Python-compatible semantic
reassembly for larger request/response Resources remains deferred.
Error codes (LXMPeer.py:24-31)
Returned by node request handlers or, for ERROR_INVALID_STAMP, as upload Link
signalling:
| Code | Name | Meaning |
|---|---|---|
| 0xF0 | ERROR_NO_IDENTITY | requester did not identify on the link |
| 0xF1 | ERROR_NO_ACCESS | requester not allowed |
| 0xF3 | ERROR_INVALID_KEY | invalid peering key |
| 0xF4 | ERROR_INVALID_DATA | malformed request |
| 0xF5 | ERROR_INVALID_STAMP | propagation stamp invalid |
| 0xF6 | ERROR_THROTTLED | rate limited (PN_STAMP_THROTTLE = 180 s) |
| 0xFD | ERROR_NOT_FOUND | requested message not found |
| 0xFE | ERROR_TIMEOUT | request timed out |
Peering key (Python reference only; not implemented)
A peering key is a proof-of-work over peer_identity_hash || node_identity_hash
with WORKBLOCK_EXPAND_ROUNDS_PEERING = 25 rounds against the node's advertised
peering_cost (LXMPeer.py:242-265; validated by validate_peering_key,
LXStamper.py:79-82). It authorizes a party to offer messages to the node.
Node transient ingest and expiry (Python reference only; not implemented)
A node stores each accepted message keyed by transient_id, validates its
propagation stamp in batches (LXStamper.py:118-121), and expires entries after
MESSAGE_EXPIRY = 30 days (LXMRouter.py:38). The on-disk layout and peer
bookkeeping are informative.
Router internals (informative)
Everything in this section is informative. It documents how the reference router behaves so an implementation can match observable timing and limits where useful, but an implementation MAY diverge from any of it without breaking wire or semantic compatibility. The normative obligations are in the preceding sections.
Delivery scheduling (LXMRouter.py:30-91)
| Constant | Value | Meaning |
|---|---|---|
MAX_DELIVERY_ATTEMPTS | 5 | retries before a message fails |
PROCESSING_INTERVAL | 4 s | jobloop tick |
DELIVERY_RETRY_WAIT | 10 s | wait between delivery attempts |
PATH_REQUEST_WAIT | 7 s | wait after a path request |
MAX_PATHLESS_TRIES | 1 | sends attempted before forcing a path request |
LINK_MAX_INACTIVITY | 600 s | idle link teardown |
P_LINK_MAX_INACTIVITY | 180 s | idle propagation link teardown |
Expiry and limits
| Constant | Value | Meaning |
|---|---|---|
MESSAGE_EXPIRY | 30 days | propagation store retention |
STAMP_COST_EXPIRY | 45 days | cached outbound stamp cost retention |
PROPAGATION_LIMIT | 256 KB | per-transfer propagation limit |
SYNC_LIMIT | 256*40 KB | per-sync cumulative limit |
DELIVERY_LIMIT | 1000 KB | direct delivery resource limit |
PN_STAMP_THROTTLE | 180 s | propagation stamp throttle window |
Peering (LXMRouter.py:43-63)
MAX_PEERS = 20, AUTOPEER = true, AUTOPEER_MAXDEPTH = 4 hops,
ROTATION_HEADROOM_PCT = 10, ROTATION_AR_MAX = 0.5, PEERING_COST = 18
(max 26), PROPAGATION_COST = 16 (min 13, flex 3). Peer selection and rotation
policy are internal.
Jobloop cadence (LXMRouter.py:871-879)
A single jobloop dispatches staggered jobs: outbound processing (1 s), deferred
stamp generation (1 s), link cleanup (1 s), transient-cache cleanup (60 s),
message-store cleanup (120 s), peer sync (6 s), peer rotation (56 * peer-ingest
interval). An implementation may use any scheduler.
Persistence
The Python reference persists to a writable storage path: the message store
(one file per propagation message, named
transient_id_timestamp_stampvalue), peers, available_tickets,
outbound_stamp_costs, locally_delivered_transient_ids,
locally_processed_transient_ids, and node_stats. The on-disk encoding
(mostly msgpack) and file naming are implementation choices; only the messages
and announces these structures drive onto the wire are normative.
leviculum-lxmf checkpoints only its client/router state through
LxmfStorage: outbound direct/opportunistic messages, delivered and processed
IDs (including downloaded transient IDs), stamp costs, tickets, ignored
destinations, and the next job deadline. It has no propagation message store,
peer records, node statistics, or node-hosting persistence.
Python peer sync state machine (not implemented; LXMPeer.py:17-22)
A propagation peer progresses IDLE -> LINK_ESTABLISHING -> LINK_READY -> REQUEST_SENT -> RESPONSE_RECEIVED -> RESOURCE_TRANSFERRING -> IDLE, with
STRATEGY_LAZY/STRATEGY_PERSISTENT (default persistent) controlling whether a
peer keeps syncing while unhandled messages remain. The states are local; the
wire payloads they produce (/offer, the sync Resource) are documented for
Python-reference completeness (Propagation); they are not
part of the Rust client implementation.
Constants reference
Every constant in the Symbol inventory, grouped, with value and
citation. Values are the genuine reference values; the derived sizes are captured
in vectors.json constants.
Application (LXMF.py)
| Constant | Value | Line |
|---|---|---|
APP_NAME | "lxmf" | 1 |
SF_COMPRESSION | 0x00 | 108 |
Message states (LXMessage.py:15-22) — informative
GENERATING 0x00, OUTBOUND 0x01, SENDING 0x02, SENT 0x04, DELIVERED
0x08, REJECTED 0xFD, CANCELLED 0xFE, FAILED 0xFF.
Representations and methods (LXMessage.py:25-33)
UNKNOWN 0x00, PACKET 0x01, RESOURCE 0x02. OPPORTUNISTIC 0x01, DIRECT
0x02, PROPAGATED 0x03, PAPER 0x05.
Unverified reasons (LXMessage.py:36-37)
SOURCE_UNKNOWN 0x01, SIGNATURE_INVALID 0x02.
Sizes (LXMessage.py:40-106)
| Constant | Value | Line |
|---|---|---|
DESTINATION_LENGTH | 16 | 39 |
SIGNATURE_LENGTH | 64 | 40 |
TICKET_LENGTH | 16 | 41 |
TIMESTAMP_SIZE | 8 | 60 |
STRUCT_OVERHEAD | 8 | 61 |
LXMF_OVERHEAD | 112 | 62 |
ENCRYPTED_PACKET_MDU | 391 | 67 |
ENCRYPTED_PACKET_MAX_CONTENT | 295 | 78 |
LINK_PACKET_MDU | 431 | 83 |
LINK_PACKET_MAX_CONTENT | 319 | 89 |
PLAIN_PACKET_MDU | 464 | 93 |
PLAIN_PACKET_MAX_CONTENT | 368 | 94 |
QR_MAX_STORAGE | 2953 | 104 |
PAPER_MDU | 2210 | 105 |
URI_SCHEMA | "lxm" | 102 |
Tickets (LXMessage.py:49-53)
| Constant | Value | Seconds | Line |
|---|---|---|---|
TICKET_EXPIRY | 21 days | 1 814 400 | 48 |
TICKET_GRACE | 5 days | 432 000 | 49 |
TICKET_RENEW | 14 days | 1 209 600 | 50 |
TICKET_INTERVAL | 1 day | 86 400 | 51 |
COST_TICKET | 0x100 (256) | — | 52 |
Fields (LXMF.py:8-51)
FIELD_EMBEDDED_LXMS 0x01 … FIELD_RENDERER 0x0F; FIELD_CUSTOM_TYPE 0xFB,
FIELD_CUSTOM_DATA 0xFC, FIELD_CUSTOM_META 0xFD; FIELD_NON_SPECIFIC 0xFE,
FIELD_DEBUG 0xFF. See Fields for the full table.
Renderers and audio modes (LXMF.py:65-102)
RENDERER_PLAIN 0x00 … RENDERER_BBCODE 0x03. AM_CODEC2_* 0x01–0x09,
AM_OPUS_* 0x10–0x19, AM_CUSTOM 0xFF.
Propagation metadata (LXMF.py:132-138)
PN_META_VERSION 0x00, PN_META_NAME 0x01, PN_META_SYNC_STRATUM 0x02,
PN_META_SYNC_THROTTLE 0x03, PN_META_AUTH_BAND 0x04, PN_META_UTIL_PRESSURE
0x05, PN_META_CUSTOM 0xFF.
Stamps (LXStamper.py:12-16)
| Constant | Value | Line |
|---|---|---|
WORKBLOCK_EXPAND_ROUNDS | 3000 | 10 |
WORKBLOCK_EXPAND_ROUNDS_PN | 1000 | 11 |
WORKBLOCK_EXPAND_ROUNDS_PEERING | 25 | 12 |
STAMP_SIZE | 32 | 13 |
PN_VALIDATION_POOL_MIN_SIZE | 256 | 14 |
Python propagation peer (LXMPeer.py:14-50; Rust server out of scope)
Paths OFFER_REQUEST_PATH "/offer" (14), MESSAGE_GET_PATH "/get" (15). States
IDLE 0x00 … RESOURCE_TRANSFERRING 0x05 (17-22). Errors ERROR_NO_IDENTITY
0xF0, ERROR_NO_ACCESS 0xF1, ERROR_INVALID_KEY 0xF3, ERROR_INVALID_DATA
0xF4, ERROR_INVALID_STAMP 0xF5, ERROR_THROTTLED 0xF6, ERROR_NOT_FOUND 0xFD,
ERROR_TIMEOUT 0xFE (24-31). STRATEGY_LAZY 0x01, STRATEGY_PERSISTENT 0x02
(33-34). MAX_UNREACHABLE 14 days, SYNC_BACKOFF_STEP 12 min,
PATH_REQUEST_GRACE 7.5 s (39-50).
leviculum-lxmf implements origin-client uploads, the /get mailbox-client
path, and response errors. It does not expose /offer, these peer-sync
states/strategies, or a propagation server. Its default /get transfer limit
is 1000 KB, which stays below Core's current 1,048,575-byte single-segment
request/response Resource ceiling.
Router (LXMRouter.py:30-91) — informative
See Router internals for the full set
(MAX_DELIVERY_ATTEMPTS, MESSAGE_EXPIRY, PROPAGATION_COST, etc.).
Coverage ledger
This is the traceability matrix from the frozen Symbol inventory to the specification. Every normative (N) symbol maps to a section and a proof. Every informative (I) and out-of-scope (X) symbol carries a reason. A normative symbol with no section and proof is a coverage gap.
This ledger preserves the original LXMF 0.9.6 source-line audit. Active Rust
compatibility is locked to the 1.1.0 canonical fixture. Entries for /offer,
propagation-node ingest, and LXMPeer describe the Python protocol only;
leviculum-lxmf implements origin uploads and /get mailbox collection, not
the server or peer state machine. The /get implementation is interoperable
within one request/response Resource segment; Python's larger split form is a
documented implementation gap rather than claimed coverage.
Post-audit implementation and API differences introduced through LXMF 1.1.0
are tracked separately in the parity status.
Proof column: vector (a [VEC-...]), computed (derivation shown), quoted
(byte/enum value cited verbatim), n/a (informative/out-of-scope).
LXMessage.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
representation UNKNOWN/PACKET/RESOURCE | 24-26 | N | 05 | quoted |
method OPPORTUNISTIC/DIRECT/PROPAGATED/PAPER | 29-32 | N | 05 | quoted |
unverified SOURCE_UNKNOWN/SIGNATURE_INVALID | 35-36 | N | 03 | quoted |
size constants DESTINATION_LENGTH … PLAIN_PACKET_MAX_CONTENT | 39-94 | N | 02 | computed (vector constants) |
ticket constants TICKET_EXPIRY/GRACE/RENEW/INTERVAL, COST_TICKET | 48-52 | N | 08 | quoted (vector constants) |
URI_SCHEMA, QR_MAX_STORAGE, PAPER_MDU | 102-105 | N | 02, 05 | computed |
ENCRYPTION_DESCRIPTION_* | 97-99 | I | — | n/a (local labels) |
QR_ERROR_CORRECTION | 103 | I | — | n/a (QR rendering) |
| state constants | 14-21 | I | 06 | n/a (local lifecycle) |
set_title/content_*, *_as_string | 190-205 | N | 03 | vector VEC-MSG-1/3 |
set_fields/get_fields | 212-218 | N | 03, 04 | vectors VEC-MSG-2/VEC-MSG-NEGATIVE-FIELD |
validate_stamp | 270 | N | 07 | vector VEC-STAMP-1 |
get_stamp | 293 | N | 07 | vector VEC-MSG-TICKET (ticket branch) |
get_propagation_stamp | 326 | N | 07, 10 | vector VEC-STAMP-PN |
pack | 352 | N | 03, 05 | vector VEC-MSG-1/2 |
__as_packet | 623 | N | 05, 06 | vector VEC-DLV-OPP/DIRECT |
__as_resource | 637 | N | 05, 06 | quoted |
packed_container, write_to_directory, unpack_from_file | 657-810 | I/N | 11 | n/a (local storage) |
as_uri | 687 | N | 05, 10 | vector VEC-PAPER-URI |
as_qr | 707 | I | — | n/a (QR rendering) |
unpack_from_bytes | 735 | N | 03 | vector VEC-MSG-3 |
send, __mark_*, __resource_concluded, timers | 460-620 | I | 06, 11 | n/a (orchestration) |
| msgpack sites 364/378/433/669/741/747 | — | N | 03, 10 | vector |
time.time() 354/433 | — | N | 03, 10 | vector (timestamp fields) |
LXMF.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
APP_NAME | 1 | N | 00 | quoted |
FIELD_* (all) | 8-41 | N | 04 | quoted (VEC-MSG-2 for one) |
AM_* audio modes | 55-79 | N | 04 | quoted |
RENDERER_* | 89-92 | N | 04 | quoted |
PN_META_* | 98-104 | N | 09 | quoted (VEC-ANN-PROPAGATION) |
SF_COMPRESSION | 108 | N | 12 | quoted |
display_name_from_app_data, stamp_cost_from_app_data | 117-152 | N | 09 | vector VEC-ANN-DELIVERY |
compression_support_from_app_data | 154 | N | 09 | quoted |
pn_name_from_app_data, pn_stamp_cost_from_app_data, pn_announce_data_is_valid | 169-217 | N | 09 | vector VEC-ANN-PROPAGATION |
LXStamper.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
WORKBLOCK_EXPAND_ROUNDS*, STAMP_SIZE | 10-13 | N | 07 | vectors VEC-STAMP-1/VEC-STAMP-PN |
stamp_workblock | 18 | N | 07 | vector VEC-STAMP-1 |
stamp_value | 31 | N | 07 | vector VEC-STAMP-1 |
stamp_valid | 42 | N | 07 | vector VEC-STAMP-1 |
validate_peering_key | 48 | N | 10 | quoted |
validate_pn_stamp | 53 | N | 10 | quoted |
generate_stamp | 92 | N | 07 | quoted |
validate_pn_stamps*, job_*, cancel_work, PN_VALIDATION_POOL_MIN_SIZE | 67-354 | I | 07 | n/a (PoW parallelism) |
Handlers.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
LXMFDeliveryAnnounceHandler.received_announce | 9-32 | N | 09 | vector VEC-ANN-DELIVERY |
LXMFPropagationAnnounceHandler.received_announce | 35-72 | N+I | 09, 11 | vector / n/a (auto-peer) |
LXMRouter.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
get_propagation_node_announce_metadata, get_propagation_node_app_data | 302-319 | N | 09 | vector VEC-ANN-PROPAGATION |
get_announce_app_data | 986 | N | 09 | vector VEC-ANN-DELIVERY |
generate_ticket, remember_ticket, get_outbound_ticket* | 1025-1086 | N | 08 | vector VEC-MSG-TICKET (FIELD_TICKET value) |
message_get_request/list_response/get_response | 1427-1591 | N | 10 | quoted |
offer_request, propagation_packet, propagation_resource_concluded, lxmf_propagation, ingest_lxm_uri | 2110-2392 | N | 10 | quoted |
lxmf_delivery | 1732 | N | 06 | quoted |
PR_* states | 62-77 | N | 10 | quoted |
request paths STATS/SYNC/UNPEER | 81-83 | I | 11 | n/a |
| delivery/expiry/peer/job constants | 30-83, 853-860 | I | 11 | n/a (scheduling) |
persistence, queues, jobloop, rotation, time.time() | various | I | 11 | n/a (internal) |
LXMPeer.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
OFFER_REQUEST_PATH, MESSAGE_GET_PATH | 14-15 | N | 10 | quoted |
ERROR_* | 24-31 | N | 10 | quoted |
generate_peering_key | 242 | N | 10 | quoted |
sync (offer payload), offer_response, resource_concluded | 267-520 | N | 10 | quoted |
state constants, strategy, timing, from_bytes/to_bytes, peer counts | 17-50, 52-175, 544-640 | I | 11 | n/a (internal/persistence) |
| msgpack 462 (sync resource) | — | N | 10 | quoted |
_version.py / Utilities/lxmd.py
| Symbol(s) | file:line | Class | Section | Proof |
|---|---|---|---|---|
__version__ | _version.py:1 | N | 00 | quoted (pin) |
| daemon/CLI | lxmd.py | X | — | n/a (not protocol) |
Result
Every normative symbol in the inventory maps to a section and a proof above; no
normative row is left with an empty section or n/a proof. Informative and
out-of-scope rows are reasoned. Coverage is therefore complete against the frozen
inventory at the pinned commit. To re-audit after a reference bump, re-enumerate
the source and diff the inventory; any new symbol appears here unclassified.
Test vectors
These are the golden vectors that prove the binary claims in this specification.
They are the genuine output of the reference code, generated by
vectors/gen_vectors.py and stored in machine-readable
form in vectors/vectors.json.
Regenerate and verify (from the repository root):
PYTHONPATH=reference/Reticulum:reference/LXMF \
python3 docs/src/appendix/lxmf/vectors/gen_vectors.py
Pinned to LXMF 795fdaa (1.1.0) and Reticulum d5e62d4 (RNS 1.3.5). Fixed
inputs: source identity private = 00010203…3f, destination identity private =
404142…7f, message timestamp = 1700000000.0.
Vector kinds
- frozen — deterministic; the hex is the proof; reproduces byte for byte.
- roundtrip — encryption uses ephemeral key material. The fixture omits the random ciphertext and ciphertext-derived hashes, retaining deterministic lengths, structure, and decrypt assertions. The complete JSON therefore reproduces byte for byte.
VEC-MSG-1 (frozen) — minimal opportunistic message
title=b"Hi", content=b"Hello", fields={}. (LXMessage.py:355-388)
packed (118 B):
cf0b2a4a8d2a0b6978b71290da7cc80e destination_hash(16)
fae321c442e3c9bdcd7a3e79d850e03c source_hash(16)
fb321978105a4c709c3b86930ff15a9d
7b53b3485517ec19e2083b39f7661e6e
531c78fb71d932f0baf13794c42234ab
9320f1ab5b7688e93eaf5960810ece00 signature(64)
94cb41d954fc40000000c4024869c405
48656c6c6f80 packed_payload (msgpack)
message_id = 9aec506b63deab21d8fa4954d9f743cf20f5adeeb1abd1c7429bb3f832dc287b
signature_valid = true
VEC-MSG-2 (frozen) — fields dict with integer key
title=b"", content=b"body text", fields={0x0F: 0x02}. (LXMessage.py:355)
packed_payload: 94cb41d954fc40000000c400c409626f64792074657874810f02
94 array(4)
cb 41d954fc40000000 timestamp 1700000000.0
c4 00 title bin(0) = ""
c4 09 626f6479... content bin(9) = "body text"
81 0f 02 fields {0x0F: 0x02}
VEC-MSG-NEGATIVE-FIELD (frozen) — negative-fixint field key
fields={-1: b"negative field"} proves that unknown signed integer keys retain
their MessagePack type. The 0xff byte below is negative fixint -1, not an
unsigned field ID or a string key. (LXMessage.py:215-219,355-387)
packed_payload:
94cb41d954fc40000000c40c6e65676174697665206b6579c404626f6479
81ffc40e6e65676174697665206669656c64
^^ key = -1
VEC-MSG-TICKET (frozen) — ticket-stamped message, five-element payload
The only deterministic stamped message: an outbound ticket yields
truncated_hash(ticket || message_id) instead of a mined proof of work, and the
carried FIELD_TICKET value is the [expires, ticket] list generate_ticket
returns. outbound_ticket = 101112…1f, carried ticket secret 202122…2f,
expires = 1700000000.0 + TICKET_EXPIRY.
(LXMessage.py:355-388,293-302; LXMRouter.py:1096-1100,1770-1772)
packed_payload:
95cb41d954fc40000000c40154c4087469636b65746564
810c92cb41d95be820000000c410202122232425262728292a2b2c2d2e2f
c41054f9096872188db445566721224cca74
95 array(5) <- stamp present
cb 41d954fc40000000 timestamp 1700000000.0
c4 01 54 title bin(1) = "T"
c4 08 7469636b65746564 content bin(8) = "ticketed"
81 0c 92 … fields {FIELD_TICKET: [expires, ticket]}
92 array(2)
cb 41d95be820000000 expires 1701814400.0
c4 10 202122…2f ticket bin(16)
c4 10 54f909…a74 stamp bin(16) = truncated_hash(ticket || message_id)
message_id = 190d522017c991b7e74bd958faadfc780dced409514904e6f7fd514b4a84d974
stamp_value = 256 (COST_TICKET)
The message ID and the signature are computed over the four-element payload; the stamp is appended afterwards and is covered by neither.
VEC-MSG-3 (frozen) — unpack + verify of VEC-MSG-1
After making the source identity recallable, unpack_from_bytes of VEC-MSG-1
yields signature_validated = true, matches_source = true, recovered
title="Hi", content="Hello". (LXMessage.py:747-822)
VEC-MSG-FOREIGN-* (frozen) — reader tolerance for payload[0]
Three payloads built with payload[0] spliced in as a non-float64 msgpack
number — uint32, float32 and a negative fixint — and signed over the bytes
the writer packed. Each vector records the reference decoder's own verdict:
reference_signature_validated = true and reference_message_id_matches = true in all three cases, with reference_timestamp_type showing int or
float as packed. They exist because LXMF's writer cannot produce these forms
(time.time() is always a Python float) while its reader accepts them, so
nothing generated from the reference's writer covers the case. See
Reader tolerance.
(LXMessage.py:747-766)
VEC-DLV-OPP (frozen) — opportunistic on-air payload
on_air = packed[16:] (destination hash omitted). (LXMessage.py:633-634)
VEC-DLV-DIRECT (frozen) — direct on-air payload
on_air = packed (full bytes over a link). (LXMessage.py:635-636)
VEC-PROP-ENVELOPE (roundtrip) — propagation envelope
Structure destination_hash(16) || destination.encrypt(packed[16:]); envelope
msgpack([timestamp, [lxmf_data]]); transient_id = full_hash(lxmf_data);
destination.decrypt(pn_encrypted) == packed[16:] holds. The random transient
ID itself is intentionally not stored. (LXMessage.py:426-436)
VEC-PAPER-URI (roundtrip) — paper URI
lxm://base64url(destination_hash(16) || destination.encrypt(packed[16:])) with
= padding stripped; prefix lxm://; decoding and decrypting recovers
packed[16:]. Random URI body bytes are intentionally not stored.
(LXMessage.py:446-451,698-713)
VEC-STAMP-1 (frozen) — stamp workblock, validity, value
material = full_hash(b"lxmf-spec-stamp-material") = 1c91877ffb9797aa6f33064586b47a3c41f6dfa75e10aa17bc24bf0ac6833712,
expand_rounds=4, target_cost=8, workblock 1024 B. Deterministic search
stamp = full_hash(material || counter_be8) finds counter=377:
stamp = 9b79689af899049accea13624a3c59221603117e81086a86a3249ce278acc35e
target = 0100000000000000000000000000000000000000000000000000000000000000 (1 << 248)
valid = true
value = 8
(LXStamper.py:49-77)
VEC-STAMP-PN (frozen) — propagation-node stamp
This vector pins the full propagation-node workblock rather than only the
common stamp algorithm: material 00..1f, expand_rounds=1000, and
target_cost=8 produce a 256000-byte workblock whose SHA-256 is
36224067c2ebd1a40b0dc8908bd34dc4b7cd84be57f68a765dbaf3298bcd9719.
Searching counter_be32 finds counter 17:
stamp = 0000000000000000000000000000000000000000000000000000000000000011
digest = 00420981fc866c09cb2f28e7503147d56bae0b3a6fc69f00478cb247f8d95b91
valid = true
value = 9
(LXStamper.py:13,53-63,122-155)
Propagation mailbox /get vectors (frozen)
These are the complete client-only exchange at /get
(LXMPeer.MESSAGE_GET_PATH): initial list, selected download,
acknowledgement/purge, and both response shapes. The implementation does not
expose /offer.
VEC-PROP-GET-LIST
[None, None] (LXMRouter.py:511-520):
request: 92c0c0
VEC-PROP-GET-DOWNLOAD
One wanted ID (20..3f), one already-held ID (40..5f), transfer limit 1000
KB (LXMRouter.py:1577-1595):
request:
9391c420202122232425262728292a2b2c2d2e2f303132333435363738393a3b3c3d3e3f
91c420404142434445464748494a4b4c4d4e4f505152535455565758595a5b5c5d5e5f
cd03e8
VEC-PROP-GET-ACK
[None, [transient_id]], which asks the node to purge the acknowledged entry
(LXMRouter.py:1625-1637):
request: 92c091c420000102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f
VEC-PROP-LIST-RESPONSE
response:
92c420202122232425262728292a2b2c2d2e2f303132333435363738393a3b3c3d3e3f
c420404142434445464748494a4b4c4d4e4f505152535455565758595a5b5c5d5e5f
VEC-PROP-GET-RESPONSE
Two binary downloaded entries, b"one" and b"two"; error responses are
ccf0 (NO_IDENTITY) and ccf1 (NO_ACCESS).
response: 92c4036f6e65c40374776f
VEC-ANN-DELIVERY (frozen) — delivery announce app_data
LXMF 1.0.1 advertises Resource compression in the third element:
msgpack([b"Alice", 8, [SF_COMPRESSION]]).
app_data: 93c405416c696365089100
93 array(3)
c4 05 416c696365 display_name bin(5) = "Alice"
08 stamp_cost = 8
91 00 supported functionality = [SF_COMPRESSION]
first_byte = 0x93 (new-format sniff)
decoded: display_name="Alice", stamp_cost=8, compression_supported=true
(LXMRouter.py:1034-1050; LXMF.py:151-200)
VEC-ANN-PROPAGATION (frozen) — propagation node announce app_data
7-element list, metadata={PN_META_NAME: b"Node"}, stamp_costs=[16,3,18],
fixed timebase:
app_data: 97c2ce6553f100c3cd0100cd2800931003128101c4044e6f6465
valid = true; pn_name = "Node"; pn_stamp_cost = 16
(LXMRouter.py:324-336; LXMF.py:202-250)
Python LXMF 1.1.0 and RNS 1.4.0 parity status
This page records implementation parity separately from the wire-format
coverage ledger. The active Python LXMF reference is release 1.1.0 at commit
795fdaa2b0777c13033787d933d1afc94a2377cb. Its package metadata requires
RNS 1.4.0 or newer. Leviculum's pinned Reticulum fixture is still the
RNS 1.3.5-based commit d5e62d4e15c5fe2e170f7bd9e120551671f21a27,
so passing the LXMF vectors does not by itself constitute an RNS 1.4.0 parity
claim.
The 1.1.0 vector regeneration changed only the recorded LXMF version and commit. The existing message, stamp, ticket, delivery-announce, propagation announce, upload, and mailbox wire bytes remained identical.
LXMF 1.1.0 client parity
| Python 1.1.0 behaviour | Leviculum status | Notes |
|---|---|---|
| Propagation announces accept numeric transfer and sync limits | Implemented | leviculum-lxmf accepts unsigned integers and finite, non-negative, integral MessagePack floats, then normalises both to u64. This matches Python's int(...) conversion, including the 1.1.0 announce observed with 1024.0. Fractional, negative, non-finite, and out-of-range values remain rejected. |
propagation_transfer_size records a /get response Resource size | Implemented | Incoming request progress carries transfer_size through PropagationTransportEvent; the mailbox runtime retains it, and WASM returns it as transferSize from lxmfPropagationStatus(). It resets with a new or cancelled sync and remains available with the completed result. |
inbound_count() and inbound_resources() expose active incoming delivery Resources | Implemented, read-only | LxmfNode tracks accepted/started Resources by Resource hash and removes them on completion, failure, or Link close. Rust exposes count and iterator APIs. WASM exposes lxmfInboundResources() snapshots containing linkId, resourceHash, transferSize, dataSize, and progress. |
cancel_inbound(resource_hash) and cancel_all_inbound() | Blocked on core | Tracking is deliberately read-only. Clean cancellation needs a core API that locates an already accepted incoming Resource, transitions it exactly once, emits/sends the Reticulum Resource cancellation (RESOURCE_RCL), and produces the normal terminal event. Merely dropping the LXMF tracking entry would leave the Reticulum transfer active. |
| Thread locks and atomic Python file replacement | Equivalent by architecture | The Rust router is single-owner mutable state. Its bounded checkpoint is handed to the embedding storage transactionally, so Python's thread/file implementation details are not copied literally. |
Configurable lxmd inbound delivery stamp cost, with daemon default 12 | Policy-supported | leviculum-lxmf accepts a configured inbound cost and advertises/enforces it. It does not implement the lxmd daemon or inherit its CLI/config default. An embedding application may deliberately choose a different default. |
Python propagation-node and peer functionality
leviculum-lxmf remains a propagation client, not a propagation-node
server. The following Python 1.1.0 changes therefore remain intentionally
unimplemented:
- corrected minimum accepted
/offerstamp cost (max(0, cost - flexibility)instead ofmin(...)); - completing peer synchronisation without sending an empty offer;
- accepted-offer Link accounting and the new offer state values;
- sequential propagation-stamp validation, optional treatment of static peers, and the maximum concurrent inbound-sync limit;
- propagation-node hosting, peer selection, transit storage,
/offer, stats, rotation, and synchronisation generally.
These are not client interoperability blockers. They become requirements if
Leviculum adds propagation-node hosting. At that point they belong in a
separate server/peer state machine with bounded queues and tests, not in
leviculum-core.
The existing client also deliberately limits a semantic Link request or response Resource to one efficient Resource segment. Python-compatible reassembly of split request/response Resources above 1,048,575 bytes remains a layer-level gap. It is independent of the LXMF 1.1.0 wire changes above.
Required leviculum-core work for the RNS 1.4.0 baseline
The items below are the core changes identified by the RNS 1.4.0 compatibility audit that affect LXMF or the security and lifecycle of primitives it uses. They must be completed before claiming the LXMF 1.1.0-required RNS baseline. This list is scoped to Leviculum's implemented packet, announce, Link, request, and Resource surfaces; a release claim still requires advancing the vendored RNS reference and running a complete differential audit.
1. Raise the discovery stamp default to 16
leviculum-core/src/discovery/stamp.rs still defines
DEFAULT_STAMP_VALUE = 14; the RNS 1.4.0 baseline uses 16.
Required work:
- change the default and its documentation;
- regenerate discovery announce vectors and update tests that assume cost 14;
- verify acceptance at the threshold and rejection below it;
- verify encrypted and plaintext discovery announces;
- confirm any application-configured overrides remain explicit.
Receiving a higher-cost RNS 1.4.0 discovery announce already works. The interoperability problem is generation: a local default-cost-14 discovery announce may be rejected by a peer enforcing the new default.
2. Bound and serialise discovery announce validation
RNS 1.4.0 adds bounded valid/invalid validation caches and serialises expensive discovery validation. Leviculum validates correctly but does not yet reproduce those resource controls.
Required work:
- add bounded caches keyed by the validation input/hash, with explicit capacities and deterministic eviction;
- cache both successful and failed validation without caching partially parsed or unauthenticated state;
- ensure only one expensive validation for the same input can be in flight;
- keep the
no_stdcore single-threaded and expose cooperative work if the validation cannot be completed within the caller's budget; - add flood, duplicate, eviction, malformed-input, and persistence-boundary tests.
This is primarily CPU and memory denial-of-service hardening, not a wire-format change.
3. Make Link identification one-time
The current LINKIDENTIFY handler replaces the Link's remote identity and emits
LinkIdentified every time a valid identify packet is received. RNS 1.4.0 only
sets the remote identity while it is unknown.
Required work:
- ignore or explicitly reject subsequent LINKIDENTIFY attempts after the first accepted identity;
- never replace an established remote identity;
- emit the identification event exactly once;
- preserve the existing blackhole check before accepting the first identity;
- add tests for duplicate-same-identity and conflicting-identity attempts.
This is a state-integrity and application-trust boundary, so it should be fixed before exposing identified-Link metadata as durable identity evidence.
4. Add active incoming Resource cancellation
The incoming Resource object has a private cancel() transition and core
already understands the RESOURCE_RCL context, but NodeCore cannot cancel an
already accepted inbound Resource by hash.
Required work:
- add a public
NodeCoreoperation scoped by Link ID and Resource hash; - locate only an active receiver-side Resource and make cancellation idempotent;
- send the correct Resource cancellation packet to the peer;
- remove the active receiver state and emit one
ResourceFailed { error: Cancelled, is_sender: false }; - define behaviour for unknown, completed, sender-side, and wrong-Link hashes;
- add loss, duplicate-cancel, Link-close, and simultaneous full-duplex Resource tests.
After this exists, leviculum-lxmf can implement Python-compatible
cancel_inbound() and cancel_all_inbound() without leaking transport state.
5. Align Link keepalive activity timing
Leviculum currently schedules proactive initiator keepalives primarily from its last keepalive timestamp. RNS 1.4.0 bases the idle interval on general outbound Link activity and keepalive echo activity. The current behaviour is interoperable but can transmit unnecessary keepalives.
Required work:
- track last outbound Link activity independently of the keepalive counter;
- reset the idle deadline for ordinary outbound Link traffic and appropriate keepalive echoes;
- retain the RTT-derived interval, stale detection, and configured override;
- test sustained one-way traffic, idle initiator/responder pairs, delayed echoes, and stale recovery.
This is lower priority than the stamp, identification, validation-cache, and Resource-cancellation items because it does not change message correctness.
Completion criteria
RNS 1.4.0 parity should only be marked complete after all required core items are implemented, the Reticulum submodule is advanced to an immutable 1.4.0 reference, affected golden vectors are regenerated, and differential tests cover malformed, duplicated, delayed, and concurrent inputs. Until then, Leviculum should describe itself as LXMF 1.1.0 wire-compatible on its implemented client surfaces, with the RNS core baseline still in progress.
LXMF reference symbol inventory (frozen)
This file is the historical, frozen enumeration of every symbol and wire
surface in the Python LXMF 0.9.6 audit baseline. The coverage ledger
(Coverage ledger) maps every entry here to a
specification section and a proof. The active implementation and canonical
fixture are pinned separately to LXMF 1.1.0 (795fdaa).
Do not edit by hand to reflect wishful coverage. To refresh, re-enumerate the source at the pinned commit and diff. A new symbol that appears unclassified is a coverage gap.
The inventory is broader than the Rust implementation. In particular,
/offer, propagation-node hosting/storage, and LXMPeer synchronisation are
reference-only; leviculum-lxmf implements origin uploads and the recipient
/get client.
Pin
| Component | Version | Submodule commit |
|---|---|---|
LXMF (reference/LXMF) | 0.9.6 | 8499729024a4cddfceb47ca07188bb5b1d11d179 |
Reticulum (reference/Reticulum) | RNS 1.3.5 | d5e62d4e15c5fe2e170f7bd9e120551671f21a27 |
Reference files under reference/LXMF/LXMF/: LXMessage.py, LXMF.py,
LXMRouter.py, LXMPeer.py, LXStamper.py, Handlers.py, _version.py,
Utilities/lxmd.py.
Classification key
- N normative: crosses the wire or is observable by a Python peer; must be specified exactly and proven.
- I informative: internal behaviour an implementer may diverge on without breaking interop; described, not byte-proven.
- X out of scope: daemon, CLI, build or test scaffolding.
LXMessage.py (827 lines) — class LXMessage (line 13)
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
state GENERATING/OUTBOUND/SENDING/SENT/DELIVERED/REJECTED/CANCELLED/FAILED | 0x00,0x01,0x02,0x04,0x08,0xFD,0xFE,0xFF | 14-21 | I |
representation UNKNOWN/PACKET/RESOURCE | 0x00,0x01,0x02 | 24-26 | N |
method OPPORTUNISTIC/DIRECT/PROPAGATED/PAPER | 0x01,0x02,0x03,0x05 | 29-32 | N |
unverified SOURCE_UNKNOWN/SIGNATURE_INVALID | 0x01,0x02 | 35-36 | N |
DESTINATION_LENGTH | 16 | 39 | N |
SIGNATURE_LENGTH | 64 | 40 | N |
TICKET_LENGTH | 16 | 41 | N |
TICKET_EXPIRY/GRACE/RENEW/INTERVAL | 21d/5d/14d/1d | 48-51 | N |
COST_TICKET | 0x100 | 52 | N |
TIMESTAMP_SIZE | 8 | 60 | N |
STRUCT_OVERHEAD | 8 | 61 | N |
LXMF_OVERHEAD | 112 | 62 | N |
ENCRYPTED_PACKET_MDU | derived | 67 | N |
ENCRYPTED_PACKET_MAX_CONTENT | 295 | 78 | N |
LINK_PACKET_MDU | RNS.Link.MDU | 83 | N |
LINK_PACKET_MAX_CONTENT | 319 | 89 | N |
PLAIN_PACKET_MDU | RNS.Packet.PLAIN_MDU | 93 | N |
PLAIN_PACKET_MAX_CONTENT | 368 | 94 | N |
ENCRYPTION_DESCRIPTION_AES/EC/UNENCRYPTED | strings | 97-99 | I |
URI_SCHEMA | "lxm" | 102 | N |
QR_ERROR_CORRECTION | "ERROR_CORRECT_L" | 103 | I |
QR_MAX_STORAGE | 2953 | 104 | N |
PAPER_MDU | 2210 | 105 | N |
Methods (wire-relevant marked N)
| Method | Line | Class |
|---|---|---|
__init__ | 113 | N (field defaults) |
set_title_from_string/bytes, title_as_string | 190-196 | N |
set_content_from_string/bytes, content_as_string | 199-205 | N |
set_fields, get_fields | 212-218 | N |
validate_stamp | 270 | N |
get_stamp | 293 | N |
get_propagation_stamp | 326 | N |
pack | 352 | N |
send | 460 | I |
determine_compression_support | 507 | N |
determine_transport_encryption | 517 | I |
__mark_delivered/propagated/paper_generated | 558-582 | I |
__resource_concluded, __propagation_resource_concluded | 594-605 | I |
__link_packet_timed_out, __update_transfer_progress | 613-620 | I |
__as_packet | 623 | N |
__as_resource | 637 | N |
packed_container | 657 | N |
write_to_directory | 672 | I |
as_uri | 687 | N |
as_qr | 707 | I |
unpack_from_bytes (static) | 735 | N |
unpack_from_file (static) | 810 | I |
msgpack sites
364 (pack payload, N), 378 (pack payload, N), 433 (propagation envelope, N), 669 (packed_container, N), 741 (unpack payload, N), 747 (re-pack for hash, N), 812 (unpack from file, I).
RNS primitives
Identity.full_hash 365/431, Identity.truncated_hash 274, source.sign 375,
identity.validate 794, Destination.encrypt 427/446, Destination.SINGLE/PLAIN/LINK/GROUP
395-548, Packet 476, Link.ACTIVE 647, Resource 651/653, Identity.recall 759/765,
size constants Identity.TRUNCATED_HASHLENGTH/SIGLENGTH 39/40, Packet.ENCRYPTED_MDU/PLAIN_MDU
67/93, Link.MDU 83.
Filesystem / time
open/write 677-679 (write_to_directory, I). time.time() 354, 433 (N: payload
and envelope timestamps).
LXMF.py (217 lines) — module
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
APP_NAME | "lxmf" | 1 | N |
FIELD_EMBEDDED_LXMS..FIELD_RENDERER | 0x01..0x0F | 8-22 | N |
FIELD_CUSTOM_TYPE/DATA/META | 0xFB-0xFD | 34-36 | N |
FIELD_NON_SPECIFIC/DEBUG | 0xFE-0xFF | 40-41 | N |
AM_CODEC2_* | 0x01..0x09 | 55-63 | N |
AM_OPUS_* | 0x10..0x19 | 66-75 | N |
AM_CUSTOM | 0xFF | 79 | N |
RENDERER_PLAIN/MICRON/MARKDOWN/BBCODE | 0x00-0x03 | 89-92 | N |
PN_META_VERSION..PN_META_CUSTOM | 0x00..0xFF | 98-104 | N |
SF_COMPRESSION | 0x00 | 108 | N |
Functions
| Function | Line | Class |
|---|---|---|
display_name_from_app_data | 117 | N |
stamp_cost_from_app_data | 141 | N |
compression_support_from_app_data | 154 | N |
pn_name_from_app_data | 169 | N |
pn_stamp_cost_from_app_data | 182 | N |
pn_announce_data_is_valid | 191 | N |
msgpack sites
123, 146, 159, 173, 186, 194 (all unpack of announce app_data, N).
LXStamper.py (396 lines) — module
Constants
| Symbol | Value | Line | Class |
|---|---|---|---|
WORKBLOCK_EXPAND_ROUNDS | 3000 | 10 | N |
WORKBLOCK_EXPAND_ROUNDS_PN | 1000 | 11 | N |
WORKBLOCK_EXPAND_ROUNDS_PEERING | 25 | 12 | N |
STAMP_SIZE | 32 | 13 | N |
PN_VALIDATION_POOL_MIN_SIZE | 256 | 14 | I |
Functions
| Function | Line | Class |
|---|---|---|
stamp_workblock | 18 | N |
stamp_value | 31 | N |
stamp_valid | 42 | N |
validate_peering_key | 48 | N |
validate_pn_stamp | 53 | N |
validate_pn_stamps_job_simple/multip, validate_pn_stamps | 67-87 | I (parallelism) |
generate_stamp | 92 | N (algorithm) |
cancel_work | 113 | I |
job_simple/linux/android | 145-260 | I (platform PoW workers) |
msgpack sites
24 (salt packb(n), N). RNS: Cryptography.hkdf 22, Identity.full_hash 24/34/44.
Handlers.py (92 lines)
| Class / method | Line | Class |
|---|---|---|
LXMFDeliveryAnnounceHandler.received_announce | 9/15 | N (announce parsing) |
LXMFPropagationAnnounceHandler.received_announce | 35/41 | N + I (auto-peer is I) |
msgpack: 46 (unpack announce, N). RNS: Transport.hops_to 71/72 (I).
LXMRouter.py (2733 lines) — class LXMRouter (line 29)
Mostly I (router internals: jobloop, queues, persistence, peer rotation, retry cadences). The N surfaces are the announce/app-data builders and the propagation request handlers and packers.
Normative surfaces
| Symbol | Line | Class |
|---|---|---|
get_propagation_node_announce_metadata | 302 | N |
get_propagation_node_app_data | 307 | N |
get_announce_app_data | 986 | N |
generate_ticket | 1025 | N |
message_get_request | 1427 | N (/get request shape) |
message_list_response | 1507 | N |
message_get_response | 1552 | N |
offer_request | 2142 | N (/offer handler) |
propagation_packet | 2110 | N |
propagation_resource_concluded | 2194 | N |
lxmf_propagation | 2310 | N (transient ingest) |
ingest_lxm_uri | 2370 | N (paper ingest) |
lxmf_delivery | 1732 | N (inbound delivery dispatch) |
Informative constants (selected; full set in Router internals)
MAX_DELIVERY_ATTEMPTS=5(30), PROCESSING_INTERVAL=4(31), DELIVERY_RETRY_WAIT=10(32),
PATH_REQUEST_WAIT=7(33), MESSAGE_EXPIRY=30d(38), STAMP_COST_EXPIRY=45d(39),
MAX_PEERS=20(43), AUTOPEER_MAXDEPTH=4(45), PEERING_COST=18(50), PROPAGATION_COST=16(54),
PROPAGATION_LIMIT=256(55), SYNC_LIMIT=256*40(56), DELIVERY_LIMIT=1000(57),
PN_STAMP_THROTTLE=180(60), PR_* states 62-77 (N: appear in /get FSM signalling),
request paths STATS_GET/SYNC_REQUEST/UNPEER_REQUEST 81-83 (I), JOB_* intervals 853-860 (I).
Persistence / time
Extensive filesystem use (message store, peers, tickets, costs, stats) — all I.
80+ time.time() calls — I except where a value is signed or hashed.
LXMPeer.py (642 lines) — class LXMPeer (line 13)
| Symbol | Line | Class |
|---|---|---|
OFFER_REQUEST_PATH="/offer", MESSAGE_GET_PATH="/get" | 14-15 | N |
state IDLE..RESOURCE_TRANSFERRING | 17-22 | I |
ERROR_NO_IDENTITY..ERROR_TIMEOUT | 24-31 | N (/offer response codes) |
STRATEGY_LAZY/PERSISTENT, DEFAULT_SYNC_STRATEGY | 33-35 | I |
MAX_UNREACHABLE=14d, SYNC_BACKOFF_STEP=12m, PATH_REQUEST_GRACE=7.5 | 39-50 | I |
from_bytes/to_bytes (peer persistence) | 52/138 | I |
generate_peering_key | 242 | N |
sync | 267 | I + N (offer payload shape) |
offer_response | 396 | N |
resource_concluded | 488 | N (sync resource payload) |
msgpack: 54/172 (peer persistence, I), 462 (sync resource packb([time, lxm_list]), N).
_version.py (2 lines)
__version__ = "0.9.6" (1) — N (reference pin).
Utilities/lxmd.py (1127 lines)
Daemon and CLI. X out of scope (not protocol). Listed for completeness only.
Deferred RNS primitives (cited into reference/Reticulum, not re-specified)
RNS.Identity.full_hash (SHA-256), truncated_hash, Ed25519 sign/validate,
RNS.Destination.encrypt/decrypt (ECDH + AES token + ratchets),
RNS.Cryptography.hkdf, and the size constants RNS.Identity.TRUNCATED_HASHLENGTH
(128), SIGLENGTH (512), HASHLENGTH (256), RNS.Packet.ENCRYPTED_MDU/PLAIN_MDU,
RNS.Link.MDU. Their exact behaviour-as-used is pinned by the test vectors.