The Core Lock Budget

The async driver owns exactly one NodeCore behind one mutex (leviculum-std/src/driver/mod.rs:1321). Every packet the node decrypts, routes, forwards or emits passes through it. It is the narrowest point in the stack, and the rule that follows from that is:

No caller holds the core lock across CPU-heavy work. Work that scales with payload size runs off the lock, between a cheap capture and a cheap commit.

This is not a style preference. It was measured.

The measurement that set the rule

NodeCore::send_resource used to run its whole build — bz2 compress, bulk token encrypt, full/map hashing — inside the mutex. For a 1 MiB incompressible payload with compression on, that was a 141 ms hold in release. On a 20-link inbound flood it cost a 32 % inbound throughput stall while round-sized sends ran; ~0 % after the fix (Codeberg #152).

The fix is the three-phase shape, and it is the house idiom for anything with the same profile:

phaselockcode
NodeCore::resource_send_paramsbriefleviculum-core/src/node/mod.rs:1417
resource::prepare_resource_sendnoneleviculum-core/src/resource/outgoing.rs:89
NodeCore::commit_resource_sendbriefleviculum-core/src/node/mod.rs:1452

Commit re-validates what could have changed while the build ran unlocked: link gone, a transfer raced in, or the link re-keyed (#66) — the last returns the retryable ResourceError::LinkStateChanged and the caller rebuilds once. The std driver calls the three phases itself (leviculum-std/src/driver/mod.rs:3602).

NodeCore::send_resource still exists as the composed single call (leviculum-core/src/node/mod.rs:1543) because no_std and FFI callers have no lock to hold and no second thread to starve. It is the composed form that is dangerous behind the driver, not the code it composes.

The numbers, restated for callers

Measured with leviculum-core/compression on (as leviculum-std builds it), comparing the composed call against the locked portion of the phased path.

The conditions, because a timing without them is an anecdote: profile release, target x86_64-unknown-linux-musl, an AMD Ryzen 9 7950X. Each cell is the median of five timed runs after one discarded warm-up, and each run builds a freshly linked pair — a second Resource to the same link cannot build at all (ResourceError::TransferInProgress), so re-using one would not repeat the measurement. Both harnesses are #[ignore]d and print every number below: measure_send_lock_costs (leviculum-lxmf/src/node.rs:2607) for the send tables, measure_deferred_tick_costs (leviculum-lxmf/tests/direct_delivery_attempts.rs:1560) for the tick table.

Every column names the bytes it was given, because the cost being reported is a compressor's and a compressor's cost is a property of its input. The three classes are generated by incompressible (leviculum-lxmf/tests/common/payloads.rs:32), compressible (leviculum-lxmf/tests/common/payloads.rs:79) and degenerate (leviculum-lxmf/tests/common/payloads.rs:104), and both harnesses below take them from the same array so neither can drift onto different bytes under the same name:

classwhat it isbz2 at 1 MiB
incompressibleseeded xorshift bytes1.00x — grows 0.5 %
compressibledictionary words with sentence breaks9.0x
degeneratevec![0x5a; n]21 845x
payloadincompressiblecompressibledegeneratephased (locked)
16 KiB1.67 ms0.91 ms0.14 ms2 µs
256 KiB15.8 ms11.0 ms1.70 ms3 µs
1 MiB (segment 1)64.6 ms48.7 ms7.61 ms2–11 µs

The first three columns are the composed call under the lock; the last is the locked half of the phased path, which does not vary with the payload because no payload passes through it.

Incompressible data is the worse case, not compressible data. An earlier revision of this page said the opposite — "bz2 does more work when it succeeds" — and the measurement above refuses it at every size: compressible text costs 0.54x to 0.75x of incompressible bytes of the same length. The mechanism is that the Burrows-Wheeler sort is paid in full either way, and a run that succeeds then has less output left to code, not more. Plan for the incompressible number: it is both the larger one and the one an attachment actually hits, since anything already compressed looks incompressible to bz2.

Without the compression feature — the embedded default — the 1 MiB build drops to 6.4 ms, which is why the same code is tolerable on an nRF52 and intolerable behind the driver. That figure is from the #152 pass and was not re-measured here.

The segment boundary sits under the 1 MiB row

RESOURCE_MAX_EFFICIENT_SIZE is 1 048 575 bytes (leviculum-core/src/resource/mod.rs:69) and the split is decided on the packed, uncompressed length (leviculum-core/src/resource/outgoing.rs:106), so a 1 MiB body is above it in every payload class — the compression ratio does not move the boundary. The 1 MiB row therefore measures segment 1 of a two-segment transfer, not a whole one. Segments 2..N are built on the receive path, under the caller's lock, and the phased path does not cover them.

That also bounds who can reach the row at all. Python LXMF refuses an incoming delivery Resource larger than DELIVERY_LIMIT × 1000, which is 1 000 000 bytes (delivery_resource_advertised, reference/LXMF/LXMF/LXMRouter.py:1977) — below the segment boundary. So no LXMF transfer that a Python peer would accept is ever a two-segment one, and the segment path is reached only by a Rust-to-Rust transfer or a non-LXMF core Resource user.

What the payload class was worth, as a number

This matters beyond bookkeeping, because a table measured on degenerate was once proposed as a correction to this page. On the same machine and profile, degenerate reads 8.5x below incompressible at 1 MiB and 12x below it at 16 KiB: bz2's run-length front end collapses a single repeated byte before the Burrows-Wheeler transform ever runs, so the number that comes out is the cost of compressing almost nothing.

Two other candidate explanations for a table reading low were tested and do not hold.

A missing warm-up is not one. The harness prints the run it discards next to the median it keeps, so this is checkable rather than assumed: every cell above 1 ms has its cold run within 2 % of its median, and the largest gap anywhere is 11 % — on degenerate at 16 KiB, the cheapest cell in the table at 154 µs cold against 139 µs. An n=1 harness on this path is imprecise; it is not biased low.

A build-profile difference is not one either. This run reproduces the figures the page carried before it — 1.7, 16.4 and 65.2 ms — to within 4 % on the incompressible column, so those were release-profile numbers taken on comparable hardware, and the column they belong to is incompressible.

Costs that do not justify phasing, measured the same way: packing a 1 MiB LXMF message is 0.8 ms, and unpacking one with signature verification is 3.2 ms. Inbound verification cannot be phased away — the bytes are already in hand — and at that magnitude it does not need to be.

The adapter-side number: LxmfRouter::tick

Measured during the #196 design pass, on the same machine and profile: one LxmfRouter::tick with 8 due 256 KiB messages holds 126.6 ms in a single uninterrupted borrow. It is the composed-send_resource cost of the table above, multiplied by a queue depth an adapter reaches routinely — the router builds each due message in turn, and nothing between them yields.

That figure is consistent with the table above: eight times the incompressible 256 KiB build is 127 ms.

LxmfRouter can hand the build out instead, under RouterConfig::defer_resource_builds (leviculum-lxmf/src/router.rs:124). What the tick then costs, for one due message, measured the same way:

payloaddeferred tickcomposed tick (incompressible)
16 KiB13.6 µs1.63 ms
256 KiB300 µs15.9 ms
1 MiB1.11 ms66.6 ms

The deferred column is flat across all three payload classes — 13.6 to 14.6 µs at 16 KiB, 282 to 300 µs at 256 KiB — which is the mechanism showing through: a deferring tick copies bytes and does not compress them, so the class it was handed cannot matter. What is left in it is the Message clone, and that is why the deferred column still grows with size at all.

One due message rather than the eight above, because a second Resource to the same link is refused before it builds, so an eight-message comparison would not be comparing like with like. The composed column carries the router's own tick work on top of the build, which is why it reads a little above the send table at the same size; the gap is under 3 % at every row.

It is recorded here rather than in the issue because this page is where a caller looks for it, and because a number that lives only in a tracker cannot be cited from the tree: PROCESSOR_TICK_BUDGET (leviculum-std/src/driver/processor.rs:181) is set against this measurement, and until it was written down the only number behind a public constant could not be traced at all.

What this binds

Any protocol adapter layered on the core. leviculum-lxmf is the current instance and leviculum-lxst will be the next. An adapter that offers only a monolithic submit call forces its host either to hold the lock for the build or to fork the driver. Adapters that are expected to run behind the async driver expose the phase split; the composed form stays for the embedded caller.

Anything the driver runs inside its event loop. The loop's dispatch_output (leviculum-std/src/driver/mod.rs:5247) routes actions to interfaces and forwards events. Work done there blocks not just the lock but interface I/O dispatch — strictly worse than the mutex case. The in-loop /status responder (leviculum-std/src/driver/remote_mgmt.rs:82) is the reference for how much is acceptable there: take the lock, build a small bundle, hand back a TickOutput, return.

Two things the loop's callees may never do. They may not .await, and they may not call back into the driver's public async API: those methods end in action_dispatch_tx.send(output).await on a bounded channel that the same loop drains, so a full channel deadlocks the node.

Those are one rule, not two, and knowing which way round matters when you have to enforce it. The second is a consequence of the first: an async fn called and not awaited builds a future and drops it, sends nothing and blocks nothing. The deadlock needs the bounded-channel send to complete, and only .await can complete it. So a callee expressed as a synchronous fn has both prohibitions closed at once, which is what the in-driver core processor (#196) is built on — see leviculum-std/src/driver/processor.rs. It follows that a runtime guard on the async API would be the wrong shape: there is nothing to guard until an .await that cannot be written.

The residue is re-entrancy, not the async API

Both prohibitions above are special cases of a plainer one, and stating them first got the emphasis wrong for two commits. The loop calls its callees with the core mutex held, and that mutex is a non-reentrant std::sync::Mutex. Any path from a callee back to it hangs the node immediately — first call, no load required.

The async route is one such path and not the instructive one. Consider the block_on case the previous wording named as the whole residue: a callee that smuggles a PacketSender and blocks on send does not reach the bounded channel at all, because PacketSender::send (leviculum-std/src/driver/sender.rs:92-107) takes the core lock in a block and releases it before its .await. It deadlocks one line earlier, on the mutex.

And no block_on is needed. The number is 58, and it is a number rather than a phrase on purpose: scripts/check-core-lock-census.py rebuilds the list of public methods that lock the core out of the sources on every just fast and pins it, name by name, in a checked-in file (TOTAL, scripts/core-lock-census.txt:30). Two earlier revisions of this page estimated the size of this set in words and were low by nearly half — which is what an estimate nothing can check is worth.

53 of them are on ReticulumNode — plain synchronous pub fns that open by locking the core, of which has_path (leviculum-std/src/driver/mod.rs:3085) is self.inner.lock_recover().has_path(dest_hash) and entirely typical. The other five are on PacketSender and LinkHandle, which matters more than the count suggests: those are the two handles a callee is most likely to have been handed in the first place.

A callee holding an Arc<ReticulumNode> deadlocks the node on its first invocation, in ordinary safe synchronous code, with no .await, no channel, and nothing a compile-fail fixture can catch.

So the rule for anything the loop calls is: hold no handle to the node you run inside. The &mut StdNodeCore the seam hands out locks nothing and is the whole intended surface. Everything else belongs on the far side of a channel.

What holds this up is not the type system. It is that registration happens on the builder, before the node exists, so a callee cannot be constructed holding a node handle — injecting one afterwards takes a deliberate OnceLock or Weak and a reference cycle. That is a construction-order barrier, and it is why the hazard stays theoretical in practice. It is not a guarantee, and this page should not be read as offering one.

The one call the seam hands out that this page forbids

NodeCore::send_resource is pub (leviculum-core/src/node/mod.rs:1747) and therefore reachable on the &mut StdNodeCore a processor hook holds. It is the 141 ms composed call this page opens with — one line, in consumer code, behind the driver and under the lock. PROCESSOR_TICK_BUDGET reports it 141 ms after the fact and cannot prevent it, and no fixture can refuse it: it compiles, because for the no_std and FFI callers it is the correct API.

A hook that has to send a resource uses the three-phase form the driver itself uses — resource_send_params, prepare_resource_send off the lock, commit_resource_send — or hands the send to the application side of a channel. This is named again in leviculum-std/src/driver/processor.rs where a consumer will meet it.

Work that is already off the core by construction

Proof-of-work stamp generation borrows only its executor, never the router and never the core (leviculum-lxmf/src/router/stamp_runtime.rs:26). The router emits a pending-stamp event, the application computes the stamp on whatever schedule it likes, and hands the result back. This matters because a peer chooses the stamp cost: an announced cost of 254 is legal and effectively unfinishable (#185). If that search could ever run under the core lock, any peer could stop the node by announcing a number. It cannot, and no seam added later may make it possible.

That pattern — emit a request, compute detached, submit the result — is the general answer whenever the cost of a step is not ours to bound.

One correction to the paragraph above, because its phrasing is wider than what holds. It is exactly true of the peer-priced search, which is the one that matters: generate_with is an async fn, so no synchronous callee of the event loop can drive it to completion at all. It is not true that the tree contains no synchronous proof-of-work. leviculum-core::discovery::stamp::generate_stamp (leviculum-core/src/discovery/stamp.rs:176) is a public synchronous brute-force loop taking a caller-supplied cost, and nothing stops a loop callee from calling it. It is not a DoS vector today because no peer picks its number: the only caller is the discovery announcer, which runs it once per discoverable interface during ReticulumNode::start() — off the loop — at the locally fixed DEFAULT_STAMP_VALUE. The invariant to keep is therefore "no peer-chosen cost is ever ground synchronously", and the async signature is what enforces it.

That fixed cost is not static across RNS versions: #328 raised it from 14 to 16 to stay visible to RNS 1.5.0 listeners, and each extra bit doubles the search. Measured on the coder host (release build, x86-64), one mint went from 39 ms mean / 198 ms max to 191 ms mean / 737 ms max over 24 samples. It is paid once per discoverable interface at wiring time and the result is reused for every re-announce, so this is startup latency, not a per-announce or per-loop cost. The budget argument is unchanged; the number it is measured against is four to five times larger.

A diagnostic write is inside the budget too (#418)

The budget above is about CPU: work whose cost scales with a payload. The miauhaus soak found the other half, and it is worse, because nothing about the call site looks expensive.

Over 49 days and 397 023 881 events, the node's 10 s PATH_TABLE liveness heartbeat missed at least one beat 2 928 times out of 421 605 intervals, with a tail to 37.0 s. During each of those the daemon emitted nothing at all — no packet, no announce, not the heartbeat — on a node that otherwise logs 30 to 150 events a second. Both long stalls pulled out of the raw log have the same shape: the hole sits between the ANN_RX of one announce and the PATH_ADD of that same destination, a span in which nothing can take seconds.

The emission can. Until #418 the event-log layer wrote each line with a blocking write(2), flushed, under a process-global mutex, on the thread that emitted it — and the event loop emits while it holds the core mutex (apply_inbound, leviculum-std/src/driver/mod.rs:4456). A write(2) to a USB disk under writeback throttling blocks for seconds, so the loop stopped, and everything that wanted the core queued behind it. That is why the symptom was total silence rather than a missing log line.

The rule that follows is the CPU rule's sibling:

No caller holds the core lock across an I/O call whose latency belongs to a device. A diagnostic that can stop the transport is a worse bug than the missing diagnostic.

The event log now inverts the trade (FileSink, leviculum-std/src/event_log.rs:1332): the emitting thread does a bounded enqueue and returns, one writer thread owns the file, and an overrun drops lines and says so with EVENT_LOG_DROPPED rather than blocking the mesh. LEVICULUM_EVENT_LOG_SYNC=1 restores the old behaviour for anyone who would rather block than lose a line, and it is what the mvr's positive-control arm runs.

What measures this

Three events, all threshold-gated so a healthy node emits none of them, and deliberately at different altitudes so they disagree informatively:

eventwheresays
ANN_SLOWhandle_announce (leviculum-core/src/transport.rs:6104)announce handling itself took ≥ 100 ms
CORE_STALLspawn_core_stall_watchdog (leviculum-std/src/driver/mod.rs:4247)an outside thread waited ≥ 250 ms for the core lock
EVENT_LOG_WRITE_SLOWwriter_loop (leviculum-std/src/event_log.rs:1426)one batch write to the log file took ≥ 50 ms

CORE_STALL without ANN_SLOW means the loop was stopped by something other than announce handling; EVENT_LOG_WRITE_SLOW alongside it names the disk. The watchdog measures the wait for the core mutex, never the hold, which is exactly "how long the loop spent not polling" without an Instant in a dozen select! arms.

A durable store append is inside the budget too (#384)

The #418 rule above is about a diagnostic, and a diagnostic can be dropped. The propagation node found the same shape in work that cannot be: storing a message somebody sent us.

lnpnd ran for 32 minutes in the public mesh on miauhaus (2026-09-21) and emitted 73 CORE_PROCESSOR_OVER_BUDGET warnings out of 324 log lines, 54 of them under hook="on_event" and 19 under hook="on_tick", with nine over 100 ms. lnsd on the same host, same uptime, same traffic: none. The worst one names its own cause in the line above it:

10:25:12 lnpnd: sync in peer 535d9c5db65bfcd4 transferred 105 (30240 B): ok
10:25:12 CORE_PROCESSOR_OVER_BUDGET hook="on_tick" elapsed_us=239016 budget_us=5000 events=0

A quarter of a second of core lock, having processed events=0. The time was not in event work; it was in the store. FilePropagationStore::append is a write, an fsync, a rename and a second fsync (leviculum-std/src/file_propagation_store.rs, "Power-cut safety, and why this store fsyncs"), and the engine stored the whole inbound batch inside the one hook that classified it.

The attribution is a measurement, not a reading of the code. 105 appends of 288-byte bodies cost 124 µs into a memory store and 127-189 ms into the file store on the coder host's ext4, over six runs — 1.1 to 1.5 ms each, worst single append 6.4 ms. Three orders of magnitude, on the same verb with the same bytes: the cost is the device, not the book-keeping. append_cost (leviculum-std/src/file_propagation_store.rs) prints both lines on whatever filesystem it is pointed at, and pointing it at a tmpfs — where fsync never reaches a device — reads 60x cheaper and proves nothing.

Two things follow, and only one of them is a fix.

The store cannot be the layer that fixes it. The fsyncs are what makes "persist before you prove" true (docs/src/concepts/propagation-node-on-a-board.md §3): dropping them trades a message for a millisecond. The cost is the device's, and the device's cost is not negotiable from this side.

The hook is. lnpnd's engine now queues validated payloads and stores PERSIST_PER_HOOK of them per hook (lnpnd/src/engine.rs), asking the driver to come straight back for the rest; a peer's batch is resumable across hooks (PeeringRuntime::advance_sync_resource, lnpnd/src/peering.rs) and still reports itself as one round. The batch still costs what it costs in wall-clock — the disk did not get faster — but the core lock is released between messages, so the node keeps answering while it absorbs a sync.

A cost that cannot be made cheap and cannot be dropped is still a cost that may not be paid all at once under the lock. Bound the work per hook and come back.

What this does not buy: one append can exceed the budget on its own — 6.4 ms measured, against a 5 ms budget — so CORE_PROCESSOR_OVER_BUDGET can still fire on a slow disk. What is gone is the multiplier, which is the part that scaled with someone else's batch size.

A diagnostic's volume is a cost too, and the cadence is the wrong lever

The same log gave the other half of the lesson. Of those 397 023 881 events, 214 594 631 — 54 % — are PATH_TABLE_ENTRY, the per-path snapshot of the path table; in the last megabyte of the live tail it is 77 %. The file is 61 GiB and grows ~2.6 GB a day.

That share is what remains after a fix. The dump used to fire every 10 s; 5ab938a4 (2026-07-13) moved it to every five minutes and cut its volume thirtyfold. It did not settle the problem, because a cadence is the wrong lever for this cost:

One snapshot costs one line per path. The miauhaus path table is 22 362 entries — the number is in the heartbeat itself, PATH_TABLE node=miauhaus size=22362. Five minutes apart, that is still ~6.4 million lines a day, and it grows with the mesh, not with anything the code chose.

Against that cost, nothing consumes the lines. The soak's analyze.py puts them on a count-only fast path and never reads dst, hops, iface, next_hop or expires_in_ms; periculum does not reference the event at all. And the one number counting them yields — how many paths there are — is already emitted every 10 s as PATH_TABLE size=.

So the dump is now off by default (TransportConfig::path_entries_dump, config key path_entries_dump under [reticulum]), reachable for a debugging session that wants the per-path fields. Nothing else moved: the 10 s PATH_TABLE heartbeat still carries liveness and the count, and PATH_ADD still records every insertion, so the path table's history survives the silence.

The rule, sibling to the one above:

A diagnostic whose volume scales with mesh state, not with a rate the code picks, is not a diagnostic you can leave on. Gate it, and keep the cheap scalar that answers the question people actually ask of it.