The Core Lock Budget
The async driver owns exactly one NodeCore behind one mutex
(leviculum-std/src/driver/mod.rs:1321). Every packet the node decrypts,
routes, forwards or emits passes through it. It is the narrowest point
in the stack, and the rule that follows from that is:
No caller holds the core lock across CPU-heavy work. Work that scales with payload size runs off the lock, between a cheap capture and a cheap commit.
This is not a style preference. It was measured.
The measurement that set the rule
NodeCore::send_resource used to run its whole build — bz2 compress,
bulk token encrypt, full/map hashing — inside the mutex. For a 1 MiB
incompressible payload with compression on, that was a 141 ms hold
in release. On a 20-link inbound flood it cost a 32 % inbound
throughput stall while round-sized sends ran; ~0 % after the fix
(Codeberg #152).
The fix is the three-phase shape, and it is the house idiom for anything with the same profile:
| phase | lock | code |
|---|---|---|
NodeCore::resource_send_params | brief | leviculum-core/src/node/mod.rs:1417 |
resource::prepare_resource_send | none | leviculum-core/src/resource/outgoing.rs:89 |
NodeCore::commit_resource_send | brief | leviculum-core/src/node/mod.rs:1452 |
Commit re-validates what could have changed while the build ran
unlocked: link gone, a transfer raced in, or the link re-keyed (#66) —
the last returns the retryable ResourceError::LinkStateChanged and
the caller rebuilds once. The std driver calls the three phases itself
(leviculum-std/src/driver/mod.rs:3602).
NodeCore::send_resource still exists as the composed single call
(leviculum-core/src/node/mod.rs:1543) because no_std and FFI callers
have no lock to hold and no second thread to starve. It is the
composed form that is dangerous behind the driver, not the code it
composes.
The numbers, restated for callers
Measured with leviculum-core/compression on (as leviculum-std
builds it), comparing the composed call against the locked portion of
the phased path.
The conditions, because a timing without them is an anecdote: profile
release, target x86_64-unknown-linux-musl, an AMD Ryzen 9 7950X.
Each cell is the median of five timed runs after one discarded
warm-up, and each run builds a freshly linked pair — a second
Resource to the same link cannot build at all
(ResourceError::TransferInProgress), so re-using one would not repeat
the measurement. Both harnesses are #[ignore]d and print every number
below:
measure_send_lock_costs (leviculum-lxmf/src/node.rs:2607) for the
send tables, measure_deferred_tick_costs
(leviculum-lxmf/tests/direct_delivery_attempts.rs:1560) for the tick
table.
Every column names the bytes it was given, because the cost being
reported is a compressor's and a compressor's cost is a property of its
input. The three classes are generated by
incompressible (leviculum-lxmf/tests/common/payloads.rs:32),
compressible (leviculum-lxmf/tests/common/payloads.rs:79) and
degenerate (leviculum-lxmf/tests/common/payloads.rs:104), and both
harnesses below take them from the same array so neither can drift onto
different bytes under the same name:
| class | what it is | bz2 at 1 MiB |
|---|---|---|
incompressible | seeded xorshift bytes | 1.00x — grows 0.5 % |
compressible | dictionary words with sentence breaks | 9.0x |
degenerate | vec![0x5a; n] | 21 845x |
| payload | incompressible | compressible | degenerate | phased (locked) |
|---|---|---|---|---|
| 16 KiB | 1.67 ms | 0.91 ms | 0.14 ms | 2 µs |
| 256 KiB | 15.8 ms | 11.0 ms | 1.70 ms | 3 µs |
| 1 MiB (segment 1) | 64.6 ms | 48.7 ms | 7.61 ms | 2–11 µs |
The first three columns are the composed call under the lock; the last is the locked half of the phased path, which does not vary with the payload because no payload passes through it.
Incompressible data is the worse case, not compressible data. An earlier revision of this page said the opposite — "bz2 does more work when it succeeds" — and the measurement above refuses it at every size: compressible text costs 0.54x to 0.75x of incompressible bytes of the same length. The mechanism is that the Burrows-Wheeler sort is paid in full either way, and a run that succeeds then has less output left to code, not more. Plan for the incompressible number: it is both the larger one and the one an attachment actually hits, since anything already compressed looks incompressible to bz2.
Without the compression feature — the embedded default — the 1 MiB
build drops to 6.4 ms, which is why the same code is tolerable on an
nRF52 and intolerable behind the driver. That figure is from the #152
pass and was not re-measured here.
The segment boundary sits under the 1 MiB row
RESOURCE_MAX_EFFICIENT_SIZE is 1 048 575 bytes
(leviculum-core/src/resource/mod.rs:69) and the split is decided on
the packed, uncompressed length
(leviculum-core/src/resource/outgoing.rs:106), so a 1 MiB body is
above it in every payload class — the compression ratio does not move
the boundary. The 1 MiB row therefore measures segment 1 of a
two-segment transfer, not a whole one. Segments 2..N are built on the
receive path, under the caller's lock, and the phased path does not
cover them.
That also bounds who can reach the row at all. Python LXMF refuses an
incoming delivery Resource larger than DELIVERY_LIMIT × 1000, which
is 1 000 000 bytes
(delivery_resource_advertised, reference/LXMF/LXMF/LXMRouter.py:1977)
— below the segment boundary. So no LXMF transfer that a Python peer
would accept is ever a two-segment one, and the segment path is reached
only by a Rust-to-Rust transfer or a non-LXMF core Resource user.
What the payload class was worth, as a number
This matters beyond bookkeeping, because a table measured on
degenerate was once proposed as a correction to this page. On the
same machine and profile, degenerate reads 8.5x below
incompressible at 1 MiB and 12x below it at 16 KiB: bz2's
run-length front end collapses a single repeated byte before the
Burrows-Wheeler transform ever runs, so the number that comes out is
the cost of compressing almost nothing.
Two other candidate explanations for a table reading low were tested and do not hold.
A missing warm-up is not one. The harness prints the run it discards
next to the median it keeps, so this is checkable rather than assumed:
every cell above 1 ms has its cold run within 2 % of its median, and
the largest gap anywhere is 11 % — on degenerate at 16 KiB, the
cheapest cell in the table at 154 µs cold against 139 µs. An n=1
harness on this path is imprecise; it is not biased low.
A build-profile difference is not one either. This run reproduces the
figures the page carried before it — 1.7, 16.4 and 65.2 ms — to within
4 % on the incompressible column, so those were release-profile
numbers taken on comparable hardware, and the column they belong to is
incompressible.
Costs that do not justify phasing, measured the same way: packing a 1 MiB LXMF message is 0.8 ms, and unpacking one with signature verification is 3.2 ms. Inbound verification cannot be phased away — the bytes are already in hand — and at that magnitude it does not need to be.
The adapter-side number: LxmfRouter::tick
Measured during the #196 design pass, on the same machine and profile:
one LxmfRouter::tick with 8 due 256 KiB messages holds 126.6 ms in
a single uninterrupted borrow. It is the composed-send_resource cost
of the table above, multiplied by a queue depth an adapter reaches
routinely — the router builds each due message in turn, and nothing
between them yields.
That figure is consistent with the table above: eight times the
incompressible 256 KiB build is 127 ms.
LxmfRouter can hand the build out instead, under
RouterConfig::defer_resource_builds
(leviculum-lxmf/src/router.rs:124). What the tick then costs, for one
due message, measured the same way:
| payload | deferred tick | composed tick (incompressible) |
|---|---|---|
| 16 KiB | 13.6 µs | 1.63 ms |
| 256 KiB | 300 µs | 15.9 ms |
| 1 MiB | 1.11 ms | 66.6 ms |
The deferred column is flat across all three payload classes — 13.6 to
14.6 µs at 16 KiB, 282 to 300 µs at 256 KiB — which is the mechanism
showing through: a deferring tick copies bytes and does not compress
them, so the class it was handed cannot matter. What is left in it is
the Message clone, and that is why the deferred column still grows
with size at all.
One due message rather than the eight above, because a second Resource to the same link is refused before it builds, so an eight-message comparison would not be comparing like with like. The composed column carries the router's own tick work on top of the build, which is why it reads a little above the send table at the same size; the gap is under 3 % at every row.
It is recorded here rather than in the issue because this page is where
a caller looks for it, and because a number that lives only in a tracker
cannot be cited from the tree: PROCESSOR_TICK_BUDGET
(leviculum-std/src/driver/processor.rs:181) is set against this
measurement, and until it was written down the only number behind a
public constant could not be traced at all.
What this binds
Any protocol adapter layered on the core. leviculum-lxmf is the
current instance and leviculum-lxst will be the next. An adapter that
offers only a monolithic submit call forces its host either to hold the
lock for the build or to fork the driver. Adapters that are expected to
run behind the async driver expose the phase split; the composed form
stays for the embedded caller.
Anything the driver runs inside its event loop. The loop's
dispatch_output (leviculum-std/src/driver/mod.rs:5247) routes
actions to interfaces and forwards events. Work done there blocks not
just the lock but interface I/O dispatch — strictly worse than the
mutex case. The in-loop /status responder
(leviculum-std/src/driver/remote_mgmt.rs:82) is the reference for how
much is acceptable there: take the lock, build a small bundle, hand
back a TickOutput, return.
Two things the loop's callees may never do. They may not .await,
and they may not call back into the driver's public async API: those
methods end in action_dispatch_tx.send(output).await on a bounded
channel that the same loop drains, so a full channel deadlocks the
node.
Those are one rule, not two, and knowing which way round matters when
you have to enforce it. The second is a consequence of the first: an
async fn called and not awaited builds a future and drops it, sends
nothing and blocks nothing. The deadlock needs the bounded-channel send
to complete, and only .await can complete it. So a callee expressed
as a synchronous fn has both prohibitions closed at once, which is
what the in-driver core processor (#196) is built on — see
leviculum-std/src/driver/processor.rs. It follows that a runtime
guard on the async API would be the wrong shape: there is nothing to
guard until an .await that cannot be written.
The residue is re-entrancy, not the async API
Both prohibitions above are special cases of a plainer one, and stating
them first got the emphasis wrong for two commits. The loop calls its
callees with the core mutex held, and that mutex is a non-reentrant
std::sync::Mutex. Any path from a callee back to it hangs the node
immediately — first call, no load required.
The async route is one such path and not the instructive one. Consider
the block_on case the previous wording named as the whole residue: a
callee that smuggles a PacketSender and blocks on send does not
reach the bounded channel at all, because PacketSender::send
(leviculum-std/src/driver/sender.rs:92-107) takes the core lock in a
block and releases it before its .await. It deadlocks one line
earlier, on the mutex.
And no block_on is needed. The number is 58, and it is a number
rather than a phrase on purpose: scripts/check-core-lock-census.py
rebuilds the list of public methods that lock the core out of the
sources on every just fast and pins it, name by name, in a
checked-in file (TOTAL, scripts/core-lock-census.txt:30). Two
earlier revisions of this page estimated the size of this set in words
and were low by nearly half — which is what an estimate nothing can
check is worth.
53 of them are on ReticulumNode — plain synchronous pub fns that
open by locking the core, of which has_path
(leviculum-std/src/driver/mod.rs:3085) is
self.inner.lock_recover().has_path(dest_hash) and entirely typical.
The other five are on PacketSender and LinkHandle, which matters
more than the count suggests: those are the two handles a callee is
most likely to have been handed in the first place.
A callee holding an Arc<ReticulumNode> deadlocks the node on its first
invocation, in ordinary safe synchronous code, with no .await, no
channel, and nothing a compile-fail fixture can catch.
So the rule for anything the loop calls is: hold no handle to the node
you run inside. The &mut StdNodeCore the seam hands out locks
nothing and is the whole intended surface. Everything else belongs on
the far side of a channel.
What holds this up is not the type system. It is that registration
happens on the builder, before the node exists, so a callee cannot be
constructed holding a node handle — injecting one afterwards takes a
deliberate OnceLock or Weak and a reference cycle. That is a
construction-order barrier, and it is why the hazard stays theoretical
in practice. It is not a guarantee, and this page should not be read as
offering one.
The one call the seam hands out that this page forbids
NodeCore::send_resource is pub
(leviculum-core/src/node/mod.rs:1747) and therefore reachable on the
&mut StdNodeCore a processor hook holds. It is the 141 ms composed
call this page opens with — one line, in consumer code, behind the
driver and under the lock. PROCESSOR_TICK_BUDGET reports it 141 ms
after the fact and cannot prevent it, and no fixture can refuse it: it
compiles, because for the no_std and FFI callers it is the correct API.
A hook that has to send a resource uses the three-phase form the driver
itself uses — resource_send_params, prepare_resource_send off the
lock, commit_resource_send — or hands the send to the application side
of a channel. This is named again in
leviculum-std/src/driver/processor.rs where a consumer will meet it.
Work that is already off the core by construction
Proof-of-work stamp generation borrows only its executor, never the
router and never the core
(leviculum-lxmf/src/router/stamp_runtime.rs:26). The router emits a
pending-stamp event, the application computes the stamp on whatever
schedule it likes, and hands the result back. This matters because a
peer chooses the stamp cost: an announced cost of 254 is legal and
effectively unfinishable (#185). If that search could ever run under
the core lock, any peer could stop the node by announcing a number.
It cannot, and no seam added later may make it possible.
That pattern — emit a request, compute detached, submit the result — is the general answer whenever the cost of a step is not ours to bound.
One correction to the paragraph above, because its phrasing is wider
than what holds. It is exactly true of the peer-priced search, which
is the one that matters: generate_with is an async fn, so no
synchronous callee of the event loop can drive it to completion at all.
It is not true that the tree contains no synchronous proof-of-work.
leviculum-core::discovery::stamp::generate_stamp
(leviculum-core/src/discovery/stamp.rs:176) is a public synchronous
brute-force loop taking a caller-supplied cost, and nothing stops a
loop callee from calling it. It is not a DoS vector today because no
peer picks its number: the only caller is the discovery announcer,
which runs it once per discoverable interface during
ReticulumNode::start() — off the loop — at the locally fixed
DEFAULT_STAMP_VALUE. The invariant to keep is therefore "no
peer-chosen cost is ever ground synchronously", and the async
signature is what enforces it.
That fixed cost is not static across RNS versions: #328 raised it from 14 to 16 to stay visible to RNS 1.5.0 listeners, and each extra bit doubles the search. Measured on the coder host (release build, x86-64), one mint went from 39 ms mean / 198 ms max to 191 ms mean / 737 ms max over 24 samples. It is paid once per discoverable interface at wiring time and the result is reused for every re-announce, so this is startup latency, not a per-announce or per-loop cost. The budget argument is unchanged; the number it is measured against is four to five times larger.
A diagnostic write is inside the budget too (#418)
The budget above is about CPU: work whose cost scales with a payload. The miauhaus soak found the other half, and it is worse, because nothing about the call site looks expensive.
Over 49 days and 397 023 881 events, the node's 10 s PATH_TABLE
liveness heartbeat missed at least one beat 2 928 times out of
421 605 intervals, with a tail to 37.0 s. During each of those the
daemon emitted nothing at all — no packet, no announce, not the
heartbeat — on a node that otherwise logs 30 to 150 events a second.
Both long stalls pulled out of the raw log have the same shape: the
hole sits between the ANN_RX of one announce and the PATH_ADD of
that same destination, a span in which nothing can take seconds.
The emission can. Until #418 the event-log layer wrote each line with
a blocking write(2), flushed, under a process-global mutex, on the
thread that emitted it — and the event loop emits while it holds the
core mutex (apply_inbound,
leviculum-std/src/driver/mod.rs:4456). A write(2) to a USB disk
under writeback throttling blocks for seconds, so the loop stopped,
and everything that wanted the core queued behind it. That is why the
symptom was total silence rather than a missing log line.
The rule that follows is the CPU rule's sibling:
No caller holds the core lock across an I/O call whose latency belongs to a device. A diagnostic that can stop the transport is a worse bug than the missing diagnostic.
The event log now inverts the trade (FileSink,
leviculum-std/src/event_log.rs:1332): the emitting thread does a
bounded enqueue and returns, one writer thread owns the file, and an
overrun drops lines and says so with EVENT_LOG_DROPPED rather than
blocking the mesh. LEVICULUM_EVENT_LOG_SYNC=1 restores the old
behaviour for anyone who would rather block than lose a line, and it
is what the mvr's positive-control arm runs.
What measures this
Three events, all threshold-gated so a healthy node emits none of them, and deliberately at different altitudes so they disagree informatively:
| event | where | says |
|---|---|---|
ANN_SLOW | handle_announce (leviculum-core/src/transport.rs:6104) | announce handling itself took ≥ 100 ms |
CORE_STALL | spawn_core_stall_watchdog (leviculum-std/src/driver/mod.rs:4247) | an outside thread waited ≥ 250 ms for the core lock |
EVENT_LOG_WRITE_SLOW | writer_loop (leviculum-std/src/event_log.rs:1426) | one batch write to the log file took ≥ 50 ms |
CORE_STALL without ANN_SLOW means the loop was stopped by
something other than announce handling; EVENT_LOG_WRITE_SLOW
alongside it names the disk. The watchdog measures the wait for the
core mutex, never the hold, which is exactly "how long the loop spent
not polling" without an Instant in a dozen select! arms.
A durable store append is inside the budget too (#384)
The #418 rule above is about a diagnostic, and a diagnostic can be dropped. The propagation node found the same shape in work that cannot be: storing a message somebody sent us.
lnpnd ran for 32 minutes in the public mesh on miauhaus (2026-09-21) and
emitted 73 CORE_PROCESSOR_OVER_BUDGET warnings out of 324 log lines,
54 of them under hook="on_event" and 19 under hook="on_tick", with nine
over 100 ms. lnsd on the same host, same uptime, same traffic: none. The
worst one names its own cause in the line above it:
10:25:12 lnpnd: sync in peer 535d9c5db65bfcd4 transferred 105 (30240 B): ok
10:25:12 CORE_PROCESSOR_OVER_BUDGET hook="on_tick" elapsed_us=239016 budget_us=5000 events=0
A quarter of a second of core lock, having processed events=0. The time
was not in event work; it was in the store. FilePropagationStore::append
is a write, an fsync, a rename and a second fsync
(leviculum-std/src/file_propagation_store.rs, "Power-cut safety, and why
this store fsyncs"), and the engine stored the whole inbound batch inside
the one hook that classified it.
The attribution is a measurement, not a reading of the code. 105 appends of
288-byte bodies cost 124 µs into a memory store and 127-189 ms into
the file store on the coder host's ext4, over six runs — 1.1 to 1.5 ms each,
worst single append 6.4 ms. Three orders of magnitude, on the same verb with
the same bytes: the cost is the device, not the book-keeping. append_cost
(leviculum-std/src/file_propagation_store.rs) prints both lines on
whatever filesystem it is pointed at, and pointing it at a tmpfs — where
fsync never reaches a device — reads 60x cheaper and proves nothing.
Two things follow, and only one of them is a fix.
The store cannot be the layer that fixes it. The fsyncs are what makes
"persist before you prove" true
(docs/src/concepts/propagation-node-on-a-board.md §3): dropping them
trades a message for a millisecond. The cost is the device's, and the
device's cost is not negotiable from this side.
The hook is. lnpnd's engine now queues validated payloads and stores
PERSIST_PER_HOOK of them per hook (lnpnd/src/engine.rs), asking the
driver to come straight back for the rest; a peer's batch is resumable
across hooks (PeeringRuntime::advance_sync_resource, lnpnd/src/peering.rs)
and still reports itself as one round. The batch still costs what it costs
in wall-clock — the disk did not get faster — but the core lock is released
between messages, so the node keeps answering while it absorbs a sync.
A cost that cannot be made cheap and cannot be dropped is still a cost that may not be paid all at once under the lock. Bound the work per hook and come back.
What this does not buy: one append can exceed the budget on its own —
6.4 ms measured, against a 5 ms budget — so CORE_PROCESSOR_OVER_BUDGET
can still fire on a slow disk. What is gone is the multiplier, which is the
part that scaled with someone else's batch size.
A diagnostic's volume is a cost too, and the cadence is the wrong lever
The same log gave the other half of the lesson. Of those 397 023 881
events, 214 594 631 — 54 % — are PATH_TABLE_ENTRY, the per-path
snapshot of the path table; in the last megabyte of the live tail it is
77 %. The file is 61 GiB and grows ~2.6 GB a day.
That share is what remains after a fix. The dump used to fire every
10 s; 5ab938a4 (2026-07-13) moved it to every five minutes and cut
its volume thirtyfold. It did not settle the problem, because a
cadence is the wrong lever for this cost:
One snapshot costs one line per path. The miauhaus path table is 22 362 entries — the number is in the heartbeat itself,
PATH_TABLE node=miauhaus size=22362. Five minutes apart, that is still ~6.4 million lines a day, and it grows with the mesh, not with anything the code chose.
Against that cost, nothing consumes the lines. The soak's analyze.py
puts them on a count-only fast path and never reads dst, hops,
iface, next_hop or expires_in_ms; periculum does not reference
the event at all. And the one number counting them yields — how many
paths there are — is already emitted every 10 s as PATH_TABLE size=.
So the dump is now off by default
(TransportConfig::path_entries_dump, config key path_entries_dump
under [reticulum]), reachable for a debugging session that wants the
per-path fields. Nothing else moved: the 10 s PATH_TABLE heartbeat
still carries liveness and the count, and PATH_ADD still records
every insertion, so the path table's history survives the silence.
The rule, sibling to the one above:
A diagnostic whose volume scales with mesh state, not with a rate the code picks, is not a diagnostic you can leave on. Gate it, and keep the cheap scalar that answers the question people actually ask of it.