Nostr Reliability: Live Relay Delivery and NIP-46 Reply Correlation
Make a fast reply observable
A remote signing request can fail from the client's perspective even when the signer answered correctly. If the client starts listening after sending, a fast reply may already be gone. If it retains an old reply in its state, a later timeout can be mistaken for success.
DamageBDD's live NIP-46 test path addresses both cases. It establishes reply listeners before publication, correlates the decrypted response with the current request ID, and clears old response state before the next request. The bunker relay adapter manages its connection workers and dispatches incoming requests asynchronously.
This matters for any integration in which a request crosses a relay and a separate response must come back. Connection success, request delivery and a completed operation are different observations. Keeping them separate makes a failure easier to diagnose.
Release-build run 20260925132037 records the live bunker scenarios passing
on 25 September 2026. That result remains tied to its tested release and
deployment; public-relay conditions and later releases need fresh checks.
Why subscription order matters
NIP-46 request and response traffic uses kind 24133. NIP-01 places that kind in
the ephemeral range, whose events are not expected to be retained. A client
which publishes a request and only then queries for the reply can miss the
response permanently. Moving the query's since value backwards cannot recover
an event the relay never stored. See NIP-01.
The corrected test path starts a reply-listener worker for each configured
relay. Each worker opens a WebSocket connection and sends a reply filter with
limit set to zero. It reports readiness when the matching EOSE arrives.
The parent publishes only after readiness has been established for at least one
listener. Listeners remain open while publication and bunker ingress are
checked, allowing a fast matching reply to wait in the parent's mailbox.
The implementation uses dedicated listener connections and a separate publication path. This also exercises delivery between connections rather than relying on one socket to observe its own traffic. Listeners remain active until the matching response arrives or the operation reaches its bounded failure condition.
Matching the current request
The response filter narrows traffic by kind, expected bunker author and the client's recipient tag. After decryption, the response ID must match the current request ID. The later assertion checks that correlation again.
Before creating each request, the function removes
last_nip46_reply_event, last_nip46_response and
last_nip46_reply_error from its live-test namespace. It replaces that namespace
with maps:put/3. The normal put_live/2 helper merges maps and therefore cannot
be used to remove these fields.
This prevents a particularly confusing failure: an earlier sign_event result
being mistaken for the answer to a later ping. A timeout now remains a timeout
instead of passing a reply-exists check against old context.
Relay adapter responsibilities
damage_nsecbunker_relay owns the public-relay connections for the bunker
bridge. It subscribes to requests addressed to the bunker and publishes signed
response events returned by the bunker path. The module does not itself load
vault secrets or perform signing.
The adapter keeps connector workers alive as Gun connection owners, tracks recently observed event IDs and exposes operational counters. Inbound dispatch is asynchronous. This avoids a call cycle in which the relay adapter waits synchronously for the client bridge while that bridge tries to publish a response back through the same adapter.
The test's readiness step examines adapter status instead of treating the
initial subscribe/0 return as proof that the subscription is active. Listener
workers monitor the parent and close their connection on completion, stop or
parent death. These controls make lifecycle failures visible, but they do not
turn a remote relay into a durable request queue.
Distinguish each layer of evidence
| Observation | What it establishes |
|---|---|
| TCP and TLS complete | A transport path to the endpoint exists |
| HTTP 101 | The endpoint accepted the WebSocket upgrade |
| Matching subscription readiness | The test has established its listener before sending |
| Relay accepts the request event | At least one relay accepted that event |
| Adapter counters and event ID advance | The bunker-side adapter observed and dispatched it |
| Correlated decrypted reply arrives | The request completed the reply path |
| Expected result and author checks pass | The particular response satisfies the scenario |
Request ingress alone does not show that response publication succeeded. A WebSocket upgrade alone does not establish kind-24133 delivery. These distinctions are why the live suite retains all of its separate assertions.
Diagnose the layer that failed
A successful WebSocket upgrade only confirms that the transport opened. If no request reaches the bunker, inspect the event kind, recipient filter, relay policy and subscription state. If the request arrives but no reply does, inspect signing outcome, response publication and client correlation.
Repeat the check from the node and network where the service will run. Use the intended identities and a current kind-24133 round trip. Public relay behaviour can change, so a previously working relay list is a starting configuration rather than evidence that today's request will complete.
Keep the failing event ID, request ID and relevant counters with the report. Those details let another engineer repeat the same observation without guessing which stage a generic timeout describes.
Evidence and remaining limits
The recorded release result includes relay canary ingress, black-box ping and
signing loops, ordinary live round trips, and policy-rejection responses as
successful. The tests use the node's existing damage_nostr identity. A
separate local policy probe verifies the requested external client's allowlist
entry; it does not replace that client's end-to-end smoke test.
The listener readiness loop can spend time waiting for other configured relays even when one becomes ready earlier. The report includes heartbeat delays, and no latency target is established by a green result. A bounded negative query for a signed article also establishes only that it was not observed on the checked relay path during the check.
Explore the implementation and recorded result
- Reply listeners, state cleanup and correlation checks
- Relay connection ownership and asynchronous dispatch
- Request and response bridge
- Live bunker acceptance scenarios
- Release-build run 20260925132037
- NIP-46 protocol specification
The code links pin the revision used for this explanation. The September report describes its recorded release; this documentation update did not repeat that live run or measure current relay availability.
