Nostr Reliability: Live Relay Delivery and NIP-46 Reply Correlation

Make a fast reply observable

A remote signing request can fail from the client's perspective even when the signer answered correctly. If the client starts listening after sending, a fast reply may already be gone. If it retains an old reply in its state, a later timeout can be mistaken for success.

DamageBDD's live NIP-46 test path addresses both cases. It establishes reply listeners before publication, correlates the decrypted response with the current request ID, and clears old response state before the next request. The bunker relay adapter manages its connection workers and dispatches incoming requests asynchronously.

This matters for any integration in which a request crosses a relay and a separate response must come back. Connection success, request delivery and a completed operation are different observations. Keeping them separate makes a failure easier to diagnose.

Release-build run 20260925132037 records the live bunker scenarios passing on 25 September 2026. That result remains tied to its tested release and deployment; public-relay conditions and later releases need fresh checks.

Why subscription order matters

NIP-46 request and response traffic uses kind 24133. NIP-01 places that kind in the ephemeral range, whose events are not expected to be retained. A client which publishes a request and only then queries for the reply can miss the response permanently. Moving the query's since value backwards cannot recover an event the relay never stored. See NIP-01.

The corrected test path starts a reply-listener worker for each configured relay. Each worker opens a WebSocket connection and sends a reply filter with limit set to zero. It reports readiness when the matching EOSE arrives. The parent publishes only after readiness has been established for at least one listener. Listeners remain open while publication and bunker ingress are checked, allowing a fast matching reply to wait in the parent's mailbox.

The implementation uses dedicated listener connections and a separate publication path. This also exercises delivery between connections rather than relying on one socket to observe its own traffic. Listeners remain active until the matching response arrives or the operation reaches its bounded failure condition.

Matching the current request

The response filter narrows traffic by kind, expected bunker author and the client's recipient tag. After decryption, the response ID must match the current request ID. The later assertion checks that correlation again.

Before creating each request, the function removes last_nip46_reply_event, last_nip46_response and last_nip46_reply_error from its live-test namespace. It replaces that namespace with maps:put/3. The normal put_live/2 helper merges maps and therefore cannot be used to remove these fields.

This prevents a particularly confusing failure: an earlier sign_event result being mistaken for the answer to a later ping. A timeout now remains a timeout instead of passing a reply-exists check against old context.

Relay adapter responsibilities

damage_nsecbunker_relay owns the public-relay connections for the bunker bridge. It subscribes to requests addressed to the bunker and publishes signed response events returned by the bunker path. The module does not itself load vault secrets or perform signing.

The adapter keeps connector workers alive as Gun connection owners, tracks recently observed event IDs and exposes operational counters. Inbound dispatch is asynchronous. This avoids a call cycle in which the relay adapter waits synchronously for the client bridge while that bridge tries to publish a response back through the same adapter.

The test's readiness step examines adapter status instead of treating the initial subscribe/0 return as proof that the subscription is active. Listener workers monitor the parent and close their connection on completion, stop or parent death. These controls make lifecycle failures visible, but they do not turn a remote relay into a durable request queue.

Distinguish each layer of evidence

Observation What it establishes
TCP and TLS complete A transport path to the endpoint exists
HTTP 101 The endpoint accepted the WebSocket upgrade
Matching subscription readiness The test has established its listener before sending
Relay accepts the request event At least one relay accepted that event
Adapter counters and event ID advance The bunker-side adapter observed and dispatched it
Correlated decrypted reply arrives The request completed the reply path
Expected result and author checks pass The particular response satisfies the scenario

Request ingress alone does not show that response publication succeeded. A WebSocket upgrade alone does not establish kind-24133 delivery. These distinctions are why the live suite retains all of its separate assertions.

Diagnose the layer that failed

A successful WebSocket upgrade only confirms that the transport opened. If no request reaches the bunker, inspect the event kind, recipient filter, relay policy and subscription state. If the request arrives but no reply does, inspect signing outcome, response publication and client correlation.

Repeat the check from the node and network where the service will run. Use the intended identities and a current kind-24133 round trip. Public relay behaviour can change, so a previously working relay list is a starting configuration rather than evidence that today's request will complete.

Keep the failing event ID, request ID and relevant counters with the report. Those details let another engineer repeat the same observation without guessing which stage a generic timeout describes.

Evidence and remaining limits

The recorded release result includes relay canary ingress, black-box ping and signing loops, ordinary live round trips, and policy-rejection responses as successful. The tests use the node's existing damage_nostr identity. A separate local policy probe verifies the requested external client's allowlist entry; it does not replace that client's end-to-end smoke test.

The listener readiness loop can spend time waiting for other configured relays even when one becomes ready earlier. The report includes heartbeat delays, and no latency target is established by a green result. A bounded negative query for a signed article also establishes only that it was not observed on the checked relay path during the check.

Explore the implementation and recorded result

The code links pin the revision used for this explanation. The September report describes its recorded release; this documentation update did not repeat that live run or measure current relay availability.