Fleetmesh fleetmesh

The fleetmesh Protocol · part 8 of 10

Running a node

Privacy threat modeling, adversary analysis, operator manifests, and reference implementation requirements.

Privacy rules

Privacy Guarantees & Threat Modeling

Publishing receipts and reports makes traffic volumes public, and publishing grant envelopes makes the existence and timing of peerings public. Those are the deliberate trades: auditability of enforcement in exchange for a coarse view of how much data moved and when relationships changed.

The cost of a public grant table is not, as it is easy to assume, mere "traffic volume leakage". Such a table exposes full peering topologies, bilateral subjective ratings and exact downgrade timestamps. Blinded commitments remove that surveillance surface entirely.

  • Counters, never addresses. Receipts and reports carry aggregate counts only. Client IP addresses, user agents, and connection fingerprints stay on the node that observed them. A relay behind a proxy resolves client addresses to enforce local rate limits. That address resolution MUST NOT reach any federated mesh surface. Exposing resolved client IP addresses across the mesh creates an unacceptable deanonymization vector.
  • Counters are bucketed. Published counts round to two significant figures. Exact byte counts across short windows fingerprint individual conversations; the reputation arithmetic does not need that resolution and MUST NOT be given it.
  • Gift wraps are not a shard. The replication rule is defined in § Replication. Refusals are strictly silent. Distinguishing between nonexistent events and forbidden events leaks private communication metadata. A federated puller requesting unauthorized #p-scoped Kind 1059 events receives the exact response given to an unauthenticated stranger.
  • Interests are coarse. An interest filter is a standing declaration of what a node cares about, readable by whoever serves it. Nodes SHOULD declare interests at kind and tag granularity rather than author granularity; a node that wants one person's events should pull the shard containing them. An anchor MAY take the full firehose within a scope, which is the least revealing option available to it.
  • Leaves hide behind mediators. Hole punching exposes IP addresses to peers and rendezvous servers. The v1 leaf profile disables direct peer connections entirely. A leaf maintains a single wss connection to its mediator. Only the mediator observes its network address. Future WebRTC extensions MUST NOT make direct peer connections a default.
  • Two mediators, not one. A mediator observes message sizes, arrival times, and connection schedules. Under the Three-Layer Operator Independence Model, leaves SHOULD maintain dual-mediation with two distinct anchors. The leaf operator selects anchors with distinct operator pubkeys and separate BGP ASNs. This distributes message metadata so no single intermediary observes complete traffic patterns.

Attack surface

Security Analysis & Failure Modes

Attack Mitigation Residual
sybil flood Probation grants are near-worthless; weight requires clean windows observed by the victim; PoW on descriptors. Noise in discovery. Descriptor storage cost, bounded by expiration.
eclipse a leaf Leaves MUST hold mediation with at least two anchors publishing different operator tags, and compare their coverage fingerprints. A weak proxy, since the tag is self-declared — but mechanically checkable, which "independently operated" is not until § Open questions settles it. A leaf with exactly one operator's anchors is eclipsable. Detectable, not preventable.
receipt forgery Receipts are signed by the subject and hash-chained per window; the issuer stores the head. None on the chain. A subject can still refuse to send receipts, which is an S2 by absence.
grant forgery Grants are signed nostr events by the issuer; the subject verifies before enforcing. An issuer can shrink a grant retroactively-looking via clock games — hence the window grace period.
laundering through a forwarder Delivery envelopes carry a signed hop path; the receiver charges every node in the path it holds a grant with. A forwarder outside the receiver's grant table is uncharged — but then it is also unknown, and probation-limited.
withholding within a claimed shard completeness: asserted is falsifiable by sampling; a missing event id is evidence for an S2. Selective withholding to one peer only is expensive to distinguish from lag. Cross-checking fingerprints across peers detects it eventually.
pull amplification Wide range requests are charged to bytes_out and sync_ranges against the puller's grant. The first oversized request is served before the budget bites. Bounded by the window's remaining capacity.
retaliation for a report Evidence immunity: a reporter whose evidence verifies keeps its weight in every other node's arithmetic, so a subject cannot make reporting cost the reporter anything network-wide. Unsolved bilaterally. The subject can zero the reporter's grant, and because grants are encrypted to their subjects nobody else can see that it happened. Anonymous reporting would fix it and is not constructible — see § Trust weighting.
curated bootstrap mirror Every anchor named in a bootstrap.toml is verified by its own signed descriptor, so a mirror cannot invent nodes. The file is unsigned and grants no authority; one --bootstrap address, an empty list, gossipsub or any nostr relay each reach the mesh without it. A mirror can still serve a list naming only anchors it controls, and every one of them verifies. A first-run node's initial view is chosen by whoever served the file. Detected by comparing coverage fingerprints across anchors, or by using any second discovery path — not prevented.
false-report campaign Reports need subject-signed evidence, weigh zero from probation nodes, and cost the reporter standing when contradicted. A trusted node can spend its standing on one damaging false report. It does so exactly once.
transport key impersonation A transport key binds to a DID only through a live signed descriptor; connections from unbound keys are strangers. A stolen transport key impersonates until the descriptor is replaced. Ed25519 keys are cheap to rotate for this reason.
clock manipulation Grace period of one tenth of a window; skew over 60 s produces a notice rather than a report. Sustained skew degrades a peer's measured conformance. Operators must run NTP; the spec cannot enforce it.
split-brain partition Replication is set union; both sides remain internally correct and reconcile on heal. Divergent deletion visibility during the partition. Inherent to the data model.
The one that is not solved

A compliant anchor node that turns hostile is difficult to detect immediately because its traffic conforms to granted limits. The protocol bounds potential damage to the granted filter scope, and remedies require only a single grant modification. Detecting subtle data corruption by an otherwise compliant peer requires cross-checking range fingerprints against an independent third node. Operators requiring strict data guarantees should maintain asserted coverage locally rather than delegating validation.

Alpha contract

Alpha Evaluation Scope & Guarantees

The mesh ships as alpha and will stay there for some time. That word is doing real work: it licenses breaking changes to nearly everything specified here, on the understanding that a short list of promises holds anyway. Operators are volunteering hardware, and volunteers who get burned do not come back, so that list is the whole basis on which anyone should be asked to run a node.

Promises that hold during alpha

no data loss, ever
Mesh participation is strictly additive to local storage. A node MUST NOT delete, overwrite, or mutate existing local events based on peer messages. The only permitted store mutations are new event insertions and verified NIP-09 author deletions.
leaving is a flag
Setting --mesh=off halts mesh participation immediately without restarts or migrations. A node retains all previously stored data upon exit.
opt-in per shard
Federation is granular. Operators publish specific shards while isolating others. New nodes publish no data until explicitly configured.
enforcement stays off by default
The reputation engine starts in shadow mode: it measures, evaluates, and logs without throttling. Operators explicitly enable active budget enforcement.
a standard relay underneath
A mesh node functions as a standard, compliant NIP-01 relay for third-party Nostr clients. Standard clients (such as nak, mobile apps, or web extensions) connect and interact without requiring mesh awareness.

What alpha explicitly permits breaking

  • Kind numbers, and every tag name inside them.
  • Grant, receipt and report schemas, including which limits exist.
  • Every constant in the window arithmetic and the trust weighting.
  • DIDComm protocol URIs, message names and state machines.
  • Transport ALPNs, gossipsub topic names, and the negotiation ladder's ordering.

Breaking changes are not announced, because nothing announces anything here. A node declares the version it speaks, peers compare digests on connect, and a mismatch declines politely rather than half-speaking — see § Governance. During alpha the declared version is fleetmesh/0.1-draft, whose whole meaning is that its digest may change on any day and nobody is owed a migration. Running alpha means accepting a re-sync, and possibly several.

Reputation resets at GA

Every grant, receipt and report from the alpha period is discarded when the mesh reaches 1.0. Standing accumulated against constants that changed underneath it is not standing, and carrying it forward would bake early participants' luck into the network permanently. Nobody should join the alpha for a head start; they should join it to find out whether the thing works.

What alpha is for measuring

Alpha is not only a warning label; it is the only chance to collect the numbers that several deferred decisions are waiting on. Nodes SHOULD record and publish, in aggregate:

  • Whether grants bind at all, and which limit binds first. This is the measurement that answers § Grants & receipts' settlement question with data instead of a guess. If capacity is never the scarce thing, the settlement field gets deleted rather than filled in.
    A first answer, from a reference sync rather than from a network. Reconciling a plain text shard, events_in is exhausted long before any byte limit: those events average 290 bytes, so a member-tier events_in ceiling of 3,000 is reached with the bytes_in ceiling 39× away, and the same holds at probation and trusted because both scale together. For text traffic the byte ceilings are decoration and events_in is the grant. Either it is too tight or the byte ceilings are far too loose; one of the two should move before anyone calls these calibrated. Blob traffic is the case that would invert it, and blob_bytes defaults to zero.
  • The distribution of divergence between claimed and measured counters among peers nobody suspects. Every constant in the ladder is calibrated against that distribution or against nothing.
  • How often the transport ladder falls through to wss, and at which rung. If hole punching succeeds rarely enough, the edge class needs rethinking rather than tuning.
  • Which retention mode operators actually choose, and whether peers leaves them unable to evaluate a candidate before peering.
  • Whether disk is ever the binding constraint. Coverage declined for capacity, shards dropped for capacity rather than age, and whether any operator raises blob_bytes above its default at all. This is the measurement that answers § Grants & receipts' storage question. If nobody ever declines a shard for want of disk, storage scarcity was imaginary, and the counters are deleted rather than acted on.

kind 21802 — ephemeral diagnostic telemetry (opt-in)

{
  "kind": 21802,
  "tags": [
    ["metric", "grant_binding", "events_in:12", "bytes_out:45", "sync_ranges:3"],
    ["metric", "divergence_hist", "s0:984", "s1:14", "s2:2", "s3:0"],
    ["metric", "transport_success", "iroh:820", "libp2p:120", "wss:60", "unix:0"],
    ["metric", "retention_mode", "peers"],
    ["metric", "shard_bytes_held", "p50:8388608", "p95:134217728"],
    ["metric", "coverage_declined", "capacity:0", "policy:3"],
    ["metric", "shard_dropped", "age:2", "capacity:0"],
    ["metric", "blob_ceiling_raised", "false"],
    ["sample_window", "86400"],
    ["expiration", "1787090000"]
  ],
  "content": ""
}

Diagnostic Collection Path: Kind 21802 is an ephemeral event (NIP-01, 20000 ≤ kind < 30000) broadcast over gossipsub or published to opted-in diagnostic monitors. To preserve absolute privacy, all counters are coarse-binned and aggregated across the whole node over a 24-hour window, with no per-peer attribution or identifying metadata.

Provenance during alpha

Data received solely through the mesh carries provisional status. A peer's claim justifies verification rather than assertion. Concretely: a node MUST NOT publish completeness: asserted for shards held only through mesh replication. The gateway MUST record mesh origin on all quarantined events. This prevents third-party replication from laundering unverified peer data into authoritative records.

Implementation

Reference Implementation Components

Very little here is new machinery. A conformant node is an ordinary NIP-01 relay with three additions: a pluggable event store, a write-policy seam, and an embedded-database backend for the small targets. Good relay software already has all three. did:nostr is specified and implementable in a few hundred lines. Blossom already addresses blobs by content hash. What has to be built is range reconciliation, the mesh layer itself, and the gateway that keeps it at arm's length from whatever else an operator runs.

New crates

This is the shape a production implementation should take, and none of it exists. The one implementation that does is Python, built around the executable reference in ref/ rather than in this layout, and its only claim is that it runs the whole of this document once over a wire and found several defects doing so. Read the list below as a plan.

crates/mesh
The library. Node DID document assembly, DIDComm v2 envelopes and the six protocols, grant and receipt types with canonical serialisation, the window evaluator, the trust weighting, and the transport ladder. No I/O policy of its own.
crates/mesh-node
The binary a stranger installs: relay plus mesh, optional blob server, embedded store, no orchestrator and no external database. Statically linked for aarch64 and x86_64, built size-optimised with LTO, because the target is a device somebody already owns rather than one they buy for this.
crates/mesh-gateway
The boundary, for operators bridging an existing estate. A separate deployment in its own isolation boundary, holding its own mesh node key and its own storage, whose only credential on the other side is a relay connection. It speaks wss:// and NIP-77 to that relay exactly as an outside client would, and links none of the operator's own code. This is the piece that makes § Two networks true rather than merely intended.
crates/mesh-mobile
The leaf, exposed through UniFFI for Swift and Kotlin. One WebSocket, a range reconciler, a bounded store and a keystore binding — no p2p stack compiled in at all. It assumes the process dies without warning and checkpoints after every batch.
crates/mesh-transports
Modular, pluggable transport adapters implementing MeshTransport. Ships with feature-gated crates: mesh-transport-iroh (QUIC & hole punching), mesh-transport-libp2p (gossipsub & discovery), mesh-transport-ws (mediated WebSockets/HTTPS), and mesh-transport-unix (zero-copy local IPC).
crates/meshctl
The declarative operator CLI and orchestration engine. Parses, validates and diffs Kubernetes-style manifests (apiVersion: fleetmesh.org/v1alpha1). Manages the local runtime daemon by non-destructive hot reconciliation. Exports to Kubernetes CRDs, Podman Quadlets and systemd service units.
crates/mesh-stores
Modular event storage backends implementing MeshStore with deterministic range fingerprinting. Ships with mesh-store-lmdb, mesh-store-sqlite, and adapter interfaces for strfry, khatru, and PostgreSQL.
crates/mesh-executors
Modular compute runtime execution drivers implementing ComputeExecutor: mesh-executor-wasmtime, mesh-executor-wasmer, and container runner adapters for Bacalhau.
crates/mesh-payment
Modular settlement rail drivers implementing PaymentRail: mesh-payment-ipd (Interledger Payment Daemon / ILP STREAM), mesh-payment-nwc (NIP-47 Nostr Wallet Connect), mesh-payment-ln (LND / Core Lightning / LDK), and mesh-payment-cashu (Chaumian ecash).
crates/mesh-inference
Modular AI/LLM inference router drivers implementing InferenceProvider: mesh-inference-routstr (Routstr OpenAI-compatible proxy with Cashu token auth), mesh-inference-nip90 (NIP-90 Data Vending Machine event router), and mesh-inference-local (Ollama / vLLM local GPU driver).

What an existing relay has to add

Every item below is scoped to a capability rather than a file, because it has to be implementable against any relay codebase. Each is also independently useful: a relay that adds them all and never joins the mesh has still gained range sync, better rate limiting and a correctness fix.

Capability What it means concretely
range fingerprints The event store gains a fingerprint over an arbitrary filter-and-time range, plus id enumeration within one. Every backend an implementation ships MUST produce byte-identical fingerprints for identical event sets, or two nodes silently fail to converge.
NIP-77 frames NEG-OPEN / NEG-MSG / NEG-CLOSE on the websocket, and 77 advertised in the NIP-11 document.
grant-shaped limits Rate limiting keyed on a subject — node DID, authenticated pubkey, or IP — rather than on IP alone. Bucket sizes come from the grant in force, not from a flat constant.
grants as write policy A grant source feeding the same accept / reject / shadow-reject decision the relay's policy layer already produces. Existing allow and deny lists become its file-backed special case; any strfry-compatible plugin seam stays untouched beside it.
peer-aware read gating The gate on gift wraps gains a peer arm: a federated puller is never the recipient, and qualifies only under the rule in § Replication.
mesh egress & pluginOut A separate outbound path off the local event broadcast, applying per-grant shaping, NIP-77 interest filtering, and the pluginOut egress-policy seam. Internal fanout remains completely untouched and isolated.
DID node profile keyAgreement, DIDCommMessaging service entries, and X25519 / Ed25519 Multikey encoding alongside the secp256k1 path the did:nostr draft already defines.
monitor probing For anyone running a NIP-66 monitor: probe mesh nodes as well as relays and publish kind 30166 for them. Reads only public data, so it crosses no boundary and needs nobody's permission.
blob transfer iroh-blobs fetch alongside the BUD-04 HTTP mirror path, and the CID mapping on the read side.
One-way dependency

The dependency arrow points one way and MUST stay that way: the mesh layer links the relay; the relay never links the mesh. Every capability above is defined so that it can be implemented with no mesh types in the relay's signatures at all. A build in which the relay imports the mesh has lost the boundary before it has served a single request, and no amount of care elsewhere gets it back.

Rollout

  • W1
    NegentropyNIP-77 in the relay, both stores, conformance tests. Useful on its own — every client with range sync benefits immediately, and nothing else here works without it.
  • W2
    BoundaryThe gateway skeleton, its own namespace, keys and quota, and the unplug test wired into CI. Built before there is anything to gate, because a boundary retrofitted after traffic exists is a boundary that already leaked.
  • W3
    IdentityNode descriptors, the DID node profile, revocation. Publishable and resolvable before anything peers, so identity can be validated in isolation.
  • W4
    Control planeDIDComm v2 over the wss fallback only. The slowest transport, chosen first: it exercises every protocol without a p2p stack in the way.
  • W5
    GrantsIssuance, receipts, window evaluation, the ladder — shadow mode only. At least four weeks against real traffic before any constant is treated as settled.
  • W6
    Transportsiroh and libp2p, the negotiation ladder, mediation. The first point at which an edge behind NAT is a real participant.
  • W7
    The binaryPackaged fleetmesh-node, ARM images, a one-command install, a published bootstrap.toml. The first release an outsider can actually run — and the point at which alpha starts meaning something to someone other than us.
  • W8
    LeavesMobile library and QR pairing, at the scoped profile only: one wss connection to a mediator, foreground sync, push to wake. Last, because a leaf depends on every other layer being real — and small, because the operating systems have already decided how small.
Shadow mode is not optional

W5 ships with enforcement disabled by default and a flag to enable it. Every parameter in the arithmetic — the 110% threshold, the divergence floor, the AIMD step, the tier weights — is a guess until it has been measured against real peer behaviour. A ladder tuned on assumption will throttle honest peers on its first bad afternoon, and the first thing a volunteer network loses when that happens is its volunteers.

Operator surface

Declarative Manifests & State Reconciliation

An operator surface that takes thirty minutes to read or relies on custom imperative scripts fails before it has served a single packet. A node's whole operational surface is authored as YAML or JSON manifests, following the cloud-native pattern (apiVersion: fleetmesh.org/v1alpha1). The standalone meshctl CLI (built in Rust) manages them and reconciles them live against running meshnode daemons via the meshnode-control/1 Unix domain socket, executing non-destructive hot-reconciliation and multi-target compilation. The Python reference node consumes the same manifests once, at meshnode init --manifest.

The declarative manifest remains the single source of truth across all deployment environments. It deploys identically on Kubernetes, rootless Podman Quadlets, headless Raspberry Pis, or embedded mobile leaf runtimes.

Core Resource Kinds

KindAPI VersionRole & Responsibility
MeshNode fleetmesh.org/v1alpha1 Declares node class (anchor, edge, leaf), DID identity, listen endpoints (Iroh, libp2p, WSS), storage backend, telemetry, and operator attribution.
Peering fleetmesh.org/v1alpha1 Declares bilateral peering relationships (static peer DID + endpoints or dynamic discovery filters), transport preference, ingress/egress filter seams, and token bucket budgets.
GrantPolicy fleetmesh.org/v1alpha1 Declares admission rules, tier capacity ceilings, AIMD recovery schedules, window durations, and NIP-44 blinded commitment secrets.
ComputeProvider fleetmesh.org/v1alpha1 Declares WASM compute executor parameters, runtime engine (wasmtime, wasmer), fuel pricing (msat_per_mfuel), max memory, and Lightning payout endpoints.
InferenceRouter fleetmesh.org/v1alpha1 Declares AI/LLM inference routing adapters (routstr, nip90, ollama, vllm), model endpoints, token pricing rates, Cashu mints, and Lightning payout targets.

Declarative Manifest Examples

MeshNode Manifest (meshnode.yaml)

apiVersion: fleetmesh.org/v1alpha1
kind: MeshNode
metadata:
  name: bristol-anchor-01
  labels:
    region: eu-west
    env: production
spec:
  class: anchor
  identity:
    did: "did:nostr:d4e287a91176b6a0ff2ff2384a6c8e5473f309a47ef0e854d92305574581eb08"
    keyAgreement: "fec01c7d23a4b9180fa49c30f40d99ef87b3a98c0b2efd1487ea0b240398f498c"
  operator:
    # Distinct from the node DID above. Layer 1 of the independence model detects
    # shadow nodes by comparing operator pubkeys; reusing the node key defeats it.
    pubkey: "9b21c0f4e8a17d3562bc04ea7f19d8c05e3a6b7419fd28ce03b5a6142d7e08f3"
    policyUrl: "https://relay.example.org/policy.txt"
  endpoints:
    - transport: iroh
      ticket: "iroh_ticket_node_7f8a9b"
      priority: 10
    - transport: libp2p
      address: "/dnsaddr/relay.example.org/p2p/12D3KooWDpJ7As7BWAwRMfu1VU2WCqNjvq387JEYKDBj4kx6nXTN"
      priority: 20
    - transport: wss
      address: "wss://relay.example.org"
      priority: 40
  storage:
    engine: lmdb
    path: "/var/lib/fleetmesh/data"
    maxBytes: 107374182400  # 100 GiB
  capabilities:
    - relay
    - blossom
    - mediator
    - negentropy

Peering Manifest (peering-bristol.yaml) — Symmetrical Ingress & Egress

apiVersion: fleetmesh.org/v1alpha1
kind: Peering
metadata:
  name: bristol-community-sync
spec:
  peer: "did:nostr:e5f398b02287c7b1003003495b7d9f65840410b58f0f1965e03416685692fc19"
  transport: iroh
  ingress:
    filter:
      kinds: [0, 1, 3, 7, 10002]
      "#t": ["bristol"]
    plugin: "/usr/local/bin/mesh-filter-in"
    maxEventsPerSec: 100
    maxBytesPerSec: 2097152
  egress:
    filter:
      kinds: [0, 1, 3, 7, 10002]
      "#t": ["bristol"]
    plugin: "/usr/local/bin/mesh-filter-out"
    maxEventsPerSec: 100
    maxBytesPerSec: 2097152
  budget:
    window: 3600
    events_in: 3000
    bytes_in: 33554432
    bytes_out: 268435456

InferenceRouter Manifest (inferencerouter-routstr.yaml) — AI Inference

apiVersion: fleetmesh.org/v1alpha1
kind: InferenceRouter
metadata:
  name: routstr-gpu-bridge
spec:
  provider: routstr
  endpoint: "http://127.0.0.1:8000/v1"
  models:
    - id: "llama-3.3-70b-instruct"
      rateMsatPerToken: 12
      contextWindow: 131072
    - id: "deepseek-r1"
      rateMsatPerToken: 18
      contextWindow: 65536
  payment:
    cashuMint: "https://mint.minibits.cash/Bitcoin"
    bolt11: true

Hot-Reconciliation & Non-Destructive Semantics

A central flaw in traditional relay configurations is that changing a filter drops all active connections. meshctl enforces non-destructive hot-reconciliation:

meshctl apply -f <manifest>
Diffs the submitted YAML/JSON manifest against the active runtime daemon. If an interest filter narrows or expands, the node sends updated NIP-77 negentropy subscription ranges over the live QUIC/TLS session without dropping the connection, resetting receipt counters, or interrupting unrelated live peerings.
Graceful Peering Teardown
Deleting a Peering resource executes a graceful teardown: the node emits a DIDComm peering/1.0/terminate message, issues final budget/1.0/receipt accounting tallies, flushes any pending gift-wrap queues, and closes the channel cleanly.
Multi-Target Export Drivers
The meshctl export command compiles manifests into native deployment configuration. --target=k8s writes Kubernetes Deployments, Services and ConfigMaps. --target=quadlet writes Podman container units. --target=systemd writes systemd service units.