Fleetmesh fleetmesh

field guide · meshnode · § Implementation · § Operator surface · § Conformance

Running a Node and Operator Manual

How to deploy, configure, peer, and monitor a meshnode instance. This guide covers service installation, Caddy TLS reverse proxying, telemetry with meshwatch, the 40-subcommand CLI reference, declarative manifests, and diagnostic playbooks.

Operator Summary

A mesh node extends a standard Nostr relay with bilateral budgets, receipt accounting, and deterministic peering.

The reference implementation in node/ runs as a Python systemd service, requiring Python 3.12 and the cryptography package. Persistent state resides in three SQLite databases. Operator commands communicate through a local UNIX control socket.

Alpha Version Contract

The protocol version string is 0-draft. Schema digests may update during alpha; nodes with differing digests still peer on the protocols they share (same major, minor at or above the floor), and meshnode peers shows which peers were built from another schema. Joining the mesh preserves existing stored data, and unpeering takes effect immediately through local configuration. See § Alpha contract.

Deployment

Node setup in four commands

Clone the repository into your preferred deployment directory (the default service unit targets /opt/fleetmesh). Python 3.12 and pip install cryptography are the only prerequisites.

# 1. Initialize node directory and identity
python3 -m node.meshnode init /var/lib/meshnode --class edge \
    --operator <operator-pubkey> --relay wss://relay.example.org --face-port 7777

# 2. Add an initial peering configuration
python3 -m node.meshnode peer /var/lib/meshnode <peer-pubkey> \
    --shard bristol --filter '{"kinds":[1],"#t":["bristol"]}'

# 3. Install the systemd service unit
install -m 0644 node/meshnode.service /etc/systemd/system/meshnode.service

# 4. Enable and start the daemon
systemctl enable --now meshnode

During init, the node generates identity keys and mines a kind 11801 descriptor with 20 bits of proof-of-work (typically 1–15 seconds on a standard cloud vCPU). The descriptor is cached on disk and reused across restarts unless endpoints or relays change.

The node directory contains:

Running Multiple Nodes on One Host

Use the template service unit node/meshnode@.service to run separate instances on different ports. Initialize a second directory (e.g. /var/lib/meshnode-b on port 7778), then run systemctl enable --now meshnode@meshnode-b. Each instance maintains its own independent keys, databases, face port, and control socket.

Network Architecture

Public relay face and TLS reverse proxy

A node without public inbound ports remains fully reachable over the Nostr transport via gift-wrapped DIDComm messages routed through mutual public relays. To serve external clients over WebSocket, bind the relay face to localhost and place a TLS reverse proxy (such as Caddy) in front:

# /etc/caddy/Caddyfile
relay.example.org {
    reverse_proxy 127.0.0.1:7777
}

Declare the public URL in your node config so peers can dial the relay endpoint directly:

"endpoints": [{"transport": "wss", "address": "wss://relay.example.org", "priority": 40}]

When endpoints change, the node automatically mines an updated descriptor. Peers discover new endpoints by comparing the endpoints digest in periodic heartbeats.

Operator Interface

Command reference (the 40 subcommands)

meshnode provides forty subcommands designed for automation, lifecycle management, bilateral peering, compute routing, transport diagnostics, and settlement. Commands reading state operate directly against local SQLite files (and query the live daemon when running). Configuration commands write to config.json, which the daemon reloads on its next tick. Action commands communicate with the running daemon over control.sock. The table is held to the parser by scripts/check-cli.py, which fails the build when a subcommand is added without a row here.

Operator commands for node inspection and control
Command Type Description
initLifecycleGenerates node keys, creates config.json, mines PoW, and publishes the initial descriptor. Supports --manifest and --peering.
runLifecycleStarts the node daemon, binds the relay face, opens control.sock, and ticks on cadence.
retire [--revoke]LifecycleLeaves the mesh in order: hands over asserted shards, withdraws claims, stops, and optionally publishes kind 11802 key revocation.
status [--json]InspectionIdentity, face port, stored events, active grants, pending changes, and last evaluated windows.
health [--json]DiagnosticsRelay connection status, drop counts, refused sessions, handler exceptions, and active throttles.
capacity [--json]DiagnosticsConfigured aggregate ceilings, current pressure, engaged load-shedding rungs, and historical peak.
subsystems [--json]DiagnosticsPer-compartment status across all 20 core compartments and whichever of the 6 optional ones (p2p, gossip, compute, bridge, dvm, updates) the config enables, including compartment refusals.
refusals [--last N]DiagnosticsStructured log of every refused request and transaction, with specific failure reasons.
events [--since N] [--topics ...] [--last N]DiagnosticsThe most recent event-bus envelopes the daemon still holds in memory (windows evaluated and repriced, sessions completed, grants issued, peers connected, reports read), with sequence numbers to resume from.
subscribe [--since N] [--topics ...]DiagnosticsHold the control socket open and print one line per bus event as it happens; --since replays what the ring still holds first, and a slow reader is told how many envelopes it missed. The surface an out-of-process plugin follows.
strain [--json]DiagnosticsSurvived environmental anomalies by direction, reserve multiples, and dominant share flag-tree warnings.
whoami [--json]InspectionDID, pubkey, npub, encryption keys, declared endpoints, and descriptor PoW status.
check [--json]VerificationValidates directory permissions (0600), config schema, descriptor digest, and SQLite database integrity.
peers [--json]PeeringAll peered nodes, active tiers in both directions, clean window streaks, and shard pulls.
grant <peer> [--json]AccountingInspects bilateral grants with a specific peer: limits, cadence, scope, and envelope history.
windows <peer> [--last N]AccountingEvaluated windows for a peer: measured vs. claimed counters, receipts sent/received, and severity rating.
scheduled [--json]InspectionDisplays every scheduled pull, when it last ran, active throttles, capacity holds, and declines.
descriptor [--republish]InspectionDisplays the local node kind 11801 descriptor event; --republish resends it.
peer <peer> --shard NAME --filter JSONConfigAdds a shard pull from a peer to config.json, on the operator's cadence. The daemon reconciles state on the next tick.
unpeer <peer> [--shard NAME]ConfigRemoves the pull from config.json. The grants stay live; terminate ends the peering.
trust [peer] [--tier]Config / InspectionInspects active trust decisions across peers, or seeds an operator trust tier in config.json.
coverage list|add|removeConfigManages shard coverage claims published across the mesh as kind 30803 events.
pull <peer> <shard>ActionTriggers an immediate Negentropy synchronization cycle with a designated peer.
terminate <peer>ActionCleanly terminates an active peering: zeroes both grants and notifies the peer.
withdraw <peer> <window>ActionWithdraws an adverse finding or report about a specific window epoch on the complaint thread.
standing <peer> <target>ActionQueries the peer local standing ledger about a third node over standing/1.0.
reports <target>ActionFetches kind 30802 dispute reports about a node from relays, weighted by the local node ledger.
tick [--json]ActionForces an immediate scheduler evaluation and housekeeping tick.
api [--json]InspectionIntrospects the control socket protocol, available commands, and dispatch schema.
watch [--every N] [--quiet]TelemetryRuns automated health and invariant checks, returning structured exit codes for monitoring agents.
bridge [action] [--relay URL] [--event JSON] [--filter JSON]BridgeInteracts with the Universal Nostr Relay Bridge: status, publish, add-relay, remove-relay, sync, and range-bounded reconcile backfill.
gossip [action] [--name N] [--member PK]... [--ttl S] [--topic ID] [--event JSON|@file]Gossipgossip/1.0 group topics: status, topics, create (signs and publishes a kind 30804 roster for <node pk>:<name>; --no-relay keeps it local), admit another admin's roster that names this node (a roster published to the relays by a peer this node holds a grant with is admitted on its own; admit is for one handed over by other means), leave, and publish a sealed NIP-59 wrap into a topic mesh. GET /gossip on the relay face is the read-only view of the same rosters.
compute [action] [--engine] [--rate] [--ask]ComputeInspects or interacts with the compute market and order book: status, show, book, offer, match, and receipts.
dvm [action] [--request JSON]ComputeInteracts with the NIP-90 Data Vending Machine Gateway: status, show telemetry, process, and delegate compute requests.
cluster [action] [--nodes N] [--topology T]ClusterManages a local multi-node testnet cluster: start, status, stop, and sync across mesh, star, ring, or line topologies.
settle [action] [--ask ID] [--amount N] [--rail R]SettlementInspects settlement rails and escrows: status, create, hold, settle, nwc-pay, cashu-swap, and cashu-melt.
strfry [action] [--url URL] [--event JSON]BridgeThe strfry relay bridge: status, the write-policy plugin (with --enforce-grants and --shadow), eval of one event against it, and export / import of JSONL over the pipe.
enclave [action] [--type p256|secp256k1] [--message M] [--sig S] [--pub P]TransportHardware enclaves and passkeys: status, generate a key, sign a message, verify a signature against a multikey. Exits 1 on a verification that fails.
iroh [action] [--host H] [--port P] [--ticket T] [--to PK] [--message M]Transportiroh QUIC transport diagnostics: status, probe and punch a remote endpoint, send a payload to a ticket or pubkey.
fips [action] [--sock PATH] [--port P] [--peer PK] [--message M]TransportFIPS Native Datagram API control: status, probe the control socket, listen on an FSP port, connect a flow to a peer, send a framed datagram, and list flows.

Declarative Manifests

Kubernetes-style manifests and non-destructive reconciliation

For fleet operations, CI/CD pipelines, and cloud-native deployments, Fleetmesh provides a declarative manifest suite modeled after Kubernetes resources under apiVersion: fleetmesh.org/v1alpha1. Operators declare desired state in version-controlled YAML or JSON files:

Core Declarative Resource Kinds
Kind API Version Operational Role
MeshNodefleetmesh.org/v1alpha1Declares node class (anchor/edge/leaf), identity DID, transport listen endpoints (Iroh, libp2p, WSS), storage paths, and operator attribution.
Peeringfleetmesh.org/v1alpha1Declares bilateral peering relationships, preferred transport, ingress/egress topic shard filters, and token bucket budgets.
GrantPolicyfleetmesh.org/v1alpha1Declares admission ceilings across tiers (probation, member, trusted, anchor), AIMD parameters, and NIP-44 blinded secrets.
ComputeProviderfleetmesh.org/v1alpha1Declares WASM / ZK compute executor pools, fuel limits, msat pricing rates, and Lightning payout endpoints.
InferenceRouterfleetmesh.org/v1alpha1Declares AI/LLM routing adapters (vLLM, Ollama, Routstr), model IDs, context windows, and Cashu/Lightning payout targets.

Fleetmesh provides two complementary software products: the meshnode sovereign daemon and the standalone meshctl declarative fleet management CLI (built in Rust).

Fleet Management with meshctl

meshctl provides declarative, non-destructive hot-reconciliation against running meshnode daemons via the meshnode-control/1 Unix domain socket. Applying a manifest dynamically updates peering filters, rate limits, and budget allocations without dropping live QUIC/TLS sessions or resetting receipt hash chains.

# Apply manifests with non-destructive live hot-reconciliation:
meshctl apply -f manifests/peering-bristol.yaml

# Dry-run validation to preview reconciliation deltas:
meshctl apply -f manifests/ --dry-run

# Diff manifest against live daemon state:
meshctl diff -f manifests/peering-bristol.yaml

# Inspect node and fleet health via the control socket:
meshctl status --dir /var/lib/fleetmesh

# Compile manifests to production deployment targets:
meshctl export -f manifests/meshnode-anchor.yaml --target k8s > k8s-node.yaml
meshctl export -f manifests/meshnode-anchor.yaml --target quadlet > ~/.config/containers/systemd/node.container
meshctl export -f manifests/meshnode-anchor.yaml --target systemd > /etc/systemd/system/meshnode.service
meshctl export -f manifests/meshnode-anchor.yaml --target compose > docker-compose.yml

# Convert legacy fleetmesh.org/v1alpha1 manifests to fleetmesh.org/v1alpha1:
meshctl convert -f legacy-manifest.yaml -o fleetmesh.yaml

# Node-local fallback init:
meshnode init --manifest manifests/meshnode-anchor.yaml --peering manifests/peering-bristol.yaml
meshnode run

Automated Monitoring

The watchdog daemon (meshwatch)

The watcher daemon (node/watch.py) runs on a systemd timer (meshwatch.timer) to execute recurring health audits against the running node. It uses clean process exit codes for seamless integration with Prometheus, Datadog, or nagios:

# Inspect health status via the CLI
python3 -m node.meshnode watch /var/lib/meshnode

Accounting In Practice

Reading window evaluations

Bilateral accountability centers on the window evaluation log. Inspect it using meshnode windows <peer>:

windows with f22c163a211c50e3  (measured/claimed, in the issuer frame)
window  closed UTC        severity  req  events_in  bytes_in      receipts in  sent
496817  2026-09-04 18:00  S0        0/-  37/37      45275/45275   2            2
496818  2026-09-04 19:00  S0        0/-  12/12      14820/14820   1            1

Key metrics to observe:

Reliability Principles

Three architectural invariants

INVARIANT 01

Environmental anomalies are treated as routine network conditions.

Peer restarts, transient relay disconnections, clock skew within tolerance, and throttled peers represent standard operating conditions. The protocol classifies timeouts, silence, and unfulfilled sync queries as network partitions.

Rule: Incomplete responses indicate network partitions.

INVARIANT 02

State is persisted to disk before network acknowledgment.

Keys, configuration, descriptors, grant states, window meters, and receipt hash chains are committed to SQLite before wire confirmation. A node restarting mid-window recovers exact meter counters.

Rule: State is persistent before public transmission.

INVARIANT 03

Failure domains are compartmentalized.

If the sync engine encounters an oversized envelope or a relay disconnects, the peering, transport, and accounting subsystems continue operating. Refusal logs record structured telemetry.

Rule: Compartment errors remain isolated.

Diagnostics

Troubleshooting checklist

A field guide to § Implementation, § Operator surface, and § Conformance. All commands and data structures reflect the active reference implementation in node/.