Operator manual · pre-production
The Musivum & Tessella manual
Two planes: a proprietary orchestrator that decides who computes what, and an open-source node that computes it. This manual is written from the code that is running today — including the parts that are missing. The site is live; the final end-to-end tests are still pending, and the project opens to the public when they close.
01 The three roles, and what you run
The mesh has one orchestrator, many nodes, and any number of clients. Only the middle one is something you install.
- Musivum — orchestrator (proprietary)
- Admits nodes, computes the layer split from the VRAM each node claims, mints a signed WorkTicket per session, keeps the credits ledger and the reputation counters. It never receives a tensor: prompt and hidden states travel node to node.
- tessella-worker — compute node (open source, Apache-2.0)
- Holds a contiguous range of whole layers of the model, prefills, decodes, and serves the hop endpoints. Every request it makes to the orchestrator is Ed25519-signed with its own peer identity. Package tessella-node, binary tessella-worker, one node per GPU.
- Client — anything that speaks HTTP
- Finds an entry node with GET /peers, then POST /v1/chat/completions. The bundled tessella-cli is one; an OpenAI-style script is another.
Everything you install and read is in the open repository. The orchestrator is run by the Musivum team — today that is the development tracker this manual points at — or by you with the official image, which is announced but not published yet. Your nodes only ever need its URL.
02 Install a node and register it
One binary, one config file, one key. The wizard writes the first and generates the second.
-
Build it. Rust 1.75+ and a CUDA toolkit for the GPU path — CUDA is the default
feature, so the plain build is the GPU build; Apple Silicon asks for Metal
explicitly.
git clone https://github.com/vitichenko/tessella cd tessella cargo build --release # Linux + NVIDIA cargo build --release --no-default-features --features metal # Apple Silicon # artifact: target/release/tessella-worker -
Answer the wizard. Started with no config file, the node asks for the essentials and
queries the orchestrator for the models it is serving
(GET /api/network/models), so you choose from what actually has
workers behind it:
Prompts: orchestrator URL, model id, VRAM claim in MB, weights directory, keypair path, API port (8080), P2P port (9080).export HF_TOKEN="hf_…" ./target/release/tessella-worker --config worker_config.toml -
Read the identity you just created. The peer id is derived from the Ed25519 key and is
what the orchestrator, the ledger and the reputation table all key on:
./target/release/tessella-worker keys show # write it down: a lost key is a lost identity — and its credit balance -
Join a mesh. The flag, the environment variable and the config key mean the same
thing:
./target/release/tessella-worker \ --tracker-url https://tracker1.musivum.dev \ --model-id Qwen/Qwen2.5-0.5B-Instruct \ --vram-gb 2.0 # --musivum-url is the older alias of --tracker-url -
Confirm from outside. Your node publishes its shard; the orchestrator publishes the
mesh:
curl -s localhost:8080/status | jq '{model, layer_ranges, gpu}' curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq length
The config file
Written by the wizard, editable by hand, overridable by environment. This is the shape the node accepts — the comment on each key is the rule the code enforces:
# Global identity and orchestrator
orchestrator_url = "https://tracker1.musivum.dev"
keypair_path = "/home/operator/.config/tessella/node.key"
default_model = "Qwen/Qwen2.5-0.5B-Instruct"
[worker]
vram_mb = 2048 # the claim the layer split is computed from
models_dir = "./models"
api_host = "0.0.0.0"
api_port = 8080
p2p_host = "0.0.0.0"
p2p_port = 9080
device = "cuda" # "cuda" | "metal" — there is no CPU path
[gpu]
oom_streak = 3 # consecutive OOMs that mean "saturated"
streak_window_secs = 30
The optional keys — public_ip, public_api_port, public_p2p_port, public_api_url, max_input_tokens, p2p_inference_timeout_secs, quantization, canary_mode, model_cache_prune, kademlia_enabled, registration_grace_secs, wal_checkpoint_interval_secs, allow_private_addrs, callback_allowed_hosts — each have environment equivalents (TESSELLA_PUBLIC_IP, TESSELLA_VRAM_MB, TESSELLA_QUANTIZATION, …) tabulated in the repository's docs/HOWTO.md.
The orchestrator URL is read from ORCHESTRATOR_URL (or its older alias TRACKER_URL), from the --tracker-url flag — canonical, with --musivum-url accepted as its alias — or from orchestrator_url in the config. MUSIVUM_URL appears in older guides and in the Docker examples but is not read by the node: set it alone and your node still points at 127.0.0.1:8080.
03 Claims, shards and the handshake
You declare VRAM; the orchestrator answers with layers. There is no telemetry negotiation — the claim in your config is your shard, so it can be honoured or refused, but never silently adjusted.
- Register. POST /register with a signed WorkerRegistration: peer id, P2P address, api_addr, VRAM claim, device, model id, total layers, the weight manifest and its parameter-contract hash, plus the Ed25519 public key and signature. Admission is evaluated before the signature, so a private mesh refuses an unlisted peer without spending a verification.
-
Receive a range, or a refusal. The answer is the assignment the node must serve plus
every peer it may dial as the next hop of the chain:
A claim that cannot host a single layer is 409; a private mesh with a peer outside the allow-list is 403 worker_not_authorized; a malformed body, a manifest the node does not own or a private address with allow_private_addrs = false is 400.{ "assignment": { "model_repo": "Qwen/Qwen2.5-0.5B-Instruct", "layer_ranges": [[0, 8]], "replica_group_id": 0, "expert_range": null }, "peers": [ { "peer_id": "12D3KooW…", "layer_ranges": [[0, 8]], "http_addr": "203.0.113.7", "queue_depth": 0, "canary": false } ] } - Load, then say so. While shards download and the model warms up the node is PendingLoad: it holds a stretch it cannot serve yet, and the coverage walk counts that stretch as a hole. POST /ready is what flips it, and from then on /peers hands it out to clients.
- Heartbeat. POST /announce refreshes liveness, role, queue depth and health. An announce that changes a node's status republishes the topology, which is how the mesh converges without anyone polling it.
- Leave cleanly. POST /deregister. A peer that never comes back is marked Offline and its stretch becomes a hole for the next node to repair.
Sizing a shard without guessing
GET /model_health?model_id=…&device=cuda publishes exactly what a repair needs: the holes nobody serves, the prefix that renders, and the allocator's own budget terms (mem_per_layer_mb, weights_per_layer_mb, kv_per_layer_mb, global_tensor_mb, usable_vram_pct, min_vram_mb). Joining a hole is then a division:
claim_mb = ceil((layers_of_the_hole * mem_per_layer_mb + global_tensor_mb) * 100 / usable_vram_pct)
global_tensor_mb is also why the split is uneven where it counts: the range that starts at layer 0 also loads the token embedding, and the last range loads the lm_head, so a card that serves a middle range comfortably can be refused for the ends. A mesh is anchored by one node whose claim has a real VRAM budget behind it — the seed node.
Measured on Qwen/Qwen2.5-0.5B-Instruct: a 24-layer model where a layer costs ~56 MB on CUDA (40 MB weights + 16 MB KV) against ~89 MB on Metal — F16 against F32. The hard floor is 256 MB, and the operator's min_vram_gb raises it from there. Health is one of three words, and the words are the allocator's own walk: healthy — a ready chain renders the whole model; degraded — a gap with shards above it (held but unreachable); incomplete — the walk runs out before the last layer, which is also what a mesh that lost its tail node answers.
04 How a token crosses the mesh
The client talks to one node — the entry node. Everything after that is hops between peers, and the tokens come back the same way they left: signed, and through the chain.
-
Client → entry node. Discover a node for the model, then talk to it like any
completion server. Skip anything with "canary": true: those are
replicas that answer with a known-wrong tensor on purpose, to catch a node that copies its
neighbour.
curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" \ | jq -r '.[] | select(.canary == false and .status == "Healthy") | .http_api_url' | head -1 - Entry node prefills. It tokenizes the prompt and runs its own layers — the embedding plus its range. From here on the unit of transport is the hidden state, not text.
- Hop. The entry node POSTs to the next peer's /v1/forward with the hidden state, the tokens so far, its own entry_node_url and the requester's tier. Each node runs the layers it owns and forwards the result — a hop body is capped at 64 MiB, and each node publishes the body size it accepts.
- Tail → tokens. The last node detokenizes and pushes tokens back to {entry_node_url}/v1/token_stream/{request_id}. That callback is a capability, not a guess: it is authorised by the single-use signature the entry node put on the payload this shard processed, so only a node that saw this pipeline's traffic can write into the stream. The entry node relays them to the client as they arrive.
- KV cache stays local. Each node keeps the KV cache of the sessions it serves, so a hop only carries the increment — not the whole context. That is the reason the chain is cheap to keep alive and expensive only to build.
The three numbers of the deal
Before sending a long prompt, read GET /status on the entry node. It publishes the ceiling it accepts, the time it waits for a hop, and how many hops it reads at once — a client that reads only the first one sizes a prompt the node will then declare dead:
- max_input_tokens
- tokens this node can prefill in one request, measured on its own VRAM, with max_input_tokens_source telling you whether it was declared or measured
- hop_timeout_secs
- how long a hop of this pipeline may take (default 10 s, p2p_inference_timeout_secs)
- max_concurrent_hop_bodies / hop_bodies_in_flight
- the quota and the live count. At the top, the node refuses with 503 too_many_hop_bodies — that is a full node, not a broken one, and retrying elsewhere pushes 70–80 MiB at every candidate
When a hop fails, the upstream node records it in peer_failures, reports it with POST /report_failure, and the orchestrator can re-shard the stretch. A node that answers with the wrong tensor is caught by sampled layer audits and canary replicas — not by trusting its own logs.
05 Your node's HTTP API
Eight routes on api_port. Four of them are the wire — the hop and the completion surfaces; the rest are for you and the orchestrator.
- POST /v1/forward
- the hop: hidden state in, hidden state out. This is what peer nodes dial, and the only route that carries the whole context in one body.
- POST /v1/chat/completions
- client surface — messages in, tokens out (streaming supported)
- POST /v1/generate
- client surface — raw prompt completion
- POST /v1/fim/completions
- fill-in-the-middle for code: prefix + suffix in, completion out
- POST /v1/token_stream/{request_id}
- callback the tail node pushes tokens into (single-use capability, rate limited)
- POST /api/v1/update_assignment
- how a new layer split reaches the running node when the mesh is re-sharded
- GET /status
- model, layer_ranges, assignment, hidden size, the three hop numbers, peer failure counters, GPU health
- GET /metrics
- counters for your own monitoring
curl -s localhost:8080/status | jq
curl -s localhost:8080/v1/chat/completions -H 'content-type: application/json' -d '{
"model": "Qwen/Qwen2.5-0.5B-Instruct",
"messages": [{ "role": "user", "content": "Hello!" }],
"max_tokens": 64, "temperature": 0.0
}'
curl -s localhost:8080/v1/fim/completions -H 'content-type: application/json' -d '{
"prefix": "fn main() {", "suffix": "}", "max_tokens": 128
}'
06 Credits, tickets and proofs
The ledger pays for work that can be pointed at. Everything else — your uptime, your good intentions, your hardware — pays nothing.
- A ticket authorises. When a session starts, the orchestrator mints a signed WorkTicket with a one-hour TTL and hands it to the requester in the same round-trip that grants the session — so authorising the payment adds no latency to the request itself.
- A proof settles. Each node writes a signed WorkProof into its own SQLite WAL before uploading it: a node that dies mid-upload still gets paid for what it served. Settlement is authorised by the proof itself — signed by the worker and by the orchestrator that minted its ticket — never by the requester's word.
- One credit, one node. A credit is paid once per (request_id, worker_id), and only for the layers the worker is actually assigned: a replay of the same ticket pays nothing the second time.
- Your tier is a priority, not a wallet. The balance of the identity behind the key decides the queue tier of its requests: > 1000 → VIP, > 0 → Standard, ≤ 0 → Leech. A busy mesh serves VIP first; a quiet one serves everyone.
With economy_enabled = false every requester is treated as VIP — the ledger keeps recording, the tier stops mattering. In deployment_mode = "local" nobody is privileged: a score cannot order a LAN of trust, so every requester gets the mesh baseline.
Cheating is priced rather than forbidden: sampled layer audits compare what a node returns against a recomputation, canary replicas answer correctly-but-wrongly so a copier is provably wrong, and the orchestrator keeps the blunt instruments — /slash, blacklisting, reputation counters.
07 The orchestrator API
Three tiers: what anyone may read, what a node must sign, and what only the operator may touch. The full reference, with measured bodies, ships in the Musivum repository (docs/API_REFERENCE.md).
Public — discovery and routing
- GET /peers?model_id=<repo>
- the routing contract: every node of a model with its dialable endpoints, layer ranges, health counters and canary flag. A model nobody serves answers [], never a 404
- GET /api/network/models
- what the mesh can render: [{ "model_id": …, "active_workers": n }]
- GET /workers
- monitoring view — peer id, model, status, VRAM claim and queue depth, with no addresses in it
- GET /model_health?model_id=&device=
- coverage, holes, and the memory terms a repair is sized with
- GET /api/v1/assignment?model_id=&vram_mb=&device=
- the range a claim would get — resolved before you install anything
- GET /weights/manifest & /weights/manifests
- the manifest a node must own to serve the model, and the manifests the mesh knows
- GET /p2p_id
- the mesh's swarm identity, so a node can pin the key it dials
- GET /federation/status & GET /metrics
- federation cursor state and operational counters
Worker-signed — the write path
- POST /register
- join the topology with a claim; this is where the layer split is computed
- POST /announce
- heartbeat: liveness, role, queue depth, health
- POST /ready
- shards loaded — I serve now
- POST /deregister
- leave cleanly
- GET /api/v1/session_grant
- the session's QoS ticket and its work ticket, in one round-trip
- POST /api/v1/audit_report
- what a sampled audit saw
- POST /work_done & POST /api/v1/work_done_batch
- settlement, authorised by the proof each request carries
- POST /report_failure
- a hop or a peer that failed, so the mesh can act on it
Operator — needs admin_token
- GET /ledger?limit= & GET /ledger/snapshot?since=&limit=
- the credit ledger and its cursor-paged snapshot — how federation syncs without a full dump
- GET /credits?peer_id= & GET /reputation?peer_id=
- one identity's balance and score (or ?all=true for the table)
- POST /slash
- penalise an identity that was caught
- POST /rebalance
- recompute the split and re-assign the mesh
- POST /weights/manifest
- publish the manifest a model must match
The public one answers at https://tracker1.musivum.dev (http://34.171.244.142:8080 directly). It is the team's development deployment: the read endpoints are open, and anything that writes needs a node key or the admin token. Try it before you build anything —
curl -s https://tracker1.musivum.dev/api/network/models
curl -s "https://tracker1.musivum.dev/model_health?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq '{health, covered_layers, holes}'
If /api/network/models answers [], the mesh you are about to join is empty — your node will be the first tile.
08 Chatting as a client
Two ways in: the bundled CLI, which discovers a node for you, or any client that speaks the completion routes.
tessella-cli
# a node's own API is the default target (http://127.0.0.1:8080)
cargo run -p tessella-cli --release -- chat "hello from the swarm"
# point it at any entry node you discovered in /peers
cargo run -p tessella-cli --release -- \
--api-url http://203.0.113.7:8080 chat "hello from the swarm"
# interactive, with a session file that keeps the conversation's KV cache alive
cargo run -p tessella-cli --release -- \
--api-url http://203.0.113.7:8080 chat --interactive --session-file ./session.json
# fill-in-the-middle, and a per-request budget
cargo run -p tessella-cli --release -- -a http://203.0.113.7:8080 fim \
--prefix "fn main() {" --suffix "}" --max-tokens 128
--api-url (short -a) and --model-id are top-level arguments: they come before the subcommand. chat takes -p/--prompt or trailing words, -m/--max-tokens (512 by default), --temperature (0.0 by default), -i/--interactive, --session-file and --timing; --mode switches it between chat and raw completion, and --prefix/--suffix do the same for fill-in-the-middle without leaving the subcommand. tessella-cli init writes the config and the client keys, and -c/--config points at a different file.
Any OpenAI-style client
The entry node's completion routes are the whole contract — no SDK, no handshake, no session token to fetch first:
curl -s http://203.0.113.7:8080/v1/chat/completions \
-H 'content-type: application/json' -d '{
"model": "Qwen/Qwen2.5-0.5B-Instruct",
"messages": [{ "role": "user", "content": "Explain a mosaic." }],
"max_tokens": 128, "stream": true
}'
stream: true is what makes the pipeline legible: tokens arrive from the tail node, hop by hop, as the entry node relays them.
09 Operations: GPU, proxies, systemd
Everything an operator actually has to decide: how the node starts, what it announces, where it keeps weights, and how much of the GPU it may take.
Reaching a node that sits behind NAT
- TESSELLA_PUBLIC_API_URL
- proxy mode: the full public URL peers should dial, e.g. https://pod-8080.proxy.runpod.net. Anything with an https:// scheme and no port is treated as a tunnel
- TESSELLA_PUBLIC_IP + TESSELLA_PUBLIC_API_PORT + TESSELLA_PUBLIC_P2P_PORT
- TCP mode: what a NAT rule maps to your private ports. Publish the API port and the P2P port — with only the API, hops work but discovery does not
- TESSELLA_ALLOW_PRIVATE_ADDRS
- set this only when the callback may legitimately target a private address (otherwise it is refused: an unvalidated callback target is an SSRF primitive)
Running it under systemd
[Unit]
Description=Tessella compute node
After=network-online.target
[Service]
User=operator
WorkingDirectory=/opt/tessella
EnvironmentFile=/etc/tessella/node.env # HF_TOKEN=…, ORCHESTRATOR_URL=…
ExecStart=/opt/tessella/tessella-worker --config /etc/tessella/worker_config.toml
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
The trade-offs you can turn
- --model-cache-prune off|dry-run|on
- whether the node deletes the shards its new assignment no longer needs. Run dry-run first: it reports what it would delete and deletes nothing
- --clean-cache
- wipe the local shard cache and re-download — the honest way out of a corrupted cache
- quantization = "Q8_0"
- 8-bit weights in memory instead of the default unquantized layout: a smaller claim, a slightly different forward
- canary_mode = "off" | "echo" | "scale"
- operator tool: the node answers with a known-wrong tensor instead of computing, which is how the mesh proves its audits catch a liar
- kademlia_enabled
- join the DHT or stay with the orchestrator's membership only
- max_input_tokens
- declare the prefill ceiling you can hold instead of letting the node measure it — clients read it before sizing a prompt
- oom_streak / streak_window_secs
- how many consecutive OOMs, inside how many seconds, mean "saturated" rather than "broken"
In a container, mount the weights directory and the key outside the image or you will re-download a model and lose your identity on every restart:
docker run -d --name tessella --gpus all \
-p 8080:8080 -p 9080:9080 \
-e HF_TOKEN="hf_…" \
-e ORCHESTRATOR_URL="https://tracker1.musivum.dev" \
-v /srv/tessella/models:/app/models \
-v /srv/tessella/key:/home/tessella/.config/tessella \
tessella-worker:latest
10 Troubleshooting & limits
Nine of every ten problems are one of these. The tenth is a bug worth reporting with the /status body attached.
- Node runs but /peers never lists it
- it is PendingLoad — shards still downloading, or the model is gated and HF_TOKEN is missing. /status shows the assignment it already owns; /ready is what publishes it
- POST /register → 409
- the claim cannot host a single layer of that model. Size it with /model_health and remember the floor: 256 MB hard, the operator's min_vram_gb above it
- POST /register → 403 worker_not_authorized
- the mesh is private and this peer is not in the allow-list. Nothing to fix on the node
- POST /register → 401
- the signature did not verify: wrong key file, a copied config, or a badly skewed clock. Regenerating the key gives you a new identity — and a new credit balance
- 503 too_many_hop_bodies
- this node is reading its quota of hop bodies at once. It is full, not broken: retrying elsewhere pushes 70–80 MiB at every candidate, so wait and retry
- 413 on a long prompt
- the body exceeds the ceiling the node publishes. Read max_input_tokens from /status instead of discovering it by truncation
- Hops die at the same timeout every time
- one node in the chain is slower than the pipeline's patience. Raise p2p_inference_timeout_secs on the slow node, or lower the prompt it has to carry
- Requests are always low priority
- the identity is Leech: balance ≤ 0. Serve work — the tier follows the balance, not the intent
- CUDA out of memory under load
- the claim is above what the GPU can hold once sessions pile up. Lower vram_mb, or quantize to Q8_0. The node declares saturation after oom_streak consecutive OOMs and stops taking work instead of thrashing
- Answers are subtly wrong
- check whether you dialled a canary node: it returns a deliberately wrong tensor. Clients must skip "canary": true
What runs today, and what is still missing
- Whole-layer sharding, hops over HTTP, signed identity, ledger, audits
- Federation by signed cursor pull, and NAT-aware addressing
- Two-dimensional split (layers × experts): designed, parked until there is hardware to measure it — MoE meshes run as chains today
- Speculative decoding and INT4 weights
- A CPU path: the node is GPU-only, and CUDA is the default build
- Public-swarm hardening at scale: Sybil resistance and relays are the open problems, and the public orchestrator is still the team's development deployment
The node, the protocol and the CLI are one repository — github.com/vitichenko/tessella — and the operator notes live in docs/HOWTO.md and README_DEPLOY.md. It is private until the final end-to-end tests close, and it opens to the public with them. If this manual and the code disagree, the code is right and the manual is a bug.