musivum.dev

Operator manual · pre-production

The Musivum & Tessella manual

Two planes: a proprietary orchestrator that decides who computes what, and an open-source node that computes it. This manual is written from the code that is running today — including the parts that are missing. The site is live; the final end-to-end tests are still pending, and the project opens to the public when they close.

  1. The three roles, and what you run
  2. Install a node and register it
  3. Claims, shards and the handshake
  4. How a token crosses the mesh
  5. Your node's HTTP API
  6. Credits, tickets and proofs
  7. The orchestrator API
  8. Chatting as a client
  9. Operations: GPU, proxies, systemd
  10. Troubleshooting & limits

01 The three roles, and what you run

The mesh has one orchestrator, many nodes, and any number of clients. Only the middle one is something you install.

Musivum — orchestrator (proprietary)
Admits nodes, computes the layer split from the VRAM each node claims, mints a signed WorkTicket per session, keeps the credits ledger and the reputation counters. It never receives a tensor: prompt and hidden states travel node to node.
tessella-worker — compute node (open source, Apache-2.0)
Holds a contiguous range of whole layers of the model, prefills, decodes, and serves the hop endpoints. Every request it makes to the orchestrator is Ed25519-signed with its own peer identity. Package tessella-node, binary tessella-worker, one node per GPU.
Client — anything that speaks HTTP
Finds an entry node with GET /peers, then POST /v1/chat/completions. The bundled tessella-cli is one; an OpenAI-style script is another.
Two planes

Everything you install and read is in the open repository. The orchestrator is run by the Musivum team — today that is the development tracker this manual points at — or by you with the official image, which is announced but not published yet. Your nodes only ever need its URL.

02 Install a node and register it

One binary, one config file, one key. The wizard writes the first and generates the second.

  1. Build it. Rust 1.75+ and a CUDA toolkit for the GPU path — CUDA is the default feature, so the plain build is the GPU build; Apple Silicon asks for Metal explicitly.
    git clone https://github.com/vitichenko/tessella
    cd tessella
    cargo build --release                                      # Linux + NVIDIA
    cargo build --release --no-default-features --features metal  # Apple Silicon
    # artifact: target/release/tessella-worker
  2. Answer the wizard. Started with no config file, the node asks for the essentials and queries the orchestrator for the models it is serving (GET /api/network/models), so you choose from what actually has workers behind it:
    export HF_TOKEN="hf_…"
    ./target/release/tessella-worker --config worker_config.toml
    Prompts: orchestrator URL, model id, VRAM claim in MB, weights directory, keypair path, API port (8080), P2P port (9080).
  3. Read the identity you just created. The peer id is derived from the Ed25519 key and is what the orchestrator, the ledger and the reputation table all key on:
    ./target/release/tessella-worker keys show
    # write it down: a lost key is a lost identity — and its credit balance
  4. Join a mesh. The flag, the environment variable and the config key mean the same thing:
    ./target/release/tessella-worker \
      --tracker-url https://tracker1.musivum.dev \
      --model-id Qwen/Qwen2.5-0.5B-Instruct \
      --vram-gb 2.0
    # --musivum-url is the older alias of --tracker-url
  5. Confirm from outside. Your node publishes its shard; the orchestrator publishes the mesh:
    curl -s localhost:8080/status | jq '{model, layer_ranges, gpu}'
    curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq length

The config file

Written by the wizard, editable by hand, overridable by environment. This is the shape the node accepts — the comment on each key is the rule the code enforces:

# Global identity and orchestrator
orchestrator_url = "https://tracker1.musivum.dev"
keypair_path     = "/home/operator/.config/tessella/node.key"
default_model    = "Qwen/Qwen2.5-0.5B-Instruct"

[worker]
vram_mb    = 2048          # the claim the layer split is computed from
models_dir = "./models"
api_host   = "0.0.0.0"
api_port   = 8080
p2p_host   = "0.0.0.0"
p2p_port   = 9080
device     = "cuda"        # "cuda" | "metal" — there is no CPU path

[gpu]
oom_streak         = 3     # consecutive OOMs that mean "saturated"
streak_window_secs = 30

The optional keys — public_ip, public_api_port, public_p2p_port, public_api_url, max_input_tokens, p2p_inference_timeout_secs, quantization, canary_mode, model_cache_prune, kademlia_enabled, registration_grace_secs, wal_checkpoint_interval_secs, allow_private_addrs, callback_allowed_hosts — each have environment equivalents (TESSELLA_PUBLIC_IP, TESSELLA_VRAM_MB, TESSELLA_QUANTIZATION, …) tabulated in the repository's docs/HOWTO.md.

One name, not two

The orchestrator URL is read from ORCHESTRATOR_URL (or its older alias TRACKER_URL), from the --tracker-url flag — canonical, with --musivum-url accepted as its alias — or from orchestrator_url in the config. MUSIVUM_URL appears in older guides and in the Docker examples but is not read by the node: set it alone and your node still points at 127.0.0.1:8080.

03 Claims, shards and the handshake

You declare VRAM; the orchestrator answers with layers. There is no telemetry negotiation — the claim in your config is your shard, so it can be honoured or refused, but never silently adjusted.

  1. Register. POST /register with a signed WorkerRegistration: peer id, P2P address, api_addr, VRAM claim, device, model id, total layers, the weight manifest and its parameter-contract hash, plus the Ed25519 public key and signature. Admission is evaluated before the signature, so a private mesh refuses an unlisted peer without spending a verification.
  2. Receive a range, or a refusal. The answer is the assignment the node must serve plus every peer it may dial as the next hop of the chain:
    { "assignment": { "model_repo": "Qwen/Qwen2.5-0.5B-Instruct",
                      "layer_ranges": [[0, 8]], "replica_group_id": 0, "expert_range": null },
      "peers": [ { "peer_id": "12D3KooW…", "layer_ranges": [[0, 8]],
                   "http_addr": "203.0.113.7", "queue_depth": 0, "canary": false } ] }
    A claim that cannot host a single layer is 409; a private mesh with a peer outside the allow-list is 403 worker_not_authorized; a malformed body, a manifest the node does not own or a private address with allow_private_addrs = false is 400.
  3. Load, then say so. While shards download and the model warms up the node is PendingLoad: it holds a stretch it cannot serve yet, and the coverage walk counts that stretch as a hole. POST /ready is what flips it, and from then on /peers hands it out to clients.
  4. Heartbeat. POST /announce refreshes liveness, role, queue depth and health. An announce that changes a node's status republishes the topology, which is how the mesh converges without anyone polling it.
  5. Leave cleanly. POST /deregister. A peer that never comes back is marked Offline and its stretch becomes a hole for the next node to repair.

Sizing a shard without guessing

GET /model_health?model_id=…&device=cuda publishes exactly what a repair needs: the holes nobody serves, the prefix that renders, and the allocator's own budget terms (mem_per_layer_mb, weights_per_layer_mb, kv_per_layer_mb, global_tensor_mb, usable_vram_pct, min_vram_mb). Joining a hole is then a division:

claim_mb = ceil((layers_of_the_hole * mem_per_layer_mb + global_tensor_mb) * 100 / usable_vram_pct)

global_tensor_mb is also why the split is uneven where it counts: the range that starts at layer 0 also loads the token embedding, and the last range loads the lm_head, so a card that serves a middle range comfortably can be refused for the ends. A mesh is anchored by one node whose claim has a real VRAM budget behind it — the seed node.

Measured on Qwen/Qwen2.5-0.5B-Instruct: a 24-layer model where a layer costs ~56 MB on CUDA (40 MB weights + 16 MB KV) against ~89 MB on Metal — F16 against F32. The hard floor is 256 MB, and the operator's min_vram_gb raises it from there. Health is one of three words, and the words are the allocator's own walk: healthy — a ready chain renders the whole model; degraded — a gap with shards above it (held but unreachable); incomplete — the walk runs out before the last layer, which is also what a mesh that lost its tail node answers.

04 How a token crosses the mesh

The client talks to one node — the entry node. Everything after that is hops between peers, and the tokens come back the same way they left: signed, and through the chain.

  1. Client → entry node. Discover a node for the model, then talk to it like any completion server. Skip anything with "canary": true: those are replicas that answer with a known-wrong tensor on purpose, to catch a node that copies its neighbour.
    curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" \
      | jq -r '.[] | select(.canary == false and .status == "Healthy") | .http_api_url' | head -1
  2. Entry node prefills. It tokenizes the prompt and runs its own layers — the embedding plus its range. From here on the unit of transport is the hidden state, not text.
  3. Hop. The entry node POSTs to the next peer's /v1/forward with the hidden state, the tokens so far, its own entry_node_url and the requester's tier. Each node runs the layers it owns and forwards the result — a hop body is capped at 64 MiB, and each node publishes the body size it accepts.
  4. Tail → tokens. The last node detokenizes and pushes tokens back to {entry_node_url}/v1/token_stream/{request_id}. That callback is a capability, not a guess: it is authorised by the single-use signature the entry node put on the payload this shard processed, so only a node that saw this pipeline's traffic can write into the stream. The entry node relays them to the client as they arrive.
  5. KV cache stays local. Each node keeps the KV cache of the sessions it serves, so a hop only carries the increment — not the whole context. That is the reason the chain is cheap to keep alive and expensive only to build.

The three numbers of the deal

Before sending a long prompt, read GET /status on the entry node. It publishes the ceiling it accepts, the time it waits for a hop, and how many hops it reads at once — a client that reads only the first one sizes a prompt the node will then declare dead:

max_input_tokens
tokens this node can prefill in one request, measured on its own VRAM, with max_input_tokens_source telling you whether it was declared or measured
hop_timeout_secs
how long a hop of this pipeline may take (default 10 s, p2p_inference_timeout_secs)
max_concurrent_hop_bodies / hop_bodies_in_flight
the quota and the live count. At the top, the node refuses with 503 too_many_hop_bodies — that is a full node, not a broken one, and retrying elsewhere pushes 70–80 MiB at every candidate

When a hop fails, the upstream node records it in peer_failures, reports it with POST /report_failure, and the orchestrator can re-shard the stretch. A node that answers with the wrong tensor is caught by sampled layer audits and canary replicas — not by trusting its own logs.

05 Your node's HTTP API

Eight routes on api_port. Four of them are the wire — the hop and the completion surfaces; the rest are for you and the orchestrator.

POST /v1/forward
the hop: hidden state in, hidden state out. This is what peer nodes dial, and the only route that carries the whole context in one body.
POST /v1/chat/completions
client surface — messages in, tokens out (streaming supported)
POST /v1/generate
client surface — raw prompt completion
POST /v1/fim/completions
fill-in-the-middle for code: prefix + suffix in, completion out
POST /v1/token_stream/{request_id}
callback the tail node pushes tokens into (single-use capability, rate limited)
POST /api/v1/update_assignment
how a new layer split reaches the running node when the mesh is re-sharded
GET /status
model, layer_ranges, assignment, hidden size, the three hop numbers, peer failure counters, GPU health
GET /metrics
counters for your own monitoring
curl -s localhost:8080/status | jq

curl -s localhost:8080/v1/chat/completions -H 'content-type: application/json' -d '{
  "model": "Qwen/Qwen2.5-0.5B-Instruct",
  "messages": [{ "role": "user", "content": "Hello!" }],
  "max_tokens": 64, "temperature": 0.0
}'

curl -s localhost:8080/v1/fim/completions -H 'content-type: application/json' -d '{
  "prefix": "fn main() {", "suffix": "}", "max_tokens": 128
}'

06 Credits, tickets and proofs

The ledger pays for work that can be pointed at. Everything else — your uptime, your good intentions, your hardware — pays nothing.

  1. A ticket authorises. When a session starts, the orchestrator mints a signed WorkTicket with a one-hour TTL and hands it to the requester in the same round-trip that grants the session — so authorising the payment adds no latency to the request itself.
  2. A proof settles. Each node writes a signed WorkProof into its own SQLite WAL before uploading it: a node that dies mid-upload still gets paid for what it served. Settlement is authorised by the proof itself — signed by the worker and by the orchestrator that minted its ticket — never by the requester's word.
  3. One credit, one node. A credit is paid once per (request_id, worker_id), and only for the layers the worker is actually assigned: a replay of the same ticket pays nothing the second time.
  4. Your tier is a priority, not a wallet. The balance of the identity behind the key decides the queue tier of its requests: > 1000 → VIP, > 0 → Standard, ≤ 0 → Leech. A busy mesh serves VIP first; a quiet one serves everyone.
Two shortcuts for operators

With economy_enabled = false every requester is treated as VIP — the ledger keeps recording, the tier stops mattering. In deployment_mode = "local" nobody is privileged: a score cannot order a LAN of trust, so every requester gets the mesh baseline.

Cheating is priced rather than forbidden: sampled layer audits compare what a node returns against a recomputation, canary replicas answer correctly-but-wrongly so a copier is provably wrong, and the orchestrator keeps the blunt instruments — /slash, blacklisting, reputation counters.

07 The orchestrator API

Three tiers: what anyone may read, what a node must sign, and what only the operator may touch. The full reference, with measured bodies, ships in the Musivum repository (docs/API_REFERENCE.md).

Public — discovery and routing

GET /peers?model_id=<repo>
the routing contract: every node of a model with its dialable endpoints, layer ranges, health counters and canary flag. A model nobody serves answers [], never a 404
GET /api/network/models
what the mesh can render: [{ "model_id": …, "active_workers": n }]
GET /workers
monitoring view — peer id, model, status, VRAM claim and queue depth, with no addresses in it
GET /model_health?model_id=&device=
coverage, holes, and the memory terms a repair is sized with
GET /api/v1/assignment?model_id=&vram_mb=&device=
the range a claim would get — resolved before you install anything
GET /weights/manifest & /weights/manifests
the manifest a node must own to serve the model, and the manifests the mesh knows
GET /p2p_id
the mesh's swarm identity, so a node can pin the key it dials
GET /federation/status & GET /metrics
federation cursor state and operational counters

Worker-signed — the write path

POST /register
join the topology with a claim; this is where the layer split is computed
POST /announce
heartbeat: liveness, role, queue depth, health
POST /ready
shards loaded — I serve now
POST /deregister
leave cleanly
GET /api/v1/session_grant
the session's QoS ticket and its work ticket, in one round-trip
POST /api/v1/audit_report
what a sampled audit saw
POST /work_done & POST /api/v1/work_done_batch
settlement, authorised by the proof each request carries
POST /report_failure
a hop or a peer that failed, so the mesh can act on it

Operator — needs admin_token

GET /ledger?limit= & GET /ledger/snapshot?since=&limit=
the credit ledger and its cursor-paged snapshot — how federation syncs without a full dump
GET /credits?peer_id= & GET /reputation?peer_id=
one identity's balance and score (or ?all=true for the table)
POST /slash
penalise an identity that was caught
POST /rebalance
recompute the split and re-assign the mesh
POST /weights/manifest
publish the manifest a model must match
The live orchestrator

The public one answers at https://tracker1.musivum.dev (http://34.171.244.142:8080 directly). It is the team's development deployment: the read endpoints are open, and anything that writes needs a node key or the admin token. Try it before you build anything —

curl -s https://tracker1.musivum.dev/api/network/models
curl -s "https://tracker1.musivum.dev/model_health?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq '{health, covered_layers, holes}'

If /api/network/models answers [], the mesh you are about to join is empty — your node will be the first tile.

08 Chatting as a client

Two ways in: the bundled CLI, which discovers a node for you, or any client that speaks the completion routes.

tessella-cli

# a node's own API is the default target (http://127.0.0.1:8080)
cargo run -p tessella-cli --release -- chat "hello from the swarm"

# point it at any entry node you discovered in /peers
cargo run -p tessella-cli --release -- \
  --api-url http://203.0.113.7:8080 chat "hello from the swarm"

# interactive, with a session file that keeps the conversation's KV cache alive
cargo run -p tessella-cli --release -- \
  --api-url http://203.0.113.7:8080 chat --interactive --session-file ./session.json

# fill-in-the-middle, and a per-request budget
cargo run -p tessella-cli --release -- -a http://203.0.113.7:8080 fim \
  --prefix "fn main() {" --suffix "}" --max-tokens 128

--api-url (short -a) and --model-id are top-level arguments: they come before the subcommand. chat takes -p/--prompt or trailing words, -m/--max-tokens (512 by default), --temperature (0.0 by default), -i/--interactive, --session-file and --timing; --mode switches it between chat and raw completion, and --prefix/--suffix do the same for fill-in-the-middle without leaving the subcommand. tessella-cli init writes the config and the client keys, and -c/--config points at a different file.

Any OpenAI-style client

The entry node's completion routes are the whole contract — no SDK, no handshake, no session token to fetch first:

curl -s http://203.0.113.7:8080/v1/chat/completions \
  -H 'content-type: application/json' -d '{
    "model": "Qwen/Qwen2.5-0.5B-Instruct",
    "messages": [{ "role": "user", "content": "Explain a mosaic." }],
    "max_tokens": 128, "stream": true
  }'

stream: true is what makes the pipeline legible: tokens arrive from the tail node, hop by hop, as the entry node relays them.

09 Operations: GPU, proxies, systemd

Everything an operator actually has to decide: how the node starts, what it announces, where it keeps weights, and how much of the GPU it may take.

Reaching a node that sits behind NAT

TESSELLA_PUBLIC_API_URL
proxy mode: the full public URL peers should dial, e.g. https://pod-8080.proxy.runpod.net. Anything with an https:// scheme and no port is treated as a tunnel
TESSELLA_PUBLIC_IP + TESSELLA_PUBLIC_API_PORT + TESSELLA_PUBLIC_P2P_PORT
TCP mode: what a NAT rule maps to your private ports. Publish the API port and the P2P port — with only the API, hops work but discovery does not
TESSELLA_ALLOW_PRIVATE_ADDRS
set this only when the callback may legitimately target a private address (otherwise it is refused: an unvalidated callback target is an SSRF primitive)

Running it under systemd

[Unit]
Description=Tessella compute node
After=network-online.target

[Service]
User=operator
WorkingDirectory=/opt/tessella
EnvironmentFile=/etc/tessella/node.env        # HF_TOKEN=…, ORCHESTRATOR_URL=…
ExecStart=/opt/tessella/tessella-worker --config /etc/tessella/worker_config.toml
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target

The trade-offs you can turn

--model-cache-prune off|dry-run|on
whether the node deletes the shards its new assignment no longer needs. Run dry-run first: it reports what it would delete and deletes nothing
--clean-cache
wipe the local shard cache and re-download — the honest way out of a corrupted cache
quantization = "Q8_0"
8-bit weights in memory instead of the default unquantized layout: a smaller claim, a slightly different forward
canary_mode = "off" | "echo" | "scale"
operator tool: the node answers with a known-wrong tensor instead of computing, which is how the mesh proves its audits catch a liar
kademlia_enabled
join the DHT or stay with the orchestrator's membership only
max_input_tokens
declare the prefill ceiling you can hold instead of letting the node measure it — clients read it before sizing a prompt
oom_streak / streak_window_secs
how many consecutive OOMs, inside how many seconds, mean "saturated" rather than "broken"

In a container, mount the weights directory and the key outside the image or you will re-download a model and lose your identity on every restart:

docker run -d --name tessella --gpus all \
  -p 8080:8080 -p 9080:9080 \
  -e HF_TOKEN="hf_…" \
  -e ORCHESTRATOR_URL="https://tracker1.musivum.dev" \
  -v /srv/tessella/models:/app/models \
  -v /srv/tessella/key:/home/tessella/.config/tessella \
  tessella-worker:latest

10 Troubleshooting & limits

Nine of every ten problems are one of these. The tenth is a bug worth reporting with the /status body attached.

Node runs but /peers never lists it
it is PendingLoad — shards still downloading, or the model is gated and HF_TOKEN is missing. /status shows the assignment it already owns; /ready is what publishes it
POST /register → 409
the claim cannot host a single layer of that model. Size it with /model_health and remember the floor: 256 MB hard, the operator's min_vram_gb above it
POST /register → 403 worker_not_authorized
the mesh is private and this peer is not in the allow-list. Nothing to fix on the node
POST /register → 401
the signature did not verify: wrong key file, a copied config, or a badly skewed clock. Regenerating the key gives you a new identity — and a new credit balance
503 too_many_hop_bodies
this node is reading its quota of hop bodies at once. It is full, not broken: retrying elsewhere pushes 70–80 MiB at every candidate, so wait and retry
413 on a long prompt
the body exceeds the ceiling the node publishes. Read max_input_tokens from /status instead of discovering it by truncation
Hops die at the same timeout every time
one node in the chain is slower than the pipeline's patience. Raise p2p_inference_timeout_secs on the slow node, or lower the prompt it has to carry
Requests are always low priority
the identity is Leech: balance ≤ 0. Serve work — the tier follows the balance, not the intent
CUDA out of memory under load
the claim is above what the GPU can hold once sessions pile up. Lower vram_mb, or quantize to Q8_0. The node declares saturation after oom_streak consecutive OOMs and stops taking work instead of thrashing
Answers are subtly wrong
check whether you dialled a canary node: it returns a deliberately wrong tensor. Clients must skip "canary": true

What runs today, and what is still missing

Read the source

The node, the protocol and the CLI are one repository — github.com/vitichenko/tessella — and the operator notes live in docs/HOWTO.md and README_DEPLOY.md. It is private until the final end-to-end tests close, and it opens to the public with them. If this manual and the code disagree, the code is right and the manual is a bug.

↑ back to the top of the manual