musivum.dev

Open-source compute · proprietary orchestration

A mosaic of
idle GPUs.

Musivum stitches scattered consumer graphics cards into one model. Each machine runs a Tessella node holding a whole contiguous range of layers, and the hidden state hops node to node until the sentence is done. It is built for private meshes — your machines, your LAN, your VPN — not a public API.

tessella (Latin) — a small tile of a mosaic; the diminutive of tessera.

Claim floor
256 MB VRAM
Tensor hop
HTTP
Membership
libp2p
Engine
Candle
One model · N Tessella nodes · tokens flow tile to tile

01 Two planes, one net

A coordination brain and a muscle of GPUs. They share data contracts and nothing else.

Musivum control plane · proprietary

The orchestrator. It admits nodes, computes the whole-layer split from the VRAM each node declares, issues a signed WorkTicket per session, and keeps the ledger of who served what. It never touches a tensor: inference hops travel worker to worker, not through it.

Closed-source, operated by the Musivum team — today that is the public development tracker; the official image for your own server is announced but not published yet.

  • Axum REST control plane
  • libp2p membership + relay
  • SQLite ledger & WAL
  • operator API

Tessella data plane · open source

The compute node. It loads only its layers, runs the forward pass and POSTs the hidden state to the next tile over plain HTTP. Every message it sends to the orchestrator is Ed25519-signed with its own peer identity.

Yours to read, build and run — everything you install lives here. GPU-only: there is no CPU fallback.

02 What this is — and what it is not

One model, cut into contiguous ranges of whole layers, one Tessella node per range. The two lists below decide whether that fits your case.

It is runs today

A mesh for machines you own or trust, and the shape it is built for is a private one: deployment_mode = "private" plus the peer ids you list, so a colleague's desktop on your LAN or a machine across your VPN is the normal case — not a market of strangers. Each node gets a layer_range sized by the VRAM it claims.

  • whole layers per node
  • heterogeneous consumer GPUs
  • claims, never telemetry
  • re-shardable by the operator

It is not do not expect

A public swarm you can depend on today: the tracker at tracker1.musivum.dev is the team's development deployment. It is not a token or a crypto project — credits are a priority queue and are never sold. It is not a CPU path, and a node alone serves nothing without an orchestrator. And it is not a way to make one small card run a big model by itself: the ends of the pipeline are the heavy ones.

  • no token sale
  • no CPU path
  • no cloud API
  • no public service yet

03 The activation travels

Prompt in at the entry node. The hidden state hops tile to tile — each one runs the layers it owns. Tokens stream back from the last tile to the entry node, and from there to you.

  1. 0Entryembed · layers 0–8
  2. 1Middlelayers 9–20
  3. 2Middlelayers 21–28
  4. 3Taillayers 29–36 · lm_head

Pipeline parallelism, whole layers. The model is cut into contiguous layer_ranges and each node gets one, sized by the VRAM it claims — never by telemetry. The split is two-dimensional (layers × experts) only on paper: the second dimension is parked until there is hardware to measure it.

If a node blinks out, the orchestrator re-shards the gap and the pipeline keeps flowing — that rebalance is a knob the operator turns, not a background surprise.

The ends are the heavy ones

The split is even in layers and uneven in memory. Whoever is assigned the first range also loads the token embedding, and whoever gets the last one loads the lm_head — the extra weight the allocator publishes as global_tensor_mb. So a mesh is anchored by one node with a real VRAM budget, and a 256 MB claim buys a middle tile, not a place at the head of the pipeline.

04 Compute for compute

No fiat. No token sale. You earn the right to use the swarm by feeding it.

VIPcredits > 1000
Standardcredits > 0
Leechcredits ≤ 0

Credits are paid only against the WorkTicket the orchestrator issued for a session, and the node's evidence is a signed WorkProof written to its own SQLite WAL before it is uploaded. Your tier is a priority queue for when the mesh is busy — not a wallet you can drain.

05 Run a Tessella node

Everything you install is open source, and it is the same binary in every case. Point it at an orchestrator — the team's development one below, or your own.

quickstart
# 1 · get Tessella (open source, Apache-2.0) and build it for your GPU
git clone https://github.com/vitichenko/tessella
cd tessella && cargo build --release
# artifact: target/release/tessella-worker

# 2 · an identity, a model and a claim — the claim, not your free VRAM, sizes your shard
export HF_TOKEN="hf_…"
./target/release/tessella-worker init        # worker_config.toml + the Ed25519 keypair
./target/release/tessella-worker keys show   # the peer id your orchestrator will see

# 3 · serve: register, receive a layer_range, wait for hops
./target/release/tessella-worker \
  --tracker-url https://tracker1.musivum.dev \
  --model-id Qwen/Qwen2.5-0.5B-Instruct \
  --vram-gb 2.0

# 4 · find a live entry node, then chat with the mesh
curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq length
./target/release/tessella-cli \
  --api-url http://<entry-node>:<api-port> \
  --model-id Qwen/Qwen2.5-0.5B-Instruct \
  chat --prompt "hello from the mesh"

Point it anywhere

--tracker-url (or ORCHESTRATOR_URL) is the only thing an operator has to hand you. Behind NAT or a cloud proxy your node also needs its public address:

--tracker-url  /  ORCHESTRATOR_URL
HTTP address of the orchestrator to join (older aliases: --musivum-url, TRACKER_URL; default 127.0.0.1:8080)
HF_TOKEN
Hugging Face token for the weight download (required the first time you load a shard)
--model-id
The Hugging Face repo the mesh is serving (Qwen2 dense or Qwen2-MoE)
--vram-gb
VRAM you claim for shards — the split is computed from this (alias --vram, or TESSELLA_VRAM_MB)
TESSELLA_PUBLIC_API_URL
Full public URL announced to peers, e.g. https://pod-8080.proxy.runpod.net (proxy mode)
TESSELLA_PUBLIC_IP
Public IP when there is no proxy (TCP mode, alias TESSELLA_EXTERNAL_IP)
TESSELLA_PUBLIC_API_PORT
Externally mapped API port (default 8080)
TESSELLA_PUBLIC_P2P_PORT
Externally mapped P2P port (default 9080)

The full walkthrough — registration, the hop protocol, the ledger, the operator API and what each error means — is in the manual.

06 Who runs the orchestrator

The node you install is identical in every case. What changes is which orchestrator it talks to — and who is allowed in.

the case this is built for

A private mesh

Your machines and the people you trust, on a LAN or a VPN. Admission is evaluated before the signature, so an unlisted peer is refused without spending a verification. Keep one machine with a real VRAM budget in the mesh: it will be the node that holds the embedding and the lm_head.

# musivum.toml — the orchestrator that admits only your peers
deployment_mode    = "private"   # or "local": a LAN where nobody is privileged
authorized_workers = ["12D3KooW…", "12D3KooW…"]   # ignored outside "private"
allow_private_addrs = true       # peers on a LAN/VPN publish private addresses

This is the deployment the protocol is best at today: no Sybil resistance is needed when you know every peer id in the mesh.

self-hosted · announced, not published

Your own orchestrator

The orchestrator is proprietary and the official image is not published yet: it ships with the self-hosted bundle for small groups, and a large multi-tenant deployment stays with the Musivum team. Below is the shape that command will take — until it exists, the team's development tracker is the only ready-made orchestrator.

docker run -d --name musivum \
  -p 8080:8080 -p 9080:9080 \
  -e MUSIVUM_EXTERNAL_IP=<public-ip> \
  -e MUSIVUM_ADMIN_TOKEN=<32+ hex chars> \
  musivum:latest

Your nodes talk to it over the open tessella-protocol, so a mesh and its orchestrator can be upgraded apart.

The public deployment is a test bench

tracker1.musivum.dev runs so the protocol can be exercised in the open, and it is the one place where you can watch a mesh work today: point a node at it and see it get a layer_range. Treat it as a development deployment — models and workers change with whatever is being tested, and nothing there is a service-level promise. Live mesh →

07 Where it stands

This is a pre-production network: the public orchestrator is the team's development deployment. Below is the honest state of the engineering, updated as items ship.

Status

This site is live, and it is the only place where Musivum and Tessella are documented today. The code runs end to end on a test mesh; the final end-to-end tests are still pending, and the repository and the self-hosted bundle open to the public when those close. Until then there is nothing to buy here, and no public service to depend on.