Musivum stitches scattered consumer graphics cards into one model. Each machine runs a
Tessella node holding a whole contiguous range of layers,
and the hidden state hops node to node until the sentence is done. It is built for
private meshes — your machines, your LAN, your VPN — not a public API.
tessella (Latin) — a small tile of a mosaic;
the diminutive of tessera.
One model · N Tessella nodes · tokens flow tile to tile
01 Two planes, one net
A coordination brain and a muscle of GPUs. They share data contracts and nothing else.
Musivum control plane · proprietary
The orchestrator. It admits nodes, computes the whole-layer split from the
VRAM each node declares, issues a signed WorkTicket per session,
and keeps the ledger of who served what. It never touches a tensor: inference hops travel
worker to worker, not through it.
Closed-source, operated by the Musivum team — today that is the public development tracker;
the official image for your own server is announced but not published yet.
The compute node. It loads only its layers, runs the forward pass and
POSTs the hidden state to the next tile over plain HTTP. Every
message it sends to the orchestrator is Ed25519-signed with its own peer identity.
Yours to read, build and run — everything you install lives here. GPU-only: there is no CPU
fallback.
One model, cut into contiguous ranges of whole layers, one Tessella node per range. The two
lists below decide whether that fits your case.
It is runs today
A mesh for machines you own or trust, and the shape it is built for is a private one:
deployment_mode = "private" plus the peer ids you list, so a
colleague's desktop on your LAN or a machine across your VPN is the normal case — not a
market of strangers. Each node gets a layer_range sized by the
VRAM it claims.
whole layers per node
heterogeneous consumer GPUs
claims, never telemetry
re-shardable by the operator
It is not do not expect
A public swarm you can depend on today: the tracker at
tracker1.musivum.dev is the team's development deployment. It is
not a token or a crypto project — credits are a priority queue and are never sold. It is
not a CPU path, and a node alone serves nothing without an orchestrator. And it is not a way
to make one small card run a big model by itself: the ends of the pipeline are the heavy
ones.
no token sale
no CPU path
no cloud API
no public service yet
03 The activation travels
Prompt in at the entry node. The hidden state hops tile to tile — each one runs the layers it
owns. Tokens stream back from the last tile to the entry node, and from there to you.
0Entryembed · layers 0–8
1Middlelayers 9–20
2Middlelayers 21–28
3Taillayers 29–36 · lm_head
Pipeline parallelism, whole layers. The model is cut into contiguous
layer_ranges and each node gets one, sized by the VRAM it claims —
never by telemetry. The split is two-dimensional (layers × experts) only on paper: the second
dimension is parked until there is hardware to measure it.
If a node blinks out, the orchestrator re-shards the gap and the pipeline keeps flowing —
that rebalance is a knob the operator turns, not a background surprise.
The ends are the heavy ones
The split is even in layers and uneven in memory. Whoever is assigned the first range also
loads the token embedding, and whoever gets the last one loads the
lm_head — the extra weight the allocator publishes as
global_tensor_mb. So a mesh is anchored by one node with a real
VRAM budget, and a 256 MB claim buys a middle tile, not a place at the head of the
pipeline.
04 Compute for compute
No fiat. No token sale. You earn the right to use the swarm by feeding it.
VIPcredits > 1000
Standardcredits > 0
Leechcredits ≤ 0
Credits are paid only against the WorkTicket the orchestrator issued
for a session, and the node's evidence is a signed WorkProof written to
its own SQLite WAL before it is uploaded. Your tier is a priority queue for when the
mesh is busy — not a wallet you can drain.
05 Run a Tessella node
Everything you install is open source, and it is the same binary in every case. Point it at an
orchestrator — the team's development one below, or your own.
quickstart
# 1 · get Tessella (open source, Apache-2.0) and build it for your GPU
git clone https://github.com/vitichenko/tessella
cd tessella && cargo build --release
# artifact: target/release/tessella-worker# 2 · an identity, a model and a claim — the claim, not your free VRAM, sizes your shard
export HF_TOKEN="hf_…"
./target/release/tessella-worker init # worker_config.toml + the Ed25519 keypair
./target/release/tessella-worker keys show # the peer id your orchestrator will see# 3 · serve: register, receive a layer_range, wait for hops
./target/release/tessella-worker \
--tracker-url https://tracker1.musivum.dev \
--model-id Qwen/Qwen2.5-0.5B-Instruct \
--vram-gb 2.0
# 4 · find a live entry node, then chat with the mesh
curl -s "https://tracker1.musivum.dev/peers?model_id=Qwen/Qwen2.5-0.5B-Instruct" | jq length
./target/release/tessella-cli \
--api-url http://<entry-node>:<api-port> \
--model-id Qwen/Qwen2.5-0.5B-Instruct \
chat --prompt "hello from the mesh"
Point it anywhere
--tracker-url (or ORCHESTRATOR_URL) is the
only thing an operator has to hand you. Behind NAT or a cloud proxy your node also needs its
public address:
--tracker-url / ORCHESTRATOR_URL
HTTP address of the orchestrator to join (older aliases: --musivum-url, TRACKER_URL; default 127.0.0.1:8080)
HF_TOKEN
Hugging Face token for the weight download (required the first time you load a shard)
--model-id
The Hugging Face repo the mesh is serving (Qwen2 dense or Qwen2-MoE)
--vram-gb
VRAM you claim for shards — the split is computed from this (alias --vram, or TESSELLA_VRAM_MB)
TESSELLA_PUBLIC_API_URL
Full public URL announced to peers, e.g. https://pod-8080.proxy.runpod.net(proxy mode)
TESSELLA_PUBLIC_IP
Public IP when there is no proxy (TCP mode, alias TESSELLA_EXTERNAL_IP)
TESSELLA_PUBLIC_API_PORT
Externally mapped API port (default 8080)
TESSELLA_PUBLIC_P2P_PORT
Externally mapped P2P port (default 9080)
The full walkthrough — registration, the hop protocol, the ledger, the operator API and what each
error means — is in the manual.
06 Who runs the orchestrator
The node you install is identical in every case. What changes is which orchestrator it talks
to — and who is allowed in.
the case this is built for
A private mesh
Your machines and the people you trust, on a LAN or a VPN. Admission is evaluated
before the signature, so an unlisted peer is refused without spending a
verification. Keep one machine with a real VRAM budget in the mesh: it will be the node
that holds the embedding and the lm_head.
# musivum.toml — the orchestrator that admits only your peers
deployment_mode = "private" # or "local": a LAN where nobody is privileged
authorized_workers = ["12D3KooW…", "12D3KooW…"] # ignored outside "private"
allow_private_addrs = true # peers on a LAN/VPN publish private addresses
This is the deployment the protocol is best at today: no Sybil resistance is needed when you know every peer id in the mesh.
self-hosted · announced, not published
Your own orchestrator
The orchestrator is proprietary and the official image is not published yet: it
ships with the self-hosted bundle for small groups, and a large multi-tenant deployment
stays with the Musivum team. Below is the shape that command will take — until it exists,
the team's development tracker is the only ready-made orchestrator.
Your nodes talk to it over the open tessella-protocol, so a mesh and its orchestrator can be upgraded apart.
The public deployment is a test bench
tracker1.musivum.dev runs so the protocol can be exercised in the
open, and it is the one place where you can watch a mesh work today: point a node at it and
see it get a layer_range. Treat it as a development deployment —
models and workers change with whatever is being tested, and nothing there is a service-level
promise. Live mesh →
07 Where it stands
This is a pre-production network: the public orchestrator is the team's development deployment.
Below is the honest state of the engineering, updated as items ship.
Status
This site is live, and it is the only place where Musivum and Tessella are documented
today. The code runs end to end on a test mesh; the final end-to-end tests are still
pending, and the repository and the self-hosted bundle open to the public when those
close. Until then there is nothing to buy here, and no public service to depend on.
Whole-layer sharding across machines (D-15)
Hidden-state hops over HTTP (D-10)
Ed25519-signed node DTOs + WAL work proofs
WorkTicket-authorized credits ledger
Sampled layer audit & canary replicas
Federation by signed cursor pull
NAT-aware addressing for cloud pods
Final end-to-end acceptance run — pending, and it is what opens this to the public