Making · 2026
One model across two laptops
Pooling the GPUs of a Windows laptop and a Kali laptop over Ethernet to load a language model that fits on neither. It works, it's slower than you'd hope, and measuring exactly why was the point.
The idea
I have two laptops: one with an RTX 4050 and 6 GB of VRAM, one with a GTX 1650 and 4 GB. Neither can load a model that wants 8 GB. Together, in principle, they can.
llama.cpp has an RPC backend that makes this possible. One machine runs the server and holds
some of the model’s layers; the other runs a worker and holds the rest. Activations cross the
network between them.
The build
| Head | Worker | |
|---|---|---|
| OS | Windows 11 | Kali Linux |
| GPU | RTX 4050, 6 GB | GTX 1650, 4 GB |
| Link | 1 Gbps Ethernet, direct | 0.9 ms round trip |
About 8.7 GB of pooled VRAM once each machine’s own overhead is subtracted.
The repository is a set of numbered scripts — detect hardware on both sides, build the worker from source with CUDA, set up the head, run the cluster, smoke-test it — because a setup that spans two operating systems and two driver stacks needs to be reproducible or it will never work twice.
What it actually buys you
Capacity, not speed. This is the part most write-ups skip.
The RPC backend does not parallelize. Layers run in sequence, one machine at a time, and every token’s activations make the trip across the wire. Splitting a model over two GPUs on gigabit Ethernet runs at very roughly a third to a half of what a single GPU would manage if the model fit on it.
So the honest guidance, which is the first thing in the README: if your model fits on one GPU, don’t do this. Splitting across a GPU and system RAM on a single machine is often faster than splitting across two machines. Reach for the cluster only when the model genuinely doesn’t fit anywhere else.
What I took from it
The older GPU is the slower half by a wide margin, and a pipeline of sequential stages runs at the speed of its slowest stage. That’s the same lesson as every data pipeline I’ve built, in a different costume.
The RPC protocol also has no authentication and no encryption. A direct cable between two machines is the intended use; anything else wants an SSH tunnel around it.