Making · 2026

One model across two laptops

Pooling the GPUs of a Windows laptop and a Kali laptop over Ethernet to load a language model that fits on neither. It works, it's slower than you'd hope, and measuring exactly why was the point.

llama.cpp · CUDA · Networking · PowerShell

The idea

I have two laptops: one with an RTX 4050 and 6 GB of VRAM, one with a GTX 1650 and 4 GB. Neither can load a model that wants 8 GB. Together, in principle, they can.

llama.cpp has an RPC backend that makes this possible. One machine runs the server and holds some of the model’s layers; the other runs a worker and holds the rest. Activations cross the network between them.

The build

HeadWorker
OSWindows 11Kali Linux
GPURTX 4050, 6 GBGTX 1650, 4 GB
Link1 Gbps Ethernet, direct0.9 ms round trip

About 8.7 GB of pooled VRAM once each machine’s own overhead is subtracted.

The repository is a set of numbered scripts — detect hardware on both sides, build the worker from source with CUDA, set up the head, run the cluster, smoke-test it — because a setup that spans two operating systems and two driver stacks needs to be reproducible or it will never work twice.

What it actually buys you

Capacity, not speed. This is the part most write-ups skip.

The RPC backend does not parallelize. Layers run in sequence, one machine at a time, and every token’s activations make the trip across the wire. Splitting a model over two GPUs on gigabit Ethernet runs at very roughly a third to a half of what a single GPU would manage if the model fit on it.

So the honest guidance, which is the first thing in the README: if your model fits on one GPU, don’t do this. Splitting across a GPU and system RAM on a single machine is often faster than splitting across two machines. Reach for the cluster only when the model genuinely doesn’t fit anywhere else.

What I took from it

The older GPU is the slower half by a wide margin, and a pipeline of sequential stages runs at the speed of its slowest stage. That’s the same lesson as every data pipeline I’ve built, in a different costume.

The RPC protocol also has no authentication and no encryption. A direct cable between two machines is the intended use; anything else wants an SSH tunnel around it.

← All making

Contact

Let's build something.

I'm looking for data science, machine learning and AI engineering roles in Phoenix or remote — and I'm always happy to talk about models, retrieval, or why yours is overfitting.