Featured field note

Running a 27B model across two desktops with llama.cpp RPC

I have two machines with a 16 GB consumer GPU in each, and a model I wanted to run that does not fit in either one on its own. llama.cpp's RPC backend lets you treat the two cards as one pool, so I put Qwen3.8-27B...

Read the field note
The archive

Latest notes

14 field notes and counting