Using hetzner's free inference in your coding agent

Hetzner is offering some free inference and since free tokens are always welcome it would be a shame not to give the models for a spin. Specially since they now offer Qwen3.8-27B for free, which is perfect for those who don’t have enough local compute to run it themselves (or those who need more tokens - and who doesnt). Below are examples for two agents I regularly Pi configuration cat ~/.pi/agent/models.json ...

August 19, 2026

262K context on 16GB VRAM because why not

I’ve been running local LLMs for a while now and the eternal struggle is always the same: you want more context, more model, more speed — and you have none of the VRAM to support any of it. So when Gemma 4 dropped with 262K context window I obviously had to try fitting the whole thing on my RTX A5000. 16GB. Turns out you can. And its actually usable. ~658 tok/s prompt eval. ~35 tok/s decode. Full 262K context window. f16 KV cache, no compression tricks. ...

June 3, 2026