Engineering notes

Technical blog

How we tune inference per accelerator, what it does to cost per token, and what we find along the way.

BenchmarkResult pending

Fastest GLM-5.2 on AMD — measured on 256k context