The same weights, for less
Cut the token bill.Keep the performance.
Run open-weight models for less, without changing how you work. On GLM-5.2, Draupnir delivers the same model at lower output-token cost than the standard hosted price — same weights, no quantization tricks.
01 — Models & pricing
Models & pricing
The same open weights, below the public hosted rate. Prices are per million tokens.
| Model | Context | Input $/M | Output $/M | Public output $/M | Status |
|---|---|---|---|---|---|
| GLM-5.2 | 500k | Pending | Pending | $3.85 | Live |
Pricing shared on API access.
Weights on Hugging FaceBenchmark figures and third-party verification pending.
OpenAI-compatible.
Change the base URL and ship.
Access details shared after review.
02 — Who it is for
Two ways the token bill hits you
01 / Serving teams
Teams serving open weights
Your margin is the gap between what a token costs you to serve and what you charge for it. Draupnir widens it without changing the model.
02 / Routers
Gateways and routers
You select providers on price and latency at a measurable target. Draupnir earns the route by occupying the better point on that surface.