MODEL SELECTION TOOLS
Find the right fit.
Compare capabilities and rates, then estimate costs for your own workload.
Your estimated workload
Calculated in each model’s billing unit. Models with different capabilities are not interchangeable.
| Model & cost | GLM 5.2 FP8 Primary LLM for chat, reasoning, and coding workloads. | Qwen3-VL 30B Vision and OCR inference for document and image understanding. |
|---|---|---|
| Estimated costUSD | $1.68Based on input & output tokens | $0.60Based on input & output tokens |
| Input / unit rate | $0.93per 1M input tokens | $0.30per 1M input tokens |
| Output rate | $3.00per 1M output tokens | $1.20per 1M output tokens |
| Context window | 131K | 33K |
| Maximum output | 16K | 4K |
| Streaming | ✓ Supported | ✓ Supported |
| Tool calling | ✓ Supported | Not listed |
| Structured output | ✓ Supported | ✓ Supported |
| Get started | Configure integration ↗ | Configure integration ↗ |
Based on published catalog rates and specifications, not performance benchmarks or a live bill. See service status for model availability. Check status ↗