Overview
Seventeen models behind one OpenAI-compatible endpoint
Console
Issue keys, watch usage and spend
API reference
curl, Python and Node, ready to paste
Templates
Browse everything you can run
Deploy
Pick a card, get a URL in 60 seconds
GPU pricing
Per-hour rates, card by card
Provider comparisons, deployment recipes, and honest cost breakdowns for independent AI developers.
A working table of VRAM needed to run Llama, Qwen, Gemma, Mistral, DeepSeek distills, gpt-oss and FLUX at FP16 and Q4, with KV-cache headroom, mapped to the GPU you should rent: RTX 3090, 4090, 5090, A6000, A100 or RTX PRO 6000.