Overview
Seventeen models behind one OpenAI-compatible endpoint
Console
Issue keys, watch usage and spend
API reference
curl, Python and Node, ready to paste
Templates
Browse everything you can run
Deploy
Pick a card, get a URL in 60 seconds
GPU pricing
Per-hour rates, card by card
Provider comparisons, deployment recipes, and honest cost breakdowns for independent AI developers.
Step by step: rent an RTX 4090 or 5090 by the minute, get Ollama running with a public OpenAI-compatible endpoint on cloudgpu.app, and call it from curl, Python or Cursor. Includes measured cold-start times and what happens under the hood.