Step by step: local or API
Choose between running models on your own machine (private, offline, free) or consuming via API from major providers.
Runs on your hardware. Zero cost per token, data never leaves.
Install Ollama, download a model, and make your first local call in under 10 minutes without sending data anywhere.
If you prefer clicking over terminal commands, LM Studio is the fastest path on Windows.
Building from source gives you the highest throughput and fine-grained GPU flags.
Set up OpenAI, Anthropic, Groq and others in minutes.
Go from zero to a production feature with GPT-5 without blowing up the bill.
Claude 4.5 Sonnet is a strong default for coding and agents. Here's how to connect it.
Llama, Qwen, and Mixtral at 500+ tokens per second. Perfect for chatbots with instant UX.