Every model below runs on ForgeAPI's own GPU infrastructure. Prices are per 1M tokens, billed from your prepaid balance. No monthly fees, no minimum spend.
| Model | Description | Context | $/1M input | $/1M output | Status |
|---|---|---|---|---|---|
| deepseek-v4-flash | Fast general chat & agentic workloads | 128K | $0.25 | $0.75 | available |
| deepseek-v4-pro | Heavy reasoning, math & code | 128K | $1.00 | $3.00 | available |
| kimi-3 | Strong long-context comprehension | 256K | $0.60 | $2.40 | available |
| qwen-3-235b | Top open-weight general model | 256K | $0.40 | $1.20 | beta |
| llama-3.3-70b | Reliable workhorse, great tool-calling | 128K | $0.28 | $0.84 | available |
| qwen-3-8b | Small, fast, perfect for prototyping | 32K | $0 | $0 | free |
Prices are placeholders for this demo — final pricing will be set against real hardware cost per token. Volume discounts apply automatically above $500/month spend. Want a model that isn't listed? Open-weight models are regularly added — request it.
Add credit with a card or crypto (USDT). Balances never expire. No subscription, no surprise invoice at month end.
Every request is metered exactly like the big providers — input and output tokens counted separately. See costs per key in real time.
Set a hard monthly cap per API key. When it hits the limit, requests fail with a 402 — no surprise overages, ever.
We run open-weight models on our own GPUs instead of paying per-token markup to a cloud provider. Open models cost what the silicon costs — we pass that on.
Yes. Every account gets daily free access to qwen-3-8b (rate-limited). Enough to build, test and benchmark before paying anything.
No. Prepaid balances never expire and are refundable within 30 days of purchase, minus usage.
Cards (Visa / Mastercard via Stripe) and USDT on major chains. Volume customers can arrange invoicing.
Fully. Use https://api.forgeapi.dev/v1 as your base_url with any OpenAI SDK — Python, Node, Go, or raw curl.