Baseten
Model deployment with fast cold starts and per-minute billing, so idle time costs nothing.
Token-metered serving for open-weight models at some of the lowest published rates, with no contracts or upfront cost. Dedicated GPUs are available by the hour.
Usage-based plans are shown as Metered and are excluded from stack totals.
| Category | Infra |
|---|---|
| Company | DeepInfra |
| Free plan | No |
| Starting price | Usage-based |
| Top plan | Not verified |
| Pricing model | usage |
| API | Yes |
| Open weights | Yes |
| Commercial use | Not verified |
| Rating | Not rated — we do not publish scores we cannot source. |
Last verified: 22 Aug 2026
Pricing source: deepinfra.com
Listing status: Unclaimed, verified
Are you the owner of DeepInfra? Claim this listing.
Model deployment with fast cold starts and per-minute billing, so idle time costs nothing.
Low-latency open-model serving with tuned deployments and function calling.
Fast serverless inference and fine-tuning for open-weight models, billed per token.
Custom silicon delivering very high tokens-per-second on open models.
Serverless GPUs for your own inference and training jobs, billed per second with a free starter allowance.
Run open models locally on your own machine.