DeepInfra
Token-metered serving for open-weight models at some of the lowest published rates, with no contracts or upfront cost.
Model deployment with fast cold starts and per-minute billing, so idle time costs nothing. Ships both dedicated deployments and hosted model APIs.
Usage-based plans are shown as Metered and are excluded from stack totals.
| Category | Infra |
|---|---|
| Company | Baseten |
| Free plan | Yes |
| Starting price | Free |
| Top plan | Not verified |
| Pricing model | free and usage |
| API | Not verified |
| Open weights | Not verified |
| Commercial use | Not verified |
| Rating | Not rated — we do not publish scores we cannot source. |
Last verified: 22 Aug 2026
Pricing source: baseten.co
Listing status: Unclaimed, verified
Are you the owner of Baseten? Claim this listing.
Token-metered serving for open-weight models at some of the lowest published rates, with no contracts or upfront cost.
Low-latency open-model serving with tuned deployments and function calling.
Fast serverless inference and fine-tuning for open-weight models, billed per token.
Managed Ray for distributed training and inference, either hosted or inside your own cloud account.
Custom silicon delivering very high tokens-per-second on open models.
On-demand GPU instances billed by the minute with no egress fees, from a company that also builds the hardware.