Baseten
Model deployment with fast cold starts and per-minute billing, so idle time costs nothing.
Fast serverless inference and fine-tuning for open-weight models, billed per token.
Usage-based plans are shown as Metered and are excluded from stack totals.
| Category | Infra |
|---|---|
| Company | Together |
| Free plan | No |
| Starting price | Usage-based |
| Top plan | Not verified |
| Pricing model | usage |
| API | Yes |
| Open weights | Not verified |
| Commercial use | Not verified |
| Rating | Not rated — we do not publish scores we cannot source. |
Last verified: 17 Aug 2026
Pricing source: together.ai
Listing status: Unclaimed, verified
Are you the owner of Together AI? Claim this listing.
Model deployment with fast cold starts and per-minute billing, so idle time costs nothing.
Token-metered serving for open-weight models at some of the lowest published rates, with no contracts or upfront cost.
Low-latency open-model serving with tuned deployments and function calling.
Custom silicon delivering very high tokens-per-second on open models.
Serverless GPUs for your own inference and training jobs, billed per second with a free starter allowance.
Run open models locally on your own machine.