Infra

Inference APIs, routers and GPUs you rent by the second.

Reset

12 products

editorially ranked — placement is never sold

Anyscale

Anyscale
✓ 22 Aug 2026

Managed Ray for distributed training and inference, either hosted or inside your own cloud account.

RayDistributedBYOC

Baseten

Baseten
✓ 22 Aug 2026

Model deployment with fast cold starts and per-minute billing, so idle time costs nothing.

InferenceDeploymentPer-minute

DeepInfra

DeepInfra
✓ 22 Aug 2026

Token-metered serving for open-weight models at some of the lowest published rates, with no contracts or upfront cost.

InferenceOpen weightsPer-token

Fireworks AI

Fireworks
✓ 17 Aug 2026

Low-latency open-model serving with tuned deployments and function calling.

InferenceLow latency

Groq

Groq
✓ 17 Aug 2026

Custom silicon delivering very high tokens-per-second on open models.

Lambda

Lambda
✓ 22 Aug 2026

On-demand GPU instances billed by the minute with no egress fees, from a company that also builds the hardware.

GPUTrainingNo egress

Modal

Modal
✓ 17 Aug 2026

Serverless GPUs for your own inference and training jobs, billed per second with a free starter allowance.

GPUServerless

OpenRouter

OpenRouter
✓ 17 Aug 2026

One API key, hundreds of models, automatic failover.

RouterMulti-model

Replicate

Replicate
✓ 17 Aug 2026

Run thousands of community models by the second — image, video, audio, text.

ModelsPer-second

RunPod

RunPod
✓ 22 Aug 2026

Per-second GPU rental across Community and Secure Cloud tiers, plus serverless inference workers.

GPUPer-secondServerless

Together AI

Together
✓ 17 Aug 2026

Fast serverless inference and fine-tuning for open-weight models, billed per token.

InferenceFine-tune