Skip to main content
CoreWeave offers three ways to serve AI models: Serverless Inference, Dedicated Inference, and Inference on CoreWeave Kubernetes Service (CKS), which you operate yourself. Use this guide to choose a starting point based on your model, traffic pattern, performance targets, and the serving infrastructure your team wants to operate.

Choose a deployment option

Start with the model and runtime your application needs, then compare the operational responsibilities of each option.

Serverless Inference

Call open-weight models from a managed catalog, or serve your own LoRA adapters on supported base models. CoreWeave manages provisioning, routing, and scaling.

Dedicated Inference

Bring your own weights (BYOW), choose GPU resources and a supported runtime, and let CoreWeave operate the serving platform.

Inference on CKS

Operate your serving stack on Kubernetes when you need control over runtimes, networking, and orchestration.
The following table compares what you configure and operate in each option:

Match the option to your workload

These recommendations are starting points for evaluation. A workload can fit more than one option. Model compatibility and measured performance determine the final choice. The following table maps common workloads to a starting option and what to evaluate before you commit: If your workload has data residency, isolation, encryption, or private connectivity requirements, verify each required control before choosing an option. Reserved GPU capacity or an Availability Zone selection doesn’t, by itself, establish that all application requirements are met.

Example deployment choices

The following hypothetical examples show how requirements affect the choice:
  • A team testing a document assistant starts with Serverless Inference because a catalog model meets its needs. It measures quality and latency before deciding whether a LoRA adapter on Serverless Inference or full custom weights on Dedicated Inference is justified.
  • A team serving a fine-tuned text model evaluates Dedicated Inference to keep control of weights and scaling without operating the serving platform. It stages the artifacts in CoreWeave AI Object Storage and benchmarks the deployment against production traffic.
  • A media application uses Inference on CKS to operate a custom model server alongside preprocessing workers and a queue. The team configures workload networking, scaling, and monitoring.
  • An application with overnight generation and interactive requests evaluates the two traffic patterns separately. It can use different deployments or options when their latency, model, or orchestration requirements differ.
For runnable examples, see Deploy Kimi K3 on Dedicated Inference, Deploy NVIDIA Dynamo on CKS, and Deploy vLLM for inference.

Compare performance, capacity, and cost

Define success using traffic representative of your application. Include typical and peak request rates, input and output lengths, and repeated context. Measure time to first token, response latency (including slow requests), throughput, and cost per completed request or token. For Dedicated Inference, tune concurrency against your latency and throughput targets. A setting that works for batch generation might not meet an interactive application’s response target. Follow the concurrency benchmarking guidance and repeat measurements when the model or traffic changes. Autoscaling adjusts replica counts within available capacity, but it doesn’t reserve GPUs. If you need reserved hardware, evaluate capacity claims. Managed claims incur charges while they exist, including when their capacity is idle. Dedicated Inference doesn’t support scale-to-zero, so account for its minimum running replicas when comparing costs.

Pricing

Compare Serverless Inference pricing with Dedicated Inference billing and CoreWeave compute pricing. Include expected utilization and the effort to operate a self-managed serving stack. Steady traffic is a reason to evaluate Dedicated Inference, but it doesn’t establish a universal cost advantage.

Key capabilities

Serverless Inference and Dedicated Inference expose OpenAI-compatible endpoints, so existing OpenAI client libraries, agents, and tooling can connect with minimal changes. On CKS, the serving software you deploy determines the API your application uses. Check the selected product’s documentation for model support, endpoint behavior, resource availability, and platform integrations. For Dedicated Inference, consult supported engines before choosing a runtime.

Access the API

You can configure Dedicated Inference resources through the Dedicated Inference API over REST, gRPC, or Connect, through the CoreWeave Intelligent CLI, or through Terraform. For more information, see the Dedicated Inference API reference and Deploy Dedicated Inference with Terraform. For Serverless Inference setup and request examples, see Serverless Inference prerequisites and usage examples.

Get started

Choose the guide for your deployment option:
Last modified on September 29, 2026