Choose a deployment option
Start with the model and runtime your application needs, then compare the operational responsibilities of each option.Serverless Inference
Call open-weight models from a managed catalog, or serve your own LoRA adapters on supported base models. CoreWeave manages provisioning, routing, and scaling.
Dedicated Inference
Bring your own weights (BYOW), choose GPU resources and a supported runtime, and let CoreWeave operate the serving platform.
Inference on CKS
Operate your serving stack on Kubernetes when you need control over runtimes, networking, and orchestration.
Match the option to your workload
These recommendations are starting points for evaluation. A workload can fit more than one option. Model compatibility and measured performance determine the final choice. The following table maps common workloads to a starting option and what to evaluate before you commit:
If your workload has data residency, isolation, encryption, or private connectivity requirements, verify each required control before choosing an option. Reserved GPU capacity or an Availability Zone selection doesn’t, by itself, establish that all application requirements are met.
Example deployment choices
The following hypothetical examples show how requirements affect the choice:- A team testing a document assistant starts with Serverless Inference because a catalog model meets its needs. It measures quality and latency before deciding whether a LoRA adapter on Serverless Inference or full custom weights on Dedicated Inference is justified.
- A team serving a fine-tuned text model evaluates Dedicated Inference to keep control of weights and scaling without operating the serving platform. It stages the artifacts in CoreWeave AI Object Storage and benchmarks the deployment against production traffic.
- A media application uses Inference on CKS to operate a custom model server alongside preprocessing workers and a queue. The team configures workload networking, scaling, and monitoring.
- An application with overnight generation and interactive requests evaluates the two traffic patterns separately. It can use different deployments or options when their latency, model, or orchestration requirements differ.