Deployment
Deployment
API, GPU worker, and serverless patterns for reliable inference in production.
deploymentintermediate
Serverless Inference
Package inference behind serverless routes with strict timeouts and retries.
serverlessdeploymentinference
3 refs
deploymentintermediate
Modal GPU Worker
Run bursty GPU inference jobs on Modal with cold-start aware workers.
deploymentgpumodal
3 refs
deploymentbeginner
FastAPI Deployment
Wrap model inference behind a typed FastAPI service.
deploymentfastapiapi
3 refs