What Is KServe?
KServe (formerly KFServing) is a Kubernetes-native platform for serving ML models. It provides:
- Serverless inference — scale to zero when idle
- Autoscaling — handle traffic spikes automatically
- Canary deployments — roll out new model versions safely
- Multi-framework support — TensorFlow, PyTorch, scikit-learn, XGBoost
Installing KServe
Prerequisites
- A running Kubernetes cluster (Kind, Minikube, or cloud)
- kubectl configured
Install with kubectl
kubectl apply -f https://github.com/kserve/kserve/releases/download/v0.12.0/kserve.yamlVerify the installation:
kubectl get pods -n kserve
kubectl get pods -n kserve | grep kserve-controllerMaster this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Creating an InferenceService
Here's a minimal InferenceService for an MLflow model:
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: wine-quality-model
spec:
predictor:
model:
modelFormat:
name: mlflow
storageUri: "gs://your-bucket/mlflow-model"Deploy it:
kubectl apply -f inference-service.yamlCheck readiness:
kubectl get inferenceservice wine-quality-modelSending Inference Requests
MODEL_NAME=wine-quality-model
SERVICE_HOSTNAME=$(kubectl get inferenceservice $MODEL_NAME \
-o jsonpath='{.status.url}' | cut -d"/" -f3)
curl -v -H "Host: ${SERVICE_HOSTNAME}" \
http://localhost:8080/v2/models/$MODEL_NAME/infer \
-d '{
"inputs": [{
"name": "input",
"shape": [1, 13],
"datatype": "FP32",
"data": [7.4, 0.7, 0.0, 1.9, 0.076, 11.0, 34.0, 0.9978, 3.51, 0.56, 9.4, 5.0, 6.0]
}]
}'Autoscaling
KServe automatically scales based on traffic:
spec:
predictor:
minReplicas: 1
maxReplicas: 10
model:
modelFormat:
name: mlflow
storageUri: "gs://your-bucket/model"Get weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Canary Deployments
Roll out a new model version to 20% of traffic:
spec:
predictor:
canaryTrafficPercent: 20
model:
modelFormat:
name: mlflow
storageUri: "gs://your-bucket/model-v2"Monitoring
Check pod health and logs:
kubectl get pods -l serving.kserve.io/inferenceservice=wine-quality-model
kubectl logs -l serving.kserve.io/inferenceservice=wine-quality-modelLearn More
Get hands-on experience deploying models with KServe in our MLflow for Kubernetes course — from training to production inference.
---
Ready to go deeper? Check out our hands-on course: MLflow for Kubernetes — practical exercises you can follow along on your own machine.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
MLflow on Kubernetes: Full Guide
Deploy MLflow on Kubernetes for production MLOps. Helm charts, experiment tracking, model registry, and model serving guide.
MLOps Pipeline Architecture Guide
Design a production MLOps pipeline: MLflow experiment tracking, model registry, CI/CD for ML, and Kubernetes deployment patterns.
MLflow for Kubernetes
Learn how to deploy and manage ML models at scale using MLflow, Kubernetes, KServe, and Docker. A comprehensive guide to production MLOps.
Kubebuilder Custom Operators Guide
Kubebuilder scaffolds Kubernetes operators in Go. Learn how to create custom controllers, define CRDs, and build operators that automate complex application.
Kubecost Kubernetes Cost Monitoring
Kubecost shows real-time cost allocation per namespace, deployment, and label in Kubernetes. Learn how to install Kubecost, identify waste, and set budgets.
kubectl Cheat Sheet for DevOps
Essential kubectl commands for pods, deployments, services, logs, and debugging Kubernetes clusters. Copy-paste ready DevOps reference.
Explore topics
Browse more articles on the topics covered here.