Running MLflow on Kubernetes gives you scalable, production-grade MLOps. This guide covers everything from deployment to model serving.
Why MLflow on Kubernetes?
MLflow handles the ML lifecycle ā experiment tracking, model versioning, and deployment. Kubernetes handles the infrastructure ā scaling, reliability, and resource management. Together they solve the full MLOps puzzle.
Benefits: - Scalable tracking server: Handle hundreds of concurrent experiments - Reliable model registry: Version models with Kubernetes-backed storage - Auto-scaling serving: Scale model endpoints based on traffic - Team collaboration: Shared tracking server for the entire ML team
Architecture Overview
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā Kubernetes Cluster ā
ā ā
ā āāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāā ā
ā ā MLflow Server ā ā PostgreSQL DB ā ā
ā ā (Tracking) ā ā (Metadata) ā ā
ā āāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāā ā
ā ā
ā āāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāā ā
ā ā MinIO / S3 ā ā KServe ā ā
ā ā (Artifacts) ā ā (Model Serve) ā ā
ā āāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāā ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāStep 1: Deploy MLflow with Helm
Create mlflow-values.yaml:
tracking:
service:
type: ClusterIP
persistence:
enabled: true
size: 10Gi
backendStore:
postgres:
enabled: true
host: postgresql.mlflow.svc.cluster.local
port: 5432
database: mlflow
user: mlflow
artifactRoot:
s3:
enabled: true
bucket: mlflow-artifacts
endpointUrl: http://minio.mlflow.svc.cluster.local:9000Deploy:
helm repo add community-charts https://community-charts.github.io/helm-charts
helm install mlflow community-charts/mlflow \
--namespace mlflow \
--create-namespace \
-f mlflow-values.yamlMaster this topic with hands-on labs
Go beyond reading ā build real projects in sandboxed environments with expert video guidance.
Browse Courses āStep 2: Experiment Tracking
Configure your Python environment to use the Kubernetes MLflow server:
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
# Point to your K8s MLflow server
mlflow.set_tracking_uri("http://mlflow.mlflow.svc.cluster.local:5000")
mlflow.set_experiment("iris-classification")
# Load data
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
with mlflow.start_run():
# Log parameters
mlflow.log_param("n_estimators", 100)
mlflow.log_param("max_depth", 5)
# Train model
model = RandomForestClassifier(n_estimators=100, max_depth=5)
model.fit(X_train, y_train)
# Log metrics
accuracy = model.score(X_test, y_test)
mlflow.log_metric("accuracy", accuracy)
# Log model
mlflow.sklearn.log_model(model, "model")
print(f"Accuracy: {accuracy:.4f}")Step 3: Model Registry
Register your best model:
# Register model from a run
result = mlflow.register_model(
model_uri=f"runs:/{run_id}/model",
name="iris-classifier"
)
# Transition to production
from mlflow.tracking import MlflowClient
client = MlflowClient()
client.transition_model_version_stage(
name="iris-classifier",
version=result.version,
stage="Production"
)Step 4: Model Serving with KServe
Deploy your MLflow model as a Kubernetes service:
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: iris-classifier
namespace: serving
spec:
predictor:
model:
modelFormat:
name: mlflow
storageUri: "s3://mlflow-artifacts/1/abc123/artifacts/model"
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1"
memory: "1Gi"Test the endpoint:
curl -X POST http://iris-classifier.serving.svc.cluster.local/v1/models/iris-classifier:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[5.1, 3.5, 1.4, 0.2]]}'Get weekly IT automation tips
Docker, Ansible, Terraform, MLOps ā curated insights delivered to your inbox. No spam.
Subscribe Free āStep 5: Monitoring
Add Prometheus metrics for your model endpoints:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: mlflow-monitor
spec:
selector:
matchLabels:
app: mlflow
endpoints:
- port: http
interval: 30s
path: /metricsKey metrics to track:
| Metric | What It Tells You |
|---|---|
| Request latency | Model inference speed |
| Error rate | Failed predictions |
| Memory usage | Model resource needs |
| Request count | Traffic patterns |
| Data drift score | Input distribution changes |
Common Pitfalls
- Artifact storage: Use S3/MinIO, not local filesystem ā pods are ephemeral
- Database backups: PostgreSQL metadata is critical ā set up automated backups
- Resource limits: ML models are memory-hungry ā set proper requests and limits
- Namespace isolation: Separate tracking, serving, and monitoring namespaces
What's Next?
Our MLflow for Kubernetes MLOps course covers all 15 steps in depth with hands-on labs ā from local development with Kind to production deployment with monitoring. The first lesson is free.
---
Ready to go deeper? Check out our hands-on course: MLflow for Kubernetes ā practical exercises you can follow along on your own machine.
Ready to learn by doing?
Stop reading tutorials ā start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
MLOps Pipeline Architecture Guide
Design a production MLOps pipeline: MLflow experiment tracking, model registry, CI/CD for ML, and Kubernetes deployment patterns.
MLflow for Kubernetes
Learn how to deploy and manage ML models at scale using MLflow, Kubernetes, KServe, and Docker. A comprehensive guide to production MLOps.
MLflow Experiment Tracking
Learn how to track ML experiments with MLflow ā log parameters, metrics, and artifacts. Compare model runs and find the best configuration.
MLflow Model Registry
Use MLflow Model Registry to manage model versions, stage transitions, and governance. Essential for production MLOps workflows.
MLServer: Test ML Models Locally
Use MLServer to serve and test MLflow models locally before deploying to Kubernetes. Quick setup guide with inference examples.
Monitoring ML Models in K8s
Monitor deployed ML models on Kubernetes ā track prediction accuracy, latency, resource usage, and detect model drift in production.
Explore topics
Browse more articles on the topics covered here.