What Is MLServer?
MLServer is an open-source inference server that implements the V2 Inference Protocol. It's the same serving layer KServe uses — meaning your local tests perfectly mirror production behavior.
Installing MLServer
pip install mlserver mlserver-mlflowMaster this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Serving an MLflow Model
After training and logging a model with MLflow:
mlflow models serve \
-m "runs:/<run-id>/model" \
--port 8080 \
--enable-mlserverOr serve from a local directory:
mlflow models serve \
-m ./mlruns/0/<run-id>/artifacts/model \
--port 8080 \
--enable-mlserverTesting Inference
V2 Protocol (same as KServe)
curl http://localhost:8080/v2/models/model/infer \
-H "Content-Type: application/json" \
-d '{
"inputs": [{
"name": "input",
"shape": [1, 13],
"datatype": "FP32",
"data": [7.4, 0.7, 0.0, 1.9, 0.076, 11.0, 34.0, 0.9978, 3.51, 0.56, 9.4, 5.0, 6.0]
}]
}'Health Check
curl http://localhost:8080/v2/health/readyModel Metadata
curl http://localhost:8080/v2/models/modelGet weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Python Client
import requests
import json
url = "http://localhost:8080/v2/models/model/infer"
payload = {
"inputs": [{
"name": "input",
"shape": [1, 13],
"datatype": "FP32",
"data": [7.4, 0.7, 0.0, 1.9, 0.076, 11.0, 34.0,
0.9978, 3.51, 0.56, 9.4, 5.0, 6.0]
}]
}
response = requests.post(url, json=payload)
prediction = response.json()
print(f"Prediction: {prediction['outputs'][0]['data']}")Why Test Locally First?
- Fast iteration — no waiting for Kubernetes deployments
- Same protocol — V2 protocol matches KServe exactly
- Debug easily — full access to logs and model internals
- Save resources — no cloud costs during development
- Catch errors early — before they hit production
Local to Production Workflow
Train model → Log to MLflow → Serve with MLServer (local)
→ Test thoroughly → Build Docker image → Deploy to KServeLearn this complete workflow hands-on in our MLflow for Kubernetes course.
---
Ready to go deeper? Check out our hands-on course: MLflow for Kubernetes — practical exercises you can follow along on your own machine.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
MLflow on Kubernetes: Full Guide
Deploy MLflow on Kubernetes for production MLOps. Helm charts, experiment tracking, model registry, and model serving guide.
MLOps Pipeline Architecture Guide
Design a production MLOps pipeline: MLflow experiment tracking, model registry, CI/CD for ML, and Kubernetes deployment patterns.
MLflow for Kubernetes
Learn how to deploy and manage ML models at scale using MLflow, Kubernetes, KServe, and Docker. A comprehensive guide to production MLOps.
Monitoring ML Models in K8s
Monitor deployed ML models on Kubernetes — track prediction accuracy, latency, resource usage, and detect model drift in production.
Neuromorphic Computing Explained
Understand neuromorphic computing with brain-inspired chip architectures, spiking neural networks, and practical applications for edge AI workloads.
Nginx Reverse Proxy Configuration
Configure Nginx as a production reverse proxy. Upstream pools, load balancing, SSL termination, caching, and rate limiting.
Explore topics
Browse more articles on the topics covered here.