AI infrastructure costs can spiral quickly. A single H100 GPU costs $2-3/hour in the cloud. Training runs and inference at scale easily reach six or seven figures monthly.
Where the Money Goes
Typical AI infrastructure cost breakdown:
- Training compute — 40-60% (GPU hours for model training)
- Inference serving — 20-35% (GPU/CPU for production predictions)
- Storage — 10-15% (datasets, model checkpoints, logs)
- Networking — 5-10% (data transfer, especially multi-region)
- Human overhead — Often underestimated (MLOps engineering time)
GPU Scheduling Optimization
Maximize GPU utilization to reduce waste:
# Kubernetes GPU time-slicing
apiVersion: v1
kind: ConfigMap
metadata:
name: gpu-sharing-config
data:
any: |-
version: v1
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4Time-slicing lets 4 workloads share one GPU — ideal for inference and development.
Multi-Instance GPU (MIG)
For H100/A100 GPUs, MIG provides hardware isolation:
# Create MIG instances
nvidia-smi mig -i 0 -cgi 9,9,9 -C
# Verify
nvidia-smi mig -i 0 -lgiEach MIG instance gets dedicated memory and compute — better isolation than time-slicing.
Model Optimization Techniques
Smaller models = cheaper inference:
| Technique | Size Reduction | Speed Improvement | Accuracy Loss |
|---|---|---|---|
| Quantization (INT8) | 2-4x | 2-3x | < 1% |
| Quantization (INT4) | 4-8x | 3-5x | 1-3% |
| Pruning | 2-10x | 2-5x | 1-5% |
| Distillation | 3-10x | 3-10x | 2-5% |
| LoRA adapters | N/A | Faster FT | Minimal |
Master this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Intelligent Model Routing
Route requests to the cheapest model that can handle them:
def route_request(prompt: str, complexity: str) -> str:
if complexity == "simple":
# Small model: $0.001/request
return call_model("llama-3-8b", prompt)
elif complexity == "medium":
# Medium model: $0.01/request
return call_model("llama-3-70b", prompt)
else:
# Large model: $0.10/request
return call_model("gpt-4o", prompt)A classifier routes 80% of requests to cheap models, saving 70%+ on inference costs.
Spot Instances for Training
Use spot/preemptible instances for fault-tolerant training:
# PyTorch checkpoint for spot instance resilience
def save_checkpoint(model, optimizer, epoch, path):
torch.save({
'epoch': epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
}, path)
# Resume from checkpoint after interruption
checkpoint = torch.load(path)
model.load_state_dict(checkpoint['model_state_dict'])
optimizer.load_state_dict(checkpoint['optimizer_state_dict'])Implement checkpointing every 10-30 minutes to minimize lost work on preemption.
Caching and Batching
Reduce redundant computation:
- Semantic caching — Cache responses for similar prompts (save 30-50% on repeated queries)
- Request batching — Batch inference requests for better GPU utilization
- KV cache optimization — Reduce memory usage for long-context inference
- CDN for embeddings — Cache embedding vectors for frequently accessed documents
Get weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Cost Monitoring Dashboard
Track these metrics:
- Cost per inference request — By model, by endpoint
- GPU utilization percentage — Target 80%+ during business hours
- Spot vs. on-demand ratio — Higher spot = lower costs
- Model efficiency — Tokens per second per dollar
- Idle GPU hours — Waste that can be eliminated with better scheduling
FAQ
What's a reasonable AI infrastructure budget? Depends on scale. Startups: $5K-50K/month. Mid-market: $50K-500K/month. Enterprise: $500K-10M+/month.
Should I use cloud GPUs or buy hardware? Cloud for variable workloads and experimentation. Own hardware when sustained utilization exceeds 60% (breakeven typically 12-18 months).
How much can optimization actually save? Typically 40-70% reduction through quantization, routing, caching, and spot instances combined.
---
Ready to go deeper?
This article is part of a hands-on learning path. Continue building your skills with our course catalog on CopyPasteLearn.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
AI Platform Engineering Explained
Learn what AI platform engineering is, why enterprises need it, and how to build production-grade GenAI infrastructure from scratch with proven DevOps.
What is Context7?
Discover Context7, the tool that gives version-specific, accurate documentation to LLMs and AI code editors like Cursor and Claude. No more hallucinated APIs.
Context7 + Cursor: Stop AI Errors
Learn how to use Context7 with Cursor AI editor for accurate, version-specific code completions. Step-by-step setup and workflow guide.
AI-Native Software Development
Explore AI-native software development practices including AI-assisted coding, automated testing, intelligent code review, and AI-driven architecture.
AI Security Platform Engineering
Build secure AI platforms with guardrails, prompt injection defense, model access controls, and observability for production LLM deployments.
AI Supercomputing Infrastructure
Explore how GPU clusters and AI supercomputing infrastructure power modern ML training with Kubernetes orchestration and cost optimization.
Explore topics
Browse more articles on the topics covered here.