Model Serving: Triton & NIM
NickPinaAbout This Course
A hands-on module for infrastructure engineers running models in production. Starting from serving fundamentals — latency vs. throughput, dynamic batching, model formats and the export step — it builds the full NVIDIA-aligned serving stack: Triton Inference Server (model repository, config.pbtxt, batching and instance groups, ensembles, perf_analyzer, Kubernetes deployment), LLM serving (prefill/decode, KV cache, continuous batching, vLLM and TensorRT-LLM, NVIDIA NIM and the OpenAI-compatible API layer, sizing and scaling), and production operations with KServe — autoscaling GPU inference, progressive delivery for models, three-tier observability, multi-model GPU density and economics, and the serving incident runbook.
Each of the four modules ends with a graded knowledge check of scenario-based questions. Best suited for DevOps/platform engineers who have completed the NCP-AIO study track and the MLOps Pipelines course, or have equivalent Kubernetes and GPU-cluster experience.