Associate Principal - Cloud Engineering
Albany - New York - USAOn-siteFull-timeCloud & DevOps
Description
100% Remote Max Salary $164K+Benefits Role Summary The NVIDIA GPU Platform Engineer builds and operates the GPU and inferenceserving stack NVIDIA AI Enterprise on EKS NIM microservices and NVIDIA Dynamo serving and owns GPU sharing scheduling integration and serving performance Onshore placement supports hands-on GPUedge coordination and close work with the architecture team Key Responsibilities Deploy and operate NVIDIA AI Enterprise NVAIE on EKS GPU Operator drivers CUDA runtime and DCGM Deploy and operate NIM microservices and NVIDIA Dynamo serving exposing the OpenAIcompatible API chatcompletions embeddings streaming batch Configure GPU sharing MIG where the SKU supports it and Run AI fractional GPU as the isolation baseline and clearly distinguish hardware vs software isolation Integrate Run AI for GPU scheduling quota and preemption and KEDA for replica autoscaling including scaletozero Instrument GPU telemetry via DCGM and drive inference performance validation and SLAtier binding Support the modelonboarding NIM packaging and inferencedeploy pipelines Tune GPUserving performance and troubleshoot CUDA driver and scheduling issues Must Have Skills Experience 8 years in infrastructure ML engineering with hands-on NVIDIA GPU operations NVIDIA GPU stack drivers CUDA GPU Operator and DCGM Model serving Triton NVIDIA Dynamo NIM Tensor RTLLM andor v LLM Kubernetes GPU workloads and device plugins GPU scheduling Run AI or equivalent and GPU partitioning MIG fractional GPU Autoscaling KEDA and OpenAIcompatible inference API patterns Niceto Have Skills NVAIE on AWSEKS and edge GPU Outposts G7 LLM RAG serving and vector databases Inference performance benchmarking and optimization
Requirements
Mandatory Skills : AWS EKS
Good to Have Skills : AWS Cloud Architecture, AWS Cloud Formation