Job Details - Find Your Perfect Career Opportunity

About the Role


Overview :


Roles & responsibilities:

● Deploy and manage ML/LLM models (<20B parameters) in production environments

● Design and execute model load testing and performance benchmarking (latency,

throughput, memory, cost)

● Build and optimize multi-node, multi-GPU training pipelines

● Configure and tune distributed training frameworks (data parallelism, model parallelism,

pipeline parallelism)

● Optimize GPU utilization, memory footprint, and inference costs

● Set up CI/CD pipelines for model deployment and retraining

● Troubleshoot GPU, networking, and performance bottlenecks

● Work across cloud platforms to ensure portability and vendor-agnostic deployments


Skill Sets:

● Strong experience with GPU workloads (NVIDIA GPUs, CUDA concepts)

● Proven expertise in model deployment on AWS, Azure, and GCP

● Hands-on experience deploying models up to 20B parameters

● Experience with distributed training (multi-node, multi-GPU setups)

● Deep understanding of load testing, stress testing, and benchmarking ML systems

Skills :

DevOps,AIOps,MLOps,kubernetes,genkins,docker,CI/CD,AWS,Azure