Job Description: Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models Automate build, test, deployment, and rollback processes Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker Manage and provision specialized computing resources, including GPU clusters, for high-performance AI worklo and model inferencing Own high-availability design in production environments Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning Champion Infrastructure as Code practices using Terraform, Ansible, and Helm Architect and refine monitoring, logging, and alerting systems using Prometheus, Grafana, and ELK/EFK stack Build the Internal Developer Platform and golden paths enabling product, model, and data-science teams to deploy without opening a ticket Collaborate with R&D, Data Science, Security, and Business teams to streamline workflows and eliminate bottlenecks Establish and enforce system stability and security standards Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure SOC2/ISO27001 compliance Lead troubleshooting, root-cause analysis, and preventative remediation during complex anomalies and major incidents Convert incident learnings into automation to prevent recurrence Requirements: Bachelor's degree or above in Computer Science, Engineering, or a related technical field 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles Expert-level knowledge of Linux operating systems and core networking principles, including TCP/IP, DNS, HTTP, Load Balancing, and VPCs Deep mastery of Docker and Kubernetes orchestration, including cluster management and production-level best practices Proficiency designing and managing infrastructure on major public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud Experience with multi-cloud and hybrid-cloud strategies Strong coding and scripting capabilities in at least one major language such as Go, Python, or Shell Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), observability paradigms, and SRE principles Exceptional problem-solving abilities and sharp technical judgment Excellent cross-team communication skills Preferred: Familiarity with MLOps practices, model serving/inferencing frameworks such as vLLM, TGI, or Triton Inference Server Preferred: Experience managing GPU clusters for AI/ML worklo Preferred: Experience with large-scale distributed systems or high-concurrency environments Preferred: Hands-on experience designing and building Internal Developer Platforms (IDP) Preferred: Familiarity with Zero Trust architecture, automated security testing (DevSecOps), SOC2, or ISO27001 Preferred: Prior experience as a Technical Lead, mentoring junior engineers, or managing DevOps teams Preferred: Experience wiring an LLM-driven code/config helper into a pipeline or strong opinions on how to Must comply with applicable work authorization and equal employment requirements in the relevant country, state, and local jurisdictions Benefits: