zscaler logo

Sr. Staff Software Development Engineer - AI Platform

zscaler Bangalore, IND; Pune, IND


No Relocation

Posted: August 4, 2026

Job Description

Role

We are looking for a Sr. Staff AI Platform Engineer to join our team. This is an On-site role based in Bangalore/Pune, reporting to the Senior Manager in the IT Data Strategy department. In this position, you will design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads. Operating within the IT Data Strategy team, you will drive infrastructure automation, observability, and platform governance while empowering AI and data engineering teams to deliver high-impact solutions reliably and securely.

What you’ll do (Role Expectations)

  • Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints

  • Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases

  • Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR

  • Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)

  • Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews

Who You Are (Success Profile)

  • You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution.

  • You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust.

  • You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges.

  • You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication.

  • You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability.

What We’re Looking for (Minimum Qualifications)

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving

  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)

  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation

  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services

  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement

  • Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer

What Will Make You Stand Out (Preferred Qualifications)

  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex

  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix

  • Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments

#LI-RR1

#LI-Hybrid

Additional Content

Role

We are looking for a Sr. Staff AI Platform Engineer to join our team. This is an On-site role based in Bangalore/Pune, reporting to the Senior Manager in the IT Data Strategy department. In this position, you will design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads. Operating within the IT Data Strategy team, you will drive infrastructure automation, observability, and platform governance while empowering AI and data engineering teams to deliver high-impact solutions reliably and securely.

What you’ll do (Role Expectations)

  • Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints

  • Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases

  • Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR

  • Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)

  • Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews

Who You Are (Success Profile)

  • You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution.

  • You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust.

  • You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges.

  • You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication.

  • You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability.

What We’re Looking for (Minimum Qualifications)

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving

  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)

  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation

  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services

  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement

  • Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer

What Will Make You Stand Out (Preferred Qualifications)

  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex

  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix

  • Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments

#LI-RR1

#LI-Hybrid