.jpg?1700169058)
Sr. Staff Software Development Engineer - AI Platform
zscaler • Bangalore, IND; Pune, IND
Posted: August 4, 2026
Job Description
Role
We are looking for a Sr. Staff AI Platform Engineer to join our team. This is an On-site role based in Bangalore/Pune, reporting to the Senior Manager in the IT Data Strategy department. In this position, you will design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads. Operating within the IT Data Strategy team, you will drive infrastructure automation, observability, and platform governance while empowering AI and data engineering teams to deliver high-impact solutions reliably and securely.
What you’ll do (Role Expectations)
-
Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
-
Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
-
Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
-
Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
-
Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews
Who You Are (Success Profile)
-
You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution.
-
You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust.
-
You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges.
-
You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication.
-
You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability.
What We’re Looking for (Minimum Qualifications)
-
Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
-
8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
-
Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
-
Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services
-
Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement
-
Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer
What Will Make You Stand Out (Preferred Qualifications)
-
Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex
-
Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix
-
Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments
#LI-RR1
#LI-Hybrid
Additional Content
Role
We are looking for a Sr. Staff AI Platform Engineer to join our team. This is an On-site role based in Bangalore/Pune, reporting to the Senior Manager in the IT Data Strategy department. In this position, you will design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads. Operating within the IT Data Strategy team, you will drive infrastructure automation, observability, and platform governance while empowering AI and data engineering teams to deliver high-impact solutions reliably and securely.
What you’ll do (Role Expectations)
-
Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
-
Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
-
Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
-
Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
-
Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews
Who You Are (Success Profile)
-
You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution.
-
You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust.
-
You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges.
-
You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication.
-
You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability.
What We’re Looking for (Minimum Qualifications)
-
Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
-
8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
-
Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
-
Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services
-
Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement
-
Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer
What Will Make You Stand Out (Preferred Qualifications)
-
Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex
-
Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix
-
Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments
#LI-RR1
#LI-Hybrid