Senior Infrastructure Engineer - GPU Compute
Boundless Networks, Inc. • United States
Posted: August 3, 2026
Job Description
Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.
You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.
What You'll Do
GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.
Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.
Bare-Metal & GPU Optimization: Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node.
Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.
Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability.
Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally di...- 5+ years of infrastructure/DevOps experience operating large-scale production systems
- Deep expertise in Kubernetes, Docker, and container orchestration at scale
- Strong Linux systems administration skills
- Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
- Track record of managing mission-critical, high-throughput systems
- Strong infrastructure-as-code background in heterogeneous environments
- Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
- Comfort navigating ambiguity with a strong bias for action
Nice to Have
- Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
- Experience operating ML training or other large-scale distributed compute infrastructure
- Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
- Familiarity with fleet access and networking tooling (Tailscale, Teleport)
- Knowledge of network optimization and topology design
- Experience with multi-region, globally distributed systems
- Proficiency in Rust or low-level systems programming
- Experience with on-premises data center operations
Additional Requirements
- Candidates must include a public GitHub profile in their application.
- The GitHub profile should demonstrate a minimum of 1 year of activity/history.
- Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.
Additional Content
Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.
You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.
What You'll Do
GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.
Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.
Bare-Metal & GPU Optimization: Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node.
Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.
Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability.
Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally di...- 5+ years of infrastructure/DevOps experience operating large-scale production systems
- Deep expertise in Kubernetes, Docker, and container orchestration at scale
- Strong Linux systems administration skills
- Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
- Track record of managing mission-critical, high-throughput systems
- Strong infrastructure-as-code background in heterogeneous environments
- Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
- Comfort navigating ambiguity with a strong bias for action
Nice to Have
- Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
- Experience operating ML training or other large-scale distributed compute infrastructure
- Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
- Familiarity with fleet access and networking tooling (Tailscale, Teleport)
- Knowledge of network optimization and topology design
- Experience with multi-region, globally distributed systems
- Proficiency in Rust or low-level systems programming
- Experience with on-premises data center operations
Additional Requirements
- Candidates must include a public GitHub profile in their application.
- The GitHub profile should demonstrate a minimum of 1 year of activity/history.
- Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.