aspenviewtech logo

Senior Site Reliability Engineer

aspenviewtech Colombia, Mexico, Costa Rica, Brazil, or Peru


No Relocation

Posted: August 18, 2026

Job Description

Why Join AspenView?

At AspenView, we’re more than a nearshore IT partner—we’re a people-first, purpose-driven company that believes great culture drives great outcomes. We’re passionate about connecting talent and technology to deliver measurable value for clients—and meaningful career paths for our people.

Here’s what you can expect:

  • Competitive base
  • Comprehensive benefits and wellness support
  • Flexible work model: hybrid, remote, or in-office
  • Real growth opportunities and leadership visibility
  • Inclusive, respectful culture that blends U.S. innovation with Colombian heart
  • A company that listens, invests in you, and celebrates wins together

About the Role

We’re looking for a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a full-time, remote opportunity open to candidates based in Colombia, Mexico, Costa Rica, Brazil, or Peru.

As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.

This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.

What You Will Do

Reliability Strategy & Metrics

  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.

Observability

  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.

Incident Management

  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

Cross-Functional Collaboration

  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.

Automation

  • Automate operational processes and reduce toil through engineering solutions.

Architecture & Capacity Planning

  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.

What You Bring

Education

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Experience

  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP environments.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Experience supporting large-scale cloud migration or modernization programs.
  • Expertise in incident management and production operations for high-availability systems.
  • Experience working in Agile, DevOps, or DevSecOps environments.

Technical Expertise

  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Experience implementing chaos engineering or resilience testing practices.
  • Knowledge of service mesh technologies such as Istio.

Certifications

  • AWS and/or GCP certifications are a plus.

Additional Content

Why Join AspenView?

At AspenView, we’re more than a nearshore IT partner—we’re a people-first, purpose-driven company that believes great culture drives great outcomes. We’re passionate about connecting talent and technology to deliver measurable value for clients—and meaningful career paths for our people.

Here’s what you can expect:

  • Competitive base
  • Comprehensive benefits and wellness support
  • Flexible work model: hybrid, remote, or in-office
  • Real growth opportunities and leadership visibility
  • Inclusive, respectful culture that blends U.S. innovation with Colombian heart
  • A company that listens, invests in you, and celebrates wins together

About the Role

We’re looking for a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a full-time, remote opportunity open to candidates based in Colombia, Mexico, Costa Rica, Brazil, or Peru.

As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.

This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.

What You Will Do

Reliability Strategy & Metrics

  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.

Observability

  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.

Incident Management

  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

Cross-Functional Collaboration

  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.

Automation

  • Automate operational processes and reduce toil through engineering solutions.

Architecture & Capacity Planning

  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.

What You Bring

Education

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Experience

  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP environments.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Experience supporting large-scale cloud migration or modernization programs.
  • Expertise in incident management and production operations for high-availability systems.
  • Experience working in Agile, DevOps, or DevSecOps environments.

Technical Expertise

  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Experience implementing chaos engineering or resilience testing practices.
  • Knowledge of service mesh technologies such as Istio.

Certifications

  • AWS and/or GCP certifications are a plus.