Resumen del rol

Software Engineer III, Site Reliability

Requisitos y responsabilidades

Contenido del rol extraído en secciones para revisar más rápido.

What you’ll be doing:

  • Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
  • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
  • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
  • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
  • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
  • Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
  • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
  • Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
  • Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
  • Coach team members and engineers across the org on reliability patterns and operational best practices

Qualifications to be successful in this role:

  • 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
  • Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
  • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
  • Proven track record leading incident response and building SLO-driven reliability practices.
  • Working fluency with observability tooling (Datadog is a plus)
  • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
  • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
  • The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
  • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
  • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
  • Chaos engineering or game-day experience is a plus
  • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

Values you’ll model

  • Be Kind and Care — build with empathy; assume positive intent; support teammates and members.
  • Live Good Health — champion healthy habits and balance in how we work and what we ship.
  • Be Data-Inspired — ground decisions in research and data; measure what matters.
  • Champion Change — lean into ambiguity; iterate, learn, and improve continuously.
  • Leave it Better than You Found It — raise quality, clarify systems, and document as you go.
  • Make It Happen — bias to action; deliver impact with craft and accountability.
Roles similares

Mantén una lista de respaldo.

Ver stack
FocoSite Reliability EngineeringÁrea del rol
Señal de senioritySeniorNivel del candidato
StackAWS, CI/CD, KubernetesSkills principales
Ubicación1 país aceptadoElegibilidad

Stack

Usa estas tags para comparar roles remotos similares.

Elegibilidad de ubicación

Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.

Tu perfilPaís no definidoInicia sesión para comparar tu país con este rol.

Flujo de contratación

WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.

1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.