Resumo da vaga

Software Engineer III, Site Reliability

Requisitos e responsabilidades

Conteúdo da vaga extraído em seções para revisão mais rápida.

What you’ll be doing:

  • Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
  • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
  • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
  • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
  • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
  • Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
  • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
  • Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
  • Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
  • Coach team members and engineers across the org on reliability patterns and operational best practices

Qualifications to be successful in this role:

  • 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
  • Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
  • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
  • Proven track record leading incident response and building SLO-driven reliability practices.
  • Working fluency with observability tooling (Datadog is a plus)
  • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
  • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
  • The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
  • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
  • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
  • Chaos engineering or game-day experience is a plus
  • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

Values you’ll model

  • Be Kind and Care — build with empathy; assume positive intent; support teammates and members.
  • Live Good Health — champion healthy habits and balance in how we work and what we ship.
  • Be Data-Inspired — ground decisions in research and data; measure what matters.
  • Champion Change — lean into ambiguity; iterate, learn, and improve continuously.
  • Leave it Better than You Found It — raise quality, clarify systems, and document as you go.
  • Make It Happen — bias to action; deliver impact with craft and accountability.
Vagas similares

Mantenha uma lista reserva.

Ver stack
FocoSite Reliability EngineeringÁrea da vaga
Sinal de senioridadeSeniorNível do candidato
StackAWS, CI/CD, KubernetesSkills principais
Localização1 país aceitoElegibilidade

Stack

Use estas tags para comparar vagas remotas similares.

Elegibilidade de localização

Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.

Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.

Fluxo de contratação

O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.

1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.