MyFitnessPal
Software Engineer III, Site Reliability
Vaga remota de Site Reliability Engineering com fit claro de localização do candidato.
Publicada18 de jul. de 2026
Países elegíveis1 país aceito
Sinal de senioridadeSenior
Modelo de trabalhoRemoto
Locais aceitos para candidatos
Estados Unidos
Resumo da vaga
Software Engineer III, Site Reliability
Requisitos e responsabilidades
Conteúdo da vaga extraído em seções para revisão mais rápida.
What you’ll be doing:
- Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
- Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
- Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
- Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
- Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
- Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
- Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
- Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
- Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
- Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
- Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
- Coach team members and engineers across the org on reliability patterns and operational best practices
Qualifications to be successful in this role:
- 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
- Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
- Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
- Proven track record leading incident response and building SLO-driven reliability practices.
- Working fluency with observability tooling (Datadog is a plus)
- Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
- Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
- The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
- Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
- Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
- Chaos engineering or game-day experience is a plus
- Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus
Values you’ll model
- Be Kind and Care — build with empathy; assume positive intent; support teammates and members.
- Live Good Health — champion healthy habits and balance in how we work and what we ship.
- Be Data-Inspired — ground decisions in research and data; measure what matters.
- Champion Change — lean into ambiguity; iterate, learn, and improve continuously.
- Leave it Better than You Found It — raise quality, clarify systems, and document as you go.
- Make It Happen — bias to action; deliver impact with craft and accountability.
Vagas similares
Mantenha uma lista reserva.
Stack
Use estas tags para comparar vagas remotas similares.
Elegibilidade de localização
Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.
Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.
Fluxo de contratação
O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.
1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.