Resumo da vaga

Senior Site Reliability Engineer

Requisitos e responsabilidades

Conteúdo da vaga extraído em seções para revisão mais rápida.

Reliability assessment and remediation

  • Assess the reliability of web applications, Flutter mobile services, APIs, backend systems, infrastructure, databases, queues, and third-party integrations.
  • Identify single points of failure, fragile dependencies, manual operational processes, and failure modes.
  • Distinguish confirmed issues from suspected risks and areas that have not yet been evaluated.
  • Design and implement shared reliability capabilities and improvements within the SRE domain.
  • Define application-specific reliability changes and work with system-owning teams to implement them; those teams retain responsibility for their code and services.
  • Transfer service-specific instrumentation, runbooks, and ongoing maintenance responsibilities to the appropriate system owners after the solution is tested and hardened.
  • Track identified risks through implementation and validation rather than stopping at recommendations.

Performance, load, and capacity

  • Establish a practical performance and capacity-testing program.
  • Work with QA and product stakeholders to identify critical workflows and realistic usage scenarios.
  • Establish baseline response times, throughput, concurrency, and resource consumption.
  • Design and execute load, stress, endurance, scalability, and failure tests.
  • Identify bottlenecks involving applications, databases, networks, queues, caches, infrastructure, and external services.
  • Implement shared SRE mitigations and coordinate application-specific mitigation work with system-owning teams, which retain responsibility for their code and services.
  • Repeat testing after changes to verify results.
  • Document tested capacity, observed constraints, and remaining unknowns.
  • Define safe operating limits and early-warning indicators.

Observability

  • Assess current logging, metrics, tracing, health checks, dashboards, and alerting.
  • Establish consistent observability standards across services.
  • Implement and harden shared health-checking and observability capabilities, transfer shared components to the designated long-term owner, and work with system-owning teams on service-specific instrumentation that they will maintain after handoff.
  • Define meaningful service-level indicators and initial reliability objectives with system owners and engineering leadership.
  • Build shared reliability dashboards and initial service-specific views, then train system owners to maintain their service-specific dashboards and alerts.
  • Ensure alerts are actionable, routed to accountable owners, and tested.
  • Identify monitoring blind spots.
  • Improve application instrumentation in collaboration with developers.
  • Ensure logs and telemetry do not expose sensitive information.

Incident readiness

  • Help establish the initial on-call and escalation model.
  • Create and test incident-response and troubleshooting runbooks.
  • Define severity levels and technical escalation paths.
  • Lead or support incident investigation.
  • Improve detection, diagnosis, mitigation, and restoration capabilities.
  • Facilitate technically focused post-incident reviews.
  • Track corrective actions and recurring failure patterns.
  • Conduct controlled failure exercises where appropriate.

SRE automation

  • Automate repetitive reliability-engineering work, including health checks, diagnostics, alert enrichment, incident triage, capacity checks, and evidence collection.
  • Develop safe automated remediation or self-healing for clearly defined and well-tested failure conditions.
  • Reduce manual diagnostic, maintenance, incident-response, and reliability-validation steps within the SRE function.
  • Define and implement the reliability checks, test logic, and SRE automation that should run through CI/CD. Work with DevOps to integrate them into the shared delivery framework, and with system-owning teams to maintain application-specific configuration after handoff.
  • Create reusable reliability tools and patterns, harden and document them, transfer shared components to the designated long-term owner, and train application teams to operate the service-specific portions they own.
  • Document automation ownership, safeguards, limitations, rollback behavior, and conditions requiring human intervention.

Collaboration

  • Work with the Disaster Recovery and Resilience Engineer on service and data recovery dependencies.
  • Work with Security Operations on the security review of SRE implementations before production adoption.
  • Provide reliability requirements for deployment safeguards and environment health, and work with DevOps and platform staff on shared integration. SRE does not own deployment automation or routine application deployments.
  • Work with QA to define realistic user scenarios and post-mitigation validation.
  • Work with application teams on system-specific code changes and instrumentation.
  • Provide factual technical findings to engineering leadership without presenting unverified assumptions as confirmed conclusions.

Initial Priorities

  • Inventory critical services and their dependencies.
  • Assess reliability risks and potential single points of failure.
  • Establish system-health and latency visibility.
  • Define critical workflows for performance testing.
  • Build the initial performance, load, and capacity-testing capability.
  • Establish technical baselines.
  • Identify and implement the highest-priority reliability and scaling mitigations.
  • Create initial runbooks and on-call recommendations.
  • Identify SRE processes that should be automated, including diagnostics, alert handling, capacity validation, incident response, and safe remediation.
  • Document tested behavior, unresolved risks, and unknowns.

Required Qualifications

  • Five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.
  • Hands-on experience supporting cloud-hosted applications and distributed systems.
  • Experience implementing reliability improvements rather than only producing assessments.
  • Experience with performance, load, stress, or capacity testing.
  • Experience diagnosing application, database, network, and infrastructure bottlenecks.
  • Strong knowledge of monitoring, logging, metrics, alerting, and incident response.
  • Experience with reliability automation and scripting.
  • Experience working with CI/CD systems.
  • Experience troubleshooting APIs, backend services, databases, queues, caches, and external integrations.
  • Understanding of timeout, retry, rate-limit, circuit-breaking, and graceful-degradation patterns.
  • Ability to work directly in unfamiliar codebases and environments.
  • Strong technical documentation and communication skills.
  • Ability to distinguish confirmed findings, suspected risks, and unknowns.
  • Comfort working in a small organization where operational practices are still being established.

Bonus Points

  • Experience preparing a platform for its first external users.
  • Experience establishing an SRE capability in a startup or small engineering organization.
  • Experience with financial, digital-asset, telemetry, equipment-monitoring, or operational platforms.
  • Experience working with Flutter-backed mobile services.
  • Experience with chaos or fault-injection testing.
  • Experience with data-intensive or event-driven systems.
  • Experience in a remote, asynchronous environment.

What Success Looks Like

  • Critical services and dependencies are inventoried.
  • Reliability risks are visible and prioritized.
  • Performance and capacity baselines exist for critical workflows.
  • High-priority reliability and scaling mitigations are implemented and tested.
  • Critical services have meaningful health checks, dashboards, alerts, and runbooks.
  • The company can identify and respond to failures without relying entirely on one person.
  • Application teams understand and maintain the reliability improvements made to their systems.
  • Remaining capacity and reliability risks are documented clearly.
Vagas similares

Mantenha uma lista reserva.

Ver stack
FocoSite Reliability EngineeringÁrea da vaga
Sinal de senioridadeSeniorNível do candidato
StackCI/CDSkills principais
Localização2 países aceitosElegibilidade

Stack

Use estas tags para comparar vagas remotas similares.

Elegibilidade de localização

Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.

Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.

Fluxo de contratação

O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.

1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.