Resumen del rol

Site Reliability / Production Engineer

Requisitos y responsabilidades

Contenido del rol extraído en secciones para revisar más rápido.

Working Pattern

  • The two hires will share an agreed rota that ensures one engineer is actively working from 17:00-01:00 UK time, seven days a week.
  • Neither person will work seven days a week; the active shifts will be divided between both hires, with appropriate rest days.
  • The two engineers will also share weekday pager coverage from 01:00-06:00 UK time. The detailed allocation of active and pager shifts will be explained during the hiring process.
  • Weekend daytime coverage is provided separately and is not an additional expectation for these roles.
  • When there are no live incidents, the active shift will be used for reliability-improvement work.
  • Meetings and collaboration with management and product teams will be arranged within the agreed working pattern.

Respond to live incidents

  • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state.
  • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code.
  • Use AI throughout triage and diagnosis while checking its conclusions against real evidence.
  • Choose and execute a proportionate mitigation, rollback, repair or bounded fix.
  • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green.
  • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication.
  • Join customer conversations occasionally when direct technical involvement is genuinely useful.

Coordinate the right response

  • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix.
  • Escalate with evidence, customer impact, actions already taken and the specific decision or help required.
  • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait.
  • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.

Improve the reliability system

  • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes.
  • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements.
  • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve.
  • Create safe, supervised automation for common operational actions.
  • Work with product teams to close observability, rollback, runbook and supportability gaps.
  • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner.
  • Make reliability and on-call performance easier for the company to understand and improve over time.

What we're looking for

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed.
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand.
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it.

What we're looking for

  • Cloud platforms such as Azure or Cloudflare.
  • Distributed application and API diagnostics.
  • Databases, queues and background-processing systems.
  • Observability, alerting and incident-management platforms.
  • Infrastructure, deployment and release automation.
  • Application development and safe production debugging.
  • AI coding agents and workflow automation.
Roles similares

Mantén una lista de respaldo.

Ver stack
FocoSite Reliability EngineeringÁrea del rol
Señal de senioritySeniorNivel del candidato
StackAzure, RESTSkills principales
Ubicación2 países aceptadosElegibilidad

Stack

Usa estas tags para comparar roles remotos similares.

Elegibilidad de ubicación

Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.

Tu perfilPaís no definidoInicia sesión para comparar tu país con este rol.

Flujo de contratación

WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.

1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.