Resumen del rol

Senior SRE

Requisitos y responsabilidades

Contenido del rol extraído en secciones para revisar más rápido.

Details

  • Sees a recurring alert or a fragile deploy path and cannot leave it alone. Excels at shipping the right fix and the right automation, not the perfect one.
  • Has run real production systems at scale — not just written runbooks for them.
  • Has built with Datadog, OpenTelemetry, incident tooling, and AI coding assistants long enough to have strong opinions about what fits our needs.
  • Can prototype an operational agent in Cursor and iterate as they go.
  • Operates with autonomy, and can carry a technical discussion on system architecture, failure modes, and tradeoffs.
  • Is genuinely curious about applying emerging AI to reliability and operations.
  • Own the reliability roadmap end to end. Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure.
  • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired.
  • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency).
  • Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop.
  • Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards.
  • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). Use Cursor as your build environment. Treat the agents as products solving specific problems.
  • Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.
  • Proven experience designing, operating, and scaling reliable production systems.
  • Deep hands-on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns.
  • Strong background defining SLIs, SLOs, and error budgets — and translating them into business-level KPIs, not just infrastructure metrics.
  • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps).
  • Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices.
  • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you've built and deployed agents.
  • Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business.
  • Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead.
  • Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.
  • Experience in insurance, fintech, or other regulated financial services industries.
  • Familiarity with insurance and finance concepts (premium, claims, settlement, reserving, monthly close) or willingness to learn them deeply.
  • Experience with streaming and event pipelines (Red Panda / Kafka) and data lineage, retention, and auditability requirements.
  • Strong working knowledge of chaos engineering, performance and load testing, and capacity planning.
  • Experience deploying AI agents on an internal AI platform or fabric (governance, eval harnesses, prompt/version management).
Roles similares

Mantén una lista de respaldo.

Ver stack
FocoSite Reliability EngineeringÁrea del rol
Señal de senioritySeniorNivel del candidato
StackAWS, SnowflakeSkills principales
Ubicación1 país aceptadoElegibilidad

Stack

Usa estas tags para comparar roles remotos similares.

Elegibilidad de ubicación

Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.

Tu perfilPaís no definidoInicia sesión para comparar tu país con este rol.

Flujo de contratación

WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.

1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.