NVIDIA
Senior Site Reliability Engineer
Rol remoto de Site Reliability Engineering con fit claro de ubicación del candidato.
Publicado18 jul 2026
Países elegibles1 país aceptado
Señal de senioritySenior
Modelo de trabajoRemoto
Ubicaciones aceptadas para candidatos
India
Resumen del rol
Senior Site Reliability Engineer
Requisitos y responsabilidades
Contenido del rol extraído en secciones para revisar más rápido.
What You Will Be Doing
- Monitor, support, and maintain the reliability, availability, and performance of large-scale GeForce NOW production services running across cloud and datacenter environments.
- Participate in production incident triage, troubleshooting, and resolution of complex infrastructure and application issues. Take part in the team's on-call rotation, including occasional weekend coverage, to ensure timely restoration of customer-facing services.
- Monitor service health using metrics, logs, traces, and dashboards, and proactively identify reliability, performance, and capacity issues before they impact customers.
- Collaborate with software engineering, platform, networking, and infrastructure teams to improve operational readiness, reliability, and service resilience.
- Drive observability initiatives by improving monitoring, alerting, dashboards, and telemetry to enable faster detection and diagnosis of production issues.
- Scale services sustainably by building automation, eliminating operational toil, and continuously improving deployment, recovery, and operational workflows.
- Lead and participate in incident response, root cause analysis, and blameless postmortems, driving corrective and preventive actions to improve long-term service reliability.
- Design and develop custom tools, automation, and self-service solutions that simplify operations, improve engineer productivity, and enhance the overall GeForce NOW platform.
- Continuously evaluate existing operational processes and identify opportunities to improve service reliability, operational efficiency, and customer experience through engineering-driven solutions.
- Contribute to the design, deployment, and operation of Kubernetes-based services, ensuring they meet scalability, reliability, and performance requirements.
What we need to see:
- BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field (or equivalent experience).
- 5+ years of experience supporting and operating mission-critical production services in a live-site environment as a Site Reliability Engineer (SRE), Production Engineer, or similar role.
- Strong understanding of containerization, microservices architecture, and Kubernetes, including Kubernetes ecosystem components and operational best practices.
- Demonstrated ability to troubleshoot complex production issues, identify root causes, and drive issues to resolution.
- Strong understanding of distributed systems and how complex production environments interact across applications, infrastructure, networking, and cloud services.
- Experience supporting production operations, including incident management, change management, postmortem reviews, and operational excellence initiatives.
- Hands-on experience developing automation using Python, Go, Bash, or similar scripting/programming languages.
- Strong understanding of SLOs, SLIs, error budgets, KPIs, and service reliability best practices.
- Experience with observability platforms such as Prometheus, Grafana, ELK/OpenSearch, and modern monitoring and alerting solutions.
- Experience operating services in public cloud environments such as AWS, Azure, GCP, or equivalent cloud platforms.
Ways to stand out from the crowd:
- Experience supporting large-scale customer-facing cloud or gaming services.
- Strong Kubernetes operational and troubleshooting expertise.
- Experience with observability platforms, including Prometheus, Grafana, ELK/OpenSearch, and OpenTelemetry.
- Strong scripting or programming skills in Python, Go, or similar languages with a focus on automation.
- Experience driving production incident response, postmortems, and operational excellence initiatives.
Roles similares
Mantén una lista de respaldo.
AWS, Kubernetes India
Lead Platform EngineerAppianVer rol AWS, Azure 6 países aceptados
Senior Technical Program ManagerMorgan StanleyVer rol AWS, Kubernetes 13 países aceptados
Senior Backend Engineer (AdTech)Leap ToolsVer rol AWS, Kubernetes 13 países aceptados
Senior Backend EngineerLeap ToolsVer rol Stack
Usa estas tags para comparar roles remotos similares.
Elegibilidad de ubicación
Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.
Tu perfilPaís no definidoInicia sesión para comparar tu país con este rol.
Flujo de contratación
WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.
1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.