Resumo da vaga

Senior Platform Engineer, Network Infrastructure- DGX Cloud

Requisitos e responsabilidades

Conteúdo da vaga extraído em seções para revisão mais rápida.

What You’ll Be Doing:

  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
  • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
  • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
  • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.

What We Need to See:

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • Proficiency in at least one general-purpose programming language, such as Go or Python.
  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
  • Experience deploying and supporting network automation or telemetry services on Kubernetes.
  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.

Ways to Stand Out From the Crowd:

  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.
  • Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.
  • Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.
  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
  • Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.
Vagas similares

Mantenha uma lista reserva.

Ver stack
FocoPlatform EngineeringÁrea da vaga
Sinal de senioridadeSeniorNível do candidato
StackCI/CD, Kubernetes, PythonSkills principais
Localização1 país aceitoElegibilidade

Stack

Use estas tags para comparar vagas remotas similares.

Elegibilidade de localização

Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.

Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.

Fluxo de contratação

O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.

1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.