Resumen del rol

Infrastructure Ops Engineer

Requisitos y responsabilidades

Contenido del rol extraído en secciones para revisar más rápido.

Details

  • The "Lost Node" Investigation: Debugging cluster-level blockers to solve why pods aren't scheduling despite available capacity
  • Regional Compliance Guard: Auditing and correcting scheduling policies to ensure customer data stays within specified geographical constraints (e.g., EU-only vs US-only)
  • High-Stakes Maintenance Orchestration: Coordinating critical maintenance cycles both externally (with vendors) and internally (with Baseten SREs) to evacuate workloads from unhealthy nodes and integrate replacement hardware with zero customer disruption
  • Fleet Maintenance: Manage daily node operations including tainting/untainting, node draining, and PVC repairs to ensure GPU fleet health and operational cost control
  • GTM & Capacity Fulfillment: Partner with Sales and account teams to scope and fulfill customer capacity requests, translating complex timelines into concrete infrastructure actions and clear ETAs
  • Process & Observability Engineering: Identify recurring gaps in the capacity lifecycle (intake, triage, comms) and drive fixes by defining lightweight processes and improving system observability
  • Technical Orchestration: Act as the operational bridge between SRE and Infra teams, executing discrete changes and verifying system status during high-stakes maintenance windows
  • Technical Documentation: Contribute to the internal knowledge base for GPU-specific issues (H100/A100/B200) to accelerate future incident resolution
  • Automation & Tooling: Identify repetitive workflows and partner with engineering to build scripts, dashboards, and internal tools that reduce manual intervention and shorten time-to-mitigation
  • Knowledge Excellence: Maintain a living database of GPU-specific intelligence (H100/B200) and market moves to accelerate incident resolution and support strategic briefings for leadership
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field
  • 2+ years of professional work experience, ideally in a customer-facing technical role or as a junior SRE/Cloud Engineer
  • Strong familiarity with Kubernetes and the lifecycle of cloud-based container orchestration
  • Strong ownership mindset and attention to detail, demonstrated through fast detection, clear communication, and reliable follow-through
  • Demonstrated ability to communicate complex technical blockers clearly to both internal engineering teams and external vendors
  • Preference for SF or NYC-based candidates to foster a close-knit "family" atmosphere in the office
  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Roles similares

Mantén una lista de respaldo.

Ver stack
FocoComputeÁrea del rol
Señal de seniorityNivel abiertoNivel del candidato
StackKubernetes, SparkSkills principales
Ubicación38 países aceptadosElegibilidad

Stack

Usa estas tags para comparar roles remotos similares.

Elegibilidad de ubicación

Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.

Flujo de contratación

WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.

1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.