Baseten
Software Engineer- Training Infrastructure
Rol remoto de Training Platform con fit claro de ubicación del candidato.
Publicado29 ago 2025
Países elegibles1 país aceptado
Señal de seniorityNivel abierto
Modelo de trabajoRemoto
Ubicaciones aceptadas para candidatos
Estados Unidos
Resumen del rol
Software Engineer- Training Infrastructure
Requisitos y responsabilidades
Contenido del rol extraído en secciones para revisar más rápido.
Details
- Overview of the product so far
- Training docs overview
- Story of the Training product
- Research we've done
- Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking)
- Partner closely with developers and research engineers to translate complex training requirements into technical solutions
- Design and architect a global training scheduler
- Design and architect reinforcement learning systems and continuous learning pipelines
- Drive long-term improvements to improve reliability of systems and velocity of development
- Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure
- Make critical architectural decisions balancing performance with system reliability
- Lead technical discussions and mentor junior engineers on infrastructure best practices
- Contribute to long-term technical strategy and infrastructure roadmap
- Bachelor’s degree or high in Computer Science or related field
- Proficiency in Go, with Python experience a plus
- Deep expertise with Kubernetes in production environments
- Extensive experience with major cloud providers (AWS, GCP) and neo-cloud providers (Crusoe, DigitalOcean, Nebius) a plus
- Advanced understanding of distributed systems concepts and performance tuning
- Proven experience designing observability systems
- Experience with ML/AI workloads and MLOps platforms highly valued
- Experience with distributed storage systems
- Experience with workload orchestration platforms like Temporal or Airflow
- Familiarity or experience with the open source training stack and frameworks (NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainier) and distributed training techniques (FSDP, DeepSpeed).
- Experience developing AI products, tooling, or agents
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Roles similares
Mantén una lista de respaldo.
AWS, Kubernetes 1 país aceptado
Lead Platform EngineerAppianVer rol AWS, Kubernetes 13 países aceptados
Senior Backend Engineer (AdTech)Leap ToolsVer rol AWS, Kubernetes 13 países aceptados
Senior Backend EngineerLeap ToolsVer rol Python, Spark USA
Senior Data EngineerTop Us Wealth Management FirmVer rol Stack
Usa estas tags para comparar roles remotos similares.
Elegibilidad de ubicación
Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.
Tu perfilPaís no definidoInicia sesión para comparar tu país con este rol.
Flujo de contratación
WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.
1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.