Resumo da vaga

Site Reliability Engineer- Platform Engineering

Requisitos e responsabilidades

Conteúdo da vaga extraído em seções para revisão mais rápida.

Infrastructure & Platform Ownership

  • Design, implement, and maintain scalable infrastructure on Google Cloud Platform to support CodeRabbit's rapidly growing user base and processing demands
  • Develop, own and operate critical platform services
  • Build and maintain Infrastructure as Code using Terraform-Terragrunt to ensure consistent, reproducible, and version-controlled infrastructure deployments

Reliability & Performance Engineering

  • Establish and maintain SLI/SLO frameworks for all critical services, ensuring we meet our reliability commitments to users
  • Implement comprehensive monitoring, alerting, observability and incident management solutions to maintain a high reliability.
  • Conduct thorough incident response, root cause analysis, and post-mortem processes to continuously improve system reliability
  • Optimize application and infrastructure performance and cost to handle millions of pull request analyses effeciently.
  • Design and implement chaos engineering practices to proactively identify and resolve system weaknesses

Automation & Developer Experience

  • Develop self-service platforms and tooling that empower engineering teams to deploy, monitor, and troubleshoot their services independently
  • Automate operational tasks including scaling, backup/recovery, security patching, and routine maintenance
  • Create and maintain infrastructure APIs and abstractions that simplify complex operations for development teams

Security & Compliance

  • Integrate security best practices into all infrastructure and platform services
  • Implement and maintain security monitoring, vulnerability scanning, and compliance reporting
  • Design secure network architectures including VPC configuration, firewall rules, and access control systems
  • Establish and maintain disaster recovery procedures and business continuity planning

Experience & Background

  • 6-8 years of hands-on experience in Site Reliability Engineering, Platform Engineering, or DevOps Engineering roles
  • Proven track record of managing production systems at scale, preferably in high-growth technology companies
  • Strong background with cloud platforms, particularly Google Cloud Platform (GCP) or Amazon Web Services(AWS) including compute, storage, networking, and managed services
  • Experience in containerization and orchestration platforms (Kubernetes, Docker)

Technical Skills

  • Programming Languages: Proficiency in Node.js and TypeScript for building automation tools, monitoring solutions, and platform services
  • Infrastructure as Code: Advanced experience with Terraform for infrastructure provisioning and management
  • Monitoring & Observability: Hands-on experience with Datadog or similar platforms (Prometheus, Grafana, ELK stack) for observability
  • Cloud Platforms: Comprehensive experience with GCP services including Compute Engine, GKE, Cloud Run, Cloud SQL, Cloud Storage, Load Balancing, and IAM

Systems & Operations

  • Strong Linux/Unix systems skills
  • Experience with network protocols, load balancing, and CDN technologies
  • Knowledge of security principles and best practices for cloud infrastructure
  • Familiarity with CI/CD tools and practices (Jenkins GitHub Actions, Circle CI)
  • Understanding of microservices architecture and distributed systems principles
  • Investigate, troubleshooting and root cause complex production issues methodically and prevent recurrance

Preferred Qualifications

  • Experience with AI/ML infrastructure and tools
  • Background in managing high-traffic web applications and API services
  • Experience with disaster recovery planning and execution
  • Familiarity with compliance frameworks (SOC 2, ISO 27001)
  • Contributions to open-source infrastructure or SRE tooling projects
  • Experience with cost optimization and FinOps practices
  • Knowledge of performance testing and capacity planning methodologies

Technical Excellence

  • Strong problem-solving skills with the ability to debug complex distributed systems issues
  • Systematic approach to troubleshooting with excellent attention to detail
  • Passion for automation and eliminating toil through intelligent tooling and processes
  • Understanding of software engineering principles and ability to write production-quality code and develop tooling

Collaboration & Communication

  • Excellent communication skills with the ability to work effectively across engineering, product, and business teams
  • Ability to translate complex technical concepts into business impact and user value
  • Strong documentation skills and commitment to knowledge sharing

Growth Mindset

  • Enthusiasm for continuous learning and staying current with emerging technologies
  • Ability to thrive in a fast-paced, rapidly evolving startup environment
  • Proactive mindset with the ability to identify and solve problems before they impact users
  • Commitment to building inclusive, diverse, and collaborative team environments
Vagas similares

Mantenha uma lista reserva.

Ver stack
FocoEngineeringÁrea da vaga
Sinal de senioridadeNível abertoNível do candidato
StackAWS, CI/CD, DockerSkills principais
Localização1 país aceitoElegibilidade

Stack

Use estas tags para comparar vagas remotas similares.

Elegibilidade de localização

Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.

Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.

Fluxo de contratação

O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.

1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.