Resumo da vaga

Senior Site Reliability Engineer (Remote, TN, IN)

Requisitos e responsabilidades

Conteúdo da vaga extraído em seções para revisão mais rápida.

Key Responsibilities

  • Delivery extensive CI/CD pipeline development and support, including Terraform-based pipelines, Azure DevOps pipelines, and multi-language build pipelines (Go, Python, .NET)
  • Build and support CI/CD platforms and custom build agent ecosystems, combining Terraform-driven infrastructure provisioning with Azure DevOps pipelines and multi-language (Go, Python, .NET) build automation for consistent and reliable releases.
  • Lead Terraform automation and infrastructure-as-code initiatives, creating reusable modules for Pub/Sub, GCS, BigQuery, Kubernetes, and Artifact repositories
  • Manage event-driven architecture, including large-scale Pub/Sub topic, subscription, and messaging configurations across environments
  • Implement security and identity solutions, including Auth0 pipeline automation, workload identity, IAM roles, and service principal migrations
  • Delivery GCP networking and integration setups, including PSC connections, service attachments, and cross-service communication enablement
  • Lead platform modernization and automation improvements, including migration of resources to Terraform and decommissioning of legacy configurations
  • Manage ADO agent lifecycle and platform engineering, including agent migrations, upgrades, vulnerability remediation, and ARM-based scaling
  • Enable data and caching solutions, including Valkey (Redis), Firestore integrations, and distributed data workflows
  • Enhance pipeline security and compliance, including Wiz scans, CVE remediation, SSL fixes, and secure base image adoption
  • Strengthening platform reliability and developer experience through troubleshooting pipelines, optimizing performance, and improving onboarding automation
  • Required to work EST hours, approximately 8:00 AM to 5:00 PM.
  • SRE team supporting mission‑critical applications; 24/7 on‑call rotation required. The shift is 9:00 AM – 9:00 PM EST.
  • Design and operate highly available, scalable cloud infrastructure in GCP and Azure.
  • Drive SRE best practices including SLOs, SLIs, error budgets, and incident management.
  • Build and evolve internal developer platforms to enable self-service and accelerate delivery.
  • Manage and optimize Kubernetes environments (GKE/OpenShift), including operators and service mesh.
  • Implement Infrastructure as Code using Terraform and Config Connector.
  • Develop CI/CD pipelines and GitOps workflows using Argo CD and Azure Pipelines.
  • Enhance observability through monitoring, logging, and tracing (Prometheus, Grafana, Dynatrace).
  • Automate operational workflows using AI, Python, shell scripting, and Ansible.
  • Implement security best practices including secrets management (Vault) and policy enforcement.
  • Incident response, reliability improvements, and postmortem analysis.

Required Qualifications

  • 5+ years experience in SRE, DevOps, or cloud engineering roles.
  • Strong expertise in Kubernetes and container platforms.
  • Experience with Terraform and infrastructure automation.
  • Proficiency in Python, Bash, or similar scripting languages.
  • Experience with CI/CD and GitOps methodologies.
  • Strong understanding of observability and monitoring tools.
  • Ability to lead cross-functional initiatives and mentor engineers.

Preferred Qualifications

  • Akamai (edge and CDN services)
  • Azure Pipelines (CI/CD) and pipeline maintenance
  • Google Config Connector
  • Terraform (Infrastructure as Code)
  • Python scripting for automation
  • Shell scripting
  • Ansible and AWX Tower
  • Linux system administration
  • Docker containerization
  • Google Cloud Platform (GCP)
  • Kubernetes administration (GKE)
  • OpenShift (OCP4) administration
  • Grafana and Loki (dashboard and log management)
  • Istio service mesh
  • Gatekeeper policy management
  • Kiali (service mesh observability)
  • Argo CD (GitOps deployments)
  • Argo Workflows
  • Apache web server configuration (URL rewrites, domain management)
  • Apigee API management
  • Google Cloud Firestore
  • Google Bigtable
  • Google AlloyDB
  • Google BigQuery
  • Google Pub/Sub messaging
  • Apache Kafka (Strimzi)
  • Confluent Kafka
  • Apache Solr
  • Zookeeper
  • HashiCorp Vault (secrets management)
  • SOPS (Secrets Operations)
  • external-secrets-operator
  • cert-manager
  • Prometheus operator
  • OpenTelemetry operator
  • Strimzi operator
  • Solr operator
  • Crane (container tooling)
  • WebMethods
  • Camunda (workflow automation)
  • Dynatrace (dashboard setup and monitoring)
  • Dynatrace operator
  • PagerDuty (alerting and incident management)
Vagas similares

Mantenha uma lista reserva.

Ver stack
FocoSite Reliability EngineerÁrea da vaga
Sinal de senioridadeSeniorNível do candidato
StackAzure, CI/CD, DockerSkills principais
Localização2 países aceitosElegibilidade

Stack

Use estas tags para comparar vagas remotas similares.

Elegibilidade de localização

Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.

Seu perfilPaís não definidoEntre para comparar seu país com esta vaga.

Fluxo de contratação

O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.

1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.