Role overview

Senior Site Reliability Engineer (Remote, TN, IN)

Requirements and responsibilities

Readable role content extracted into sections for faster review.

Key Responsibilities

  • Delivery extensive CI/CD pipeline development and support, including Terraform-based pipelines, Azure DevOps pipelines, and multi-language build pipelines (Go, Python, .NET)
  • Build and support CI/CD platforms and custom build agent ecosystems, combining Terraform-driven infrastructure provisioning with Azure DevOps pipelines and multi-language (Go, Python, .NET) build automation for consistent and reliable releases.
  • Lead Terraform automation and infrastructure-as-code initiatives, creating reusable modules for Pub/Sub, GCS, BigQuery, Kubernetes, and Artifact repositories
  • Manage event-driven architecture, including large-scale Pub/Sub topic, subscription, and messaging configurations across environments
  • Implement security and identity solutions, including Auth0 pipeline automation, workload identity, IAM roles, and service principal migrations
  • Delivery GCP networking and integration setups, including PSC connections, service attachments, and cross-service communication enablement
  • Lead platform modernization and automation improvements, including migration of resources to Terraform and decommissioning of legacy configurations
  • Manage ADO agent lifecycle and platform engineering, including agent migrations, upgrades, vulnerability remediation, and ARM-based scaling
  • Enable data and caching solutions, including Valkey (Redis), Firestore integrations, and distributed data workflows
  • Enhance pipeline security and compliance, including Wiz scans, CVE remediation, SSL fixes, and secure base image adoption
  • Strengthening platform reliability and developer experience through troubleshooting pipelines, optimizing performance, and improving onboarding automation
  • Required to work EST hours, approximately 8:00 AM to 5:00 PM.
  • SRE team supporting mission‑critical applications; 24/7 on‑call rotation required. The shift is 9:00 AM – 9:00 PM EST.
  • Design and operate highly available, scalable cloud infrastructure in GCP and Azure.
  • Drive SRE best practices including SLOs, SLIs, error budgets, and incident management.
  • Build and evolve internal developer platforms to enable self-service and accelerate delivery.
  • Manage and optimize Kubernetes environments (GKE/OpenShift), including operators and service mesh.
  • Implement Infrastructure as Code using Terraform and Config Connector.
  • Develop CI/CD pipelines and GitOps workflows using Argo CD and Azure Pipelines.
  • Enhance observability through monitoring, logging, and tracing (Prometheus, Grafana, Dynatrace).
  • Automate operational workflows using AI, Python, shell scripting, and Ansible.
  • Implement security best practices including secrets management (Vault) and policy enforcement.
  • Incident response, reliability improvements, and postmortem analysis.

Required Qualifications

  • 5+ years experience in SRE, DevOps, or cloud engineering roles.
  • Strong expertise in Kubernetes and container platforms.
  • Experience with Terraform and infrastructure automation.
  • Proficiency in Python, Bash, or similar scripting languages.
  • Experience with CI/CD and GitOps methodologies.
  • Strong understanding of observability and monitoring tools.
  • Ability to lead cross-functional initiatives and mentor engineers.

Preferred Qualifications

  • Akamai (edge and CDN services)
  • Azure Pipelines (CI/CD) and pipeline maintenance
  • Google Config Connector
  • Terraform (Infrastructure as Code)
  • Python scripting for automation
  • Shell scripting
  • Ansible and AWX Tower
  • Linux system administration
  • Docker containerization
  • Google Cloud Platform (GCP)
  • Kubernetes administration (GKE)
  • OpenShift (OCP4) administration
  • Grafana and Loki (dashboard and log management)
  • Istio service mesh
  • Gatekeeper policy management
  • Kiali (service mesh observability)
  • Argo CD (GitOps deployments)
  • Argo Workflows
  • Apache web server configuration (URL rewrites, domain management)
  • Apigee API management
  • Google Cloud Firestore
  • Google Bigtable
  • Google AlloyDB
  • Google BigQuery
  • Google Pub/Sub messaging
  • Apache Kafka (Strimzi)
  • Confluent Kafka
  • Apache Solr
  • Zookeeper
  • HashiCorp Vault (secrets management)
  • SOPS (Secrets Operations)
  • external-secrets-operator
  • cert-manager
  • Prometheus operator
  • OpenTelemetry operator
  • Strimzi operator
  • Solr operator
  • Crane (container tooling)
  • WebMethods
  • Camunda (workflow automation)
  • Dynatrace (dashboard setup and monitoring)
  • Dynatrace operator
  • PagerDuty (alerting and incident management)
Similar roles

Keep a backup shortlist.

Browse stack
FocusSite Reliability EngineerRole area
Seniority signalSeniorCandidate level
StackAzure, CI/CD, DockerPrimary skills
Location2 accepted countriesEligibility

Stack

Use these tags to compare similar remote roles.

Location eligibility

Candidates should apply only when their profile country is listed here.

Your profileCountry not setSign in to check your country against this role.

Hiring flow

WithMira shows the role, then sends candidates to the company application.

1Check role fit, stack, and location eligibility in WithMira.
2Open the company application page from the tracked apply link.
3Save the role or subscribe for similar opportunities before leaving.