NTT DATA
Senior Site Reliability Engineer (Remote, TN, IN)
Rol remoto de Site Reliability Engineer con fit claro de ubicación del candidato.
Publicado19 jul 2026
Países elegibles2 países aceptados
Señal de senioritySenior
Modelo de trabajoRemoto
Ubicaciones aceptadas para candidatos
IndiaTúnez
Resumen del rol
Senior Site Reliability Engineer (Remote, TN, IN)
Requisitos y responsabilidades
Contenido del rol extraído en secciones para revisar más rápido.
Key Responsibilities
- Delivery extensive CI/CD pipeline development and support, including Terraform-based pipelines, Azure DevOps pipelines, and multi-language build pipelines (Go, Python, .NET)
- Build and support CI/CD platforms and custom build agent ecosystems, combining Terraform-driven infrastructure provisioning with Azure DevOps pipelines and multi-language (Go, Python, .NET) build automation for consistent and reliable releases.
- Lead Terraform automation and infrastructure-as-code initiatives, creating reusable modules for Pub/Sub, GCS, BigQuery, Kubernetes, and Artifact repositories
- Manage event-driven architecture, including large-scale Pub/Sub topic, subscription, and messaging configurations across environments
- Implement security and identity solutions, including Auth0 pipeline automation, workload identity, IAM roles, and service principal migrations
- Delivery GCP networking and integration setups, including PSC connections, service attachments, and cross-service communication enablement
- Lead platform modernization and automation improvements, including migration of resources to Terraform and decommissioning of legacy configurations
- Manage ADO agent lifecycle and platform engineering, including agent migrations, upgrades, vulnerability remediation, and ARM-based scaling
- Enable data and caching solutions, including Valkey (Redis), Firestore integrations, and distributed data workflows
- Enhance pipeline security and compliance, including Wiz scans, CVE remediation, SSL fixes, and secure base image adoption
- Strengthening platform reliability and developer experience through troubleshooting pipelines, optimizing performance, and improving onboarding automation
- Required to work EST hours, approximately 8:00 AM to 5:00 PM.
- SRE team supporting mission‑critical applications; 24/7 on‑call rotation required. The shift is 9:00 AM – 9:00 PM EST.
- Design and operate highly available, scalable cloud infrastructure in GCP and Azure.
- Drive SRE best practices including SLOs, SLIs, error budgets, and incident management.
- Build and evolve internal developer platforms to enable self-service and accelerate delivery.
- Manage and optimize Kubernetes environments (GKE/OpenShift), including operators and service mesh.
- Implement Infrastructure as Code using Terraform and Config Connector.
- Develop CI/CD pipelines and GitOps workflows using Argo CD and Azure Pipelines.
- Enhance observability through monitoring, logging, and tracing (Prometheus, Grafana, Dynatrace).
- Automate operational workflows using AI, Python, shell scripting, and Ansible.
- Implement security best practices including secrets management (Vault) and policy enforcement.
- Incident response, reliability improvements, and postmortem analysis.
Required Qualifications
- 5+ years experience in SRE, DevOps, or cloud engineering roles.
- Strong expertise in Kubernetes and container platforms.
- Experience with Terraform and infrastructure automation.
- Proficiency in Python, Bash, or similar scripting languages.
- Experience with CI/CD and GitOps methodologies.
- Strong understanding of observability and monitoring tools.
- Ability to lead cross-functional initiatives and mentor engineers.
Preferred Qualifications
- Akamai (edge and CDN services)
- Azure Pipelines (CI/CD) and pipeline maintenance
- Google Config Connector
- Terraform (Infrastructure as Code)
- Python scripting for automation
- Shell scripting
- Ansible and AWX Tower
- Linux system administration
- Docker containerization
- Google Cloud Platform (GCP)
- Kubernetes administration (GKE)
- OpenShift (OCP4) administration
- Grafana and Loki (dashboard and log management)
- Istio service mesh
- Gatekeeper policy management
- Kiali (service mesh observability)
- Argo CD (GitOps deployments)
- Argo Workflows
- Apache web server configuration (URL rewrites, domain management)
- Apigee API management
- Google Cloud Firestore
- Google Bigtable
- Google AlloyDB
- Google BigQuery
- Google Pub/Sub messaging
- Apache Kafka (Strimzi)
- Confluent Kafka
- Apache Solr
- Zookeeper
- HashiCorp Vault (secrets management)
- SOPS (Secrets Operations)
- external-secrets-operator
- cert-manager
- Prometheus operator
- OpenTelemetry operator
- Strimzi operator
- Solr operator
- Crane (container tooling)
- WebMethods
- Camunda (workflow automation)
- Dynatrace (dashboard setup and monitoring)
- Dynatrace operator
- PagerDuty (alerting and incident management)
Roles similares
Mantén una lista de respaldo.
Stack
Usa estas tags para comparar roles remotos similares.
Elegibilidad de ubicación
Candidatos deberían aplicar solo cuando el país del perfil aparece aquí.
Flujo de contratación
WithMira muestra el rol y luego envía candidatos a la aplicación de la empresa.
1Revisa fit del rol, stack y elegibilidad de ubicación en WithMira.
2Abre la página de aplicación de la empresa desde el link rastreado.
3Guarda el rol o suscríbete a oportunidades similares antes de salir.