NTT DATA
Senior Site Reliability Engineer (Remote, TN, IN)
Vaga remota de Site Reliability Engineer com fit claro de localização do candidato.
Publicada19 de jul. de 2026
Países elegíveis2 países aceitos
Sinal de senioridadeSenior
Modelo de trabalhoRemoto
Locais aceitos para candidatos
ÍndiaTunísia
Resumo da vaga
Senior Site Reliability Engineer (Remote, TN, IN)
Requisitos e responsabilidades
Conteúdo da vaga extraído em seções para revisão mais rápida.
Key Responsibilities
- Delivery extensive CI/CD pipeline development and support, including Terraform-based pipelines, Azure DevOps pipelines, and multi-language build pipelines (Go, Python, .NET)
- Build and support CI/CD platforms and custom build agent ecosystems, combining Terraform-driven infrastructure provisioning with Azure DevOps pipelines and multi-language (Go, Python, .NET) build automation for consistent and reliable releases.
- Lead Terraform automation and infrastructure-as-code initiatives, creating reusable modules for Pub/Sub, GCS, BigQuery, Kubernetes, and Artifact repositories
- Manage event-driven architecture, including large-scale Pub/Sub topic, subscription, and messaging configurations across environments
- Implement security and identity solutions, including Auth0 pipeline automation, workload identity, IAM roles, and service principal migrations
- Delivery GCP networking and integration setups, including PSC connections, service attachments, and cross-service communication enablement
- Lead platform modernization and automation improvements, including migration of resources to Terraform and decommissioning of legacy configurations
- Manage ADO agent lifecycle and platform engineering, including agent migrations, upgrades, vulnerability remediation, and ARM-based scaling
- Enable data and caching solutions, including Valkey (Redis), Firestore integrations, and distributed data workflows
- Enhance pipeline security and compliance, including Wiz scans, CVE remediation, SSL fixes, and secure base image adoption
- Strengthening platform reliability and developer experience through troubleshooting pipelines, optimizing performance, and improving onboarding automation
- Required to work EST hours, approximately 8:00 AM to 5:00 PM.
- SRE team supporting mission‑critical applications; 24/7 on‑call rotation required. The shift is 9:00 AM – 9:00 PM EST.
- Design and operate highly available, scalable cloud infrastructure in GCP and Azure.
- Drive SRE best practices including SLOs, SLIs, error budgets, and incident management.
- Build and evolve internal developer platforms to enable self-service and accelerate delivery.
- Manage and optimize Kubernetes environments (GKE/OpenShift), including operators and service mesh.
- Implement Infrastructure as Code using Terraform and Config Connector.
- Develop CI/CD pipelines and GitOps workflows using Argo CD and Azure Pipelines.
- Enhance observability through monitoring, logging, and tracing (Prometheus, Grafana, Dynatrace).
- Automate operational workflows using AI, Python, shell scripting, and Ansible.
- Implement security best practices including secrets management (Vault) and policy enforcement.
- Incident response, reliability improvements, and postmortem analysis.
Required Qualifications
- 5+ years experience in SRE, DevOps, or cloud engineering roles.
- Strong expertise in Kubernetes and container platforms.
- Experience with Terraform and infrastructure automation.
- Proficiency in Python, Bash, or similar scripting languages.
- Experience with CI/CD and GitOps methodologies.
- Strong understanding of observability and monitoring tools.
- Ability to lead cross-functional initiatives and mentor engineers.
Preferred Qualifications
- Akamai (edge and CDN services)
- Azure Pipelines (CI/CD) and pipeline maintenance
- Google Config Connector
- Terraform (Infrastructure as Code)
- Python scripting for automation
- Shell scripting
- Ansible and AWX Tower
- Linux system administration
- Docker containerization
- Google Cloud Platform (GCP)
- Kubernetes administration (GKE)
- OpenShift (OCP4) administration
- Grafana and Loki (dashboard and log management)
- Istio service mesh
- Gatekeeper policy management
- Kiali (service mesh observability)
- Argo CD (GitOps deployments)
- Argo Workflows
- Apache web server configuration (URL rewrites, domain management)
- Apigee API management
- Google Cloud Firestore
- Google Bigtable
- Google AlloyDB
- Google BigQuery
- Google Pub/Sub messaging
- Apache Kafka (Strimzi)
- Confluent Kafka
- Apache Solr
- Zookeeper
- HashiCorp Vault (secrets management)
- SOPS (Secrets Operations)
- external-secrets-operator
- cert-manager
- Prometheus operator
- OpenTelemetry operator
- Strimzi operator
- Solr operator
- Crane (container tooling)
- WebMethods
- Camunda (workflow automation)
- Dynatrace (dashboard setup and monitoring)
- Dynatrace operator
- PagerDuty (alerting and incident management)
Vagas similares
Mantenha uma lista reserva.
Stack
Use estas tags para comparar vagas remotas similares.
Elegibilidade de localização
Candidatos devem aplicar apenas quando o país do perfil estiver listado aqui.
Fluxo de contratação
O WithMira mostra a vaga e depois envia candidatos para a aplicação da empresa.
1Confira fit da vaga, stack e elegibilidade de localização no WithMira.
2Abra a página de aplicação da empresa pelo link rastreado.
3Salve a vaga ou assine oportunidades similares antes de sair.