Stord
Senior Site Reliability Engineer
Remote Site Reliability Engineering role with clear candidate location fit.
PostedJul 19, 2026
Eligible countries41 accepted countries
Seniority signalSenior
Work settingRemote
Accepted candidate locations
Role overview
Senior Site Reliability Engineer
Requirements and responsibilities
Readable role content extracted into sections for faster review.
Why This Role
- You'll own high-impact infrastructure work directly, in a small team where your contributions are visible and your judgment is trusted. The path from decision to production is short.
- The surface area is broad and real: GKE, Cloud Run, GCP core services, AlloyDB, and the CI/CD and observability tooling that ties it together.
- You'll shape reliability and automation practices during a period of real growth, working alongside engineers who care about doing it well.
Infrastructure & Platform
- Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
- Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
- Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
- Drive down cost and toil through better defaults, right-sizing, and automation rather than manual intervention.
Reliability & Observability
- Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do.
- Define the reliability signals that matter for the services you own, and hold the line on them.
- Develop and maintain disaster recovery and business-continuity strategies, and prove they work.
Automation & Delivery
- Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
- Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows.
- Build custom tooling and scripts that remove recurring operational pain.
Collaboration & Incident Response
- Partner with data and development teams to improve deployment practices and application reliability.
- Provide escalation support for production incidents, help lead post-incident reviews, and turn findings into durable fixes.
- Participate in technical design reviews and offer architectural input across teams.
- Help improve SRE and infrastructure best practices across the team, and participate in on-call for critical systems.
Required
- 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision.
- GCP depth: strong hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM). You know how these fit together in production, not just in a certification.
- Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads.
- Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code.
- A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts.
- Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one.
- Distributed systems fundamentals: failure modes, consistency, and how systems break at scale.
- Git and collaborative development workflows: you work in shared codebases and review others' changes well.
- Incident management: you've run incidents and post-mortems and can stay calm and methodical when production is on fire.
Required Soft Skills
- Ownership & Accountability: You own features end-to-end and take pride in what you ship. You follow through from design to production and don't drop things.
- Strong Communication: You can explain technical decisions and trade-offs to engineers, PMs, and stakeholders. You ask good questions and listen well.
- Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates.
- Production Mindset: You prioritize reliability and user impact. You think about failure modes, monitoring, and operational concerns as part of your design process.
- Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools. You stay current without chasing hype.
- Directed AI-Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality and your own technical judgment.
Strongly Preferred
- Database operations depth: PostgreSQL internals (logical replication, vacuuming, lock contention) or experience with migrations and database scaling. Familiarity with Redis, ClickHouse, or analytical stores is a plus.
- Event-driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and the operational realities of streaming at scale.
- Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability.
Nice to haves
- GCP certifications (Cloud Architect, Cloud DevOps Engineer) or demonstrably equivalent depth.
- Cloudflare: experience with Workers and other Cloudflare services.
- Multi-cloud or hybrid architecture exposure.
What Success Looks Like
- 30 days: You've ramped on our GCP environment and Terraform setup, shipped your first infrastructure or automation change to production, and are contributing in incident and design discussions.
- 90 days: You independently own a meaningful slice of the infrastructure, have improved a reliability, cost, or automation pain point that was slowing the team down, and teammates lean on you in your area.
- 6 months: You're a trusted owner of your part of the infrastructure, consistently delivering improvements that make the platform more reliable and efficient.
Similar roles
Keep a backup shortlist.
Content Classification, English 9 accepted countries
Taxonomy Analyst (German Speaker)IndeedView role Content Classification, English 9 accepted countries
Taxonomy Analyst (Spanish Speaker)IndeedView role Content Classification, English 9 accepted countries
Taxonomy Analyst (Dutch Speaker)IndeedView role Content Classification, English 9 accepted countries
Taxonomy Analyst (French Speaker)IndeedView role Stack
Use these tags to compare similar remote roles.
Location eligibility
Candidates should apply only when their profile country is listed here.
Your profileCountry not setSign in to check your country against this role.
View all 41 accepted countries
AlbaniaArmeniaAustriaBelarusBelgiumBulgariaCroatiaCyprusCzechiaDenmarkEstoniaFinlandFranceGermanyGreeceHungaryIcelandIrelandItalyLatviaLithuaniaLuxembourgMaltaMoldovaMontenegroNetherlandsNorth MacedoniaNorwayPolandPortugalRomaniaSerbiaSlovakiaSloveniaSpainSwedenSwitzerlandTurkeyUkraineUnited KingdomUSA
Hiring flow
WithMira shows the role, then sends candidates to the company application.
1Check role fit, stack, and location eligibility in WithMira.
2Open the company application page from the tracked apply link.
3Save the role or subscribe for similar opportunities before leaving.