Role overview

Site Reliability Engineer

Requirements and responsibilities

Readable role content extracted into sections for faster review.

1. Infrastructure Design and Automation:

  • Collaborate with software engineering and operations teams to design, build, and maintain cloud-based infrastructure using AWS and Terraform.
  • Implement and enhance infrastructure-as-code (IaC) practices using Terraform to ensure reproducibility and scalability of infrastructure components.

2. Monitoring and Incident Management:

  • Develop and maintain monitoring solutions to proactively identify performance bottlenecks, system outages, and other potential issues.
  • Participate in incident response and root cause analysis efforts to drive continuous improvement and prevent future incidents.

2. Monitoring and Incident Management:

  • Optimise system performance, reliability, and cost efficiency through continuous monitoring, performance tuning, and capacity planning.
  • Identify opportunities to automate manual processes and improve system resilience.

4. Scripting and Automation:

  • Utilise Python or Bash scripting to create and maintain automation tools for various operational tasks and deployments.
  • Implement and improve continuous integration and continuous deployment (CI/CD) pipelines.

5. Security and Compliance:

  • Collaborate with security teams to implement best practices for securing cloud infrastructure and services.
  • Ensure compliance with relevant industry standards and regulations.

6. Deployment and Release Management:

  • Support CI/CD pipelines for application deployments and updates.
  • Contribute to the design and implementation of deployment strategies that promote zero-downtime releases.

7. Documentation and Knowledge Sharing:

  • Maintain clear and up-to-date documentation for infrastructure configurations, processes, and incident resolution procedures.
  • Participate in knowledge sharing with team members to enhance overall expertise and skill sets.

1. Education and Experience:

  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • Proven experience as a Site Reliability Engineer or similar role.

2. Technical Skills:

  • Extensive experience with Amazon Web Services (AWS) and its core services (EC2, S3, RDS, IAM, etc.).
  • Strong proficiency in infrastructure-as-code (IaC) tools, with a focus on Terraform.
  • Proficient in scripting with Python or Bash for automation and operational tasks.
  • Solid understanding of networking principles and protocols.
  • Knowledge of CI/CD pipelines and related tools.

2. Technical Skills:

  • Ability to diagnose and resolve complex technical issues in a fast-paced environment.
  • Analytical mindset to proactively identify potential system weaknesses and performance bottlenecks.

4. Collaboration and Communication:

  • Strong teamwork and collaboration skills to work effectively with cross-functional teams.
  • Excellent verbal and written communication skills.
Similar roles

Keep a backup shortlist.

Browse stack
FocusSite Reliability EngineeringRole area
Seniority signalSeniorCandidate level
StackAWS, CI/CD, PythonPrimary skills
Location1 accepted countryEligibility

Stack

Use these tags to compare similar remote roles.

Location eligibility

Candidates should apply only when their profile country is listed here.

Your profileCountry not setSign in to check your country against this role.

Hiring flow

WithMira shows the role, then sends candidates to the company application.

1Check role fit, stack, and location eligibility in WithMira.
2Open the company application page from the tracked apply link.
3Save the role or subscribe for similar opportunities before leaving.