About this role
The team is seeking a Senior Cloud Site Reliability Engineer (SRE) for a remote position. The ideal candidate will have extensive experience in cloud infrastructure and a strong background in Site Reliability Engineering, focusing on building, automating, and operating highly available and scalable AWS environments.
Key Responsibilities:
- Design and manage AWS cloud-native and hybrid infrastructure.
- Develop event-driven integrations using API Gateway, Lambda, SNS, SQS, EventBridge, Step Functions, and Glue.
- Build CI/CD pipelines using Jenkins, GitHub Actions, AWS CodePipeline, and other relevant tools.
- Implement Infrastructure as Code (IaC) using tools like Terraform or CloudFormation.
- Monitor system performance and reliability, troubleshooting issues as they arise.
- Collaborate with development teams to ensure seamless deployment and integration of applications.
- Maintain documentation of systems, processes, and procedures related to cloud operations.
Required Skills & Qualifications:
- Proven experience with AWS services and cloud architecture.
- Strong understanding of SRE principles and practices.
- Proficiency in scripting languages such as Python, Bash, or similar.
- Experience with containerization technologies such as Docker and orchestration tools like Kubernetes.
- Familiarity with monitoring tools like CloudWatch, Prometheus, or Grafana.
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience:
- Minimum of 5 to 8 years of relevant experience in cloud engineering and site reliability.
What we offer:
- Opportunity to work in a dynamic and innovative environment.
- Flexible remote work arrangements.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.