About this role
The team is seeking a highly autonomous and technically proficient Senior Site Reliability Engineer (SRE) to own their observability strategy and critical platform services. In this role, you will not just monitor systems; you will engineer their reliability, performance, and security. You will serve as the primary authority for the monitoring stack and the operational lead for core middleware layers.
Key Responsibilities:
- Design and implement observability solutions to enhance system reliability and performance.
- Lead the operational management of critical platform services and middleware layers.
- Collaborate with development teams to integrate reliability and performance best practices.
- Proactively identify and resolve system issues before they impact users.
- Develop and maintain documentation for system architecture and processes.
- Mentor junior engineers and contribute to a culture of continuous improvement.
Required Skills & Qualifications:
- Strong experience with observability tools and practices, including monitoring, logging, and tracing.
- Proficient in cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with containerization and orchestration technologies like Docker and Kubernetes.
- Solid understanding of networking, security, and system architecture.
- Excellent problem-solving skills and ability to work in a fast-paced environment.
- Strong scripting skills in languages such as Python, Bash, or similar.
Experience:
- 5-8 years of experience in Site Reliability Engineering or related fields, with a focus on observability and platform systems.
What we offer:
- A dynamic work environment with opportunities for professional growth and development.
- The chance to work on cutting-edge technology and contribute to innovative projects.
- A collaborative team culture that values autonomy and initiative.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.