About this role
The team is seeking a highly skilled and intelligent Senior Site Reliability Engineer (SRE) to join their Platform Site Reliability Engineering (PSRE) team. This team is vital in ensuring the stability, reliability, and availability of mission-critical production applications.
Key Responsibilities:
- Develop and implement observability, monitoring, logging, and tracing systems to proactively detect and prevent issues.
- Build and maintain tools that enhance the reliability and performance of the platform.
- Collaborate with software engineering teams to improve system performance and reliability.
- Manage incident response and post-mortem processes to ensure continuous improvement.
- Automate manual processes to improve operational efficiency and reduce downtime.
Required Skills & Qualifications:
- Strong experience in Site Reliability Engineering or a related field.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with container orchestration tools like Kubernetes or Docker.
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, or ELK stack.
- Solid programming skills in languages such as Python, Go, or Java.
- Excellent problem-solving abilities and a proactive mindset.
Experience:
- A minimum of 5-8 years of relevant experience in Site Reliability Engineering or a similar role.
What we offer:
- Opportunity to work on cutting-edge technologies in a dynamic environment.
- Collaborative team culture focused on innovation and continuous learning.
- Professional development opportunities to advance your career.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.