Senior SRE

clera

📍 Remote🌐 Remote💼 fulltime🕐 1mo ago🔗 arbeitnow

Job Description

### About the Role We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and infrastructure analytics. Our engineering team is distributed across Europe and North America, and we're looking for a **Senior Site Reliability Engineer** to take ownership of our cloud infrastructure and elevate our DevOps and reliability practices. In this role, you'll evolve our **Google Cloud Platform (GCP)** infrastructure, mature our observability platform, drive incident management processes, and partner closely with Product & Engineering teams to ship reliable, high-quality software. You'll be a key voice in championing SLOs, error budgets, and DORA metrics across the organisation. ### What You'll Do * Design, evolve, and scale our cloud infrastructure on GCP. * Build tooling and automation that promote team autonomy and reduce toil. * Advance our observability platform, improving mean time to recovery (MTTR) and system visibility. * Build transparency into infrastructure costs and drive cost optimisation initiatives. * Champion reliability best practices including SLOs/SLIs, error budgets, and post-incident reviews. * Lead on-call rotations and incident management, fostering a blameless culture. * Help engineering teams leverage GCP effectively and govern usage at scale. ### What We're Looking For **Required:** * 3+ years of Site Reliability Engineering or production SRE experience. * Strong proficiency with **Google Cloud Platform (GCP)**, including cost optimisation and governance. **Nice to Have / Additional Skills:** * Hands-on experience with **Kubernetes** for cluster and workload management. * Infrastructure as Code experience — **Terraform**, Deployment Manager, or similar. * Scripting and automation skills in **Python**, **Bash**, or **Go**. * Strong observability stack experience: **Prometheus**, **Grafana**, **OpenTelemetry**, logging, and tracing. * Proven ability to define and implement SLOs/SLIs and error budgets. * Experience with incident management, post-incident reviews, and on-call rotations. ### Location This is a **fully remote** role, open to candidates based in the **EU, UK, or North America**. The primary hub is in the **Netherlands**. Please note that **visa sponsorship is not available** for this position. ### Compensation & Benefits Compensation details were not provided for this role. Our team spans multiple countries and we offer competitive, location-adjusted packages. Further details will be discussed during the interview process. Find [Jobs in Germany](https://www.arbeitnow.com) on Arbeitnow