Senior SRE
clera
📍 Remote🌐 Remote💼 fulltime🕐 1mo ago🔗 arbeitnow
Job Description
### About the Role
We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and infrastructure analytics. Our engineering team is distributed across Europe and North America, and we're looking for a **Senior Site Reliability Engineer** to take ownership of our cloud infrastructure and elevate our DevOps and reliability practices.
In this role, you'll evolve our **Google Cloud Platform (GCP)** infrastructure, mature our observability platform, drive incident management processes, and partner closely with Product & Engineering teams to ship reliable, high-quality software. You'll be a key voice in championing SLOs, error budgets, and DORA metrics across the organisation.
### What You'll Do
* Design, evolve, and scale our cloud infrastructure on GCP.
* Build tooling and automation that promote team autonomy and reduce toil.
* Advance our observability platform, improving mean time to recovery (MTTR) and system visibility.
* Build transparency into infrastructure costs and drive cost optimisation initiatives.
* Champion reliability best practices including SLOs/SLIs, error budgets, and post-incident reviews.
* Lead on-call rotations and incident management, fostering a blameless culture.
* Help engineering teams leverage GCP effectively and govern usage at scale.
### What We're Looking For
**Required:**
* 3+ years of Site Reliability Engineering or production SRE experience.
* Strong proficiency with **Google Cloud Platform (GCP)**, including cost optimisation and governance.
**Nice to Have / Additional Skills:**
* Hands-on experience with **Kubernetes** for cluster and workload management.
* Infrastructure as Code experience — **Terraform**, Deployment Manager, or similar.
* Scripting and automation skills in **Python**, **Bash**, or **Go**.
* Strong observability stack experience: **Prometheus**, **Grafana**, **OpenTelemetry**, logging, and tracing.
* Proven ability to define and implement SLOs/SLIs and error budgets.
* Experience with incident management, post-incident reviews, and on-call rotations.
### Location
This is a **fully remote** role, open to candidates based in the **EU, UK, or North America**. The primary hub is in the **Netherlands**. Please note that **visa sponsorship is not available** for this position.
### Compensation & Benefits
Compensation details were not provided for this role. Our team spans multiple countries and we offer competitive, location-adjusted packages. Further details will be discussed during the interview process.
Find [Jobs in Germany](https://www.arbeitnow.com) on Arbeitnow