Senior SRE – Kubernetes

software mind americas

📍 Remote🌐 Remote🕐 26d ago🔗 devopsjobs

Job Description

Company Description ------------------- We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!   **About the Client** Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies. You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production. **Contract Duration:** Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance. Job Description --------------- **About the Role** This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack. This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership. **What You’ll Do** * Support the deployment, operation, and reliability of production services running on Kubernetes. * Monitor service health and investigate production incidents across distributed applications. * Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements. * Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams. * Support CI/CD, GitOps-based deployments, observability, and production monitoring. * Work within a client-directed backlog and established priorities. Qualifications -------------- **Required Qualifications** * 5+ years of experience in **Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering**, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services. * 3+ years of hands-on production **Kubernetes** experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting * Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene * **Splunk** experience for log aggregation, search, and production troubleshooting * **Prometheus and Grafana** experience, specifically building alert rules and dashboards, not only using existing dashboards * **CI/CD** and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux * Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking * **Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment.** Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation. * Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication Additional Information ---------------------- **Nice to Have** * Web Components / Lit experience, to perform first-level debugging of UI-related issues * Server-side rendering or isomorphic runtime experience * Canary rollout / multi-version production operations * Distributed tracing and request-context correlation * KEDA or event-driven autoscaling * Experience with enterprise platform integration layers **What We Offer** * Competitive salary and laptop * Professional development and training opportunities * Work with cutting-edge cloud and container technologies * Flexible work arrangements and collaborative team environment * Impact on organization-wide digital transformation initiatives