Lead SRE Incident Management

qlys_in qualys security techservices private

📍 pune india🕐 1mo ago🔗 workday

Job Description

Come work at a place where innovation and teamwork come together to support the most exciting missions in the world! **THE OPPORTUNITY** The Lead Site Reliability Engineer – Incident Management provides technical leadership for production reliability, major incident management, automation, and operational excellence across Qualys' global SaaS platform. This role owns complex cross-functional initiatives, drives engineering best practices, mentors engineers, and partners with architecture and product teams to improve reliability, scalability, and customer experience. **WHERE THIS ROLE SITS** Partners with Engineering, Infrastructure, DevOps, Cloud Operations, Database Engineering, Security, Product Engineering, Customer Support, and Executive Leadership. **WHAT YOU WILL DO** Reliability & Incident Management · Lead enterprise-wide critical incident response and serve as Incident Commander. · Own end-to-end restoration strategy for complex production outages. · Drive executive communications and customer-impact assessments. · Lead root cause analysis and systemic reliability improvements. **Reliability Engineering & Automation** · Design self-healing platforms and automation. · Improve observability, SLOs, SLIs, and error budgets. · Reduce MTTD and MTTR through engineering improvements. **Operational Excellence** · Lead capacity planning and operational readiness. · Review architecture for reliability and scalability. · Drive cross-functional operational standards. **Leadership & Collaboration** · Mentor Senior and Lead SREs. · Drive SRE strategy and reliability roadmap. · Influence engineering priorities and reliability culture. **WHAT GOOD LOOKS LIKE** · Drives measurable improvements in availability and resiliency. · Builds a culture of automation and operational excellence. · Leads major incidents with confidence. · Influences engineering decisions across organizations. **DISTINGUISHING EXPECTATION** · Enterprise-wide reliability improvements · Reduction in critical incidents · Higher automation adoption · Leadership across SRE organization **REQUIRED QUALIFICATIONS** · Bachelor's degree or equivalent. · 8–12+ years in SRE/Production Operations. · Strong Linux, Kubernetes, Cloud, Networking and Databases. · Expertise in Python, Go, Bash or Java. · Experience leading major incidents and mentoring engineers. · Excellent communication and leadership skills. **PREFERRED QUALIFICATIONS** · Large-scale SaaS experience · Multi-cloud expertise · Chaos Engineering · Cloud/Kubernetes certifications · ITIL certification **WORK ENVIRONMENT** · Full-time · 24x7 production support · On-call leadership · Hybrid/Remote · Cross-functional collaboration