Lead SRE Incident Management
qlys_in qualys security techservices private
📍 pune india🕐 1mo ago🔗 workday
Job Description
Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!
**THE OPPORTUNITY**
The Lead Site Reliability Engineer – Incident Management provides technical leadership for production reliability, major incident management, automation, and operational excellence across Qualys' global SaaS platform. This role owns complex cross-functional initiatives, drives engineering best practices, mentors engineers, and partners with architecture and product teams to improve reliability, scalability, and customer experience.
**WHERE THIS ROLE SITS**
Partners with Engineering, Infrastructure, DevOps, Cloud Operations, Database Engineering, Security, Product Engineering, Customer Support, and Executive Leadership.
**WHAT YOU WILL DO**
Reliability & Incident Management
· Lead enterprise-wide critical incident response and serve as Incident Commander.
· Own end-to-end restoration strategy for complex production outages.
· Drive executive communications and customer-impact assessments.
· Lead root cause analysis and systemic reliability improvements.
**Reliability Engineering & Automation**
· Design self-healing platforms and automation.
· Improve observability, SLOs, SLIs, and error budgets.
· Reduce MTTD and MTTR through engineering improvements.
**Operational Excellence**
· Lead capacity planning and operational readiness.
· Review architecture for reliability and scalability.
· Drive cross-functional operational standards.
**Leadership & Collaboration**
· Mentor Senior and Lead SREs.
· Drive SRE strategy and reliability roadmap.
· Influence engineering priorities and reliability culture.
**WHAT GOOD LOOKS LIKE**
· Drives measurable improvements in availability and resiliency.
· Builds a culture of automation and operational excellence.
· Leads major incidents with confidence.
· Influences engineering decisions across organizations.
**DISTINGUISHING EXPECTATION**
· Enterprise-wide reliability improvements
· Reduction in critical incidents
· Higher automation adoption
· Leadership across SRE organization
**REQUIRED QUALIFICATIONS**
· Bachelor's degree or equivalent.
· 8–12+ years in SRE/Production Operations.
· Strong Linux, Kubernetes, Cloud, Networking and Databases.
· Expertise in Python, Go, Bash or Java.
· Experience leading major incidents and mentoring engineers.
· Excellent communication and leadership skills.
**PREFERRED QUALIFICATIONS**
· Large-scale SaaS experience
· Multi-cloud expertise
· Chaos Engineering
· Cloud/Kubernetes certifications
· ITIL certification
**WORK ENVIRONMENT**
· Full-time
· 24x7 production support
· On-call leadership
· Hybrid/Remote
· Cross-functional collaboration