Head of Technology and Service Operations
2038 cubic transportation india private
📍 ind hyderabad aparna india🕐 20d ago🔗 workday
Job Description
**Business Unit:**
==================
Cubic Transportation Systems
**Company Details:**
====================
When you join Cubic, you become part of a company that creates and delivers technology solutions in transportation to make people’s lives easier by simplifying their daily journeys, and defense capabilities to help promote mission success and safety for those who serve their nation. Led by our talented teams around the world, Cubic is committed to solving global issues through innovation and service to our customers and partners.
We have a top-tier portfolio of businesses, including Cubic Transportation Systems (CTS) and Cubic Defense (CD). Explore more on Cubic.com.
**Job Details:**
================
**Position Summary**
The Head of Technology and Service Operations is a senior executive responsible for the global operations, service delivery, and reliability of mission-critical transit payment systems used by millions of passengers daily. The role leads more than 500 staff across two global 24x7 operations centers (UK and India), localized in-country teams, and multiple customer contact centers.
**Key Responsibilities**
**Service Reliability & Availability:**
* _Deliver Five Nines Availability (99.999%):_ Build and enforce architectural and operational practices that ensure global transit payment systems achieve and sustain ultra-high uptime. This includes proactive monitoring, high-availability design enforcement, and automated failover systems across cloud and on-premises platforms.
* _Govern Service Level Objectives (SLOs) & Error Budgets_: Define, track, and report SLOs for all critical services, ensuring error budgets are managed responsibly to balance reliability with change velocity.
* _Transaction Performance Management_: Guarantee low-latency, high-throughput processing across all environments, actively tuning Oracle databases, Kubernetes clusters, UCS fabrics and cloud services for peak demand conditions (e.g., major city events, commuter surges).
* _Preventative Maintenance Program_: Own a proactive, structured maintenance strategy (firmware updates, DB patching, load balancing, failover readiness) to ensure reliability and reduce risk of unplanned outages.
**Incident, Problem & Change Management**
* _Global Incident Oversight_: Lead the 24x7 incident response process across UK and India operations centers, ensuring ≥95% SLA compliance for response and resolution.
* _Root Cause & Blameless Postmortems_: Mandate blameless postmortems for every significant incident, driving systemic fixes to prevent recurrence and sharing lessons learned across all regions.
* _Problem Management_: Establish a structured problem management process to identify trends, reduce repeat incidents, and address root technical or process flaws.
* _Change Governance:_ Oversee change management to balance service stability with innovation. Implement automated pipelines where possible to reduce manual error and ensure changes are tested for resilience before release.
* _Recovery Metrics_: Continuously track and drive down Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) with quarterly improvement goals.
**Observability, Automation & SRE Practices**
* _Full-Stack Observability_: Deploy and govern monitoring solutions across metrics, logs, and traces, enabling real-time visibility into system health and early anomaly detection.
* _Automation of Toil:_ Drive a culture of automation by eliminating repetitive manual tasks in patching, scaling, failover, and incident response. Track progress with explicit automation coverage targets.
* _Chaos Engineering & Resilience Testing:_ Institutionalize chaos testing and DR drills (“game days”) to validate RTO/RPO readiness and system recovery under stress conditions.
* _Capacity & Performance Engineering_: Oversee predictive capacity planning to ensure the system scales automatically to meet load and maintain performance even under extreme demand conditions.
* _Proactive Reliability Improvements_: Fund and prioritize engineering initiatives aimed at improving long-term service resilience, not just reactive firefighting.
**Global Operations & Customer Support**
* _Global Ops Center Leadership_: Direct two 24x7 global operations centers (UK and India) as the backbone of global service delivery. Ensure staffing, shift rotations, and runbooks meet the highest standards of responsiveness and reliability.
* _Localized Team Oversight_: Manage in-country support teams that ensure local compliance, and customer-specific responsiveness.
* _Customer Contact Centers_: Own the operations of customer-facing contact centers, ensuring tight integration with back-end service teams to provide consistent, rapid, and high-quality customer experience.
* _Standardization Across Regions_: Drive consistency of ITIL-aligned service management processes globally, ensuring customers experience the same high standards regardless of geography.
**Compliance, Risk & Audit Readiness**
* _Global Compliance Ownership_: Ensure operations remain compliant with ISO27001, PCI DSS v4.1, Essential Eight, Cyber Essentials, GDPR, and all other local privacy regulations in customer jurisdictions.
* _Audit Readiness & Evidence_: Maintain continuous audit readiness with robust evidence trails across all systems, processes, and controls, avoiding last-minute remediation before audits.
* _Corporate CISO & Service Assurance Partnership_: Work in lockstep with the Corporate CISO and Service Assurance functions to embed security controls, risk frameworks, and governance processes into day-to-day operations, aiming for regulatory compliance excellence.
* _Risk Reporting_: Provide transparent reporting to the COO and Senior Leadership Teams on compliance posture, vulnerabilities, audit results, and remediation progress.
* _Vulnerability Management_: Ensure vulnerabilities are tracked, prioritized, and remediated on schedule, with operational accountability for timely fixes.
**Financial Management & Vendor Oversight**
* _Budget Accountability:_ Own the operations budget, ensuring disciplined financial management, proactive forecasting, and cost optimization across global operations.
* _Abatements & Liquidated Damages_: Minimize exposure to penalties by delivering against contractual SLAs, proactively addressing risks that might result in service-level breaches.
* _Vendor & Partner Management_: Oversee the performance of critical vendors and partners — including telcos, cloud providers, and managed service providers. Hold them accountable to contractual SLAs and enforce remedies when necessary.
* _Commercial Strategy:_ Negotiate favorable terms for renewals and expansions, aligning vendor performance with global service reliability goals.
* _Cost Optimization_: Continuously evaluate opportunities to optimize spending without compromising reliability or compliance.
**Leadership & Governance**
* _Executive Leadership:_ Act as a key member of the Senior Leadership Team, reporting to the COO, with direct accountability for global technology and service operations.
* _Cross-Functional Collaboration:_ Partner with the CTO and Software Delivery leaders to align infrastructure improvements, operational enhancements, and vulnerability remediation with technology roadmaps.
* _Cadence of Accountability_: Establish and enforce disciplined operational cadences (daily incident reviews, weekly ops reviews, monthly resilience drills, quarterly DR game days, annual audit cycles).
* _Culture of Excellence_: Foster a global culture of reliability-first thinking, preventative maintenance, process rigor, and continuous improvement.
* _Talent Leadership_: Build and develop global operational leadership talent, ensuring succession planning, skills development, and team engagement across all regions.
**KPIs**
* _Service Reliability:_ Sustain 99.999% uptime across mission-critical systems, with consistently low latency and high throughput.
* _Incident Outcomes:_ ≥95% SLA compliance for incident response and resolution; demonstrate quarterly improvement in MTTR/MTTD.
* _Preventative Maintenance & Resilience_: 100% completion of scheduled maintenance; successful quarterly DR drills with RTO/RPO close to zero.
* _Automation & Efficiency:_ Year-on-year reduction of manual toil, with measurable increases in automation coverage.
* _Compliance Excellence:_ Zero major findings in ISO27001, PCI DSS, GDPR, Essential Eight, and local regulatory audits; evidence and control logs always audit-ready.
* _Financial Stewardship:_ Deliver operations budget to plan; minimize abatements/liquidated damages; ≥98% vendor SLA compliance.
* _Stakeholder Satisfaction:_ Positive CSAT/NPS with transit authority clients and strong feedback from COO, CISO, CTO, and board stakeholders.
**Candidate Profile**
* _Leadership_: 15+ years in global technology operations, with 10+ years in senior leadership of large-scale, 24x7 organizations.
* _Technical Breadth:_ Strong working knowledge of Azure, AWS, Kubernetes, Oracle databases, UCS fabrics, and enterprise networks; proven track record delivering highly available, resilient platforms.
* _Service & Reliability Engineering:_ Deep experience with ITIL service management integrated with SRE practices (SLOs, error budgets, observability, automation, chaos testing).
* _Compliance & Risk:_ Demonstrated success managing ISO27001, PCI DSS, GDPR, and local privacy obligations; experienced in working directly with CISOs and Service Assurance functions.
* _Financial & Vendor Management:_ Proven ability to manage large budgets, minimize penalties, and hold global/cloud/telco vendors accountable to SLAs.
* _Leadership Attributes:_ Process-driven operator with a reliability-first mindset; skilled communicator with executive presence; committed to building high-performing, globally distributed teams.
**Worker Type:**
================
Employee
_We are committed to creating an inclusive workplace and welcome applications from people of all backgrounds. We do not discriminate based on any protected characteristic under applicable law._