Senior SWE

in01 nvidia graphics bengaluru

📍 india bengaluru india🕐 15d ago🔗 workday

Job Description

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. **What You Will Be Doing:** * Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. * Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. * Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. * Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. * Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. * Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. * Provide technical leadership and mentorship, influence engineering standards and architecture decisions, and deliver measurable improvements in reliability, MTTR, operational toil, engineering productivity, and infrastructure efficiency. **What We Need To See:** * Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience, with 10+ years of software engineering, SRE, infrastructure, or distributed-systems experience and demonstrated technical leadership. * Strong software engineering expertise in Go, Python, or equivalent languages, with experience designing and building production-grade distributed systems, platform services, APIs, and automation. * Proven experience owning complex software/platform initiatives across multiple teams or infrastructure domains, from architecture and implementation through adoption and measurable impact. * Deep understanding of distributed systems, event-driven architectures, microservices, APIs, and high-throughput data processing, including technologies such as Kafka, NATS, gRPC, or equivalent. * Strong experience with modern observability and telemetry platforms, including OpenTelemetry, Prometheus, VictoriaMetrics, Vector, Loki, Grafana, ClickHouse, or equivalent technologies. * Strong SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, networking, and/or cloud environments, with experience using Terraform, Ansible, or equivalent automation technologies. * Demonstrated ability to solve ambiguous problems, influence technical direction without direct authority, mentor engineers, establish engineering standards, and deliver measurable operational and business outcomes. **Ways To Stand Out From The Crowd:** * Experience building software and reliability platforms for large-scale on-premises infrastructure, particularly Storage, Compute, Networking, VMware, and Kubernetes/OpenShift. * Deep understanding of storage and infrastructure telemetry, including IOPS, latency, NVMe health, SAN/NAS topology, block/object storage, and infrastructure failure domains. * Experience building self-healing systems, automated remediation, predictive operations, or autonomous SRE capabilities. * Production experience applying Generative AI, AIOps, LLMs, or Agentic AI to incident triage, RCA, operational intelligence, debugging, or remediation; experience with LangChain, LlamaIndex, AutoGen, or equivalent is a plus. * Experience building high-performance platform services using FastAPI, gRPC, or equivalent technologies.
Senior SWE at in01 nvidia graphics bengaluru | MergeJobs