Senior Hpc Platform Architect

in01 nvidia graphics bengaluru

📍 india bengaluru india🕐 28d ago🔗 workday

Job Description

NVIDIA is looking for an exceptional engineer to grow and thrive alongside our HPC Infrastructure team that designs, evaluates, and optimizes the compute inrastructure powering NVIDIA's next-generation silicon design and AI workloads. You will be the primary architecture reviewer and performance champion for new data center clusters being built across multiple sites globally. **What you'll be doing:** * Own data center architecture reviews for new HPC clusters — evaluating compute, storage,networking, and cooling decisions and challenging assumptions before clusters are built. * Serve as the BDC representative in cluster build meetings across multiple simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments), driving architecture alignment independently. * Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology — and surface risk and tradeoff recommendations to leadership. * Lead performance benchmarking and profiling of HPC cluster infrastructure, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early. * Drive infrastructure optimization at multiple layers: scheduler-level tuning (LSF/Slurm),hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge page configuration, kernel image selection). * Collaborate with platform and operations teams on cluster health, capacity planning, and operational readiness for tapeout milestones. * Partner with vendors and internal teams to evaluate new hardware, storage systems, and networking fabrics; produce architecture recommendations backed by data. * Continuously improve infrastructure observability, benchmarking frameworks, and architecture documentation to raise the bar across the BDC HPC platform. **What we need to see:** * B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale. * Deep understanding of data center architecture fundamentals: compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand,Ethernet, NVLink). * Proven ability to evaluate and challenge infrastructure design decisions — including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology. * Experience with OS and kernel-level performance tuning: NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization. * Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and/or Slurm, including job scheduling optimization and resource utilization analysis. * Strong Linux/Unix system administration skills and proficiency in scripting (Python, Bash, or Perl) for automation and analysis. * Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks. * Excellent problem-solving and communication skills — able to synthesize complex architectural tradeoffs and present clear recommendations to engineering leadership.
Senior Hpc Platform Architect at in01 nvidia graphics bengaluru | MergeJobs