Senior Hpc Platform Architect
in01 nvidia graphics bengaluru
📍 india bengaluru india🕐 28d ago🔗 workday
Job Description
NVIDIA is looking for an exceptional engineer to grow and thrive alongside our HPC Infrastructure team that designs, evaluates, and optimizes the compute inrastructure powering NVIDIA's next-generation silicon design and AI
workloads. You will be the primary architecture reviewer and performance champion for new data center clusters being built across multiple sites globally.
**What you'll be doing:**
* Own data center architecture reviews for new HPC clusters — evaluating compute, storage,networking, and cooling decisions and challenging assumptions before clusters are built.
* Serve as the BDC representative in cluster build meetings across multiple simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments), driving architecture alignment independently.
* Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology — and surface risk and tradeoff recommendations to leadership.
* Lead performance benchmarking and profiling of HPC cluster infrastructure, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early.
* Drive infrastructure optimization at multiple layers: scheduler-level tuning (LSF/Slurm),hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge page configuration, kernel image selection).
* Collaborate with platform and operations teams on cluster health, capacity planning, and operational readiness for tapeout milestones.
* Partner with vendors and internal teams to evaluate new hardware, storage systems, and networking fabrics; produce architecture recommendations backed by data.
* Continuously improve infrastructure observability, benchmarking frameworks, and architecture documentation to raise the bar across the BDC HPC platform.
**What we need to see:**
* B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
* Deep understanding of data center architecture fundamentals: compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand,Ethernet, NVLink).
* Proven ability to evaluate and challenge infrastructure design decisions — including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology.
* Experience with OS and kernel-level performance tuning: NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization.
* Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and/or Slurm, including job scheduling optimization and resource utilization analysis.
* Strong Linux/Unix system administration skills and proficiency in scripting (Python, Bash, or Perl) for automation and analysis.
* Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
* Excellent problem-solving and communication skills — able to synthesize complex architectural tradeoffs and present clear recommendations to engineering leadership.