Senior Modeling Architect, Performance Benchmarking
Core
Own performance and energy metrics for the T100 optical inference accelerator, producing benchmark numbers across modeling fidelities and measured hardware to guide architecture and product decisions.
Role type
Senior IC performance engineer (optical inference accelerator)
Builds
Performance models, benchmark harnesses, and measurement reports for silicon photonics inference chips
Domain
AI hardware / Silicon photonics / High-performance computing
Required skills
GPU performance engineering, accelerator benchmarking, roofline analysis, limiter analysis, analytical performance modeling, NVIDIA Nsight Systems/Compute, LLM inference stacks (vLLM, SGLang, TensorRT-LLM), Python, Linux, cloud GPU operations
Preferred skills
RTL/Verilator simulation correlation, CUDA/CUTLASS/Triton kernel work, PyTorch internals, distributed inference (NCCL, NVLink, InfiniBand), MLPerf pipelines
Technologies
NVIDIA Nsight Systems, Nsight Compute, Hugging Face, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Python, Linux, AWS, GCP, Azure
Responsibilities
Produce benchmark numbers for workloads across roofline models, architecture models, RTL simulation, and measured competitor hardware; Maintain consistent workload definitions across fidelities; Measure competing GPUs/accelerators end-to-end including cloud/lab setup; Report TTFT, ITL, tokens/sec, and energy metrics; Document discrepancies between simulation and measurement; Maintain a reviewed internal benchmark suite
Seniority
Senior, hands-on IC