All open positions

Open role · Hardware Research

Member of Technical Staff, GPU & ASIC Performance Modeling

Understand external GPUs and ASICs through portable measurements and predictive performance models.

Develop portable measurements and predictive performance models that explain how externally available GPUs and ASICs behave under representative workloads, including compute, memory, data movement, and contention. Turn that evidence into calibrated capability profiles that help compute-world-model, compiler, and runtime teams reason about unlike targets—without owning kernel implementation, compiler lowering, placement policy, or silicon design.

What you’ll work on

  • Design controlled microbenchmarks and workload-level experiments that isolate compute throughput, memory-hierarchy behavior, data movement, and contention on external GPU and ASIC targets.
  • Build repeatable characterization protocols that record workload, hardware, software, configuration, and measurement conditions so comparisons across targets are traceable.
  • Develop and validate performance models against measured behavior; quantify prediction error, uncertainty, and the conditions under which a model transfers or ceases to transfer.
  • Use profiler evidence and end-to-end measurements to distinguish compute, memory, communication, and interference limits rather than treating peak specifications as attainable performance.
  • Publish decision-ready capability profiles and methodological caveats for compute-world-model, compiler, and runtime partners, including the evidence behind each conclusion.
  • Partner with GPU Systems & Fabric on multi-device measurements and incorporate relevant data-movement characteristics into profiles, while leaving collective and fabric optimization to that team.

What you bring

  • Experience in accelerator architecture or performance modeling grounded in measured GPU, inference-ASIC, or comparable heterogeneous-system behavior.
  • Ability to design benchmarks that combine focused microkernels with representative workloads, control confounding variables, and make results reproducible.
  • Demonstrated model-to-measurement correlation work, including calibration, error analysis, and clear treatment of uncertainty under workload or architecture shift.
  • Fluency with performance-profiling methods and the interaction of arithmetic intensity, memory hierarchy, occupancy or utilization, and data movement.
  • Vendor-neutral hardware–software judgment: able to compare capabilities through evidence rather than feature lists, and communicate limits to systems and ML collaborators.

What progress looks like

  • A traceable characterization suite produces repeatable capability profiles across selected external targets, with workload and environment provenance sufficient for another engineer to reproduce the comparison.
  • Validation reports compare model estimates with held-out measurements, quantify error and uncertainty by workload regime, and document transfer limits rather than obscuring mismatches.
  • Compiler, runtime, and research partners receive evidence-backed profiles that identify the relevant bottleneck or compatibility caveat for a target and cite the underlying measurements.

Where this role fits

This is characterization of external silicon, not RTL chip design or ownership of a vendor roadmap.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email