Open role · GPU Systems
Member of Technical Staff, GPU Systems & Fabric
Improve performance and reliability across multi-GPU systems and their interconnects.
Make the communication paths between accelerators measurable, dependable, and useful to the rest of General Diffusion’s heterogeneous-compute stack. You will characterize how topology, collectives, data movement, and contention shape multi-GPU execution on GD-X and translate that evidence into actionable interfaces for performance models and runtime decisions. This role owns fabric behavior between devices—not per-device kernel implementation or production placement policy.
What you’ll work on
- Design repeatable experiments that measure collective latency, bandwidth, overlap, and tail behavior across supported multi-GPU and multi-node topologies.
- Profile end-to-end execution to distinguish communication and topology limits from per-device kernel, memory-capacity, or host-side bottlenecks.
- Develop and validate topology-aware collective configurations and execution strategies under representative workload shapes, message sizes, and contention.
- Build telemetry and diagnostic artifacts that connect observed fabric paths, transport choices, failures, and performance outcomes to reproducible runs.
- Investigate communication stalls, degradation, and configuration mismatches methodically; document bounded mitigations, failure signatures, and reversal criteria.
- Provide measured fabric capability and cost signals to Compute World Models, Runtime & Placement, and Measurement & Data Infrastructure, with stated assumptions and confidence limits.
- Maintain regression benchmarks that catch changes in collective performance or reliability across software versions, hardware configurations, and load conditions.
What you bring
- Hands-on experience improving or operating multi-GPU or multi-node workloads that rely on collectives such as all-reduce, all-gather, reduce-scatter, or all-to-all.
- Deep practical knowledge of GPU communication stacks—such as NCCL or an equivalent—and the behavior of PCIe, NVLink-class links, RDMA-capable NICs, and network fabrics.
- Ability to construct controlled performance experiments and interpret traces, counters, and timelines rather than attributing slowdowns to a single layer by default.
- Experience reasoning about topology-aware communication, rank-to-device mapping, rail affinity, synchronization, and the effects of concurrent workloads.
- Sound systems judgment in Linux-based distributed environments, including debugging hangs, timeouts, throughput regressions, and configuration-dependent failures.
What progress looks like
- Produces a versioned baseline of collective and data-movement behavior across the available testbed topologies, including methodology, variability, and known measurement limits.
- Delivers a reproducible diagnosis workflow that separates fabric-bound regressions from kernel, memory, host, and runtime causes, demonstrated on representative failure or slowdown cases.
- Establishes evidence-backed fabric features and cost signals that downstream performance-model and runtime teams can consume, with validation against held-out workload and topology conditions.
Where this role fits
Owns communication and fabric behavior between devices; Kernels owns per-device primitives and Runtime owns placement.
Apply
Interested in this role?
Email your background and a short note about the work you would like to do.
careers@generaldiffusion.com