All open positions

Open role · GPU Systems

Member of Technical Staff, GPU Systems & Fabric

Improve performance and reliability across multi-GPU systems and their interconnects.

Make the communication paths between accelerators measurable, dependable, and useful to the rest of General Diffusion’s heterogeneous-compute stack. You will characterize how topology, collectives, data movement, and contention shape multi-GPU execution on GD-X and translate that evidence into actionable interfaces for performance models and runtime decisions. This role owns fabric behavior between devices—not per-device kernel implementation or production placement policy.

What you’ll work on

  • Design repeatable experiments that measure collective latency, bandwidth, overlap, and tail behavior across supported multi-GPU and multi-node topologies.
  • Profile end-to-end execution to distinguish communication and topology limits from per-device kernel, memory-capacity, or host-side bottlenecks.
  • Develop and validate topology-aware collective configurations and execution strategies under representative workload shapes, message sizes, and contention.
  • Build telemetry and diagnostic artifacts that connect observed fabric paths, transport choices, failures, and performance outcomes to reproducible runs.
  • Investigate communication stalls, degradation, and configuration mismatches methodically; document bounded mitigations, failure signatures, and reversal criteria.
  • Provide measured fabric capability and cost signals to Compute World Models, Runtime & Placement, and Measurement & Data Infrastructure, with stated assumptions and confidence limits.
  • Maintain regression benchmarks that catch changes in collective performance or reliability across software versions, hardware configurations, and load conditions.

What you bring

  • Hands-on experience improving or operating multi-GPU or multi-node workloads that rely on collectives such as all-reduce, all-gather, reduce-scatter, or all-to-all.
  • Deep practical knowledge of GPU communication stacks—such as NCCL or an equivalent—and the behavior of PCIe, NVLink-class links, RDMA-capable NICs, and network fabrics.
  • Ability to construct controlled performance experiments and interpret traces, counters, and timelines rather than attributing slowdowns to a single layer by default.
  • Experience reasoning about topology-aware communication, rank-to-device mapping, rail affinity, synchronization, and the effects of concurrent workloads.
  • Sound systems judgment in Linux-based distributed environments, including debugging hangs, timeouts, throughput regressions, and configuration-dependent failures.

What progress looks like

  • Produces a versioned baseline of collective and data-movement behavior across the available testbed topologies, including methodology, variability, and known measurement limits.
  • Delivers a reproducible diagnosis workflow that separates fabric-bound regressions from kernel, memory, host, and runtime causes, demonstrated on representative failure or slowdown cases.
  • Establishes evidence-backed fabric features and cost signals that downstream performance-model and runtime teams can consume, with validation against held-out workload and topology conditions.

Where this role fits

Owns communication and fabric behavior between devices; Kernels owns per-device primitives and Runtime owns placement.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email