All open positions

Open role · Kernels

Member of Technical Staff, Kernels

Make critical operators fast, portable where possible, and correct by construction and test.

Build the critical operator paths that make General Diffusion’s heterogeneous-compute research measurable in real execution. You will turn profiler evidence into carefully tuned kernels, then make the performance claim inseparable from numerical-correctness and regression evidence. The work advances the company’s silicon-neutral direction by making target-specific assumptions explicit and preserving portability wherever measurements support it.

What you’ll work on

  • Profile representative workloads to identify the operators and execution regimes where kernel work can materially affect end-to-end performance; distinguish a kernel bottleneck from a compiler, runtime, or fabric bottleneck before optimizing.
  • Design, implement, and tune priority kernels in an appropriate low-level accelerator stack (for example CUDA, Triton, or an equivalent), with deliberate choices around data layout, tiling, launch configuration, and fusion.
  • Use hardware counters and repeatable experiments to reason about memory traffic, cache behavior, register pressure, occupancy, instruction mix, and synchronization rather than optimizing from timing alone.
  • Build and maintain numerical-correctness coverage against clear reference implementations across meaningful shapes, dtypes, layouts, boundary conditions, and tolerance policies.
  • Maintain versioned microbenchmarks and workload-level regressions that record the environment, inputs, baselines, latency or throughput distributions, and known trade-offs behind each optimization.
  • Document each target’s supported kernel paths, constraints, and safe fallbacks, and provide compiler and runtime partners the operator-level capability information they need without owning compiler-wide lowering or production placement.

What you bring

  • Demonstrated low-level accelerator-kernel engineering in CUDA, Triton, or a comparable environment, including work on performance-sensitive tensor, reduction, normalization, attention, or matrix-multiplication-style operators.
  • Fluency with profiler-led performance analysis and the ability to connect measurements to the GPU memory hierarchy, parallel execution model, and instruction-level behavior.
  • A rigorous approach to correctness: experience constructing reference comparisons, numerical-tolerance policies, edge-case tests, and regression gates for optimized code.
  • Judgment about when specialization is warranted, how to state a kernel’s supported envelope, and how to retain a reliable fallback when assumptions do not hold.
  • Ability to communicate compact, reproducible evidence to compiler, runtime, research, and verification collaborators, separating an operator-level finding from a broader systems conclusion.

What progress looks like

  • A reproducible priority-kernel corpus exists for representative workloads, pairing baseline and optimized results with recorded software/hardware context, profiler evidence, and a clear statement of the applicable input regime.
  • Critical kernel changes are covered by automated reference comparisons and numerical-regression cases that exercise relevant shapes, layouts, dtypes, and boundary conditions; failures identify the affected configuration rather than merely reporting a generic mismatch.
  • Compiler and runtime partners can consume an up-to-date, evidence-backed description of kernel capabilities, limits, and fallback conditions for supported targets, making operator choices auditable without transferring ownership of lowering, placement, or fleet behavior.

Where this role fits

Owns kernel hot paths, not compiler-wide lowering or multi-node fleet orchestration.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email