All open positions

Open role · Research

Member of Technical Staff, Compute World Models

Model what happens to a workload before a placement decision reaches a real machine.

Build action-conditioned models that forecast how a workload will behave before a placement reaches a real machine—across latency, throughput, contention, and available headroom. You will own the predictor and the evidence for when it transfers or fails to transfer; RL & Compute Environments owns policy learning, Heterogeneous Runtime & Placement owns execution, Measurement & Data Infrastructure owns evidence plumbing, and Safety & Formal Verification independently determines which actions are allowed.

What you’ll work on

  • Develop action-conditioned dynamics models that predict compute-system behavior across hardware and workload timescales from observed execution traces and measured outcomes.
  • Define target variables and uncertainty estimates for latency, throughput, contention, and headroom, with diagnostics that make miscalibration and failure modes legible to systems partners.
  • Design rigorous evaluation splits that withhold architectures, workload mixes, and load regimes; measure transfer, calibration, and degradation rather than relying on aggregate in-distribution accuracy.
  • Use replay and shadow-mode experiments to test predicted consequences of candidate placement actions before those predictions inform a production decision.
  • Work with Measurement & Data Infrastructure to specify the trace fields, provenance, and reproducible training/evaluation slices the model needs, without taking ownership of the shared data platform.
  • Publish a clear predictor contract—forecasts, uncertainty, assumptions, and known limits—for RL & Compute Environments and Heterogeneous Runtime & Placement, while leaving policy selection, execution, and action permissioning to their respective owners.

What you bring

  • Demonstrated work in world models, system identification, probabilistic forecasting, learned dynamics, or model-based reinforcement learning, including action-conditioned state prediction.
  • A record of evaluating models under distribution shift, especially with held-out environments, hardware configurations, workload regimes, or time horizons; able to reason about calibration as well as point error.
  • Experience instrumenting or analyzing real compute systems through execution traces, telemetry, profiler data, or performance counters, with practical intuition for latency, throughput, queueing, and contention.
  • Strong experimental engineering skills: can turn messy measured behavior into reproducible datasets, training runs, ablations, and failure analyses that other engineers can inspect.
  • Judgment about the boundary between predicting system outcomes and choosing or authorizing actions; comfortable supplying evidence to policy, runtime, and safety teams without collapsing those responsibilities.

What progress looks like

  • A versioned evaluation suite reports forecast error, calibration, uncertainty coverage, and failure slices for latency, throughput, contention, and headroom on held-out architectures and load patterns.
  • Replay and shadow-mode results document where the model's action-conditioned forecasts remain reliable, where they break under shift, and the evidence-based conditions for limiting their use.
  • An integration contract and reproducible test cases let policy and runtime teams consume forecasts and uncertainty while demonstrating that independent safety checks—not model confidence—remain the authority gate.

Where this role fits

Owns the predictor and its evidence, not the production policy or the independently enforced boundary.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email