All open positions

Open role · Research

Member of Technical Staff, RL & Compute Environments

Turn compute predictions into constrained, measurable decisions without delegating safety to a reward function.

Build the decision-learning environments that turn General Diffusion’s predictions of heterogeneous compute behavior into measurable placement and resource-allocation choices. You will define constrained RL problems, make offline evidence credible enough to inform bounded shadow or online evaluation, and keep the policy’s optimization work distinct from the independent controls that authorize actions. The work connects compute world models to policy optimization across unlike hardware without owning production execution, telemetry infrastructure, or safety enforcement.

What you’ll work on

  • Design representative compute-environment abstractions: state, feasible placement actions, transition assumptions, rewards, and explicit constraint signals for heterogeneous workloads.
  • Build reproducible offline evaluation and counterfactual-analysis workflows that compare candidate policies against baselines using held-out workload, architecture, and drift conditions.
  • Develop policy-learning and planning experiments that use action-conditioned performance predictions while reporting calibration limits and uncertainty-sensitive failure cases.
  • Define shadow-evaluation and narrowly scoped online-experiment protocols with predeclared metrics, stop conditions, and rollback criteria; partner with Runtime on execution but do not own the production placement path.
  • Measure how policy quality changes with workload mix, hardware capability differences, delayed effects, and distribution shift, separating reward improvement from operational reliability.
  • Work with Measurement & Data Infrastructure to specify the environment/action/outcome evidence needed for reproducible training and evaluation slices, without taking ownership of the underlying data platform.
  • Communicate decision evidence and known limitations to World Models, Runtime, and Safety & Verification; independent controls—not this role’s reward or policy—determine whether an action is permitted.

What you bring

  • Demonstrated work in reinforcement learning, contextual bandits, planning, or off-policy evaluation where state distributions change and decisions have delayed system effects.
  • Ability to turn a systems problem into a falsifiable experimental environment, including action semantics, reward trade-offs, baselines, held-out conditions, and variance-aware measurement.
  • Strong programming and experimental-systems practice: build reliable research pipelines, inspect failures, and make experiments repeatable rather than optimizing only a headline metric.
  • Practical scheduling or resource-allocation intuition, including the trade-offs among latency, throughput, utilization, contention, and feasibility across heterogeneous resources.
  • Sound judgment about safe empirical work: distinguish simulation or offline evidence from online evidence, set bounded experiment conditions, and state uncertainty plainly.
  • Clear cross-functional communication with ML researchers, runtime engineers, data/measurement partners, and independent safety reviewers.

What progress looks like

  • A documented compute-environment and evaluation contract exists for priority placement decisions, with explicit action feasibility, reward trade-offs, constraint signals, baselines, and known modeling assumptions.
  • Candidate policies can be compared reproducibly on held-out workload and hardware conditions, with decision-quality, variance, drift, and failure-mode results that make limits visible rather than obscuring them in aggregate rewards.
  • For experiments that progress beyond offline evaluation, evidence packages define predeclared success measures, bounded authority, stop/rollback conditions, and results from shadow or other appropriately controlled evaluation; Safety & Verification retains permissioning authority.

Where this role fits

Owns policy learning and evidence; independent controls decide whether an action is permitted.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email