All open positions

Open role · Infrastructure

Member of Technical Staff, Distributed Systems & Fleet

Keep a multi-architecture testbed reproducible, observable, and reliable.

Run the operational foundation for GD-X, General Diffusion’s bare-metal, multi-architecture testbed, so experiments on unlike silicon are repeatable and diagnosable rather than one-off machine events. You will make fleet state, experiment conditions, failures, and recovery visible to the researchers and systems teams learning how workloads behave across heterogeneous compute—without owning the learning policy, kernels, or placement implementation.

What you’ll work on

  • Own machine inventory, bring-up, provisioning, retirement, and readiness checks across the heterogeneous testbed; make the state of each usable system explicit and auditable.
  • Build a repeatable experiment lifecycle—from reservation and environment preparation through execution, cleanup, and machine return—so comparable runs do not depend on undocumented operator steps.
  • Capture operational context around real-world traces and failures, including the machine and experiment conditions needed to distinguish fleet effects from workload behavior; partner with Measurement & Data Infrastructure on the shared evidence contract.
  • Develop fleet-health signals, dashboards, and alerts that connect host, network, accelerator, and experiment symptoms to actionable diagnosis while keeping paging paths understandable.
  • Engineer safe recovery paths for failed or suspect machines, including isolation, reprovisioning, validation, and documented return-to-service criteria.
  • Maintain practical incident runbooks and lead operational follow-through: triage, evidence preservation, root-cause analysis, and improvements validated through representative failure exercises.

What you bring

  • Experience operating production distributed systems on Linux, with ownership that spans deployment, diagnosis, and reliable day-two operation.
  • Hands-on experience with bare-metal infrastructure or a heterogeneous fleet, including hardware lifecycle, provisioning, configuration drift, and failure isolation.
  • Strong observability judgment: you can combine metrics, logs, traces, and targeted health checks into signals that support both rapid triage and retrospective analysis.
  • A rigorous approach to reproducibility in performance-sensitive or experimental systems, including recording the operational conditions that make results comparable.
  • Demonstrated incident-response discipline—clear runbooks, low-noise escalation, recovery procedures, and post-incident changes that are tested rather than merely documented.

What progress looks like

  • An auditable inventory and repeatable provisioning/readiness workflow provides evidence of each supported machine’s operational state before it enters an experiment.
  • Priority testbed runs yield comparable operational traces with sufficient machine and failure context to investigate whether observed differences arise from the workload or the environment, in collaboration with the data-infrastructure owner.
  • Dashboards, alerts, recovery runbooks, and recorded failure exercises demonstrate that representative fleet faults can be detected, contained, investigated, and returned to service through a defined operational path.

Where this role fits

Owns physical-system evidence and operation, not RL algorithms or accelerator kernels.

Apply

Interested in this role?

Email your background and a short note about the work you would like to do.

careers@generaldiffusion.com

Apply via Email