Open role · Data Systems
Member of Technical Staff, Measurement & Data Infrastructure
Make observations across unlike machines comparable, traceable, and useful for learning.
General Diffusion is building compute world models across unlike machines; this role supplies the measurement layer that lets a workload run, environment, placement action, and observed outcome be compared with their full context intact. You will turn GD-X testbed and runtime observations into versioned, reproducible evidence for research and systems teams—without owning model training, placement policy, or the independent safety decision.
01 / The work
What you’ll work on
- Define and evolve versioned data contracts for workload shape, hardware and software environment, experiment or placement action, and observed outcome, including units, sampling semantics, and compatibility rules.
- Build ingestion and canonicalization paths that make measurements from heterogeneous machines comparable while retaining the raw records, calibration context, and derivation history needed to interpret them.
- Attach end-to-end provenance to datasets and derived slices: instrumentation and schema versions, workload revision, environment configuration, transformations, and quality status.
- Produce reproducible training and evaluation snapshots with explicit cohort definitions across architecture, workload, and time, so analyses can be rerun and compared without hidden data leakage.
- Implement automated checks for missing, late, duplicate, conflicting, or out-of-range records; surface schema breaks, instrumentation changes, and distribution shifts before they silently affect downstream conclusions.
- Publish clear producer and consumer interfaces with Runtime, Fleet, Research, and Safety teams, and make evidence fitness and known measurement limits visible rather than implicit.
02 / The background
What you bring
- Experience building telemetry, event, or data infrastructure that serves multiple producers and consumers in a distributed systems environment.
- Strong schema and data-contract judgment, including backward-compatible evolution, stable identifiers, units, and semantics that remain interpretable as systems change.
- Hands-on practice with lineage, dataset versioning, manifests, and reproducible data or ML evaluation workflows.
- Measurement rigor: you can reason about sampling, clock alignment, aggregation, calibration, confounders, and the difference between an observation and an inference.
- Ability to implement data-quality validation and operational diagnostics using tools such as SQL and Python, then explain the resulting limits clearly to systems and research partners.
03 / The evidence
What progress looks like
- A representative cross-architecture run can be traced from its raw records through its canonical dataset slice, with a versioned manifest that identifies the workload, environment, instrumentation, transformations, and applicable quality checks.
- Research and systems partners can independently reconstruct a selected training or evaluation cohort from a declared snapshot and obtain documented completeness, validity, and comparability results rather than relying on an informal export.
- The quality framework demonstrably catches seeded and observed failure modes—such as missing fields, incompatible units, schema changes, or conflicting measurements—and records the alert, triage context, and disposition.
04 / In the system
Where this role fits
Owns evidence infrastructure, not model training, policy learning, or safety sign-off.
Apply
Interested in this role?
Email your background and a short note about the work you would like to do.
careers@generaldiffusion.com