Governance background

Warning

Pre-implementation. This page describes proposed contracts. Behavior is subject to change before code lands.

The deeper rationale and policy for EMMA (Electromagnetic Multi-task Model Assessment) governance. The mental model, why neutrality is the moat and how the ring-fence is shown to hold, lives in Concepts / Neutrality; this page covers the mechanism, the maturation path, and the geopolitical tradeoff that shapes both.

Why governance is first-class

Superpose builds signal foundation models and operates EMMA, so the benchmark must be governed independently of the model team. A model-builder that grades its own model on its own private test set produces a result no competitor is bound by, and once the conflict exists it cannot be undone retroactively: no later ring-fence can erase the fact that the operator saw the held-out before any external model did. This is the single property no model vendor can replicate. Every model-builder-as-evaluator incumbent (PReD-Bench, LWM, the DeepSig paywalled successors) is structurally barred from it.

The ring-fence mechanism

Four instruments make the ring-fence auditable rather than asserted.

  • Operational separation. An independent steering committee, ring-fenced from Superpose’s foundation-model team, with a published conflict-of-interest (COI) policy. The committee charter is recorded as the proposed COICharter.

  • Third-party re-scoring. Superpose’s own model is re-scored on the private held-out set by an external co-steward on every release, with published logs. Each re-scoring run is a ReScoringLog, produced by HoldoutScorer and written by the proposed ReScoringLogWriter.

  • Pre-registration. The seed backbone’s architecture and training are declared before the held-out is unsealed, clinical-trial-style, and recorded as a PreRegistration in the proposed PreRegistrationStore. A posted operator score without a pre-registration is rejected by SubmissionValidator.

  • Visible vulnerability. The seeded baseline is deliberately beatable. A leaderboard where the operator sweeps every task is the surest credibility-killer, so the design assumes an external model will win early columns. This is also the venture-deciding signal tracked in Open questions.

Maturation path

Governance matures with adoption, following the sequencing every benchmark that became a standard actually followed.

  • v0.1: operational ring-fence. A named external co-steward physically holds the test-set seed and re-scores every Superpose submission; an auditable COI firewall sits between the model team and EMMA ops.

  • v1: steering committee. A cross-community committee is seated, with one elder per sub-community (communications, radar, spectrum, drone-RF, standards, testbeds). A formal COI charter takes over from the v0.1 operational agreement.

  • v1+: independent non-profit. Once a healthy set of external submitters exists, EMMA is operated by an independent non-profit under a Linux Foundation or Joint Development Foundation (JDF) vehicle, as MLCommons and SPEC (the Standard Performance Evaluation Corporation) are. Superpose becomes a member, not the operator.

The v1+ transition is gated on at least five external submitters, mirroring the roughly two-year lag between MLPerf’s launch and MLCommons’ formalization. Reversing the sequence, formalizing governance before a non-empty board, produces rules with nothing to protect.

The geopolitical double-edge

Defense and government legitimacy is a recruiting asset and an adoption risk. A US-led neutral alternative is a natural counter to a Chinese multi-task benchmark and is useful for DARPA (Defense Advanced Research Projects Agency) and AFRL (Air Force Research Laboratory) funding, and the DARPA RFMLS (Radio Frequency Machine Learning Systems) program (2017 to 2021) is the cautionary tale: a well-funded program that built capability but left no public governed benchmark because its data stayed classified. EMMA’s stated reason to exist is to not repeat that isolation failure.

The double-edge is that positioning EMMA as US-aligned risks making it look like a government project, which suppresses the international, especially European and Asian academic, adoption a cross-tribal benchmark needs. The mitigation is structural: the data license is CC-BY-4.0 (permissive, deliberately replacing RadioML’s NonCommercial license), submission is open, and the held-out is a frozen secret rather than a controlled artifact. The design must read as jurisdiction-neutral even while it recruits US defense legitimacy onto the steering ring. Open RF captures can hit US export controls (BIS, the Bureau of Industry and Security) and FCC (Federal Communications Commission) or NTIA (National Telecommunications and Information Administration) spectrum regulation, so the public board is seeded by academic and open-submission entries, not by classified models.

References

  • Mattson et al., “MLPerf Training Benchmark,” MLSys 2020, arXiv:1910.01500. MLCommons neutral governance formalized roughly two years after MLPerf launched; the sequencing precedent for the v1+ transition.

  • Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” IJCV 2015, arXiv:1409.0575. Seeded with baselines; governance followed adoption.

  • DARPA RFMLS program, 2017 to 2021. A funded RF fingerprinting program that produced no reusable public benchmark; the isolation-failure cautionary tale. (verify program record)

  • DeepSig, “RadioML datasets.” NonCommercial license incompatible with commercial foundation-model work; the CC-BY-4.0 data-license decision inverts it. (verify URL)

  • Han et al., “PReD” and PReD-Bench, 2026, arXiv:2603.28183. The operator-as-evaluator cautionary tale EMMA’s ring-fence is the structural answer to.

See Also