Neutrality¶
Pre-launch / design spec
The governance structure, licensing terms, and re-scoring contracts below are proposed. They finalize with the v0.1 release and the seated external co-steward.
EMMA (Electromagnetic Multi-task Model Assessment) is stewarded by Superpose and built on data from
rfgen.
Why neutrality is the moat¶
Superpose also builds signal foundation models - so the benchmark must be governed independently of the model team, or its scores are not credible. Neutrality cannot be manufactured retroactively: once a model-builder has graded its own model on its own private test set, no later “ring-fence” can undo the conflict. This is the single property no model vendor can replicate, and the reason EMMA exists as a separate, governed evaluation surface rather than another self- reported result.
How the ring-fence works¶
Independent steering committee, ring-fenced from Superpose’s foundation-model team, with a published conflict-of-interest (COI) policy.
Third-party re-scoring: Superpose’s own model is re-scored on the private held-out set by an external co-steward on every release, with published logs (ReScoringLog) scored by HoldoutScorer.
Pre-registered baselines: the seed backbone’s architecture and training are declared before the held-out is unsealed, clinical-trial-style, and recorded as a PreRegistration.
Visible vulnerability: the seeded baseline is deliberately beatable. A leaderboard where the operator sweeps every task is the surest credibility- killer - so the design assumes an external model will win early columns.
Governance matures with adoption¶
v0.1 - operational ring-fence: a named external co-steward physically holds the test-set seed and re-scores every Superpose submission; an auditable COI firewall between the model team and EMMA ops.
v1 - steering committee: a cross-community committee seated, with one elder per sub-community; a formal COI charter.
v1+ - independent non-profit: once a healthy set of external submitters exists, EMMA is operated by an independent non-profit (under a Linux Foundation / Joint Development Foundation vehicle, as MLCommons and SPEC are). Superpose becomes a member, not the operator.
Licensing (planned)¶
Data: CC-BY-4.0 (permissive - and a deliberate replacement for RadioML’s NonCommercial license, which is incompatible with commercial foundation-model work).
Evaluation code: Apache-2.0.
References¶
Mattson et al., “MLPerf Training Benchmark,” MLSys 2020, arXiv:1910.01500. MLCommons neutral governance formalized roughly two years after MLPerf launched; the sequencing precedent for EMMA’s v1+ independent-operator transition.
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” IJCV 2015, arXiv:1409.0575. Seeded with baselines; governance followed adoption.
DeepSig, “RadioML datasets.” Published under a NonCommercial license incompatible with commercial foundation-model work; the licensing decision CC-BY-4.0 inverts. (verify URL)
DARPA RFMLS program, 2017 to 2021. The capability-without-public-benchmark cautionary tale: a funded RF fingerprinting program that produced no reusable public benchmark. (verify program record)
See Also¶
Leaderboard: the reporting surface the ring-fence protects.
Reproducibility: the held-out secret seed and auditable result triple.
Sim-to-real gap: the honesty metric that makes self-reported synthetic scores checkable.
Signal as a modality: why governance is one of the four divergences from SUPERB.