Neutrality

Pre-launch / design spec

The governance structure, licensing terms, and re-scoring contracts below are proposed. They finalize with the v0.1 release and the seated external co-steward.

EMMA (Electromagnetic Multi-task Model Assessment) is stewarded by Superpose and built on data from rfgen.

Why neutrality is the moat

Superpose also builds signal foundation models - so the benchmark must be governed independently of the model team, or its scores are not credible. Neutrality cannot be manufactured retroactively: once a model-builder has graded its own model on its own private test set, no later “ring-fence” can undo the conflict. This is the single property no model vendor can replicate, and the reason EMMA exists as a separate, governed evaluation surface rather than another self- reported result.

How the ring-fence works

  • Independent steering committee, ring-fenced from Superpose’s foundation-model team, with a published conflict-of-interest (COI) policy.

  • Third-party re-scoring: Superpose’s own model is re-scored on the private held-out set by an external co-steward on every release, with published logs (ReScoringLog) scored by HoldoutScorer.

  • Pre-registered baselines: the seed backbone’s architecture and training are declared before the held-out is unsealed, clinical-trial-style, and recorded as a PreRegistration.

  • Visible vulnerability: the seeded baseline is deliberately beatable. A leaderboard where the operator sweeps every task is the surest credibility- killer - so the design assumes an external model will win early columns.

Governance matures with adoption

  • v0.1 - operational ring-fence: a named external co-steward physically holds the test-set seed and re-scores every Superpose submission; an auditable COI firewall between the model team and EMMA ops.

  • v1 - steering committee: a cross-community committee seated, with one elder per sub-community; a formal COI charter.

  • v1+ - independent non-profit: once a healthy set of external submitters exists, EMMA is operated by an independent non-profit (under a Linux Foundation / Joint Development Foundation vehicle, as MLCommons and SPEC are). Superpose becomes a member, not the operator.

Licensing (planned)

  • Data: CC-BY-4.0 (permissive - and a deliberate replacement for RadioML’s NonCommercial license, which is incompatible with commercial foundation-model work).

  • Evaluation code: Apache-2.0.

References

  • Mattson et al., “MLPerf Training Benchmark,” MLSys 2020, arXiv:1910.01500. MLCommons neutral governance formalized roughly two years after MLPerf launched; the sequencing precedent for EMMA’s v1+ independent-operator transition.

  • Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” IJCV 2015, arXiv:1409.0575. Seeded with baselines; governance followed adoption.

  • DeepSig, “RadioML datasets.” Published under a NonCommercial license incompatible with commercial foundation-model work; the licensing decision CC-BY-4.0 inverts. (verify URL)

  • DARPA RFMLS program, 2017 to 2021. The capability-without-public-benchmark cautionary tale: a funded RF fingerprinting program that produced no reusable public benchmark. (verify program record)

See Also