Design decisions¶
Warning
Pre-implementation. This page describes proposed contracts. Behavior is subject to change before code lands.
The choices below are load-bearing: changing any one of them changes what EMMA (Electromagnetic Multi-task Model Assessment) measures. Each is stated with the alternatives considered and why they were rejected. The mental models these decisions rationalize live in Concepts; the exact contracts they produce live in Reference.
EMMA rests on four design pillars: generalization is the score; the canonical input is raw, phase-coherent, multi-antenna I/Q (in-phase and quadrature); the benchmark is reproducible by construction; and governance is neutral by design. The decisions below make those pillars concrete.
Localization is the flagship pillar¶
AoA (angle-of-arrival) / DoA (direction-of-arrival) regression is the v0.1 flagship task, not merely the first task in the list. Three properties make it load-bearing.
First, localization physically requires inter-antenna phase coherence. The phase difference between antennas is exactly what direction-finding reads, so the moment the input becomes a magnitude spectrogram or a CSI (channel state information) tensor, AoA and beam prediction become unevaluable from the input. Localization is therefore the justification for the entire raw multi-antenna I/Q contract, not just one task among five.
Second, localization is legible across communities. Communications cares about positioning; radar cares about direction-finding; the drone-RF (radio-frequency) community cares about localizing emitters. A flagship that only one sub-community recognizes cannot pull adoption across EMMA’s audiences.
Third, localization is unclaimed. Every public RF dataset is single-antenna (RadioML, TorchSig) or a derived representation (DeepMIMO at the CSI level), so array direction-finding is literally unevaluable in the existing literature.
Alternatives considered:
AMC (automatic modulation classification) as flagship. Rejected: it is saturated, single-antenna-solvable, and its publisher states the data has known errata and is not used in its products. Leading with it would put EMMA behind on day one.
Beam management as flagship. Deferred to v0.2: it pulls in the 3GPP (3rd Generation Partnership Project) audience but does not, by itself, require the multi-antenna phase coherence argument to be made first.
Drone detection as flagship. Kept as a headline identity task, not the flagship: it is a near-monopoly for EMMA but it does not physically require the array the way AoA does.
See Tasks for the pillar layout and Tasks validation for the construct-validity argument.
The canonical input is raw multi-antenna I/Q¶
Every EMMA task reads the same record: raw, complex, multi-antenna I/Q captured simultaneously across antenna elements, preserving inter-antenna phase coherence. The alternative representations each destroy the one property localization and beam tasks need.
A magnitude spectrogram discards phase. Inter-antenna phase differences are exactly what direction-finding reads, so once the spectrogram is taken, AoA and beam prediction are physically unevaluable from the input.
A CSI tensor collapses the time-domain waveform into a frequency-domain channel description. The CSI camp (LWM, DeepMIMO) operates here natively; EMMA offers CSI only as a bridge task (
E-CH-CSI), never as the primary input.A single-antenna capture has no inter-antenna phase to begin with. Every existing RF dataset is single-antenna (RadioML, TorchSig, DroneRF) or a derived representation (DeepMIMO), so the raw multi-antenna I/Q layer is the one no incumbent natively provides.
The canonical record is therefore raw I/Q, and derived views are computed inside a readout head if a model wants them. The input layer never makes that lossy choice for the model. See Data model for the tensor contract.
Alternatives considered:
Adopt the CSI camp’s representation. Rejected as primary: it discards exactly the phase coherence that justifies the array, and it concedes the benchmark surface to one sub-community.
Standardize on spectrograms. Rejected: the entire RF-language frontier (RF-GPT, RF-Analyzer, PReD-Bench) already does this and discards phase; EMMA’s moat is the representation they dropped.
Modulation classification is a continuity column, excluded from the aggregate¶
E-ID-AMC ships at v0.1 but is reported, never aggregated. It exists so the AMC community can map from RadioML to EMMA, not to compete with it as a headline. Three reasons keep it out of the OOD average.
AMC is saturated: apparent model gaps on RadioML collapse to a few percentage points under matched hyperparameter search, so the column carries little signal. AMC is single-antenna-solvable: it does not require the array, so scoring it alongside localization would dilute the multi-antenna property the benchmark exists to measure. AMC carries a publisher caveat: DeepSig states its RadioML datasets carry known errata and are not used in its products, so a benchmark that led with AMC would inherit a legitimacy vacuum.
The column stays so the AMC community has an on-ramp; the aggregate excludes it so the headline measures what saturated benchmarks cannot. See Tasks for the exclusion note and Aggregation for the exact rule.
Alternatives considered:
Drop AMC entirely. Rejected: the AMC community is large, and a continuity column costs nothing while removing it would look like exclusion.
Include AMC in the aggregate. Rejected: a saturated, single-antenna metric would dilute the transfer signal and reward models that memorize RadioML.
OOD transfer is the headline, not in-distribution accuracy¶
The leaderboard headline is a normalized OOD (out-of-distribution) average, computed under a controlled transfer axis per task (unseen channel environment, unseen device, unseen frequency band, unseen SNR (signal-to-noise ratio) regime). In-distribution accuracy is reported only as context.
In-distribution accuracy is what every saturated single-task dataset already measures. A model can score high on it while failing the moment the channel, device, or band changes. Scoring only transfer forces the benchmark to measure the property that survives deployment. This is the single divergence from SUPERB (Speech processing Universal PERformance Benchmark), which scores in-distribution downstream accuracy because speech has abundant real captures; RF does not.
See Generalization for the protocol and OOD protocol for the fold-construction contract.
Alternatives considered:
Report in-distribution accuracy as the headline. Rejected: it would put EMMA behind RadioML and MSTAR (Moving and Stationary Target Acquisition and Recognition) on day one, since both already optimize that quantity.
Report transfer only for some tasks. Rejected: a partial OOD contract lets a model hide overfitting on the tasks where the axis is missing. Every scored task carries an axis.
EMMA consumes rfgen; it does not generate¶
EMMA is a benchmark library. It owns dataset recipes (pinned rfgen configurations), loaders, label extraction, the evaluation harness, the leaderboard, and governance. It does not own emitters, channels, propagation, antenna patterns, or any signal generation. Those live in rfgen, and EMMA references rfgen’s config schema by name rather than redefining it.
The boundary exists for two reasons. First, it keeps the benchmark honest about scope: a benchmark is a dataset plus correct documentation plus an evaluation protocol, and generation is a separate engineering surface with its own validation burden. Second, it keeps the generator replaceable and auditable: a dataset page pins the exact rfgen commit, config, and seed, so the generator, not a static file, is the test set. A bug is fixed by shipping a new content-hashed release without invalidating provenance.
See Datasets and rfgen for the boundary and the smell test in PRINCIPLES for the rule that catches a crossed boundary.
Alternatives considered:
Generate I/Q inside EMMA. Rejected: it would couple the benchmark to one physics implementation, blur the benchmark-or-library question, and duplicate work rfgen already owns and validates.
Vendor a frozen snapshot of rfgen output as static files. Rejected: static-file benchmarks inherit frozen label errors, dying eval servers, and leaking test sets, the exact failure modes the generator-as-test-set design avoids.
References¶
Schmidt, “Multiple emitter location and signal parameter estimation,” IEEE Trans. Acoustics, Speech, Signal Processing 1986, DOI:10.1109/TASSP.1986.1164830. The MUSIC algorithm; grounds that AoA is physically a continuous angular quantity read from inter-antenna phase.
Yang et al., “SUPERB: Speech processing Universal PERformance Benchmark,” Interspeech 2021, arXiv:2105.01051. The structural template EMMA follows and diverges from on the headline metric.
Koh et al., “WILDS: A Benchmark of in-the-Wild Distribution Shifts,” ICML 2021, arXiv:2012.07421. The generalization-first precedent and the leave-one-X-out framing.
Hanna and Hussain, “Robust Low-SNR Modulation Classification,” 2026, arXiv:2605.27673. Apparent model gaps on RadioML collapse under matched hyperparameter search; grounds the AMC continuity-column exclusion.
DeepSig, “RadioML datasets,” datasets page (deepsig.ai/datasets). Published under a NonCommercial license; the publisher’s public note that RML2016 has known errata and is not used in DeepSig products. Grounds the AMC and licensing decisions. (verify URL)
NVIDIA Sionna: Hoydis et al., “Sionna: An Open-Source Library for Next-Generation Physical Layer Research,” 2023, arXiv:2203.11854. The simulator rfgen builds on; the source of the generation EMMA consumes but does not own.
See Also¶
Landscape: the incumbents and the gap these decisions exploit.
Signal as a modality: the four divergences from SUPERB in concept form.
Data model: the raw multi-antenna I/Q tensor contract.
Generalization: the OOD axes and the transfer score.
Datasets and rfgen: the benchmark-or-library boundary.
Open questions: the unresolved decisions behind these choices.