Literature¶
The master reading list for EMMA (Electromagnetic Multi-task Model Assessment). Every page that makes a load-bearing claim cites one of these sources in its ## References section. Entries carry a verifiable identifier (DOI, arXiv, standard, or PyPI version) and an intra-source location where applicable. Entries I have not yet verified against the primary source carry (verify); they must be checked before any are presented as fact in a validation report.
This page is a living index. New entries are added as the dataset recipes and validation reports require them.
Benchmark methodology and precedents¶
Yang et al., “SUPERB: Speech processing Universal PERformance Benchmark,” Interspeech 2021. The structural template EMMA follows: one frozen upstream backbone, lightweight downstream readouts, one aggregated score. arXiv:2105.01051.
Wang et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” EMNLP 2018. arXiv:1804.07461. Saturation signaled the move to SuperGLUE; the lesson EMMA applies by not leading with saturated tasks.
Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” NeurIPS 2019. arXiv:1905.00537.
Koh et al., “WILDS: A Benchmark of in-the-Wild Distribution Shifts,” ICML 2021. arXiv:2012.07421. A generalization-first precedent; the leave-one-X-out framing.
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” IJCV 2015. arXiv:1409.0575. Seeded with baselines; governance followed adoption.
Mattson et al., “MLPerf Training Benchmark,” MLSys 2020. arXiv:1910.01500. MLCommons neutral governance formalized roughly two years after MLPerf launched; the sequencing precedent.
Srivastava et al., “Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models (BIG-bench),” 2023. arXiv:2206.04615. The tribe-contributes-a-task adoption mechanism.
Angle-of-arrival and direction-finding¶
Schmidt, “Multiple emitter location and signal parameter estimation,” IEEE Trans. Acoustics, Speech, Signal Processing 1986. The MUSIC algorithm. DOI:10.1109/TASSP.1986.1164830. Canonical array-signal direction-finding; grounds that AoA is physically a continuous angular quantity.
Roy and Kailath, “ESPRIT: Estimation of signal parameters via rotational invariance techniques,” IEEE Trans. Acoustics, Speech, Signal Processing 1989. DOI:10.1109/29.32276.
Van Trees, “Optimum Array Processing,” Wiley 2002. The textbook of record for array processing.
OOD and distribution shift¶
Koh et al., WILDS (see above).
Ovadia et al., “Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,” NeurIPS 2019. arXiv:1906.02530.
Automatic modulation classification and RadioML¶
O’Shea, Corgan, and Clancy, “Convolutional Radio Modulation Recognition Networks,” 2016. arXiv:1602.04105.
O’Shea, Roy, and Clancy, “Over-the-Air Deep Learning Based Radio Signal Classification,” IEEE J-STSP 2018. The RadioML 2018 lineage.
DeepSig, “RadioML datasets,” datasets page. The publisher’s public note that RML2016 has known errata and is not used in DeepSig products. (verify URL: errata claim not confirmable from the rendered page)
Hanna and Hussain, “Robust Low-SNR Modulation Classification,” 2026. arXiv:2605.27673. Apparent model gaps on RML2018.01A collapse under matched hyperparameter search; grounds AMC is saturated.
RF fingerprinting and specific emitter identification (SEI)¶
DARPA RFMLS program, 2017 to 2021. The capability-without-public-benchmark cautionary tale. (verify program record)
Bihl, Bauer, and Temple, “Feature representation for RF fingerprinting,” 2017 and the broader SEI literature. (verify: venue not pinned)
“No Radio Left Behind,” 2019. RF device identification across capture sessions; the leave-one-unit-out motivation. (verify: authors and venue not confirmable)
Beam management and CSI feedback¶
3GPP TR 38.843, “Study on Artificial Intelligence (AI)/Machine Learning (ML) for NR,” current release. Ships AI/ML use cases (CSI feedback, beam management, positioning) with no dataset or baseline. (verify revision-year)
Alkhateeb, “DeepMIMO: A Generic Deep Learning Dataset for Millimeter Wave and Massive MIMO Applications,” 2019. arXiv:1902.06435. CSI-level channel datasets; the data source for the CSI camp.
Djordjevic, Ali, and Alkhateeb, “LWM: A Foundation Model for Wireless Channel Data,” 2024. arXiv:2411.08872. The CSI-native foundation-model lineage.
Drone and UAV detection¶
Media Inhof et al. (DroneRF), 2019. The de-facto drone-RF benchmark; single SDR. (verify: authorship not confirmable against primary source)
Coluccia et al. (DroneDetect / RFUAV), 2019 to 2025. Competing single-antenna drone-RF datasets. (verify: multiple datasets, no single citation pinned)
“Against the Monolithic Wireless World Model,” 2026. arXiv:2605.16689. The field’s position paper naming the missing shared RF substrate.
Radar, micro-Doppler, and SAR¶
Skolnik, “Introduction to Radar Systems,” McGraw-Hill. The radar textbook of record. (verify edition)
Chen et al., “Micro-Doppler Effect in Radar,” 2003 and the micro-Doppler literature. (verify: year ambiguous, 2003 IEE Proc. vs 2006 IEEE TAES)
MSTAR (Moving and Stationary Target Acquisition and Recognition) dataset, 1990s. The saturated, leaky SAR-ATR benchmark EMMA intends to move beyond. (verify: program/dataset, no citable identifier)
Signal-to-text and RF language¶
Debbah and the MBZUAI group, “RF-GPT,” 2026. arXiv:2602.14833. Spectrogram-based RF-language model.
“RF-Analyzer,” 2026. arXiv:2605.04676. Concedes sim-to-real generalization breaks in low-SNR and OOD regimes; the load-bearing sim-to-real caveat.
Han et al., “PReD: Pretrained Remote-sensing foundation model” and PReD-Bench, 2026. arXiv:2603.28183. Same group builds model, dataset, and benchmark and reports SOTA on its own bench; the operator-as-evaluator cautionary tale.
Mashaal and Abou-Zeid, “IQFM: A Raw IQ Foundation Model,” 2025. arXiv:2506.06718. The closest existing analog to the model EMMA evaluates; one raw I/Q backbone read across modulation, AoA, beam, and fingerprinting.
Leaderboard integrity, pre-registration, and anti-gaming¶
Tramer, Zhang, Juels, Reiter, and Ristenpart, “Stealing Machine Learning Models via Prediction APIs,” USENIX Security Symposium 2016. arXiv:1609.02943. Adversarial extraction of a private model from its prediction outputs; the residual risk EMMA’s query auditing monitors.
International Committee of Medical Journal Editors (ICMJE), “Clinical Trial Registration.” The pre-registration analogy for declaring a model architecture before a held-out is unsealed. (verify: no specific citable document pinned)
Blum et al., “PrivateFingerprinting,” and the broader membership-inference and attribute-inference literature. Probing a secret test set through repeated scoring queries. (verify: citation not confirmable)
Spectrum sensing¶
Yucek and Arslan, “A Survey of Spectrum Sensing Algorithms for Cognitive Radio Applications,” IEEE Commun. Surveys Tuts. 2009. The spectrum-sensing construct and detection taxonomy behind
E-SC-SENSE.ChangShuoRadioData / CSRD2025, 2025. arXiv:2508.19552. A large-scale open-source synthetic radio dataset for spectrum sensing (SISO/MISO/MIMO, IQ to spectrogram). Single-task, single-org, unhosted; validates the generator-as-test-set bet while showing the multi-task governed gap.
Metrics¶
Le Roux, Wisdom, Erdogan, and Hershey, “SDR: Half-baked or Well Done?,” ICASSP 2019. arXiv:1811.02508. SI-SDR for source separation; the metric for
E-SC-SEP.Papineni et al., “BLEU: a Method for Automatic Evaluation of Machine Translation,” ACL 2002. The BLEU (Bilingual Evaluation Understudy) captioning metric.
Banerjee and Lavie, “METEOR: An Automatic Metric for MT Evaluation,” ACL 2005.
Vedantam et al., “CIDEr: Consensus-based Image Description Evaluation,” CVPR 2015.
Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” EMNLP 2016. arXiv:1606.05250. Exact-match and F1 for QA.
Davis and Goadrich, “The relationship between Precision-Recall and ROC curves,” ICML 2006. AUC (area under the receiver operating characteristic curve) and AUROC grounding.
Brodersen et al., “The balanced accuracy and its posterior distribution,” ICPR 2010. Balanced accuracy.
Real-capture testbeds¶
The public testbeds behind the EMMA-REAL-OOD real-capture OOD (out-of-distribution) subset and the sim-to-real gap column.
Colosseum (Northeastern University), “Colosseum wireless network emulator.” The large-scale RF network emulator behind part of the real-capture subset. (verify citation)
POWDER (Powder Platform for Open Wireless Data-driven Experimental Research), University of Utah. The programmable city-scale wireless testbed behind part of the real-capture subset. (verify citation)
COSMOS (Cloud-enhanced Open Software defined Mobile wireless testbed for city-scale deployment), NYU / Rutgers / Columbia. The city-scale software-defined wireless testbed behind part of the real-capture subset. (verify citation)
Libraries and tools¶
torchmetrics >= 0.11. Classification, regression, and detection metrics (AUC, AUROC, F1, BLEU, etc.).scipy >= 1.11. Signal-processing primitives.numpy. Numerical primitives.torch >= 2.0. Tensors and autograd.pydantic >= 2.0. Schema validation.nltk. Natural Language Toolkit; the METEOR (Metric for Evaluation of Translation with Explicit ORdering) reduction forE-S2T-CAPcaptioning. (verify version)Apache Parquet, “Parquet File Format.” Columnar on-disk format for prediction rows and manifests. (verify: no pinned spec version)
Typer. The CLI (command-line interface) framework the
emmaharness adopts. (verify version)NVIDIA Sionna. Hoydis et al., “Sionna: An Open-Source Library for Next-Generation Physical Layer Research,” 2023. arXiv:2203.11854. The ray-tracing and link-level simulator rfgen builds on.
TorchSig. The signal-generation library rfgen composes for benchmark-compatible modulations.
NIST, “Secure Hash Standard (SHS),” FIPS PUB 180-4. The SHA-256 digest construction behind the EMMA content hash.
See Also¶
Landscape: the competitive map this reading list distills.
Validation methodology: the six-lens framework that consumes these citations.
Design decisions: the load-bearing choices each citation grounds.