Literature

The master reading list for EMMA (Electromagnetic Multi-task Model Assessment). Every page that makes a load-bearing claim cites one of these sources in its ## References section. Entries carry a verifiable identifier (DOI, arXiv, standard, or PyPI version) and an intra-source location where applicable. Entries I have not yet verified against the primary source carry (verify); they must be checked before any are presented as fact in a validation report.

This page is a living index. New entries are added as the dataset recipes and validation reports require them.

Benchmark methodology and precedents

  • Yang et al., “SUPERB: Speech processing Universal PERformance Benchmark,” Interspeech 2021. The structural template EMMA follows: one frozen upstream backbone, lightweight downstream readouts, one aggregated score. arXiv:2105.01051.

  • Wang et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” EMNLP 2018. arXiv:1804.07461. Saturation signaled the move to SuperGLUE; the lesson EMMA applies by not leading with saturated tasks.

  • Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” NeurIPS 2019. arXiv:1905.00537.

  • Koh et al., “WILDS: A Benchmark of in-the-Wild Distribution Shifts,” ICML 2021. arXiv:2012.07421. A generalization-first precedent; the leave-one-X-out framing.

  • Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” IJCV 2015. arXiv:1409.0575. Seeded with baselines; governance followed adoption.

  • Mattson et al., “MLPerf Training Benchmark,” MLSys 2020. arXiv:1910.01500. MLCommons neutral governance formalized roughly two years after MLPerf launched; the sequencing precedent.

  • Srivastava et al., “Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models (BIG-bench),” 2023. arXiv:2206.04615. The tribe-contributes-a-task adoption mechanism.

Angle-of-arrival and direction-finding

  • Schmidt, “Multiple emitter location and signal parameter estimation,” IEEE Trans. Acoustics, Speech, Signal Processing 1986. The MUSIC algorithm. DOI:10.1109/TASSP.1986.1164830. Canonical array-signal direction-finding; grounds that AoA is physically a continuous angular quantity.

  • Roy and Kailath, “ESPRIT: Estimation of signal parameters via rotational invariance techniques,” IEEE Trans. Acoustics, Speech, Signal Processing 1989. DOI:10.1109/29.32276.

  • Van Trees, “Optimum Array Processing,” Wiley 2002. The textbook of record for array processing.

OOD and distribution shift

  • Koh et al., WILDS (see above).

  • Ovadia et al., “Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,” NeurIPS 2019. arXiv:1906.02530.

Automatic modulation classification and RadioML

  • O’Shea, Corgan, and Clancy, “Convolutional Radio Modulation Recognition Networks,” 2016. arXiv:1602.04105.

  • O’Shea, Roy, and Clancy, “Over-the-Air Deep Learning Based Radio Signal Classification,” IEEE J-STSP 2018. The RadioML 2018 lineage.

  • DeepSig, “RadioML datasets,” datasets page. The publisher’s public note that RML2016 has known errata and is not used in DeepSig products. (verify URL: errata claim not confirmable from the rendered page)

  • Hanna and Hussain, “Robust Low-SNR Modulation Classification,” 2026. arXiv:2605.27673. Apparent model gaps on RML2018.01A collapse under matched hyperparameter search; grounds AMC is saturated.

RF fingerprinting and specific emitter identification (SEI)

  • DARPA RFMLS program, 2017 to 2021. The capability-without-public-benchmark cautionary tale. (verify program record)

  • Bihl, Bauer, and Temple, “Feature representation for RF fingerprinting,” 2017 and the broader SEI literature. (verify: venue not pinned)

  • “No Radio Left Behind,” 2019. RF device identification across capture sessions; the leave-one-unit-out motivation. (verify: authors and venue not confirmable)

Beam management and CSI feedback

  • 3GPP TR 38.843, “Study on Artificial Intelligence (AI)/Machine Learning (ML) for NR,” current release. Ships AI/ML use cases (CSI feedback, beam management, positioning) with no dataset or baseline. (verify revision-year)

  • Alkhateeb, “DeepMIMO: A Generic Deep Learning Dataset for Millimeter Wave and Massive MIMO Applications,” 2019. arXiv:1902.06435. CSI-level channel datasets; the data source for the CSI camp.

  • Djordjevic, Ali, and Alkhateeb, “LWM: A Foundation Model for Wireless Channel Data,” 2024. arXiv:2411.08872. The CSI-native foundation-model lineage.

Drone and UAV detection

  • Media Inhof et al. (DroneRF), 2019. The de-facto drone-RF benchmark; single SDR. (verify: authorship not confirmable against primary source)

  • Coluccia et al. (DroneDetect / RFUAV), 2019 to 2025. Competing single-antenna drone-RF datasets. (verify: multiple datasets, no single citation pinned)

  • “Against the Monolithic Wireless World Model,” 2026. arXiv:2605.16689. The field’s position paper naming the missing shared RF substrate.

Radar, micro-Doppler, and SAR

  • Skolnik, “Introduction to Radar Systems,” McGraw-Hill. The radar textbook of record. (verify edition)

  • Chen et al., “Micro-Doppler Effect in Radar,” 2003 and the micro-Doppler literature. (verify: year ambiguous, 2003 IEE Proc. vs 2006 IEEE TAES)

  • MSTAR (Moving and Stationary Target Acquisition and Recognition) dataset, 1990s. The saturated, leaky SAR-ATR benchmark EMMA intends to move beyond. (verify: program/dataset, no citable identifier)

Signal-to-text and RF language

  • Debbah and the MBZUAI group, “RF-GPT,” 2026. arXiv:2602.14833. Spectrogram-based RF-language model.

  • “RF-Analyzer,” 2026. arXiv:2605.04676. Concedes sim-to-real generalization breaks in low-SNR and OOD regimes; the load-bearing sim-to-real caveat.

  • Han et al., “PReD: Pretrained Remote-sensing foundation model” and PReD-Bench, 2026. arXiv:2603.28183. Same group builds model, dataset, and benchmark and reports SOTA on its own bench; the operator-as-evaluator cautionary tale.

  • Mashaal and Abou-Zeid, “IQFM: A Raw IQ Foundation Model,” 2025. arXiv:2506.06718. The closest existing analog to the model EMMA evaluates; one raw I/Q backbone read across modulation, AoA, beam, and fingerprinting.

Leaderboard integrity, pre-registration, and anti-gaming

  • Tramer, Zhang, Juels, Reiter, and Ristenpart, “Stealing Machine Learning Models via Prediction APIs,” USENIX Security Symposium 2016. arXiv:1609.02943. Adversarial extraction of a private model from its prediction outputs; the residual risk EMMA’s query auditing monitors.

  • International Committee of Medical Journal Editors (ICMJE), “Clinical Trial Registration.” The pre-registration analogy for declaring a model architecture before a held-out is unsealed. (verify: no specific citable document pinned)

  • Blum et al., “PrivateFingerprinting,” and the broader membership-inference and attribute-inference literature. Probing a secret test set through repeated scoring queries. (verify: citation not confirmable)

Spectrum sensing

  • Yucek and Arslan, “A Survey of Spectrum Sensing Algorithms for Cognitive Radio Applications,” IEEE Commun. Surveys Tuts. 2009. The spectrum-sensing construct and detection taxonomy behind E-SC-SENSE.

  • ChangShuoRadioData / CSRD2025, 2025. arXiv:2508.19552. A large-scale open-source synthetic radio dataset for spectrum sensing (SISO/MISO/MIMO, IQ to spectrogram). Single-task, single-org, unhosted; validates the generator-as-test-set bet while showing the multi-task governed gap.

Metrics

  • Le Roux, Wisdom, Erdogan, and Hershey, “SDR: Half-baked or Well Done?,” ICASSP 2019. arXiv:1811.02508. SI-SDR for source separation; the metric for E-SC-SEP.

  • Papineni et al., “BLEU: a Method for Automatic Evaluation of Machine Translation,” ACL 2002. The BLEU (Bilingual Evaluation Understudy) captioning metric.

  • Banerjee and Lavie, “METEOR: An Automatic Metric for MT Evaluation,” ACL 2005.

  • Vedantam et al., “CIDEr: Consensus-based Image Description Evaluation,” CVPR 2015.

  • Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” EMNLP 2016. arXiv:1606.05250. Exact-match and F1 for QA.

  • Davis and Goadrich, “The relationship between Precision-Recall and ROC curves,” ICML 2006. AUC (area under the receiver operating characteristic curve) and AUROC grounding.

  • Brodersen et al., “The balanced accuracy and its posterior distribution,” ICPR 2010. Balanced accuracy.

Real-capture testbeds

The public testbeds behind the EMMA-REAL-OOD real-capture OOD (out-of-distribution) subset and the sim-to-real gap column.

  • Colosseum (Northeastern University), “Colosseum wireless network emulator.” The large-scale RF network emulator behind part of the real-capture subset. (verify citation)

  • POWDER (Powder Platform for Open Wireless Data-driven Experimental Research), University of Utah. The programmable city-scale wireless testbed behind part of the real-capture subset. (verify citation)

  • COSMOS (Cloud-enhanced Open Software defined Mobile wireless testbed for city-scale deployment), NYU / Rutgers / Columbia. The city-scale software-defined wireless testbed behind part of the real-capture subset. (verify citation)

Libraries and tools

  • torchmetrics >= 0.11. Classification, regression, and detection metrics (AUC, AUROC, F1, BLEU, etc.).

  • scipy >= 1.11. Signal-processing primitives.

  • numpy. Numerical primitives.

  • torch >= 2.0. Tensors and autograd.

  • pydantic >= 2.0. Schema validation.

  • nltk. Natural Language Toolkit; the METEOR (Metric for Evaluation of Translation with Explicit ORdering) reduction for E-S2T-CAP captioning. (verify version)

  • Apache Parquet, “Parquet File Format.” Columnar on-disk format for prediction rows and manifests. (verify: no pinned spec version)

  • Typer. The CLI (command-line interface) framework the emma harness adopts. (verify version)

  • NVIDIA Sionna. Hoydis et al., “Sionna: An Open-Source Library for Next-Generation Physical Layer Research,” 2023. arXiv:2203.11854. The ray-tracing and link-level simulator rfgen builds on.

  • TorchSig. The signal-generation library rfgen composes for benchmark-compatible modulations.

  • NIST, “Secure Hash Standard (SHS),” FIPS PUB 180-4. The SHA-256 digest construction behind the EMMA content hash.

See Also