STUDY / 02 · Discover Artificial Intelligence, 2026 · article in press (supplied version)
Attention-QELM-GWO
Attention-guided quantum-inspired extreme learning machine with grey wolf optimisation for rare-attack detection in IIoT networks.
Contribution: conceptualization, methodology design, implementation, and initial framework development.
Quantum-inspired here means classical trigonometric random features used to initialise a fixed hidden layer. No quantum hardware is used and no quantum advantage is claimed.
Interactive
Model explorer
Feature tokens, attention, fixed random hidden layers and solved output weights, with the grey wolf search drawn as an offline step beside training.
Section 4.2, Table 5
Feature tokens
48 features enter a linear layer with GELU, a 64-dimensional bottleneck, and are then laid out as 8 tokens of 32 dimensions with a learned feature-index embedding. The published text describes a reshape from 64 to 8 by 32, which does not reshape; the repository adds an explicit expansion and logs it as an ambiguity.
Offline search, beside training
Grey wolf optimisation trace
The search runs before the closed-form solve and chooses hidden width, layer count, activation, regularisation and the initialisation angle. It is not part of inference. The curve below is the repository's run on its fixture.
Loading trace
Phase 1 encoder training
Learning curves
Loading log
Published protocol
- Edge-IIoTset: 1,909,671 processed records, 15 classes, 48 retained features after dropping constant, duplicate and metadata columns.
- Stratified 80/20 split (1,527,736 training, 381,935 held out). Train-fit median and mode imputation; min-max scaling to [0, 1].
- Grey wolf search fitness: 5-fold cross-validated macro-F1 on the training partition; 20 wolves, at most 50 iterations, early stop on under 0.0001 improvement over 10 iterations.
- Seeds 7, 13, 21, 42 and 99; results reported as means.
- The split is random rather than temporal or device-grouped, so shortcut features and near-duplicate flows are not excluded by the protocol itself; the repository's audit lists identifier-like columns in its fixture for this reason.
Ambiguity log
Places where the supplied text could not be followed literally. Each entry names the choice the repository made so it can be reversed if the authors clarify.
- Impossible 64 to 8 x 32 reshape
- The encoder text describes a 64-dimensional bottleneck reshaped into 8 tokens of 32 dimensions (256 values). The repository inserts an explicit 64 to 256 expansion and names it in the trace as expand_64_to_256.
- L1 naming
- The term the paper calls lambda1 (L1) appears in the closed-form solution as an additive identity term, which is a ridge penalty. The repository keeps the paper's name and documents that it is not an L1 penalty.
- Single versus three-layer ELM
- Prose describes a single hidden layer; the final configuration table gives L = 3 with 512 units. Both are implemented and each demo run states which it used.
- Global angle
- The initialisation angle theta appears both as a per-weight uniform draw and as a single searched value (theta = 1.84). The repository treats the searched value as a phase offset added to the per-weight draw and records this choice.
- Feature manifest
- The 48 retained features are not listed by name in the supplied version. The repository uses a documented synthetic manifest of 48 columns and cannot claim column-level fidelity.
The supplied version also reports two different Friedman statistics for the same comparison; neither is reproduced or repeated here.