Benchmark Arena
Every claim should be executable
Compare the published cases within their stated dataset and metric boundaries. Filtering, model compilation and advisory control use different tests. The tables preserve recorded results, including zeros and non-applicable cases.
Choose a challenge
Three evidence tracks, three direct product routes
Real-waveform denoising with known injected noise, three latency budgets and explicit identity abstention.
Open filtration table →MODELfree-runCompiled nonlinear dynamics against a declared stable baseline on a held-out chronological tail.
Open NDC table →CONTROLfull-modelSDRE against a frozen linear comparator under the same model, state and command boundary.
Open control result →How to read this page
Six terms, once
dB (RMSE-reduction) — how much error the filter removed vs the noisy input; higher is better. Rule of thumb: +6 dB cuts the error roughly in half; a negative number means the filter made the signal worse than doing nothing.
SNR — signal-to-noise ratio of the input: 30 dB is nearly clean, 0 dB is noise as loud as the signal.
on / flt / OFF (on product pages: instant = A1 = on · low-latency = A2 = flt · offline = A3 = OFF) — the three outputs of the same engine at three latency budgets: on = zero-lag (sees past + current sample only), flt = 120 samples of delay, OFF = offline batch that re-processes the retained signal tail.
NRMSE & free-run — normalized prediction error during a rollout using the model’s own previous predictions after initialization. Free-run does not establish an independent test split; check whether the evaluation data were also used in model selection.
Settling time & effort — response time and the specified control-effort metric. A controller supplied with true equations is a reference comparator; calling it an oracle does not prove it is globally optimal.
Constraint violation & ms/step — the recorded limit violation and time per control step. A rounded 0.000 alone is not proof of no violations; inspect the tolerance and event count in the protocol.
1 · Signal filtration
Real-waveform service acceptance
Sixteen real waveform datasets from three source archives were evaluated at 20 dB and 5 dB with reproducible injected Gaussian noise: 32 planned cases, 32 completed, zero failures. Gain is RMSE reduction relative to the noisy input; 0 dB is an explicit Identity/no-regression abstention.
| Output | Evaluated | Mean gain | Median | Maximum |
|---|---|---|---|---|
| A1 online | 32 | 0.827 dB | 0.000 dB | 3.736 dB |
| A2 causal delayed | 30 | 2.808 dB | 3.367 dB | 5.558 dB |
| A3 offline | 32 | 6.460 dB | 4.864 dB | 21.031 dB |
A2 has a declared logical delay of 120 samples. Two short Nile records were not applicable for A2 because they contain fewer than the required 121 samples; they were not counted as passes. This protocol measures denoising on real-shaped waveforms with known injected noise; it is not a commercial-competitor comparison.
All 32 cases — the full table
| Dataset | SNR, dB | Online A1, dB | Min-latency A2, dB | Offline A3, dB |
|---|---|---|---|---|
| Cascaded_Tanks_test | 20 | 0.411 | 3.635 | 8.643 |
| Cascaded_Tanks_test | 5 | 2.811 | 4.067 | 10.337 |
| Cascaded_Tanks_train | 20 | 0.510 | 4.086 | 7.565 |
| Cascaded_Tanks_train | 5 | 3.015 | 4.562 | 12.589 |
| EMPS_test | 20 | 2.792 | 5.034 | 17.993 |
| EMPS_test | 5 | 3.736 | 5.555 | 19.193 |
| EMPS_train | 20 | 2.717 | 5.254 | 18.856 |
| EMPS_train | 5 | 3.718 | 5.558 | 21.031 |
| Silverbox_test1 | 20 | 0.000 | 0.000 | 1.388 |
| Silverbox_test1 | 5 | 0.753 | 3.037 | 4.644 |
| Silverbox_test2 | 20 | 0.000 | 0.000 | 2.738 |
| Silverbox_test2 | 5 | 0.000 | 2.402 | 4.514 |
| Silverbox_test3 | 20 | 0.000 | 0.000 | 1.139 |
| Silverbox_test3 | 5 | 0.000 | 2.716 | 4.288 |
| Silverbox_train | 20 | 0.000 | 0.000 | 1.562 |
| Silverbox_train | 5 | 0.489 | 3.098 | 5.083 |
| WienerHammer_test | 20 | 0.000 | 4.475 | 6.729 |
| WienerHammer_test | 5 | 2.222 | 4.821 | 8.573 |
| WienerHammer_train | 20 | 0.000 | 3.884 | 7.145 |
| WienerHammer_train | 5 | 2.662 | 5.335 | 9.292 |
| BoucWen_test1 | 20 | 0.000 | 5.395 | 8.207 |
| BoucWen_test1 | 5 | 0.208 | 4.890 | 14.161 |
| BoucWen_train | 20 | 0.000 | 1.305 | 2.923 |
| BoucWen_train | 5 | 0.409 | 3.678 | 7.155 |
| nile | 20 | 0.000 | N/A | 0.000 |
| nile | 5 | 0.000 | N/A | 0.000 |
| sunspots | 20 | 0.000 | 0.000 | 0.000 |
| sunspots | 5 | 0.000 | 0.000 | 0.000 |
| EURUSD_logret | 20 | 0.000 | 0.000 | 0.000 |
| EURUSD_logret | 5 | 0.000 | 0.387 | 0.181 |
| GSPC_logret | 20 | 0.000 | 0.000 | 0.000 |
| GSPC_logret | 5 | 0.000 | 1.063 | 0.801 |
In this published suite, the source labels the zero-gain cases as Identity abstentions. In general, numerical 0 dB does not imply unchanged output; use the explicit status. The two 100-sample Nile records are N/A for A2 and are not counted as passes. Dataset identifiers are shown in the rows above.
2 · Model compiler
Free-run results on public measured-system cases
Each case comes from a public nonlinear system-identification benchmark suite and is evaluated in free-run on a held-out tail. The NRMSE ratio is compiler ÷ the declared stable classical baseline on the same dataset; below 1 means the recorded compiler result is lower. The evidence package identifies which values were reproduced directly from current artifacts and which remain hash-bound to accepted source records.
| Benchmark case | Compiler — free-run NRMSE | Best classical baseline — NRMSE · method | Ratio (ours ÷ baseline) | |
|---|---|---|---|---|
| PUB-001 | 0.0204 | 0.0547 · lightweight nonlinear state-space | 0.37× | WIN |
| PUB-002 | 0.0253 | 0.0959 · lightweight nonlinear state-space | 0.26× | WIN |
| PUB-004 | 0.2168 | 0.3174 · EDMD/DMD | 0.68× | WIN |
| PUB-005 | 0.0561 | 0.5577 · EDMD/DMD | 0.10× | WIN |
| PUB-006 | 0.2583 | 0.2838 · RLS/Kalman tracker | 0.91× | WIN |
| PUB-008 | 0.1103 | 0.1794 · EDMD/DMD | 0.61× | WIN |
PUB-001…PUB-008 are case identifiers from the bound NDC benchmark pack. Directly reproduced and carried-forward values are distinguished in the evidence package; this table does not imply that every case was recompiled in the current release.
3 · Control
Advisory control against a frozen linear baseline
The controller is evaluated with a commissioned model, declared state information and command constraints. The frozen LQR comparator uses the linearized version of the same model and information boundary.
| Controller | Mean normalized score ↓ | Release interpretation |
|---|---|---|
| Automatic B1 full-model SDRE | 0.434471 | Deployable causal control policy |
| Future-selected frozen linear LQR | 0.761707 | Diagnostic lower-information model family |
The exact-linear parity test agrees at floating-point precision. The nonlinear result comes from recomputing the state-dependent control law from the commissioned model. These simulations are algorithm evidence, not a guarantee of physical-system safety or site-specific performance.
4 · Wide systems and MIMO
Up to 4,096 channels through deterministic sharding
A physical stream supports up to 264 synchronized response channels. A logical stream group partitions up to 4,096 channels into deterministic shards, fans ingest out, checks tick alignment and merges filtered rows back in the original column order.
| Capability | Release contract |
|---|---|
| Filtering and analytics | Group-wide read/write surface, up to 4,096 channels |
| Cross-channel model | Within each physical shard |
| Control | Shard-local; cross-shard coordinated control withheld |
Wide stream groups preserve filtering and analytics across deterministic shards. Control remains local to one physical shard; the service does not combine independent shard commands into a coordinated wide-system controller.
Filtering, NDC and Control use different protocols and metrics. Results are evidence for those protocols, not universal performance promises. Live use must remain within the product and artifact limits shown for the customer configuration.
Versioned evidence, not universal claims
Filtering evidence is tied to the published waveform/noise protocol. NDC evidence uses free-run NRMSE and a separate benchmark pack. Control evidence is model-conditional simulation with a fixed information and actuator contract. Production activation adds tenant isolation, live PostgreSQL, managed billing and identity, installation-specific envelopes and, for autonomous actuation, HIL acceptance.