CYCLONE SPECIFICATION DOC NO. CY-SPEC-001 · REV 2026-09-12

Same case hash.
Same test result. Every time.

The deterministic XiL scenario testing kernel

ABSTRACT Cyclone turns ADAS / robotics regression testing into one pipeline: replay → fault injection → statistical verdict → evidence chain. A PASS is no longer "it went green once" — it is reproducible, auditable machine proof.
FIG. 0-1 · Live run (recomputable)
$ cyclone run cases/aeb_cipv_50kph.yaml --out out/ verdict: PASS $ shasum -a 256 out/AEB_CIPV_50kph_log.jsonl 2e075b42f19031aa… $ cyclone run cases/aeb_cipv_50kph.yaml --out out2/ # run it again $ shasum -a 256 out2/AEB_CIPV_50kph_log.jsonl 2e075b42f19031aa… # same seed → bitwise identical
Fig. 0-1 — Two runs of the same case, bitwise-identical SHA-256. Machine-checkable.

261 tests green · bitwise reproducible · Wilson · SPRT verdicts · rules_cyclone prototype · single binary planned

Book a POC demo → App.A · Get the scenario pack

§01 · PROBLEM STATEMENT

Problem statement: old problems, new answers

§1.1 A failure: product bug or test noise? You can't tell. The deterministic kernel makes "same input, different result" structurally impossible. Red means red.
§1.2 The report says PASS, with no machine-verifiable evidence. Every case emits a case hash + SHA-256 digest; bitwise-identical reruns are machine-checkable proof.
§1.3 CI integration costs a week of glue code. Native JUnit / Bazel / BES output — not "integratable", a first-class citizen.
§1.4 Why is the acceptance line 0.98 and not 0.95? One question in review, and the room goes silent. Wilson lower bounds replace surface pass rates, and thresholds are derived by three-step margin analysis — a threshold with a derivation is a gate; without one, decoration.

§02 · SYSTEM CHARACTERISTICS

System characteristics

§2.1 Deterministic Replay Supported

Virtual-time replay with content addressing: identical seeds produce bitwise-identical evidence logs.

§2.2 Statistical Verdicts Supported

Wilson score intervals + SPRT sequential testing — statistical significance instead of hand-tuned thresholds.

§2.3 Evidence Chain Supported

The four-piece set: case hash + JSONL SHA-256 digest + failure snapshot + JUnit XML — every CI record traces back to the case version and raw data.

§2.4 CI-Native

JUnit XML out of the box Supported; rules_cyclone makes scenario tests first-class bazel test citizens Prototype; BES-ready Prototype.

Full characteristics, architecture & integrations → Deep dives: Deterministic fault injection · ADAS scenario regression testing · Auditable test evidence chain

APPENDIX A · SCENARIO PACK

Appendix A · Scenario pack

22 ADAS scenario YAMLs, every verdict matching expectations. One YAML is one regression case — tweak parameters to derive new scenarios.

Active safety ×13 Driving assist ×5 Parking ×1 Shared driving ×2 L3 admission ×1

The catalog is public; the full pack of 22 YAMLs is sent on request:

Email us for the full pack

All in-scenario figures are demo data, not measured on a real SUT; thresholds must be calibrated against your SUT baseline (a margin-analysis service item).

How scenarios become CI regression: ADAS scenario regression testing →

APPENDIX C · TECHNICAL NOTES

Appendix C · Technical notes

Statistical verdicts, architecture notes, and cross-industry methodology. For official sources of testing standards, see the standards directory.

TN-22 · 2026-09-12 Engineering Infrastructure: Abstraction Layers, ALM, and Benchmarks The abstraction layer decides whether SIL investments appreciate or evaporate in the HIL era; ALM keeps bidirectional traceability; benchmarks answer 'how fast'. TN-21 · 2026-09-12 Three Views of Functional Safety: ASIL, the Safety MCU, and NCAP ASIL sets how strict the rules are, the safety MCU backstops the domain controller, and NCAP translates 'safety' into scenarios the market understands. TN-20 · 2026-09-12 Programmable Fault Injection: Turning Abnormal Inputs into Test Assets Fault injection grows from a switchboard-operator craft into versioned, regression-ready test assets. TN-19 · 2026-09-12 The Scenario Trilogy: ODD, OpenSCENARIO 2.0, and the Economics of Sampling ODD carves the test boundary, OpenSCENARIO 2.0 describes scenarios declaratively, and sampling economics answers the combinatorial explosion. TN-18 · 2026-09-12 Talking and Listening on the Bus: RX/TX, Restbus Simulation, and XCP RX/TX direction is always relative to someone; restbus simulation completes the environment's other half, XCP opens observability depth. TN-17 · 2026-09-12 HIL Real-Time: A Missed Cycle Isn't Slow, It's Wrong HIL 'real-time' isn't fast — it's on time, every cycle. FPGAs wire logic in code to guarantee microsecond determinism. TN-16 · 2026-09-12 PIL (Processor-in-the-Loop): Why Passing SIL Isn't Enough Four walls stand between SIL and the target: compiler, numerics, resources, timing. Back-to-back comparison is PIL's core method. TN-15 · 2026-09-12 One AEB Feature from Model to Production: The V-Model's Six Levels and the TTC Metric Walk one AEB feature through the V-model's six test environments, then examine TTC, its classic trigger metric, and its five flaws. TN-14 · 2026-09-12 Classic vs Adaptive AUTOSAR: Not a Succession, a Division of Labor Classic AUTOSAR owns body/chassis determinism; Adaptive owns compute-heavy domain controllers — not a succession, a division of labor. TN-13 · 2026-09-12 The Automotive Test Toolchain: CAPL and ECU-TEST CANoe's embedded CAPL and ECU-TEST, the de-facto test automation layer — two industry playbooks: platform-embedded scripting vs a standalone automation layer. TN-12 · 2026-09-12 The Exit of Test Results: JUnit XML and the Verdict Engine How JUnit XML became the de-facto test report standard, and how a read-only verdict engine advances expectations through a state machine. TN-11 · 2026-09-12 Verification on a Budget: Vertical Slices, Test Doubles, and Stimulus Injection Vertical slices test the riskiest assumption first; stubs and fixtures replace dependencies; stimulus injection feeds the inputs — three ways to buy confidence cheaply. TN-10 · 2026-09-12 Flaky Tests: Why They're Poison and How to Catch Them A test that passes and fails at random isn't just annoying — it zeroes out the credibility of the whole suite. TN-09 · 2026-09-12 Why Nightly Full Regression Is a Luxury in the HIL Era Nightly full regression was a given in software testing; once the test environment is a million-dollar HIL rig, it becomes an expensive luxury. TN-08 · 2026-09-12 Test Tiering: The Cost Logic Behind Smoke, Regression, and Full Suites Smoke, regression, and full suites aren't three test sets — they're three cost-budgeted ways to run one case library. TN-07 · 2026-09-10 Determinism, Fault Injection, and the Evidence Chain: One "Exam Philosophy" Shared by Three Industries Chips, AI agents, and automotive electronics: the SUT is too complex to exhaust and failure too costly to improvise — so all three grew the same methodology. TN-06 · 2026-09-09 The XiL Ladder, Side by Side: Automotive and Silicon Are Drawing the Same Picture Automotive says SiL/HiL; silicon says ISS/cycle-approx — on paper, they are the same ladder from fast-and-fake to slow-and-real. TN-05 · 2026-09-08 Dual-Mode Clock: How to Have Both Offline Reproducibility and Online Realism One engine, two physics: the dual-mode clock gives you both offline reproducibility and online realism — virtual clock for replay, real clock for the field. TN-04 · 2026-09-07 Margin Analysis in Three Steps: Thresholds Are Derived, Not Picked Why is the acceptance line at 0.98? Baseline, margin, two-way validation — turn thresholds from decoration into evidence. TN-03 · 2026-09-06 SPRT and LLR: Sequential Testing That Delivers the Verdict as Soon as the Evidence Is In How many runs are enough? SPRT gives the verdict a scoreboard: stop as soon as the evidence is in, cutting execution volume by 30–50% on average. TN-02 · 2026-09-05 Wilson Lower-Bound Verdicts: An Honest Ruler for Pass Rates Wald, Clopper-Pearson, Wilson: three rulers, three floor scores. Why acceptance promises must stand on the lower bound, with a cheat table and 5-line code. TN-01 · 2026-09-04 30 Runs, 27 Passes: Why That Doesn't Count as Verified A 90% pass rate meets the 90% acceptance line, so why can't you sign off? Coin flips, a spoonful of soup, and one sample-size table explain point estimates, confidence intervals, and floor scores.

DOCUMENT CONTROL

Run it on your own scenario

A 30-minute demo: AEB regression → inject a fault → open the evidence report.

Book a 30-min demo Get a quote Product sheet PDF Roadmap PDF