The Scenario Trilogy: ODD, OpenSCENARIO 2.0, and the Economics of Sampling
Automotive testing starts from a despair-inducing premise: the world is infinite. Weather × lighting × road × traffic participants × … — any finite test budget thrown into this combinatorial space disappears without a sound. The industry’s answer is not to muscle through, but three moves in sequence: first ODD carves a boundary into the infinite world, then OpenSCENARIO 2.0 gives scenarios a declarative language, and finally the economics of sampling decides which few points the limited budget buys. This article strings the three into one line: boundary → description → selection.
1. ODD: The First Cut into an Infinite World
ODD (Operational Design Domain) = the set of conditions under which the system is designed to operate: road type, speed range, weather, lighting, geographic region, types of traffic participants… one boundary per dimension. The concept comes from SAE J3016 and is the foundation of L3+ automated driving: the system only commits to working inside its ODD; once outside, it must degrade or request a handover (MRM, Minimal Risk Maneuver).
A concrete example: one highway NOA system’s ODD = structured highways + 0–130 km/h + light rain or less + illuminated (daytime, or nights with street lighting).
Without an ODD, the Test Scope Is Infinite
The fundamental difficulty of ADAS validation is the “mileage disaster”: proving statistically that the system is safer than a human driver takes billions of test kilometers — real-world scenario combinations are infinite, and any fixed budget scattered into them is effectively zero.
The ODD is the first cut that turns an infinite problem into a bounded one: commit to safety only inside the ODD, verify behavior only inside the ODD. The test scope shrinks from “the whole world” to “a finite-dimensional parameter space” — and that is something you can plan, measure, and sign off.
Four Mappings from ODD to Test Scope
- Inside the ODD = a positive list. Each ODD dimension becomes a coordinate axis of the scenario space, and the three scenario levels (functional → logical → concrete) sample that bounded space for coverage. The meaning of coverage changes: not “how much of the world is covered” but “how much of the ODD is covered” — and that is computable (the grid fill rate).
- The ODD boundary = the highest-weighted test region. The boundary is the most dangerous place: the system must recognize “I’m about to leave my ODD” right at the edge. So the boundary gets two-sided testing — the inner edge (say 129 km/h) must work normally; the outer edge (131 km/h) must be correctly detected, with degradation and a handover request. Uniform sampling is waste; densifying the boundary is the right move — the same intuition as boundary-value analysis in unit testing, just with more dimensions.
- Outside the ODD = not “no testing,” but “test something else.” Behavior outside the ODD is not required to be “functionally correct,” but it is required to “fail safely.” So out-of-bounds scenarios stay in scope, testing boundary detection: does the system know it’s outside its ODD? And how fast does it know? This is the core concern of SOTIF (ISO 21448, Safety of the Intended Functionality) — ISO 26262 covers “system faults,” SOTIF covers “no fault, but insufficient performance” (say, perception rates dropping to unusable against backlight glare), and the foundation of a SOTIF argument is ODD + scenario coverage.
- An ODD change = a trigger for re-estimating test scope. Product iteration is ODD expansion (highway first, then urban; clear weather first, then light rain). Every expansion fires the same chain: new dimensions/boundaries → scenario-space growth → case-gap analysis → test-scope change — and that chain goes through the ALM tool’s impact-analysis workflow. The ODD is declared at the requirement layer and annotated at the case layer, with the traceability chain connecting the two.
Landing It: ODD in Case Management
The minimal engineering implementation is an odd field in the case manifest, with each case annotating the ODD coordinates it covers:
- id: TC-AEB-0042
req: [REQ-AEBS-003]
odd: {road: highway, speed: 50kph, weather: clear, light: day}
“ODD coverage” then becomes a report you can generate: fill every tested case’s parameters into the ODD grid, and the empty cells are the test gaps — the minimal implementation of scenario-coverage measurement. Note that the report’s real purpose is not “fill every cell” (with many parameters, the cell count is astronomical) but ranking the empty cells by risk — which cells the limited budget fills first is exactly the third topic of this article.
2. OpenSCENARIO 2.0: A Declarative Language for Scenarios
Boundary established — but how do you write the scenarios? OSC2 = ASAM OpenSCENARIO 2.0, version 2.0 of the scenario-description standard, officially released in 2022. Its origin is worth noting: the Israeli company Foretellix (a team from the semiconductor-verification industry) open-sourced its scenario language M-SDL (Measurable Scenario Description Language) and donated it to ASAM as the basis for 2.0; once the standard shipped, Foretellix retired M-SDL and folded everything into OSC2.
First, clear up a common misunderstanding: 1.x and 2.0 are not an upgrade-and-replace pair. The two are developed in parallel (with a merger planned for the future). 1.x is XML, imperative, aimed at “concrete scenarios”; 2.0 is a declarative language built from scratch.
The Essential Differences from 1.x
| OpenSCENARIO 1.x | OpenSCENARIO 2.0 | |
|---|---|---|
| Form | XML | Textual DSL (python-like syntax) |
| Paradigm | Imperative: “how to execute” | Declarative: “what should happen” |
| Levels | Concrete scenarios (fixed parameter values) | Abstract / logical / concrete — all three |
| Parameter space | Basically unsupported | Native: ranges + constraints + coverage goals |
| Ecosystem | Large install base (esmini, CARLA) | Early days; Foretify is the first native platform |
The “all three levels” row deserves expansion: the three scenario levels (functional → logical → concrete) become language constructs in OSC2 — one file can be refined all the way from a functional scenario down to concrete ones, while 1.x can only hold the “concrete” level.
Three Keywords
OSC2’s key design is just three keywords — see the snowstorm example from the official concept document:
scenario env.snowstorm:
storm_data: storm_data with:
keep(it.storm == snow) # hard constraint: must be snow
keep(soft it.wind_velocity >= 30kph) # soft constraint: satisfy if possible
cover(it.wind_velocity, unit: kph) # coverage goal: fill this dimension
keep= hard constraint, carving out the legal parameter space;keep(soft)= soft constraint, expressing a preference the solver tries to satisfy;cover= coverage goal — telling the tool “fill this dimension systematically.”
Concrete scenarios are no longer hand-written; a constraint solver generates them against the cover goals.
The Killer Feature: Dual Interpretation
The same scenario description can both drive a simulation (make the scenario happen) and serve as a monitor (judge whether the scenario really happened). The value for testing is direct: stimulus and checking come from one declaration, and the class of mistakes from “writing it twice, with the two sides disagreeing” simply disappears.
Its Place in the ASAM OpenX Family
- OpenDRIVE: the static road network (the map);
- OpenCRG: road-surface detail (potholes, friction);
- OpenSCENARIO: dynamic content (the behavior of traffic participants) ← this article;
- OSI: the sensor interface (the format of simulator → perception-algorithm input).
What OSC2 means for a test platform mirrors the environment abstraction of XiL: the scenario description is decoupled from the simulator — CARLA today, Prescan tomorrow, and the scenario files don’t change.
3. The Economics of Sampling: Why Full Factorial Is Not the Answer
The boundary is set (ODD), the language exists (OSC2) — one question remains: when a logical scenario is expanded into concrete ones, how do you pick the points? The easiest answer is full factorial — take every value of every parameter, pair them all up, run them all. A one-sentence review:
Full factorial is the theoretical answer — “perfect coverage, infinite budget.” The real engineering problem is “allocate a finite budget by risk density,” and full factorial happens to allocate budget by spatial volume. The direction is simply wrong.
Do the Explosion Math First
Take a cut-in scenario. A logical scenario’s typical parameters: ego speed, target speed, cut-in distance, cut-in duration, lane width, curvature, weather, lighting, road friction, target type… 8–15 parameters is perfectly normal.
Take just 10 values per dimension: 10¹⁰ = ten billion concrete scenarios. Even at SIL speeds (100 accelerated scenarios per second) that’s over 3 years; on HIL or a real vehicle (minutes per scenario and up), it’s physically impossible. That is the curse of dimensionality.
You can compute the inflation curve yourself: two parameters (ego speed in 6 steps × cut-in distance in 9 steps) = 54 scenarios, SIL finishes that in no time; stack on 3 more parameters at 10 values each = 54,000, about 9 minutes; stack on 2 more = 5.4 million, a 15-hour run; fill out 10 parameters at 10 values each and you’re back to the ten billion, the three years. The knee of the curve is right in front of you, not in theory.
Deadlier Still: Finishing It Buys You Nothing
Suppose you really did have ten billion scenario-hours. Full factorial is still the wrong investment:
- Most combinations are invalid: physically impossible (130 km/h + high curvature + icy friction at once), semantically contradictory (cut-in distance 50 m but duration 0.5 s), or outside the ODD (points outside the domain shouldn’t be tested for functional correctness). Full factorial knows none of this and runs them anyway;
- Dangerous scenarios are sparse: the “problem-triggering” region of parameter space is tiny — it usually hugs boundary combinations. Full factorial spreads budget uniformly by spatial volume, so 99.99% burns on scenarios “destined to pass,” and the information yield per unit cost approaches zero;
- Coverage is an illusion: running ten billion points gets you a “spatial coverage” number, but regulators and auditors want a “risk argument” — full factorial hands you a number, not an argument.
The Industry’s Replacement Strategies (Four Families)
- Rational reduction: the “smart subset” of full factorial. Boundary-value analysis + equivalence classes: instead of 10 values per parameter, take “the boundary ± one representative value.” Pairwise combinatorial testing: don’t require all combinations — only that every pair of parameters’ value combination appears once. NIST research (Kuhn et al.) shows the vast majority of faults are triggered by interactions of 1–2 parameters, so coverage size drops from exponential to logarithmic while the interaction coverage you actually get barely suffers.
- Risk-weighted sampling: by density, not volume. Sample by real-world distribution (Naturalistic Driving Data, NDD): common scenarios get sampled more. Weight by risk: rare-but-deadly ones (a pedestrian darting out from occlusion, a lead car’s emergency brake) get densified instead. Regulatory test matrices are a ready-made example — Euro NCAP’s CCRs speed points (a set of speed points chosen by risk) are essentially a risk-weighted selection an authority made for you; borrowing it directly is free-riding on their sampling strategy.
- Guided search: let simulation results command the sampling. Falsification: treat “find dangerous scenarios” as an optimization problem (genetic algorithms, Bayesian optimization) whose objective is “how far from failure,” drilling toward the almost-failing direction. Coverage guidance (the fuzzing idea): monitor the internal state coverage of the system under test and steer parameters toward unexplored state space.
- Real-data replay. Mine “combinations that really happened” from naturalistic driving data, crash databases, and shadow mode — no advantage in probability, unbeatable in relevance.
When Full Factorial Is Legitimate
Full factorial is not always wrong: with few parameters (2–3), few and discrete values, and cheap execution (SIL), it is exactly the right strategy — think configuration-switch matrices or mode combinations. The test in one sentence: you can count the total, you can run the total, and every combination means something — missing any one of the three means switch strategies.
4. The Trilogy in Concert: Boundary → Description → Selection
Each of the three owns one segment; together they form a pipeline:
the real world (infinitely many scenario combinations)
│ ODD: commit to safety inside the domain only, densify both sides of the boundary
▼
a bounded parameter space (gridded ODD coordinates)
│ OSC2: keep carves the legal space, cover declares fill goals
▼
logical scenarios (parameter ranges + constraints, all three levels in one language)
│ sampling: the solver picks points by cover goals and risk density — full factorial never appears
▼
a concrete scenario set that finishes — and can be signed off
A few meshing interfaces deserve to be named separately:
- OSC2 turns sampling economics into grammar:
coverdeclares the coverage goals, the solver samples against them — full factorial does not even exist at the language level; - ODD is the common input of both: OSC2 constraints can be generated under ODD guidance, with the tool automatically biasing toward ODD boundaries and coverage gaps; the sampling strategy’s “cell ranking” likewise comes from the ODD report;
- The three scenario levels are the trilogy’s spine: functional scenario (natural language) → logical scenario (OSC2’s expression layer) → concrete scenario (the points the solver picked) — one scenario asset, refined from abstract to concrete.
Looking back at the chain: ODD answers “where to test” — cutting the infinite world into a finite grid; OSC2 answers “how to write it” — declaring with keep/cover “what should happen” rather than “how to execute”; sampling answers “how to choose” — with a finite budget, divide the money by risk density, not by spatial volume. The order of the three cuts cannot be reversed: without a boundary, description has no domain; without description, sampling has nothing to operate on; without sampling, the first two are just good ideas on paper. Boundary, description, selection — that is the entire main line of scenario-testing methodology.
Appendix: Glossary (in order of appearance)
| Term | Plain explanation |
|---|---|
| ODD (Operational Design Domain) | The set of conditions (road, speed, weather, lighting…) under which the system commits to working — one boundary per dimension |
| SAE J3016 | The driving-automation taxonomy standard (L0–L5); the source of the ODD concept |
| MRM (Minimal Risk Maneuver) | The fallback action when the system leaves its ODD or fails — degrade, pull over, request handover |
| NOA | Navigate on Autopilot-style systems: advanced assistance that follows a navigation route, including lane changes and ramps |
| Mileage disaster | The dilemma that statistically proving an automated system safer than a human needs billions of test kilometers — impossible to drive in reality |
| Three scenario levels | The layering of scenarios: functional (natural language) → logical (parameter ranges) → concrete (fixed parameter values) |
| Boundary-value analysis / equivalence classes | Classic test-design techniques: take values near boundaries plus a few representatives, don’t tile the whole domain |
| SOTIF / ISO 21448 | Safety of the Intended Functionality: covers risks from “no fault, but insufficient performance” |
| ISO 26262 | The automotive functional-safety standard: covers risks from system faults |
| ALM | Application Lifecycle Management tooling; owns requirement–case–result traceability |
| Manifest | A case-metadata file kept separate from code; case tags and ODD coordinates live in it |
| ASAM | The Association for Standardization of Automation and Measuring Systems; maintains the OpenX simulation-standard family |
| OpenSCENARIO 2.0 / OSC2 | ASAM’s scenario-description standard, 2.0: a declarative DSL released in 2022; 1.x (XML) exists in parallel |
| M-SDL | Foretellix’s Measurable Scenario Description Language, open-sourced and donated to ASAM as the basis of OSC2 |
| Foretellix | The Israeli scenario-verification company; the team comes from the semiconductor-verification industry |
| DSL (domain-specific language) | A small language tailored to one domain, as opposed to a general-purpose programming language |
| Imperative / declarative | Imperative describes “how to execute step by step”; declarative only states “what should happen” and leaves execution to the tool |
| esmini / CARLA | An open-source scenario player / an open-source autonomous-driving simulator, both supporting OpenSCENARIO |
| Foretify | Foretellix’s verification platform, the first commercial tool with native OSC2 support |
| keep / keep(soft) / cover | OSC2’s three keywords: hard constraint, soft constraint, coverage goal |
| Constraint solver | The engine that automatically picks parameter values inside the constraint-bounded legal space; concrete scenarios are generated by it |
| Dual interpretation | One scenario description can both drive a simulation and monitor whether the scenario really happened |
| OpenDRIVE / OpenCRG / OSI | ASAM OpenX family members: static road network / road-surface detail / sensor interface |
| Prescan | A Siemens autonomous-driving simulator |
| SIL / HIL | Software-in-the-Loop / Hardware-in-the-Loop: the former runs pure software, cheap and fast; the latter connects a real ECU to a rig, expensive and slow |
| Curse of dimensionality | The phenomenon where a few more parameter dimensions explode the full-factorial count past anything physically runnable |
| Pairwise testing | A sampling method that only guarantees every 2-parameter value combination appears once, cutting size from exponential to logarithmic |
| NIST | The US National Institute of Standards and Technology; Kuhn et al.’s combinatorial-testing research is pairwise’s theoretical basis |
| NDD (Naturalistic Driving Data) | Real-world driving data collected by production fleets in everyday driving |
| Pedestrian dart-out (“ghost probe”) | The high-risk scenario of a pedestrian or cyclist suddenly emerging from occlusion |
| Euro NCAP / CCRs | The European new-car safety assessment; CCRs is its car-to-car rear-stationary scenario, with speed points chosen by risk |
| Falsification | Formalizing “find dangerous scenarios” as an optimization problem: the objective is “how far from failure,” and the search drills toward near-failure |
| Bayesian optimization | A global-optimization method that uses a surrogate model to steer the search; suited to expensive objectives like simulation |
| Coverage guidance / fuzzing | Borrowed from software fuzzing: monitor internal state coverage of the system under test and steer inputs toward new state space |
| Shadow mode | A new feature runs silently on production vehicles — logging without actuating — to mine scenarios that really happened |
Next step:
View all notes