The Exit of Test Results: JUnit XML and the Verdict Engine

A test run finishes — hundreds of cases, tens of thousands of simulation ticks — and two questions are waiting: in what format do the results leave, and who decided pass or fail. This article puts the two halves of the exit together: the first half is about the report format — JUnit XML, which has no specification and no schema yet locks down the entire ecosystem; the second half is about the verdict engine — a read-only “referee” that looks at the signal snapshot once per tick and advances every expectation through a state machine.

The two meet at the end: the verdict engine’s summary output is exactly those <testcase> elements in the JUnit XML.

1. JUnit XML Is Not a “Standard” at All

JUnit XML has no specification, no schema, and no organization that owns it. It is simply the TEST-*.xml file format that Apache Ant’s <junit> task spat out in the early 2000s. Later, Hudson/Jenkins parsed it natively to draw trend graphs — and tools worldwide discovered: if I emit this XML too, I get Jenkins’ report UI, trend curves, and failure emails for free. The format is so simple that any language can write an emitter for it in a day.

2. The Positive-Feedback Network Effect (the Real Mechanism Behind Its Reign)

CI servers support it (Jenkins / GitLab CI / Azure DevOps / CircleCI)
        ↑↓
Test frameworks emit it (pytest --junitxml, gtest --gtest_output=xml, Catch2 -r junit...)
        ↑↓
Downstream tools consume it (report aggregation, ALM reporting, flake detection)

Three sides reinforce each other into a moat: replacing it means replacing the whole toolchain at once, and no single point has an incentive to move first. That is the essential difference between a de-facto standard and a formal one (like ASAM XIL) — XIL was designed by a committee and pushed out to everyone; JUnit XML became a “standard” only after everyone was already using it.

3. The Format Design Happens to Sit on the Sweet Spot

Survival wasn’t pure luck; the structure really is “just enough”:

4. Why the Automotive Industry Adopted It Too

5. Limitations and Successors

The report format is settled. The next question: the “pass/fail” written into the report — who judged it?

6. Where the Verdict Engine Stands: The Very End of the Data Path

A typical verdict engine design looks like this: it reads and never writes; once per tick it looks at a snapshot of the signal world and advances the declared expectations through state machines, one by one. Broken down:

Bus frames → decode (FlatBuffers) → SignalStore (signal snapshot) ──▶ VerdictEngine.evaluate(tick)
                                        ▲                                  ▲
FaultPipe → FaultEvent counters ────────┘                          YAML assertion rules

Two input sources:

The engine itself is a pure consumer: it modifies no signal and sends no frame. This read-only constraint matters — a referee may never touch the match.

7. The Core Design: Every Assertion Is a Small State Machine

The easiest way to get a verdict engine wrong is “run the test, then scan the log and judge once.” The sound approach is incremental evaluation every tick — every assertion is instantiated before the simulation starts and advanced one step per tick:

// Condition: the assertion's "body", a closed variant set
struct EqSignal   { SignalId sig; double expected; };
struct InRange    { SignalId sig; double lo, hi; };
struct SeqAfter   { SignalId trig_sig; double trig_val;   // after A happens
                    SignalId cons_sig; double cons_val;   // B must happen within T
                    uint32_t window_ticks; };
struct FaultHits  { uint32_t rule_id; uint32_t min_hits; };
using Condition = std::variant<EqSignal, InRange, SeqAfter, FaultHits>;

// Assertion: condition + time window + state
struct Assertion {
    const char*  name;             // string pool
    Condition    cond;
    Tick         window_start;     // start listening at this tick
    Tick         deadline;         // "by 2500ms" → deadline tick
    State        state = State::Pending;   // Pending/Satisfied/Violated/TimedOut
    // failure-scene snapshot (for the report)
    double       last_actual = 0;
    Tick         last_change = 0;
};

class VerdictEngine {
    std::array<Assertion, kMaxAssertions> assertions_{};
    size_t count_ = 0;
public:
    void evaluate(Tick now, const SignalStore& sig, const FaultStats& faults);
    CaseResult summarize() const;   // roll up into the JUnit report
};

The state-transition rules (taking expect: brake_request == true, by: 2500ms as the example):

Pending ──condition true──────────────▶ Satisfied (latched; never judged again)
   │
   └── tick passes deadline, still false ──▶ TimedOut (record the scene: last_actual=0, last change time)

always-type assertions go the other way: the moment the condition turns false inside the window, they latch into Violated. All terminal states latch — an assertion is judged exactly once in its life, which guarantees that replaying the same log always yields the same verdict.

8. The Per-Tick Evaluation Loop

void VerdictEngine::evaluate(Tick now, const SignalStore& sig, const FaultStats& fs) {
    for (size_t i = 0; i < count_; ++i) {
        auto& a = assertions_[i];
        if (a.state != State::Pending) continue;      // terminal state: skip
        if (now < a.window_start) continue;           // not yet in the listening window

        bool ok = std::visit([&](const auto& c){ return check(c, sig, fs, a); }, a.cond);
        if (ok) {
            a.state = State::Satisfied;
        } else if (now >= a.deadline) {
            a.state = State::TimedOut;
            a.last_actual = readActual(a.cond, sig);  // capture the scene
            a.last_change = sig.lastChangeTick(signalOf(a.cond));
        }
    }
}

Note how three disciplines show up:

9. Mapping YAML to State Machines

verdict:
  - expect: brake_request == true      # → EqSignal
    by: 2500ms                          # → deadline = 2500 ticks
  - expect: vehicle_speed in [0, 5]     # → InRange
    by: 4000ms
  - expect: fault_injected(radar_dropout, min_hits: 200)  # → FaultHits
    by: 1500ms
  - expect: target_valid == false then dtc_reported == true within 500ms
                                        # → SeqAfter, a two-stage state machine

At load time (the one phase where heap allocation is allowed), every expect is compiled into an Assertion: signal name → SignalId, time string → tick count, action name → rule_id. Every check that can be done at compile time is done at load time (nonexistent signals, inverted time windows, undefined rule_ids), so the RT loop is left with nothing but integer comparisons.

SeqAfter deserves its own look — it is a two-stage state machine:

WaitTrigger ──trig condition true──▶ Watching (record trig_tick)
                                        │ cons true within trig+T ──▶ Satisfied
                                        └── still false past trig+T ──▶ TimedOut

With it, temporal-causal assertions like “after the fault occurs, the SUT must report a DTC within 500ms” become declarable — the single most common sentence pattern in functional-safety verification.

10. What to Capture on Failure: A Verdict Needs a “Scene”

A report that says only “expected brake_request==true, timed out” is useless. When a failure latches, capture along with it:

Scene information Source What it looks like in the report
Last actual signal value SignalStore brake_request stayed 0
Last change time of the signal SignalStore’s change-tick no change after t=380ms
Fault-injection state at that tick FaultStats radar_dropout active at the time
Last N values of related signals optional ring history attached to the failure details

11. Summary Output: The Mapping to JUnit

CaseResult VerdictEngine::summarize() const {
    // each assertion → one JUnit <testcase>, or one <testcase> for the whole case;
    // failed assertions → <failure message="...">, the message built from the scene snapshot with fmt;
    // fault-hit statistics → <properties>
}

A failure in the report looks roughly like this:

FAILED: aeb_radar_dropout / brake_request == true by 2500ms
  actual: 0 (unchanged since t=380ms)
  context: radar_dropout active [1000ms,1300ms), injected 240 times

Whoever reviews the report can locate the direction of the problem — did the SUT fail to decide, or did the frame never go out — without digging through bus logs.

12. A Free Capability: Offline Re-verdict

Because the engine consumes only SignalStore + FaultStats, and both can be rebuilt from recorded bus logs, the very same engine can re-judge a recorded log offline: change an assertion threshold and you don’t need to re-run the simulation — just feed it the log. This is also the dividend of the “execution plane / evidence plane separation” architecture — verdict logic is fully decoupled from stimulus execution.

13. Landing Priorities

Phase Content
P0 EqSignal + by deadline + JUnit output (covers 80% of cases)
P1 InRange, always window assertions, failure-scene snapshots
P2 FaultHits (plugging into the fault-injection evidence chain), SeqAfter timing assertions
P3 Offline re-verdict mode, ring buffers for signal history

One sentence to sum up: a verdict engine = a set of assertion state machines compiled at load time + one pure-function evaluation per tick + failure-scene latching + the JUnit mapping. It doesn’t aim to be clever; it aims for every verdict to be deterministic, reproducible, and auditable — which is exactly what the end of a safety evidence chain should look like.


Back to the two halves of the exit: JUnit XML answers “in what format do results leave” — it was crowned by the network effect and stays seated by a just-enough design; the verdict engine answers “who judged the results” — read-only, per-tick, terminal states latched, turning every verdict into reproducible, auditable evidence. One is an ecosystem problem, the other a design problem, but the two meet inside <testcase>: every pass/fail line in the report is the footprint left by a state machine that walked to its terminal state.


Appendix: Glossary (in order of appearance)

Term Plain explanation
JUnit XML The common test-result reporting format: born from Ant’s <junit> task, a de-facto standard with no specification
emitter The code module that produces and writes out a file format; here, the report writer that emits JUnit XML
CI Continuous Integration: the pipeline that builds and tests automatically on every commit
ALM Application Lifecycle Management tooling; owns requirement–case–result traceability
flake An unstable test that passes and fails without any code change
ASAM XIL A formal automotive simulation-testing standard: defines a generic interface for test scripts to access rigs, designed by a committee
the four states (pass/fail/error/skipped) JUnit XML’s result classification: assertion failed / environment errored / skipped because preconditions weren’t met
<properties> JUnit XML’s official pocket for custom key-value pairs — requirement IDs and ASIL levels ride here
ASIL Automotive Safety Integrity Level (QM/A/B/C/D, D strictest), a core ISO 26262 concept
ASAM ATX A standard format for exchanging test-case descriptions across the supply chain
bus trace A timestamped, complete record of every frame on the bus
CTRF Common Test Report Format: a next-generation, JSON-based universal test report format
verdict engine The “referee” of a test system: read-only over the signal snapshot, advancing expectations through state machines to reach pass/fail
tick The smallest time step of a simulation; at a 1kHz cadence, 1 tick = 1 ms
SignalStore The table of current signal values updated every tick — essentially a flat SignalId → double array
FaultEvent Event counters from the fault-injection module, backing assertions like fault_injected(...)
closed variant set C++’s std::variant: all types listed at compile time, visited without virtual functions
latching (terminal states) Once the state machine enters a terminal state it never changes — an assertion is judged exactly once in its life
zero-allocation The discipline of doing no heap allocation in the real-time loop, via fixed-size arrays and string pools
bitwise reproducible Re-running the same input yields bit-identical output; not touching wall-clock time is one prerequisite
SeqAfter A two-stage timing assertion: “after A happens, B must happen within T”
SUT System Under Test
DTC Diagnostic Trouble Code: the standardized code a controller records and reports when it detects a fault
failure-scene snapshot The context captured when a failure latches: last actual value, last change time, fault state, etc.
offline re-verdict Re-judging a recorded log with the same verdict engine — change thresholds without re-running the simulation

Next step:

View all notes