Programmable Fault Injection: Turning Abnormal Inputs into Test Assets

Ten thousand kilometers of normal driving without a single false brake from AEB proves nothing — everything functional safety truly cares about hides in the abnormal: radar frames dropped for 300 ms, the vehicle-speed signal frozen at a stale value, a CRC that suddenly doesn’t check out… These inputs are rare, dangerous, and nearly impossible to wait for naturally during testing. Fault injection is the means of manufacturing them on purpose, and “programmable” is what decides whether it stays a one-off craft or becomes a test asset that lives in git and runs in nightly regression. This article covers two things: why fault injection is worth making programmable, and what a programmable fault-injection engine looks like.

1. What Are “Abnormal Inputs”?

Abnormal inputs = the bad inputs an SUT may receive in the real world but can hardly reproduce naturally in testing. Classified by layer:

Layer Typical abnormality Real-world source
Data Wild values, out-of-range values, NaN, negative vehicle speed Sensor failure, calibration error
Communication Dropped frames, duplicated frames, reordering, bit flips, CRC errors Bus interference, gateway failure
Timing Delay, jitter, timeout, cycle drift Sender ECU overloaded, rebooting
Protocol Rolling counter not incrementing, wrong message length Software bug on the peer node
State Diagnostic session jumping around, SecurityAccess brute-forcing Misbehaving tester, attacks

What these inputs share: rare, dangerous, and exactly what functional safety cares about most. AEB behaving well on normal inputs is no achievement — whether it false-brakes when it receives “distance = 500 m, then 5 m three frames later” is what ISO 26262 wants verified.

2. What Does “Programmable” Mean?

Compare the two approaches.

The non-programmable (hard-coded) approach — fault logic baked into C++ code:

// Want to test frame drops? Edit code, recompile, edit it back afterwards
void onCanFrame(const CanFrame& f) {
    static int cnt = 0;
    if (f.id == 0x123 && ++cnt >= 100 && cnt <= 102) return; // drop 3 frames
    forward(f);
}

The problems are obvious: every new fault scenario means changing platform code and recompiling — like an old telephone switchboard operator crawling back into the code to re-plug the wires; fault logic gets tangled up with business logic; and no test asset ever accumulates — the next project has to dig through git history for that if.

The programmable approach — faults are declared rules, and the platform provides an engine to interpret them:

# Declared in the test case; not a single line of C++
faults:
  - name: radar_track_dropout
    at: 1000ms          # when to inject
    duration: 300ms     # how long it lasts
    channel: can0
    match: { id: "0x2A0..0x2AF" }    # which frames it applies to
    action: drop                     # action: discard

  - name: speed_freeze
    at: 2000ms
    duration: 500ms
    channel: can0
    match: { id: "0x123" }
    action: freeze                   # freeze the vehicle_speed field — manufacture "stale data"

On the platform side there is just one generic fault-execution engine: match rules → trigger at the right moment → apply the action to the data stream. A test engineer adds a new scenario = writes YAML, without touching platform code. That is “programmable” — fault behavior becomes data, not code. (The example is simplified — the full syntax for parameterized actions like field-level freezing is in Section 7.)

3. Why It’s Worth Making Programmable

  1. Fault scenarios become test assets: they can go into git, be reviewed, be reused, and join the nightly regression suite — just like normal stimulus, all orchestrated in YAML;
  2. Aligned with the functional-safety evidence chain: ISO 26262 Part 6’s software-testing method tables list fault-injection testing as recommended for ASIL C/D — to verify that safety mechanisms actually work, e.g. whether the SUT enters its safe state after the E2E protection detects a CRC error. Every YAML fault rule + verdict result is a traceable piece of safety-verification evidence;
  3. It composes with normal stimulus into complex scenarios: normal stimulus drives the SUT toward a dangerous condition (target approaching) while an anomaly is injected at the critical frame (drop exactly 3 frames right now) — “what if the SUT goes blind for 300 ms at the very moment of decision?” Only programmability makes such combinations possible.

One sentence: abnormal inputs are the SUT’s “stress exam questions,” and programmability lets those questions be declared, stored, reused, and replayed like ordinary test cases — fault injection turns from a one-off code hack into a first-class platform capability. Now let’s look at the concrete design of a programmable fault-injection engine.

4. Where It Lives: An Adapter Decorator

The core idea in one sentence: fault injection is not a new module — it is a “rule pipeline” (FaultPipe) added to the data path of the bus adapter (IBusAdapter) — rules loaded from a YAML DSL, driven by the Time Master’s tick, and written under real-time-path C++ discipline throughout.

Don’t let fault logic grow into the SocketCAN implementation. Wrap the existing adapter in a decorator so the injection point stays transparent to the business logic:

┌──────────────────┐   ┌─────────────────┐   ┌──────────────────┐   ┌─────┐
│ YAML rules       │──▶│ FaultPipe       │──▶│ IBusAdapter      │──▶│ SUT │
│ (parsed at load) │   │ (rule pipeline) │   │ (SocketCAN etc.) │   │     │
└──────────────────┘   └─────────────────┘   └──────────────────┘   └─────┘
                                ▲
                        Time Master tick

5. Core Data Structures: variant + Preallocation

The action set is closed, so use std::variant; the rule container is fixed-size, filled once at load time — zero heap allocation on the RT path:

// fault_rule.h — POD-ish structures only, no member that allocates
struct Drop         {};
struct Delay        { uint32_t delay_ticks; };
struct CorruptField { uint8_t byte_offset; uint8_t mask; uint8_t value; };  // simplified
struct Freeze       {};   // freeze the field at its previous value
struct CorruptCrc   {};

using FaultAction = std::variant<Drop, Delay, CorruptField, Freeze, CorruptCrc>;

struct FaultRule {
    uint32_t    id;
    const char* name;          // points into a string pool allocated at load time, living for the whole run
    uint8_t     channel_mask;  // bit0 = can0, bit1 = can1...
    uint32_t    id_min, id_max;// frame-ID filter interval
    Tick        start_tick;    // the Time Master's tick count
    Tick        end_tick;
    FaultAction action;

    // Mutable run state (Freeze's last value, Delay's pending-queue pointer, etc.)
    // reset at the start of every run to guarantee reproducibility
    bool active(Tick now) const { return now >= start_tick && now < end_tick; }
};

class FaultPipe {
    std::array<FaultRule, kMaxRules> rules_{};   // kMaxRules = 64 or so
    size_t rule_count_ = 0;
    RingBuffer<PendingFrame, 256> delayed_;      // preallocated ring buffer for the Delay action
public:
    std::optional<CanFrame> process(const CanFrame& f, Tick now);
    void releaseDelayed(Tick now, FrameSink& out);  // release due delayed frames every tick
    void resetAll();                                 // clear state between runs
};

The key point: YAML parsing happens at load time (where heap allocation is harmless); afterwards the rules are copied into the fixed-size array and the strings into the string pool. Once the simulation loop starts, there is not a single new on this path.

The action library only needs the four or five most common ones first: drop, delay, corrupt_field (modify a value), freeze (hold the old value), corrupt_crc — these five cover 80% of fault-injection needs in functional-safety verification.

6. Runtime: A Three-Stage Pipeline on the Tick

A single-threaded 1 ms scheduler makes the processing order inside each tick strictly deterministic:

// Inside each tick, the scheduler executes in a fixed order:
void SimLoop::tick(Tick now) {
    faultPipe_.releaseDelayed(now, busSink_);   // 1. release due delayed frames first
    replayEngine_.emit(now, busSink_);          // 2. CSV normal stimulus (through the FaultPipe)
    adapters_.pollAll();                        // 3. receive real bus frames (through the FaultPipe)
    verdict_.evaluate(now);                     // 4. verdict
}

// FaultPipe::process is a pure matching pipeline, rule by rule in declaration order:
std::optional<CanFrame> FaultPipe::process(const CanFrame& f, Tick now) {
    for (size_t i = 0; i < rule_count_; ++i) {
        auto& r = rules_[i];
        if (!r.active(now) || !matches(r, f)) continue;
        logFaultEvent(r, f, now);              // key: every hit enters the evidence chain
        return apply(r, f, now);               // Drop → nullopt; Delay → into the ring buffer
    }
    return f;                                   // no hit — pass through unchanged
}

The “three stages” are the FaultPipe’s three hook points in each tick (steps 1–3: release delayed frames, emit replay stimulus, receive real bus frames — all passing through the FaultPipe); step 4, the verdict, runs after the pipeline and is not part of the data path.

Why does a fixed order matter? The same frame may hit both a “drop” rule and a “corrupt value” rule, and which one wins must be a hard rule (YAML declaration order) — otherwise two runs diverge and bitwise reproducibility is broken.

7. The YAML DSL: Faults and Normal Stimulus Live in the Same Case File

test_case: aeb_radar_dropout
duration: 5000ms

stimulus:                        # normal stimulus
  - replay: scenarios/aeb_approach.csv
    channel: can0

faults:                          # abnormal input, programmable
  - name: radar_dropout
    at: 1000ms
    duration: 300ms
    channel: can0
    match: { id: "0x2A0..0x2AF" }
    action: drop

  - name: speed_freeze
    at: 2000ms
    duration: 500ms
    channel: can0
    match: { id: "0x123" }
    action: freeze

verdict:                         # verdict as usual
  - expect: brake_request == true
    by: 2500ms
  - expect: dtc_reported == true   # 300 ms of dropped frames should trip the timeout diagnostic
    by: 2000ms

A test engineer adds a scenario = adds a few lines of YAML, with zero platform-code changes. That is the final form of “programmable.”

8. The Evidence Chain: Injection Events Themselves Must Be Observable

This is what many people miss — if the fault injection never happened, a passing verdict counts for nothing (say the rule’s ID interval was written wrong, no frame was ever dropped, and the test ran for nothing). So every rule hit emits a FaultEvent onto the bus log (FlatBuffers):

// Record: which tick, which rule, on which frame, what action
FaultEvent{ now, r.id, f.id, actionTag(r.action) }

The verdict engine then supports a new kind of assertion:

verdict:
  - expect: fault_injected(radar_dropout)          # confirm the fault really went in
  - expect: fault_hit_count(radar_dropout) >= 240  # ~240 frames should be dropped within 300 ms

The JUnit XML report writes the injection statistics into <properties>, so whoever reviews the report sees at a glance that “this case’s faults actually took effect.” This is exactly the evidence shape ISO 26262 fault-injection testing asks for.

9. The Fault Injector Itself Must Be Tested Too

The injector is part of a safety-verification tool and may one day be classified as a TCL3 tool — its erroneous output can directly contaminate safety evidence without being easily detected. So give it unit tests:

TEST_F(FaultPipeTest, DropRuleSuppressesMatchingFramesOnly) {
    // Fixture: fixed fake clock + 3 rules + a golden input sequence of 20 frames
    // Assert: the output sequence matches the golden file byte for byte
}

One fixture per action type, with golden frame sequences for both input and output — the injector is deterministic pure logic, inherently easy to test.

10. Suggested Rollout Order

Phase Content Rationale
P0 drop + delay actions + FaultEvent logging Get the structure running; covers the most common drop/timeout scenarios
P1 corrupt_field + freeze + corrupt_crc Data-layer anomalies, paired with E2E verification scenarios
P2 RX/TX bidirectional injection points + the verdict’s fault_injected assertion Close the evidence-chain loop
P3 Sequenced faults (a “fault train” like “drop 3 frames → 5 normal frames → drop 2 more”) Covers intermittent-fault scenarios

P0 can be up and running in two or three days: the tick, adapters, YAML, and log bus are all off-the-shelf infrastructure in a mature XiL platform — the FaultPipe is just a layer of pure logic stringing them together.


Looking back along the chain: “programmable abnormal input” = a fixed-size rule array + variant actions + a tick-driven matching pipeline + YAML declarations + FaultEvent evidence. Write the platform code once, and fault scenarios are data from then on — into git, into review, into nightly regression, on equal footing with normal stimulus.


Appendix: Glossary (in order of appearance)

Term Plain explanation
AEB Autonomous Emergency Braking; the recurring example SUT function in this article
Functional safety The engineering discipline of preventing personal harm from E/E system failures; the automotive standard is ISO 26262
Fault injection The test technique of deliberately manufacturing abnormal inputs for the system under test
SUT System Under Test — the object this test aims at
NaN Not a Number, the “not-a-number” value in floating point, often from illegal arithmetic or invalid sensor data
Calibration The work of assigning values to ECU parameters (thresholds, gains, etc.); calibration error is a typical source of data-layer anomalies
CRC Cyclic Redundancy Check: the checksum field in a message that detects corrupted data
Rolling counter A per-frame incrementing counter in the message, against replay attacks; a wrong sequence gets judged “data not trustworthy”
SecurityAccess The security-unlock service (0x27) in UDS diagnostics; “brute-forcing” means an attacker repeatedly guessing the key
ISO 26262 The automotive functional-safety standard; explicitly recommends fault-injection testing for ASIL C/D
YAML A human-readable data-description format; this article’s test cases and fault rules are all declared in it
ASIL Automotive Safety Integrity Level (QM/A/B/C/D, D the strictest)
E2E protection AUTOSAR’s end-to-end communication protection: rolling counter + CRC and friends, keeping messages trustworthy
Evidence chain The complete set of records that lets a third party recompute and verify a verdict
Verdict The pass/fail conclusion the test platform gives a test case
Decorator A design pattern that adds functionality by wrapping a layer outside, without changing the original interface
SocketCAN The Linux kernel’s CAN protocol stack and driver framework
Time Master The platform module that advances simulation time uniformly
Tick The smallest time step of the simulation scheduler; the scheduler here runs one tick per 1 ms
std::variant C++17’s type-safe union: a closed “one of these few types”
RT path (real-time path) A code path with hard constraints on latency and memory allocation
Heap allocation Requesting memory from the heap at runtime; normally forbidden on the RT path
DSL (domain-specific language) A small description language built for a specific domain; this set of YAML fault rules is the fault-injection DSL
Bitwise reproducible Two runs produce bitwise-identical output
FlatBuffers A zero-copy serialization library, used here as the encoding format of the bus log
JUnit XML The universal test-result reporting format, understood by CI and test-management tools
TCL3 One of ISO 26262’s tool confidence levels: the tool’s erroneous output can contaminate safety evidence without being easily detected
Golden file A pre-frozen reference-answer file; output is compared against it byte for byte
Fault train A scripted fault sequence alternating “fault/normal,” used to simulate intermittent faults

Next step:

View all notes