Programmable Fault Injection: Turning Abnormal Inputs into Test Assets
Ten thousand kilometers of normal driving without a single false brake from AEB proves nothing — everything functional safety truly cares about hides in the abnormal: radar frames dropped for 300 ms, the vehicle-speed signal frozen at a stale value, a CRC that suddenly doesn’t check out… These inputs are rare, dangerous, and nearly impossible to wait for naturally during testing. Fault injection is the means of manufacturing them on purpose, and “programmable” is what decides whether it stays a one-off craft or becomes a test asset that lives in git and runs in nightly regression. This article covers two things: why fault injection is worth making programmable, and what a programmable fault-injection engine looks like.
1. What Are “Abnormal Inputs”?
Abnormal inputs = the bad inputs an SUT may receive in the real world but can hardly reproduce naturally in testing. Classified by layer:
| Layer | Typical abnormality | Real-world source |
|---|---|---|
| Data | Wild values, out-of-range values, NaN, negative vehicle speed | Sensor failure, calibration error |
| Communication | Dropped frames, duplicated frames, reordering, bit flips, CRC errors | Bus interference, gateway failure |
| Timing | Delay, jitter, timeout, cycle drift | Sender ECU overloaded, rebooting |
| Protocol | Rolling counter not incrementing, wrong message length | Software bug on the peer node |
| State | Diagnostic session jumping around, SecurityAccess brute-forcing | Misbehaving tester, attacks |
What these inputs share: rare, dangerous, and exactly what functional safety cares about most. AEB behaving well on normal inputs is no achievement — whether it false-brakes when it receives “distance = 500 m, then 5 m three frames later” is what ISO 26262 wants verified.
2. What Does “Programmable” Mean?
Compare the two approaches.
The non-programmable (hard-coded) approach — fault logic baked into C++ code:
// Want to test frame drops? Edit code, recompile, edit it back afterwards
void onCanFrame(const CanFrame& f) {
static int cnt = 0;
if (f.id == 0x123 && ++cnt >= 100 && cnt <= 102) return; // drop 3 frames
forward(f);
}
The problems are obvious: every new fault scenario means changing platform code and recompiling — like an old telephone switchboard operator crawling back into the code to re-plug the wires; fault logic gets tangled up with business logic; and no test asset ever accumulates — the next project has to dig through git history for that if.
The programmable approach — faults are declared rules, and the platform provides an engine to interpret them:
# Declared in the test case; not a single line of C++
faults:
- name: radar_track_dropout
at: 1000ms # when to inject
duration: 300ms # how long it lasts
channel: can0
match: { id: "0x2A0..0x2AF" } # which frames it applies to
action: drop # action: discard
- name: speed_freeze
at: 2000ms
duration: 500ms
channel: can0
match: { id: "0x123" }
action: freeze # freeze the vehicle_speed field — manufacture "stale data"
On the platform side there is just one generic fault-execution engine: match rules → trigger at the right moment → apply the action to the data stream. A test engineer adds a new scenario = writes YAML, without touching platform code. That is “programmable” — fault behavior becomes data, not code. (The example is simplified — the full syntax for parameterized actions like field-level freezing is in Section 7.)
3. Why It’s Worth Making Programmable
- Fault scenarios become test assets: they can go into git, be reviewed, be reused, and join the nightly regression suite — just like normal stimulus, all orchestrated in YAML;
- Aligned with the functional-safety evidence chain: ISO 26262 Part 6’s software-testing method tables list fault-injection testing as recommended for ASIL C/D — to verify that safety mechanisms actually work, e.g. whether the SUT enters its safe state after the E2E protection detects a CRC error. Every YAML fault rule + verdict result is a traceable piece of safety-verification evidence;
- It composes with normal stimulus into complex scenarios: normal stimulus drives the SUT toward a dangerous condition (target approaching) while an anomaly is injected at the critical frame (drop exactly 3 frames right now) — “what if the SUT goes blind for 300 ms at the very moment of decision?” Only programmability makes such combinations possible.
One sentence: abnormal inputs are the SUT’s “stress exam questions,” and programmability lets those questions be declared, stored, reused, and replayed like ordinary test cases — fault injection turns from a one-off code hack into a first-class platform capability. Now let’s look at the concrete design of a programmable fault-injection engine.
4. Where It Lives: An Adapter Decorator
The core idea in one sentence: fault injection is not a new module — it is a “rule pipeline” (FaultPipe) added to the data path of the bus adapter (IBusAdapter) — rules loaded from a YAML DSL, driven by the Time Master’s tick, and written under real-time-path C++ discipline throughout.
Don’t let fault logic grow into the SocketCAN implementation. Wrap the existing adapter in a decorator so the injection point stays transparent to the business logic:
┌──────────────────┐ ┌─────────────────┐ ┌──────────────────┐ ┌─────┐
│ YAML rules │──▶│ FaultPipe │──▶│ IBusAdapter │──▶│ SUT │
│ (parsed at load) │ │ (rule pipeline) │ │ (SocketCAN etc.) │ │ │
└──────────────────┘ └─────────────────┘ └──────────────────┘ └─────┘
▲
Time Master tick
- RX direction (bus → SUT): injecting here feeds the SUT abnormal input (dropped frames / wild values / timeouts);
- TX direction (SUT → bus): injecting here tampers with the SUT’s output, to verify downstream protection mechanisms (e.g. whether the receiver’s E2E check catches a broken CRC).
5. Core Data Structures: variant + Preallocation
The action set is closed, so use std::variant; the rule container is fixed-size, filled once at load time — zero heap allocation on the RT path:
// fault_rule.h — POD-ish structures only, no member that allocates
struct Drop {};
struct Delay { uint32_t delay_ticks; };
struct CorruptField { uint8_t byte_offset; uint8_t mask; uint8_t value; }; // simplified
struct Freeze {}; // freeze the field at its previous value
struct CorruptCrc {};
using FaultAction = std::variant<Drop, Delay, CorruptField, Freeze, CorruptCrc>;
struct FaultRule {
uint32_t id;
const char* name; // points into a string pool allocated at load time, living for the whole run
uint8_t channel_mask; // bit0 = can0, bit1 = can1...
uint32_t id_min, id_max;// frame-ID filter interval
Tick start_tick; // the Time Master's tick count
Tick end_tick;
FaultAction action;
// Mutable run state (Freeze's last value, Delay's pending-queue pointer, etc.)
// reset at the start of every run to guarantee reproducibility
bool active(Tick now) const { return now >= start_tick && now < end_tick; }
};
class FaultPipe {
std::array<FaultRule, kMaxRules> rules_{}; // kMaxRules = 64 or so
size_t rule_count_ = 0;
RingBuffer<PendingFrame, 256> delayed_; // preallocated ring buffer for the Delay action
public:
std::optional<CanFrame> process(const CanFrame& f, Tick now);
void releaseDelayed(Tick now, FrameSink& out); // release due delayed frames every tick
void resetAll(); // clear state between runs
};
The key point: YAML parsing happens at load time (where heap allocation is harmless); afterwards the rules are copied into the fixed-size array and the strings into the string pool. Once the simulation loop starts, there is not a single new on this path.
The action library only needs the four or five most common ones first: drop, delay, corrupt_field (modify a value), freeze (hold the old value), corrupt_crc — these five cover 80% of fault-injection needs in functional-safety verification.
6. Runtime: A Three-Stage Pipeline on the Tick
A single-threaded 1 ms scheduler makes the processing order inside each tick strictly deterministic:
// Inside each tick, the scheduler executes in a fixed order:
void SimLoop::tick(Tick now) {
faultPipe_.releaseDelayed(now, busSink_); // 1. release due delayed frames first
replayEngine_.emit(now, busSink_); // 2. CSV normal stimulus (through the FaultPipe)
adapters_.pollAll(); // 3. receive real bus frames (through the FaultPipe)
verdict_.evaluate(now); // 4. verdict
}
// FaultPipe::process is a pure matching pipeline, rule by rule in declaration order:
std::optional<CanFrame> FaultPipe::process(const CanFrame& f, Tick now) {
for (size_t i = 0; i < rule_count_; ++i) {
auto& r = rules_[i];
if (!r.active(now) || !matches(r, f)) continue;
logFaultEvent(r, f, now); // key: every hit enters the evidence chain
return apply(r, f, now); // Drop → nullopt; Delay → into the ring buffer
}
return f; // no hit — pass through unchanged
}
The “three stages” are the FaultPipe’s three hook points in each tick (steps 1–3: release delayed frames, emit replay stimulus, receive real bus frames — all passing through the FaultPipe); step 4, the verdict, runs after the pipeline and is not part of the data path.
Why does a fixed order matter? The same frame may hit both a “drop” rule and a “corrupt value” rule, and which one wins must be a hard rule (YAML declaration order) — otherwise two runs diverge and bitwise reproducibility is broken.
7. The YAML DSL: Faults and Normal Stimulus Live in the Same Case File
test_case: aeb_radar_dropout
duration: 5000ms
stimulus: # normal stimulus
- replay: scenarios/aeb_approach.csv
channel: can0
faults: # abnormal input, programmable
- name: radar_dropout
at: 1000ms
duration: 300ms
channel: can0
match: { id: "0x2A0..0x2AF" }
action: drop
- name: speed_freeze
at: 2000ms
duration: 500ms
channel: can0
match: { id: "0x123" }
action: freeze
verdict: # verdict as usual
- expect: brake_request == true
by: 2500ms
- expect: dtc_reported == true # 300 ms of dropped frames should trip the timeout diagnostic
by: 2000ms
A test engineer adds a scenario = adds a few lines of YAML, with zero platform-code changes. That is the final form of “programmable.”
8. The Evidence Chain: Injection Events Themselves Must Be Observable
This is what many people miss — if the fault injection never happened, a passing verdict counts for nothing (say the rule’s ID interval was written wrong, no frame was ever dropped, and the test ran for nothing). So every rule hit emits a FaultEvent onto the bus log (FlatBuffers):
// Record: which tick, which rule, on which frame, what action
FaultEvent{ now, r.id, f.id, actionTag(r.action) }
The verdict engine then supports a new kind of assertion:
verdict:
- expect: fault_injected(radar_dropout) # confirm the fault really went in
- expect: fault_hit_count(radar_dropout) >= 240 # ~240 frames should be dropped within 300 ms
The JUnit XML report writes the injection statistics into <properties>, so whoever reviews the report sees at a glance that “this case’s faults actually took effect.” This is exactly the evidence shape ISO 26262 fault-injection testing asks for.
9. The Fault Injector Itself Must Be Tested Too
The injector is part of a safety-verification tool and may one day be classified as a TCL3 tool — its erroneous output can directly contaminate safety evidence without being easily detected. So give it unit tests:
TEST_F(FaultPipeTest, DropRuleSuppressesMatchingFramesOnly) {
// Fixture: fixed fake clock + 3 rules + a golden input sequence of 20 frames
// Assert: the output sequence matches the golden file byte for byte
}
One fixture per action type, with golden frame sequences for both input and output — the injector is deterministic pure logic, inherently easy to test.
10. Suggested Rollout Order
| Phase | Content | Rationale |
|---|---|---|
| P0 | drop + delay actions + FaultEvent logging |
Get the structure running; covers the most common drop/timeout scenarios |
| P1 | corrupt_field + freeze + corrupt_crc |
Data-layer anomalies, paired with E2E verification scenarios |
| P2 | RX/TX bidirectional injection points + the verdict’s fault_injected assertion |
Close the evidence-chain loop |
| P3 | Sequenced faults (a “fault train” like “drop 3 frames → 5 normal frames → drop 2 more”) | Covers intermittent-fault scenarios |
P0 can be up and running in two or three days: the tick, adapters, YAML, and log bus are all off-the-shelf infrastructure in a mature XiL platform — the FaultPipe is just a layer of pure logic stringing them together.
Looking back along the chain: “programmable abnormal input” = a fixed-size rule array + variant actions + a tick-driven matching pipeline + YAML declarations + FaultEvent evidence. Write the platform code once, and fault scenarios are data from then on — into git, into review, into nightly regression, on equal footing with normal stimulus.
Appendix: Glossary (in order of appearance)
| Term | Plain explanation |
|---|---|
| AEB | Autonomous Emergency Braking; the recurring example SUT function in this article |
| Functional safety | The engineering discipline of preventing personal harm from E/E system failures; the automotive standard is ISO 26262 |
| Fault injection | The test technique of deliberately manufacturing abnormal inputs for the system under test |
| SUT | System Under Test — the object this test aims at |
| NaN | Not a Number, the “not-a-number” value in floating point, often from illegal arithmetic or invalid sensor data |
| Calibration | The work of assigning values to ECU parameters (thresholds, gains, etc.); calibration error is a typical source of data-layer anomalies |
| CRC | Cyclic Redundancy Check: the checksum field in a message that detects corrupted data |
| Rolling counter | A per-frame incrementing counter in the message, against replay attacks; a wrong sequence gets judged “data not trustworthy” |
| SecurityAccess | The security-unlock service (0x27) in UDS diagnostics; “brute-forcing” means an attacker repeatedly guessing the key |
| ISO 26262 | The automotive functional-safety standard; explicitly recommends fault-injection testing for ASIL C/D |
| YAML | A human-readable data-description format; this article’s test cases and fault rules are all declared in it |
| ASIL | Automotive Safety Integrity Level (QM/A/B/C/D, D the strictest) |
| E2E protection | AUTOSAR’s end-to-end communication protection: rolling counter + CRC and friends, keeping messages trustworthy |
| Evidence chain | The complete set of records that lets a third party recompute and verify a verdict |
| Verdict | The pass/fail conclusion the test platform gives a test case |
| Decorator | A design pattern that adds functionality by wrapping a layer outside, without changing the original interface |
| SocketCAN | The Linux kernel’s CAN protocol stack and driver framework |
| Time Master | The platform module that advances simulation time uniformly |
| Tick | The smallest time step of the simulation scheduler; the scheduler here runs one tick per 1 ms |
| std::variant | C++17’s type-safe union: a closed “one of these few types” |
| RT path (real-time path) | A code path with hard constraints on latency and memory allocation |
| Heap allocation | Requesting memory from the heap at runtime; normally forbidden on the RT path |
| DSL (domain-specific language) | A small description language built for a specific domain; this set of YAML fault rules is the fault-injection DSL |
| Bitwise reproducible | Two runs produce bitwise-identical output |
| FlatBuffers | A zero-copy serialization library, used here as the encoding format of the bus log |
| JUnit XML | The universal test-result reporting format, understood by CI and test-management tools |
| TCL3 | One of ISO 26262’s tool confidence levels: the tool’s erroneous output can contaminate safety evidence without being easily detected |
| Golden file | A pre-frozen reference-answer file; output is compared against it byte for byte |
| Fault train | A scripted fault sequence alternating “fault/normal,” used to simulate intermittent faults |
Next step:
View all notes