Verification on a Budget: Vertical Slices, Test Doubles, and Stimulus Injection

There is never enough verification budget — rig hours are queued for, CI time has a cap, and engineers’ patience is the scarcest resource of all. The question is: with the same budget, how do you buy the most confidence? This article covers three moves that happen to form one pipeline:

  1. Vertical slices answer the strategy question — what to verify first? The riskiest assumption;
  2. stubs and fixtures answer the mechanics — how to cut the dependencies of the code under test, and how to keep the test environment clean;
  3. stimulus injection answers the execution question — through which entry point, and at which moment, input signals get fed in.

All three share one idea: buy confidence cheaply.

1. Vertical Slices: Test the Riskiest Assumption First

A slice is the method of cutting one small piece out of a big system and finishing it first. The key is telling the two cutting directions apart:

Horizontal (cut by layer): finish the whole data layer, then the whole logic layer, then the whole UI layer — the risk is discovering at the very end that the layers don’t mesh: all of integration hell lands in the second half.

Vertical (cut by feature): one cut through all the layers, building one feature that is narrow but complete — instead of “the full backend plus the full frontend,” you first build “sign-up and login,” working end to end from database to UI.

Horizontal: ████████ data        Vertical: █
            ████████ logic                 █  sign-up/login
            ████████ UI                    █  (every layer pierced)
            (runs only at the end)         (runs in week one)

Three iron rules judge whether a slice is cut well:

  1. Pierce every layer: miss any one layer and it isn’t a slice — it’s a half-finished layer;
  2. Runnable end to end: demo-able and acceptable on its own — “it runs” is the defining property of a slice;
  3. Cut along the riskiest seam: where to cut is not random — aim at the seam where the architectural assumptions are most dubious. The point of slicing is not saving labor; it’s testing the most dangerous assumption at the smallest cost.

Take a domain-controller SIL platform as the example. Where does the first cut go? Not the easiest feature — the AEB closed loop: clock → scheduler → bus → SUT → verdict → report, every layer included. It was chosen not because it’s simple — quite the opposite: time-base abstraction, bus abstraction, the measurement chain, the SUT lifecycle, and case verdicts — five high-risk architectural assumptions all live on this chain. The pure-workload items (more frame formats, more bus protocols) are low-risk and go to the back of the queue. Once the loop runs end to end — positive, negative, and determinism acceptance lines all green — the architectural bet has paid off.

That is the essence of slicing: a green SIL closed loop means the architecture bet was right, and hardware money can follow; if SIL can’t even run smoothly, what you lose is a dozen person-weeks, not a rig. A slice is an option that caps your worst-case loss.

A close relative is the walking skeleton (coined by Alistair Cockburn): stand up the smallest working skeleton of the system, make it walk a few steps, then grow flesh on it.

2. Stubs: The Dummy Loads of Software

Once the slice decides “what to build first,” the next problem is: the code in the slice depends on modules and hardware that don’t exist yet — how do you run it anyway? The answer is a test double, and the most commonly used kind is the stub: a fake part with pre-recorded canned answers that replaces a real dependency during tests.

The most fitting analogy is from hardware: the dummy load. When you test a power-supply module, you don’t hook up a real device — you connect a resistor box. It isn’t a real load, but it makes the supply think it’s loaded, exposing the supply’s own problems. A stub is the software dummy load: the code under test doesn’t know the other side is fake.

// Real dependency: reads a real vehicle-speed sensor (no such hardware in the test env)
class ISpeedSensor { public: virtual float speed_kmh() = 0; };

// Stub: a fake test sensor that returns a preset value
class StubSpeedSensor : public ISpeedSensor {
public:
    explicit StubSpeedSensor(float v) : v_{v} {}
    float speed_kmh() override { return v_; }   // canned answer: always says 60
private:
    float v_;
};

// Test: plug the fake sensor into the AEB logic, check the braking decision at 60 km/h
TEST(AebTest, brakes_at_60kmh) {
    StubSpeedSensor stub{60.0f};   // ← the stub stands in for real hardware
    AebLogic aeb{stub};
    EXPECT_TRUE(aeb.should_brake());
}

Three traits: minimal implementation (a few lines), predictable behavior (fixed return values), test-only (never ships in production code).

The double family: don’t confuse stub with mock

There are five test doubles. The industry mixes the names up all the time, but the precise distinctions are:

Double What it does What it cares about
Dummy Only fills out a parameter list; never called Nothing
Stub Preset canned answers (“ask, get 60”) State: steer the code under test onto a path
Fake A working simplified version (an in-memory DB standing in for the real one) Works, but simplified
Spy A wiretap that records the calls Interactions: how many calls, with what arguments
Mock Presets “how it should be called”; fails otherwise Expected interactions (strongest, most brittle)

A one-line mnemonic: a stub owns the answers (canned answers); a mock owns the questions (expectations) — a stub verifies state (“given 60, did it brake?”), a mock verifies behavior (“did it call the brake function, with the right arguments?”).

Two kinds of stub in the embedded world

The industry uses the same word in two places, and both are worth recognizing:

  1. The unit-test stub: as above, replacing function/module dependencies. When testing ECU software, swap CAN transceiving and ADC reads for stubs, and the logic runs on a PC — this is exactly what lets SIL leave the hardware behind;
  2. Restbus simulation: the part of a CANoe setup that simulates “the other 20 nodes on the bus” is essentially a whole car full of stubs — each simulated node sends preset frames on schedule per the DBC. When you test a single domain controller, the entire vehicle network is one big stub.

Push the second idea one notch further and you get a stub SUT: a fake device-under-test with a minimal behavioral model — it takes injected bus/OSI data and replies per preset logic (e.g. UDS 0x10/0x22/0x27 responses). Because it’s fully software-controlled and fake, you can make it fail at will: return the wrong NRC, answer 500 ms late, emit a malformed CAN frame — the stub now doubles as a fault injector. Its purpose is not to test itself, but to run the test platform through before the real SUT arrives — the walking skeleton’s “skeleton” stands up on stubs.

One sentence: a stub is a “controllable illusion” — to test A, replace all of A’s neighbors with stubs and let A show its true colors in a sterile environment.

3. Fixtures: Clamp It In, Start from the Same Line Every Time

Stubs solve “the dependencies are fake.” One question remains: “is the environment the same for every test?” That’s the job of the fixture. The word is borrowed from hardware testing: on a production line, an ECU board gets clamped into a test jig — probes press on the test points, power is supplied, loads are connected; after the test, the board is released and the next one goes in. The jig guarantees one thing: every unit under test gets an identical, repeatable test environment.

Software testing borrowed the same concept: a fixture = the fixed environment shared by a group of tests, set up before each one and torn down after. In Google Test it looks like this:

class AebTest : public ::testing::Test {   // the fixture class: defines the shared environment
protected:
    void SetUp() override {                // auto-called before each case: build the environment
        sensor_ = new StubSpeedSensor{60.0f};
        aeb_    = new AebLogic{*sensor_};
    }
    void TearDown() override {             // auto-called after each case: tear it down
        delete aeb_;
        delete sensor_;
    }
    StubSpeedSensor* sensor_;              // fixture members: objects shared by all cases
    AebLogic*        aeb_;
};

TEST_F(AebTest, brakes_at_60kmh) {         // TEST_F: run with this fixture
    EXPECT_TRUE(aeb_->should_brake());
}
TEST_F(AebTest, no_brake_when_clear) {
    StubSpeedSensor slow{0.0f};           // swap the input inside the case body: a 0 km/h stub
    AebLogic aeb2{slow};
    EXPECT_FALSE(aeb2.should_brake());
}

The key mechanism (many people use GTest for a year without noticing it): every TEST_F gets a brand-new fixture object — SetUp → test body → TearDown, then it’s destroyed. So there’s zero state residue between cases: if the last case mangled aeb_’s internals, the next one still gets a clean one. That is the fixture’s core value: isolation between cases.

Fixture vs stub: different jobs

The two often show up together, but they do different things:

  Stub Fixture
What it replaces The SUT’s dependencies (fake sensor, fake bus) The test’s own preparation (build objects, feed initial values, clean up)
Analogy The dummy load The jig itself
Relationship A fixture often holds stubs The fixture is the container; the stub is a prop inside it

In the code above, StubSpeedSensor is the stub, and the whole AebTest class is the fixture — the fixture “clamps the unit into the jig”; the stub “plays the fake part mounted on the jig.”

By the way, nearly all of software-testing vocabulary is borrowed from hardware testing: fixture ← the jig/rig (a HIL rig is essentially a physical fixture: fixed power, fixed harness, fixed load box); stub ← dummy load; harness ← wiring harness; golden sample ← the reference part in metrology. pytest’s fixtures share the same root with a fancier mechanism — dependency-injection style, reusable at session/module/function scope.

Rules of use

  1. Never share state between cases: static members and globals break the fixture’s isolation promise;
  2. SetUp holds only what every case needs; case-specific setup goes in the case body — don’t bloat the fixture;
  3. Sink heavy resources to the suite level: expensive operations like starting a simulator belong in SetUpTestSuite() (once per suite), not in per-case SetUp;
  4. TearDown must be symmetric: every resource acquired must be returned — in embedded work, a fixture that forgets to reset hardware state causes the haunted bug “different case order, different result.”

One sentence: a fixture guarantees “every test starts from the same starting line” — half of a test’s credibility lives in the code under test, the other half in whether the fixture is clean.

4. Stimulus Injection: The Test Platform’s Hands

The slice sets the scope; stubs and fixtures build the environment; the last step is making the SUT “act out” the scenario you want. Stimulus injection is the testing term — common especially in embedded and XIL testing:

Stimulus = the input signals fed to the object under test; injection = “punching” that signal into the system at a chosen moment, through a chosen entry point, replacing or driving its original real inputs.

The SUT won’t act on its own — it’s an input → process → output pipeline. To make it show a behavior, you inject the matching inputs from upstream and then watch whether the outputs meet expectations. Testing a domain controller’s AEB function:

t=0ms:     inject CAN frame → lead target at 50 m, relative speed -10 m/s
t=1000ms:  inject update frame → distance 30 m (target closing in)
t=2000ms:  distance 15 m → observe: did the SUT issue a brake request?

Every frame here is one stimulus injection. The CSV-replay feature of a test platform is essentially an offline-scripted stimulus-injection sequence — the CSV defines the stimuli, and the replay engine injects them one by one on the master clock’s beat.

Choosing the injection point

From far from the SUT to close to it, injection points come in layers:

Layer What gets injected Example Fidelity
Physical/electrical Voltage, current, resistance, PWM Simulating a wheel-speed sensor signal on a HIL rig Highest (needs real hardware)
Bus CAN/CAN FD/automotive Ethernet frames Sending frame 0x123 over SocketCAN High
Sensor data Raw point clouds, image frames, IMU data Replaying a lidar point-cloud packet Medium-high
Software interface DDS topics, API calls, function arguments Mocking the positioning module’s return value Medium (fast, highly controllable)

The principle: the closer the injection point is to the SUT’s real input boundary, the truer the test; the deeper inside, the cheaper the injection — but the more likely you’re “testing a mannequin.” This is the same question as where stubs live from the previous section — a stub is a stimulus source sitting on an injection point.

Relations to neighboring concepts

Three key engineering questions

  1. Timing determinism: a stimulus must land at its intended moment, not “as soon as possible.” In virtual-time mode, the 500 ms stimulus must land on the 500th tick — this is a fundamental reason a test platform wants deterministic scheduling and a dual-mode clock;
  2. Multi-channel synchronization: real scenarios have concurrent stimuli (CAN + point cloud + a diagnostic request at once); the injection engine must dispatch multiple channels in dependency order within one tick;
  3. Reproducibility of stimuli: run the same stimulus definition twice, and the SUT must see a bit-for-bit identical stream — otherwise the verdict is untrustworthy.

One sentence: stimulus injection is the test platform’s “hands” — the scenario is acted out by the system under test, and the way you make it act is to feed inputs through a chosen injection point at precise moments.


Looking back along the line: the slice picks the direction — test the most dangerous assumption first; stubs and fixtures do the decoupling — swap dependencies for controllable illusions and keep the environment clean every time; stimulus injection does the execution — feed inputs through the chosen point at the exact moment. All three serve one goal: before the real money goes into hardware, verify the riskiest assumptions at the smallest cost. Confidence doesn’t have to be expensive — what matters is where you spend it.


Appendix: Glossary (in order of appearance)

Term Plain explanation
vertical slice A minimal end-to-end feature that pierces every architectural layer, built to test the riskiest assumption first
AEB Autonomous Emergency Braking; the canonical ADAS safety feature used as this article’s example
SUT System Under Test — the object being tested
SIL Software-in-the-Loop: the controller software runs fully simulated on a PC/server
walking skeleton Alistair Cockburn’s concept: stand up the smallest working skeleton of a system, then grow flesh on it; a close relative of the slice
test double The umbrella term for fake objects that replace real dependencies in tests
stub The double with preset canned answers; stands in for a dependency
dummy load The hardware prototype of the stub: a resistor box that stands in for a real load when testing power supplies
Dummy / Fake / Spy / Mock The other four doubles: parameter filler; working simplified version; call recorder; interaction-expectation enforcer
restbus simulation Simulating “all the other nodes on the bus” so a single ECU thinks it’s on a real vehicle network
CANoe A widely used automotive bus-simulation and test tool
DBC The CAN database file describing which nodes send which frames and signals
UDS Unified Diagnostic Services (ISO 14229); 0x10 session switch, 0x22 read data, 0x27 security unlock are its service IDs
NRC Negative Response Code — the reason code a diagnostic service returns when it refuses a request
fault injection Feeding abnormal inputs (dropped frames, CRC errors, timeouts, wild values) to stress robustness
fixture The fixed environment a test group shares: set up before each case, torn down after; borrowed from hardware test jigs
GTest (Google Test) Google’s C++ test framework; its TEST_F + SetUp/TearDown is the canonical fixture mechanism
HIL Hardware-in-the-Loop: a real ECU wired to a simulation rig
harness / golden sample More testing words borrowed from hardware: the wiring harness; the metrology reference part
pytest The Python test framework whose fixtures share the same idea with a fancier mechanism (scoped at session/module/function level)
stimulus injection Feeding input signals into the SUT at chosen moments through chosen entry points
XIL X-in-the-Loop: the family name for MIL/SIL/PIL/HIL-style closed-loop test environments
SocketCAN The Linux kernel’s CAN protocol-stack interface
DDS Data Distribution Service — a real-time publish/subscribe middleware

Next step:

View all notes