PIL (Processor-in-the-Loop): Why Passing SIL Isn't Enough
PIL is the most frequently challenged link in the XiL family: “SIL already passed — why not go straight to HIL?” No, and the reasons are concrete. Passing SIL only proves that the algorithm logic of the source code behaves correctly under a PC compiler; the production controller runs a different binary in a different numeric environment. This article covers PIL end to end: what it is, the four walls between SIL and the target machine, its core method of back-to-back comparison, why these problems must not be left for HIL to catch, and how PIL is morphing in the domain-controller era.
1. What PIL Is
PIL = Processor-in-the-Loop. The software under test is compiled with the target chip’s compiler into machine code and run on a real processor (an evaluation board) or an instruction-set simulator; the rest of the loop — the plant, the environment model, the test framework — stays simulated on the PC.
What distinguishes the XiL levels is “what gets placed inside the loop”:
| Level | What’s in the loop | Where it runs |
|---|---|---|
| MIL | The model itself (block diagram) | PC, pure math |
| SIL | The x86 binary compiled from source | PC, Docker / native |
| PIL | Target machine code from the target compiler | Real processor or instruction-set simulator |
| HIL | The whole ECU (processor + memory + IO + connectors) | Real-time rig + real buses |
The closed loop of one PIL test:
PC test framework Target processor (eval board)
───────────────── ─────────────────────────────
Load test vectors
│ ① Download target binary (JTAG / Ethernet)
│ ──────────────────▶
│ ② Send stimulus for step i (input signal values)
│ ──────────────────▶ ③ Function under test runs one step on target code
│ ④ Read back outputs + internal state
│ ◀──────────────────
Compare against expected / SIL baseline
│ ⑤ Next step...
The PC is the “director,” the processor is the “actor” — everything is driven and sampled by the test framework.
Two forms: a real evaluation board (JTAG flashing, highest fidelity, and you get execution time and stack high-water marks for free); or an instruction-set simulator (ISS) — pure software executing target machine code instruction by instruction (QEMU is the free close relative; commercial options like Synopsys VDK fit into CI and can be cycle-accurate).
What it measures / doesn’t measure: target-code correctness (compiler differences, float/fixed-point precision, word-length alignment), execution time, and stack usage. It does not measure real IO electricals, real sensors, or bus physical layers — that’s HIL territory. At the PIL level, inputs and outputs are “numbers,” not voltages.
2. Why It Exists: Four Walls Between SIL and the Target
Passing SIL only proves that the algorithm logic of the source code behaves correctly under a PC compiler. What a production ECU runs is a different story:
- Different compiler. Automotive-grade MCUs use TASKING / Green Hills / specific cross GCCs — their code generation is a different species from a PC compiler. Same C source, different compiler, different
-O2, and undefined behavior (overflow, strict aliasing) can surface on the spot. - Different numeric environment. Many classic MCUs have no double-precision FPU, or are entirely fixed-point. A control law that computes happily in
doubleunder SIL becomesfloator even Q15 fixed-point on target — quantization error and overflow are simply invisible to SIL. - Different resources. On a PC, the stack is all-you-can-eat; on an MCU, one task gets 2KB. SIL can never detect a stack overflow.
- Different timing. On x86, a 10ms step is trivial; on a 160MHz Tricore, does the control function take 800µs or 8ms per step? SIL has zero awareness of execution time.
3. Back-to-Back: PIL’s Core Method
Back-to-back testing is PIL’s core method: drive both sides with exactly the same test vectors used in SIL, and compare the outputs point by point (within tolerance). If SIL ≈ PIL, the algorithm has survived the “code generation → compilation → target processor” chain intact; if they diverge, the problem is in the chain, not in the algorithm — the suspects are the code generator, the compiler, and the target runtime, while the control logic can be ruled out first.
This is exactly the value of XiL tiering: each level introduces only one new class of variables, so a failure comes with a bounded suspect list.
4. Deep Dive: Two Sources of Back-to-Back Divergence
When both sides use IEEE double with the same compiler family (say, x86 versus an ARM64 dev board), back-to-back divergence is ≈ 0 — beautifully boring. In real projects, divergence appears once optimization techniques enter the picture. The two most common sources are fixed-point and SIMD intrinsics — they are how you make code run at all, and run fast, on the target chip; the price is numeric divergence from the x86 reference implementation, and back-to-back comparison is what measures that divergence.
(a) Fixed-point: pretending decimals are integers
What it is: scale values by a fixed factor and store them as integers. In Q15 format, an int16_t stores “true value × 32768” — 0.5 is stored as 16384. Addition and subtraction are plain integer operations; multiplication needs care with the scale — Q15×Q15 yields Q30, so the result must be shifted back with >>15.
Why it exists: classic automotive MCUs (the AUTOSAR CP side) have no floating-point unit (FPU), or only a feeble single-precision one — a float multiplication emulated in software is tens of times slower. Every chip has a fast, deterministic integer ALU. Fixed-point = trading “the programmer manages scaling by hand” for “the hardware needs no floating point.”
The cost (measured with g++ on x86):
0.1 → stored as 3276 in Q15 → restored as 0.0999755859, error 2.44e-05 ← quantization: 0.1 can't be stored exactly at all
0.1 × 0.2: true value in double is 0.02, Q15 computes 0.0199890137 ← every multiplication adds another rounding
- Quantization error accumulates inside integrators;
- Every multiplication’s
>>15rescaling loses precision; - Overflow risk: two large Q15 values multiply into Q30 — a single multiplication always fits in int32, but −1×−1 shifted back to Q15 yields +1, which exceeds the int16 range (hence saturation), and a running multiply-accumulate (MAC) chain can still blow int32.
(b) SIMD intrinsics: one instruction, a row of numbers
What it is: SIMD = Single Instruction Multiple Data — the CPU’s vector unit computes 4/8/16 values with one instruction (NEON on ARM, SSE/AVX on x86). An intrinsic is a “pseudo-function” the compiler provides, one per machine instruction — e.g. NEON’s vaddq_f32 (four floats added in parallel) or SSE’s _mm_add_ps.
Why it exists: AP domain controllers run perception and signal processing, where element-by-element loops can’t reach the required throughput — vectorization is routine.
The cost: floating-point addition is not associative (measured):
(x + big) - big = 0 (x=1e-9, big=1e9: the big number swallows the small one)
x + (big - big) = 1e-09 (reassociate, and the result is correct)
A SIMD reduction sums in a necessarily different order than a scalar loop; add FMA (fused multiply-add — one rounding versus two) and the different handling of denormals (flush-to-zero) between x86 and ARM, and it’s normal for the x86 scalar reference and the ARM NEON version of the same algorithm to disagree bit for bit.
(c) How to judge divergence: a tolerance band, not bit-identity
- CP side (MCUs without FPU): the algorithm lands as Q15 fixed-point, diverging from the x86 double reference by quantization;
- AP side (domain-controller SoCs): performance hot spots use NEON intrinsics, diverging from the scalar reference in floating point.
Only now does back-to-back turn from a “formality” into a “tool”: the question it answers is not “are both sides bit-identical” (they never are), but “does the divergence stay inside the tolerance band.” Report max abs error / max rel error, and only out-of-band counts as a bug — the idea that “pass/fail thresholds scale with the object under test” is universal in testing; here it judges numeric precision rather than time.
5. Why Not Leave It for HIL to Catch
- HIL is the most expensive resource: a rig worth millions is queued up across the whole company; letting compiler differences and stack overflows — low-level problems — consume HIL rig hours is burning gold like iron. PIL filters target-side problems first, using a board that costs a few thousand, or a pure-software ISS;
- HIL observability is terrible: the whole cabinet is a black box, and when something dies you only see the bus-level symptoms; PIL runs with a debugger attached — single-step, inspect memory, measure target-side coverage and stack high-water marks;
- Fault-localization cost: the essence of each XiL level is narrowing the suspect list — MIL checks the model, SIL checks the algorithm, PIL checks the target chain, HIL checks system integration. Skip a level and the search space at failure doubles;
- Standards require it explicitly: for ASIL C/D, ISO 26262 requires unit/integration verification in the target environment or a representative one, with target-side coverage. HIL is system-level testing and cannot substitute for that.
6. How PIL Is Morphing in the Domain-Controller Era
On the ADAS domain-controller track, the gap between target and development environments is shrinking: Linux + ARM64/x86 + same-family GCC + IEEE floating point throughout — PIL degenerates into “run SIL once more in a target-isomorphic environment (QEMU, containerized arm64, a dev board).” Traditional PIL’s home turf is the classic MCUs running AUTOSAR CP, fixed-point, and TASKING compilers (body / chassis / powertrain).
So whether PIL is a must depends on the project’s target environment: on classic MCU projects it’s a hard requirement; on domain-controller projects it shrinks to a lightweight step. But the principle doesn’t change — as long as “the development-environment binary” and “the target-environment binary” are not the same artifact, you need a step that proves the two behave the same.
Looking back at the chain: SIL proves the algorithm is right; PIL proves the algorithm is still right on the target chain. In between stand four walls — compiler, numerics, resources, timing — and back-to-back is the rope ladder over them: run the same vectors on both sides, compare point by point, and only out-of-band divergence counts as a bug. Skip this step and leave it to HIL, and what you save is a board costing a few thousand; what you spend is rig hours worth millions, plus the search space for localizing the problem.
Appendix: Glossary (in order of appearance)
| Term | Plain explanation |
|---|---|
| PIL (Processor-in-the-Loop) | Run machine code built by the target compiler on a real processor or instruction-set simulator, with the rest of the loop simulated on a PC |
| ISS (instruction-set simulator) | A pure-software simulator that executes target machine code instruction by instruction; QEMU is the free close relative, commercial options like Synopsys VDK can be cycle-accurate |
| MIL / SIL / HIL | Model- / Software- / Hardware-in-the-Loop; PIL’s sibling levels in the XiL family |
| JTAG | A chip debug/flashing interface; the usual channel for downloading the target binary onto an eval board |
| stack high-water mark | The historical maximum depth of a task’s stack usage, used to assess overflow risk |
| cross compiler | A compiler that runs on a PC but produces machine code for the target chip; automotive projects commonly use TASKING, Green Hills, or specific cross GCCs |
| undefined behavior (UB) | C/C++ constructs the standard doesn’t define (e.g. signed overflow, strict-aliasing violations); they can surface when the compiler or optimization level changes |
| FPU (floating-point unit) | Hardware for floating-point arithmetic; classic automotive MCUs have none, or only a feeble single-precision one |
| back-to-back testing | Drive two implementations (e.g. SIL and PIL) with the same test vectors and compare outputs point by point against a tolerance band |
| fixed-point / Q15 | A numeric representation storing “true value × fixed factor” as integers; Q15 uses int16 to store “true value × 32768” |
| quantization error | The rounding error produced when a true value lands on the representable grid of fixed/floating point; it accumulates inside integrators |
| saturation | Clamping an out-of-range result to the max/min representable value instead of letting it wrap around |
| SIMD | Single Instruction Multiple Data: CPU vector capability that computes a row of values with one instruction |
| intrinsic | A compiler-provided “pseudo-function,” one per machine instruction — e.g. NEON’s vaddq_f32, SSE’s _mm_add_ps |
| NEON / SSE / AVX | The names of the SIMD instruction sets on ARM and x86 respectively |
| FMA (fused multiply-add) | One instruction computing a×b+c with a single rounding; differs from multiply-then-add, which rounds twice |
| denormals / flush-to-zero | Floating-point values extremely close to zero; x86 and ARM handle them differently (some flush them to zero), a source of cross-platform bit-level differences |
| ISO 26262 / ASIL | The automotive functional-safety standard and its safety integrity levels (QM/A/B/C/D, D strictest); ASIL C/D requires target-environment verification |
| AUTOSAR CP / AP | The two branches of the automotive software platform: Classic Platform (classic MCUs) and Adaptive Platform (domain-controller SoCs) |
Next step:
View all notes