Why Nightly Full Regression Is a Luxury in the HIL Era

“Trigger it when you leave work; read the report at dawn” — nightly full regression was standard CI in the internet/SIL era, so ordinary that nobody felt it needed defending.

In the HIL era, it becomes a luxury. The act quietly assumes three things — elastically scalable compute, compressible time, and cheap failure — and HIL happens to have none of them. This article does the math: why nightly full runs are expensive, how the industry schedules around it, and how global teams coordinate when there is no shared “night.”

1. Do the Math First

Suppose a case library of 5,000 cases at 3 minutes each on average → one full pass = 250 rig-hours. One cabinet is usable for about 10 hours a night → you need 25 cabinets in parallel to see results by morning. A single HIL cabinet runs from a few hundred thousand to over a million RMB — tens of millions in fixed assets, just to squeeze a full regression into one night.

It’s not that it can’t be done — it’s just expensive enough that only the biggest players can afford it (top OEMs do run “cabinet farms” for exactly this). That’s what “luxury” means in arithmetic.

2. The Three Assumptions HIL Doesn’t Have

2.1 Compute doesn’t scale elastically

2.2 Time doesn’t compress

2.3 Failure isn’t cheap

3. Two More Hidden Costs

4. So What Does the Industry Do?

Tier Cadence Why
SIL Full run every night Containers on demand, logical-clock speedup, reproducible failures — negligible cost
HIL A curated smoke/must-run set every night (tens to hundreds of cases) Rig time is expensive; burn it only on the most valuable cases
HIL full Weekends / pre-release / sharded across cabinets Only attempted once enough rig-hours have been saved up

The corollary: the scope: smoke / nightly / regression / release case-tag dimension exists precisely for the resource reality that “HIL can’t afford full runs” — full regression is pushed down to the SIL tier, and HIL uses tags to filter out subsets. “SIL first” isn’t just good practice; it’s an economic necessity.

5. Which Software Needs Nightly Regression?

Step back from HIL and look across the industry. Nightly regression exists for exactly one reason: the full test suite doesn’t fit into the minutes-long pipeline budget of every commit, so it gets batched into the night. Asking “which kinds of software need it” is really asking “whose complete verification is slow, expensive, and mandatory.” Walk the list:

Type Why it must run at night What runs at night
Embedded / automotive software Depends on scarce hardware (only a few HIL rigs); scenario libraries in the thousands Full SIL scenario regression, HIL rigs queuing smoke + core cases, cross-ECU variant matrices
OS / drivers / compilers Test suites measured in hours (Linux LTP/kselftest, the GCC/LLVM suites); multi-arch cross matrices Nightly full build of main + all-arch regression (KernelCI and the LLVM buildbot are exactly this pattern)
Databases / storage systems Crash recovery, fault injection, and long stress tests can’t run inside a PR Power-loss recovery, Jepsen-style consistency verification, multi-hour soaks
EDA / chip simulation The birthplace of regression-farm culture — one simulation takes hours, cases in the tens of thousands Nightly regression-farm queueing, failure reports in the morning (standard in Synopsys/Cadence flows)
Game engines / graphics software Cross-GPU/platform rendering comparisons, a combinatorial explosion of shaders Pixel-exact image regression, performance-baseline comparisons
Large web / microservice systems Full E2E in the thousands × a browser matrix Full Selenium/Playwright runs, cross-browser, performance baselines
Mobile apps Device fragmentation: dozens of real devices × multiple OS versions Nightly batches on device farms (the Firebase Test Lab model)
Finance / trading systems Correctness of end-of-day batch processing; regulatory compliance evidence Ledger reconciliation replays, recomputation against historical data
AI/ML models and platforms One eval-set run takes tens of minutes to hours; data and models change daily Nightly benchmark evals, accuracy regression, inference-performance drift monitoring
Medical device / aviation software Compliance-driven: every build must leave full verification evidence Full regression + evidence archiving (for certification audits)

6. The Common Traits — Does Your Software Need It?

Meet two or more of these, and nightly regression is basically a necessity:

  1. Full suite runs longer than 15 minutes: the MR pipeline can’t wait (industry consensus: once MR feedback exceeds 10 minutes, developers start working around it)
  2. Depends on scarce resources: rigs, GPU clusters, license-limited simulators — these must be scheduled, and queuing naturally suits the night
  3. State-space explosion: platform matrices (chip × OS × variant), scenario libraries, parameter combinations — full coverage can only be batched
  4. High quality bar: safety-critical / regulated industries need “full pass” evidence for every build
  5. External dependencies move daily: upstream data, third-party services, model weights — green last night, broken today; only the nightly catches it

7. The Standard Industry Layering

Every commit/MR   → smoke (≤5–10 min): build + static checks + 20–50 core cases
Every night       → full functional regression: entire case library + platform matrix + perf baseline
Weekly/on-demand  → stability soak: hours of stress, memory leaks, long-run stability
Pre-release       → scarce-resource full run: all HIL scenarios, certification evidence package

The key point: nightly regression is not “the only test” — it is the middle layer of the tiered pyramid. Smoke catches 90% of the cheap mistakes (fast feedback); the nightly backstops the rest (full coverage). The two layers don’t overlap in duty.

ADAS domain-controller software actually hits several rows of that table at once: embedded (scarce rigs) + safety-critical (evidence requirements) + scenario explosion (the scenario library) + AI components (perception-model drift). Which is why the ADAS industry’s test architecture is uniformly the three-layer “MR smoke → nightly SIL full → pre-release HIL” structure.

8. Global Teams: The Planet Has No Shared “Night”

Global development breaks the assumption hidden inside the phrase “nightly regression”: the planet has no common night. 2 a.m. in Shanghai is 8 p.m. in Munich (they’re still committing) and 11 a.m. in California (they just sat down). So the first thing a global team must internalize is a cognitive shift:

The essence of “nightly” isn’t “runs at night” — it’s “snapshot regression on a fixed cadence.” Coordination shifts from “what time does it run” to “which snapshot does it run, who reads the results, who owns the resources.”

Three mainstream coordination patterns

Pattern How it works Who it suits The cost
Single anchor One fixed reference time (e.g. UTC 02:00) for everyone Teams centered on one time zone, with small overseas groups Someone’s workday always starts with “half-finished” results; unfair to balanced global teams
Regional windows Each region cuts a snapshot at its end of day (Asia EOD, Europe EOD, Americas EOD — 2–3 rounds a day) Large balanced multi-site teams, especially with rig constraints Scheduling and result-management complexity doubles
Continuous regression Abolish “night”: merge queue + run on arrival, regression streamed by capacity Resource-rich teams with mature automation Requires a solid merge queue and fast smoke first

The automotive industry’s practical answer is usually regional windows — because HIL rigs are physical assets: a rig runs the night of whichever time zone it sits in (Shanghai rigs run Shanghai nights, Munich rigs run Munich nights), with data flowing into a unified dashboard. Pure SIL/GPU resources, on the other hand, can follow the sun around the clock — global distribution becomes a utilization advantage: the same GPU farm serves Asia after hours, then Europe, then the Americas — 24 hours without idling.

Five mechanisms that make all three patterns work

1. Cut by commit, not by clock (the most important) “Tonight’s build” must be a concrete git SHA (e.g. “the merge point of all MRs that went green before 18:00”), not “whatever happens to be on main when the run starts.” Otherwise a commit a German colleague pushes at 21:00 slips into Asia’s regression, and the results get pinned on the wrong change. Regression reports are always nailed to a commit — which is why a well-built test platform makes code version and artifact hash first-class citizens of the run manifest.

2. A merge queue (merge train) keeps trunk green With merges landing around the clock, trunk gets broken constantly. A merge queue makes every MR line up and pass the gate before merging, so trunk is green at any moment — and regression can safely “take a snapshot anytime.” This is the bridge from “windowed” to “continuous” regression.

3. Automatic attribution and automatic disposition When the nightly run fails, the failure cannot wait 8 hours for its owner to wake up. Mature practice:

4. A follow-the-sun handoff for results Define a cross-time-zone relay SLA: Asia’s on-call does the first triage in the morning; failures belonging to European modules get @’d to their European owners, so Europe starts the day not with “a wall of red” but with “preliminarily attributed, awaiting your confirmation.” The dashboard is the single source of truth — no failure conclusions over email or IM.

5. Globally unified flaky governance The cancer that time zones amplify most is the flaky test: Asia marks a case “known flaky,” Europe doesn’t know, and burns two more hours on it. You need a globally unified quarantine regime (N consecutive unstable runs → out of the trunk suite, run separately, fix by a deadline).

A typical global setup

The classic global-OEM configuration: a China team + European HQ + a North American software center. The usual combination:


Looking back: a nightly full run is standard issue on SIL and a luxury on HIL — the industry’s answer is layering: push the full run down to SIL, let HIL filter subsets by tags, and let global teams split “night” into regional windows. And global coordination, which looks like a time-zone problem, is really three things — snapshot governance, failure ownership, and resource scheduling: time zones only decide when the window opens; the other 90% of the work is making sure that anyone, in any time zone, can open the dashboard in the morning and answer three questions — which version ran? Who broke it? Whose turn is it to fix?


Appendix: Glossary (in order of appearance)

Term Plain explanation
HIL (Hardware-in-the-Loop) A test environment where a real ECU is wired to a simulation rig; closest to the vehicle, but rigs are expensive and few
SIL (Software-in-the-Loop) A test environment where the controller software runs fully simulated on a PC/server; pure software, massively parallelizable
CI (Continuous Integration) The practice of automatically building and testing every commit
OEM The vehicle maker at the top of the automotive supply chain
ECU Electronic Control Unit — the embedded controllers in a car
AEB Autonomous Emergency Braking, a canonical ADAS safety feature often used as the example test scenario
soak test A stability test lasting hours or more, built to catch memory leaks and long-run degradation
Jepsen A well-known consistency-verification framework for distributed systems, built to catch “consistent in theory, not in practice”
EDA Electronic Design Automation — the umbrella term for chip-design and simulation toolchains
E2E (end-to-end test) A test that walks the entire chain from user entry to exit
MR Merge Request; pipeline gates usually hang on it
merge queue / merge train A mechanism where MRs line up and pass the gate before merging, keeping trunk green at all times
follow-the-sun A global collaboration pattern where work is relayed across time zones so resources never idle
quality gate An automated checkpoint on the pipeline: fail it and you don’t merge
bisect Binary-searching the commit history to find the commit that introduced a problem
CODEOWNERS A repo file declaring who owns which code; automation routes issues by it
triage The first-pass sorting of failures: whose module, what kind, who it goes to
flaky test A case that passes and fails with no code change

Next step:

View all notes