Concepts¶
This page defines the model’s entities and how the traceability matrix turns a
model plus test evidence into verdicts. The rules here are exactly what
rules_requirements.trace implements.
Entities¶
Kind |
Default prefix |
What it is |
Question the report answers |
|---|---|---|---|
User need |
|
What a user must be able to do. |
Is it validated? |
Requirement |
|
A verifiable statement the product must meet. |
Is it verified, at the rigor it demands? |
Risk |
|
A hazard → hazardous situation → harm chain, with estimated severity and likelihood. |
Is it mitigated? |
Mitigation |
|
A risk control measure. |
Is its implementation verified? |
Test method |
|
A named verification procedure at a rigor level. |
Which requirements rely on it, and do they pass? |
Ids are <prefix>-<number> by default (REQ-7, REQ-0007); prefixes and the
id pattern are configurable per project, so an existing PR-12 scheme can be
kept (The config: section). Ids are unique across all kinds.
Every entity also carries optional description, status, owner, tags
and notes. Notes are how gaps found in review — by a person or an agent —
are attached to the object they concern; open notes appear in the reports.
References¶
References always point from the more specific object to the more general one, so each trace has exactly one source of truth:
Field |
From → to |
Meaning |
|---|---|---|
|
requirement → user need |
The requirement is (part of) how the need is met. |
|
requirement → requirement |
Decomposition, e.g. system → software requirement. One parent per requirement: refines forms a tree ( |
|
requirement → test method or level |
The verification rigor the requirement demands. |
|
requirement, mitigation → test cases |
The cases that verify it (claims, see below); a user need’s are |
|
mitigation → risk |
The control acts on this risk. |
|
mitigation → requirement |
The requirements that realise the control. A requirement that implements a mitigation has no other parent ( |
|
risk → mitigation |
Optional back-reference; if given it must agree with |
The reverse views (“which requirements satisfy UN-1?”, “which risks does
REQ-4’s mitigation control?”) are computed. A requirement that implements a
mitigation needs no satisfies — it traces through the risk instead.
Verification levels¶
A level is a rung of verification rigor. The default ladder is:
Level |
Rank |
Meaning |
|---|---|---|
|
1 |
Static argument: derivation, review of a proof, static analysis. |
|
2 |
Host-side unit or simulation test. |
|
3 |
Software-in-the-loop: the integrated software against simulated I/O. |
|
4 |
Hardware-in-the-loop: a component on real hardware. |
|
5 |
Full system on real hardware, end to end. |
|
— |
Manual or visual sign-off recorded as evidence. |
Ranked levels are ordered: evidence at hitl satisfies a hil demand, but
simulation evidence does not. Unordered levels such as inspection are
incomparable: an unordered demand is met only by evidence at exactly that
level, and unordered evidence never meets a ranked demand.
A requirement’s method names either a level directly (method: hil) or a
test method (method: TM-2), whose level then applies.
Without a method the requirement demands default_level (simulation).
Evidence names the level it provides (a level property on the test case);
evidence without one provides default_provided_level (simulation). Both
defaults, and the ladder itself, are configurable.
Test methods¶
A test method gives a verification procedure a name, a level and a description
of how it is carried out (procedure). Pointing several requirements at the
same test method keeps their demanded rigor consistent and lets the report list
which requirements depend on which procedure. A test method’s own status is the
rollup of the requirements that use it.
Evidence and attribution¶
Evidence is a set of test cases, each with a status (passed, failed,
error, skipped), the level it provides, optionally the identity of the
artifact it exercised, and the ids it declares — its tags. Every case is
filed under a key, <target>#<path> (Case keys): retries, repeated
runs, shards and several evidence roots of one case merge into one result.
A test case verifies at most one requirement; a set of test cases may
together verify one (One test case, one requirement explains the rule
and why it cannot be bypassed). Which entity a case verifies — its owner —
is decided in one place, rules_requirements.attribution.attribute(),
from two inputs:
- Claims (the model)
verified_byon requirements and mitigations andvalidated_byon user needs name the cases of a target an entity claims — per case, by selector, or the whole target (Claims: which test cases verify an entity). A case selected by exactly one entity’s claims is owned by that entity.- Tags (the evidence)
A
requirementproperty on a JUnit<testcase>(written by the hooks), a record’srequirement, an[rr:ID]name tag. Withconfig.attribution: hybrid(the default) a single tag owns a case that no claim covers; withmodela tag never owns anything and only cross-checks the claims (tag-mismatch,unclaimed-tag).
Ambiguity fails closed: a case is quarantined when its evidence names more
than one id (multi-tag), when claims of more than one entity select it
(attribution-conflict, also a static shared-case error), or when the same
test code — the same source file and case path, or targets declared in
config.variants — is owned by different entities in different targets
(same-code-multiple-owners; a source file is compared in one spelling, so
./x.py, a/../x.py and the absolute path a harness started outside the
workspace records are one file). Equal case paths with different owners whose
source is unknown or recorded differently get the same-path-multiple-owners
warning — or, when one of them is an absolute path outside the workspace that
ends with two recorded relative paths (it may be either file), the
ambiguous-source error: whether two owners share test code is then unknown,
so it fails closed. A quarantined case owns nothing, and every
entity it names reads INVALID until it has one owner. A tag naming a risk or a
test method is misdirected-evidence and one naming an undefined id
unknown-id; neither owns anything.
A result about a whole target run — an exit status after passing cases, a load
error, a report that cannot be read (rr.scope=target) — is never a case of
anyone. It taints the target: every member claimed on it reads error, so
each requirement fails through its own members.
Verdicts¶
Verification sets¶
Each requirement, user need and mitigation has one verification set:
the cases it owns (through its claims, or its tag in hybrid mode);
the cases it expects: each literal selector’s case, and each entry of the verification-set lock (The verification-set lock) naming it;
a member for each selector that matched nothing, and for each quarantined case that names it.
A member’s state is the result of its case (passed, failed, error,
skipped), or: missing (its target ran without it — a renamed or deleted
test, a filter), not-run (no evidence for its target at all — another lane,
a target that never built), moved (a lock entry whose case now has another
owner or none) or quarantined. A member’s level is its case’s own level,
else its claim’s, else default_provided_level; when the case and the claim
disagree the lower one counts (level-mismatch), and when either level is
unordered (inspection, say: it has no rank) the case’s own level counts, so
a selector can never lend a case a level it did not provide.
Requirements¶
A requirement’s verdict is the first that applies to its set:
Any member quarantined → INVALID.
Any member failed or errored (a taint included), or a member passed only on a retry under
config.flaky: fail→ FAILED.No members at all, or none of them ran → UNVERIFIED.
Any member missing, not run, skipped or moved (or the set mixes builds under
config.set_consistency: enforce) → INCOMPLETE.The whole set passed, but a member is stale, or passed only on a retry (
config.flaky: under-verify, the default), or the best level of the set is below the demanded one → UNDER-VERIFIED.Otherwise → VERIFIED: the whole set passed together, at the demanded rigor; it provides the best level among its members, the cheaper members being the pyramid’s base.
For an unordered demand (inspection) the set must include a passing member at
exactly that level. config.flaky decides what a retry-masked pass is worth:
accept (a pass), flag (a pass and a flaky gap), under-verify or
fail. Only an earlier failed or errored attempt makes a member flaky.
Refinement. When other requirements refine a requirement, their verdicts
roll up into it — verdicts, never cases: a parent’s set holds only its own
claims. Each verdict says what it rests on: basis is own (its own set),
derived (other entities’ verdicts, listed in derived_from) or
own+derived.
the children’s rollup is FAILED if any child failed or is INVALID, VERIFIED if all are verified, PARTIAL if at least one is verified, under-verified, partial or incomplete, and UNVERIFIED otherwise;
the parent is INVALID if its own set is, and FAILED if its own set or any child failed;
a parent with claims of its own needs its own set complete too: while it is incomplete (or did not run), the parent is INCOMPLETE;
when every child is VERIFIED, the parent is VERIFIED only if its own demand is met — by its own set, or because every child’s best evidence is at least as rigorous as the parent demands (a stale or flaky own set is not made good by the children). A
hitlsystem requirement is not proven by simulation-verified software requirements: it stays UNDER-VERIFIED until system-level evidence athitlexists;otherwise a parent whose own set passed, or with partly verified children, is PARTIAL; a parent with neither stays UNVERIFIED.
Rollups¶
Entity |
Rolls up |
Verdicts |
|---|---|---|
User need |
requirements that |
VALIDATED · PARTIAL · FAILED · INVALID · UNVALIDATED |
Mitigation |
requirements it is |
VERIFIED · PARTIAL · FAILED · INVALID · UNVERIFIED |
Risk |
mitigations that |
MITIGATED · PARTIAL · FAILED · OPEN |
Test method |
requirements whose |
VERIFIED · PARTIAL · FAILED · UNVERIFIED |
Module |
requirements listing it in |
VERIFIED · PARTIAL · FAILED · UNVERIFIED |
Every rollup uses the same rule: no children → the “none” verdict (UNVALIDATED, UNVERIFIED, OPEN); any child FAILED or INVALID → FAILED; every child fully verified → the “all good” verdict; at least one child verified, under-verified, partial or incomplete → PARTIAL; otherwise the “none” verdict. An under-verified or incomplete requirement therefore makes its user need and its mitigation PARTIAL, never VALIDATED or VERIFIED, and an INVALID one makes them FAILED. A user need or mitigation that is itself named by a quarantined case is INVALID.
Direct validation evidence¶
A user need’s validated_by and a mitigation’s verified_by claim cases of
their own — a usability study that validates UN-2, say (in hybrid mode a
case tagged with the need’s id works too). Such a set is not graded by level
(any pass counts) and joins the rollup as one more child: a failing usability
study makes the need FAILED even if every requirement is verified. A user
need’s or mitigation’s verdict is therefore always a rollup: an own set that
is INCOMPLETE (a skipped or missing member) or UNDER-VERIFIED makes it
PARTIAL, never INCOMPLETE, while the set’s own incomplete gap is still
raised; only INVALID (a quarantined case names it) is carried through as is.
Staleness¶
“It passed once” and “it passes for the build we are shipping” are different
claims. Evidence can record the identity of what it exercised as
artifact.<key> properties — a firmware build id, a board revision, a DUT git
SHA. When a report is built with the current identity
(--current-build KEY=VALUE, or current_build in Bazel), a passing case
whose recorded identity differs on any key both sides have is stale.
Evidence without an identity is never stale.
A verification set must hold on the current build as a whole: a set that passed
but has any stale member is UNDER-VERIFIED and flagged stale (the reports show
a STALE badge), and the gap queue carries a stale item for it. Passed
members stamped with different values of one key (two dut_git_sha values in
one set) are mixed builds: a mixed-builds gap with
config.set_consistency: warn (the default), INCOMPLETE with enforce.
The cost pyramid¶
Physical verification is expensive; cheap verification should back it rather
than be skipped. A requirement that demands at least pyramid_min_level
(default hil) and has passing physical evidence (at pyramid_min_level or
above), none of which is backed by evidence at one of the
pyramid_cheap_levels (default analysis, simulation), is a cost-pyramid
violation: a physical result nobody sanity-checked cheaply. (Evidence below
the demand without any physical result is simply under-verified.) Violations are
listed in the reports and as pyramid gaps; rr report --pyramid-policy
decides whether they only warn (default), fail the command, or are ignored.
High-severity risks¶
A risk whose severity is one of high_severities (default high,
critical) and whose verdict is anything but MITIGATED is hoisted into a banner
at the top of the reports and becomes a high-risk-open gap. Mitigation is not
elimination: a high risk controlled only by an under-verified requirement is a
headline, not a checkmark. The check uses the risk’s initial severity, not its
residual severity.
Gaps and routing¶
The matrix ends with a list of gaps — everything between the model and a
complete verification and validation argument — which rr report --queue-out
writes as a machine-readable work queue:
Gap kind |
Raised for |
Route |
|---|---|---|
|
a quarantined case (one gap per case, naming every claim’s origin and the ids its evidence declares); first in the queue |
autonomous |
|
an entity a quarantined case names |
autonomous |
|
a requirement with a failed or errored member (also a user need or mitigation whose own set has one), or failing refinements |
by demanded level — reproducing a bench failure needs the bench |
|
a requirement whose set is incomplete: |
human-gate if a member that did not run or was skipped is above |
|
a selector or lock entry whose case its target did not report, with the nearest case it did report |
autonomous |
|
a requirement whose set passed with a stale member |
by demanded level |
|
a requirement whose set passed with a retry-masked member ( |
by demanded level |
|
a set whose passed members were stamped with different builds ( |
by demanded level |
|
a requirement with no members ( |
by demanded level; by the members’ levels when they did not run |
|
a requirement whose set is below its demand |
by demanded level |
|
a requirement whose refinements are only partly verified |
by demanded level |
|
a cost-pyramid violation |
autonomous |
|
a requirement (without refinements) that no source annotation implements — only when sources were scanned |
autonomous |
|
evidence tagged with a risk or test-method id (they are not verified by tests — tag the requirement) |
autonomous |
|
a failing case no entity owns, or a target-scope failure that affects no member ( |
autonomous |
|
a tag that disagrees with the claim owning its case, or ( |
autonomous |
|
one case key reported twice in a run; a whole-target claim on a target with per-case results |
autonomous |
|
the lock disagrees with the attribution (The verification-set lock) |
autonomous |
|
equal case paths in two targets with different owners and no common recorded source ( |
autonomous |
|
no lock: the entities whose sets have glob, whole-target or tag-owned members |
autonomous |
|
a high-severity risk that is not MITIGATED |
human-gate |
|
evidence tagged with an id the model does not define |
autonomous |
|
( |
autonomous |
|
an open note of that kind on any entity (questions route to a human) |
by demanded level |
Every attribution issue is one of these gaps, warnings included, so
rr report --fail-on gaps fails on any of them (Attribution issues are gaps lists
each one and what to do about it).
Routing splits the work between agents and people. A gap whose
requirement demands a level at or below autonomous_max_level (default sil)
is autonomous: an agent can write the missing unit, simulation or
software-in-the-loop test itself. Anything demanding more — hil, hitl, or an
unordered level such as inspection — is a human-gate: it needs a bench, a
device or a signature, and an agent may prepare the harness but must not
synthesise the evidence.
What is not modelled¶
Benefit–risk analysis (ISO 14971 §7.4) and the overall residual risk judgement (§8) are decisions, not computations; record them in the risk’s
residualnote and in your risk management file.Software safety classification (IEC 62304 §4.3) is not a field; record it in
project:metadata and express the rigor it implies through levels and test methods.The tool records and checks traceability. Whether a test actually proves a requirement is a review question — which the gap queue and notes help you ask, but do not answer.