Migrating to one requirement per test case¶
The path from 0.2 to 0.3
This guide takes a project from 0.2’s union rule to 0.3’s model mode in
steps that each keep CI green (a CI that runs --strict, rr_model(strict = True) or rr_report(strict = True), or gates on --fail-on gaps, needs the
extra steps the pin bump below lists): the first half runs on 0.2 (rr cases,
rr migrate plan, rr migrate apply --stage tags, and since v0.2.1
rr migrate verify, none of which changes a verdict); the second pins 0.3
in hybrid mode, writes the claims with rr migrate apply --stage model and
locks the sets. One test case, one requirement explains the rule
itself.
A test case should verify at most one requirement. A set of test cases may verify one requirement together, but a single test that counts toward two requirements makes both verdicts depend on one result, and a reviewer cannot tell which requirement the test was written for.
Today a case counts toward every id it is tagged with, plus every requirement
whose verified_by names its target. That is the union rule. Three common
patterns make one test count twice:
a module-level
pytestmark = pytest.mark.requirements("PR-13", "PR-29"), or any marker or@rr.verifies(...)that names two ids;ids that accumulate across scopes, for example a module marker naming PR-11 and a function marker naming PR-13 (from v0.3.0 the hooks record only the nearest scope’s id, here PR-13);
one target listed in the
verified_byof two requirements, or one target whose cases are tagged for one requirement and listed whole by another.
From v0.3.0 such evidence is quarantined (it counts for nobody) and a target claimed by two requirements is a model error. Migrating on v0.2 first keeps CI green throughout: every step of the first half is valid under 0.2’s rules, so the 0.3 pin bump only has to fix what 0.2 cannot express.
The steps¶
Collect evidence. Run the tests that feed your traceability report (
bazel test //..., plus HITL runs if you have them) and keepbazel-testlogs, includingtest_attempts/.Plan.
rr migrate planwrites a worksheet listing every evidence unit that counts toward two or more entities.Decide. The owners of the requirements involved record one owner per case on the worksheet.
Rewrite the tags.
rr migrate apply --stage tagsrewrites Python test markers to the decided owners.Edit the model. Remove the
verified_byreferences the decisions leave unused; the tool lists them.Check. Re-run the tests, check their evidence against the worksheet with
rr migrate verify, and re-run the plan. What is left is the work for v0.3.0’s case selectors.
On 0.3:
Pin 0.3 in hybrid mode. Bump the pin. Turn every target two entities still share into per-case selectors, so the model is valid; keep
attribution: hybrid(the default), so the single-id tags keep owning their cases. Confirm withrr attribution --checkthat nothing is quarantined before merging.Move to model mode.
rr migrate apply --stage modelwrites an explicit selector for every case each entity owns today and proves the owner table unchanged; then setconfig.attribution: model.Lock.
rr sets lock --writepins every set’s members; gate CI onrr sets checkandrr check-report.
Case keys¶
Every decision is made about a case key: <target>#<path>, for example
//pi/server:server_test#pi.server.tests.test_handler::test_configure. A
target that produced only Bazel’s generated report has one case, [target].
See Case keys for the exact rules. To list the keys in your evidence:
$ rr cases --evidence bazel-testlogs
$ rr cases --evidence bazel-testlogs --target //web:clocksync_test --json
Planning: rr migrate plan¶
$ rr migrate plan --model requirements/ --evidence bazel-testlogs hitl-testlogs \
--out requirements/attribution.rrplan --md attribution.md
178 of 271 attributed evidence unit(s) count toward two or more entities; 7 target(s) shared between requirements; 0 decided, 178 open
--out writes the worksheet as YAML (a .rrplan file). --json writes the
same document as JSON, and --md writes a Markdown rendering for review. With
no output flag the YAML goes to stdout. Use an extension other than
.yaml/.yml (such as .rrplan): when --model names a directory, every
YAML file in it is read as part of the model, and the worksheet must stay out.
The worksheet groups cases by target and test module or class:
schema: rules_requirements/attribution-worksheet/v1
summary: {units: 489, attributed: 271, contested_units: 178, shared_targets: 7, decided: 0, open: 178}
targets:
- target: //web:clocksync_test
claimed_by: [PR-13, PR-29] # verified_by of both
shared: true
cases: 1
per_case_output: false # only Bazel's generated report: one [target] case
groups:
- target: //pi/server:server_test
group: pi.server.tests.test_handler
counts_toward: [PR-11, PR-13]
tags: [PR-11, PR-13] # where the ids come from: tags ...
owner: '?' # <- the decision for every case below
cases:
- path: pi.server.tests.test_handler::test_configure_renegotiates_mid_capture
status: passed
- path: pi.server.tests.test_handler::test_resumes_session
status: passed
owner: PR-22 # <- a per-case override
- target: //web:clocksync_test
group: ''
counts_toward: [PR-13, PR-29]
target_claims: [PR-13, PR-29] # ... or verified_by
owner: '?'
cases:
- path: '[target]'
status: passed
Field |
Meaning |
|---|---|
|
Every target named in |
|
The entities the cases count toward today. Only requirements, user needs and mitigations count; tags naming risks, test methods or unknown ids are listed per case as |
|
Whether those ids come from the tests’ tags, from |
|
A hint where the model or the evidence determines an owner. A case tagged with a requirement and its refinement proposes the refinement, because the parent’s verdict rolls up from it. A case with its own single tag inside a target claimed whole by another requirement proposes the tag. Everything else is left open. A proposal is never a decision. |
|
The decision: one of the ids the case counts toward, |
|
Whole-target results, such as an |
An --evidence path that holds no evidence file is reported as a warning
([no-evidence]), and no evidence at all stops the plan (exit status 2). A
target named in verified_by that the evidence has no result of — a HITL-only
or manual target planned from software evidence, say — is listed under
Targets without evidence: nothing is decided about it, and its references
are not unused.
The JSON Schema is schema/worksheet.schema.json
(https://studio-fug.github.io/rules_requirements/schema/worksheet.schema.json).
Evidence changes while decisions are made. Re-run the plan with
--merge requirements/attribution.rrplan to keep every decision whose case
is still contested.
Deciding¶
Decide per group first, then override per case. Some rules of thumb:
The owner is the requirement whose acceptance criterion the test asserts, not every requirement the code under test helps with. A protocol round-trip test verifies the protocol requirement, not the feature that uses the protocol.
noneis a valid answer. Helper and plumbing tests often verify no requirement. They still run and still gate CI; they just do not count as evidence.A requirement with no test of its own is a gap, not a reason to share. If deciding leaves a requirement without evidence, the honest result is UNVERIFIED, and the remedy is a test written for it.
Parametrized cases can go to different owners. The codemod cannot split one test function between owners; split those by hand (see below).
Shared whole targets (
per_case_output: false) can only be given to one owner as they are. To split them, give the target per-case output first (rr_node_testfor node:test,rr_case.hfor plain-assert C/C++,JUnitWriter/CheckPlanfor scripts), re-run the plan, and decide per case.Hardware phases that check several things in one case should become one case per check (
CheckPlan), each with one owner.
The PR owner of each requirement decides. Agents and heuristics only propose.
The collection check: the guarantee¶
A rewrite’s promise is that each test keeps collecting exactly the ids it
should — every decided case ends with its owner, and nothing else moves. No
reading of the syntax tree can prove that: Python collects tests pytest never
sees in the source — installed by a factory, a metaclass, __init_subclass__,
setattr, exec, or an import call. So, by default, rr migrate apply
proves the rewrite against pytest itself.
After the rewrite passes the static guards (below) in memory, and whenever it
would write anything (with --partial and --dry-run too), apply:
copies the tree under
--rootto a temporary directory and writes the rewritten files there, never in place. It skips version control and cache directories, virtualenvs (a directory holdingpyvenv.cfg) and Bazel’sbazel-*convenience symlinks, and passes an--ignorefor each skipped path to both collections, so the two trees are collected over the same files. Other symlinks are never skipped silently, because a real pytest run from--rootwould follow them: a symlink whose target resolves inside the tree is recreated in the copy (so both collections follow it and the check covers whatever it reaches, refusing if its attribution changes); a symlink to any other file outside the tree (a sharedpytest.ini,pyproject.toml,setup.cfgortox.ini, a data file) is copied with its content, so both collections run under the same config; and a symlink that reaches a test outside the tree — a directory, or a.pyfile — makes apply refuse and name the path (the copy cannot cover it) whenever the original collection shows pytest would reach it. A link pytest never recurses into is left alone: one outside the paths it collects (its arguments, elsetestpaths) or under a directorynorecursedirsexcludes (.*by default, so.direnv’s flake-input links). Remove or redirect a link that refuses, or pass--no-collect-check;runs
python -m pytest --collect-onlyin the original tree and in the copy, from--rootwith no path argument (so the project’s ini,testpathsincluded, decides what is collected, exactly as a plainpytestrun there would), with a tiny plugin that records, for every collected item, its nodeid and the ids the JUnit hook would give it. Neither run leaves__pycache__in your tree;compares: each decided case the written files settle must end with exactly its owner (no id for
none), every other collected test must keep exactly the ids it had, and the set of collected nodeids must not change. Decided cases the write does not settle — cases left toverified_by(their test declares no id), cases in files held back or refused, cases outside--only— are held to their before-ids like undecided ones.
Any difference — including a settled decided case that no collected test
matches — makes apply refuse, write nothing, and name each offending item
with its before, after and expected ids (--dry-run exits 1 and reports the
files as held back). This is what catches the dynamic shapes the static guards
cannot: a setattr-installed method that would silently lose a shared marker,
a factory-built subclass, an aliased test. If either collection records the
same nodeid twice (a conftest that builds items by hand), apply refuses and
names it: the check cannot tell the two cases apart.
Apply never writes through a symlink: just before writing, it re-checks
each file it would rewrite, and if one is (or lies under) a symlink — even one
swapped in after the check ran — it refuses, names the path and where it
resolves, and writes nothing. This holds with --no-collect-check too.
Run apply where python -m pytest --collect-only works for the project — the
same interpreter and dependencies the tests need. For a plain project that is
its virtualenv; for a Bazel project, make a virtualenv with the test
dependencies (the same ones the py_test targets use) and run apply from the
source tree. Where that is impractical — py_test targets that import
through their runfiles, generated code, toolchain-provided modules — apply
with --no-collect-check (and --trust-main-guard only if needed), push,
then check the rewrite against the CI test evidence with rr migrate verify
before merging (see Verifying with test evidence: rr migrate verify). If collection fails in either tree (an import error, a non-zero
exit, no tests found while the worksheet has decided cases), apply refuses and
prints the collection output, so a broken environment never passes silently.
The guarantee covers what collects in the check’s environment — the
--python interpreter, its installed packages and the current environment
variables. A test that only exists elsewhere (a class built only when HW=1 is
set, a module that imports a hardware library only the bench has) is not
collected by the check, so it cannot be verified. Collectors pytest skips at
collection time (pytest.importorskip, a module-level pytest.skip) are listed
in a warning for that reason. Run apply in the environment your tests really
run in (set the same variables, install the same packages), or check those
tests by hand.
--python PATHpicks the interpreter whose pytest and project dependencies collect the tests (default: the interpreter runningrr). Point it at the project’s venv whenrrruns from another; onlyrules_requirementsitself is added to itsPYTHONPATH, never the rest ofrr’s own environment.--pytest-args "..."passes extra arguments to the check’s pytest, one shell-quoted string, for a project that needs them to collect:-c pytest.ini,--rootdir .,-p my_project.pluginfor a plugin the tests rely on,--ignore=scriptsfor a directory that does not import here, or a path to collect beyondtestpaths. A value that starts with a dash works either way:--pytest-args "--ignore=x"or--pytest-args=--ignore=x. Any plugin your runner loads with-pmust be repeated here (for example--pytest-args "-p my_project.plugin"): a plugin that mutates markers at collection time changes what a real run attributes, and the check is blind to it unless it loads the same plugin. (Onlyrr migrate applytakes--pytest-args; it is passed through to the check’s pytest unchanged.) The check’s plugin registers rr’srrmarkers itself, so--strict-markerscollects even when rr’s pytest plugin is loaded with-ponly in the real run (as the Bazel runner does).--no-collect-checkskips the check and writes on the static guards alone. It prints a loud warning: the guards are best-effort and cannot see what pytest collects dynamically, so a rewrite they accept can still move a shared marker off a test you did not mean to touch. It also gives up everything the check adds: the before/after collection comparison, the refusal on symlinks that reach tests outside the tree, and the refusal on duplicate nodeids. Use it only where pytest cannot collect the project at all, and check the result against real test evidence afterwards withrr migrate verify(Verifying with test evidence: rr migrate verify).--trust-main-guard(off by default) leaves the code only anif __name__ == "__main__":block runs out of the static guards (below). That exclusion is best-effort, so only use it together with a definitive check: the collection check, orrr migrate verifyagainst fresh test evidence before merging. apply prints a warning saying so.
The static guards: a first, conservative line¶
The guards below run first, entirely in memory. They are deliberately
conservative — they refuse anything they cannot read or follow at import or
class-creation time (module and class bodies, decorators, metaclasses,
__init_subclass__, __new__, and anything called at module or class scope),
because that is what decides collection. They no longer refuse code in the
bodies of tests, fixtures (@pytest.fixture), setUp / tearDown /
setUpClass / setup_method-style hooks: that code runs after collection
and cannot change it, and the collection check vouches for the result either
way. A function only counts as such a hook while nothing that may run at
import time refers to it: a setUp or setup_module the module calls itself
(directly, through a helper, getattr, or as a decorator) is judged like any
other import-time code.
By default they also judge the code only an if __name__ == "__main__":
block runs, like any other code: static analysis cannot prove what Python
runs at import, and the guards fail closed. A script-style test whose
main() loads a helper with importlib.util.spec_from_file_location is
therefore refused, and with it every file that import call may import.
--trust-main-guard (opt-in) leaves out the body of an
if __name__ == "__main__": block (either operand order, either quote; not
its else) and the bodies of module-level functions reachable only from
it: when pytest itself imports a test module, it imports it under its module
name, so that code does not run at collection. The exclusion is
best-effort. It follows the names it can see, and a name computed at run
time can get past it (a builtins alias, __getattribute__,
runpy.run_module(..., run_name="__main__") or a "__main__" module spec in
another file), so only use it together with a definitive check: the
collection check, or rr migrate verify against fresh test evidence before
merging. With the flag, such a function is still judged as usual when anything that may run
at import time names it (a module-level call, an alias, a getattr /
globals() string, a test, a default argument), when it is decorated,
rebound or named like a test or a hook, or when another file under --root
imports it, passes its module around, reads it from sys.modules, or names
it on a pytest item’s .module / .obj. Nothing in the file is excluded —
the block included — when the file may rebind __name__, when code that may
run at import looks names up dynamically (globals(), vars(), eval,
getattr with a computed name, inspect.getmembers, __dict__,
sys.modules, a frame’s globals), or when the block launches the tests
in-process (pytest.main(), unittest.main()): a Bazel py_test whose
main is the test file runs the block, and whatever it runs before the
launch runs before collection, in the same interpreter. Nor is anything
excluded anywhere when another file may reach any module’s functions
without importing it (a computed sys.modules read, a frame, eval /
exec, a pytest_pycollect_makeitem hook, or a computed lookup on a pytest
item’s module). A file is refused and left unchanged, never
half-migrated, when:
a test in it names several ids and has no decision (
?, or a test the evidence never ran). With--unassigned dropsuch tests lose their tags instead.the parametrizations of one test were decided differently. Split it by hand:
@pytest.mark.parametrize( "mode", [ pytest.param("fast", marks=pytest.mark.requirements("PR-1")), pytest.param("slow", marks=pytest.mark.requirements("PR-3")), ], )
a declaration is not something it can read statically: ids, levels or artifacts that are not literals,
**kwargs, a marker insidepytest.param(...), or apytestmarkthat shares its line with other code.a decided test, or a class or module around it, carries a decorator or a
pytestmarkelement that may hold ids the codemod cannot see: a variable holding a marker (AB = pytest.mark.rr("A", "B"), then@AB), an alias such asm = pytest.markwith@m.rr(...), a helper’s decorator (@tagged("A"), or any decorator from a third-party library), apytest.param(..., marks=AB), or apytestmarkthat is augmented, annotated, assigned twice or bound inside a block. What it reads: therr/requirements/verifiesdeclarations, any otherpytest.mark.NAMEmarker, the rest ofpytest(pytest.fixture),staticmethod/classmethod/property, and whatunittestandmockprovide (mock.patch).one scope declares two different levels.
a change would reach a test the codemod cannot see: a test class that inherits tests or declarations from another class, in the file or in another module; a test defined inside an
if,try,withor loop block; a decided case of the module that the file does not define (inherited from another module, or generated); or a test name bound other than by adeforclassstatement, whether or not it is on the worksheet. That covers every way of binding a name in a module or class body: assignments (tuple, starred, annotated and augmented ones too,test_b = test_a),forandwith ... astargets,:=anywhere in the scope,from m import ...(alsoas, and star imports, which may bring tests),del,global,except ... as,matchcaptures,Class.test_x = ...,globals()["test_x"] = ...andsetattr(..., "test_x", ...), andexec/eval/globals()/setattrwith a computed name at the top of a module or class. Bindings from inside a function body, which may run at import time, count too, and then nothing in the file is changed:global test_x; an attribute store,setattr/delattror__setattr__naming a test or with a computed name, on anything but a method’s ownself(TestK.test_x = ...,sys.modules[__name__].test_x = ...); any use ofglobals(),vars(x)orx.__dict__(exceptself.__dict__);execandeval. A class deriving from another may be aunittest.TestCase, which pytest collects whatever its name, soAlias = Checkscounts as a test name bound by an assignment too. A name is a test name when pytest’spython_functionsorpython_classesmatch it (testandTestby default; the codemod also reads the patterns configured inpytest.ini,pyproject.toml,tox.iniorsetup.cfgunder--root). A literal (test_cases = [...]) and a plainimport m as test_m(a module) are not tests. When what such a name holds cannot be traced to adefof the file (x = test_a, thentest_b = x), nothing in the file is changed; otherwise nothing that reaches it is. Migrate the file by hand, or rename the name if it does not hold a test.another file imports or subclasses a class or test whose declarations would change. pytest collects an imported test class or function again in the importing module, and a subclass inherits its base’s tests and markers, so the change would reach tests in that file unseen. Both files are refused. The codemod indexes every Python file under
--rootfor this, also those outside--only:from m import C(with or withoutas),import mwithm.C(alsopkg.sub.Cwith onlypkgimported), relative imports and star imports,importlib.import_module("m")and__import__("m")with a literal module name (also relative ones), and a base class it cannot trace to an import by its name. A package passed around as an object (getattr(pkg, ...),__import__("pkg")) reaches every module below it. Files are read in their own encoding (a PEP 263 coding cookie or a BOM, as Python reads them) and written back in it.a file under
--rootcannot be read or parsed, while some declaration would change: it may import or subclass anything. That file is refused by name, together with every file that would change. A file holding a decided case that cannot be read or parsed (an unknown coding cookie, or syntax newer than the Python runningrr) is refused by name too.a test module, or a file a test module imports, imports a module by a call whose module the codemod cannot name:
importlib.import_module(name)with a computed name,importlib.util.spec_from_file_location,runpy.run_path,execof a computed string or of a literal naming an import, and the like. pytest collects what it imports there, which may be any test or class, so when any declaration in the tree would change, that file and every file that would change are refused.a class has a base the codemod cannot resolve statically: a name bound by an assignment or a
def(B = importlib.import_module("m").C,B = __import__(...),try: from m import B/except: B = object) or an expression (getattr(m, "C")). It may be any class, so when the declarations of any class in the tree would change, its file and every file whose classes would change are refused, and its own decided tests are refused. This counts for classes pytest may collect (named like a test class, or in a test module:python_files,test_*.pyand*_test.pyby default), classes another class subclasses, and classes a test module imports.a multi-line declaration it would rewrite has comments inside it (they would be lost).
The command also lists, without changing anything:
decided cases whose module it found but whose test is not defined there. These make it exit 1: it cannot vouch for them.
decided cases with no Python test (googletest
RR_VERIFIES, Rustrr::verifies!,JUnitWriterharnesses, records, and whole-target[target]results). Edit those by hand so that each case names one id. With--only, cases outside the given paths are only counted.decided cases whose test declares no id in its source, every decorator and
pytestmarkelement around it being one the codemod reads. They count only throughverified_by: see the edits below.the
verified_byedits the decisions imply:remove(no case of the target is left for that requirement) orsplit(the requirement keeps some of the target’s cases while others went elsewhere). A whole-target reference cannot express a split; that waits for case selectors in v0.3.0.
Markers added by a conftest.py (item.add_marker) are invisible to the
codemod; check them by hand.
It is all or nothing: when a file is refused or a decided case’s test was
not found in its module, nothing is written, and the files it would have
rewritten are listed as held back, each with the cause. Fix those, or pass
--partial to write each rewritten file that holds no such case and is not
linked to a refused file or to one holding such a case. Two files are linked
when one star-imports the other or passes it around as a module object, or
imports, references (m.name) or subclasses a class of it, a test-named
name, or a name not defined at its top level. Importing a plain helper
function or constant does not link files: changing a file never changes what
such a name holds. A file that cannot be read, a class whose base cannot
be resolved, and an import call whose module cannot be named are linked to
the files refused with them.
It exits 0 when every file it had to change was rewritten, 1 when a file was
refused or a decided case’s test was not found in its module (with or without
--partial), and 2 when the
worksheet is unreadable or an owner is not one of the ids its case counts
toward (or, with a model, not a requirement, user need or mitigation of it).
Without --model it checks the owners against the model named in the
worksheet’s inputs, when that is still there. Files keep their line endings
and their encoding.
--only PATH (repeatable) limits the rewrite to files below a path, so several
pull requests can migrate disjoint directories in parallel from one worksheet.
Verifying with test evidence: rr migrate verify¶
The exact check of a rewrite is the evidence of a real test run. For a Bazel
project, where collecting the tests locally is impractical (py_test
targets that import through their runfiles, as most do), this is the
verification path: apply with --no-collect-check (adding
--trust-main-guard only if needed), push, then rr migrate verify against
the CI evidence before merging. --trust-main-guard is needed only when
apply refuses a file for code that only its __main__ block runs (a
script-style test whose main() imports a helper by path, say); its
exclusion is best-effort, and the verify run is what checks the result.
$ rr migrate apply requirements/attribution.rrplan --stage tags --no-collect-check
$ # only if apply refused a file for code only its __main__ block runs:
$ rr migrate apply requirements/attribution.rrplan --stage tags --no-collect-check --trust-main-guard
$ git push # CI runs the tests; download its bazel-testlogs into ci/
$ rr migrate verify --worksheet requirements/attribution.rrplan \
--evidence ci/bazel-testlogs/ --baseline bazel-testlogs/
Keep each evidence tree in a directory named bazel-testlogs (or
testlogs): a test.xml’s build target comes from its path
(bazel-testlogs/<package>/<name>/test.xml), so the same files under
ci-testlogs/ would be keyed by suite name and match no case of the
worksheet. verify exits 2 with that hint when the worksheet’s cases belong
to build targets and the evidence names none.
It checks, case by case (by case key, so pytest, node,
rr_case.h, googletest and every other evidence rr ingests alike):
every case the worksheet decides declares exactly its owner in the
--evidence(no id at all fornone). A decided case with no result there is an error. A decided case that declared no id before the rewrite and declares none after counts only throughverified_by, which apply does not edit (it lists thekeep/remove/splitedits): it is listed in a note, not an error. With--model(or the model named in the worksheet’sinputs, when that is still there) the note says whether the model’s claims select it for exactly its owner (selected as attribution selects them: a case selector reaches only the cases it matches, a whole-target claim every case of its target), or the case is still pending a model edit (a target split between owners needs case selectors:rr migrate apply --stage model); a decided case that declares exactly its owner while claims of other entities select it is noted too: from v0.3.0 a claim decides a case’s owner over its tag (tag-mismatch), and claims of two entities quarantine it. Without a model,verifychecks declared ids only. “Before” is the--baselinewhen it has the case, else the worksheet (a group with notagshas no tagged case); a tagged case that ends with no id lost its tag and is an error;with
--baseline(the evidence the worksheet was planned from, or any run from before the rewrite), every other case — undecided, still open, or not on the worksheet — declares the same ids as before (in any order), and no case of the baseline is missing from the new evidence. A target-scope result (a whole run’s exit status) carries the union of its cases’ ids, which the decisions change: it must still be there, its ids are not compared. Cases only the new evidence has are counted, not refused;--allow-missing(a HITL or manual target CI does not run) makes a case with no result — decided or from the baseline — a warning, not verified, when its target has no result file at all in the new evidence. A target with any result file ran: a case, a target-scope or Bazel synthetic result, or atest.xmlthat holds no testcase (pytest collected nothing). Outside a testlogs tree such an empty report ran the targets its cases would be keyed under:suite:<name>for each of its<testsuite>names, the file stem only for an unnamed suite or none. A case missing from a target that did run stays an error: it was renamed or lost by the rewrite.
It prints one line per offending case, with the ids expected and found, and exits 1 on any:
//pi/server:server_test#pi.server.tests.test_handler::test_configure: a decided case does not declare exactly its owner (expected [PR-11], found [PR-11, PR-13])
//pi/server:server_test#pi.server.tests.test_session::test_resume: a case of the baseline has no result in the evidence (expected [PR-13], found no result)
Otherwise it exits 0 with a summary on stderr. It exits 2 when the worksheet
cannot be read or an owner is not one of the ids its case counts toward (with
--model, or the model named in the worksheet’s inputs when that is still
there, also when it is not an entity of the model), when --evidence or
--baseline holds no evidence at all, or when the worksheet’s cases belong
to build targets and the evidence names none (above). --format and --ingestor work as for
rr cases.
Checking the result¶
Re-run the tests, verify them (above) and re-run the plan:
$ bazel test //...
$ rr migrate plan --model requirements/ --evidence bazel-testlogs --merge requirements/attribution.rrplan --md attribution.md
Everything still listed is either decided but not yet applied (a non-Python
test, a split you have to make by hand) or a split target waiting for case
selectors. rr report verdicts change honestly as tags move. A requirement
that was VERIFIED only through a shared test now shows what it really has.
On 0.3: pin the release in hybrid mode¶
0.3 changes semantics on purpose: a case whose evidence names two ids is
quarantined (it counts for nobody, every id it names reads INVALID, and
rr report exits 3), and a target that two entities claim — two
requirements listing one target in verified_by, say — is a shared-case
model error. The first half of this guide removed the multi-id tags; the pin
bump fixes the shared targets, which 0.2 cannot express per case.
Bump the pin (
bazel_dep(name = "rules_requirements", version = "0.3.0")and its override) and runrr validate requirements/. Everyshared-caseerror names the two entities, both selectors and an example case. A requirement thatrefinestwo or more parents is amulti-parent-refineserror (refines must form a tree, so each case’s evidence rolls up one chain of requirements: Derived verdicts are not ownership): keep one parent and split the child into one requirement per parent, or setconfig.rules: {multi-parent-refines: warning}until you do. A requirement that implements a mitigation and has another parent (a second mitigation, or a requirement it refines) is amulti-parent-implementserror for the same reason: merge the mitigations into one that mitigates every risk, let the parent requirement implement the mitigation, or split the requirement; or setconfig.rules: {multi-parent-implements: warning}until you do.Replace each shared whole-target reference with the cases each entity owns, as the worksheet decided them.
rr cases --evidence bazel-testlogs --target //web:clocksync_testlists the exact case paths:# before (0.2): both requirements list the whole target # - id: PR-13 # verified_by: [//web:improv_provision_test] # - id: PR-29 # verified_by: [//web:improv_provision_test] - id: PR-13 verified_by: - target: //web:improv_provision_test cases: - "improv_provision::provisionViaBle: sends the correct wifi-settings wire and returns the redirect" - "improv_provision::provisionViaBle: surfaces a device error notification as a rejection" - id: PR-29 verified_by: - target: //web:improv_provision_test cases: ["improv_provision::provisionViaBle: survives Android's first-attempt GATT flake via retry"]
Other 0.2 references (a bare label,
{target, level}) still parse as whole-target claims with abare-target-referencewarning; leave them for step 8, unless CI runs strict:--strict,rr_model(strict = True)andrr_report(strict = True)escalate that warning (and the other new 0.3 warnings:unknown-id,coarse-claim,suite-level-requirement,unscoped-evidence,same-path-multiple-owners, …) to errors. Then convert them in this change, setconfig.rules: {bare-target-reference: warning}(strict still escalates it), or drop strict until step 8. Two more effects of the bump: a whole-target claim of a target with no results in this evidence (another lane’s HIL target, say) is now an expected not-run member, so its entity reads INCOMPLETE and--fail-on unverifiedfails (pass that lane’s evidence too, or report it with--lane/--lane-targets); and--fail-on gapsfails on more gaps. Withoutconfig.sets_lockthe report carries anunpinned-setsgap (setsets_lockand runrr sets lock --write, which works in hybrid mode), and every attribution issue is a gap, warnings included, so locking alone may not make--fail-on gapspass. The attribution issues that are gaps:duplicate-case,unscoped-evidence,suite-level-requirement,coarse-claim,tag-mismatch,unclaimed-tag,misdirected-evidence,unknown-id,same-path-multiple-owners,ambiguous-source,level-mismatch,unlocked-member,lock-stale,lock-owner-changed,lock-invalidandmulti-verifies-annotation(under--scan). Move JUnit written outside a testlogs tree into one first (unscoped-evidence, which no rule configures: write it totestlogs/<pkg>/<name>/test.xml), since that changes its case keys; then lock for the four lock findings (a lock written before the move keeps the oldsuite:entries aslock-staleuntilrr sets lock --write --allow-removalsdrops them); resolve the rest (Attribution issues are gaps says how for each); or set a configurable ruleoffunderconfig.rules; or stop gating on gaps until step 8.Keep
attribution: hybrid, the 0.3 default: a single-id tag still owns a case that no claim covers, so every module whose tags you rewrote keeps its owners. Setconfig.main_repoif other modules refer to your targets as@your_repo//....Before merging, check the attribution over fresh evidence from every lane (software and HITL):
$ rr attribution --model requirements/ --evidence bazel-testlogs hitl-testlogs --check
It exits 1 on any quarantine, missing case or error-level issue.
rr reportwould exit 3 on the same quarantines.
Verdicts change honestly at the bump. A requirement that was VERIFIED only through a test it shared now shows what it has on its own; a requirement whose set spans the software and HITL lanes reads INCOMPLETE in each lane’s report and VERIFIED only in a report over both lanes’ evidence. Announce it as a correction, not a regression.
Moving to model mode: rr migrate apply --stage model¶
In hybrid mode a tag can still own a case, and a test deleted or renamed silently leaves its requirement’s set. Model mode makes the model the single record: every owner comes from a claim, and tags only cross-check them.
Once the tags carry one id each and fresh evidence shows the decided
owners, rr migrate apply PLAN.rrplan --stage model [--compress] [--dry-run] writes the claims: one selector per case each entity owns
through a tag today (with --compress, a * glob where it selects exactly
that entity’s cases, none of them skipped, and overlaps no other claim; a
synthetic-only target gets whole: true with a reason). Before writing, it
proves that the owner table over the evidence given is unchanged under
attribution: model and that check_claims passes; it refuses a quarantine
and any worksheet decision the evidence contradicts. A case the worksheet
decided but the evidence given does not hold (another lane’s test, say)
keeps the worksheet’s owner: it gets a literal selector of that entity, and
the stage is refused unless exactly that entity’s claims select it (no
claim, for none); a --compress glob never reaches such a case of another
owner. Cases in neither the evidence nor the worksheet are not seen: pass
every lane’s evidence, or re-run rr attribution --check over it afterwards.
$ rr migrate apply requirements/attribution.rrplan --stage model --compress \
--model requirements/ --evidence bazel-testlogs hitl-testlogs --dry-run
$ rr migrate apply requirements/attribution.rrplan --stage model --compress \
--model requirements/ --evidence bazel-testlogs hitl-testlogs
Then switch the mode and name the lock:
config:
attribution: model
sets_lock: verification.rrlock
Verdicts over every lane’s evidence are identical to the hybrid ones by
construction. A single lane’s report now shows a set that spans lanes as
INCOMPLETE (its other lane’s cases are not-run members), where hybrid mode,
whose tags only own the cases present, ignored them: report lanes with
--lane/--lane-targets, or gate per lane on the combined report
(Reports). Keep the
single-id tags as cross-checks (a tag that disagrees with the model is a
tag-mismatch); new tests need none. rr attribution --suggest prints the
selector for any case still owned only by a tag, or tagged but unclaimed
(unclaimed-tag) afterwards.
Locking the sets¶
$ bazel test //...
$ rr sets lock --model requirements/ --evidence bazel-testlogs hitl-testlogs --write
$ git diff requirements/verification.rrlock
The lock maps every case to the one entity whose set holds it and adds
those cases as expected members: from now on a deleted, renamed or filtered
test makes its requirement INCOMPLETE (missing-case) instead of quietly
shrinking its set. It never creates ownership. Review its diff like a golden
file, and re-lock with every intended change (The verification-set lock).
In a hermetic Bazel project, rr_model(lock = ...) checks it statically and
rr_sets_lock_test against rr_evidence; bazel run :<name>.update
re-locks (the thermostat example does this).
Finally gate on it. In CI, after the report:
- name: Attribution checks (one test case, one requirement)
run: |
bazel query 'tests(//...)' > "$RUNNER_TEMP/targets.txt"
bazel run @rules_requirements//python:rr -- validate requirements --known-targets "$RUNNER_TEMP/targets.txt"
bazel run @rules_requirements//python:rr -- sets check --model requirements --evidence "$(readlink -f bazel-testlogs)"
bazel run @rules_requirements//python:rr -- check-report traceability-report.json
Then tighten the model: bare-target-reference: error and
whole-target-reference: error under config.rules once the remaining
whole-target claims are per case (or carry a reason), and config.variants
for targets that run the same test code. Record the policy and why you chose
it in your development plan (One test case, one requirement: stricter than required).
What 0.4 changes¶
v0.4.0 makes attribution: model the default (hybrid warns
hybrid-mode) and multi-id authoring impossible: a collection, import or
compile error. A project that has finished this guide needs no changes.
Smaller 0.3 changes¶
Smaller v0.3.0 changes a script reading the reports or the Python API may notice (the release notes list every change):
A case is named by its case key,
<target>#<path>, where 0.2 wrote<target> <classname>::<name>: in the JSON report’sunknown_evidenceand in theunknown-idgap.summary.test_casescounts case keys (retries, runs, shards and evidence roots merged; target-scope results not counted), not raw results.TestCase.requirementsis a deprecated alias ofTestCase.declared; it still reads, writes and works indataclasses.replace(case, requirements=...), butdataclasses.asdict()names the fielddeclared, and passing bothdeclared=andrequirements=to the constructor is aTypeError(two sets of ids for one case: neither may silently win), evendeclared=()or another case’sdeclared.