Test hooks¶
A hook lets a test declare which entities it verifies, and at what level. Every
hook produces the same thing: JUnit XML in which each <testcase> carries
its traces as properties. That keeps the report independent of the test
framework, and lets anything that can write JUnit take part (Evidence and ingestors).
One test case, one requirement
A test case verifies at most one requirement; a set of test cases may
verify one requirement. Declare one id per test case. A hook only declares
the id; which requirement a case verifies is decided by attribution, never by a
hook. The older multi-id forms (@pytest.mark.rr("REQ-1", "REQ-2"),
@rr.verifies("REQ-1", "REQ-2"), a list passed to JUnitWriter) still record
every id in 0.3, but are deprecated and warn with MultipleRequirementsWarning;
the case is then quarantined and counts for none of them — see
Deprecated: several ids per test case.
With config.attribution: model the model’s verified_by claims decide which
requirement a case verifies, and the id a hook writes is only a cross-check
(tag-mismatch when it disagrees, unclaimed-tag on a case no claim
selects); a test the model claims needs no hook call at all. See
One test case, one requirement.
Framework |
Declare |
Bazel |
Without Bazel |
|---|---|---|---|
pytest |
|
|
pip-installed plugin, |
unittest |
|
|
|
googletest |
|
|
|
plain-assert C++ |
|
|
|
Rust |
|
|
|
node:test |
|
|
— |
hardware runs (steps + checks) |
|
|
same |
anything else (Python) |
|
|
write the XML yourself |
shell scripts |
|
|
|
runners writing their own JUnit |
(their own) |
|
|
The JUnit conventions¶
Property |
Meaning |
|---|---|
|
The id the test case declares — one per case. (Repeated properties, and comma- or whitespace-separated values, are read as several ids (whitespace since 0.3; 0.2 read |
|
0.2’s name for comma-separated ids (googletest’s hook wrote it); still read, like |
|
The source file of the test code that produced the case, relative to the workspace (written by the pytest and unittest hooks, |
|
|
|
|
|
The verification level the case provides. Default: the model’s |
|
One key of the identity of the artifact exercised (firmware build id, board revision, git SHA…), for staleness checks. |
The properties may be <property name=... value=.../> children of the
<testcase>’s <properties>, attributes of the <testcase> element (older
googletest), or <properties> of an enclosing <testsuite>, which then apply
to every case in the suite — a convenient place for a suite-wide level or
artifact identity.
<testcase classname="tests.test_controller" name="test_heats_below_band" time="0.001">
<properties>
<property name="requirement" value="REQ-1"/>
<property name="level" value="simulation"/>
<property name="artifact.firmware_build_id" value="fw-2026.09.30-3"/>
</properties>
</testcase>
pytest¶
import pytest
pytestmark = pytest.mark.rr("REQ-10") # every test in the module
@pytest.mark.rr("REQ-1")
def test_heats_below_band(): ...
@pytest.mark.rr("REQ-5", level="hil", artifact={"board_rev": "C"})
def test_cutoff_on_the_bench(): ...
@pytest.mark.requirements("REQ-3") # alias
def test_parses_units(): ...
The positional argument is the id. Several ids — several arguments, a comma- or whitespace-separated string, a list, or several markers at one scope naming different ids — are deprecated: every id is still recorded (so the case is quarantined), and the plugin warns once per declaring test (all its parameters), class or module (details).
level=names the level the test provides;artifact=a mapping of artifact identity keys.The nearest scope wins. Only the nearest scope that names an id is recorded, in this order: a
pytest.param(..., marks=...)mark; the test function (its markers, a conftest’sitem.add_marker,@rr.verifies); its class (a subclass’s own declaration before its base class’s); the module’spytestmark, then a package’s. A nearer id replaces a farther one: a module-levelpytestmark = pytest.mark.rr("REQ-10")is a default that a test’s own marker overrides, never an id added to it. (0.2 recorded the ids of every scope.)Sibling base classes are one scope. A class that declares no id gets the ids of every base class that no nearer class overrides, together:
class TestD(MixA, MixB)withMixAnamingREQ-1andMixBnamingREQ-2records both, warns (RR-E101), and is quarantined. It is never resolved by MRO order. A base reached along two paths (a diamond) counts once, and a base that another contributing base subclasses is overridden by it.The nearest marker naming a level wins (
rrandrequirementsmarkers alike), and artifact keys resolve nearest-first.Every case also records
rr.file, the test file relative to the workspace.A raw
record_property("requirement", ...)(or"requirements") bypasses these rules: the plugin drops the property and fails the test with RR-E102. Declare the id with the marker.record_testsuite_property("requirement", ...)is written on the<testsuite>; ingest gives a suite-level requirement to no case and warns (suite-level-requirement).unittest.TestCasemethods collected by pytest honour@rr.verifies(...)(below); the decorator’s level applies only when no marker names one.
Installation. With pip, the plugin registers itself through the pytest11
entry point and pytest --junitxml=results.xml is all you need. Under Bazel,
use rr_py_test (Bazel rules), which runs pytest over the listed files, loads
the plugin and writes the JUnit to $XML_OUTPUT_FILE; add your workspace’s
pytest to its deps. A hand-written Bazel main can do the same:
from rules_requirements.hooks.pytest_runner import main
if __name__ == "__main__":
raise SystemExit(main(__file__)) # pytest over this file's directory
The properties are recorded when the test is collected, so they reach the
JUnit case however the test ends — including tests skipped by
@pytest.mark.skip / skipif (a skipped test never counts as passing evidence,
but it is listed under its requirement). The nearest declaration naming a
level wins: a method’s own marker or @rr.verifies beats its class’s, which
beats a module-level pytestmark.
unittest¶
import unittest
from rules_requirements import rr
@rr.verifies("REQ-9", level="sil") # applies to every test in the class
class InterlockTest(unittest.TestCase):
@rr.verifies("REQ-5")
def test_trips_at_limit(self): ...
if __name__ == "__main__":
rr.unittest_main()
rr.verifies(id, level="", artifact=None) records the trace on the function
or class. The nearest declaration wins: a method’s decorator replaces its
class’s, and a subclass’s replaces its base class’s (a level or artifact key
the nearer declaration does not set is still inherited). As under pytest,
sibling base classes that no nearer class overrides are one scope: two naming
different ids record both, so the case is quarantined, with a warning at the
class definition when its first test starts. A setUpClass /
setUpModule error and a failing subtest carry the same single id as their
test, and every case records rr.file. Extra ids, or stacked decorators
naming different ids, are deprecated: every id is still recorded (so the case
is quarantined), with a warning at the decorated definition
(details).
rr.unittest_main() replaces unittest.main(): it runs the module’s tests and
writes JUnit to $XML_OUTPUT_FILE, so a plain Bazel py_test with
@rules_requirements//python in its deps needs nothing else. Outside Bazel, pass --junit-xml:
$ python -m my_tests --junit-xml results.xml # a module that calls rr.unittest_main()
rules_requirements.hooks.unittest.main() also runs discovery, for a
small runner script of your own: it accepts --discover DIR, --pattern GLOB
(default test*.py), --junit-xml PATH and -q, and returns the exit status.
Expected failures are recorded as passed, unexpected successes as failed, skips
as skipped. When the module runs as __main__, its test classes are named after
the file.
googletest¶
#include "rr_gtest.h"
TEST(Interlock, TripsAtLimit) {
RR_VERIFIES("REQ-5"); // the ONE id this test verifies
RR_LEVEL("sil"); // optional
RR_ARTIFACT("firmware_build_id", kBuildId); // optional, per key
...
}
In Bazel, add @rules_requirements//cc:gtest to the cc_test’s deps next to
your own googletest (the library is header-only and deliberately does not
depend on googletest, so it never adds googletest to your module graph).
googletest writes JUnit to $XML_OUTPUT_FILE under
bazel test, so a plain cc_test needs no wrapper; elsewhere run the binary
with --gtest_output=xml:results.xml.
The id is recorded as the test’s requirement property (0.2 wrote
requirements; both are read). Call it once, with one id: several ids, or
several calls naming different ids, are deprecated — RecordProperty keeps one
value per key, so they are recorded as one comma list, which is quarantined.
RR_VERIFIES in SetUpTestSuite, an Environment or main is still recorded
on the suite, but no test case inherits a suite-level requirement. The helpers
are also available as functions — rules_requirements::Verifies({...}),
Level(...), Artifact(key, value).
Plain-assert C++: rr_case.h¶
Many C and C++ tests are a main() that calls test functions full of
assert(): the first failure aborts the binary, and Bazel can only report the
whole target. rr_case.h turns such a binary into one JUnit case per test
function, without googletest. The header is C++ (11 or later); a C-style
assert() test uses it when compiled as C++:
#include "rr_case.h"
RR_CASE(wifi_settings_vector) { // one case
RR_CHECK(Encode(kSettings) == kWireVector); // like assert(), kept under NDEBUG
}
RR_CASE(rejects_truncated_frame, "REQ-7") { // optional: the one id it verifies
assert(!Decode(kTruncated)); // plain assert() works too
}
int main(int argc, char** argv) { return rr::RunCases(argc, argv, "improv_codec"); }
An existing main converts without moving its test functions — list them:
int main(int argc, char** argv) {
return rr::RunCases(argc, argv, "improv_codec", {
{"wifi_settings_vector", test_wifi_settings_vector},
{"rejects_truncated_frame", test_rejects_truncated_frame, "REQ-7"},
});
}
In Bazel, add @rules_requirements//cc:case to the cc_test’s deps; the
library is header-only and has no dependencies. Under bazel test the JUnit
goes to $XML_OUTPUT_FILE; elsewhere pass --rr_junit=results.xml.
Isolation. On POSIX each case runs in its own forked child, with core dumps suppressed. A failing
assert()orRR_CHECK, a crash, an uncaught exception or a non-zeroexit()fails that case only — the message says how (terminated by SIGABRT: codec_test.cc:31: RR_CHECK(n == 4) failed,terminated by SIGSEGV,exited with status 1: uncaught exception: ...) — and the next case still runs. The binary exits 1 if any case failed, so the target fails exactly as it did before. Withoutfork(Windows) the cases run in-process, and a failingassert()ends the binary there.No shared state. Each case starts from the parent as it was before the first case: a global, heap object or
chdirset by one case is gone in the next, so a case must not depend on an earlier one (in a plainmainit could). Start no threads beforerr::RunCases—forkcopies only the calling thread, so a lock another thread held stays locked in the child.Child lifetime. A child ends with
_exit, soatexithandlers and static destructors run once, in the runner, never per case. A hung case does not outlive its runner: aSIGTERM,SIGINTorSIGHUPto the runner (anrr_evidencetimeout sendsSIGTERM, thenSIGKILLafter 2 s) kills the running case’s child first, and on Linux the child also dies with a runner killed bySIGKILL. The signal then goes to what the program had for it: by default it ends the runner as the signal would have; a handler the program installed beforerr::RunCasesruns as it would have (if it returns, the runner goes on and reports the killed case as failed); a signal the program ignores stays ignored. A case starts with the runner’s own signal mask and signal dispositions, as they were beforerr::RunCases. Tests run byrr_evidencestay in the action’s process group, so a cancelled build reaches them too.Coverage (Linux and other ELF targets). A child starts its case’s counts from zero and writes them before it ends, so a line run in a case is counted once per run of it, and a line run before the cases once, in the test executable and in every instrumented shared library it loads (by default
bazel coveragelinks acc_test’scc_librarydeps as shared libraries). That takes libgcov’s__gcov_resetand__gcov_dump, which gcc links only on request:@rules_requirements//cc:caseadds-Wl,-u,__gcov_dump -Wl,-u,__gcov_resetto the link underbazel coverageon Linux. A toolchain without a gcov runtime (no libgcov) cannot link those: pass--@rules_requirements//cc:coverage_hooks=false(defaulttrue) tobazel coverageand//cc:caseadds no link options; the cases’ coverage is then counted as described next for a build without them. A--coveragebuild of your own outsidebazel coverageneeds the same two link options; without them, code in a shared library that only cases run is not counted at all, and unless the test’s own source is instrumented (--instrument_test_targets), code run before the cases is counted once more per case. Verified with gcc 13: shared and static links, the test’s own source instrumented or not, andbazel coveragewith Bazel 7.7.1 and 8.8.1. Not tested: clang, with--coverage(compiler-rt defines__gcov_dumpand__gcov_resetitself) or with-fprofile-instr-generate(__llvm_profile_reset_counters,__llvm_profile_write_file;LLVM_PROFILE_FILEneeds%por%m, as Bazel sets it). On macOS the children’s coverage is lost.Leaks. Under LeakSanitizer (ELF targets) a case that leaks memory fails (
exited with status 1: rr_case: LeakSanitizer found memory leaked by this case). The runner checks once before the first case: memory leaked before any case ran (inmainor a static initializer) fails the binary without failing any case (rr_case: LeakSanitizer found memory leaked before any case ran), and the cases’ own checks are then skipped, since a child’s leak could no longer be told apart from it. On macOS the children’s leak checks are lost.Case keys. Each case is
<testcase classname="<suite>" name="<case>">, reported as<suite>::<case>(improv_codec::wifi_settings_vector). Case names must be unique within a suite; a duplicate is reported as an error.One requirement per case. A case names at most one id, recorded as its
requirementproperty.RR_CASE(name, "REQ-1", "REQ-2")does not compile (RR-E101; with 16 or more ids the error is an unrelated one), nor doesRR_CASE(name, ("REQ-1", "REQ-2")). In the list form that comma expression compiles, with no warning at all unless-Wallor-Wunused-valueis on, and records justREQ-2, so build such tests with-Werror=unused-value. An id that is empty or holds anything but ASCII letters, digits,_,-and.makes the case an error without running it (RR-E104).rr_case.haccepts ids matching[A-Za-z0-9_.-]+only: a project whoseconfig.id_patternallows other characters must use another hook. There is no call to add ids from inside a case. The id is optional: a case without one is still a test case in the report, and a whole-targetverified_byreference covers it.Evidence and annotations.
RR_CASE(name, "REQ-1")written on one line is also an annotation forrr scan; the list form’s ids ({"name", fn, "REQ-7"}) are evidence only, like anRR_CASEthat a formatter splits across lines.RR_CASEalso records where the case is defined, as the<testcase>’sfileandline.Killed runs. The JUnit is rewritten before each case with that case recorded as an error, so a binary killed mid-case (a Bazel timeout) still reports the cases that finished and names the one that did not.
Selection.
bazel test --test_filter=GLOB[,GLOB...]runs the cases whose name orsuite::casekey matches (*,?), andshard_countis honoured.
Flag |
Meaning |
|---|---|
|
Print every case key ( |
|
Run one case in-process, without fork or JUnit — for a debugger. |
|
Write the JUnit here instead of |
Other arguments are left to the test. rr_case.h records no level: a tagged
case provides the model’s default_provided_level.
Rust¶
#[test]
fn cutoff_at_limit() {
rr::verifies!("REQ-5"); // the ONE id this test verifies
assert!(interlock::trips(35.0));
}
#[test]
fn cutoff_is_logged() {
rr::verifies!("REQ-6"; level = "sil"); // ...with a level
assert!(interlock::logs_trip(35.0));
}
The rr crate is @rules_requirements//rust:rr in Bazel. With Cargo, add it
as a development dependency — the package is rules_requirements in the
repository’s rust/ directory, and its library is named rr:
[dev-dependencies]
rules_requirements = { git = "https://github.com/Studio-Fug/rules_requirements" }
libtest has no stable machine-readable output, so the macro records traces out
of band: each call appends one JSON line — the test’s name, taken from the
thread libtest runs it on, plus the id ("requirement":"REQ-1") and level — to
the file named by $RR_TRACE_FILE. A deprecated call naming several ids writes
0.2’s list form ("requirements":[...], still read), and two calls naming
different ids write two lines: either way the case is quarantined. Without that variable the macro does nothing, so plain
cargo test is unaffected. The wrapper sets the variable, runs the tests,
parses libtest’s standard output and writes JUnit with the traces attached:
in Bazel, rr_rust_test (a
rust_testplus the wrapper), or rr_wrapped_test around any libtest binary;elsewhere,
rr wrap --junit-xml rust.xml -- <test binary or command>(see below).
A trace recorded on a thread the test spawns itself carries that thread’s name, matches no test and is dropped (with a warning); call the macro from the test’s own thread.
node:test¶
A JavaScript or TypeScript test file run by Node’s built-in test runner
(node:test: test(), it(), describe()) becomes one JUnit case per test
with rr_node_test. Without it, a rules_js js_test runs
the file as a plain script and Bazel records one synthetic result for the whole
file — every test() in it shares one verdict.
load("@aspect_rules_js//js:defs.bzl", "js_test")
load("@rules_requirements//rr:defs.bzl", "rr_node_test")
rr_node_test(
name = "clocksync_test",
rule = js_test, # your rules_js js_test
test = ":dist-test/tests/clocksync.test.js", # compiled from TypeScript
data = [":web_tests_js", ":dist_test_pkg_json"],
)
The file runs exactly as js_test runs it — node <file>, with rules_js’s
node flags — and the target exits with node’s own exit code, so the runner
never turns a failing test green or a passing one red. Two reporters are
attached: spec writes the usual log, and rules_requirements’ reporter
records the cases, which are written to $XML_OUTPUT_FILE once node exits:
node:test |
JUnit case |
|---|---|
a test without subtests ( |
|
a |
no case of its own — its leaves are the cases |
passed / failed |
|
|
|
|
one |
So describe("bestSample", ...) around it("keeps the min-RTT sample", ...)
in clocksync.test.js is the case clocksync > bestSample::keeps the min-RTT sample. Two tests with one name stay two cases. Every case carries an
rr.file property: the workspace-relative file that defines the test — the
test file, or a helper module it requires. For TypeScript compiled to
JavaScript, that is the compiled file (dist-test/...), not the .ts source.
Declaring traces. A test declares what it verifies with a diagnostic:
const { test } = require("node:test");
// rr_node_test sets RR_NODE_VERIFIES; the fallback keeps the file loadable
// elsewhere (`node --test`, an IDE), without the helper's guards.
const { verifies } = process.env.RR_NODE_VERIFIES
? require(process.env.RR_NODE_VERIFIES)
: { verifies: (t, id) => t.diagnostic(`rr.requirement=${id}`) };
test("bestSample keeps the min-RTT sample", (t) => {
verifies(t, "REQ-13"); // one id; optional level: verifies(t, "REQ-13", "sil")
// the same, by hand:
// t.diagnostic("rr.requirement=REQ-13");
// t.diagnostic("rr.level=sil");
// t.diagnostic("rr.artifact.board_rev=C");
});
rr.requirement=, rr.level= and rr.artifact.<key>= diagnostics become the
case’s requirement, level and artifact.<key> properties (other
diagnostics are left alone). They belong to the test that writes them — also
under describe(..., { concurrency }) — and a diagnostic that cannot be tied
to a test, such as one from a before() / after() hook, is never guessed
onto one: it is reported as a warning in the test log. (A beforeEach() /
afterEach() hook’s t is the test’s own context, so its diagnostics do
belong to that test.) The level attribute of rr_node_test is the
default for cases that declare none.
verifies(t, id, level?) (@rules_requirements//js:verifies.cjs; under
rr_node_test its path is in $RR_NODE_VERIFIES) adds two guards to the
diagnostic: the id must be one id — no comma, no whitespace (RR-E104) — and a
test that already verifies one requirement cannot claim a second
(RR-E101); either mistake throws inside the test, which then fails. A test
case verifies at most one requirement: if raw diagnostics name several ids
anyway, every one is written (never a silent pick) and the test log carries an
RR-E101 warning.
Failures outside any test are recorded as error cases with the property
rr.scope=target — they belong to the whole target, not to a test:
Case |
When |
|---|---|
|
node exited non-zero before reporting any test: the file threw while loading, or the process died. |
|
node exited non-zero (or was killed) although no test failed — e.g. an unhandled rejection after the tests. |
|
a root-level |
|
a |
|
node exited 0 after node:test started reporting but before it finished — |
The other root-level hooks fail tests instead: a failing root before() fails
every top-level test with the hook’s error, and every top-level describe
gets a <hooks> case carrying it (its tests are cancelled); a failing root
beforeEach() / afterEach() fails every test it runs for.
A file that registers no test at all writes an empty suite (tests="0").
Two tests that report as one case key — a describe("a > b") next to a
describe("a") holding a describe("b"), or :: inside a name — get an
rr_node_test: warning in the test log naming both; rename one.
Because these cases name no requirement, they fail every whole-target
verified_by reference to the target, not the requirements its individual
tests name.
Node versions. The reporter needs --test-reporter, so Node 20 or newer
(rules_js’s default toolchain is Node 22). On older Node, or with
RR_NODE_TEST_PLAIN=1 in the test’s environment, the file runs plainly and
the runner writes one result for the whole target with the property
rr.synthetic=true — what a plain js_test gives. rules_requirements’ CI runs
the runner on Node 18 (the fallback), 20, 22 and 24.
The runner never decides the verdict itself: if it cannot write the report
(an unwritable $XML_OUTPUT_FILE directory, a full disk) it warns in the log
and still exits with node’s code, and it passes SIGTERM, SIGINT and
SIGHUP on to the test process, so a timeout or an interrupt never leaves
that process running.
Hand-rolled harnesses: JUnitWriter¶
Hardware-in-the-loop and end-to-end runners are often plain programs rather than
framework test suites. JUnitWriter
gives them the same output:
import os
from rules_requirements.hooks.junit_writer import JUnitWriter
report = JUnitWriter(
"bench_e2e",
default_level="hitl", # what this harness provides
artifact={"firmware_build_id": build_id, "dut_git_sha": sha},
)
try:
with report.case("flash_and_boot", requirement="REQ-13"):
flash(dut) # raises -> recorded as failed, then re-raised
with report.case("provision", requirement="REQ-29", level="hil"):
provision(dut)
report.add("ota_update", requirement="REQ-30", status="skipped", message="no OTA server on this rig")
finally:
report.write(os.environ.get("XML_OUTPUT_FILE", "bench_e2e.xml"))
case(name, requirement=None, level="", artifact=None, *, classname="", file=None)is a context manager that records the case as passed, or as failed with the exception’s text if the block raises (the exception propagates, so control flow is unchanged).add(name, requirement=None, status, message, duration, level, artifact, classname, *, file=None)records a result directly;statusispassed,failed,errororskipped.requirementis one id, orNone. A string that is not one id — a comma or whitespace inside it, or empty — raisesValueError(RR-E104). The pre-0.2 list form (report.case("x", ["REQ-13", "REQ-21"]), or therequirements=keyword) still records every id it holds, verbatim, with aDeprecationWarning(aMultipleRequirementsWarningwhen it names more than one: the case is then quarantined).casesis a read-only tuple of frozen cases (case.requirement, andcase.requirements, a read-only alias of the declared ids): a case cannot be removed or re-attributed once recorded. Record another case instead.not_reached(names, reason, classname="", *, tags=None)records planned cases a device failure kept from running, each as its own failed casenot reached: <reason>;tagsmaps a name to the one id it verifies.The writer’s
artifactidentity is stamped on every case (merged with any per-caseartifact), which is what makes stale bench results detectable.The writer’s
file— by default the running script,sys.argv[0], relative to the workspace (the*.runfiles/<repo>/orbazel-out/<cfg>/bin/prefix stripped) — is written as each case’srr.fileproperty; passfile=""to write none, orfile=per case to override it.write(path, append=False)writes the JUnit; withappend=Truethe cases are added to the file already atpath(to the suite of the same name, else as a new suite), replacing it atomically; on POSIX systems concurrent appends are serialised by a lock on the file — or, where the file itself cannot be locked (a read-only file on NFS, a dangling symlink), on a sidecar.<name>.locknext to it, removed again. A symlink to a file is followed for the lock, then replaced by the new file like the file itself.
Hardware runs: CheckPlan¶
An on-hardware run is a sequence of steps — flash, boot, provision, connect —
each followed by assertions. A step is an action: it verifies nothing by
itself. A check is one assertion, recorded as one JUnit case that verifies at
most one requirement. CheckPlan
records a run in those terms, including how it stopped:
import os
from rules_requirements.hooks.checkplan import CheckPlan
from rules_requirements.hooks.junit_writer import JUnitWriter
STEPS = {
"flash_boot": ("ble_advertising",),
"improv_provision": ("provisioned",),
"websocket_checks": ("ws_connect", "build_info", "rename"),
"run": ("completed",),
}
TAGS = { # "<step>.<check>" -> ONE id
"flash_boot.ble_advertising": "REQ-13",
"improv_provision.provisioned": "REQ-13",
"websocket_checks.ws_connect": "REQ-13",
"websocket_checks.build_info": "REQ-35",
"websocket_checks.rename": "REQ-13",
"run.completed": "REQ-23",
}
report = JUnitWriter("hitl_e2e", default_level="hitl")
plan = CheckPlan(report, STEPS, tags=TAGS, is_infrastructure=is_rig_trouble)
try:
with plan.run():
reserve_rig() # setup: any failure here is rig or setup trouble
plan.setup_done()
with plan.step("flash_boot"):
flash(dut) # the action: a failure here is the device's
with plan.check("ble_advertising"):
assert BLE_MARKER in serial_log()
with plan.step("improv_provision"):
provision(dut)
plan.passed("provisioned")
with plan.step("websocket_checks"):
...
with plan.step("run"):
plan.passed("completed")
finally:
report.write(os.environ["XML_OUTPUT_FILE"])
Each check becomes the case <suite>.<step>::<check> —
hitl_e2e.websocket_checks::rename — carrying its tag, if any. check(name)
records a pass, or a failure with the exception’s text (and re-raises);
passed, failed and skipped record a result directly, by the check’s name
in the current step or as "<step>.<check>". Inside the check’s own
with plan.check(name): block, such a result is the check’s only case — the
block’s end adds no pass, and a later exception no failure:
with plan.check("board_caps"):
if descriptor is None:
plan.skipped("board_caps", "no capability descriptor on this board")
else:
assert descriptor.ok
Step names cannot contain ., which separates the step from the check.
How a run is recorded when it stops:
The run |
Recorded |
|---|---|
ends normally |
Every check as it went. A planned check never recorded is an |
stops on a device failure: any exception after |
The check that raised, as failed; every check not run yet — the rest of the step and all later steps — as failed, |
stops on rig or setup trouble: an exception before |
One untagged |
is interrupted or exits cleanly: |
As rig trouble, whatever |
hits a harness bug: an unknown step or check name, a check outside a step, a check recorded twice |
An untagged |
The exception is re-raised in every case, so the harness exits as it would
without the plan. is_infrastructure defaults to claiming nothing else: pass
the harness’s own classifier (reservation errors, ssh’s own exit 255…).
Note
Rig trouble and partial runs. A requirement is verified only when every case in its verification set passed, so after rig trouble the skipped checks leave a requirement that also has passed checks INCOMPLETE: neither VERIFIED nor FAILED. The passed checks keep their tags. (0.2, without verification sets, withdrew the tags of those passed checks instead; 0.3 dropped that.)
Shell harnesses: rr case¶
rr case appends one test case to a JUnit file — $XML_OUTPUT_FILE by default
— creating it if needed, so a shell script can record its own results:
rr case --name "flash ok" --status passed --requirement REQ-21 --level hitl
rr case --name "boot banner" --status failed --message "no banner after 30 s" \
--requirement REQ-22 --artifact firmware_build_id="$BUILD_ID" --file "$0"
Option |
Meaning |
|---|---|
|
The case’s name (required). |
|
|
|
The one id the case verifies; given twice it is an error (RR-E101), as is a malformed id (RR-E104). |
|
Default: the suite, which defaults to the name part of |
|
As for the other hooks. |
|
Failure or skip message; seconds. |
|
The JUnit file (default |
|
The test code’s source file, recorded as |
rr case exits 0 whatever the case’s status — the script’s own exit status
still decides whether the test passed — and 2 when it cannot record the case
(no output file, more than one id, a malformed id or --artifact, an
unreadable existing file). The file is replaced atomically, and on POSIX
systems concurrent appends (rr case ... &) are serialised by a lock on it
(or on a sidecar .<name>.lock, as for JUnitWriter), so none is lost; on
Windows they are not.
rr wrap¶
rr wrap (also python -m rules_requirements.hooks.wrap) runs a command,
echoes its output, converts it to traceability JUnit and exits with the
command’s own exit code, so it never turns a failing test green or a passing one
red:
$ rr wrap --junit-xml rust.xml -- ./target/debug/deps/setpoint-1a2b3c --test-threads=4
$ rr wrap --format junit --junit-in out/junit.xml -- ./run_bench.sh
Option |
Meaning |
|---|---|
|
The command’s output: |
|
With |
|
Where to write JUnit (default |
|
Suite name (default: the name part of |
|
The test’s label, used to name the suite. |
|
Level for cases that did not declare one. |
The wrapped command runs with $RR_TRACE_FILE set (in $TEST_TMPDIR when
available) and without $XML_OUTPUT_FILE, so it cannot overwrite the JUnit the
wrapper writes. If no test results can be parsed — the binary crashed before
running tests, say — the wrapper records one synthetic case carrying the exit
code and the tail of the output, so the failure is visible in the report.
If the command exits non-zero although every reported test passed (a
sanitizer, a crash after the last test), an exit-status error case is added.
It declares no requirement — not the ids the run traced — and is
target-scope (rr.scope=target): it taints every case claimed on the target,
so each requirement fails through its own cases. A test that traced and then
died without reporting a result is recorded as an error carrying its own
declared id. rr_evidence’s test.exit.xml works the same way.
With --format junit the runner’s report is copied to the output as it is.
--level becomes the default level of its suites (a case’s own level wins).
If the runner exits non-zero although no case in its report failed, the same
exit-status error case is added as for libtest; if it wrote no report at all,
or a well-formed file that is not JUnit, one error case says so, whatever its
exit code. A report without a single case from a run that exited 0 is
replaced by one synthetic passed case, as for a libtest binary that printed
nothing, so the run still leaves evidence.
Deprecated: several ids per test case¶
A test case verifies at most one requirement. In 0.3 every hook still accepts
the older multi-id forms and records every id, but warns:
the Python hooks with
MultipleRequirementsWarning (a
DeprecationWarning), the others with an RR-E101 line on stderr:
Hook |
Deprecated form |
Warns |
|---|---|---|
pytest |
a marker with several ids, several markers at one scope naming different ids, a marker and |
once per declaring test (all its parameters), class or module, at the marker’s line, when the first test it applies to sets up; listed in pytest’s warnings summary |
unittest |
|
at the decorated definition; for sibling bases, at the class definition when its first test starts (escalated to an error, printed to stderr, so the run goes on) |
|
a list, tuple or other iterable naming several ids, positionally or as |
at the |
googletest |
|
on stderr (in the test log), when the test gains its second id |
Rust |
|
|
Ids at several scopes — a module-level pytestmark plus a function’s own
marker, a pytest.param mark plus the function’s marker, a class decorator
plus a method decorator, or a subclass’s @rr.verifies plus its base class’s —
are not a multi-id declaration and do not warn: the nearest one wins, and only
its id is recorded.
A case whose evidence names several ids is quarantined: it counts for no
requirement, and every requirement it names reads INVALID. 0.4 rejects
multi-id declarations outright. Split such a test into one test per requirement, or
keep the one id it really verifies. The thermostat example did both for its
Rust test requires_a_unit, which called rr::verifies!("REQ-3", "REQ-4"):
its assertions check the syntax, so it keeps REQ-3, and REQ-4 gained a test
of its own that checks the range after converting °F,
checks_the_range_after_converting. To find every remaining use, turn the
warning into an error: pytest -W error::DeprecationWarning, or
python -W error::DeprecationWarning for a script. Under pytest a marker
declaration then errors the first test it applies to, at setup, and the rest of
the session runs; but @rr.verifies warns when it decorates, at import, and a
JUnitWriter call made at import time warns there too, so an escalated
warning from either is a collection error that interrupts the whole session.
The hooks name the problem with a stable code:
Code |
Meaning |
|---|---|
RR-E101 |
One case names more than one id. |
RR-E102 |
A raw |
RR-E103 |
|
RR-E104 |
Malformed id: a comma, whitespace, or empty. |