Your GRC platform holds a control library, a risk register, and a stack of completed assessments — and none of it can prove that a single control actually operated last quarter. That is not a configuration problem. It is a category problem: GRC systems were built to document design, and examiners test operation.
Key Takeaways:
- Design effectiveness asks whether a control would work as described; operating effectiveness asks whether it actually ran every required period, and examiners test the latter
- A GRC repository can produce the control description and its mapping, but not the execution records — owners, timestamps, artifacts — that an evidence request demands
- Evidence reconstructed in the weeks before an exam is both expensive and less credible than evidence captured when the work happened
- Five diagnostic questions reveal whether your current stack can prove operation or only describe intent
Design Effectiveness vs. Operating Effectiveness: What Examiners Actually Test
Auditors and examiners draw a sharp line between two properties of a control. Design effectiveness asks: if this control operates as described, would it satisfy the requirement? Operating effectiveness asks: did it operate as described, every period it was supposed to, and can you show me?
GRC platforms are built for the first question. They hold control descriptions, map controls to risks and regulations, and record the results of periodic self-assessments. All of that is design documentation. It establishes intent.
Examination is built around the second question. The FFIEC BSA/AML Examination Manual directs examiners to test transactions, review completed monitoring output, and evaluate whether the program functions in practice — not to read the program description and move on. The FDIC Consumer Compliance Examination Manual takes the same approach for consumer regulations: examiners sample actual work product. A well-written control that did not run is, for examination purposes, a finding.
What a Request for Six Months of Operating Evidence Looks Like
Consider a routine request from an exam document list: "Provide evidence that the bank performed its monthly review of fintech partner marketing materials for the period January through June, including reviewer, date, materials reviewed, issues identified, and disposition."
A complete answer contains, for each of the six months:
- The work record showing the review occurred, with a date and a named reviewer
- The inventory of marketing pieces in scope that month
- The artifacts of the review itself — checklists, annotated materials, screenshots
- Any issues identified, and the remediation task each issue produced
- Evidence that each remediation closed, with dates
Now ask what your GRC platform contributes to that answer. It can produce the control description ("Bank reviews partner marketing monthly") and its mapping to UDAAP risk. It cannot produce the six work records, the six evidence sets, or the remediation trail, because it never generated or captured them. Those live in inboxes, shared drives, and the memory of whoever did the reviews — if they were done at all.
Why a Repository Cannot Answer an Operating Evidence Request
The gap is structural. A repository stores descriptions that change rarely; operating evidence is generated continuously by the work itself. A system that is not in the path of the work cannot capture proof of the work. Uploading artifacts into the GRC system after the fact just relocates the manual scramble.
Reconstruction is also a credibility problem. Evidence assembled in the weeks before an exam looks like evidence assembled for the exam, and experienced examiners recognize it: file creation dates cluster suspiciously, artifacts are missing for exactly the months when staff turned over, and narratives are written in the past tense from memory. Records with contemporaneous timestamps and attribution — created because the system generated the work and captured its output — read differently. We walk through a full request in When the Examiner Asks for Proof.
Five Diagnostic Questions for Your Own Program
Ask these of your current stack, honestly:
- Can you trace every control to its exact source requirement? Not to a regulation name — to the specific section that makes the control necessary.
- When a policy changes, can you identify every affected workflow? If the answer involves a person reading the redline and remembering what depends on it, the answer is no.
- Can you prove each control operated every required period? Every period, not most periods — gaps are what samples find.
- How much evidence is collected manually before an exam? Weeks of collection means the evidence exists in people's heads and inboxes, not in a system.
- Who is chasing the artifacts? If your most senior compliance people spend exam prep season emailing colleagues for screenshots, the platform is storing the program rather than running it.
Institutions with mature GRC deployments routinely fail three or more of these, because the questions test execution and GRC is a governance tool.
What Proving Operation Actually Requires
Proof of operating effectiveness has three ingredients, and all of them must be produced by the work itself rather than reconstructed afterward.
Execution records with timestamps and owners. Every required period generates a record: this control ran, on this date, performed by this person. The record exists because the system scheduled the work and logged its completion.
Artifacts generated as work happens. The reviewed file, the signed checklist, the monitoring output — attached at the moment of completion, not hunted down eighteen months later.
Lineage back to the requirement. Each record links through the control to the obligation and the source, per 31 CFR § 1020.210 or whichever provision applies, so a single evidence package answers both "what do you do?" and "why do you do it?"
This is the operating definition of a compliance execution platform: a system positioned in the path of the work, generating the schedule, capturing the output, and maintaining the chain from source to artifact.
How Modern Teams Close the Operating Evidence Gap
Teams that have been burned by an evidence scramble stop treating exam preparation as a project and start treating it as a byproduct. In Canarie, each obligation drives recurring work with a named owner and a cadence; evidence is captured when the work completes; attestations are built on that evidence rather than on memory. When the request for six months of support arrives, the six months of records already exist, timestamped and linked to the requirement they satisfy.
The GRC register described the control. The execution record proves it operated. Examiners grade the second one.
Prove every control operated, every period →
Frequently Asked Questions
What is the difference between design effectiveness and operating effectiveness?
Design effectiveness means a control, as described, would satisfy its requirement if performed — it is evaluated by reading the control and the requirement together. Operating effectiveness means the control actually performed as designed over a period, and it is evaluated by examining execution records and artifacts. A control can be perfectly designed and still fail an exam because no one can show it ran.
Why can't we just upload evidence into our GRC platform?
You can, but uploading is itself manual work that happens after the fact, which recreates the reconstruction problem inside a different tool. The system holding the evidence was not the system that scheduled the work, so nothing guarantees completeness — missed periods simply have no upload, and no one notices until an examiner does. Capture has to happen at the point of execution to be reliable.
What do examiners consider acceptable operating evidence?
Records that are contemporaneous, attributable, and complete: a dated execution record for every required period, a named performer, and the underlying artifact (report, checklist, approval, screenshot) produced by the work. Examiners also expect the evidence to connect to the requirement it satisfies, and they sample across the full period rather than accepting the best three examples.
How long does it take to respond to a six-month evidence request without an execution system?
Institutions that assemble evidence manually typically spend days to weeks per request, involving multiple staff who search email, shared drives, and ticketing tools — and exam document lists contain dozens of such requests. The elapsed time matters less than the failure rate: manual assembly routinely discovers missed periods and missing artifacts at the worst possible moment, during the exam itself.