Operational guidance, not legal advice. This page distills named public sources (regulator guidance and industry practice). It is not a legal determination, not a notification decision, and not a substitute for your counsel, insurer, or a retained DFIR firm. Verify applicability and current deadlines for your facts and jurisdiction.

How do you prove a security control actually works?

Updated

Show two things separately: that the control is designed to prevent or detect the risk, and that it operated that way throughout the period. The second needs a complete population, a sample drawn from it, and no exceptions you cannot explain.

Guidance, last verified 10 September 2026 against the AICPA Trust Services Criteria and the AICPA's guidance on design and operating effectiveness and on Type I and Type II reports, ISO/IEC 27001:2022 Clause 9.1, and PCAOB AS 2201 on testing operating effectiveness. This page is not legal advice, does not determine YOUR obligations, does not decide which frameworks apply to you, does not start a clock, does not file with an auditor, and does not issue an auditor's opinion. Only your auditor can conclude that a control is effective.

What this page is, and what it is not

Audience: a security lead who has assembled the evidence, been told a control was not effective, and cannot see the difference between what they provided and what was wanted.

This page is not legal advice and it is not an assurance opinion. It does not determine YOUR obligations, does not decide which frameworks apply to your organisation, does not start a clock, does not file with an auditor, and cannot conclude that any of your controls are effective. That conclusion belongs to the practitioner performing the engagement. Last verified 10 September 2026.

The distinction that resolves most confusion is between two questions asked in sequence. Would this control, as designed, prevent or detect the thing it is meant to address? That is design effectiveness, and configuration plus a narrative can answer it. Did it actually do so, every time, for the whole period? That is operating effectiveness, and only a complete population plus a sample can answer it. Teams that fail an audit having done real security work have usually answered the first question thoroughly and the second not at all.

  • Design is about logic; operation is about consistency. Different evidence answers each, and one does not substitute for the other.
  • The population is the foundation. Without a complete population, a sample proves nothing, because nobody can tell what it was drawn from.
  • One unexplained exception can undo a period of good operation, because it shows the control was bypassable.
  • Preventive controls are cheaper to evidence than detective ones. A configuration that makes the failure impossible needs no sampling.

Design effectiveness compared with operating effectiveness

First column is the dimension, then how each is established, then what auditors most often find missing. Last verified 10 September 2026. Not legal advice.

Design effectiveness compared with operating effectiveness (not an assurance opinion; not legal advice; last verified 10 September 2026)
DimensionDesign effectivenessOperating effectivenessWhat is usually missing
Question answeredIf it operates as described, would this control address the risk?Did it operate as described, consistently, throughout the period?Nothing. This is the part teams do well.
TimeframeA point in time — usually the end of the period, or the report dateA period of time, with the whole period coveredEvidence for the first and last month only, leaving the middle unevidenced.
Primary evidenceThe control narrative, the policy or procedure, and the configuration showing enforcementThe complete population of instances, a sample tested against the criteria, and the exceptions with explanationsThe population. It is the single most common gap in a first audit.
Typical artifactA branch-protection configuration export, a multi-factor policy export, an alert-rule definitionEvery merged change in the period with its approvals; every joiner and leaver with provisioning and revocation timestamps; every quarter's access review with reviewer decisionsPopulations reconstructed from memory or from partial exports whose completeness cannot be shown.
Where it failsThe narrative describes a control the configuration does not enforceThe control was real but enforcement had a gap, a quarter was skipped, or an administrator bypassed it onceThe record of the exception. Auditors handle disclosed exceptions routinely; they cannot handle ones they discover.
How to strengthen it cheaplyMake the control preventive in configuration, so the failure is impossible rather than discouragedAutomate the population export, and run the control on a fixed cadence with the run recordedNothing prevents both. Preventive design plus a machine-generated population is the cheapest defensible combination.
Report type it supportsA SOC 2 Type I report, and the design half of a Type IIA SOC 2 Type II report, ISO/IEC 27001 certification-audit testing, and PCAOB AS 2201 operating-effectiveness testingCompanies expecting a Type II having only assembled Type I evidence.

Populations and sampling

Sample sizes are the auditor's judgement, not yours — they depend on control frequency, risk, whether the control is manual or automated, and the standards the practitioner applies. What is entirely yours is the population. These are illustrative population definitions with the sampling shape auditors commonly apply; treat the sizes as indicative rather than as a rule, and confirm with your practitioner. Last verified 10 September 2026. Not legal advice.

  • Completeness has to be demonstrable, not asserted. A platform export with its own timestamp beats a spreadsheet somebody assembled.
  • Reconcile populations across systems where you can — human-resources records against identity accounts, asset inventory against scanner coverage. The reconciliation is itself strong evidence.
  • Do not curate. Providing a filtered population is the fastest way to turn a control finding into a question about integrity.
  • Disclose exceptions before they are found, with the cause and the correction. An exception you raised is a working control with a hiccup; one the auditor found is a control nobody was watching.
  • For low-frequency controls, the calendar is the control. Four quarterly reviews with four dated records is the entire population, so a missed quarter cannot be recovered.
Population and sampling shape by control frequency (indicative, not an audit methodology; not legal advice; last verified 10 September 2026)
Control frequencyWhat the population isSampling shape commonly appliedHow to make it defensible
Continuous or automated (for example, enforced required review on merge)Every instance in the period — every merged changeOften tested by examining the configuration plus a small number of instances, because a correctly configured automated control operates uniformlyShow the configuration was in force for the whole period, including its change history, and show the bypass list was empty or that each bypass was recorded.
Many times per day (for example, deployments)The full deployment record for the periodA sample, with the size set by the practitionerExport the population from the platform rather than compiling it by hand, and keep the export with its date.
Event-driven (for example, joiners and leavers)Every person who joined, changed role, or left in the periodA sample of each event typeReconcile the human-resources list against identity-provider accounts, so the population is provably complete rather than asserted.
Monthly (for example, patch or vulnerability review)Twelve occurrences for a twelve-month periodA subset of months, typically a handfulRecord each occurrence with its date and owner as it happens. A missed month is an exception you cannot backfill.
Quarterly (for example, access review)Four occurrences for a twelve-month periodFrequently all of them, because the population is smallWith so few instances, every occurrence must exist and be dated. One skipped quarter is a twenty-five per cent failure rate on that control.
Annual (for example, risk assessment, plan test, policy review)One occurrenceThat one occurrenceIt has to have happened inside the period and be dated inside it. An assessment dated two months after period end does not cover the period.
Incident-driven (for example, incident response)Every incident and near miss in the periodA sample, or all of them if fewInclude low-severity events. A register showing only major incidents suggests the detection control is not operating.

SOC 2 Type I compared with Type II

The difference is exactly the design-versus-operation distinction, expressed as two report types. It is worth stating plainly because Type I is often bought in the belief that it is a lighter Type II, when it answers a different question. Last verified 10 September 2026. Not legal advice.

SOC 2 Type I compared with Type II (not a recommendation of either for your situation; not legal advice; last verified 10 September 2026)
DimensionType IType IIPractical consequence
What is examinedThe description of the system and the suitability of control design as at a specified dateThe same, plus the operating effectiveness of controls throughout a specified periodType II is the one that requires populations and sampling. Type I generally does not.
Evidence neededNarrative, policies, and configuration as at the dateAll of that, plus period populations, samples, and exception explanationsThe effort difference is concentrated entirely in the operating-effectiveness evidence.
PeriodA single dateA period, commonly three to twelve monthsYou cannot produce a Type II for a period that has already passed without evidence that existed during it. That is why starting collection early matters.
What buyers usually wantOccasionally accepted as an interim stepUsually the report enterprise buyers ask forCheck what your prospects actually require before scoping. A Type I that nobody accepts is a cost with no revenue attached.
Common misunderstandingTreated as a cheaper Type IITreated as a longer Type IThey answer different questions. Type I says the design is suitable; only Type II says the controls operated.
Bridge to the next reportDesign work carries forwardThe period must be continuous with the prior report, or a bridge letter covers the gapPlan the period boundaries deliberately, or you will need to explain a gap to every customer who reads the report.

Which kind of authority each expectation carries

Sampling and effectiveness talk is full of numbers people quote as rules. Label each source. Last verified 10 September 2026.

Authority behind effectiveness expectations (not legal advice; last verified 10 September 2026)
StatementWhich kind of authorityWhat it does not mean
PCAOB AS 2201 requires the auditor to obtain evidence about the operating effectiveness of controls, with the nature, timing and extent of testing based on risk.Legal requirement in substance for issuers, through Sarbanes-Oxley obligations implemented in auditing standards, where systems are relevant to financial reporting.Does not apply to a private company with no issuer obligations, and does not fix a sample size. Risk drives extent.
AICPA Trust Services Criteria, and AICPA guidance on design and operating effectiveness and on Type I and Type II reports.Professional attestation criteria and professional guidance, applying because you sought a report.Do not publish a table of mandatory sample sizes for you to follow. The practitioner determines the extent of testing.
ISO/IEC 27001:2022 Clause 9.1 monitoring, measurement, analysis and evaluation, with Clauses 9.2 and 9.3.Certification requirement where you seek or hold certification.Requires that you evaluate performance on your stated cadence. It does not prescribe sampling or effectiveness testing in the attestation sense.
Specific sample sizes circulated in compliance communities — for example, a fixed number of items per control frequency.Industry convention only, derived from common practice and from audit-guidance heuristics.Not a rule, and not binding on your practitioner. Use conventions to prepare, then confirm the actual extent with your auditor.
Provide the complete population and let the auditor sample from it.ShipReady Metrics recommendation, and the expectation implicit in every attestation and audit standard cited here.Not optional in practice. Selecting your own sample and providing only that undermines the whole test.
Prefer preventive controls enforced in configuration over detective controls that rely on someone looking.ShipReady Metrics recommendation, and common industry best practice.Not a requirement. It is the cheapest route to operating effectiveness, because a control that cannot be bypassed needs far less sampling.
Disclose exceptions yourself, with cause and correction.ShipReady Metrics recommendation.Does not guarantee the exception is accepted. It changes the finding from a control nobody monitored into a monitored control with a recorded failure.

Checklist

Run this per control, not per framework. Not an assurance opinion and not a determination that any framework applies to you. Last verified 10 September 2026. Not legal advice.

  • For each in-scope control, can we state the risk it addresses and how it addresses it, in a written narrative?
  • Does the configuration actually enforce what the narrative claims — checked, not assumed?
  • Is the control preventive where it could be, rather than relying on someone noticing?
  • For each control, what is the population for the period, and can we export it from a system rather than assembling it by hand?
  • Can we demonstrate that the population is complete, ideally by reconciling it against a second source?
  • Does our evidence cover the whole period, including the middle months, rather than the start and the end?
  • For low-frequency controls, does every scheduled occurrence exist with a date inside the period?
  • Do we have a list of exceptions we already know about, each with a cause and a correction?
  • Was the control's own configuration changed during the period, and if so do we hold the change history?
  • Where a control is manual, is the person's judgement recorded — not just that a review happened, but what was decided?
  • Have we avoided curating any population we provide?
  • Do we know whether our buyers need a Type I or a Type II, before scoping the engagement?

What to do now

Ordered so the gap that fails most first audits closes first. None of these steps issues an assurance opinion, determines YOUR obligations, or files anything with an auditor.

  • Pick your three highest-risk controls and write the population definition for each. If you cannot define it, you cannot evidence operation.
  • Automate each population export and run it now, for the current period. Machine-generated populations with their own timestamps are the strongest form.
  • Reconcile one population against a second source — human-resources records against identity accounts is the usual starting point — and keep the reconciliation.
  • Convert one detective control into a preventive one in configuration. It permanently reduces the evidence burden on that control.
  • Export the configuration history of your enforcing controls, so you can show enforcement held for the whole period rather than on the day you looked.
  • Assemble the known-exception list yourself, with cause and correction per item, before fieldwork begins.
  • Fix the calendar for low-frequency controls and record each occurrence as it happens, since a missed quarter cannot be backfilled.
  • Make manual controls record the decision, not just the event — a reviewer's per-account decision rather than a note that a review occurred.
  • Confirm with your buyers which report type they need, and with your practitioner what extent of testing to expect.
  • Re-verify the cited criteria annually against the primary sources below and record the date. Ours says 10 September 2026.

Where this shows up in ShipReady Metrics

Only shipped behaviour is described here, including what it deliberately does not do. This product cannot conclude that a control is effective. It does not perform audit testing, does not select statistical samples on an auditor's behalf, does not issue an opinion, and does not file with an auditor.

If you already have a session: signed-in app → Compliance holds evidence collection for control-mapped artifacts across starter control subsets, and evidence review. Evidence review is where the human overlay lives, and it is worth being precise about it: a named human accepting a manual row renders that row met, and rejecting it renders it a gap, with a timestamp. That is an accept-or-reject record by a person, deliberately not an automated verdict — because ingestion can establish that an artifact exists, and only a person can judge whether it evidences the control. It produces a timestamped compliance artifact, not a forensic chain of custody, not a downloadable evidence binder, not a regulator filing pack, and not an auditor's opinion.

For the automated side, connectors give you materials rather than conclusions: read-only GitHub ingest across pull requests, Actions runs, and the Dependabot, code-scanning and secret-scanning alert feeds; GitLab security-findings ingest; vulnerability ranking using CISA KEV, EPSS and CVSS with cross-source deduplication and blast-radius search over captured dependencies; and Delivery Health DORA posture from connected sources. Those help you assemble populations. They do not test them.

Readiness figures are internal indicators. They are not certification, not CE marking, not an assurance opinion, and they must not be read as a statement that a control operated effectively.

Primary sources (last verified 10 September 2026)

Each source is labelled by the kind of authority it carries.

AICPA Trust Services Criteria are professional attestation criteria, and the AICPA's SOC 2 guidance covers design and operating effectiveness and the Type I and Type II distinction as professional guidance. ISO/IEC 27001:2022 Clauses 9.1, 9.2 and 9.3 are certification requirements. PCAOB AS 2201 is the auditing standard governing operating-effectiveness testing in an audit of internal control over financial reporting, reaching issuers through Sarbanes-Oxley obligations. Sample-size figures circulating in compliance communities are convention, not authority. Not a complete list, and not legal advice.

The repository page in this cluster covers where populations and samples live, the continuous-collection page covers keeping them fresh, and the SOC 2 framework guide covers the report types at a level above the evidence set.

Frequently asked questions

Is this legal advice?

No, and it is not an assurance opinion either. It is an operational explanation of how effectiveness is evidenced, with each source labelled by the kind of authority it carries. Only your auditor can conclude that a control is effective. This page does not determine YOUR obligations, does not start a clock, does not file anything with an auditor, and cannot issue an opinion.

What is the difference between design and operating effectiveness?

Design asks whether the control, as described and configured, would prevent or detect the risk — answered at a point in time with a narrative and a configuration export. Operating asks whether it actually did so throughout the period — answered with a complete population, a sample drawn from it, and explanations for any exceptions. Most first-audit failures are design work with no operating evidence.

How large a sample will the auditor take?

That is the practitioner's judgement, driven by control frequency, risk, and whether the control is manual or automated. Sample-size tables circulating in compliance communities are convention rather than authority, and no standard cited here publishes sizes for you to follow. Your job is the complete population; confirm the expected extent of testing with your auditor rather than guessing.

One exception happened. Is the control now ineffective?

Not necessarily, and how you present it matters. An exception you identified, explained, and corrected is a monitored control with a recorded failure, which auditors handle routinely. An exception they discover in a population you provided suggests nobody was watching, which is a harder position. Assemble the known-exception list yourself, with cause and correction, before fieldwork.

How does the human accept in ShipReady Metrics relate to effectiveness?

It records a judgement, not a conclusion about effectiveness. A named human accepting a manual evidence row renders it met and rejecting it renders it a gap, with a timestamp. That design is deliberate: ingestion can show an artifact exists, but only a person can judge whether it evidences the control. It is a timestamped compliance artifact, not an auditor's opinion, and readiness figures are internal indicators rather than assurance.

Published by ShipReady Metrics, an evidence-based technology and compliance intelligence platform. This guide is educational and vendor-neutral.