Product roadmap: where we are, and where we are headed

Updated

This is a public, honest roadmap. It marks what is live in ShipReady Metrics today, what we are building, and what we are exploring — deliberately as directions, not dated commitments. Everything past Live is subject to change, and we would rather say so than promise a calendar.

We publish it this way on purpose. The product exists to distinguish what is measured from what is merely asserted; a roadmap that quietly promised features on dates it could not honor would violate the same principle. So we tell you where the foundation is real and shipped, where work is genuinely in progress, and where we are still asking open questions we have not answered yet.

How to read this page: directions, not commitments

A roadmap is a statement of intent, and intent changes as evidence arrives. We would rather be honest about that than set an expectation we cannot honor. So nothing on this page is a promise of a feature on a date. There are no delivery dates here, and that is deliberate — a dated commitment we later missed would cost more trust than a direction we described plainly and then adjusted.

Everything below carries one of three statuses. Live means it is shipped and running in the product today, and the claims are limited to capability that actually exists. Building means work is genuinely underway, framed as a direction rather than a guarantee. Exploring means we consider the problem important and are studying it, with open questions we have not yet resolved — reading an Exploring row as a commitment would be reading it wrong.

One more framing that governs the whole page: a direction earns a place on this roadmap only if we can imagine sourcing it honestly. We do not add a lens or a score we cannot measure, because the product's core rule is that anything unmeasurable reads Not Measured rather than a fabricated grade. A roadmap item we could only fake is not a roadmap item.

What is live today: the shipped foundation

The directions further down build on a foundation that is real and running, not aspirational. It is worth being concrete about what Live means, because the rest of the roadmap only makes sense as an extension of it.

At the base is evidence-based scoring: a family of 0-100 scores — nine today, spanning engineering and readiness lenses — governed by a scoring-honesty invariant, so a subscore never contradicts the data printed beside it and anything the system cannot measure reads Not Measured rather than a low or invented grade. On top of that sits a canonical-control compliance model: read-only connector evidence, readiness derived only from evidence a named human has accepted, an append-only trail, a SHA-256 chain-of-custody linking each artifact to the period and system it came from, and an auditor bundle that exports the binder, hashes, and manifest.

The compliance surface goes deep on SOX specifically — a control register, testing and sampling, IT general controls, and segregation of duties with a recorded grantor and automated conflict detection over the duty population, all append-only. It carries AI provenance and attestation: an AI-produced artifact is routed to a named human attester, the record captures what that attester saw with a frozen provenance binding, and decision records are append-only. Around all of it run continuous control monitoring, a publishable Trust Center, and risk, third-party (TPRM), and policy platforms. That is the ground the roadmap stands on.

The roadmap at a glance

The table gathers the whole picture in one place: what is shipped, what is in progress, and what we are still studying — each with the honest caveat or open question that comes with it. Statuses are current as of the date on this page and are subject to change.

ShipReady Metrics roadmap — directions, not commitments; status and honest caveat for each (subject to change)
DirectionStatusWhat it meansThe honest caveat or open question
Evidence-based 0-100 scoring with an honesty invariantLiveNine evidence-based scores across engineering and readiness lenses; anything unmeasurable reads Not MeasuredA score is only as good as its connected sources; unconnected scope reads as unmeasured, by design
Canonical-control compliance engineLiveRead-only connector evidence, named-human acceptance, append-only trail, SHA-256 chain-of-custody, auditor bundleIt evidences a program; it does not certify compliance — that determination stays with your auditor
Deep SOX module including segregation of dutiesLiveControl register, testing, ITGC, and SoD with a recorded grantor and conflict detection over the duty populationConflict dispositions and acceptances are recorded human decisions, not automated away
AI provenance and human attestationLiveRoute an AI artifact to a named human attester; frozen provenance binding; append-only decision recordsA named person owns the attestation; automation informs it, it does not sign it
Continuous monitoring, Trust Center, risk / TPRM / policyLiveStanding control monitoring, a publishable trust page, and risk, third-party, and policy platformsCoverage depends on what you connect; gaps are shown as gaps
Deeper metric coverage across every lensBuildingExtending measured scoring past engineering into compliance, CFO cost-and-value, and board-level viewsWe add a lens only when it can be sourced honestly; one we cannot measure reads Not Measured, not shipped
Provenance and separation of duties for AI-authored codeBuilding / ExploringBinding a recorded author, reviewer, and conflict check to AI-generated changes, extending the live SoD engineReal independence between an AI author and an AI reviewer is an open design problem, not a solved feature
AI Test Data ManagementExploringCompliant, privacy-preserving synthetic test data so teams stop testing against copies of productionDirection only; synthetic-data fidelity and re-identification risk are the questions we are studying
Continuous, machine-verifiable assuranceExploringEvidence a machine can both produce and independently check, with a named human still accountableThe hard part is the independent checker; we will not ship a reviewer that shares the author's blind spots

Direction: deeper metric coverage across every lens

Today's scores lean toward the engineering picture — security readiness, delivery, technical debt, cloud and agent health, modernization, and lifecycle — with compliance readiness alongside them and an early cost view already live in the remediation-cost estimator and Value-at-Stake headline. The direction we are building toward is coverage that speaks to more than one reader without losing the measured-only discipline: the same underlying evidence rendered for the CFO as cost and value at stake, and for the board as a defensible readiness posture, not just for the engineering lead as delivery metrics.

The measurement doctrine we extend from is deliberately multi-dimensional, because single metrics get gamed. The DORA delivery metrics — deployment frequency, lead time for changes, change failure rate, and time to restore service — balance speed against stability so that shipping faster cannot masquerade as shipping better. The SPACE framework (Forsgren, Storey, and colleagues, 2021) spans satisfaction, performance, activity, communication, and efficiency for the same reason: so that gaming one dimension shows up as damage in another. Extending coverage means adding lenses that hold that balance, not adding a single number an optimizer can chase.

The constraint on this direction is the honesty invariant, and it is a real constraint rather than a slogan. A CFO or board lens is worth building only to the extent its inputs can be traced to their source. Where a figure cannot be sourced honestly, the correct behavior is to render it Not Measured — so parts of a lens may stay unmeasured on purpose rather than be filled with a confident estimate. That is the difference between a metric you can defend to an auditor and a dashboard that reads green regardless of the evidence beneath it.

Direction: provenance and separation of duties for AI-authored code

Control theory has one rule that predates computing: the maker cannot be the checker. Segregation of duties, tester independence, and the auditor's arm's-length stance all encode it, and IT general controls over change management exist to make sure the person who wrote a change is not the only person who approved it. When an AI writes the code and an AI reviews it, that separation does not disappear on its own — it collapses quietly, and a reviewer that fails precisely where the author fails is not a control, it is a mirror.

The direction we are building is to bring the discipline the platform already applies to human change management to AI-authored changes: a recorded author, a recorded reviewer, and a conflict check binding the two, extending the live SoD engine — which already records a grantor and detects conflicts over a duty population — into the AI software-development lifecycle. The aim is that an AI-generated change carries the same provenance and independence evidence a human-authored one would, rather than arriving as an approval with no checker behind it.

We mark this Building and Exploring on purpose, because the hard part is genuinely unsolved. Two systems that share a base model, a prompt scaffold, or a training corpus can be confidently, identically wrong, and recording who reviewed what does not by itself make the reviewer independent of the author. Engineering that independence — a different model lineage, differently sourced data, or a human placed at the decision point — is an open problem we are working through rather than a feature we have finished. We write about the underlying question at length in a separate essay on separation of duties when AI writes code.

Direction: AI Test Data Management

Testing is where a lot of sensitive data quietly leaks its containment. Teams under pressure to reproduce a bug or load-test a feature reach for the fastest realistic dataset available, which is often a copy of production — real customer records sitting in a lower environment with weaker controls, a pattern that sits crosswise to the data-minimization and access expectations behind regimes like the GDPR and HIPAA. The problem gets sharper as AI-assisted development speeds up the tempo of testing.

The direction we are exploring is compliant test data management: privacy-preserving synthetic data that preserves the shape and edge cases a test needs without carrying the real identities a test does not. Done well, it lets teams stop testing against production copies while keeping the fidelity that makes a test meaningful, and it produces evidence — of how test data was generated and governed — that fits the same append-only, provenance-first spine the rest of the platform runs on.

This one is firmly Exploring, and the caveats are the substance rather than a footnote. Synthetic data trades on a hard tension: too faithful and it can leak or allow re-identification of the very records it was meant to protect; too abstracted and it stops exercising the behavior a test exists to catch. Where that line sits, and how to evidence that a given synthetic set is both useful and safe, are exactly the questions we are studying before we would commit to a shape for it. We treat the topic in a dedicated guide on AI test data management.

Direction: continuous, machine-verifiable assurance

The deepest direction is also the one furthest out. Assurance is moving from a snapshot an auditor signs on a calendar toward a standing, queryable state that tracks the live system — the trajectory behind continuous control monitoring, which the platform already runs. The endpoint of that trajectory is machine-verifiable assurance: controls whose operation emits signed, timestamped, reproducible evidence that a reviewer can check without re-interviewing the people who ran it. The live SHA-256 chain-of-custody and append-only trail are the first pieces of that; the direction is to make more of the evidence both machine-produced and machine-checkable.

The word that matters here is checkable, and it is where the hard problem lives. Evidence a machine can produce is not yet evidence a machine can independently verify — if the same system that generated an artifact also vouches for it, an approval is an assertion dressed as evidence. Real machine-verifiable assurance needs an independent checker: build-provenance and supply-chain attestation approaches like the open SLSA specification point at part of the answer, letting a downstream party verify what produced an artifact without trusting the producer's word for it.

Whatever this becomes, one property is non-negotiable and shapes the exploration: a named human stays accountable. Both the EU AI Act's human-oversight requirement for high-risk systems (Regulation (EU) 2024/1689) and the accountable-owner idea in management-system standards like ISO/IEC 42001:2023 insist that automation may inform a decision but a person must own it. So the direction is not to remove the human — it is to give that human evidence strong enough that their accountability means something, evidence an independent party could re-derive rather than take on faith.

How we choose what to build — and what we will not promise

A few principles decide what earns a place here, and they are the same ones that govern the product itself. They are worth stating plainly, because they explain both what is on the roadmap and what is deliberately absent from it.

  • We measure outcomes, not activity: a direction earns priority when it moves a number a customer relies on — readiness, evidence, defensibility — not because it lengthens a feature list.
  • We do not add what we cannot source honestly: any lens or score has to trace to real evidence, or it reads Not Measured rather than getting a fabricated grade.
  • We keep a named human accountable: automation gathers and organizes evidence, but a person owns each control and each material decision, consistent with the EU AI Act and management-system standards.
  • We favor append-only, provenance-first evidence: new capability should extend the signed, timestamped, chain-of-custody spine, not bolt on a claim with no traceable source.
  • We publish directions, not dates: we would rather describe a direction plainly and adjust it than promise a delivery we cannot guarantee, so nothing here is a dated commitment.

Frequently asked questions

Is this roadmap a commitment to deliver these features?

No. It is a statement of direction, and directions change as evidence arrives. Only the items marked Live are shipped and running today. Building and Exploring items describe where we intend to go and what we are studying — they are deliberately not promises, and there are no delivery dates because a dated commitment we later missed would cost more trust than a direction described plainly and then adjusted.

What do Live, Building, and Exploring mean?

Live means shipped and running in the product today, with claims limited to capability that actually exists. Building means work is genuinely underway, framed as a direction rather than a guarantee. Exploring means we consider the problem important and are studying it, with open questions we have not yet resolved. Reading an Exploring row as a commitment would be reading it wrong.

What is actually shipped today versus aspirational?

Shipped and Live: evidence-based 0-100 scoring with a scoring-honesty invariant; the canonical-control compliance engine with read-only connector evidence, named-human acceptance, an append-only trail, a SHA-256 chain-of-custody, and an auditor bundle; a deep SOX module including segregation of duties with a recorded grantor and conflict detection; AI provenance and human attestation; continuous control monitoring; a Trust Center; and risk, TPRM, and policy platforms. Everything under Building or Exploring is a direction, not a shipped feature.

Why won't you give delivery dates?

Because the product exists to separate what is measured from what is merely asserted, and a roadmap that promised features on dates it could not honor would violate that same principle. We would rather tell you honestly where the foundation is real, where work is in progress, and where we are still asking open questions, than set a calendar expectation we cannot guarantee to meet.

What is AI Test Data Management and why is it only Exploring?

It is the idea of compliant, privacy-preserving synthetic test data, so teams can stop testing against copies of production. It is Exploring rather than Building because the caveats are the substance: synthetic data that is too faithful can allow re-identification of the records it was meant to protect, while data that is too abstracted stops exercising the behavior a test exists to catch. Where that line sits, and how to evidence that a synthetic set is both useful and safe, are the questions we are studying before committing to a shape for it.

See what is live today, and hold the rest to this page

Connect your systems read-only and see the shipped foundation — evidence-based scoring, canonical-control compliance, AI provenance, and continuous monitoring — with anything unmeasured shown as Not Measured, never a green estimate.