- Understanding Maturity Assessment and Its Purpose
- Key Benefits of Conducting Maturity Assessments
- How Maturity Assessment Is Executed in Your Organization
- Common Applications: From Procurement to Digital Transformation Maturity
- Interpreting Results and Prioritizing Improvement Initiatives
- Maturity Assessment Tools and Implementation Best Practices
- How Orca Supplies the Evidence a Cloud Security Maturity Score Needs
- Frequently Asked Questions About Maturity Assessments
Key Takeaways
- A maturity level typically reflects how systematically and repeatably a capability is performed, rather than how well it performed once. The Department of Energy’s C2M2 says practices at its lowest performed level may still be ad hoc.
- CMMI separates two things that are easy to confuse. Capability levels run 0 to 3 and apply to a single practice area. Maturity levels run 0 to 5 and apply to a predefined set of them.
- Evidence separates an assessment from a survey. ISO/IEC 33002 sets the minimum requirements that make results objective, consistent, repeatable, and representative.
- Calibration is the step teams skip. The UK’s public-sector commercial framework assigns each organization a peer reviewer whose stated job is consistent interpretation of the criteria.
- Level claims about a cloud security capability rest on artifacts covering assets, configuration, identity, and data. Orca SideScanning collects those without agents.
A maturity assessment scores a defined capability against a documented set of levels or criteria, using evidence rather than opinion. The method belongs to no single field. Procurement teams, data teams, and security teams all run one. This guide covers what a level asserts, what evidence a level claim requires, and how the resulting gaps become an order of work.
One security capability runs through the article as a worked example, because a method explained in the abstract is hard to check. The method itself stays domain-independent.
Readers who arrive skeptical have a fair reason. The standing objection is that scores are self-reported and that organizations rate themselves generously. That objection has a mechanical answer, and the sections below give it.
Understanding Maturity Assessment and Its Purpose
A maturity assessment measures a capability against a scale that someone else defined and published. The scale describes states rather than scores out of ten. Each state carries conditions, and an organization occupies a state only when it meets them. The instrument answers one question: how dependably does this capability produce its result when the person who usually does it is unavailable.
That question is different from the one a security posture assessment answers. Posture describes the current state of an environment. Maturity describes the reliability of the process that produces that state. An organization can hold a clean posture on Tuesday because two experienced engineers worked the weekend. It can still sit at the bottom of every maturity scale.
Levels, Dimensions, and Evidence
Levels are cumulative and conditional. The Department of Energy’s Cybersecurity Capability Maturity Model sets four levels, MIL0 through MIL3. They apply independently to each of its ten domains. An organization can sit at MIL1 in one domain and MIL3 in another. To earn a level in a domain, it must perform every practice at that level and at the level below.
The level definitions themselves show what a level asserts. At MIL1, initial practices are performed but may be ad hoc. At MIL2, practices are documented and adequate resources support them. At MIL3, policies guide the activity, and responsibility and authority are assigned to named personnel. Those personnel have the skills required, and the effectiveness of the activity is evaluated and tracked. Nothing in that ladder rates quality. It rates repeatability.
Dimensions are the second axis. A model divides a broad capability into domains or practice areas so a single score cannot hide an uneven organization. C2M2 uses ten domains. The UK Government’s commercial framework uses eight themes containing 27 practice areas. Splitting the capability is what makes the result usable. An average across ten domains tells a program owner nothing about which one to fix.
Evidence is the third element, and it is the one that decides whether the exercise is worth running. ISO/IEC 33002:2015 was confirmed as current in January 2026. It defines the minimum requirements for an assessment “that will ensure assessment results are objective, consistent, repeatable, and representative of the assessed processes.” Repeatable means a second assessor reaches the same rating from the same artifacts. Without artifacts, nothing is repeatable.
What a Maturity Score Is Not
A maturity score is not a risk measurement. A cyber risk assessment estimates exposure and potential loss. A maturity assessment scores the capability that manages exposure. A mature capability can still carry high residual risk, and an immature one can look calm because nobody has attacked it yet.
It is also not a compliance status. Meeting the conditions for a level is not evidence of conformance with any regulation. A maturity rating generally does not substitute for evidence that a specific compliance control has been met. Two further distinctions are worth stating once. A model is the published scale. An assessment is the act of measuring against it. And a level is not a grade. An organization that stops at MIL1 in a low-consequence domain has usually made a choice rather than an error.
Key Benefits of Conducting Maturity Assessments
The exercise earns its keep by forcing specific claims into the open and attaching each one to a document somebody can read. The number itself matters least. Four outcomes follow, and the fourth is the one that survives a change of leadership.
- A shared vocabulary. Two teams stop arguing about whether patching is “good” and start comparing their answers against the same written conditions.
- A gap list with owners. Every unmet condition identifies something concrete that is missing or incomplete, such as an artifact, process, role, metric, or control, and gives the program owner something specific to address.
- A defensible funding case. A level claim tied to named evidence survives the budget review that a subjective assessment does not.
- A baseline that outlasts people. Re-running the same instrument in a year measures change instead of measuring the current leader’s optimism.
How Maturity Assessment Is Executed in Your Organization
Running one well is mostly a scoping and evidence problem. The worked example in this section is a cyber maturity assessment of one capability. The subject is how the organization handles known vulnerabilities in its cloud estate.
Scoping the Capability
Pick the smallest unit that still tells you something. C2M2 calls this unit the function, and defines it loosely on purpose. Its examples include a line of business, a facility, a network security zone, or assets residing in the cloud. Its own guidance is to pick a scope managed homogeneously, so two groups are not forced to merge two different answers into one.
The model also recommends identifying the fewest assessments that still give an accurate picture. Scoping a whole enterprise as one unit produces a score nobody can act on. Scoping every team separately produces a filing cabinet.
For the example, the scope is vulnerability handling for production cloud workloads in one business unit. That gives a named owner, a bounded asset set, and a group of people who all work the same way.
Collecting Evidence Instead of Opinions
The evidence question is concrete: what artifact would have to exist for this level claim to be true. Answering it before the workshop is what turns an interview into an assessment. A practical evidence set can be grouped into four categories:
- Definition artifacts. A written procedure, a policy, or a service level that states what is supposed to happen and by when.
- Record artifacts. Ticket histories, change records, and scan output showing the activity happened on dates somebody can check.
- Assignment artifacts. Named ownership per asset class, and a register of exceptions with expiry dates rather than open-ended ones.
- Coverage artifacts. Proof that the activity reached everything in scope, which requires current asset coverage and context rather than a spreadsheet from last quarter.
Apply those to the example. A claim that vulnerability handling is documented needs the remediation standard. A claim that it is measured needs time-to-remediate by severity for a stated period. A claim that it is owned needs the per-asset-class owner list. A claim that it is complete needs coverage figures for the whole scope. That is the hardest of the four, because a scanner covers what it was pointed at.
The C2M2 practice text for its vulnerability objective shows the same progression. At its first level, vulnerability assessments are performed “at least in an ad hoc manner.” At the next, they are performed periodically and according to defined triggers such as system changes, and identified vulnerabilities are analyzed and prioritized. At the highest, they are performed by parties independent of the operations of the function. Monitoring also includes review confirming that the actions taken were effective.
Scoring, Calibration, and Disagreement
Scoring is where self-assessment earns its bad reputation, and the fix is procedural. C2M2 asks workshop participants to choose from four responses per practice: Not Implemented, Partially Implemented, Largely Implemented, or Fully Implemented. It then counts a practice as performed only at Largely or Fully. A defined scale with a defined threshold removes most of the room to round upward.
Calibration is the step that gets skipped. The UK Government’s Commercial Continuous Improvement Assessment Framework builds it in. Each participating organization is matched with a peer reviewer from another organization. The guidance states the reviewer’s role plainly: to ensure consistency of interpretation of the assessment criteria. Its own attainment wording is evidence-first, separating “no or isolated evidence” from “significant but inconsistent evidence.” The same framework marks its mid-cycle re-baseline as provisional, because that step is not peer reviewed.
Disagreement between assessors is useful, and burying it wastes the most informative output of the day. Two people rating the same practice differently usually means the scope was ambiguous or the artifact does not cover what one of them assumed. Record which practices split the room and why. That list is often more actionable than the score. Note also that this exercise is not running a point-in-time cloud security assessment, which looks for findings in the environment. Here the subject is the capability, and the findings are missing artifacts.
Common Applications: From Procurement to Digital Transformation Maturity
The method transfers across fields with the vocabulary intact. What changes is the capability under measurement and the published scale used to measure it.
| Domain | Capability being scored | Published model most often used |
| Cybersecurity (cyber maturity assessment) | Vulnerability handling, incident response, identity and access control | DOE C2M2 v2.1, ten domains at MIL0 to MIL3. The NIST Cybersecurity Framework is used too, though its Tiers rate rigor rather than maturity |
| Procurement (procurement maturity assessment) | Commercial planning, contracting, supplier and contract management | UK Government CCIAF, 27 practice areas across 8 themes |
| Digital transformation (digital transformation maturity assessment) | Digital strategy, readiness, and adoption of automation | The European Commission’s Digital Maturity Assessment framework |
| Data management (data maturity assessment) | Data quality, architecture, and lifecycle control | DCAM v3, 34 required capabilities and 101 sub-capabilities |
| Data governance (data governance maturity assessment) | Ownership, stewardship, and policy enforcement over data | DCAM v3 governance content, alongside data security posture management tooling for the evidence |
| Artificial intelligence (ai maturity assessment) | Model inventory, evaluation practice, and risk controls | No standards-body model. IEEE-USA published one in July 2024 built on the NIST AI Risk Management Framework |
| Agile delivery (agile maturity assessment) | Planning cadence, release frequency, and feedback loops | No standards-body model. Published scales come from vendors, coaches, and consultancies |
| IT and process generally | Any defined practice area | CMMI V3.0, published by ISACA in April 2023 |
Two rows deserve a note. The non-IT domains are not borrowing loosely from software. The CCIAF requires named evidence per practice area and gives it to a peer reviewer. That is stricter than most internal security assessments manage. DCAM v3, from the EDM Association, defines its capabilities across objectives, questions, artifacts of evidence, and criteria for scoring. Each is rated on a six-point scale whose stated minimum goal is the fifth point.
The newer domains are where the method strains. An ai maturity assessment has no scale carrying the authority of CMMI or C2M2. Published models exist, including a flexible maturity model for AI governance from the IEEE-USA AI Policy Committee, and teams also build their own. Either way, the scale must state its conditions in writing before anyone scores against it. Otherwise the exercise measures confidence.
Interpreting Results and Prioritizing Improvement Initiatives
People ask for an average. What changes anything is the ordered list of unmet conditions. C2M2 puts the point directly: it is not typically optimal for an organization to strive to achieve the highest level in all domains. The organization sets a target level per domain, compares it against the current result, and works only on the differences.
Order the resulting gaps by dependency first and by consequence second. Four inputs decide the sequence.
- Dependency. Some gaps block others. Nothing else in the example advances while asset coverage is unknown, because every other claim is scoped to assets nobody has counted.
- Consequence. Weight each domain by what it protects, which is where ordering work by risk does the work a maturity score cannot.
- Cost and available effort. C2M2 names cost of implementation and resource availability among its prioritization criteria. It also recommends comparing the cost of a level against its benefit.
- Durability. A gap closed by hiring one person reopens when that person leaves. A gap closed by a written standard and a measured service level does not.
The example makes the dependency rule concrete. Suppose the cyber maturity assessment returns a documented remediation standard, per-asset-class owners, and measured time-to-remediate by severity. Coverage figures exist only for accounts onboarded to the scanner. Every higher-level claim about the phases of vulnerability management inherits that hole. Closing the coverage gap is not the most visible initiative on the list, and it is the one that has to go first.
Maturity Assessment Tools and Implementation Best Practices
Tooling for this is thinner than the market suggests, and the free options are better than they look. The DOE publishes self-evaluation tools for C2M2 at no cost, designed so a single function can be evaluated in one day without lengthy preparation. The UK’s commercial framework runs on a government-hosted scoring platform. A spreadsheet works if the conditions and the evidence pointers sit next to the ratings, which is the only feature that actually matters.
Five practices separate an assessment that holds up from one that gets filed.
- Choose a published model before scoping. Writing your own levels while you assess produces conditions shaped to the answers you already have.
- Set the target profile per domain. Deciding the target after seeing results turns the exercise into justification.
- Record the artifact, not the assertion. A rating with a document reference beside it can be re-checked in a year. A rating alone cannot.
- Assign the calibration role explicitly. A peer reviewer or an independent assessor is what the two strongest public frameworks both build in.
- Re-run the same instrument. Changing models between rounds can break trend comparability unless the old and new criteria are explicitly mapped.
Two boundaries keep this exercise from turning into something else. It is not the same activity as sequencing a cloud security program, which decides which capabilities to build in what order. And a low score in one domain often signals unclear ownership rather than weak execution. When a domain scores badly for reasons nobody can explain, check decision rights in a security organization first. In the running example, both boundaries meet in one question. Who approves a patch management exception, and where is that written down.
How Orca Supplies the Evidence a Cloud Security Maturity Score Needs
Orca does not run maturity assessments or assign maturity levels. Its Orca Security Score measures cloud security posture, which is a different measurement. A posture score reflects the current state of cloud risk, while a maturity assessment evaluates how consistently the organization can produce and maintain that state.
The connection to this article is the artifact requirement. Every level claim about a cloud security capability rests on four records. They cover which assets exist, how they are configured, who and what can reach them, and where sensitive data sits. Orca SideScanning collects all four without agents. It reads workload runtime block storage instead of running code on the workload, and covers virtual machines, containers, and serverless functions. Orca states that it discovers and monitors assets across the estate as new ones are added. That turns the coverage claim from an assertion into a record with a date on it, which is what a level claim needs. Orca delivers this as a cloud-native application protection platform.
The division of labor stays clean. Orca supplies the evidence. The assessor still sets the scope, picks the published model, calibrates the ratings against a second reader, and decides the order of work. Teams that want to see what their coverage evidence actually looks like before their next assessment can start there.
Frequently Asked Questions About Maturity Assessments
What is the difference between a maturity model and a maturity assessment?
The model is the published scale, including its domains, its levels, and the conditions for each level. The assessment is the act of measuring one organization against that scale on a given date. C2M2 and CMMI are models. What a team runs on a Tuesday in March is an assessment. The distinction matters when a vendor offers “a maturity model” that turns out to be a questionnaire with no published conditions behind it.
How long does a maturity assessment take?
Less time than most estimates, if the scope is one capability. C2M2 was designed so a single function can be self-evaluated in one day. The preparation is the longer part. The preparation time depends on scope, evidence availability, model complexity, and whether the assessment is self-led or independently reviewed. Skipping that is what makes assessment days run long and produce soft answers.
Should an internal team or an external firm run the assessment?
Both work, and the honest answer is that it depends on what the result will be used for. An internal team knows where the artifacts are and costs nothing. An external assessor is harder to lean on. C2M2 asks for independence at MIL3 in two places: vulnerability assessments, and review of cybersecurity activities. If the score will support a funding request or reach a board, independence is worth paying for. If it is a working document for the team that owns the capability, it usually is not.
How often should an organization reassess?
Published cycles run longer than people expect. The UK’s commercial framework uses a 24-month cycle with a re-baseline at month 12. C2M2 recommends reassessing after major changes in business, technology, market, or threat environments, rather than waiting for the calendar. One caution on the mid-cycle check. If it skips the calibration step, treat the result as provisional, which is how that framework labels its own unreviewed re-baseline.
Can a maturity level be used as evidence for an audit?
No. A level records that a capability meets a set of conditions in a published model. An audit tests specific controls against a specific standard, and the two rarely map cleanly. The artifacts gathered for an assessment are often reusable as audit evidence, which is a real efficiency. The level itself is not.
- Understanding Maturity Assessment and Its Purpose
- Key Benefits of Conducting Maturity Assessments
- How Maturity Assessment Is Executed in Your Organization
- Common Applications: From Procurement to Digital Transformation Maturity
- Interpreting Results and Prioritizing Improvement Initiatives
- Maturity Assessment Tools and Implementation Best Practices
- How Orca Supplies the Evidence a Cloud Security Maturity Score Needs
- Frequently Asked Questions About Maturity Assessments
