Skip to content

What we check

The assessment groups its checks into ten categories. Each category carries a share of the overall score.

Category Weight Checks defined Scored today
Security and privacy 20% 7 7
Feature completeness 12% 3 2
System design 12% 8 1
Data integrity and lifecycle 10% 4 0
Reliability 10% 4 3
Operability, including observability 10% 10 6
Agent readiness 8% 12 11
Scalability and performance 8% 4 1
Maintainability and configuration 5% 4 3
User experience and interfaces 5% 3 1
Total 100% 59 35

The assessment defines 59 checks. Today, 35 of them are scored from source.

  • 22 checks are not assessed by this version. They appear in your result as Not assessed by this version and are excluded from the score. They are never shown as passes.
  • 2 checks need evidence from outside the source, for example from a running system. They appear as Optional follow-up and do not lower the score.

More checks move from defined to scored as the assessment grows.

  • Security and privacy. Identity, authorization, input handling, credentials, and sensitive data.
  • Feature completeness. Stubs, fake success, and missing failure states behind features that look finished.
  • System design. The boundaries between your code and the services it depends on.
  • Data integrity and lifecycle. No checks in this category are scored yet.
  • Reliability. Error handling, retries, and what happens when a dependency is down.
  • Operability, including observability. Health checks, deployment and rollback, configuration, and whether failures leave a useful signal.
  • Agent readiness. Entry points, agent instructions, setup, and the build, lint, type, and test commands an agent uses to check itself.
  • Scalability and performance. Unbounded inputs and work that grows without limit.
  • Maintainability and configuration. Supported dependencies, consistent configuration, and tests that cover what is claimed.
  • User experience and interfaces. Input and accessibility basics on the surfaces people use.

Every check ends in one of six outcomes.

Outcome Meaning
Pass The check found what it looks for.
Needs attention The check found a gap. This produces a finding.
Not verified The check could not establish the result. It never means pass. It earns no credit and counts against coverage.
Not applicable The check does not apply to this repository.
Not assessed by this version This version cannot assess the check. It is excluded from the score.
Optional follow-up The check needs evidence from outside the source. It does not lower the score.