What we check
The assessment groups its checks into ten categories. Each category carries a share of the overall score.
Categories
Section titled “Categories”| Category | Weight | Checks defined | Scored today |
|---|---|---|---|
| Security and privacy | 20% | 7 | 7 |
| Feature completeness | 12% | 3 | 2 |
| System design | 12% | 8 | 1 |
| Data integrity and lifecycle | 10% | 4 | 0 |
| Reliability | 10% | 4 | 3 |
| Operability, including observability | 10% | 10 | 6 |
| Agent readiness | 8% | 12 | 11 |
| Scalability and performance | 8% | 4 | 1 |
| Maintainability and configuration | 5% | 4 | 3 |
| User experience and interfaces | 5% | 3 | 1 |
| Total | 100% | 59 | 35 |
Why “defined” and “scored” differ
Section titled “Why “defined” and “scored” differ”The assessment defines 59 checks. Today, 35 of them are scored from source.
- 22 checks are not assessed by this version. They appear in your result as Not assessed by this version and are excluded from the score. They are never shown as passes.
- 2 checks need evidence from outside the source, for example from a running system. They appear as Optional follow-up and do not lower the score.
More checks move from defined to scored as the assessment grows.
What the categories cover
Section titled “What the categories cover”- Security and privacy. Identity, authorization, input handling, credentials, and sensitive data.
- Feature completeness. Stubs, fake success, and missing failure states behind features that look finished.
- System design. The boundaries between your code and the services it depends on.
- Data integrity and lifecycle. No checks in this category are scored yet.
- Reliability. Error handling, retries, and what happens when a dependency is down.
- Operability, including observability. Health checks, deployment and rollback, configuration, and whether failures leave a useful signal.
- Agent readiness. Entry points, agent instructions, setup, and the build, lint, type, and test commands an agent uses to check itself.
- Scalability and performance. Unbounded inputs and work that grows without limit.
- Maintainability and configuration. Supported dependencies, consistent configuration, and tests that cover what is claimed.
- User experience and interfaces. Input and accessibility basics on the surfaces people use.
Check outcomes
Section titled “Check outcomes”Every check ends in one of six outcomes.
| Outcome | Meaning |
|---|---|
| Pass | The check found what it looks for. |
| Needs attention | The check found a gap. This produces a finding. |
| Not verified | The check could not establish the result. It never means pass. It earns no credit and counts against coverage. |
| Not applicable | The check does not apply to this repository. |
| Not assessed by this version | This version cannot assess the check. It is excluded from the score. |
| Optional follow-up | The check needs evidence from outside the source. It does not lower the score. |