DIMENSIONS
Score the capabilities that matter to the decision.
Dimensions may include core capability, reliability, deployment flexibility, security, integration depth, operability, developer experience, transparency, documentation, support, value, and ecosystem maturity. Every dimension is not equally relevant to every category.
WEIGHTING
Weights follow the category—not a universal template.
We publish or explain material weighting choices. A reliability tool may weight operational depth more heavily than interface polish; a memory layer may weight retrieval quality, deployment, and integration differently.
CONFIDENCE
Evidence strength changes how firmly a score should be read.
Confidence considers testing access, documentation quality, version certainty, reproducibility, sample size, source quality, and unresolved questions. Limited evidence is disclosed instead of converted into false certainty.
INTERPRETATION
A high score can still be wrong for a specific team.
Readers should use the narrative verdict, fit segments, limitations, and deployment assumptions alongside the score. We avoid declaring a universal winner where the evidence supports conditional choices.
NO FALSE PRECISION
Decimal places cannot manufacture certainty.
Scores are rounded consistently, material methodology changes are documented, and historical comparisons account for changed versions or criteria.