Original Evidence · Reproducible Benchmark

ProjectOps360 Evidence & Benchmarks

Product evidence should show what the system actually does — and what it does not prove. This benchmark runs controlled Friction Radar v1 evidence through the product read-model contracts and publishes the result, fixture description and limitations.

Measured result

5/5 controlled evidence checks passed

Version 2026-08-23. Controlled deterministic code-level benchmark of Friction Radar v1 evidence handling. This is not a customer-outcome, recall, precision, or latency benchmark.

2/2

Repeated runs identical

Same fixture, same complete read model.

1/1

Foreign-project signals excluded

Critical evidence from another project did not leak into the scoped result.

1/1

Expected explicit clusters

Only signals sharing the explicit entity were correlated.

0

Unvalidated global scores published

Global score and severity remain null in Friction Radar v1.

0/8

Category aggregate scores published

No category aggregation is presented before validation.

8

Defined friction categories

Process, resource, dependency, schedule, cost, risk, decision, and quality.

Control-by-control evidence

Every published result has an explicit expectation

Deterministic read model

PASS
Measured
2/2 repeated runs identical
Expected
2/2 repeated runs identical
Evidence rule
The same controlled signal fixture is evaluated twice and the complete read model is compared byte-for-byte after JSON serialization.

Cross-project isolation

PASS
Measured
1/1 foreign-project signal excluded
Expected
1/1 foreign-project signal excluded
Evidence rule
The fixture contains one critical signal from another project. The read model must retain only signals scoped to the requested organization and project.

Explicit-entity cluster specificity

PASS
Measured
1/1 expected cross-category cluster
Expected
1/1 cluster from the shared entity only
Evidence rule
Dependency and decision signals share one task and should correlate. An unrelated schedule signal must remain outside that cluster.

No unvalidated aggregate score

PASS
Measured
0 global and 0/8 category aggregate scores published
Expected
0 global and 0/8 category aggregate scores published
Evidence rule
Friction Radar v1 intentionally keeps global severity/score and category aggregation null until an aggregation policy is validated.

Honest empty state

PASS
Measured
0 signals; score=null; severity=null; trend=unknown
Expected
0 signals; score=null; severity=null; trend=unknown
Evidence rule
With no scoped evidence, the model must not manufacture a score, severity, trend, or cluster.

Reproducibility

The benchmark can be inspected from fixture to result

The controlled fixture contains 3 in-project signals and 1 foreign-project signal. Two signals share an explicit entity, one signal is unrelated, and the read model is executed 2 times for determinism checking.

Evidence chain

Controlled events

Friction signals with evidence refs

Scoped Friction Radar read model

Expected-control assertions

Public JSON + stated limitations

What this benchmark does not claim

Controlled evidence is not a substitute for field validation

The benchmark uses a controlled synthetic fixture so expected behavior is known in advance.
It does not measure real-world precision, recall, false-positive rate, causal accuracy, model quality, or customer outcomes.
It does not compare ProjectOps360 against third-party products.
Passing these controls demonstrates deterministic evidence handling and governance behavior for the tested Friction Radar v1 contracts only.

Evidence standard

ProjectOps360 separates evidence from interpretation

Product evidence is strongest when a reader can distinguish what was observed, what was inferred, what is only predicted, and what remains unknown. The benchmark follows the same discipline used by Project Friction Intelligence.

Observed

Direct output or event-backed fact produced by the controlled fixture.

Inferred

Interpretation supported by evidence, but not presented as automatic proof of cause.

Predicted

Forward-looking exposure that must stay labeled as prediction rather than fact.

Unknown

Missing evidence stays unknown instead of being silently converted into a score.

FAQ

Evidence benchmark questions

What does this ProjectOps360 benchmark prove?

It demonstrates the tested Friction Radar v1 contracts for deterministic read-model generation, project isolation, explicit-entity correlation, restraint around unvalidated aggregate scores, and honest empty-state behavior.

Does this benchmark prove real-world AI accuracy?

No. The benchmark is intentionally controlled and synthetic. It does not measure real-world precision, recall, false-positive rate, causal accuracy, model quality, or customer outcomes.

Can the benchmark be independently inspected?

The measured result and its stated methodology are published as machine-readable JSON. The implementation repository is private, so this page does not claim that the source code or regression test is publicly accessible.

Why does ProjectOps360 publish null aggregate scores in this benchmark?

Friction Radar v1 deliberately avoids publishing a global friction score, global severity, or category aggregate score until an aggregation policy has been validated. Missing validation is represented honestly rather than replaced with a fabricated number.

ProjectOps360

From evidence to execution intelligence.

Process Mining reconstructs what happened. Friction Radar surfaces evidence-backed friction. Living Graph shows what the problem can affect next.