The same product graded against itself — one implementation's UI measured against its own engine.
Of everything the engine can do, how much can a user actually reach?
12 of 12 capabilities fully present · baseline Engine capabilities (Dpp.Nano) → target UI + MCP surface (WinApp) · 12 compared
assessment v20 · 2026-08-02 · authored match
Comparison revision after 86bb1b0: UI49 closes E12 with resumable official speedrun stages, configurable scaling-laws and miniseries matrices, and atomic post-train BPB/CORE sweep evaluation; portable CPU training remains covered by UI48. Validation was a zero-warning build, 12 focused UI/parser/speedrun-resume tests, and 3 bounded executable scaling/miniseries/sweep tests. No full-scale production run is claimed. All 12 high-level NanoChat engine capabilities now have an aligned WPF UI coverage path; 8 intentional UI-only additions remain target-only.
The work queue: capabilities the baseline has and the target does not.
None.
Started, not finished — present but not equivalent.
None.
Not defects. Parity is directional, so what the target adds beyond the baseline shows up here rather than counting against anyone.
Comparison revision after 86bb1b0: UI49 closes E12 with resumable official speedrun stages, configurable scaling-laws and miniseries matrices, and atomic post-train BPB/CORE sweep evaluation; portable CPU training remains covered by UI48. Validation was a zero-warning build, 12 focused UI/parser/speedrun-resume tests, and 3 bounded executable scaling/miniseries/sweep tests. No full-scale production run is claimed. All 12 high-level NanoChat engine capabilities now have an aligned WPF UI coverage path; 8 intentional UI-only additions remain target-only.