The synthesizer refuses to average away the argument
The final agent reads all 14 risk blocks, consolidates roughly 71 raw entries into a 21-row register, preserves five disagreements between agents verbatim, and writes a one-page verdict on my idea.
The verdict
00-summary.md · synthesizer"Instrument first (Phase 0 / Experiment 1), then run the retention A/B — do NOT fund the full build until both clear. The problem is plausible and the PRD is an unusually honest de-risk-first plan, but the two findings that gate the investment are both unresolved: the high-ARPU cohort the premise rests on has never been measured, and the retention lift the whole bet depends on is hypothesized, not tested."
The engine looked at my own feature idea and said "don't build this yet, and here is exactly what would change that." A nav-drawer tweak got a one-week instrumentation task as its next step, not a quarter of roadmap. Proportionate output is the credibility test.
Where the agents disagreed
verbatim, not averagedThis is the section that separates the engine from "I asked a chatbot for a PRD." Three of the five recorded disagreements:
Build it, or don't fund it at all?
Fund ONLY the cheap validation, not the full build, on current evidence: a power-user minority, no moat, and success metrics that are all UX proxies. "None of them is a dollar."
Endorse de-risk-first build sequencing: gate the build behind instrumentation and an experiment, then fund it if the cohort is real.
Synthesis take: both are right about different questions. The sequencing de-risks the UX, not the business. Add a churn/renewal metric before the build gate, and don't let a UX experiment auto-authorize a 7 to 12 eng-week beta.
The design intent vs. what actually got built
The accessibility intent is sound, and the spec's contrast math independently re-verifies as accurate.
Implements almost none of it: no keyboard route, no screen-reader route, and focus drops on every edit action.
Synthesis take: treat the non-drag path as a release-gated workstream. The prototype is the proof that a naive build drops it.
Free-drag reorder, or pin-to-top?
Recommends free-drag: a full edit mode where all five destinations reorder.
Recommend pin-to-top: the only shape any comparator has actually validated, cheaper to build, simpler accessibility, same data model.
Synthesis take: pin-to-top for the beta. It upgrades to free-drag later with no data migration.
The other two, covered in the full assessment: whether "extend Starred" is code reuse or a metaphor (it's a metaphor, and pricing it as reuse would blow the estimate), and who should hold the veto over hiding monetized surfaces (not the feature's own team).
The risk register
7 high · 11 medium · 3 lowThe seven high-severity rows, uncleaned, because the unflattering entries are the proof this is a process and not a highlight reel:
| # | Risk | Severity | Flagged by |
|---|---|---|---|
| 1 | Business bet hypothesized, not validated. All three success metrics are UX proxies; the retention A/B that would settle it isn't designed yet. | High / High | business-viability |
| 2 | Accessible reorder path unbuilt. Zero keyboard, screen-reader or switch route; focus dropped on every edit action. WCAG 2.2 and ADA exposure. | High / High | 4 agents |
| 3 | No support tooling. Nothing lets support see or reset a user's saved nav state. "My menu changed" becomes an unfixable ticket. | High / High | 3 agents |
| 4 | Hide-policy is one decision from a dark pattern. If "hide" is metric-gated to protect upsell exposure, the control conceals in name only. | High / Med | 5 agents |
| 5 | Telemetry governance. The instrumentation is a bigger risk surface than the feature: no retention window, no consent confirmation for EU/UK. | High / Med | 3 agents |
| 6 | Client-side-only invariant. Without server-side write validation, a bad client can durably strand every synced device from chat. | High / Med | privacy-security |
| 7 | Opportunity cost. The same segment has documented Dispatch reliability bugs: a better-evidenced retention lever no one compared against. | High / Med | 2 agents |
The register also states openly where it demoted two high findings after the scope cut addressed them, with the original ratings printed so nothing is hidden. All 21 rows are in the full assessment.
The deliverable
Read what a stakeholder would actually receive
The full assessment package: the one-page verdict, the complete 21-row register, the PRD, design spec, tech spec, and the embedded prototype. Every artifact ends with open questions for the team that would own it, because these are inputs to collaboration, not replacements for it.