Case study
The PM Intelligence Tool
A multi-agent system Product ran every week, and the evidence rules that turned confident output into output worth acting on.
- Used weekly by Product for PRDs and business reviews
- PMs report ~60% less customer-research time
- First PRD draft: hours → minutes
Nobody asked for this
This tool wasn't on a roadmap. From my seat on the data side, I watched customer intelligence scatter across Gong calls, HubSpot records, Slack threads, Notion docs, and Intercom tickets, with no owner and no way to see it whole. Before writing a PRD, a PM reconstructed reality by hand, one search box at a time.
I pitched the build myself. Seeing the gap from the data side and turning it into a product is the founding-hire version of customer discovery.
The real cost
The hours were the visible cost. The invisible one was worse: PRDs were only as good as whichever fragments a PM happened to find. Decisions inherited the blind spots of a manual search across four tools.
The design decision that mattered
Automate end to end, so the agent writes the PRD
PassedThe tempting demo, and the wrong product. A PRD is a decision document; automating the decision removes the accountability that makes PMs trust it, and tools that replace judgment get quietly resisted.
Agents draft, PMs decide
ChosenAgents own the grunt work of signal detection across the four sources, analysis, and a first PRD draft. The PM owns the judgment about what matters and what ships. The boundary is the product.
Why the boundary is the product
The hardest call wasn't technical. It was deciding what not to automate. Tools that eliminate grunt work get adopted; tools that replace judgment get resisted. Drawing the line at “draft, don't decide” is why it became a weekly habit instead of stalling at the demo.
What I built
Three agent roles over a shared skill library, reading Gong, HubSpot, Slack, Notion, and Intercom. An analyst turns raw signal into patterns worth a PM's attention. A scriber drafts the documents, PRDs and weekly business reviews and deep dives. A strategist handles prioritization, discovery, and pain points. Output lands in a file format on SharePoint, so Product, Sales, CX, GTM, and leadership all reach the same thing. The PM enters at the judgment step with the reconstruction already done.
What it got wrong, and the rules that fixed it
The first version was confident and partly wrong, in two distinct ways. The strategist flagged an account as at risk. My per-cycle validation against the source records caught it before Sales moved, which is the only reason CX didn't run save outreach on a perfectly healthy customer. Root cause: one signal, nothing corroborating it, and that was enough to trigger a flag.
The second failure was quieter and more annoying. The analyst kept reporting pain points that were real conversations but wrong conclusions, customers describing something the team was already shipping that month. Genuine input, useless output, and a lot of noise to wade through.
The fix wasn't prompt tuning, it was evidence rules. Every insight now cites its sources. Single-source findings are suppressed or surfaced as low confidence, and a risk flag needs a second source. Direct human conversation on Slack weighs heaviest, Gong calls support it, HubSpot and Intercom are context. A signal older than the recency window doesn't count as current risk. And before anything surfaces as a gap, the agent checks Linear and the release docs in Notion, because something already in this month's deliverables is not a gap.
The concrete payoff: the agents surfaced a deduplication problem in one of our ledgers, where the same provider charge appeared twice because provider names weren't normalised across bills and notes. Multiple customers had mentioned it. It went from background noise to a prioritised fix.
Results
Adoption is observed; the time savings are self-reported by the PMs. I'd rather label the number honestly than dress it up as telemetry. Quality was checked the same way, by validating agent output against the source records every cycle, which is how both false-positive classes surfaced.