Case study
Marquee: The 58% Experiment
The north-star metric moved 58% after a two-week UX fix with zero model changes. A case study in knowing when AI isn't the answer.
- +58% north-star metric in the weeks after launch
- 2× downstream engagement
- Cycle time roughly halved
- The metric became the company's north star
Context
Supio shipped product weekly, and engagement metrics were everywhere — but the metric tying product output to customer value wasn't tracked at all. I run one loop in situations like this: find what's missing, instrument it, diagnose the friction, ship the smallest experiment that could move it.
What the instrumentation showed
Once the metric existed, the funnel was blunt: users weren't finding new features. Launches lived levels deep inside existing flows — a feature shipped on Monday was invisible by Friday. User research confirmed what the funnels implied: customers wanted capabilities they didn't know existed.
The bottleneck wasn't the product's intelligence. It was discoverability.
The fork in the road
Add AI — a recommendation / surfacing model
PassedAt an AI company, the reflex answer. But the data made the debate unnecessary: users weren't rejecting features, they weren't finding them. A model would optimize which hidden thing users continue not to see.
Give launches a permanent home
ChosenA beta-releases tab, pinned bottom-left: a dedicated, always-visible place showing each week's new features. Two weeks of engineering. No model changes.
Why no-AI won
The principle is boring and I stand by it: when UX is the bottleneck, don't add a model. The instrumentation showed the friction so clearly that nobody had to argue — which is itself the lesson. Good instrumentation doesn't just inform decisions; it dissolves debates before they start.
Product judgment at an AI company includes knowing when AI isn't the answer. This was the cheapest credible experiment that could move the metric, so it went first.
What shipped
I scoped the spec — the placement, what qualifies as a launch, and the weekly cadence the tab commits the company to — and an engineering team shipped it in two weeks. Every new feature now has a guaranteed moment of visibility instead of hoping users stumble into it.
Results
Measured before/after rather than as a controlled experiment — so I claim the association, not clean causality. The durability of the lift and the doubling of downstream engagement are what convinced the company to adopt the metric as its north star.