← All projects

Case study

Marquee: The 58% Experiment

The north-star metric moved 58% after a two-week UX fix with zero model changes. A case study in knowing when AI isn't the answer.

Role
PM — instrumentation, user research, experiment design
Timeline
2 weeks from scoped spec to ship
Team
Engineering team built to my scoped spec
Outcomes

Context

Supio shipped product weekly, and engagement metrics were everywhere — but the metric tying product output to customer value wasn't tracked at all. I run one loop in situations like this: find what's missing, instrument it, diagnose the friction, ship the smallest experiment that could move it.

What the instrumentation showed

Once the metric existed, the funnel was blunt: users weren't finding new features. Launches lived levels deep inside existing flows — a feature shipped on Monday was invisible by Friday. User research confirmed what the funnels implied: customers wanted capabilities they didn't know existed.

The bottleneck wasn't the product's intelligence. It was discoverability.

The fork in the road

Add AI — a recommendation / surfacing model

Passed

At an AI company, the reflex answer. But the data made the debate unnecessary: users weren't rejecting features, they weren't finding them. A model would optimize which hidden thing users continue not to see.

Give launches a permanent home

Chosen

A beta-releases tab, pinned bottom-left: a dedicated, always-visible place showing each week's new features. Two weeks of engineering. No model changes.

Why no-AI won

The principle is boring and I stand by it: when UX is the bottleneck, don't add a model. The instrumentation showed the friction so clearly that nobody had to argue — which is itself the lesson. Good instrumentation doesn't just inform decisions; it dissolves debates before they start.

Product judgment at an AI company includes knowing when AI isn't the answer. This was the cheapest credible experiment that could move the metric, so it went first.

What shipped

I scoped the spec — the placement, what qualifies as a launch, and the weekly cadence the tab commits the company to — and an engineering team shipped it in two weeks. Every new feature now has a guaranteed moment of visibility instead of hoping users stumble into it.

Results

+58%
north-star metric in the weeks following launch
downstream engagement
cycle time, roughly halved

Measured before/after rather than as a controlled experiment — so I claim the association, not clean causality. The durability of the lift and the doubling of downstream engagement are what convinced the company to adopt the metric as its north star.

What I'd do differently

Design the measurement as a staged rollout so the causal claim is airtight — the lift was real, but a before/after leaves room for argument that a control group would have closed.

And the tab created a contract I hadn't fully priced in: a dedicated home for weekly launches only works if you launch weekly. The UX fix quietly became an operating commitment for the whole product org — worth making, but worth making explicitly.