← All projects

Case study

Marquee: The 58% Experiment

Week-one feature activation went from 30.0% to 47.4% after a navigation change with zero model changes, proven with a firm-randomised A/B. A case study in knowing when AI isn't the answer.

Role
PM for instrumentation, user research, and experiment design
Timeline
Four-week test across 300 customer firms
Team
Engineering team built to my scoped spec
Outcomes

Context

Supio shipped product weekly, and engagement metrics were everywhere, but nothing measured whether customers ever found what we shipped. I run one loop in situations like this. Find what's missing, instrument it, diagnose the friction, ship the smallest experiment that could move it. So I instrumented three things in Fullstory and PostHog: discovery, meaning the feature was encountered; activation, meaning first real use; and engagement, meaning it stuck.

What the instrumentation showed

Once the metric existed, the funnel was blunt. Activation of new features was low and slow. User research found the mechanism, and it was almost mundane: the release popover only fired on login, and our users kept sessions open for days, so most of them never saw it. The announcement email got lost. The features weren't the problem, the announcement channel was.

We tried telling users where the popovers appeared. It didn't stick, because they still had to log in to see them. The bottleneck wasn't the product's intelligence. It was discoverability.

The fork in the road

Add AI, a recommendation or surfacing model

Passed

At an AI company, the reflex answer. But the data made the debate unnecessary: users weren't rejecting features, they weren't finding them. A model would optimize which hidden thing users continue not to see.

Give launches a permanent home

Chosen

A persistent release tab in the left nav listing the last month of features: a dedicated, always-visible place users could find on their own schedule. No model changes.

Why no-AI won

The principle is boring and I stand by it. When UX is the bottleneck, don't add a model. The instrumentation showed the friction so clearly that nobody had to argue, which is itself the lesson. Good instrumentation doesn't just inform decisions; it dissolves debates before they start.

Product judgment at an AI company includes knowing when AI isn't the answer. This was the cheapest credible experiment that could move the metric, so it went first.

What shipped

I scoped the spec, which set the placement, what qualifies as a launch, and the weekly cadence the tab commits the company to. Then I designed the test to settle it rather than argue it: a 50/50 A/B randomised at the firm level, not the user level, because colleagues tip each other off and intra-firm contamination would have washed the result out. 300 firms and roughly 5,000 users, stratified into four size buckets so the arms were balanced on the thing that most drives baseline adoption and so the lift could be read per segment. Four weeks, long enough to cover complete enterprise login cycles and give late-entering cohorts a full exposure period. Primary metric was the share of users activating a new feature within their first week of exposure, first real use rather than discovery, because seeing the tab would have counted itself. Core engagement and support volume rode along as guardrails.

Results

47.4%
week-one activation, up from 30.0%
2×
repeat use of newly released features
~½
time from release to first use

Significance was computed cluster-adjusted, since the firm is the unit of randomisation and a user-level test would have overstated power. Every stratum lifted independently, so no single large account carried the result, and the guardrails held: no rise in support volume, no drop in core engagement. The activation lift is causal. The link onward to renewal and expansion is a correlation from our trial and renewal analytics, which is why leadership adopted feature adoption as the north star, and it is worth saying plainly that the second claim is weaker than the first.

What I'd do differently

Follow each release cohort for a quarter after the test. Four weeks proves the mechanism works; it doesn't separate a durable discovery habit from a novelty bump. I'd also pre-register the downstream read-through, repeat use and then renewal, so the experiment gets judged on revenue-linked behaviour instead of activation alone.

And the tab created a contract I hadn't fully priced in. A dedicated home for weekly launches only works if you launch weekly. The UX fix quietly became an operating commitment for the whole product org, worth making, but worth making explicitly.