Frederick Sona
HomeCase Studies › Conversion rate optimization
Methodology Playbook · CRO Playbook

Conversion rate optimization

How I run CRO as a compounding program: hypothesis pipeline, test structure that respects statistics, sequencing that avoids waste, and the discipline that keeps wins in production instead of quietly disappearing.

Type: Methodology playbook Discipline: Optimization Updated: 2026-07-23
Playbook, not a single engagement. This is how I run conversion rate optimization across ecommerce, SaaS, and lead generation programs. The frame, the tools, the traps.

TL;DR

CRO is a compounding program of prioritized hypotheses tested with statistical rigor against a shipped baseline. A working program runs two to four tests per month, keeps a 60 to 75 percent test-to-launch ratio for winners, and lifts primary conversion by 15 to 40 percent in the first year. The failure mode is one-off tests that never accumulate into a system, run without sample size math, and stop the moment they look like a win.

The playbook, in one paragraph

Build a research foundation (analytics, session recordings, heatmaps, user interviews) before writing a hypothesis, prioritize the hypothesis backlog with a real framework (I use PIE or ICE with adjustments), design tests with pre-calculated sample size and stop rules, ship each test in a documented pipeline (hypothesis, treatment, primary metric, guardrails, sample size, duration, result), analyze results only after the pre-committed duration, and archive every test outcome so the program compounds insight rather than repeating past mistakes. The wins come from the discipline, not from any single test.

Where this fits in the modern discovery layer

CRO is the leverage on every acquisition channel. A doubled conversion rate cuts the effective cost of every visitor from every source in half. In the Search Everywhere Optimization frame, CRO is the surface-agnostic multiplier: the same test that lifts organic search conversion lifts paid search conversion, email conversion, and referral conversion in the same session.

Of the 19 surfaces in the Playbook, CRO touches all of them indirectly and three directly. Core Web Vitals (CWV) has direct conversion effects because performance is a conversion input. E-E-A-T interacts with CRO because trust signals (named authors, real reviews, credentials) tested well lift conversion. And reputation platforms interact because on-page review displays are testable elements with meaningful lift potential.

The strategic point: CRO does not replace acquisition. It multiplies it. A brand spending $50K a month on paid acquisition and running no CRO program is leaving 20 to 60 percent of that spend unclaimed. The CRO program pays for itself, then continues paying, then compounds.

The five levers

1. Research before hypothesis

Every hypothesis worth testing is grounded in evidence: an analytics anomaly (drop off at step three of checkout), a session recording pattern (users pinching to zoom on mobile pricing), a support ticket theme (repeated questions about shipping), a user interview quote (the customer said "I could not find the return policy"). Hypotheses without evidence are guesses, and guesses have a 10 percent hit rate. Evidence-based hypotheses hit at 30 to 50 percent.

2. Prioritization with a real framework

PIE (Potential, Importance, Ease) or ICE (Impact, Confidence, Ease). Each hypothesis scored on the same three axes. Backlog sorted by total score. The top three tests each month come from the top of the sorted list, not the founder's morning meeting. Frameworks are boring, and they are how programs stay honest.

3. Statistical rigor as a non-negotiable

Every test has a pre-calculated sample size based on baseline conversion, minimum detectable effect, and desired power (typically 80 percent). Every test has a pre-committed duration (two business cycles minimum). Every test has stop rules that are not "it is winning." Peeking at a test in progress and stopping when it looks like a winner is how programs manufacture false positives. Discipline here is what separates real lift from noise the team believed.

4. Sequencing that avoids interference

Two tests running on the same page interfere. Two tests running on adjacent pages in a funnel can interfere. The test calendar sequences tests so overlapping traffic does not confound results. On a low traffic site, this means fewer tests running at once. On a high traffic site, tests can parallelize with holdout groups.

5. Post-test discipline

Winning tests get shipped to production with the same code path the winning treatment used. Losing tests get archived with the reasoning documented (this is where the program compounds). Neutral tests get analyzed for segment level effects before being called neutral. The archive of past tests prevents the team from retesting variations that already lost, which is how CRO programs quietly waste half their bandwidth.

First 30 / 60 / 90 days

Days 1 to 30: research and instrumentation

Analytics audit. Is GA4 properly configured. Are events fired for every meaningful action. Are conversion goals defined and measurable. Are segments defined for new versus returning, mobile versus desktop, paid versus organic. Broken instrumentation makes every downstream decision unreliable.

Behavior research. Hotjar or FullStory session recordings on the top five converting pages. Heatmaps on the priority landing pages. Scroll depth analysis. Rage clicks, dead clicks, exit points.

Quantitative funnel analysis. Full conversion funnel mapped with drop-off percentages at each step. The step with the largest absolute drop-off becomes a testing candidate.

Qualitative research. Five to eight user interviews or a survey to 300 recent buyers and 300 non-buyers. What almost stopped them from converting. What almost stopped them from returning.

Hypothesis backlog seeded with the first 20 to 40 test ideas, each grounded in a specific piece of evidence. Prioritized with PIE.

Metric moving in month one: a documented baseline conversion rate for the primary funnel, a hypothesis backlog with priority scores, and the CRO tool configured. No tests running yet on purpose. Rushed testing on broken instrumentation produces false wins.

Days 31 to 60: first tests and process

Two to three tests running by end of month two. Each test documented with hypothesis, treatment, primary metric, guardrails, calculated sample size, and duration. Tests run to completion. No peeking, no early stops.

Process shipped. Weekly CRO review meeting with product, engineering, and marketing. Test ideas reviewed for evidence and priority. Running tests reviewed for guardrail health, not for winning-ness. Completed tests reviewed for archive-ready documentation.

QA process shipped. Every test QA'd across browsers, devices, and account states before launch. A test with a rendering bug in the treatment produces unusable results and burns the sample.

Metric moving in month two: first tests reaching statistical significance. Even a losing first test that is properly run is a program win because the process is now working.

Days 61 to 90: cadence and compounding

Testing cadence at three to five tests per month. First winners shipped to production with implementation validated against the test treatment.

Segment analysis on neutral tests. A neutral test at the overall level often shows meaningful segment effects (mobile lift and desktop drop, for example). Segment insights become hypotheses for subsequent tests.

Program review with stakeholders. What was tested, what won, what lost, cumulative lift on the primary metric, projected annualized revenue impact. This is where the CRO program earns its next quarter of investment or loses it.

Deliverable at day 90: an active pipeline of three to five tests per month, first shipped winners with measured lift, an archive of tested hypotheses, a process the team can maintain without me, and a clear roadmap for months four through twelve (deeper segmentation, personalization tests, checkout redesign, category page redesign, mobile-specific optimization).

Tools I use

VWO for ecommerce clients running 4 or more tests per month. Solid editor, heatmaps and session recordings included, statistical engine is defensible.

Optimizely for enterprise clients with engineering teams that want to run tests server-side or on custom stacks. More expensive than VWO. Worth it above a certain scale.

GA4 experiments for lightweight testing on smaller sites where the traffic and test volume do not justify a dedicated CRO tool. The interface is thin but the statistical foundation is sound.

Hotjar for heatmaps, session recordings, and on-page surveys. This is where hypotheses come from. Session recording review is the single most productive hour of my week on any active CRO account.

FullStory for enterprise clients that need session recording at scale with search across sessions. More expensive than Hotjar. The searchability pays back at enterprise scale.

GA4 for the conversion baseline and post-test validation. Every test result gets sanity checked against GA4 conversion data before the win gets called.

Google Sheets or Notion for the hypothesis backlog and test archive. The tool matters less than the discipline of using it every week.

Sample size calculators from Evan Miller or VWO's built-in calculator. Every test gets its sample size pre-calculated. This is the two minute step that separates real CRO from theater.

What kills the program

1. Peeking and stopping early

The test looks like it is winning at day four. The team ships the win. Two weeks later production conversion is unchanged. The early stop turned normal variance into a false positive. Every mature CRO program has learned this lesson the expensive way. Pre-committed duration is not a suggestion.

2. No sample size calculation

The team runs a test on a page with 400 weekly visitors and calls a 12 percent lift after two weeks. The confidence interval on that estimate spans zero. There was no signal, only noise the team believed. Sample size math costs two minutes and saves months of misplaced strategy.

3. HiPPO tests

The highest-paid person in the room wants to test their idea. The idea is ungrounded in evidence and low priority under any real framework. The team runs the test to keep the peace. The test loses. Two weeks of the calendar are gone. HiPPO tests are a program tax and worth resisting politely.

4. Interference

Two tests running on the same page. A promo email launches during the test window and changes the traffic mix. A homepage redesign ships in the middle of a checkout test. All three contaminate results. The test calendar prevents the first, communication prevents the second, and change freezes during priority tests prevent the third.

5. No archive

The team runs 40 tests over a year. Twenty won. Twelve lost. Eight were neutral. Nobody documented anything. Six months later a new team member proposes a test the previous team already ran and lost. The archive is where the program compounds. Skipping documentation is skipping the compounding.

6. Testing tactics on strategy problems

The site converts poorly because the positioning is unclear and the pricing is not competitive. The team runs button color tests. Button color tests will never fix a strategy problem. The right work is repositioning or repricing, not a fifty-first CRO test.

KPIs that matter

Cumulative conversion rate lift. The primary program metric. Percentage lift on the primary conversion metric attributable to shipped winners over the trailing 12 months. Target 15 to 40 percent in year one, 8 to 20 percent in year two.

Test velocity. Tests launched per month. Target two to five for most programs. Below one, the program is not really running. Above eight, the program is likely testing without discipline.

Win rate. Percentage of tests that reach statistical significance in the positive direction. Target 20 to 35 percent. Higher win rates usually mean tests are underpowered or optimistically called.

Ship rate. Percentage of winners that actually shipped to production. Target 90 percent or higher. A shipped rate under 70 percent means winners are being lost between test and production.

Revenue per visitor lift. Ecommerce specific. Winners often lift conversion but hurt AOV or reduce revenue per visitor. Revenue per visitor is the honest metric.

Guardrail health. Every test has secondary metrics that should not move (bounce rate, cart abandonment, return rate). Guardrails preventing tests from winning on conversion while damaging downstream metrics.

Program payback. Annualized revenue impact of shipped winners against the program cost. Target 5x or better in year one, 10x or better at maturity.

FAQ

How much traffic do we need to run A/B tests?

For a page converting at 2 percent and a 20 percent minimum detectable effect, you need about 17,000 visitors per variant to reach 95 percent confidence. Lower baselines and smaller effects require more traffic. Below 10,000 monthly visitors on the target page, qualitative testing is usually a better investment than A/B tests.

How long should a test run?

Minimum two full business cycles, which is usually two weeks. This covers weekday versus weekend behavior. Never stop a test early because it looks like it is winning. Early stops are the number one source of false positives in CRO.

What is the difference between A/B testing and multivariate testing?

A/B testing compares two versions of one element. Multivariate testing compares combinations of multiple element changes simultaneously. A/B needs a fraction of the traffic and gives clean answers. Multivariate needs more traffic and rarely gets used well outside enterprise programs.

Do we need a CRO tool or is our developer enough?

A CRO tool (VWO, Optimizely, or GA4 with the built-in experiment feature) speeds test setup, handles bucketing, and provides statistical analysis. A developer can do the same thing. The tool is worth it once you are running two or more tests per month. Below that, dev implementation is fine.

What is a good conversion rate for our site?

There is no useful benchmark. Ecommerce ranges from 1 to 8 percent by category and traffic quality. B2B ranges from 2 to 15 percent for content, 0.5 to 4 percent for demos. Compare against your own historical baseline, not against industry averages that mix incomparable businesses.

Why did our winning test not hold up in production?

Three usual reasons. The test stopped early before reaching significance. The sample was contaminated (bot traffic, internal traffic, promotional traffic during test window). Or the winning treatment interacted with a change that shipped elsewhere on the site during the test. All three are preventable with process discipline.

If your funnel is leaking, tell me which step and I will tell you where I would start.

Start a conversation
← Back to case studies