Multivariate Testing Needs Enough Traffic

Imagine a landing page whose headline feels vague, whose hero image shows the wrong use case, and whose button says only “Learn more.” The team could replace all three and compare the new page with the old one. If signups rise, that comparison helps answer whether to ship the new page. It cannot tell the team whether the headline, image, button, or a particular pairing did the work. That is the question multivariate testing is built to explore: it compares combinations of selected elements within one experience, with visitors assigned to those combinations during the experiment. Adobe’s description of multivariate tests makes the combination, rather than an isolated edit, the unit a visitor sees.

multivariate testing: two equal unmarked page screens with different abstract layouts side by side, traffic tokens, variation swatches, calculator, notebook, desk lamp, coffee cup

The method is attractive because it promises a more useful answer than “the new page won.” Its cost is that every added option creates more combinations that need visitors. A sensible test therefore starts with a decision: which page elements might change, what outcome matters, and whether enough traffic will reach each version to make the result useful. More options do not automatically mean more learning.

Each extra option creates another page to compare

Suppose the landing page has two possible headlines and two possible hero images. Call the headlines A and B, and the images 1 and 2. A full factorial test includes all four pairings:

CombinationHeadlineHero image
A1A1
A2A2
B1B1
B2B2

Each visitor assigned to this illustrative test sees one pairing. The four cells let the team compare complete combinations and also ask how the headline performs across both images. In a full factorial design, every selected setting of one factor appears with every selected setting of the others. “Factor” is the element being changed, such as the headline; “level” is one option for that element, such as headline A. The four-cell design tests only those two headlines and two images. It says nothing directly about other headlines, other images, or a different page.

The arithmetic is simple but consequential. Multiply the number of levels for each factor: two headlines × two images = four combinations. Add two button labels and the same design has eight combinations. Give each of three elements three options and there are 27. Adobe’s testing guide uses that three-by-three-by-three example to show how quickly the count grows. The count describes versions to build and assign; it is not a count of visitors needed to learn from them.

This is why I would resist putting every proposed copy, image, color, and button into one launch. First decide which choices could plausibly change the decision you need to make. If the team would not act differently based on the image result, the image may not belong in this test. Removing a factor reduces the number of cells, but it also gives up the chance to learn whether that factor interacts with the others. The smaller test is a deliberate trade, not a free efficiency gain.

Use A/B testing when the decision is about whole pages

An A/B test can compare two complete experiences. Version B may have a new headline, image, and button at once. If B performs better than A, the team has a comparison between packages, but the result cannot separate the contribution of each edit. Optimizely’s overview of multivariate testing distinguishes that question from examining several components together. The useful distinction is the decision being made, not a rule that an A/B variant may change only one thing.

If the team is choosing between its current landing page and a finished redesign, I would run the simpler page-level comparison. Splitting the redesign into component tests would answer a different question and might force combinations that no designer intended to publish. If the layout is settled and the team wants to choose a headline and hero image for that layout, a multivariate test can be worth the additional cells. In that case, the team needs both the best observed pairing and an understanding of whether a choice works only with its partner.

There is also a useful sequence when the questions arrive at different times. The team can compare layouts as complete pages, then test content within the selected layout. Adobe describes this order in its guidance on A/B and multivariate tests. A layout change can alter where a headline or button appears, so treating layout and copy as unrelated options in one matrix may produce combinations the team would never use. Keeping the experiment close to the actual publishing choice makes its result easier to act on.

Calculate traffic per combination before launch

The number of visitors to the page is not the number available to each combination. For an illustration, suppose 24,000 visitors are eligible during the planned period and are assigned equally across eight combinations. That gives an average of 3,000 visitors per cell. With sixteen combinations, the same 24,000 visitors average 1,500 per cell. If the outcome occurs for about 4% of visitors, those averages correspond to roughly 120 and 60 outcomes per cell respectively. These are planning assumptions, not observed results or a claim that either design has enough data.

The amount needed depends on the outcome and the size of the difference worth detecting. A small change in a rare signup event is harder to distinguish with the same traffic than a larger change in a common action. Adobe notes that more combinations require more traffic or time, and it recommends starting from the page’s usual impressions and conversions. Before commissioning eight or sixteen versions, estimate eligible visitors over the actual run period, the expected outcome rate, and the smallest change that would matter to the business decision. Then estimate the run length or required sample with a method suited to the chosen metric. A fixed rule such as “one thousand visitors per variation” cannot replace those inputs.

The same arithmetic becomes tighter if the team later wants a separate answer for a segment. Eight combinations across all visitors do not each receive one eighth of mobile visitors unless mobile assignment and volume actually support that split. The full factorial count also excludes any further division by audience. If the purchase decision concerns mobile visitors, plan for that audience from the start instead of discovering after launch that its cells are too small.

When the projected run would outlast the decision window, reduce the options. Test two headlines rather than four, or set aside the button label until the headline and image decision is made. For example, two headlines × two images × two buttons creates eight cells; removing the button leaves four and doubles the average visitors available to each remaining cell under equal assignment. NIST’s account of full factorial designs shows the multiplication behind that choice. The cost is clear: the four-cell test cannot tell you whether a button label changes the effect of a headline.

Design combinations people could actually use

Before building variants, write down the complete list of combinations and read each as a page. A headline that promises one use case paired with an image of another can be a poor visitor experience even if each option looks good on its own. Either revise the options so all pairings are publishable, or ask a narrower question. A full factorial design only helps if its cells are meaningful choices.

Choose one primary outcome that matches the decision. If the page exists to generate completed signups, a button click answers only whether more people clicked the button; it does not by itself answer whether more people completed signup. Optimizely lists clicks, conversions, and engagement as possible test metrics. The team can watch supporting actions, but it should state before launch which result decides the page choice. It should also define the eligible audience and the measurement period so the denominator of each reported rate is clear.

Review all the built combinations before assigning traffic. In an eight-cell test, checking only each headline and each image in isolation misses the actual pairings a visitor will see. Verify that the selected content renders together and that the intended action remains usable in every cell. Adobe advises allowing additional quality-check time as the number of experiences grows. It also advises planning the design before launch rather than editing a live test. Changing a level midway would make “headline A” mean different things across the measurement period and muddle the comparison.

Record the options, the assignment plan, and the outcome definition in plain language before looking at results. That discipline is practical: it lets the team distinguish a finding about the planned headline-and-image combinations from a post hoc story about whichever small subgroup happens to look strongest. The aim is to make the eventual choice legible to the people who will publish the page.

Read interactions as well as the best observed cell

Consider an illustrative two-by-two result with exactly 1,000 assigned visitors in each cell. The outcome is a completed signup, counted once per visitor. The numbers below are invented solely to show how the interpretation works:

Headline and imageSignups / visitorsSignup rate
A140 / 1,0004%
A260 / 1,0006%
B170 / 1,0007%
B250 / 1,0005%

Looking at the headline alone, A produces 100 signups among 2,000 visitors, or 5%, while B produces 120 among 2,000, or 6%. That is the headline’s observed overall pattern across these two images. Looking at the images alone, image 1 and image 2 each produce 110 signups among 2,000 visitors, or 5.5%. The image averages appear identical.

The combinations tell a more interesting story. With headline A, switching from image 1 to image 2 changes the observed rate from 4% to 6%. With headline B, that same switch changes it from 7% to 5%. The image’s apparent effect reverses across the two headlines and disappears when averaged over them. A factorial model can represent both individual factor effects and interaction terms. Here, the interaction is the joint pattern that the separate headline and image averages miss.

Headline B plus image 1 has the highest observed rate in this invented table. That does not establish that it will remain best: the table contains only 70 signups in that cell, and sampled rates vary. Nor does the pattern show why the pairing worked. It suggests which combination deserves attention within the tested page, audience, and period. To make a publishing decision, compare the size and uncertainty of the differences with the improvement the team said would matter. A model term or a rank order alone is not a business decision.

There is a related trap in picking the “best” level of each element separately and combining them afterward. In the illustration, headline B looks better overall and neither image looks better overall. Yet B1 is higher than B2 in the observed cells. Main effects summarize across partners; they do not guarantee that independently chosen levels produce the strongest pairing. The interaction question is one reason to use a multivariate design in the first place.

Let the result decide only what the test covered

A clean result can support a narrow, useful choice: among these configured combinations, for these assigned visitors, one pairing performed better on the chosen outcome during the measurement period. It does not establish that the headline is universally better, that the same image will work on another page, or that untested alternatives would lose. The selected factors and levels define the reach of the conclusion. If a team wants to keep exploring, the next test should be chosen from the uncertainty that matters to the next publishing decision, not from an ambition to test every possible element.

Sometimes the honest result is that the test did not separate the combinations well enough to justify a change. That can happen when the differences are small relative to the available observations, or when too many cells have divided the traffic. The response is to examine the original decision: could a longer run deliver enough relevant visitors, or would a smaller design answer the more important question sooner? Optimizely recommends reducing variables or allowing more time when results are inconclusive. If the page must ship now, a straightforward A/B comparison of the two complete experiences under consideration may be the more useful experiment.

Multivariate testing pays off when the team has a stable page, a small set of publishable component options, a specific outcome, and enough traffic for the combinations. I would use it to learn whether a headline and image belong together, not as a way to place every idea on one crowded test plan. The best design is the one whose combinations the team can build, measure, interpret, and actually choose among.

Run your growth team from one screen.

Invite only