Skip to main content

Why Analytics Look Different in Shogun A/B Testing vs. GA4 or Shopify

Shogun A/B Testing metrics may not match GA4 or Shopify. Learn how Shogun attributes sessions, variants, and orders only for experiment participants, why conversions can differ, and when to contact support.

Written by James Power

Why Analytics Look Different in Shogun A/B Testing vs. GA4 or Shopify

When you run an experiment in Shogun A/B Testing, you might notice that the numbers in your Shogun dashboard don't exactly match what you see in Shopify or Google Analytics 4 (GA4).

Some differences are expected because each platform tracks and attributes shopper activity differently. Shopify and GA4 can still provide useful points of comparison, but how closely the numbers should align depends on whether you're comparing the same group of shoppers.

In Shogun, the population reported for an experiment can be affected by:

  • the type of experiment you're running

  • where and when a shopper becomes eligible for the experiment

  • other experiments running at the same time

  • your experiment distribution method

  • any audiences applied to the experiment

  • whether an eligible shopper actually encounters the experience being tested

Understanding these factors is important before deciding whether a difference between Shogun and another analytics platform is expected or needs further investigation.

Shogun Measures Experiment Participants

Shogun experiment analytics are specific to shoppers participating in an experiment.

That doesn't always match the population represented in all sessions reported by Shopify or GA4.

For some experiment types, the populations can be very similar. For others, shoppers only become eligible once they reach a particular page or experience, which means only part of your overall store traffic can participate.

Other experiment settings can narrow that population further.

Because of this, the first question to ask when comparing analytics isn't:

"Why don't these numbers match?"

It's:

"Are these reports measuring the same shoppers?"

Experiment Type Affects Who Can Participate

Different experiment types can become eligible at different points in the shopper journey.

Experiments such as theme, price, product, checkout, and shipping tests can be eligible broadly across the store experience. As a result, Shogun and Shopify will often observe broadly similar session populations for these tests, assuming no other experiment settings restrict participation.

Shopify can therefore provide a useful point of comparison for these experiments.

Other experiment types — such as page, template, and URL redirect tests — depend on the shopper encountering a particular tested surface.

A shopper might therefore browse your store for some time before becoming eligible for one of these experiments, or they may never encounter the tested surface at all.

This can make overall Shopify traffic a less direct comparison.

Landing Page and Experiment Entry Aren't Always the Same Thing

This is particularly important when using Shopify or GA4 reports filtered by landing page.

A landing page tells you where a shopper's session started.

Experiment participation can happen later.

For example:

Google → Homepage → Collection → Experiment page → Checkout

The shopper's landing page is the homepage.

However, they can still become eligible for and participate in an experiment when they later visit the experiment page.

A Shopify or GA4 report filtered to sessions that landed on the experiment page wouldn't include this shopper, even though Shogun may correctly report them as an experiment participant.

A landing-page report therefore doesn't necessarily represent everyone who encountered an experiment during their session.

URL Redirect Tests Are an Important Example

This distinction is especially important for URL redirect experiments.

In a URL redirect test, the shopper encounters the original URL. Shogun can then dispatch the experiment and, depending on their variant, redirect them to an alternate URL.

For example:

Homepage → Original experiment URL → Redirected to Variant B URL → Checkout

The shopper participated in Variant B.

However:

  • their landing page is still the homepage

  • they didn't participate in Variant B because they landed on the Variant B URL

  • visiting the Variant B URL was a result of the experiment redirect

Because of this, comparing Variant B sessions in Shogun with sessions that landed on the Variant B URL in Shopify or GA4 isn't an equivalent comparison.

Why Shogun May Show More Sessions for a URL Redirect Test

Shogun may report more participating sessions for a URL redirect experiment than a Shopify report filtered to sessions that landed on the original experiment URL.

This doesn't necessarily mean Shogun is duplicating or overcounting sessions.

Consider three shoppers:

Shopper 1

Google → Original experiment URL → participates in experiment

Shopper 2

Google → Homepage → Original experiment URL → participates in experiment

Shopper 3

Email → Collection → Original experiment URL → participates in experiment

All three can participate in the Shogun experiment.

However, a Shopify report filtered to sessions whose landing page was the original experiment URL would only include Shopper 1.

Shoppers 2 and 3 entered the experiment after their sessions had already started elsewhere.

In this situation, the Shopify landing-page report can still provide useful context, but it represents only a partial view of the population that can participate in the Shogun experiment.

Concurrent Experiments Can Affect Participation

If you're running multiple Shogun experiments at the same time, this can also affect how many shoppers participate in each experiment.

A shopper is only assigned to one experiment at a time.

This means that simply being potentially eligible for an experiment doesn't guarantee the shopper will ultimately participate.

How shoppers are allocated across concurrent experiments depends on your distribution method.

Greedy Distribution

With Greedy distribution, experiments that are eligible broadly across the shopper journey — such as price, product, checkout, shipping, and theme tests — can take the shopper when they land on the store. That assignment then sticks.

Experiments that only become eligible once the shopper reaches a particular surface — such as page, template, and URL redirect tests — may therefore receive fewer participants when other experiments are running.

For example:

Shopper lands → assigned to an eligible price test → later visits a page with a URL redirect test

By the time the shopper reaches the URL redirect test, they're already participating in another experiment.

So even though they visited the URL being tested, you shouldn't necessarily expect them to appear as a participant in that URL redirect experiment.

Even Distribution

With Even distribution, shoppers are allocated across concurrent experiments according to their traffic allocation.

However, being allocated to an experiment doesn't guarantee that the shopper will ultimately encounter the experience being tested.

For example, a shopper could be allocated to a page experiment but never visit that page during their session.

In that situation, they won't produce an exposure for that experiment.

This is another reason why overall Shopify session totals shouldn't automatically be expected to match the sessions or exposures reported for an individual Shogun experiment.

Audiences Can Narrow the Population Further

If you've configured an audience for an experiment, only shoppers who meet those audience conditions are eligible to participate.

For example, imagine an experiment is targeted to a particular segment of your traffic.

Comparing the resulting Shogun experiment sessions with all store sessions in Shopify wouldn't be an equivalent comparison because Shopify's total includes shoppers who were never eligible for the experiment.

Before comparing Shogun with Shopify or GA4, always check whether an audience is configured and, where possible, make sure your external report represents a similar population.

Assignment and Exposure Aren't Always the Same Thing

Another useful distinction is that being allocated or assigned to an experiment doesn't always mean a shopper will actually encounter the experience being tested.

This is most noticeable with experiments that depend on the shopper reaching a particular page or surface.

For example:

Shopper allocated to Page Experiment → browses Homepage → Collection → leaves store

If the shopper never visits the tested page, they never encounter the experiment experience.

This distinction can become especially important when multiple experiments are running concurrently.

When comparing analytics, consider not only which shoppers could be allocated to an experiment, but which shoppers could actually become exposed to the experience being tested.

How Orders Get Counted in a Shogun Experiment

An order appearing in Shopify doesn't necessarily mean it should also appear as an order attributed to a particular Shogun experiment.

The shopper needs to have participated in that experiment, and the purchase needs to be attributable to their experiment session and variant.

Some Shopify orders may come from shoppers who:

  • weren't eligible for the experiment

  • didn't meet its audience conditions

  • were participating in another concurrent experiment

  • never encountered the surface being tested

  • completed their purchase without participating in the experiment at all

Tracking factors such as ad blockers, consent settings, site performance, and competing scripts can also create differences between analytics platforms.

For these reasons, total Shopify orders shouldn't automatically be expected to equal the orders attributed to a specific Shogun experiment.

Why Conversion Rates Can Differ

Conversion rate depends on both:

the number of converting sessions ÷ the number of sessions included in the calculation

If Shogun and Shopify are measuring different groups of shoppers, their conversion rates can differ even when both systems are functioning correctly.

For example, comparing:

Shopify conversions from sessions that landed on a particular URL ÷ Shopify sessions that landed on that URL

with:

Shogun conversions attributed to an experiment variant ÷ Shogun participating sessions for that variant

isn't necessarily an apples-to-apples comparison.

Before comparing conversion rates, first establish whether the underlying populations are comparable.

When Should Shopify and Shogun Be Comparable?

Shopify can be a valuable comparison source when the populations represented by both systems should be broadly similar.

For example, an experiment such as a theme test that applies broadly across the storefront may include a large proportion of the same sessions Shopify observes.

In situations like this, Shopify's overall session data can provide a strong point of comparison with Shogun.

You still shouldn't necessarily expect an exact match. Differences can occur because of ad blockers, consent settings, session definitions, tracking behavior, and other technical factors.

For experiments whose participation depends on a shopper encountering a particular page or surface, Shopify may instead provide only a partial comparison.

Concurrent experiments and audiences can further change which shoppers are eligible to participate.

The key question is therefore:

Should these two reports represent approximately the same population?

How to Compare Shogun With Shopify or GA4

Before comparing numbers between platforms, check the following:

  • What type of experiment are you running? Does it apply broadly across the shopper journey, or does the shopper need to encounter a particular page or surface?

  • Are other experiments running concurrently? A shopper can only be assigned to one experiment, so other active experiments can affect participation.

  • Which distribution method are you using? Greedy and Even distribution can affect which experiment a shopper participates in and whether they ultimately encounter the tested experience.

  • Is an audience configured? If so, not all Shopify traffic represents the same eligible population.

  • Is your Shopify or GA4 report filtered by landing page? Remember that shoppers can land elsewhere and enter some experiments later.

  • For a URL redirect test, are you comparing experiment participation with sessions that simply landed on a variant URL? These aren't equivalent populations.

  • Are the date ranges the same? Account for the exact time the experiment started and stopped.

  • Are the metrics defined similarly? Similarly named metrics across analytics platforms don't necessarily use identical tracking or attribution rules.

Only after establishing these factors should you use the size of a discrepancy to determine whether something looks unexpected.

What's Normal vs. When to Reach Out

Some differences between Shogun, Shopify, and GA4 are expected, but there isn't a single percentage threshold that determines whether a discrepancy is normal.

A difference can be caused by:

  • comparing different shopper populations

  • the type of experiment

  • concurrent experiments

  • distribution settings

  • audience targeting

  • shopper navigation

  • ad blockers or consent settings

  • differences in tracking and attribution

This doesn't mean you should dismiss every discrepancy as a difference between analytics platforms.

If you've accounted for these factors and you're still seeing a consistent or unexplained difference between populations that should be similar, please contact Shogun Support.

Our team can investigate whether your experiment is being dispatched, tracked, and attributed as expected and determine whether the difference is expected or requires further investigation.

How Quickly Shogun Data Updates

Shogun data updates every 5 minutes. This includes sessions, clicks, and orders.

Orders can sometimes feel like they take longer to appear because they need to be tied back to the shopper's experiment session and variant before appearing in your dashboard.

If a shopper completes a purchase that can be attributed to an experiment, you can generally expect it to appear within a few minutes.

Did this answer your question?