Why Analytics Look Different in Shogun A/B Testing vs. GA4 or Shopify
When you run an experiment in Shogun A/B Testing, you might notice that the numbers in your Shogun dashboard don't exactly match what you see in Shopify or Google Analytics 4 (GA4).
Some differences are expected because each platform tracks and attributes shopper activity differently. Shopify and GA4 can still provide useful points of comparison, but how closely the numbers should align depends on whether you're comparing the same group of shoppers.
In Shogun, the population reported for an experiment can be affected by:
the type of experiment you're running
where and when a shopper becomes eligible for the experiment
other experiments running at the same time
your experiment distribution method
any audiences applied to the experiment
whether an eligible shopper actually encounters the experience being tested
Understanding these factors is important before deciding whether a difference between Shogun and another analytics platform is expected or needs further investigation.
Shogun Measures Experiment Participants
Shogun experiment analytics are specific to shoppers participating in an experiment.
That doesn't always match the population represented in all sessions reported by Shopify or GA4.
For some experiment types, the populations can be very similar. For others, shoppers only become eligible once they reach a particular page or experience, which means only part of your overall store traffic can participate.
Other experiment settings can narrow that population further.
Because of this, the first question to ask when comparing analytics isn't:
"Why don't these numbers match?"
It's:
"Are these reports measuring the same shoppers?"
Experiment Type Affects Who Can Participate
Different experiment types can become eligible at different points in the shopper journey.
Experiments such as theme, price, product, checkout, and shipping tests can be eligible broadly across the store experience. As a result, Shogun and Shopify will often observe broadly similar session populations for these tests, assuming no other experiment settings restrict participation.
Shopify can therefore provide a useful point of comparison for these experiments.
Other experiment types — such as page, template, and URL redirect tests — depend on the shopper encountering a particular tested surface.
A shopper might therefore browse your store for some time before becoming eligible for one of these experiments, or they may never encounter the tested surface at all.
This can make overall Shopify traffic a less direct comparison.
Landing Page and Experiment Entry Aren't Always the Same Thing
This is particularly important when using Shopify or GA4 reports filtered by landing page.
A landing page tells you where a shopper's session started.
Experiment participation can happen later.
For example:
Google → Homepage → Collection → Experiment page → Checkout
The shopper's landing page is the homepage.
However, they can still become eligible for and participate in an experiment when they later visit the experiment page.
A Shopify or GA4 report filtered to sessions that landed on the experiment page wouldn't include this shopper, even though Shogun may correctly report them as an experiment participant.
A landing-page report therefore doesn't necessarily represent everyone who encountered an experiment during their session.
URL Redirect Tests Are an Important Example
This distinction is especially important for URL redirect experiments.
In a URL redirect test, the shopper encounters the original URL. Shogun can then dispatch the experiment and, depending on their variant, redirect them to an alternate URL.
For example:
Homepage → Original experiment URL → Redirected to Variant B URL → Checkout
The shopper participated in Variant B.
However:
their landing page is still the homepage
they didn't participate in Variant B because they landed on the Variant B URL
visiting the Variant B URL was a result of the experiment redirect
Because of this, comparing Variant B sessions in Shogun with sessions that landed on the Variant B URL in Shopify or GA4 isn't an equivalent comparison.
Why Shogun May Show More Sessions for a URL Redirect Test
Shogun may report more participating sessions for a URL redirect experiment than a Shopify report filtered to sessions that landed on the original experiment URL.
This doesn't necessarily mean Shogun is duplicating or overcounting sessions.
Consider three shoppers:
Shopper 1
Google → Original experiment URL → participates in experiment
Shopper 2
Google → Homepage → Original experiment URL → participates in experiment
Shopper 3
Email → Collection → Original experiment URL → participates in experiment
All three can participate in the Shogun experiment.
However, a Shopify report filtered to sessions whose landing page was the original experiment URL would only include Shopper 1.
Shoppers 2 and 3 entered the experiment after their sessions had already started elsewhere.
In this situation, the Shopify landing-page report can still provide useful context, but it represents only a partial view of the population that can participate in the Shogun experiment.
Concurrent Experiments Can Affect Participation
If you're running multiple Shogun experiments at the same time, this can also affect how many shoppers participate in each experiment.
A shopper is only assigned to one experiment at a time.
This means that simply being potentially eligible for an experiment doesn't guarantee the shopper will ultimately participate.
How shoppers are allocated across concurrent experiments depends on your distribution method.
Greedy Distribution
With Greedy distribution, experiments that are eligible broadly across the shopper journey — such as price, product, checkout, shipping, and theme tests — can take the shopper when they land on the store. That assignment then sticks.
Experiments that only become eligible once the shopper reaches a particular surface — such as page, template, and URL redirect tests — may therefore receive fewer participants when other experiments are running.
For example:
Shopper lands → assigned to an eligible price test → later visits a page with a URL redirect test
By the time the shopper reaches the URL redirect test, they're already participating in another experiment.
So even though they visited the URL being tested, you shouldn't necessarily expect them to appear as a participant in that URL redirect experiment.
Even Distribution
With Even distribution, shoppers are allocated across concurrent experiments according to their traffic allocation.
However, being allocated to an experiment doesn't guarantee that the shopper will ultimately encounter the experience being tested.
For example, a shopper could be allocated to a page experiment but never visit that page during their session.
In that situation, they won't produce an exposure for that experiment.
This is another reason why overall Shopify session totals shouldn't automatically be expected to match the sessions or exposures reported for an individual Shogun experiment.
Audiences Can Narrow the Population Further
If you've configured an audience for an experiment, only shoppers who meet those audience conditions are eligible to participate.
For example, imagine an experiment is targeted to a particular segment of your traffic.
Comparing the resulting Shogun experiment sessions with all store sessions in Shopify wouldn't be an equivalent comparison because Shopify's total includes shoppers who were never eligible for the experiment.
Before comparing Shogun with Shopify or GA4, always check whether an audience is configured and, where possible, make sure your external report represents a similar population.
Assignment and Exposure Aren't Always the Same Thing
Another useful distinction is that being allocated or assigned to an experiment doesn't always mean a shopper will actually encounter the experience being tested.
This is most noticeable with experiments that depend on the shopper reaching a particular page or surface.
For example:
Shopper allocated to Page Experiment → browses Homepage → Collection → leaves store
If the shopper never visits the tested page, they never encounter the experiment experience.
This distinction can become especially important when multiple experiments are running concurrently.
When comparing analytics, consider not only which shoppers could be allocated to an experiment, but which shoppers could actually become exposed to the experience being tested.
How Orders Get Counted in a Shogun Experiment
An order appearing in Shopify doesn't necessarily mean it should also appear as an order attributed to a particular Shogun experiment.
The shopper needs to have participated in that experiment, and the purchase needs to be attributable to their experiment session and variant.
Some Shopify orders may come from shoppers who:
weren't eligible for the experiment
didn't meet its audience conditions
were participating in another concurrent experiment
never encountered the surface being tested
completed their purchase without participating in the experiment at all
Tracking factors such as ad blockers, consent settings, site performance, and competing scripts can also create differences between analytics platforms.
For these reasons, total Shopify orders shouldn't automatically be expected to equal the orders attributed to a specific Shogun experiment.
Why Conversion Rates Can Differ
Conversion rate depends on both:
the number of converting sessions ÷ the number of sessions included in the calculation
If Shogun and Shopify are measuring different groups of shoppers, their conversion rates can differ even when both systems are functioning correctly.
For example, comparing:
Shopify conversions from sessions that landed on a particular URL ÷ Shopify sessions that landed on that URL
with:
Shogun conversions attributed to an experiment variant ÷ Shogun participating sessions for that variant
isn't necessarily an apples-to-apples comparison.
Before comparing conversion rates, first establish whether the underlying populations are comparable.
When Should Shopify and Shogun Be Comparable?
Shopify can be a valuable comparison source when the populations represented by both systems should be broadly similar.
For example, an experiment such as a theme test that applies broadly across the storefront may include a large proportion of the same sessions Shopify observes.
In situations like this, Shopify's overall session data can provide a strong point of comparison with Shogun.
You still shouldn't necessarily expect an exact match. Differences can occur because of ad blockers, consent settings, session definitions, tracking behavior, and other technical factors.
For experiments whose participation depends on a shopper encountering a particular page or surface, Shopify may instead provide only a partial comparison.
Concurrent experiments and audiences can further change which shoppers are eligible to participate.
The key question is therefore:
Should these two reports represent approximately the same population?
How to Compare Shogun With Shopify or GA4
Before comparing numbers between platforms, check the following:
What type of experiment are you running? Does it apply broadly across the shopper journey, or does the shopper need to encounter a particular page or surface?
Are other experiments running concurrently? A shopper can only be assigned to one experiment, so other active experiments can affect participation.
Which distribution method are you using? Greedy and Even distribution can affect which experiment a shopper participates in and whether they ultimately encounter the tested experience.
Is an audience configured? If so, not all Shopify traffic represents the same eligible population.
Is your Shopify or GA4 report filtered by landing page? Remember that shoppers can land elsewhere and enter some experiments later.
For a URL redirect test, are you comparing experiment participation with sessions that simply landed on a variant URL? These aren't equivalent populations.
Are the date ranges the same? Account for the exact time the experiment started and stopped.
Are the metrics defined similarly? Similarly named metrics across analytics platforms don't necessarily use identical tracking or attribution rules.
Only after establishing these factors should you use the size of a discrepancy to determine whether something looks unexpected.
What's Normal vs. When to Reach Out
Some differences between Shogun, Shopify, and GA4 are expected, but there isn't a single percentage threshold that determines whether a discrepancy is normal.
A difference can be caused by:
comparing different shopper populations
the type of experiment
concurrent experiments
distribution settings
audience targeting
shopper navigation
ad blockers or consent settings
differences in tracking and attribution
This doesn't mean you should dismiss every discrepancy as a difference between analytics platforms.
If you've accounted for these factors and you're still seeing a consistent or unexplained difference between populations that should be similar, please contact Shogun Support.
Our team can investigate whether your experiment is being dispatched, tracked, and attributed as expected and determine whether the difference is expected or requires further investigation.
How Quickly Shogun Data Updates
Shogun data updates every 5 minutes. This includes sessions, clicks, and orders.
Orders can sometimes feel like they take longer to appear because they need to be tied back to the shopper's experiment session and variant before appearing in your dashboard.
If a shopper completes a purchase that can be attributed to an experiment, you can generally expect it to appear within a few minutes.