01How Should You Compare A/B Testing Software?
Every A/B testing comparison you can find is built the same way: a grid of tools against a list of features, with checkmarks. That format stopped being useful around the time every vendor shipped the same feature set. Visual editor, audience targeting, custom goals, segment reports, multivariate testing, integrations: the grid is now almost solid green, which tells you nothing about which tool will make your program work.
The dimensions below are the ones we use when a client asks us to run a selection. They are ordered by how often they turn out to be decisive, and the fifth is the one we added after our own program moved from a 27% win rate in 2024 to 55% in the most recent quarter without changing tools.
- Statistics engine: frequentist, Bayesian, or undocumented. This decides how long a test runs, when you are allowed to look, and how often the tool will tell you a flat test is a winner.
- Data ownership and location: who processes visitor data, in which jurisdiction, and whether the tool keeps its own tracking or reads from your analytics. This is the dimension legal can veto.
- Editor model: visual, code-first, or server-side. The question is not which is better, it is which one matches the people who will actually build variants next quarter.
- Pricing model: flat rate, traffic tiers, visitor credits, per impression, self-hosted, or quoted. The entry price is marketing. The model is the cost.
- Idea selection: does the product help you decide which tests to run? Almost no tool answers this, and it is where most of the money in a testing program is won or lost.
02Which Statistics Engine Does a Tool Use, and Does It Change Your Decisions?
The statistics engine is the least visible part of a testing tool and the part that decides whether your results are real. Two families exist in this market, and they fail in different ways.
A frequentist engine asks you to calculate a sample size before the test starts, then judge the result once that sample is reached. It is strict, sometimes inconvenient, and hard to fool. A Bayesian engine reports the probability that a variant beats the control, updated continuously. VWO’s SmartStats is the clearest example in this list, and its output is genuinely easier for non-statisticians to read. GrowthBook supports both and lets the team choose.
We run frequentist at DRIP, and the reason is behavioral rather than mathematical. A continuously updating probability invites the single most expensive habit in e-commerce testing: stopping a test on the day the numbers look good. A pre-calculated sample size makes that stop visible as a violation instead of a decision. Both engines can be used correctly. Only one of them is hard to abuse.
What to ask about the engine
- Is the engine frequentist, Bayesian, or a sequential variant, and is that documented publicly?
- Does the tool calculate a required sample size before the test starts?
- What does the interface show on day three, and does it warn against acting on it?
- How are multiple metrics handled, and is any correction applied when you track many goals?
- Can we export raw assignment and conversion data to check a result ourselves?
The last question matters more than it looks. A tool that lets you export raw data can be audited. A tool that only shows you its own verdict cannot, and you are then buying the vendor’s statistical judgment along with the software.
03Where Does Your Test Data Live, and Who Owns It?
For a European online shop, this dimension can end an evaluation in one meeting. It has three distinct layers, and vendors tend to answer only the first.
- Processing location: where visitor data is stored and computed. ABlyft processes on German servers, Kameleoon and AB Tasty run on European infrastructure, VWO offers an EU data center as an option rather than a default, and Optimizely is US-headquartered with Standard Contractual Clauses and a Data Processing Agreement in place.
- Own tracking or your analytics: most tools set their own identifiers and compute their own conversions. Varify.io deliberately does not, reading results from Google Analytics 4, Matomo, or Piwik Pro instead. GrowthBook reads directly from your warehouse, whether that is BigQuery, Snowflake, Postgres, or ClickHouse.
- Architecture or configuration: a tool that is compliant by architecture cannot be misconfigured into non-compliance. A tool that is compliant by configuration leaves that burden with your team, and your team changes over time.
There is a commercial consequence that legal teams rarely mention. A tool that needs consent before it can activate only ever sees the consented share of your traffic. Every visitor who declines is invisible to the test, so your sample shrinks and your test duration grows. Cookieless and server-side architectures avoid that, and on a low-traffic shop the difference decides how many decisions per year you can make.
04Does the Editor Model Match the Team That Will Run the Tests?
The editor question is usually asked as “which editor is best”. The better question is who will build the next twenty variants, and what happens to your program when that person is busy.
| Model | Who can operate it | The trade-off |
|---|---|---|
| Visual, on-page | Marketers and CRO specialists, after a one-time snippet | Fast to first test, heavier script, limits on deep DOM changes |
| Code-first | Frontend developers | Unlimited flexibility and minimal page weight, needs developer time |
| Server-side | Backend developers | Reaches pricing, logic, and headless surfaces, no visual shortcut |
| Managed execution | An external team, with your approval | Removes the QA and analysis risk, costs more than self-service |
Two failure modes recur. A marketing team buys a code-first tool because it benchmarks well on performance, then ships four tests a year because every variant waits in an engineering queue. A developer team buys a visual editor, then works around it for every non-trivial test, carrying its script weight for no benefit. Both are avoidable by asking one question before the demo: who builds variant B?
Page weight belongs in this dimension too, because the editor model largely determines it. ABlyft keeps its client-side script under 5 KB by running the editor as a Chrome extension so no editor runtime loads on your storefront. VWO, Kameleoon, and AB Tasty sit in the 30 to 40 KB range because the editor and behavioral analytics ship to the browser. Optimizely’s client-side snippet is the heaviest in this comparison at roughly 80 KB. On a shop where Core Web Vitals already affect rankings, that is a real cost of the convenience.
05How Do the Major Tools Compare on These Dimensions?
Where a cell below says “not publicly documented”, we did not fill the gap with an assumption. An empty cell is a question for the vendor, and how they answer it is part of the evaluation.
| Tool | Statistics engine | Data processing | Editor model | Pricing model |
|---|---|---|---|---|
| Apex by DRIP | Frequentist, fixed horizon with Holm-Bonferroni correction and an SRM gate, optional always-valid sequential analysis (mSPRT) | UK (London) and Ireland, one first-party visitor ID, no fingerprinting, no IP storage | In-shop testing tool plus managed execution | Book a call |
| ABlyft | Not publicly documented | German servers, cookieless by default | Code-first, visual editor as Chrome extension | Traffic tiers from €79/mo, free plan |
| Varify.io | Reads results from your analytics | German company, EU hosting, no own tracking | Visual, on-page | Flat rate, €149 or €249/mo |
| VWO | Bayesian (SmartStats) | EU data center available as an option | Visual, plus server-side via FME | Traffic tiers, $139 to $775/mo, free tier |
| Kameleoon | Not publicly documented | French-built, EU processing, CNIL-aligned | Visual, plus 10+ server-side SDKs | Quote, market data €25,000 to €50,000/yr |
| AB Tasty | Not publicly documented | European infrastructure, ISO 27001 | Visual, strongest here, plus Flagship SDKs | Visitor credits from roughly €15,000/yr |
| Optimizely | Multi-armed bandits available | US-headquartered, SCCs and DPA, SOC 2 Type II | Visual, plus full-stack and edge | Per impression, typically $36,000+/yr |
| GrowthBook | Frequentist or Bayesian, your choice | Self-hosted, no third-party processor | Code-first, no visual editor | Free self-hosted, cloud from $99/mo |
Read the pricing column as a forecast rather than a price. Flat rate means traffic growth is free. Traffic tiers step up predictably if you know your sessions. Visitor credits are consumption-based and more forecastable than pure traffic pricing, but expensive at volume. Per impression means a viral week moves your bill. Self-hosted keeps the licence flat and converts the cost into engineering time. Quoted pricing, ours included, means you cannot compare without a conversation, which is a real friction cost.
06Which Dimension Do Comparison Tables Leave Out?
The first four dimensions all measure execution: how results are computed, where data goes, who can build a variant, what it costs to run. Not one of them touches the question that decides the return on the whole program, which is whether the test was worth running.
The arithmetic is unforgiving. Industry-wide, about 1 in 5 tests produces a real winner. Each of the other four still consumes design time, developer time, QA, and two to four weeks of traffic on a page you could have been fixing instead. A better editor makes those three cycles faster. It does not make them fewer.
This is the part only volume can produce, so here is our own record. DRIP has run 4,000+ experiments for 50+ e-commerce brands. In 2024 our win rate was 27%. In the most recent quarter it was 55%. We did not change testing tools and we did not double our test volume. The improvement came from the ideas we rejected before anyone built them, and the reliable signal was never the element being changed.
- Change class beats element: “Reduce decision cost on the product page” has a track record. “Make the button green” does not and never will, because the same element wins on one shop and loses on the next.
- Context decides the sign: the same change often flips direction between a high-consideration, high-price catalog and an impulse catalog. Prior outcomes on comparable shops carry that information. A hypothesis document does not.
- The best output is a rejection: the most valuable thing a test memory returns is the list of ideas you do not run. Nobody celebrates it, and it is where the win rate actually comes from.
How Apex fits this framework
Apex is our A/B testing platform for online shops, and it exists for this fifth dimension only. It has three layers. The test memory holds 4.3 million A/B tests from 151,000 shops, collected over eight years, and scores every idea before launch. The testing tool builds, launches, and evaluates tests directly in the shop, so each prediction is checked against the result. Managed execution means tests are built, QA’d, launched, and analyzed with the DRIP team.
Judged on the other four dimensions, Apex is not the winner here. Our pricing is quoted rather than published, we have no public review profile yet, our script weight is not published, and our hosting sits in the UK (London) and Ireland rather than in the EU, which ABlyft and Varify.io answer with EU hosting. A prediction is also a prior, not a guarantee, and our win rate is our program’s record rather than a promise for any single shop. If your team already picks tests well, the fifth dimension is worth little to you and a cheaper self-service tool is the rational purchase.
So use the framework rather than the ranking. Score the shortlist on statistics, data, editor, and pricing model against your own constraints. Then ask each vendor the fifth question: what in your product helps me decide which tests to run? Most will answer with a hypothesis template. We answer with 4.3 million past outcomes, and you should hold us to that answer as strictly as any other.
Want Apex to score your test ideas before you build them? See if your shop is a fit→



