Hypotheses grounded in evidence
Tests come from analytics, heatmaps, recordings, surveys and sales feedback, not random ideas.
Replace opinions with evidence: well-designed experiments that show which changes genuinely improve conversions, and which only look like they do.
Every website change is a bet. A new headline, a shorter form, a different pricing layout or a redesigned checkout might improve results, make no difference or quietly make things worse. Without testing, businesses often judge changes by comparing before and after, which is easily distorted by seasons, campaigns, festivals, news and random variation. Decisions end up being made by whoever has the strongest opinion in the room.
A/B testing replaces guesswork with controlled experiments. Visitors are randomly split between the current version (A) and one or more variations (B), and results are compared over the same period, so external factors affect both groups equally. When designed properly, with a clear hypothesis, enough traffic, a pre-defined sample size and honest statistical analysis, an A/B test tells you whether a change caused a real improvement, and by roughly how much.
Our A/B testing services cover the full experimentation process: developing hypotheses from research, calculating whether your traffic can support a test, choosing the right testing tool and method, building variations, quality-checking them across devices, running tests for the right duration, analysing results responsibly, and documenting learnings in a shared library. We follow search engine guidance so tests do not harm SEO, and we are honest when a site does not have enough traffic for reliable testing, recommending alternative approaches through our conversion rate optimization service.
Last updated:
Tools & technologies
Tests come from analytics, heatmaps, recordings, surveys and sales feedback, not random ideas.
Sample size, duration and success metrics are decided in advance, preventing premature or biased conclusions.
We report uncertainty, avoid peeking and cherry-picking, and treat inconclusive results as useful learning.
Secondary metrics such as revenue per visitor, lead quality and refunds ensure a win does not hide a loss elsewhere.
Experiments follow search engine guidance on redirects, canonical tags and avoiding cloaking.
Every variation is tested across devices and browsers before launch to avoid broken experiences.
A searchable library of tests, results and insights guides future design and marketing decisions.
Typically 1 week
Traffic, conversions, tracking accuracy and testing tools reviewed.
Typically 1 week
Research-based hypotheses documented and prioritised.
Typically 1–2 weeks per test
Variations designed, developed and tested across devices.
Typically 2–6 weeks
Test runs for the planned duration with health monitoring.
Typically 3–5 days
Results analysed, decisions made and learnings recorded.
A/B testing, also called split testing, compares two or more versions of a page, element or experience by randomly showing each to a portion of visitors and measuring which performs better on a chosen goal. Because both versions run at the same time with similar audiences, differences in results can be attributed to the change rather than to seasons, campaigns or chance, provided the test is designed and analysed properly.
Different test types suit different situations.
A good hypothesis states what you will change, for whom, what you expect to happen and why, based on evidence. For example: because recordings show mobile visitors scrolling past the delivery information, moving delivery dates above the add-to-cart button on mobile product pages will increase add-to-cart rate. Clear hypotheses make results easier to interpret, whether the test wins or loses.
Required traffic depends on your current conversion rate, the size of improvement you want to detect and the confidence you require. Detecting small improvements on low conversion rates needs large samples; detecting larger improvements needs fewer. We calculate sample sizes before each test. If the numbers show a test would take many months, we recommend testing bigger changes, testing on higher-traffic pages, using a higher-funnel metric or choosing research-led improvements instead.
Statistical significance indicates how unlikely a result would be if there were no real difference between versions. It helps separate genuine effects from random noise, but it is often misunderstood. It does not tell you the size of the effect or guarantee future results. We report confidence intervals, showing a plausible range for the true effect, alongside significance, so decisions consider both certainty and magnitude.
Checking results repeatedly and stopping as soon as one version looks ahead, known as peeking, greatly increases the chance of false winners. Early results are volatile, and many apparent wins disappear with more data. Tests should run until the planned sample size is reached and, ideally, for full weekly cycles, since behaviour differs between weekdays and weekends. Some statistical methods are designed to allow ongoing monitoring, and we use them where appropriate.
Most tests should run at least one to two full weeks to capture weekly patterns, and until the planned sample size is reached. Very long tests have their own problems, such as cookie deletion and changing marketing activity. Tests should also avoid major disruptions, such as big sale events or festivals, unless the test is specifically about those periods.
A variation might increase one metric while harming another. A pop-up could increase email sign-ups but reduce purchases; a shorter form could increase leads but lower their quality; a discount could raise conversion rate but cut revenue. Guardrail metrics, such as revenue per visitor, average order value, qualified lead rate, refunds or page speed, are monitored so wins are genuine business improvements.
If a test is set to split traffic evenly but one version receives noticeably more visitors than expected, something is wrong: a redirect failing, bots, tracking errors or browser issues. This sample ratio mismatch can invalidate results. We check for it during and after tests and investigate before trusting any result.
Search engines support testing when it is done correctly. Good practice includes not showing search engines different content from users, using temporary redirects for split URL tests, adding canonical tags pointing to the original page on variation URLs, and ending tests and removing variations once they are complete. We follow these practices so experiments do not harm search visibility. Our SEO services team reviews tests that affect important organic pages.
Client-side testing tools change pages in the visitor’s browser using JavaScript. They are easy to set up and suit visual and content changes, but can cause flicker, where the original version appears briefly before the variation, and can affect performance. Server-side testing delivers variations from the server, avoiding flicker and supporting deeper changes such as pricing logic, search algorithms and app features, but it requires development work. We choose the method based on what is being tested.
Common tools include VWO, Optimizely, AB Tasty and Convert for websites, and Firebase A/B Testing or feature flag platforms for apps. Google Optimize has been discontinued, so businesses that relied on it need an alternative. We recommend tools based on traffic, budget, technical setup and the kinds of tests you plan to run.
The best tests address evidence-backed problems on high-traffic, high-value pages: headlines and value propositions on landing pages, product page information, pricing page layouts, form length and design, checkout steps, calls to action and trust signals. Tiny cosmetic changes, such as button colours, rarely produce meaningful results unless they fix a genuine visibility problem.
Segment analysis, such as mobile versus desktop or new versus returning visitors, can reveal useful patterns but also creates false positives, because looking at many segments makes some differences appear by chance. We define key segments before a test starts and treat unexpected segment results as ideas for future tests rather than definitive conclusions.
Inconclusive results are common and valuable. They show that the change did not matter enough to detect, which saves you from investing in changes that do not help, and they may indicate that the underlying assumption was wrong. Documenting these outcomes prevents repeating the same ideas and helps focus on bigger opportunities.
Yes. Google Ads experiments and Meta’s testing tools can compare bidding strategies, audiences, creatives and landing pages with controlled splits. These are useful for measuring the true effect of campaign changes. Our paid media team runs platform experiments, while website tests focus on on-site experience.
Over time, tests build knowledge about what your customers respond to. A library recording each hypothesis, variation, result, confidence and learning helps new team members, informs redesigns and marketing messages, and stops teams re-running old tests. It turns experimentation into a lasting asset.
These mistakes are common.
Returning visitors sometimes react to anything new simply because it is different, producing a temporary lift or dip that fades as they get used to the change. Running tests long enough, and comparing results for new and returning visitors, helps distinguish lasting improvements from short-term novelty.
Price testing can be valuable but needs care. Showing different prices to different customers for the same product can damage trust and raise fairness and legal concerns. Safer approaches include testing how prices are presented, such as plan layouts, bundles, payment options or annual versus monthly framing, rather than the prices themselves.
Costs depend on the number of tests per month, the complexity of variations and development, client-side or server-side implementation, quality assurance effort, analysis depth and documentation. Testing tool subscriptions are separate and usually priced by traffic volume.
A/B testing runs as a monthly experimentation programme, usually alongside CRO research and implementation. Testing tool subscriptions are paid directly to vendors.
Hypothesis development, sample size planning, tool setup, variation design and development, QA, monitoring, analysis and documentation.
It depends on conversion rates and the size of improvement you want to detect. We calculate this before recommending tests.
Usually two to six weeks, running full weekly cycles until the planned sample size is reached.
Not when done correctly. We follow search engine guidance on redirects, canonical tags and avoiding cloaking.
Options include VWO, Optimizely, AB Tasty and Convert. We recommend based on traffic, budget and needs.
We focus on research-led improvements, bigger changes or higher-funnel metrics, since small tests may never reach a conclusion.
Yes, using Firebase A/B Testing or feature flag tools, with app releases planned accordingly.
The winning version is implemented permanently, the test is removed and learnings are recorded for future work.
Tell us about your goals for A/B Testing. We will reply with a clear recommendation, timeline and written proposal.
Share your requirements and we will send a tailored proposal.