

A/B testing compares two versions of a page (a control and a variant) to see which one produces more of a chosen result, whether that’s sign-ups, purchases, or clicks. It’s worth running when you have enough traffic to reach roughly 350 to 1,000 conversions per variant, or when a wrong decision would be expensive to reverse. If neither applies, start with a five-user interview or a heatmap instead.
TL;DR:
- A/B testing requires at least 350 to 1,000 conversions per variant for reliable results; lower traffic typically favors qualitative methods.
- Focus testing on elements directly affecting revenue or user intent, such as headline copy, call-to-action wording, and form length, rather than cosmetic details.
- Confirm you have enough traffic before running a test and ensure the change is hard or costly to undo and would significantly impact revenue if wrong.
- Use visual editing tools and track primary metrics carefully, logging hypotheses, sample sizes, and test duration to maintain accuracy and compliance.
- Always verify accessibility compliance and segment results by device and source to avoid misleading conclusions and ensure a fair evaluation.
A/B testing, sometimes written as split testing, means showing half your visitors the original page and half a changed version, then measuring which one performs better against a single metric. Say you’re testing a checkout page: version A keeps your current “Buy Now” button, version B changes it to “Add to Basket.” Whichever version produces more completed purchases wins.
It differs from heatmaps and user interviews in one key way: those tools tell you what people do or why they hesitate, while A/B testing tells you which specific version actually converts better. Heatmaps generate hypotheses. A/B tests confirm or kill them.
A/B testing typically resolves questions like:
Not every idea deserves a live experiment. Before building one, ask three questions: do you have enough traffic, is the change reversible, and what happens if you guess wrong?
Use this checklist to decide:
If you answered yes to most of these, test it. If traffic is thin and the change is low risk (moving a testimonial block, tweaking a headline), just ship it and watch what happens, or run a small qualitative check first.
Pro Tip: Reversible, low stakes changes rarely justify the setup time of a formal test. Save A/B testing for decisions you’d genuinely regret getting wrong.
Start with elements that touch money or intent directly, not cosmetic details. In our experience with client sites, message level changes (what you’re saying, not how it looks) move the needle far more than colour or font tweaks.
Priority order for most SMB sites:
If you’re short on ideas, our guide on boosting website conversions walks through a wider set of tactics worth trying first.
This is where most small business tests fail before they even start. A common CRO rule of thumb suggests you need roughly 1,000 conversions per variant for a confident result, and around 350 for directional guidance that’s useful but not statistically bulletproof.

The maths behind this involves your baseline conversion rate, the minimum detectable effect (MDE) you care about, and standard statistical power and significance settings. Smaller expected improvements need larger samples to detect reliably; a test looking for a 2% lift needs far more traffic than one looking for a 20% lift.
Practical guidance for run time and traffic:
Statistic: Jakob Nielsen’s usability research found that testing with just five users typically uncovers around 85% of usability problems, which makes small qualitative studies a genuinely efficient alternative when your traffic can’t support a proper split test.
Two tool categories matter here: visual editors for building variants, and feature flagging platforms for more technical control. For most SMBs, visual editors remain the easier entry point because they don’t require developer time for every change.
A practical, budget-friendly combination pairs Microsoft Clarity for qualitative session data with a lightweight experiment platform like Convert, PostHog, or GrowthBook for the actual split testing. Worth noting: Google Optimize was discontinued in 2023, so any guide still recommending it is out of date.
Before launching anything, get the basics right:
A “winning” variant on your primary metric isn’t automatically a green light to ship it. Check secondary and revenue related metrics first, because a variant can lift sign ups while quietly damaging average order value or increasing support tickets.
Segmentation matters here too. A variant can win overall but lose for mobile users or a specific traffic source, so break results down by device and channel before declaring victory.
Once your test concludes, follow a consistent process:
Running variants without checking accessibility is a common blind spot. A new CTA button, a redesigned form, or a restructured navigation can all quietly break keyboard access or contrast ratios for users relying on assistive technology.
Watch for these basics during any live test:
Automated scanners like WAVE, ARC Toolkit, and similar tools catch a meaningful share of issues but not all of them, so pair automated checks with manual keyboard testing. Institutional guidance in Lithuania also points to the W3C Web Accessibility Evaluation Tools List as a solid starting reference. Our own accessibility guide covers manual checks in more depth.
Every test we help set up starts with a one line hypothesis: “Changing X will increase Y because Z.” That single sentence forces clarity before anything gets built.
Seven items to check before launch:
Three tests most SMBs can run immediately: shortening a contact form from eight fields to four, testing a visible “from” price on a services page, and testing guest checkout against forced account creation on an ecommerce checkout flow. We track downstream impact on revenue and repeat purchase rate, not just the immediate conversion, and log every test in a written registry so nothing gets repeated by accident six months later.
Pro Tip: If you can’t articulate the hypothesis in one sentence before building anything, the test isn’t ready yet.
Small sites chase statistical significance they’ll never reach. If your site gets 200 visitors a month, running a formal split test on button colour is a waste of a quarter. Spend that time on five user interviews instead.
Done.lu leans toward speed for low traffic clients and rigour for high traffic ones. That balance, not blind faith in testing, is the real skill. One takeaway: calculate your monthly conversions before writing a single hypothesis.
— Thomas
Done.lu is the practical route for SMBs in Luxembourg who want experiments run properly without hiring a dedicated CRO analyst. We’ve built this into how we handle digital marketing work generally: audit first, then test design, then tracking setup, then GDPR-aware deployment that doesn’t leak personal data into third-party dashboards you haven’t vetted.

A typical engagement starts with a short audit of your traffic and current conversion rates to work out whether a formal test is even viable, or whether user interviews would serve you better first. From there we build the hypothesis, wire the tracking correctly the first time (a mistake here invalidates weeks of data), and interpret results against secondary metrics before recommending a rollout.
If your site sits on our Website as a Service plans, starting at €195 a month for the One Pager tier, testing infrastructure and tracking can be built into ongoing maintenance rather than billed as a separate project each time. Get in touch through our consulting page to talk through whether your traffic supports testing yet, or whether a quicker qualitative pass makes more sense first.

For sample size logic, FastStrat’s guide is a solid start. For tooling and duration rules, see Kolonell’s method breakdown. Partner reading: Baby Love Growth’s CRO guide and Save Your App for qualitative research methods.
A common rule of thumb is around 1,000 conversions per variant for confident decisions and roughly 350 for directional guidance. Below that, treat results as a hint, not a verdict, and consider qualitative testing instead.
Run tests for at least one full business cycle, a minimum of two weeks and ideally four, so weekday and weekend traffic patterns both get represented. Stopping early because early numbers look good is one of the most common mistakes in split testing.
Skip formal A/B testing and run five user interviews or a short usability session instead, since testing with just five users typically surfaces around 85% of usability problems. This gives you actionable direction without needing statistical significance you can’t reach.
Yes, Done.lu designs hypotheses, sets up tracking, and deploys experiments as part of its digital marketing and website services. Pricing for managed website plans starts at €195 a month; get in touch through the consulting page for a project-specific quote.
Yes. Automated scanners like WAVE and Lighthouse catch only part of accessibility issues, so pair them with manual keyboard navigation checks on every new variant before it goes live to the full audience.