
Designing a Shopify Store That AI Agents Can Actually Understand and Buy From

Here’s what nobody tells you before your first call with a conversion rate optimization agency: you probably don’t know what a good answer sounds like yet. You’ll ask about A/B testing. They’ll say yes, of course, we do A/B testing. Every agency says yes to that question. It tells you nothing.

The mistake isn’t picking the wrong price tier or missing a red flag on some checklist. It’s walking into the conversation without knowing which questions actually separate a real experimentation partner from an agency that runs a handful of shallow tests and calls it a program. We see this on the buyer side constantly, across ten years and thousands of these conversations. Brands come in ready to evaluate cost and turnaround time. Almost nobody comes in ready to evaluate rigor.
So this isn’t a comparison table of agency logos. It’s the list of questions we’d want a friend to ask before they signed anything, in the order they actually matter.
A CRO agency runs structured experiments, usually A/B and multivariate tests, to increase the percentage of visitors who take a desired action on your site or app: buying, signing up, completing a form. The good ones don’t stop at testing. They audit funnels, research user behavior, and build a prioritized roadmap so every test ships because of a hypothesis, not a hunch.
That’s the textbook version. In practice, the range between agencies is enormous, and it’s not always visible from the homepage.
Before you take a single call, pull together what you already have: past test results (wins and losses both), session recordings, funnel drop-off data, customer support tickets, survey responses, anything that hints at where visitors get stuck.
Here’s why this matters more than people expect. An agency that starts from zero has to spend the first month guessing at hypotheses. An agency handed six months of prior test history and a stack of session recordings can start testing in week one, because the friction points are already half-mapped. We’ve had discovery calls where a prospect showed up with nothing, not even Google Analytics access sorted out, and the honest answer was: let’s fix that first, testing later. The agencies who tell you that upfront are worth more than the ones who promise to start testing next Monday regardless.
If you have nothing yet, that’s fine. Just don’t expect week-one results, and be suspicious of any agency that promises them anyway.
Every agency runs A/B tests. The number that actually predicts whether you’ll see results is how many tests ship per month, and what fraction of them are meaningful (not a button color change, but a hypothesis that could move revenue).

Ask directly: how many experiments will be live at any given time, and how many ship in a typical month? A team running two tests a month on a mid-traffic site is going to take a long time to compound into anything. A team running eight to ten, backed by proper prioritization, learns faster and adjusts faster. Speed compounds. Slow experimentation doesn’t just delay wins, it delays the learning that makes every subsequent test smarter.

Low velocity isn’t automatically a dealbreaker. If your traffic is low, fewer, better-designed tests might be the right call. But the agency should be able to explain the number, not just state it.
Ask this one directly: what’s your significance threshold, and how do you decide when a test has enough data to call it? If the answer is vague, that’s the flag to notice.

Peeking at results early and calling a test the moment it looks good is one of the most common ways agencies (and in-house teams) fool themselves. A test that looks like a 15% lift on day three can regress to nothing by day fourteen. Ask whether they pre-register sample sizes before launching a test, or whether they’re checking daily and stopping whenever the number looks good. The second approach produces exciting reports and unreliable results.
This is also where “does the agency understand your business economics” shows up. A test that lifts conversion rate but tanks average order value isn’t a win. Good agencies will tell you that upfront, sometimes before you even ask.
This is the split that matters most and gets talked about least: does the agency design and develop the winning variations themselves, or do they hand you a document and wait for your dev team to implement it?

The second model sounds cheaper. It usually isn’t, because your internal team now owns a second backlog, competing with product roadmap and everything else on their plate. We built OptiPhoenix with in-house build and server-side testing capability specifically because we watched too many “insights-only” engagements die in a client’s dev queue, sometimes for months, by which point the market had moved and the insight was stale.
Ask what happens after a test wins. If the answer involves “then we hand it off to your team,” ask how long that handoff typically takes in practice, not in theory.
Conversion rate is the easiest number to report and the easiest one to game. Free shipping above a low threshold will lift conversion rate. It might also tank your margin. Revenue per visitor, and ideally profit per visitor, is the number that actually reflects whether a test helped the business.

Ask what metric sits at the top of their reporting dashboard. If it’s conversion rate alone, dig further. If it’s revenue per visitor with conversion rate as a supporting metric, that’s a team thinking about your P&L, not just their case study slide.
Pricing across the industry is wide. Smaller retainers start around $3,000 to $5,000 a month for narrower scopes; more comprehensive programs, especially ones bundling research, design, and development, commonly run $10,000 a month and up, with enterprise engagements well into six figures annually. None of that tells you what you should pay. It tells you to be suspicious of any number that’s dramatically below the range for the scope being promised.

The more useful question isn’t “what’s your monthly fee,” it’s “what’s included in that fee, and what gets billed separately?” QA, development hours, design, and analytics setup are sometimes bundled and sometimes not. Get the full scope in writing before comparing two agencies on price alone. Two quotes that look ten thousand dollars apart might be identical once you account for what’s actually included.
An agency makes sense when you need testing volume and specialized skill (statistics, research methods, CRO-specific design patterns) faster than you can hire and train for internally. It also makes sense as a way to test the discipline itself before committing to a full-time hire or team.
In-house makes more sense once you have enough traffic and enough tests running that a dedicated hire pays for themselves in speed and institutional knowledge. A lot of the brands we work with land somewhere in between: an agency running the program with a client-side owner who feeds in business context and clears internal roadblocks. That hybrid model tends to outperform either pure version, because neither pure in-house nor pure outsourced gets both speed and business context by default.
A few patterns worth naming directly, because they show up often enough to be worth a specific warning:
An agency that promises a specific lift percentage before running discovery. Nobody knows your conversion rate will jump 20% before they’ve seen your data. That’s a sales number, not a forecast.
An agency that can’t show you a real test, win or loss, with an explanation of why it worked or didn’t. Case studies with only wins and no losses usually mean the losses aren’t being shown, not that they don’t happen. Every real testing program has losses. That’s what testing means.
An agency that wants a long contract with no early checkpoint. A short trial period, or at minimum a 90-day review point with a clear exit if velocity or quality isn’t there, protects you from a slow, expensive mismatch.
Once you’ve asked through the list above, you’ll usually know which agency actually thought about your business versus which one gave you the same pitch they give everyone. Weight the answers, not just the price: velocity, rigor, build capability, and reporting focus tell you more about six months from now than any homepage case study will.

If two agencies come out close, the tiebreaker worth trusting is how they handled your discovery questions. An agency that pushed back, asked follow-ups, and admitted where they’d need more information is usually the one that will tell you the truth six months into the engagement too, including the parts you don’t want to hear.
Retainers typically range from $3,000 to $5,000 a month for narrower scopes up to $10,000+ a month for comprehensive programs with research, design, and development bundled in. Enterprise engagements can run into six figures annually. Get the full scope in writing before comparing quotes.
An agency is usually the faster path to testing volume and specialized skill. In-house makes more sense once your traffic and test count justify a dedicated hire. Many brands land on a hybrid: an agency running the program alongside a client-side owner.
That depends heavily on how much prior testing data and traffic you’re starting with. An agency with a head start on your data can begin shipping tests within the first few weeks. One starting from zero needs time to build the roadmap first. Be wary of anyone promising fast, specific lifts before they’ve seen your funnel.

