Your “We Should Test That” List Is Worth 5 to 15% of Revenue. Capacity Is Not Why You Never Get To It

Published:
July 22, 2026

The revenue in your “we should test that” list is real, but capacity is not the constraint. Most Shopify brands between $500K and $10M never ship those tests because no item on the list has an owner, a decision date, or a rule for when it dies.

Quick Decision Framework

  • Who This Is For: Shopify founders and operators between $500K and $10M in annual revenue who keep a running list of untested ideas (bundles, pricing, shipping thresholds, win backs) that has not gotten shorter in six months.
  • Skip If: You are under $250K with fewer than roughly 2,000 lifetime orders, because most of these tests need order volume to read. Also skip if you already run a weekly test calendar where every live test has a named owner and a decision date. You have solved this.
  • Key Benefit: A one hour triage that cuts a twenty item backlog to three funded tests, plus the four tests that pay first for most brands at this stage and the numbers that decide each one.
  • What You’ll Need: Twelve months of Shopify order history, your current AOV and free shipping threshold, one hour of protected calendar time, and one person willing to have their name attached to a decision.
  • Time to Complete: 12 minutes to read, 60 minutes to run the triage, 30 to 45 days before your first test returns a decision.

A backlog with no owner is not a backlog. It is a list of things you have already decided not to do, written in a format that lets you avoid noticing.

What You’ll Learn

  • Why your test list keeps growing even after you add headcount, an agency, or an AI agent to work it
  • How to cut twenty backlog items down to three funded tests using reversibility and reach, in under an hour
  • What the four highest frequency tests are for $500K to $10M Shopify brands, and the specific number that decides each one
  • When to kill a test rather than extend it, and the rule that stops zombie experiments from consuming your calendar
  • Where automation genuinely removes work, and where it quietly makes the backlog problem worse

On the eCommerce Fastlane podcast this spring, Apoorva Modi put a number on something most operators feel but never measure: 5 to 15% of annual revenue sitting unclaimed inside a store’s existing order data. On a $10M brand that is $500,000 to $1.5M. Treat it as a vendor estimate, because it is one. Modi builds the software that finds it, and I want you reading that number with the same skepticism you would bring to any figure supplied by the person selling the solution.

The number is arguable. The list is not. Every operator I talk to above $500K has one, and it reads roughly the same in every business: the bundle nobody built, the pricing nobody refreshed, the free shipping threshold set two years ago and untouched since, the lapsed VIP segment somebody exported once and never emailed. Ask when that list was last shorter than it is today and you usually get a long pause.

The standard explanation is capacity, and the standard fix follows from it. Hire someone. Retain an agency. Install an agent. I have watched that sequence fail often enough to stop recommending it as a first move. Capacity is not the thing that is broken. What is broken is that nothing on the list has a name attached to it, a date on which it gets decided, or a rule for when it dies. Fix that first, and a meaningful number of brands discover they had the capacity the whole time.

What Is Actually Sitting on Your “We Should Test That” List

The typical Shopify test backlog between $500K and $10M holds twelve to thirty items, and four of them carry most of the recoverable revenue: the free shipping threshold, a bundle built from real co-purchase data, compare-at pricing across the catalog, and a win back aimed at lapsed high value customers. That distribution is illustrative rather than surveyed, drawn from merchant conversations rather than a dataset, but the shape is consistent enough that I now say it out loud on calls and watch people nod.

The trouble starts with everything else on the list. Backlogs fill with items written at the same level of abstraction, which makes them look like the same size of decision. “Test a subscription offer” sits on the line directly above “raise threshold from $50 to $65” in the same font, and the two read as equivalent tasks. One is a settings change you can reverse in ninety seconds. The other is a quarter of product, billing, and support work that permanently changes your P&L structure. Written identically, they become unrankable, and an unrankable list is a list nobody starts.

Before you do anything else, open the list and write the true size of each item beside it in hours. Not the size of the idea, the size of the work. A Shopify admin bulk edit is twenty minutes. A price test through a tool like Intelligems is a two hour setup and a three week wait. A subscription program is a quarter. Most operators are shocked to find that six or seven items on a twenty item list are under an hour of actual work, and that those items have been sitting there since last summer anyway. That is the first real signal that capacity was never the problem.

Why the List Never Shrinks, and Why Capacity Is the Wrong Diagnosis

Your list never shrinks because every item on it is optional, and optional work loses to operational work every single week, in every business, regardless of how much labor you throw at it. A stockout is not optional. A chargeback is not optional. A supplier who missed a container is not optional. The bundle you have been meaning to build has no external party generating pressure, so it survives another Monday.

The clearest evidence that insight is not the constraint comes from the people who sell insight. On episode 466, Thomas Gleeson of StoreHero described what his team sees in the first ten minutes of an onboarding call: usually three to ten straightforward margin improvements, things like mispriced international shipping, unaccounted duties, and discount codes that have been quietly live for weeks. Then he said the part that matters most for this article. The onus is on the founder to implement them.

Sit with that. The diagnosis takes ten minutes and costs almost nothing. The implementation still does not happen. If the constraint were genuinely a shortage of information or a shortage of hands, a free ten minute diagnosis followed by three obvious fixes would close the loop. It routinely does not. What is missing is a person who has agreed, in front of other people, that this specific change is theirs and that it gets decided by a specific date.

This is the same failure mode I have watched sink brands at the $500K to $2M stage for years, wearing a different costume. Premature complexity does not usually arrive as one bad decision. It arrives as thirty reasonable ideas, none of which are refused, all of which stay technically alive, consuming attention as a background process. The list is not a plan. The list is a record of decisions you have deferred, and deferral has a cost you never see on a P&L.

The Three Test Rule: Cut the List Before You Try to Fund It

Cut your backlog to three tests before you add any capacity to it, because an untriaged list will absorb every hour, every dollar, and every agent you point at it and still not get shorter. This is the whole intervention, and it takes about an hour.

Put the list somewhere everyone can see it. For each item, write three things: the person who owns it, the date the decision gets made, and the condition under which you kill it. If you cannot supply all three, the item does not survive. Do not soften this by creating a “later” bucket, because a later bucket is the same list with a new heading. Delete them. You can always rewrite an idea. What you cannot recover is the attention that twenty half alive ideas consume every time someone opens the document.

Three is the number for a reason. At $2M to $10M, three concurrent tests is roughly what one operator can genuinely hold while also doing their actual job, and it is few enough that you can name each one from memory in a Monday meeting. Below $2M the number is one. I know that sounds punishingly small to a founder with a twenty item list and real ambition, but one test that finishes and produces a decision beats six that all sit at 60% complete while you fly to a trade show.

The kill criterion is the piece people skip, and it is the piece that makes the system work. Write it before the test launches, when you are still honest. “If the threshold change has not lifted AOV by at least 4% after 30 days, we revert and we do not revisit it this year.” Without that sentence, losing tests do not end. They get extended, then reinterpreted, then quietly folded into someone’s ongoing responsibilities, and your three test slots are full of ghosts by March.

How to Rank What Stays: Reversibility, Reach, and Read Time

Rank the surviving tests on three axes: how fast you can reverse the change, how much of your order volume it touches, and how long it takes before the result is readable. Reversibility should carry the most weight at every stage below $10M, because a cheap reversal is what lets you act without a committee.

Reach is the axis operators get wrong most often. A test that touches 4% of orders needs an enormous effect size to produce a readable result, which is why so many product page experiments at $500K stores end in a shrug. A free shipping threshold touches every cart. Compare-at pricing touches every product view. Start where the traffic already is, not where the idea is most interesting.

Test Type
Reversibility
Time to a Read
Free shipping threshold
Instant, one settings change
21 to 30 days
New bundle
Instant, unpublish the product
30 to 45 days
Compare-at pricing
Instant, bulk edit reverses
14 to 21 days
Lapsed customer win back
None, the email already sent
7 to 14 days
Sitewide price change
Slow, trust outlasts the revert
60 to 90 days

Read the bottom row carefully, because it is the one that should stay on the shelf. A sitewide price change is technically a settings change and emotionally a one way door. Customers who noticed the increase do not un-notice it when you revert. Rank it accordingly, and run the reversible tests first while you build the confidence to touch it properly.

The Four Tests That Pay First for Most Shopify Brands

For most Shopify brands between $500K and $10M, the four tests worth funding first are the free shipping threshold, one bundle built from actual co-purchase data, compare-at pricing across the catalog, and a win back to lapsed high value customers. All four are reversible, all four touch meaningful order volume, and all four can be decided inside 45 days.

Start with the threshold, because the underlying behavior is the best documented thing in ecommerce. Baymard Institute’s checkout research finds that 39% of shoppers abandon a checkout because of extra charges like shipping, taxes, and fees, the single largest fixable cause they measure. If your threshold sits below your AOV, you are absorbing shipping cost on orders that were going to clear anyway and buying nothing with it. Set it modestly above your average order value, not your best order, and give it 30 days. I wrote a mechanical walkthrough of the threshold math a while back that still holds up on the arithmetic.

The bundle test fails when brands invent the bundle in a meeting. Pull your last twelve months of orders, find the two products that already appear together most often, and package exactly that. Rebuy and similar merchandising apps will surface co-purchase pairs for you, though a Shopify order export and a pivot table gets you there for free at this stage. Compare-at pricing is the least glamorous of the four and often the fastest. In a scan I walked through in my Revenue Agent review, a ten year old chocolate brand turned out to have 89% of its catalog with no compare-at price set at all, part of roughly $38,000 in annual opportunity surfaced across eleven actions.

The win back is the one to run first if you want an early win, because Klaviyo will hold the segment and the result reads in a week. One caution on the offer inside it. Steeper discounts do not reliably win: on episode 477, Claspo’s data across billions of popup views showed what a 10% offer does against a 15% offer, with the smaller offer producing better revenue per visitor because the larger one pulls in discount hunters. Test the offer depth, not just the send.

What Tools Actually Fix, and What They Quietly Make Worse

Tools compress the distance between a question and an answer, and they do not create the willingness to decide, which is the constraint we started with. That distinction should drive your entire buying decision in this category, and it is the reason I want to be specific about who does what.

Intelligems handles price, shipping threshold, and offer testing with the statistical rigor most stores cannot build themselves. Shoplift covers theme and page level A/B testing on Shopify. Rebuy handles bundles and upsells as a merchandising layer. Klaviyo owns the retention flows. StoreHero answers the margin question underneath all of it, which matters because a test that lifts AOV while destroying contribution margin is a loss you will book as a win. Revenue Agent sits in a newer category, scanning order data across eight levers and executing approved changes directly in the Shopify admin. In the interest of full disclosure: I have published a positive review of Revenue Agent and I have an affiliate relationship with them, which is exactly why I am telling you the next part.

Autonomous execution raises throughput on your backlog. It does not triage it. Point an agent at an untriaged twenty item list and you get twenty changes shipped faster, some of which conflict, none of which anyone owns, and a store that is harder to reason about in ninety days than it was before. Chad Rubin made the adjacent point on episode 470 when he described earning trust before you hand an agent the keys: every account starts in approval mode, and autonomy expands only as the work proves itself.

The honest sequence is triage, then tools. Cut to three, name the owners, set the kill criteria, and run one cycle by hand. If after that cycle your genuine complaint is “we know exactly what to do and we physically cannot execute it fast enough,” buy something. That complaint is rare, and it is the only version of the capacity argument I have ever seen hold up.

What This Looks Like at Your Stage

At $500K to $2M you should run one test at a time; at $2M to $10M, three concurrently; above $10M the constraint stops being capacity and becomes coordination between people who each own a different lever. The mechanics of the triage do not change across stages, but who signs their name to each item does, and that is the part that determines whether any of it survives contact with a busy week.

Stage
Concurrent Tests
Who Owns the Decision
$500K to $2M
One at a time
Founder, no delegation available
$2M to $10M
Three maximum
Named operator, weekly review
$10M and above
Six to eight
Function leads, monthly arbitration

If you are between $500K and $2M, the hardest sentence in this article is that the owner is you. There is no operator to delegate to, and pretending otherwise is how the list grew to twenty items in the first place. Pick one, put a date on it, and let the other nineteen go until that one finishes.

Between $2M and $10M, the failure is different. You have people, so the list gets distributed, and distributed without named ownership means shared, and shared means nobody’s. Name one person per test in writing. Not a team. A person.

Above $10M you will have several people running levers in parallel, and the new problem is that they will collide: a pricing test and a promotion test running on overlapping traffic in the same three weeks produce a result nobody can interpret. That is what the monthly arbitration exists for. The list still needs to be cut. It just gets cut by more people, with a calendar that prevents two of them from stepping on the same cohort. If your list has not gotten shorter in six months, none of this is a tooling problem, and no amount of software is going to make the decision for you.

Frequently Asked Questions

How do I decide which Shopify test to run first?

Run the most reversible test that touches the largest share of your orders. For most Shopify brands that is the free shipping threshold, because it can be changed in the admin in under a minute, reverted just as fast, and it affects every cart rather than a slice of traffic. Rank your remaining candidates on three things: how quickly you can undo the change, what percentage of your order volume it touches, and how many days before the result is readable. Anything that is slow to reverse, such as a sitewide price increase, should wait until you have run two or three cheap tests and built confidence in your own read of the data.

What should my free shipping threshold be on Shopify?

Set your free shipping threshold modestly above your average order value, not below it, and revisit it at least once a year. A threshold sitting under your AOV subsidizes shipping on orders that were already going to clear, which buys you nothing. Setting it slightly above nudges customers to add one more item to qualify. The common failure is not the starting number, it is that the number never gets revisited: product costs rise, AOV drifts, and a threshold set two years ago silently becomes wrong. Baymard Institute’s research finds extra charges at checkout are the largest fixable cause of cart abandonment, so make sure the threshold and remaining spend are visible in the cart, not revealed at the final step.

How long should I run a Shopify A/B test before making a decision?

Set the decision date before you launch, and honor it. For most Shopify brands under $2M, running to conventional statistical significance on a small effect is not realistic, because the traffic is not there. Rather than pretending otherwise, pick a minimum effect that would actually change your behavior, such as a 4% AOV lift, and a window long enough to cover at least two full purchase cycles, typically 21 to 30 days. If the effect you needed has not appeared by that date, revert and move on. Tests that get extended “just another week” are how three test slots fill up with experiments nobody is willing to end.

Can AI tools actually execute revenue tests in my Shopify store?

Yes, several tools now execute approved changes directly in the Shopify admin rather than only recommending them, including building bundles, updating shipping thresholds, and creating customer segments. That capability is real and it is new. What it does not do is decide which tests deserve to run. Pointing autonomous execution at an untriaged backlog ships more changes faster without resolving the underlying problem, and it produces a store that is harder to reason about because multiple overlapping changes went live with no single owner. Triage first, run one cycle manually, and only then evaluate whether execution speed is genuinely your constraint. For most brands under $5M, it is not.

Why does my ecommerce test backlog keep growing?

Your backlog grows because every item on it is optional work, and optional work loses to operational work every week. Stockouts, chargebacks, and supplier problems generate external pressure. The bundle you meant to build generates none, so it survives indefinitely. The second cause is that backlog items get written at the same level of abstraction, so a ninety second settings change and a quarter long product project appear on the list as equivalent tasks, which makes the list impossible to rank and therefore impossible to start. The fix is to attach three things to every item: an owner by name, a decision date, and a condition under which you kill it. Items that cannot carry all three get deleted.

FIND US ONLINE

WEEKLY DTC INSIGHTS

TRUSTED BY THOUSANDS

TRUSTED PARTNERS

Shopify Growth Strategies for DTC Brands | Steve Hutt | Former Shopify Merchant Success Manager | 460+ Podcast Episodes | 50K Monthly Downloads

Choose a language