Most new-flavor decisions in scoop shops come down to gut feeling and a couple of loud customers at the counter. Someone tries the honey-lavender, says "oh my god this is amazing," and suddenly you've committed a full production run to a flavor that moves four scoops a day and freezer-burns before it clears. The tub sits. You mark it down or toss it. The "win" becomes a quiet 30–40% loss on the batch.
The problem isn't that you're testing flavors — it's that you're doing it with no protocol, no clean comparison, and no rule that tells you when to pull the plug. A proper flavor test doesn't need to be complicated, but it does need structure. Otherwise you're just guessing with a nicer story attached.
What follows is the exact setup we'd use to test a new flavor against a known baseline, log the data without slowing the line down, and land a stop-or-scale decision inside two weeks with minimal wasted product.
Why casual flavor testing burns so much product
The mechanics of gut-feel testing are pretty straightforward. You drop a new flavor in the case, it competes for the same limited display slots as your proven sellers, and it cannibalizes attention instead of adding to it. A customer who would've bought your salted caramel now "tries the new thing" — so total ticket doesn't grow, you just shuffled the same sale onto a riskier flavor with no track record.
Then there's the batch-size trap. Minimum viable batches for most shops run one to two full tubs. If you make a full 3-gallon tub to "see how it does," you've already committed roughly $18–$30 in raw materials plus labor before a single scoop moves. If the flavor underperforms, you're not evaluating it anymore — you're managing dead stock and hoping it clears before quality drops.
The last issue is measurement. Without a holdout or a baseline, you literally cannot tell whether the new flavor sold "well." Fifteen scoops a day sounds fine until you realize your average flavor does twenty-two, and this one is eating a premium display slot to underperform. You need something to compare against — same shop, same weather, same foot traffic — otherwise the numbers are meaningless.
The core idea: small batch, fixed window, clean comparison
The whole protocol rests on three things: keep the test batch small enough that failure is cheap, hold the test window fixed so weather and traffic roughly average out, and always run the new flavor alongside a known baseline at the same time.
Keep your ice cream shop running smoothly.
Cremyly helps you manage every order, stock level, and staff shift with precision and ease.
- Real-time inventory tracking
- Order management dashboard
- Staff scheduling & shift coordination
No credit card required
You're not trying to prove the flavor is good in some absolute sense. You're answering one question: does this flavor earn its display slot better than what it's replacing? That's the only decision that matters, because case space is your real constraint — not your imagination.
The baseline should be a mid-tier proven seller, not your best one. Testing against your #1 flavor sets an impossible bar. Testing against a dog makes everything look like a winner. Pick something in the middle of your sales pack. That's your honest benchmark.
Step-by-step protocol
1. Set your batch size before anything else
Make the smallest batch you can actually produce and still display properly. For most shops that's a single half-tub or one tub max. The goal is to cap your downside. If the whole thing gets tossed, you want that to sting for a minute, not wreck the week.
A single-tub test at roughly $20–$28 in ingredients is a number you can afford to lose. A four-tub "soft launch" is not a test — it's a commitment you're pretending is a test.
2. Pick your baseline and your placement
Choose one mid-tier baseline flavor. Both the test flavor and the baseline need to sit in comparable case positions — same eye level, same distance from the register, same rough traffic flow. Placement massively distorts flavor sales. A flavor in the leftmost slot where people start reading the case will outsell an identical flavor buried in the back-right corner by a noticeable margin.
If you can, swap the two flavors' positions at the midpoint of the test — say day 4 of a 7-day window. That cancels out placement bias, which is the single most common thing that fakes out informal tests.
3. Set a fixed window: 7 days minimum, 14 preferred
Run the test for a full week at minimum so you capture both weekday and weekend traffic. Two weeks is better because it smooths out one weird rainy Saturday or a random slow Tuesday. Do not stop early because it "looks good" on day 2 — early spikes are almost always novelty, not real demand.
Novelty inflation is real. New flavors get a curiosity bump in the first 48 hours that has nothing to do with whether people actually want it again. The back half of your window is where the truth shows up.
4. Log the data without slowing the line
You need four numbers per flavor per day, and nothing more:
-
Scoops sold (pull from POS if your flavors are tagged)
-
Waste/melt/reject for that flavor
-
Approximate weather (hot / mild / rainy — a one-word tag)
-
Total shop traffic or total scoops for the day
That last one matters. If Wednesday was dead across the board, you don't penalize the test flavor for low absolute numbers — you look at its share of scoops that day. Consistent POS tagging is what makes this painless; if your flavors aren't tagged cleanly, fix that first. It's the backbone of any real comparison and ties directly into a broader operational data taxonomy for scoop shops — clean tags in, clean decisions out.
Tag flavors in the POS before starting the test so daily logging is painless and accurate.
5. Track a "holdout" perspective on demand
You don't need a true statistical holdout to make this work — you need a reference point. Your baseline flavor is your holdout. It tells you what normal demand looks like under today's conditions. When a rainy Tuesday tanks total sales, both flavors drop together, so the ratio between them stays meaningful even when the raw counts don't.
If your shop already runs a demand-forecasting framework tied to weather, overlay the test window onto your expected daily volume. That way you know whether a "slow" test day was slow because of the flavor or just because it was a slow day.
Here's a quick visual of the workflow.
This visual walks through batch prep, placement, daily logging, mid-test swap, and the final stop-or-scale decision.
A simple daily log template
Here's the whole thing on one sheet. A clipboard by the dipping cabinet or a shared note works fine.
| Day | Weather | Total scoops (shop) | Test flavor scoops | Baseline scoops | Test waste | Test share of day |
|---|---|---|---|---|---|---|
| Mon | Hot | 210 | 24 | 27 | 0 | 11.4% |
| Tue | Mild | 165 | 18 | 25 | 1 scoop | 10.9% |
| Wed | Rainy | 120 | 12 | 19 | 0 | 10.0% |
| Thu | Hot | 230 | 22 | 30 | 0 | 9.6% |
| Fri | Hot | 260 | 26 | 34 | 2 scoops | 10.0% |
| Sat | Hot | 340 | 31 | 45 | 0 | 9.1% |
| Sun | Mild | 290 | 25 | 38 | 1 scoop | 8.6% |
The "test share of day" column is the one to watch. Notice the trend above: starts at 11.4% on Monday — novelty bump — and drifts toward 9% by the weekend. That's the real story. Absolute scoop counts alone would've hidden that completely.
The stop / scale decision rules
After your 7–14 day window, run the flavor through fixed rules. Decide on these thresholds before you start so you don't rationalize a favorite flavor into staying on.
-
Scale it if the test flavor hits 90%+ of baseline and the trend is flat or rising in the back half. It's earning its slot and demand is holding past the novelty stage.
-
Extend the test if it lands between 75–90% of baseline with a lot of day-to-day noise. Run another 7 days, possibly in a different case position, before deciding.
-
Kill it if it's below 75% of baseline, or if the back-half trend is clearly falling off after the novelty bump. It doesn't justify displacing a proven mid-tier flavor.
Two override rules regardless of scoop counts:
-
Waste override — if reject/melt losses on the test flavor run noticeably higher than your other flavors (doesn't hold shape, melts too fast, freezer-burns quick), lean toward killing it even if sales are okay. High-maintenance flavors quietly eat margin. Managing that ties back to keeping freeze/thaw windows and portion yields tight.
-
Margin override — if the flavor costs meaningfully more to produce (expensive mix-ins, imported ingredients), require it to clear the scale threshold, not just squeak by. A pricier flavor at 90% of baseline is a worse deal than a cheap one at 90%.
These overrides matter more than most owners expect. A flavor that's technically "passing" the scoop threshold but melting faster than anything else in your case is still a problem — you just can't see it until you're writing off melt losses at the end of the week.
A real scenario
A two-location scoop shop in a beach town wanted to add a pistachio-rosewater flavor a customer kept asking for. Instead of committing a full production run, they tested a single half-tub against their "brown butter pecan" baseline over 11 days, swapping case positions on day 6.
The pistachio-rose opened strong — first three days it edged out the baseline slightly, and the owner was ready to lock it in. But the log told a different story once novelty wore off. Over the full window it landed at about 71% of baseline scoops, and the back-half trend was clearly dropping. Waste was also higher — the flavor got requested, licked once, and half-abandoned more often than usual, which showed up as extra melt in the case.
Total downside from the test: roughly $24 in ingredients and maybe an hour of labor. A typical four-tub soft launch would've cost closer to $90–$120 in product plus the markdown drag of clearing slow tubs before quality dropped. They killed the flavor on the rules, kept the beloved baseline, and tested a peach variety the following week that came in at 94% and got scaled. Same case slot, very different outcome — and they only knew because they measured both the same way.
When this protocol makes sense — and when it doesn't
It makes sense when you have clean flavor-level POS data (or are willing to tally scoops by hand), a reasonably stable traffic pattern, and you're testing flavors that compete for existing case slots. That covers most shops most of the time.
It breaks down when you're launching a flavor for a specific event or holiday where novelty and one-time volume are literally the whole point. A Valentine's rose flavor doesn't need to beat baseline over 14 days — it needs to sell hard for a weekend. Different game, different math.
Skip the formal protocol entirely if you're a brand-new shop still figuring out your core menu. Below a certain daily volume the numbers are too noisy to be meaningful. Get your baseline sellers established first, then start running controlled tests against them.
Common mistakes that fake out the results
These common mistakes will often mislead you if you don't watch for them.
-
Stopping early on a novelty spike. The first 48 hours lie. Hold the window.
-
Testing against your best flavor. Nothing beats your #1, so everything looks like a failure. Use a mid-tier baseline.
-
Ignoring placement. A great flavor in a bad slot loses to a mediocre flavor in a prime slot. Swap positions mid-test.
-
Comparing raw counts across different-traffic days. Use share-of-day, not absolute scoops, when weather swings.
-
Testing two new flavors at once. They compete with each other and you can't isolate what actually worked. One test at a time.
Avoid these pitfalls and your tests will give you a meaningful signal instead of noise.
Bringing it together
Casual flavor testing keeps costing shops money not because the flavors are wrong, but because there's no fixed method. When every test uses a different batch size, no baseline, and no stop rule, you can't compare anything — and you end up making expensive calls based on curiosity spikes and counter-side compliments.
Set the pieces once — small batch, mid-tier baseline, fixed 7–14 day window, four logged numbers a day, position swap at the midpoint, pre-set stop/scale thresholds — and every future flavor test becomes a cheap, repeatable decision instead of a gamble. You'll kill the losers before they cost you a full production run, and you'll scale the winners with actual evidence behind them, not just a good feeling on a busy Saturday.
Set the pieces once — small batch, mid-tier baseline, fixed 7–14 day window, four logged numbers a day, position swap at the midpoint, pre-set stop/scale thresholds — and every future flavor test becomes a cheap, repeatable decision instead of a gamble. You'll kill the losers before they cost you a full production run, and you'll scale the winners with actual evidence behind them, not just a good feeling on a busy Saturday.
Ready to scoop up efficiency and grow your shop?
Join hundreds of ice cream shops using Cremyly to boost productivity, reduce waste, and delight customers with faster service.