Changes to how this tool calculates, newest first. Only the ones that move a recommendation or let you do something new: if your answer would be different today, it is on this page.
Most of what is here came from people outside the project reading the method and finding it wrong: a demand planner, a wholesaler and a private-label seller, each of whom used the tool and then argued with it. That is the intended way for this to work, and publishing the corrections is part of the deal. The reasoning behind every assumption is on where the curves come from.
New capability
The check can now grade your own forecast and your own order
Five optional columns put your forecast and the order you placed into the comparison. On our sample your side wins one of the two.
The backtest can now read five more columns, if your file has them: the forecast you were working from, the units you actually ordered on the first purchase order, and the MOQ, lead time and review time that applied when you placed it. Every one is optional and they work one at a time, so a file with two of them gets two more comparisons and a file with none is scored exactly as it was yesterday.
What they unlock is two things. Your forecast is scored for accuracy beside ours, on the same launches, with the same error and bias figures. And the order you placed is set against the order the tool would have recommended under your own minimum, so it is a like-for-like comparison rather than ours against a number that had a supplier constraint on it. Where the file records your lead time and review time, that launch is also scored over the window it really had to cover instead of one shared assumption.
The order comparison is in modelled cost under assumptions you set before the run and that are printed with the result. Nobody observed that money. It is what your stated horizon, salvage rate and leftover treatment imply the order would have cost, scored against the sales in your file, and the wording never says more than that. There is no return-on-investment figure here and there never will be.
The columns mean what you knew before the order went out, not what you learned later. A forecast revised in week six, or the lead time that turned out to be true rather than the one quoted, would put knowledge into the check that nobody had on the day, and the result would flatter or punish a decision using information that did not exist when it was made. The file-format page says this against every column, because it is the one way to get the block wrong.
The two questions are judged separately and neither is allowed to cover for the other. That is not a hedge: it is what the sample library does. On it, the nearest comparable launch predicted more accurately than SKUZero, 112 units of average error against 166. On the same run, SKUZero's order held to the recorded minimum came out at 1,358 euro of modelled cost against 4,858 for the orders that were actually placed. Better forecast, more expensive order. Both sentences are on the screen, the same size, and nothing adds them together into a winner.
One more refusal worth naming: your forecast is scored as a forecast and is never run through the order logic. Turning a single number into an order quantity needs an assumption about how uncertain it is, and your file does not contain one. We would have to invent it, and then we would be scoring our own invention while calling it yours.
What to do: Add whichever of the five columns you can get out of your system and re-upload. Each one is optional, none is needed for anything else in the tool, and the screen tells you before the run which comparisons your file can support.
New capability
You can now test the method on your own launches
Load your launch history and the tool will grade itself on it, including when a spreadsheet average does better.
Every launch in your library is now predicted using the rest of it. The one being predicted is held out completely: its own sales never reach its own forecast, not through the comparable search, not through the simulation, not through the averages. Then the prediction is compared with what that launch actually sold, and the same is done for the next one, and so on through the file.
The first thing it reports is how often the range held. If SKUZero said P10 to P90, your outcomes should land inside that band about eight times in ten, and the result tells you what it actually was. That is the figure worth having, because it says how far to trust the next range the tool hands you. On the sample library it comes back at 76% against a target of 80%, which means the band has been running slightly narrow.
Then it puts the built-in forecast next to the two things a buyer would otherwise do with the same file: take the category average, or take the nearest comparable launch. Those alternatives can win, and when they do the result says so in the same words and the same size it uses for the other answer. They do win, too. On the shipped sample the answer flips depending on one assumption: with leftover stock selling on at a carrying cost SKUZero comes out ahead, and with leftovers written down at the end of the cover period the nearest comparable launch comes out ahead. The screen says that as well.
What it is not: a comparison against your own forecasting process. It cannot be. What you forecast for a launch and what you actually ordered are not in the upload file, so nothing here can speak to how your current method performs. This tests SKUZero's built-in forecast against simple alternatives, and no verdict claims more than that. Comparing your own forecasts and orders is the next thing we build, and it needs a few more columns in the file you upload. Those stay in your browser like everything else you load here.
Two more limits, said here rather than left to be discovered. Launches that went out of stock during the window being scored are left out of the result and listed by name, because their recorded sales understate what demand was, and grading ourselves against a number we had corrected would be marking our own homework twice. And each launch is predicted from the rest of the library, including launches that happened after it, which is how you get a usable answer from a library of a few dozen rather than a noisy one from half of them.
The order decisions are scored under a horizon, a salvage rate and a leftover treatment that you set before the run, because the upload has none of them. They are shown before it starts, applied to every launch, and printed with the result. Purchase cost and selling price come from your file, and a launch without a cost gets a forecast score and no order score rather than an invented one.
All of it runs in your browser, on a file that never leaves the tab, and the rows are downloadable as a CSV so you can re-add them in Excel and disagree with us.
What to do: Open Company data, load your launch history, and press Backtest SKUZero on your launches. Eight scored launches is the fewest it will name a winner on.
Changes the numbers
Four sector patterns corrected for the European trade
HVAC, building materials, private-label ecommerce and the Q4 season all moved. If you sized a buy in one of those, run it again.
The sector patterns were built from published research on adjacent industries, because nobody has published week-by-week launch curves for products sold through distribution. That gets you the right shapes and the wrong calendar. A review against the working knowledge of European technical distribution found four places where the model described a business nobody in the trade would recognise.
HVAC ran its best months in May and June and its worst in January. That is a cooling-led year, and in north-west Europe the trade is heating-led: stock and installation build through September to November, the quiet stretch is late spring and high summer, and January is busy with replacement work. The peak has moved to the autumn. A launch month that used to look strong may now look weak, and the other way round.
Building materials had one trough, in winter. It has two now. Sites across continental Europe close for two to four weeks in July and August, the Belgian bouwvak, the French congé and the German Bauferien, and the wholesalers behind them go quiet with them. A product launched in early July meets that second trough inside its coverage horizon.
Private-label ecommerce ramped over two weeks. It now ramps over four, from a lower start. A new listing with no reviews converts poorly however much traffic is bought at it, and the curve only straightens once the first reviews land. If you are relaunching inside an established storefront, with the ratings already there, the old faster ramp was closer to right and this one will read slightly conservative.
The Q4 gifting season was shaped like a shop window, with December towering over November. Black Friday pulled that spending into November years ago, and delivery cut-offs mean the second half of December sells almost nothing that has to arrive in time. November and December are now close to level. The quarter's share of the year is unchanged: this moves demand within Q4 rather than adding any.
One more, for anyone who answers the sell-in question: the dip after an initial stocking order is now shallower and about three weeks longer. Branches order to a shelf quantity and then reorder as it empties, so the trough is a slope rather than a cliff. How deep it goes depends on how many branches you filled, which the tool does not ask, so treat it as a middle case.
To be exact about who did that review: an AI persona built to hold a European distributor's experience, not a person from the trade. Every change here is its judgment rather than a measurement, and the sign-off sheet labels each one that way. They replace our own reasoning, which was worse, and they are first in line to be replaced again by backtesting against real launches.
What to do: If you sized a first buy for HVAC, building materials, a private-label ecommerce listing, or with the Q4 season switched on, run it again before you order.
New capability
FirstBuy AI is now SKUZero
New name, same engine. No recommendation moves, and a launch library you saved here is still here.
The tool is called SKUZero from today, and the address is skuzero.com. The old name had two problems. It was generic enough that half a dozen unrelated products share it, and it ended in "AI", which is the opposite of what this thing is: a Monte Carlo simulation and a newsvendor calculation, both written out on screen, with a seed so the same inputs always give the same answer. A name that promised a machine learning black box was selling the one thing we are not.
Nothing about the method changed. Same sector profiles, same censoring correction, same economics, same numbers. If you sized a purchase order here last week, the quantity still stands and you do not need to run it again.
If you had uploaded a launch history, it is stored in your own browser and it survived the rename untouched: open the tool and it will still greet you with your library. Bookmarks to the old address keep working.
Changes the numbers
Your weekly figure now means a typical week
Sector estimates went up by about a fifth. A first buy this tool sized before today was too small.
When you type 40 units a week, you mean a normal week: some better, some worse, that one in the middle. The simulation was treating your number as an average instead, and because a few launches do very well and none can do worse than nothing, an average sits above the typical week. The effect was that entering 60 a week produced a middle case of about 50, while the panel underneath told you your estimate had not been adjusted.
Both halves of that were wrong: the number moved, and the page said it had not. Your figure is now the middle of the range, with half the simulated launches above it and half below. The standalone calculators already worked this way, so the two parts of the tool now agree with each other as well.
A demand planner reviewing the tool found this by entering a number and checking what came back out of it, which is a good way to test any forecasting tool, including this one.
What to do: If you sized a buy here before today and have not ordered yet, run it again. The new answer will be higher.
New capability
Uploads can now say how many days a week you were in stock
Better stock-out correction on your own launch history, if your ERP can export it.
A week where you sold 12 units while the item was only available for two days is not a weak week. It is a week that was cut short, and real demand was closer to 40. Until now the file format only accepted a yes or no for each week, so the tool had to guess the missing part from the weeks either side.
You can now add a column with the number of days in stock, 0 to 7, per week. Where you supply it, the correction uses your actual sales scaled to a full week rather than an average of the neighbours. The old yes or no column still works exactly as before, and both are optional.
This came out of published research on censored demand which finds that knowing when stock ran out recovers nearly as much as knowing true demand. It is the cheapest accuracy improvement available to anyone whose system records it.
What to do: Check whether your ERP can export days in stock per week. The file format page has the column name and an example.
Changes the numbers
Two corrections to how the range is calculated
Ranges from your own launch history are narrower and more honest. Recommendations for items you restock came down slightly.
The first: when building a demand range from your uploaded launches, the tool was counting the spread between those launches twice. Eight launches that varied by 45% came out as though they varied by 63%, so every range was wider than your own history supported, and a wider range means a more cautious order. Shape and size are now drawn separately and the spread matches what your launches actually show.
The second: for an item you keep restocking, leftover stock is not written off, it sells in the following weeks. The tool was working out how fast that happens using one middling rate for every scenario. But the launch that comes in under forecast is the same one whose leftovers clear slowest, and treating those as independent made over-ordering look cheaper than it is. Each simulated launch now clears its own leftovers at its own rate.
Both were found by an independent methodology review, not by us. They are recorded in full, with the numbers before and after, in the project's own notes.