Whitepapers & ebooks
How to Choose Inventory Optimization Software (and How to Test It Before You Sign)
Most software selections are decided in the demo, on the vendor's data. How to check your own data first, design a backtest that can fail, and when not to buy.
08. september 2026
13 min

Most inventory software decisions are made in the demo, and the demo runs on the vendor's data. The only test that tells you anything is a backtest on your own history, against your current process, with the success criteria fixed before it starts. This guide covers what to check before you talk to vendors, how to design a pilot that is allowed to fail, and when the right decision is not to buy at all.
Do you need inventory optimization software at all?
Not every company does, and the ones that buy too early pay twice: once for the licence and once for the project that stalls because the organisation was not ready for it.
Planning maturity matters more than company size, and it comes in three levels. At the first, ordering runs on spreadsheets and experience. At the second, the ERP does the arithmetic with min–max levels or reorder points that somebody set once and rarely revisits. At the third, a specialised tool forecasts demand per item and location, converts it into order proposals that respect pack sizes, minimum order quantities and lead times, and leaves people to handle exceptions. These are the same three steps the Veritico STOCK ROI calculator asks about first, because the gap between where you are and where you would land determines most of the benefit.
Skipping a level rarely works. A company ordering from spreadsheets today will not get value from a forecasting engine next quarter, because the data discipline the engine depends on does not exist yet. The typical purchase is the move from level two to level three, and it pays off when several things hold at once: active item–location combinations run into tens of thousands, a meaningful share of demand comes from promotions or seasonality, the long tail of slow and intermittent items is too large for one reorder rule, and more than a handful of people spend their days creating orders by hand. These are not hard thresholds; they are the pattern we see where the tool earns its keep.
We have told prospective clients to wait. The reason is usually not the software but that nobody in the company owns the master data, or that nobody on the shop floor trusts the stock records. The honest advice then is to fix the foundations and run the selection a year later, with a much better chance of success.
What to check in your own data before you write the RFP
Three things decide whether a pilot can produce a meaningful result: the accuracy of your stock records, the state of your master data, and the quality of your demand history. A tool repairs none of them. It converts them into order proposals faster.
Stock record accuracy comes first because every replenishment engine orders against the record, not against the shelf. The best‑known study of the problem, by DeHoratius and Raman, examined close to 370,000 inventory records across 37 stores of one retailer and found that 65 percent did not match the physical count. That was 2008, and the retailers we work with today are better than that, but rarely by as much as they assume. Where a category shows negative stock in the system, or cycle counts keep finding phantom inventory, no algorithm will lift availability. It will do the opposite: the system sees stock that does not exist, stops ordering, and the stockout persists until somebody corrects the record. This is why better forecasting alone does not fix stockouts.
Master data is the second check, and in most projects the larger piece of work. Lead times in the ERP that say seven days when deliveries take twelve. Pack sizes of six where the supplier now ships twelve. Minimum order quantities negotiated years ago and never updated. Each makes the engine compute correctly on a wrong input, and the result looks like a software failure when it is a data failure. Before the RFP, take one representative category and reconcile these fields against reality. The share you have to correct tells you how much preparation the pilot needs.
Demand history is the third. Twelve months is the minimum so the model sees a full seasonal cycle; eighteen to twenty‑four is better because it leaves room for a holdout period. More important than length is what the history is missing. Promotions never flagged as promotions look like demand spikes the model will try to forecast. Weeks of stockout look like weeks of zero demand, which teaches the model the item does not sell. Both are fixable, but they have to be found before the pilot, not explained away after it.
Which criteria actually separate vendors, and which don't
The name of the algorithm does not separate vendors. How the tool behaves at the edges of your data does.
The criteria that matter show up in the hard cases. How does the tool forecast an item that sells twice a month? How does it separate a promotional uplift from the baseline demand, and what does it do with the dip after the promotion? Does it treat a stockout week as zero demand or as censored demand to be estimated? How many proposals does a planner have to open, does the tool explain why a proposal looks the way it does, and does it record what the planner changed? Who owns master data once the tool is live, and how do changes flow between it and the ERP? How much of the time from signature to first system‑generated order is data preparation on your side? And how does the price behave when you open ten more stores?
The criteria that do not matter fill most vendor decks. “AI‑driven” describes every product on the market. The number of models in the library says nothing about how the tool chooses between them. Dashboards are pleasant, and pleasant is not a criterion when three planners will use the tool eight hours a day. The exception is worth stating: if forty store managers will use it rather than three planners, usability becomes decisive, because adoption is the constraint. The criteria have to follow the operating model you intend to run, not the one the vendor demonstrates. Total cost belongs in the comparison too; we have covered what build versus buy really costs separately, and the licence is rarely the largest line.
Design the pilot so it is allowed to fail
A pilot that cannot fail is a demo with your logo on it. The pilot that tells you something is a backtest on your own history, against your current process, with the success criteria fixed before anyone sees a result.
The mechanics are simple to describe and easy to skip. Take eighteen to twenty‑four months of history for the pilot scope. Cut off the last three to six months as a holdout the tool never sees. Let the tool generate order proposals week by week across the holdout, using only the data it would have had at each point in time. Then compare those proposals against what your company actually ordered in the same weeks, and against what actually sold. The comparison is not “tool versus naive forecast”; that is a test the tool always wins. It is “tool versus the people and rules you run today”, because that is the decision you are making.
Fix the KPIs before the backtest starts, in writing: availability or service level, days of stock or stock value, forecast bias by category, and the share of proposals a planner would have overridden. Vendors will offer to add metrics after the fact. Do not.
Choose the scope so that it can hurt: one high‑volume category with heavy promotional activity, plus a sample from the long tail of slow movers. Not the category with the cleanest data, and not the one the vendor suggests. Six to eight weeks including data preparation is realistic. When a pilot runs longer, the cause is almost always in the data, and that is itself a finding.
Vendor demos look better than pilots because demo data has no stockouts, no assortment changes, no unflagged promotions and no supplier who delivered late for three weeks in November. Your history has all of them. That is the point.
We insist on this design because we have been on the other side of it. Backtests on client data have come back weaker than the business case assumed, and the right response was to recalculate the case, narrow the scope or postpone, rather than explain the gap away. A pilot that can produce that answer is the only one worth running.
Reading the results and turning them into money
A good backtest shows potential, not a guarantee. The difference is made by things the backtest cannot contain: planners overriding proposals in live operation, suppliers delivering late, store staff who trust or distrust the numbers.
Two outcomes are typical. The tool matches your current availability with noticeably less stock, or it lifts availability at roughly the same stock level. Both are worth money through different lines: the first releases cash from inventory and cuts holding cost, the second recovers margin on sales lost to empty shelves. The third line is people. If the backtest shows proposals planners would rarely need to touch, the time spent creating orders today becomes time spent on exceptions and suppliers, which is where planners add value.
This is the moment to put your own numbers into the ROI calculator. It takes five inputs: segment, planning maturity, annual revenue, inventory value and the number of people working on orders. It returns an annual benefit estimate, payback period, released cash and a breakdown of where the benefit comes from. It is an estimate, not a business case, but it tells you within minutes whether the backtest result is worth turning into one.
For scale, the published outcomes from our own projects sit in this range. At Dr.Max, a network of 490 pharmacies, the deployment of Veritico STOCK is associated with revenue up 5 percent, availability up 4 percent and two hours a day saved per pharmacy on ordering. At SIKO, shelf availability moved from 97 to 98.5 percent while inventory value fell by 1.65 million euros year on year. At Košík, availability reached 97 percent while waste fell by 75 percent. These are different businesses with different starting points, which is exactly why a backtest on your own data beats reading anyone's case studies, including ours.
Contract and rollout traps that show up after the signature
Most disputes after signature are about scope, not software. Three questions settle most of them, and they belong in the contract, not in the kickoff. Who cleans the master data, and who keeps it clean after go‑live? Who builds and maintains the integrations to the ERP and the warehouse system, and who pays when the ERP is upgraded? And what leaves with you if you leave: demand history, tuned parameters and exception rules, in a format another system can read?
Then check two numbers against your growth plan. A price per item or per location looks cheap at signature and grows with every store opening; ask for the figure at twice your current footprint. A go‑live in eight weeks is possible only when data preparation is excluded from the eight weeks. Our own implementations of Veritico STOCK take around four months including that preparation, and we would rather say so than discover it together in month three. Finally, read what the service level agreement covers: uptime is easy to guarantee, quality of proposals is what you are buying.
You do not have to run the selection alone or take the vendor's word for any of this. We regularly act on the client's side in selecting and implementing planning and supply systems, including systems that are not ours, and the questions above are the ones we ask on the client's behalf.
When the right answer is not to buy
Sometimes the backtest gives a clear result, and the result is that software is not your problem.
Three signals point that way. Stock record accuracy below the level where any engine can work, typically when a sample count in the pilot category shows a large share of records off by more than a pack. Master data with no owner: if nobody is responsible for lead times and pack sizes today, the tool will run on stale inputs within three months of go‑live, and the project will be blamed for it. And an override rate that stays high for reasons the tool cannot see, such as informal supplier arrangements or a planning process that happens in meetings and never reaches the system. As a rule of thumb from our projects, when planners would have changed roughly a third of the proposals or more, the process needs fixing before the tool does.
In each case the tool moves the problem rather than solving it. The better sequence is process and data first, then a second pilot a few months later on the same scope and the same KPIs, so the two results can be compared. Companies that take this route usually end up buying, and buying with a business case they can defend.
Frequently asked questions
How long should an inventory software pilot take? Six to eight weeks including data preparation is realistic for a backtest on one or two categories. If it runs much longer, the delay is almost always in data extraction and cleaning rather than in the tool, which is useful information in itself.
Should we pilot two vendors at the same time? Yes, if both get the same scope, the same history and the same holdout period. Two pilots on different categories or time windows cannot be compared, and you end up choosing on impressions after all.
Can we run a backtest without handing over sensitive data? Largely, yes. Item codes and locations can be anonymised and prices indexed. What has to stay intact is demand history at item–location–week level, stock levels, promotion flags and supplier parameters, because those are what the tool is being tested on.
What if our promotion history was never flagged? It can be partly reconstructed from price and volume changes, but reconstruction has limits and the vendor should tell you how they handled it. Expect promotion‑heavy categories to test weaker than they would with clean flags, and read the long‑tail results as the more reliable part of the pilot.
Run your own numbers first. The ROI calculator takes five inputs and returns an annual benefit estimate, payback and released cash. If the number is worth a conversation, talk to an expert and we will design the backtest with you.
More supply chain insights

Supply chain glossary
Dead Stock
Inventory that has stopped selling and will not move at its current price, place or form – and where the line between dead stock and slow movers is drawn.
11. september 2026
3 min

Supply chain glossary
Cycle Stock
The part of inventory that covers demand between two deliveries — how order quantity and order period set it, and why it rarely equals Q/2.
10. september 2026
3 min

Supply chain glossary
XYZ Analysis
Classification of items by demand variability — the coefficient of variation behind it, the ABC/XYZ matrix and what promotions and stockouts do to the result.
09. september 2026
3 min