Skip to main content
38 min readstore clusteringstore clustering assortment planning

Store Clustering and Localization by Vertical

Store clustering groups selling locations that behave alike so one plan can serve many doors. How the clustering dimension, the cluster count and the re-clustering cadence change across ten retail verticals.

Store clustering is the practice of grouping selling locations that respond to the same assortment in the same way, so that one plan can be built for the group rather than one plan for every door — and localization is what the clusters are for: matching option list, depth, size or shade distribution and delivery timing to the demand a group of doors actually has. A cluster is a planning unit rather than a description of a store. Doors belong together when giving them the same assortment would produce the same outcome, which is a different question from whether they look alike, sit in the same region or turn over the same revenue.

It belongs to the by-vertical series alongside the merchandise hierarchy, which defines the level each category plans at; assortment planning by vertical, which builds the option list a cluster receives; allocation and replenishment by vertical, which places units inside the cluster; demand forecasting by vertical, which sizes the season the clusters will carry; and the planning calendar by vertical, which dates the review. Store clustering for apparel brands is the single-vertical deep dive this guide generalizes, and localized assortment planning covers the assortment work that follows a cluster decision.

What store clustering is

Three things separate a cluster from a label. A cluster is defined by a decision that differs — a different option list, a different depth profile, a different size or shade curve, a different flow date. If nothing in the plan would change between two groups, they are the same group with two names. A cluster is auditable, which means the rule that put each door into it is written down and can be re-run: "these doors carry the wide fit" is a rule, "these are the good doors" is not. And a cluster is stable across the season it plans, because a group whose membership moves mid-season cannot be scored against the plan that was built for it.

The alternative approaches fail in opposite directions. One assortment everywhere is cheap to plan and guarantees the same mistake in every door: the sizes, shades or price bands a trade area does not buy still arrive, and the ones it does buy run out. One plan per door is accurate in principle and unbuildable in practice, because the planning work scales with the door count while the evidence under each plan gets thinner, so a hundred door-level plans are a hundred forecasts built on a hundredth of the data. Clustering exists to buy most of the accuracy of door-level planning at a fraction of its cost, and the whole craft is in deciding where on that trade-off a given fleet sits.

Clustering, grading and single-door localization

Three practices get called clustering and only one of them is. They are complements rather than competitors, and a fleet running all three is normal.

Store clusteringStore gradingSingle-door localization
Groups onThe pattern of demand across the assortmentOne ranked axis, usually volumeNothing; the door is the unit
ProducesUnordered groupsAn ordered ladder, A to CA plan per door
DecidesWhich options, sizes, shades and flow datesHow deep, and who receives newness firstEverything, per door
Evidence neededEnough door-weeks per cluster to read a patternSales history aloneMore history than a single door usually has
Fails whenGroups are built on volume or geographyIt is used as a substitute for clusteringThe fleet is large enough that the work never finishes

The substitution that costs the most is using grading where clustering is needed. A grade-only fleet sends every small door a scaled-down copy of the biggest door's assortment, so a small door with a distinctive demand pattern is permanently served a shrunken version of somebody else's. That is how a door that sells wide fits never receives them, and how a beauty door with deep demand at the ends of the shade ladder is sent a shallower version of the middle.

Single-door localization still has a place, and it is a narrow one: flagship doors that carry an exclusive, doors whose trade area is genuinely unlike anything else in the fleet, and the handful of accounts large enough to justify their own plan. Everywhere else it is a promise that quietly turns into "whatever the allocator did last week".

Choosing the clustering dimension

The dimension is chosen from the decision, not from the data

The order of work is the part most often reversed. The temptation is to run a clustering algorithm across every attribute available and then find out what the groups mean. That produces groups that are real and useless: statistically separable, operationally identical. The order that works starts at the other end. Name the decisions that could be made differently by group — the option list, the depth profile, the size or shade curve, the flow date, the pack configuration, the carryover policy. For each one, ask which observable property of a door predicts it. Cluster on that property. A clustering dimension earns its place by changing a decision, and the test is a specific, named decision rather than a general sense that the doors are different.

Most fleets find that two or three dimensions carry almost everything, and they are usually not the ones with the best data. Climate and the week a season opens are cheap to observe and move flow dates by weeks. Selling space and fixture capacity are measurable, rarely maintained, and hard-cap the option count a door can physically hold. Trade-area composition predicts the shape of a size, shade or age-band ladder. Revenue predicts depth, which is what grading is for, and predicts almost nothing about what the door should carry.

How many clusters a fleet can carry

Two limits bind, and they bind from opposite directions.

The evidence limit is a floor under cluster size. A cluster exists to be read: its sell-through tells you whether the assortment built for it worked. Below roughly one unit per option per week across the whole cluster the read is zeros and ones, and a cluster sitting there is refitted on noise every year, which shows up as membership that churns without any door's trade area having changed. The data readiness guide covers what has to be true of the underlying record before any of this is measurable.

The execution limit is a ceiling on cluster count. Every cluster is another option list to build, another depth profile to buy, potentially another size or shade curve, another pack configuration to specify and another set of allocation rules to maintain. Those costs are paid in planner hours, in supplier minimums and in DC complexity, and they are paid whether or not the cluster turns out to have been worth splitting. The cost is real and mostly invisible until the season it breaks.

Between the two sits a simple merge test that settles most arguments: take any two clusters and ask what would actually differ in next season's plan if they were one. If the answer is a different depth, that is grading and not a cluster. If the answer is nothing anyone would act on, they are one cluster. If the answer is a different option list, different widths, a different shade ladder or a different flow date, they are two.

The data you need, and what to do when you do not have it

The minimum useful record is sell-through by option by door by week, with receipts and on-hand beside it so availability can be reconstructed. Without receipts, sales alone cannot distinguish a door that did not sell something from a door that never had it, which is the single most damaging ambiguity in this work. In-stock rate is what separates the two, and lost sales is what sizes the gap.

Where that record does not exist, cluster on observable attributes and say so. Trade-area characteristics, climate zone and historical season start, selling space and fixture count, door type and channel, competitive set, and the demand pattern of a comparable door already trading are all usable. An attribute-based cluster is a hypothesis with a test date, not a measurement, and the discipline is to write down which doors were assigned by proxy so that their first full season is read as evidence about the proxy as well as about the door. New doors are permanently in this position by definition; planning for new store openings covers the assignment at opening and the re-test after a comparable period.

One correction has to be made before any history is clustered on. A door's sales record is a record of what it was given, not of what its customers wanted. A door that was never allocated the extended range, the second width or the deep shades has no demand history for them, and a cluster fitted on that record concludes the door has no demand for them. It then confirms itself next season by allocating none. Censored history does not just distort a cluster, it makes the cluster self-fulfilling, which is why the availability correction comes before the grouping rather than after it.

Illustrative example: what a fleet's option capacity actually is

The figures below are illustrative, chosen because they divide cleanly. They are not benchmarks, not targets, and not drawn from any brand.

A 60-door fleet plans one class for a 26-week season. The class buy is 9,000 units across 60 options. Minimum sellable presentation is 6 units per option per door, one per size across a six-size run, so a door either holds a full run of an option or does not carry it.

StepArithmeticResult
Average buy per option9,000 ÷ 60 options150 units
Doors one option can cover at minimum presentation150 ÷ 6 units25 doors
Total door-option placements available60 options × 25 doors1,500
Options each door can hold1,500 ÷ 60 doors25 options

The fleet cannot carry all 60 options in all 60 doors; at this budget each door holds about 25 of them. That is the real question clustering answers: which 25, in which doors. A fleet that never does this arithmetic tends to plan 60 options as though every door will carry them, then discovers the shortfall at allocation, where it resolves as thin coverage everywhere rather than as a decision.

The same figures set the evidence floor, read at full sell-through — the most generous case. Each door holds 150 units of the class for the season, which sold out over 26 weeks is about 5.8 units per week across the class, and spread over its 25 options is roughly 0.23 units per option per door-week. A five-door cluster therefore sees about 1.2 units per option per week and a ten-door cluster about 2.3. A five-door cluster in this fleet is at the edge of readability before the season even starts, and at any sell-through below full it is already under the floor rather than at it — which is the arithmetic behind the advice to keep cluster count low, not a preference for simplicity.

Both numbers move with the buy, the presentation minimum and the option count, and that is the point: cluster count is downstream of the open-to-buy envelope and the option count, not an independent choice. Open-to-buy by vertical holds the budget and what an option costs prices the breadth decision.

Apparel: climate band, price band, and the size curve underneath

Apparel clusters on two dimensions that pull in different directions and a third that sits beneath both. Climate band decides timing before it decides content: the week outerwear starts selling differs by weeks across a national fleet, and a single flow date sends heavyweight outerwear to warm doors while cold doors are still holding transitional weight. Price band decides content: the same style-color sells at full price in one trade area and only on promotion in another, and a door whose sales are concentrated in the promotional weeks is telling you its price ladder sits lower, not that the style failed.

Underneath both sits the size curve, which is where cluster-level planning pays for itself most reliably. Size demand varies by trade area in ways that are stable enough to plan on and invisible at fleet level, because the fleet curve is an average of curves that do not look like it. The condition for fitting a cluster curve is strict: it is built only from door-weeks in which the full run stood on the floor, since a curve measured across broken runs describes what was left rather than what was wanted. Where a cluster has too few complete door-weeks, it inherits the fleet curve and is marked as inherited rather than measured. Size and pack optimization carries the method, and planning an extended size range covers the bands that need their own curve rather than a stretched version of the core one.

Channel cuts across all of this. DTC doors, wholesale doors and outlet doors are usually separate clusters rather than points on one ladder, because their price architecture, markdown cadence and carryover policy differ structurally — planning an outlet channel sets out why outlet doors distort a full-price read when they are mixed into the same cluster. The frequent mistake is clustering by region, which feels natural and is usually wrong: a downtown door has more in common with a downtown door four hundred miles away than with a suburban door in its own metro. See store clustering for apparel brands for the full apparel treatment, including DTC and wholesale door clustering.

Footwear: width fit, run depth, and model-year carryover

Footwear has a clustering dimension that most fleets never use and should: width fit is binary at the door and continuous in the demand. A door either stocks a second width or it does not, and a door that has never carried one has no sales history for it, which means a history-only cluster will reliably conclude that the demand is not there. Width demand has to be estimated from returns, from special orders, from the doors that do carry the width, and from trade-area composition before any grouping is done.

The second dimension is the depth of the size run a door can sustain. Every extra width and half size divides the same units across more positions, so a door that cannot hold a complete run at sellable depth should carry fewer models in full runs rather than more models in broken ones. That is a cluster decision, not an allocation decision, because it changes the option list rather than the quantity. A broken run is worse than a narrower assortment: the core sizes sell out, the ends mark down, and the model's recorded sell-through understates what it would have done.

The third is model-year carryover policy. Some doors clear prior-year product at a discount and depend on it for traffic; others must be current-season only because prior-year pricing sitting beside new-season pricing suppresses the new model. Those are different clusters even at identical volume, and mixing them makes both reads wrong — the carryover door's discount inflates its unit velocity while depressing its margin, and the current-only door looks slow by comparison. Planning a model year changeover covers the transition itself. Athletic, comfort and dress mix is a fourth axis where a brand spans them, since the same trade area can be strong in one and weak in another. See assortment planning for footwear brands.

Accessories & bags: the host program, hero colors, and gifting doors

Accessories cluster badly on their own history and well on the program they attach to. A belt, a scarf or a small leather good that sells alongside an apparel or footwear program is driven by the attach rate on that program in that door, so the right clustering unit for attached accessories is the host's cluster, not the accessory's own thin sales record. A door that carries a deep denim program is a different accessories door from one that does not, whatever the two look like on accessories revenue alone.

The second dimension is gifting exposure. Airport, tourist, transit and mall doors concentrate demand into weeks around the gifting calendar, and their annual total is a poor description of what they need: they need depth arriving before a peak and very little afterwards. Everyday doors run flatter and want breadth over depth. Clustering those two together produces a plan that is too shallow for the peak and too deep for the rest of the year, and the residue shows up as post-peak markdown. Planning the gifting calendar dates the peaks.

The third is colorway capacity. Hero colors — the black leather, the house hardware, the evergreen core — carry across the fleet and behave like replenished product. Seasonal fashion colors are a capacity decision: how many a door can hold without any of them getting enough facings to sell. A cluster with fixture capacity for three fashion colors and an assortment built for six is a cluster that will merchandise four and stock two in the back. Tannery and hardware minimums then round the answer, so the number of clusters that can each get their own fashion color is capped by the minimum order rather than by the demand — the constraint planning accessories lines develops.

Home & furniture: floor-set capacity and delivery radius

Home and furniture clusters on two constraints that barely exist in soft goods. The first is floor-set capacity, which is physical and absolute: a showroom holds a countable number of vignettes and slots, and an assortment that exceeds them is not a stretch, it is undeliverable. A floor set is the unit here, not a facing count, so square footage and slot count are the primary clustering attributes and revenue is secondary. Two showrooms at the same revenue with different footprints need different assortments, because the larger one can stand a collection and the smaller one can only stand a hero piece from it.

The second is delivery radius. A door's quoted delivery date is a function of its distance from the stocking location and of what is held there, and a long quoted date suppresses conversion on exactly the models a customer will not wait for. Doors that share a delivery radius therefore share a conversion profile and belong together, while a door outside the radius needs a different mix weighted toward what it can hold and hand over. That makes the stocked-versus-special-order split a cluster decision: which models stand on the floor to be taken away, and which stand to be ordered against a quoted date.

Container economics then constrain the cluster count directly. Quantities round to container cube at landed cost, and a container is normally a mixed load rather than a single finish, so each additional finish-level split has to earn a viable share of one. Container economics cap how many finish-level splits a fleet can economically carry, usually to a small number — and it is a constraint settled at the buy rather than on the cluster map. Dealer prebooks add a third axis where the brand sells through dealers, since a prebooking dealer and a stocking dealer commit on different calendars. See merchandise planning for home and furniture brands.

Outdoor: climate zone, the week the season opens, and the dealer fleet

Outdoor is the clearest case of a vertical where the cluster decides dates before it decides content. The week a season actually opens moves by weeks across elevation, latitude and micro-climate, and a fleet on one national flow date is early in half its doors and late in the other half. The clustering attribute is not the region name but the historical week demand for the category turned on, anchored to weather rather than to the calendar, because a prior year re-read on its own weather anchor tells you the shape of the season while a prior year read on calendar weeks tells you when the weather happened to arrive that year.

Terrain and use case form the second dimension. Doors serving alpine, coastal, desert or urban-commuter customers need different specs from the same model line, and spec is not a depth question. Counter-seasonal lines cut across again, since a door that trades through a shoulder season needs a different flow than one that closes down.

The third is the split between dealer doors and the brand's own doors and DTC. Dealer prebook units are a commitment rather than a sale, and a cluster built on prebook volume is clustering on how a dealer buys rather than on how the dealer's customers shop. Sell-through-reporting dealers are the only ones whose record can be clustered directly; non-reporting dealers are assigned by proxy from the reporting dealers in their zone, and the proxy is recorded. A dealer fleet clustered on sell-in is clustered on the dealer's cash position, not on demand, which is why the two reads are kept apart. See merchandise planning for outdoor brands.

Sporting goods: sport mix, the school calendar, and team cycles

Sporting goods clusters on sport mix first, and sport mix is local in a way that resists national planning. Hockey, baseball, soccer, lacrosse, basketball and the outdoor field sports have participation that varies sharply by market, and a door in a hockey market and a door in a soccer market are different stores wearing the same fascia. The cluster decides which sports get floor space in which weeks, which is a sequencing decision as much as an assortment one, because the same square footage has to turn over between sports across the year.

The school calendar is the second dimension and it moves dates rather than quantities. Term start varies by region, and with it the back-to-school window that carries a large part of the year's equipment and team-wear demand. A fleet on one national back-to-school flow date is shipping into an empty market in the regions that start late and arriving after the decision in the regions that start early. Clustering on term start is cheap — the dates are published — and it is one of the few clustering dimensions that needs no sales history at all.

Team and league cycles form the third. Doors serving organised team business run on order windows that close, and a closed window is a hard stop rather than a slow week, so a team-heavy door's sales record has structural zeros in it that must not be read as weak demand. Consumables — balls, grips, shuttles, wax, tape — behave like replenishment rather than like seasonal assortment and are clustered on velocity and reorder cadence rather than on sport mix, so most fleets end up with a consumables cluster map that does not match the hardgoods one. That is normal and worth keeping separate: one door can sit in different clusters for different parts of the assortment, and forcing a single map across a mixed catalogue is what makes clustering feel useless in this vertical. Prebook against the dealer or team channel then sets how far ahead each cluster's decision has to be made. See merchandise planning for sporting goods brands.

Health & beauty: shade demand, the reset calendar, and the tester

Beauty clusters on the shape of demand across the shade ladder, and it is the vertical where getting this wrong is most visible on the shelf. Doors differ in where their demand sits along a foundation or complexion ladder, and the ends of the ladder are thin everywhere, so a fleet-average ladder sent to every door leaves the ends empty in exactly the doors whose customers needed them and sitting unsold in the doors that did not. A shade cluster is a decision about how far into the ends of the ladder a door goes, not about how much foundation it gets.

The correction that has to come first is the tester. A shade-week that ran without a working tester is not evidence about that shade, and a shade that was never shipped to a door produces a zero that records an allocation decision rather than a customer decision. Both are cleaned out before the ladder is fitted, and the ladder is fitted on the doors that carried every shade. Without that step the cluster encodes the last cluster's mistake — the mechanism described above, and the one that makes shade coverage gaps persistent rather than self-correcting.

The second dimension is the retailer's reset calendar, which the brand does not control. Gondola resets set the dates on which a door's planogram can actually change, so doors on the same reset cycle can be planned to act together and doors on different cycles cannot, whatever their demand looks like. A cluster whose members cannot change their planogram in the same week is a cluster that cannot execute its own plan. Shelf life is the third: the retailer's remaining-life requirement at delivery, computed against the product's unopened shelf life, caps how much cover a slow door can hold, so low-velocity doors need shallower depth on the same ladder rather than a narrower ladder — the arithmetic in dating rules and weeks of supply. Regimen attach, where a serum's demand follows the cleanser ahead of it, means franchise clusters travel together across steps. See merchandise planning for health and beauty brands.

Toys & games: the shape of the gifting peak and licensed windows

Toys clusters on the shape of the peak rather than on the annual total, because the annual total is close to meaningless as a planning unit in a category where a large share of the year's demand lands in a few weeks. Two doors with identical annual revenue can have very different peaks — one that builds through the autumn and one that does almost nothing until the last three weeks — and they need different arrival dates and different depth. The clustering attribute is the weekly curve, normalised, not the total under it.

The second dimension is proximity to a licensed window. A property with a theatrical, streaming or event release has a demand window with dates, and doors differ in how sharply they respond to it: some trade areas move hard on a release and others barely register it. That is a genuine cluster distinction because it changes the option list — which properties a door carries at all — and not merely the depth. Regional franchise affinity behaves the same way where a sports or local property is involved. Planning a licensed product window covers the window itself, including what happens to the assortment after it closes.

The third is the price ladder and format mix. Impulse and pocket-money formats near the till, mid-price boxed gifts and high-price hero items sit at different points on a door's ladder, and a door's ladder is a function of its trade area rather than of its size. A cluster built on revenue puts a high-volume impulse door and a high-volume hero-gift door together and sends both a blend that suits neither. The hard constraint behind all of it is the factory cut-off before the peak: the quantity closes months ahead, so cluster assignments for the peak have to be settled before the buy, not discovered during it. See merchandise planning for toy and game brands.

Baby & juvenile: age-band mix, registry doors, and certification changeover

Baby and juvenile clusters on age-band mix, which tracks local birth cohorts and the age profile of the trade area and is stable enough to plan on. Newborn, infant, toddler and preschool are different assortments with different velocities, different sizing logic and different gifting behaviour, and a door weighted to newborn is not a smaller version of a door weighted to preschool. Treating "kids" as one band is the error that makes every other decision in this vertical approximate, because the bands turn over at different rates and graduate customers out of themselves on a fixed clock.

The second dimension is registry activity. Doors with an active registry business see demand that is committed ahead of the purchase and concentrated on specific configurations, and a registry item shown as unavailable is usually swapped rather than waited for, so the sales record understates the original intent. Registry-heavy doors need depth on the registry core and coverage of the configurations the registry lists; walk-in doors need breadth. Planning with registry demand covers how that intent is read and what it does to depth.

The third is certification and model-year changeover. A configuration change driven by certification puts a hard date on the old version, and doors differ in how fast they clear it — a high-turn door is current within weeks while a slow door is still holding the prior configuration when the new one arrives. Those doors need different exit plans, which makes changeover behaviour a legitimate clustering axis rather than a one-off exception. Gifting depth cuts across again, since gift-weighted doors peak on the same calendar as toys while everyday replenishment doors run flat. See merchandise planning for baby and juvenile brands.

Jewelry & watches: price-band ladder and single-piece depth

Jewelry and watches cluster on the price-band ladder, and the reason is structural rather than stylistic. Depth is usually one piece: a door either has a piece in a band or it does not, so the assortment question is which bands are represented rather than how deep each one runs. Two doors can reach identical revenue with completely different ladders — many pieces low on the ladder, or few pieces high on it — and grouping them on revenue produces a cluster whose plan is wrong for both.

The precious-metal cost base moves the ladder from underneath. When metal cost moves, the band a given design sits in moves with it unless the design changes, so a cluster defined on last year's retail bands can quietly describe different product this year. The clustering attribute that survives that is the door's position on its own ladder — its share of units in its top, middle and entry bands — rather than an absolute price range. Planning margin on a moving cost base covers the costing side.

Single-piece depth also forces a different reading of history. One sale empties the position, so a band's record at a door is assembled from the door-weeks in which that band actually had a piece present, and a door that sold its only piece in week one contributes one week of evidence rather than a season of zeros. Appointments, try-on requests and size enquiries are demand the empty weeks could not record and belong in the band's read. Pieces out on memo are displayed somewhere other than where the stock is owned and come out of the owning door's selling record entirely. Watches add model year on top, with prior-year pricing behaving the same way it does in footwear, and engagement versus fashion mix is a fourth axis where a brand spans both. See merchandise planning for jewelry and watch brands.

The clustering dimension by vertical, side by side

VerticalDimension carrying most variationWhat the cluster decidesThe correction to make firstHow it breaks
ApparelClimate band and price bandOption list, size curve, flow dateRefit curves only on complete-run door-weeksClustering by region instead of demand
FootwearWidth fit and sustainable run depthWhich widths ship, run completeness, carryover policyEstimate width demand where no width was ever stockedGrading used in place of clustering
Accessories & bagsHost program and gifting exposureFashion colorway count, peak depthRead attach rate only on host in-stock weeksClustering on the accessory's own thin history
Home & furnitureFloor-set capacity and delivery radiusFloor-set mix, stocked vs special orderAdd back special orders as unmet floor demandClustering on revenue when space is the constraint
OutdoorClimate zone and the week the season opensFlow dates, then spec mixRe-phase prior year on its own weather anchorClustering dealers on sell-in rather than sell-through
Sporting goodsSport mix and term start dateWhich sports hold floor in which weeksTreat closed order windows as structural zerosOne national back-to-school flow date
Health & beautyShade demand shape and reset cycleHow far into the ladder ends a door goesDrop shade-weeks with no working testerFleet-average ladder sent everywhere
Toys & gamesShape of the gifting peak, licensed windowPrice ladder, property list, arrival datesNormalise the weekly curve before comparing doorsClustering on annual totals
Baby & juvenileAge-band mix and registry activityBand share, registry core depthAdd back registry items swapped when unavailableTreating all age bands as one
Jewelry & watchesPosition on the price-band ladderWhich bands get a pieceBuild band history from piece-present door-weeksGrouping two different ladders on equal revenue

Two patterns run through the table. The correction column is almost always the same operation in different clothing: the record has to be cleared of the previous assortment decision before it can be used to make the next one. And the failure column is almost always a proxy standing in for the real dimension — region for climate, revenue for space, volume for pattern.

How clusters feed allocation and size curves

A cluster sets what a group of doors receives. It does not set how much any individual door receives, and confusing the two is where cluster-level planning most often collapses into either rigidity or noise.

The division of labour is clean. The cluster owns the option list — which styles, models, shades, properties or price bands are carried — and the profile: the size, shade or age-band distribution and the presentation minimum. Allocation owns the quantity, sized per door on its own rate of sale, its space and its current position, which is what fair share allocation and door-level demand compute. A door in a large cluster is not allocated the cluster average; it is allocated its own number inside the cluster's assortment. Allocation and replenishment by vertical covers the placement work, and allocation and replenishment best practices the operating rules.

Size curves follow the same split. The curve is a cluster property because it describes shape, and the units are a door property because they describe volume; size curve allocation carries the arithmetic. The condition for a cluster to have its own measured curve is that it has enough complete-run door-weeks to fit one — otherwise it inherits the fleet curve, flagged as inherited. An inherited curve that is never re-tested becomes a measured curve by forgetting, which is worth guarding against explicitly in the record.

Two operating rules make clusters useful in season rather than just at plan time. Transfers and reallocations run within a cluster before they run across clusters, because a unit moving between doors that share a demand pattern has a much better chance of selling than one crossing into a different one. And the in-season read is taken at cluster level as well as at fleet level, since a style that is flat on the fleet is often strong in one cluster and dead in another — the distinction that decides whether a chase is worth placing and where the resulting units should go. A fleet-only read averages those two into a decision that is wrong for both.

Re-clustering: cadence, triggers, and what to keep

Clusters are reviewed once a year, on the planning calendar, timed so that assignments are settled before the next season is planned. The reason for a fixed cadence rather than continuous re-fitting is that a cluster that moves while it is being planned for cannot be scored, and a plan nobody can score is a plan nobody learns from.

Between annual reviews, a door moves on a structural trigger rather than on results. The triggers are specific: a remodel or relocation, a change in selling space or fixture count, a channel change, a competitor opening or closing in the trade area, a catchment change, or a change in the door's own operating model such as adding a service or a click-and-collect volume. A door that simply had a weak season stays where it is, because one season of one door's variation is not evidence that its demand pattern changed — and the cost of moving it is that the cluster it left and the cluster it joined both become harder to read.

Three things are worth keeping in the record. The cluster membership version by season, so that a hindsight scores each season against the clusters that were actually live when the buy was made rather than against today's map. The reason each door sits where it does, in a form that can be re-run rather than recalled — an attribute rule beats a remembered argument. And the proxy flags on doors assigned without their own history, so their first full season is read as a test of the proxy. Hindsight analysis is where the cluster map earns or loses its standing for next year.

A useful annual output is not a new map but a short list: which doors the rule would move, which of those moves anyone believes, and which clusters failed to produce a readable signal. A clustering review that moves nothing is a valid result, and it is a much better one than a map that churns every year while the trade areas stand still.

How clusters break

Six failure modes account for most of it, and all six are visible before the season if anyone looks.

Too many clusters to read. Each one falls below the units-per-option-per-week floor, so next year's re-fit is fitted to noise and membership churns without any door changing. The symptom is a map that reshuffles annually and a planning team that no longer trusts it.

Clusters that encode history rather than demand. A door that was never allocated the widths, shades, extended sizes or price bands shows no history for them, the cluster records that as no demand, and next season allocates none. This is the self-fulfilling failure described earlier, and it is the one that hides longest because the data keeps confirming it. The availability correction is not optional.

Clusters nobody allocates to. The plan is built at cluster level and then overridden door by door at allocation, usually for reasons that are individually defensible. After a season of that the cluster map is documentation rather than a decision, and the assortment that actually shipped is nobody's plan. The test is mechanical: compare what each cluster was planned to receive with what its doors actually received.

Clusters built on volume. Volume is an outcome, and clustering on it groups doors that arrived at the same revenue by different routes. This is grading wearing a clustering label, and it produces the scaled-down-copy problem in every vertical above.

Clusters that became identities. A door labelled flagship or trend-forward five years ago keeps the label after the trade area changed, because the label stopped being a rule and became a description. Attribute rules that can be re-run are the defence; names that cannot be re-derived are the vulnerability.

Clusters the supply chain cannot honour. A cluster whose assortment needs a pack configuration, a colorway or a container share that no one can actually ship is a plan that will be silently rounded at allocation. Case packs and planned depth and pre-pack versus singles cover where that rounding happens. The resulting mismatch is a common source of inventory distortion — overstocked doors beside stocked-out ones inside the same cluster — and it is usually blamed on allocation rather than on the cluster that could never be executed.

Free Template

Assortment Planning Template

Clusters become a plan at the breadth-and-depth step, and this is the working file for it: set the category mix with target shares, option counts and average selling prices, and the sheet derives planned units per category, splits them by size curve and carries a margin bridge beside them. Choosing the clustering dimension and the cluster count stays with you; the file is where the chosen assortment turns into options, depth and units.

  • Emailed to you
  • No credit card
  • Work email required
  • Yours to keep

For brand, retailer, wholesaler and manufacturer teams. We email the link to your work address; requests from companies we can’t verify are reviewed first.

Where RetailNorthstar fits

RetailNorthstar is a merchandise planning platform — OTB, assortment and line planning, buy planning, allocation, sizing, PO and WIP tracking and analytics on a shared data model. As localized assortment planning and store clustering for apparel brands set out, it supports cluster-level planning: the assortment is built at cluster level, locations are assigned to clusters, and allocation places inventory across the doors inside each cluster. Because the plan, the buy and the allocation share one data model, a cluster assignment that changes reaches the assortment and the financial roll-up rather than sitting in a separate file.

Choosing the clustering dimension, setting the cluster count and deciding when a door moves remain judgement calls that belong to the planning team. Assortment planning is where the cluster assortment is built, allocation is where it is placed, and OTB planning holds the budget that caps how many clusters the fleet can afford to carry.

Related resources

See how RetailNorthstar supports cluster-level planning — build the assortment at cluster level, assign locations to clusters, and allocate across the doors inside each one.

Start Free Trial →

No credit card. No commitments.

Common questions

What is store clustering?

Store clustering is the practice of grouping selling locations that respond to the same assortment in the same way, so one plan can be built for the group instead of one plan per door. A cluster is a planning unit, not a description of a store: doors belong together when giving them the same option list, the same depth profile and the same size or shade distribution would produce the same outcome. The cluster decides what is offered and how deep it runs; allocation then sizes each door inside the cluster on its own demand and its own space. Clustering exists because the two alternatives both fail at scale — one assortment everywhere ignores real demand differences, and one plan per door multiplies the planning work by the door count without adding accuracy.

How many store clusters should a retailer have?

As many as the fleet can support with evidence and the operation can actually execute, which is usually far fewer than the number that looks meaningful in the data. Two limits bind. The evidence limit: a cluster has to produce enough units per option per week to read, and below roughly one unit per option per week across the whole cluster the sell-through signal is zeros and ones, so a small cluster is refitted on noise every year and its membership churns for no reason. The execution limit: every cluster is a separate option list, a separate depth profile, a possible separate size or shade curve and a separate pack configuration, all of which have to be bought, packed and allocated. A practical test is to ask what decision would change if two clusters were merged. If the answer is nothing that anyone would act on, they are one cluster.

What is the difference between store clustering and store grading?

Grading ranks doors on one axis, almost always volume, and produces an ordered ladder — A, B, C — where the difference between grades is how much, not what. Clustering groups doors on the pattern of demand and produces unordered groups where the difference is which product, which sizes or shades, and when. They answer different questions and most fleets need both. Grading sets depth and the order in which doors receive newness; clustering sets the assortment those doors receive in the first place. Substituting one for the other is the most common failure: a grade-only fleet sends a scaled-down version of the biggest door's assortment to every small door, which is why a small door with a distinctive demand pattern never gets the widths, shades or price bands its customers actually buy.

How do you cluster stores when you do not have enough sales history?

Cluster on the attributes you can observe rather than on the history you do not have, and be explicit that the result is a starting hypothesis to be tested rather than a measurement. Usable attributes include trade-area characteristics, climate zone and season start date, selling space and fixture capacity, channel and door type, competitive set, and the demand pattern of a comparable door already trading. A new door has no history by definition, so it is assigned to a cluster by proxy at opening and re-tested after a full comparable selling period. The discipline that matters more than the method is recording that the assignment was a proxy, because the first season of a proxy assignment is evidence about the proxy as well as about the door.

How often should store clusters be reviewed?

Once a year on the planning calendar, timed so the new assignments are settled before the next season is planned, and never in the middle of a season the current clusters are already buying for. Between annual reviews, a door moves only on a structural trigger rather than on a run of results: a remodel or relocation, a change in selling space, a channel change, a new competitor opening or closing in the trade area, or a catchment change. A door that simply had a bad season stays where it is, because the ordinary variation of a single door over one season is not evidence that its demand pattern changed. Keep the version history of cluster membership, since a hindsight that scores this season against next season's clusters is scoring a plan nobody made.

Why does store clustering work differently by vertical?

The method holds everywhere; what changes is which dimension carries most of the variation between doors, and that is a category fact rather than a planning preference. In apparel it is usually climate and price band, with the size curve underneath. In footwear it is width fit and the depth of the size run, because a door either carries a width or it does not. In home and furniture it is floor-set capacity and delivery radius, since a showroom is limited by square footage and a long quoted delivery date suppresses conversion. In beauty it is the shape of shade demand and the retailer's reset calendar. In toys it is the shape of the gifting peak and proximity to a licensed window. Cluster ten verticals on the same dimension and the arithmetic is right while the groups are meaningless.

RetailNorthstar Editorial Team
RetailNorthstar ·

Share this guide with your team

Copy a link or a pre-written message for Slack, Teams, or email.

// Know where your operation stands

Apply this to your planning operation.

The free Apparel Planning Maturity Assessment benchmarks your operation and tells you exactly which gaps to fix first.

Take the assessment →

Apply these insights with RetailNorthstar.

See how modern apparel brands use RetailNorthstar to put this planning framework into practice.

No credit card. No commitments.

Connected merchandise planning — live in weeks, not quarters.