Skip to content

Week 9 | Session 1: Intelligent Decision Tools - DC Location Problem & K-means Clustering (Intro)

Course: Supply Chain Digitization - Module 3: Analytics in SCM



Background: A pharmaceutical company wants to expand into a new region and has collected location data (latitude & longitude) for 811 prospective customers.

Latitude →Longitude
The 811 prospective customers, plotted by latitude and longitude. (Illustrative sample - the shape of the point cloud, not the exact coordinates.)

Need to decide: where to open Distribution Centers (DCs) to serve these 811 customers.

  1. Decision 1: Where to locate DCs? (How many? At what coordinates?)
  2. Decision 2: Which customers served by which DC? (How to segment?)

As the number of DCs increases, customers per DC decreases → faster service → responsiveness improves. However, fixed + variable cost per DC × more DCs → total cost increases.

No. of DCsCustomers per DCResponsivenessTotal Cost
1All 811Low ↓Low ↓ (cheapest)
3-4~200 each↑↑↑↑
6~135 eachHigh ↑↑↑High ↑↑↑ (expensive)

Need to find the optimal K (number of clusters) that balances responsiveness with cost.

123456Number of DCsCostFacilityTransportTotaloptimal K
The DC-count tradeoff: adding DCs raises facility cost but lowers transport cost (better responsiveness). Total cost is U-shaped - the optimal number of DCs sits at its minimum.

★★DC1DC2Latitude →
K-means with K = 2 - every customer is assigned to the nearest of two clusters, and each cluster’s centroid (★) becomes the proposed DC location. (Illustrative.)

What is K-means Clustering? An unsupervised machine learning technique for grouping data points into K clusters.

  • Input: dataset with coordinates (lat, long) + value of K.
  • Output: (1) cluster assignment for each data point, (2) centroid of each cluster.
ConceptExplanation
K (input)Number of clusters = number of DCs to open. User defines K.
Output 1 - SegmentationEach of 811 customers assigned to exactly 1 cluster.
Output 2 - CentroidCenter point of each cluster = proposed DC location (lat, long).

Centroid = geometric center of all customer coordinates in that cluster. Distance from centroid to each cluster member is minimized on average. Placing the DC at the centroid → minimum average travel distance to all assigned customers.


  • Decision 1 - Customer Segmentation: K-means automatically assigns each of 811 customers to exactly 1 cluster based on spatial proximity.
  • Decision 2 - DC Location: The centroid of each cluster gives the latitude and longitude of the proposed DC.
★★★★DC1DC3DC4DC2Latitude →
With K = 4, the customers split into four tighter clusters, each served by its own centroid DC - cutting average travel distance further, at the cost of a fourth facility. (Illustrative.)

K-means solves BOTH decisions for a given K - but what is the right K? Is K = 1, 3, 4, 5, or 6 optimal? This is answered using the Elbow Method (covered in the next session).


  • Problem: 811 customer locations (lat/long) - where to open DCs, how to assign customers.
  • Tradeoff: More DCs → better responsiveness, higher cost.
  • K-means Clustering: unsupervised ML technique - groups customers into K clusters.
  • Output 1: Segmentation (customer assigned to 1 DC).
  • Output 2: Centroid (proposed DC location).
  • Why centroid? Minimizes average distance from DC to all customers in that cluster.