Online Retail Customer Segmentation | DKD
MACHINE LEARNING • MLOPS

Online Retail Customer Segmentation using Machine Learning & MLOps

End-to-End Machine Learning • Customer Analytics • FastAPI • Streamlit • MLflow • Docker • CI/CD

01

Project Overview

Developed an end-to-end Machine Learning Operations (MLOps) solution for customer segmentation using transactional retail data. The system converts raw customer purchase history into meaningful behavioural segments using RFM analysis and K-Means clustering.

The project extends beyond model development into a production-oriented workflow that includes feature engineering, model serialization, real-time prediction, prediction logging, monitoring, data-drift detection, automated retraining, CI/CD, API deployment, and an interactive Streamlit dashboard.

LIVE DEMONSTRATION

Explore the deployed customer segmentation dashboard

Open Streamlit App ↗
02

Business Problem

Retail businesses generate large volumes of transactional data, but raw transaction records do not directly reveal which customers are valuable, newly acquired, disengaged, or at risk of becoming inactive.

Traditional customer segmentation can require significant manual analysis and may be difficult to maintain as new transactions arrive.

01

Customer Value

Identify customers generating significant monetary value and purchasing frequently.

02

Engagement

Understand how recently and frequently customers interact with the business.

03

Retention

Identify behavioural patterns associated with inactive or at-risk customers.

03

Project Objectives

01

Automate customer segmentation using machine learning.

02

Transform transactional data into meaningful RFM features.

03

Deploy a real-time customer segmentation API.

04

Provide an interactive analytics dashboard.

05

Log production predictions for monitoring.

06

Detect production data drift.

07

Support automated model retraining.

08

Implement a reproducible MLOps deployment workflow.

04

Dataset & Data Preparation

The project uses the Online Retail transactional dataset. Each record represents a retail transaction containing information such as invoice number, product, quantity, unit price, customer identifier, country, and transaction date.

Raw Transactions
→
Cleaning
→
Customer Aggregation
→
RFM Features
→
Clustering

Data Preparation Steps

05

Feature Engineering

Customer transactions were transformed into behavioural features using RFM analysis. Additional derived features were incorporated to provide a richer representation of customer purchasing behaviour.

Recency

Number of days since the customer's most recent purchase.

Recency

Frequency

Number of purchases or transactions associated with the customer.

Frequency

Monetary

Total monetary value generated by the customer.

Monetary

Average Order Value

Average monetary value associated with each customer transaction.

Monetary / Frequency

Customer Value

Derived customer-level value used to capture purchasing contribution.

Frequency × AverageOrderValue
06

Machine Learning Solution

K-Means clustering was selected to group customers according to similarities in their purchasing behaviour.

ML
UNSUPERVISED LEARNING

K-Means Clustering

The model groups customers into clusters based on their engineered behavioural features. The resulting cluster assignments are mapped to business-oriented customer segments.

07

Customer Segments

01

VIP Customers

Customers demonstrating strong purchasing activity and high monetary contribution.

02

New Customers

Customers with relatively recent purchasing activity and developing engagement.

03

At-Risk Customers

Customers showing behavioural patterns that may indicate declining engagement.

04

Inactive Customers

Customers with low or historically distant purchasing activity.

08

MLOps Architecture

The project was designed as a complete machine learning lifecycle rather than a standalone notebook model.

01 Data Raw transactions
→
02 Preprocessing Cleaning & validation
→
03 Features RFM engineering
→
04 Training K-Means model
→
05 Deployment FastAPI
→
06 Monitoring Logs & metrics
→
07 Drift Data monitoring
→
08 Retraining Model refresh
09

Interactive Streamlit Dashboard

A business-facing Streamlit application was developed to make the machine learning system accessible without requiring users to interact directly with the underlying Python code.

Customer Prediction

Users can enter customer RFM information and receive a predicted customer segment.

Customer Persona

Converts model output into an interpretable customer profile.

Business Recommendations

Provides segment-oriented actions for customer engagement and retention.

Prediction History

Displays historical prediction activity using logged production requests.

Monitoring

Tracks prediction counts and cluster distribution from production logs.

Drift Detection

Compares production feature distributions with training data to identify potential drift.

10

FastAPI Prediction Service

A FastAPI backend provides programmatic access to the trained segmentation model and exposes endpoints for prediction, health checks, monitoring, and drift detection.

Endpoint Purpose
GET / API status and welcome response
GET /health Service health check
POST /predict Predict customer segment
GET /monitor Production monitoring metrics
GET /drift Data drift analysis
11

Production Monitoring

Every prediction request can be logged with the input features, engineered values, timestamp, predicted cluster, and customer segment. These logs provide the foundation for production monitoring.

Prediction Count Total production predictions
Latest Prediction Most recent prediction record
Cluster Distribution Distribution of predicted clusters
Prediction Logs Historical production requests
12

Data Drift Detection

Production feature statistics are compared against the training dataset to identify changes in customer behaviour. The drift analysis examines the same engineered features used by the production model.

Recency Frequency Monetary Average Order Value Customer Value
Training Data Reference distribution
VS
Production Data Recent prediction inputs
→
Drift Report Feature-level comparison
13

Automated Retraining

The monitoring workflow is designed to identify significant changes between the training and production feature distributions. When the configured drift condition is satisfied, the retraining workflow can be triggered to refresh the segmentation model.

MONITOR → DETECT → DECIDE → RETRAIN

This closes the machine learning lifecycle by connecting production monitoring with model maintenance.

14

Technology Stack

Python Pandas NumPy Scikit-learn Streamlit FastAPI Plotly MLflow Docker GitHub Actions Evidently AI
15

Deployment

The application is structured for cloud deployment with separate user-facing dashboard and API components. The FastAPI service provides the model inference layer, while Streamlit provides the interactive business interface.

FRONTEND

Streamlit

Interactive customer segmentation and analytics dashboard.

BACKEND

FastAPI

REST API for model inference, monitoring, and drift detection.

CONTAINERIZATION

Docker

Reproducible application environment for deployment.

CI/CD

GitHub Actions

Automated testing and deployment workflow.

16

Business Impact

01

Automated Segmentation

Reduces the need for manual customer grouping.

02

Customer Targeting

Provides interpretable customer groups for targeted marketing strategies.

03

Retention Analysis

Highlights customers whose purchasing behaviour may require retention-focused actions.

04

Production Monitoring

Makes prediction behaviour observable after deployment.

17

Key Achievements

← Back to Portfolio