End-to-End Machine Learning • Customer Analytics • FastAPI • Streamlit • MLflow • Docker • CI/CD
Developed an end-to-end Machine Learning Operations (MLOps) solution for customer segmentation using transactional retail data. The system converts raw customer purchase history into meaningful behavioural segments using RFM analysis and K-Means clustering.
The project extends beyond model development into a production-oriented workflow that includes feature engineering, model serialization, real-time prediction, prediction logging, monitoring, data-drift detection, automated retraining, CI/CD, API deployment, and an interactive Streamlit dashboard.
Retail businesses generate large volumes of transactional data, but raw transaction records do not directly reveal which customers are valuable, newly acquired, disengaged, or at risk of becoming inactive.
Traditional customer segmentation can require significant manual analysis and may be difficult to maintain as new transactions arrive.
Identify customers generating significant monetary value and purchasing frequently.
Understand how recently and frequently customers interact with the business.
Identify behavioural patterns associated with inactive or at-risk customers.
Automate customer segmentation using machine learning.
Transform transactional data into meaningful RFM features.
Deploy a real-time customer segmentation API.
Provide an interactive analytics dashboard.
Log production predictions for monitoring.
Detect production data drift.
Support automated model retraining.
Implement a reproducible MLOps deployment workflow.
The project uses the Online Retail transactional dataset. Each record represents a retail transaction containing information such as invoice number, product, quantity, unit price, customer identifier, country, and transaction date.
Customer transactions were transformed into behavioural features using RFM analysis. Additional derived features were incorporated to provide a richer representation of customer purchasing behaviour.
Number of days since the customer's most recent purchase.
Recency
Number of purchases or transactions associated with the customer.
Frequency
Total monetary value generated by the customer.
Monetary
Average monetary value associated with each customer transaction.
Monetary / Frequency
Derived customer-level value used to capture purchasing contribution.
Frequency × AverageOrderValue
K-Means clustering was selected to group customers according to similarities in their purchasing behaviour.
The model groups customers into clusters based on their engineered behavioural features. The resulting cluster assignments are mapped to business-oriented customer segments.
Customers demonstrating strong purchasing activity and high monetary contribution.
Customers with relatively recent purchasing activity and developing engagement.
Customers showing behavioural patterns that may indicate declining engagement.
Customers with low or historically distant purchasing activity.
The project was designed as a complete machine learning lifecycle rather than a standalone notebook model.
A business-facing Streamlit application was developed to make the machine learning system accessible without requiring users to interact directly with the underlying Python code.
Users can enter customer RFM information and receive a predicted customer segment.
Converts model output into an interpretable customer profile.
Provides segment-oriented actions for customer engagement and retention.
Displays historical prediction activity using logged production requests.
Tracks prediction counts and cluster distribution from production logs.
Compares production feature distributions with training data to identify potential drift.
A FastAPI backend provides programmatic access to the trained segmentation model and exposes endpoints for prediction, health checks, monitoring, and drift detection.
GET /
API status and welcome response
GET /health
Service health check
POST /predict
Predict customer segment
GET /monitor
Production monitoring metrics
GET /drift
Data drift analysis
Every prediction request can be logged with the input features, engineered values, timestamp, predicted cluster, and customer segment. These logs provide the foundation for production monitoring.
Production feature statistics are compared against the training dataset to identify changes in customer behaviour. The drift analysis examines the same engineered features used by the production model.
The monitoring workflow is designed to identify significant changes between the training and production feature distributions. When the configured drift condition is satisfied, the retraining workflow can be triggered to refresh the segmentation model.
This closes the machine learning lifecycle by connecting production monitoring with model maintenance.
The application is structured for cloud deployment with separate user-facing dashboard and API components. The FastAPI service provides the model inference layer, while Streamlit provides the interactive business interface.
Interactive customer segmentation and analytics dashboard.
REST API for model inference, monitoring, and drift detection.
Reproducible application environment for deployment.
Automated testing and deployment workflow.
Reduces the need for manual customer grouping.
Provides interpretable customer groups for targeted marketing strategies.
Highlights customers whose purchasing behaviour may require retention-focused actions.
Makes prediction behaviour observable after deployment.