House Price Prediction using Machine Learning, Docker & Kubernetes

An end-to-end machine learning application that predicts residential property prices and demonstrates the complete journey from data preprocessing and model development to API serving, containerization, Kubernetes orchestration and cloud deployment.

Machine Learning XGBoost FastAPI Docker Kubernetes Minikube Supabase Render
House Price Prediction Machine Learning Project

01. Project Overview

House prices depend on multiple factors such as location, property size, number of bedrooms, number of bathrooms, amenities and other property characteristics.

The objective of this project was to build a complete machine learning system capable of learning relationships between property attributes and their corresponding prices.

Instead of stopping at model training, the project was developed as an end-to-end production-style application. The final system exposes the trained model through FastAPI, packages the application using Docker, runs the application through Kubernetes and prepares the service for cloud deployment.

02. Problem Statement

Property buyers, sellers and real-estate platforms need reliable estimates of property prices. Traditional price estimation can involve manual research, comparison with similar properties and subjective judgement.

The goal of this project was to create a machine learning solution that automatically estimates the expected price of a property based on its available characteristics.

Objective

Build a deployable machine learning API that accepts property characteristics as input and returns an estimated house price.

03. Project Goals

04. Data & Features

The machine learning pipeline uses property-level information to learn the relationship between property characteristics and selling price.

Location

Geographic or locality information used to capture differences in property value across areas.

Property Size

Represents the physical size of the property and its relationship with market price.

Bedrooms

Number of bedrooms available in the property.

Bathrooms

Number of bathrooms representing property capacity and potential value.

Amenities

Additional property characteristics that may influence buyer demand and price.

Target Variable

The property price that the regression model learns to predict.

05. Data Preprocessing

Raw machine learning data cannot always be directly supplied to a model. The preprocessing stage converts the source data into a consistent representation suitable for training.

Data Cleaning

The dataset is inspected for missing values, inconsistent records and unsuitable data types before model training.

Feature Transformation

Numerical and categorical variables are transformed into representations that machine learning algorithms can process.

Encoding

Categorical information such as location is converted into numerical representations as required by the modelling pipeline.

Feature Scaling

Where applicable, numerical features are transformed so that models sensitive to feature magnitude can operate effectively.

06. Exploratory Data Analysis

Exploratory data analysis was used to understand the structure of the dataset before selecting the final model.

The analysis focused on relationships between property characteristics and price, distributions of numerical variables, categorical patterns and potential outliers.

Key EDA Questions
  • Which property characteristics have the strongest relationship with price?
  • How does property size affect the target variable?
  • Are there significant differences between locations?
  • Are there extreme values that could affect model training?
  • Which features should be transformed or encoded?

07. Machine Learning Approach

Since house price prediction is a continuous-value prediction problem, regression algorithms were considered.

Linear Regression

Used as a baseline model to establish a simple relationship between input features and property price.

Random Forest

An ensemble tree-based model capable of capturing nonlinear relationships between property characteristics.

Gradient Boosting

Sequentially builds decision trees to improve prediction performance by learning from previous model errors.

XGBoost

A gradient-boosting implementation designed for strong predictive performance, regularization and efficient training on structured datasets.

08. Model Selection

Multiple regression approaches were considered instead of relying on a single algorithm from the beginning.

Tree-based ensemble models are particularly useful for this problem because house prices can have nonlinear relationships with variables such as size, location and amenities.

XGBoost was selected as the main model because gradient boosting can model complex feature interactions while providing strong predictive performance on structured datasets.

Final Model

XGBoost regression model trained on the processed house-price dataset.

09. Model Evaluation

The regression models were evaluated using standard regression metrics to understand how accurately the predicted prices matched the actual target values.

MAE

Mean Absolute Error measures the average absolute difference between actual and predicted prices.

RMSE

Root Mean Squared Error gives greater weight to larger prediction errors.

R² Score

R² measures the proportion of target variation explained by the model.

Replace this section with the exact metrics from your final training run. Do not use estimated values.

10. End-to-End Machine Learning Pipeline

01

Data

Collect and inspect house-price data.

↓
02

Preprocessing

Clean, transform and encode input features.

↓
03

Training

Train and evaluate regression algorithms.

↓
04

XGBoost

Select and use the final gradient-boosting model.

↓
05

FastAPI

Expose the trained model through a REST API.

↓
06

Docker

Package the API and dependencies into a container.

↓
07

Kubernetes

Deploy and manage application replicas.

↓
08

Cloud Deployment

Make the service accessible through the deployment environment.

11. FastAPI Model Serving

The trained machine learning model is exposed through a FastAPI application.

Instead of requiring users to interact directly with the Python model, the API provides a clean interface where clients send property information and receive a prediction.

API Flow

  1. Client sends property information.
  2. FastAPI validates the request.
  3. Input features are transformed using the required preprocessing logic.
  4. XGBoost generates the prediction.
  5. The API returns the predicted house price.

API Endpoints

GET  /docs
POST /predict

FastAPI automatically provides interactive Swagger/OpenAPI documentation through the /docs endpoint.

12. Docker Containerization

The API application is containerized using Docker so that the same application environment can be reproduced across development, testing and deployment environments.

Why Docker?

Container Flow

Source Code
     ↓
Dockerfile
     ↓
Docker Image
     ↓
Container
     ↓
FastAPI Application

13. Kubernetes Deployment

Kubernetes was used to orchestrate the containerized API. The application was deployed locally using Minikube during Kubernetes development and testing.

Deployment

The Kubernetes Deployment manages the desired number of application replicas and ensures that the required pods remain available.

Service

A Kubernetes Service provides a stable networking layer in front of the application pods.

ConfigMap

Non-sensitive configuration can be separated from the application image through Kubernetes configuration resources.

Secrets

Sensitive values such as Supabase credentials and API keys are handled through Kubernetes Secrets rather than being hard-coded into application source code.

Kubernetes Resources
  • Deployment
  • Service
  • ConfigMap
  • Secret

14. Supabase Integration

Supabase is used as part of the application's backend and data infrastructure.

Credentials are supplied through environment variables instead of being embedded directly into the Python source code.

SUPABASE_URL
SUPABASE_SERVICE_KEY
API_KEY

These values are injected into the application environment through deployment configuration and Kubernetes Secrets.

15. Cloud Deployment

The application was prepared for public deployment, allowing the FastAPI service to be accessed outside the local development environment.

The deployment process uses the containerized application together with environment variables for sensitive configuration.

Deployment Architecture

User / Client
      │
      ▼
Public API
      │
      ▼
FastAPI
      │
      ▼
Input Validation
      │
      ▼
Preprocessing
      │
      ▼
XGBoost Model
      │
      ▼
House Price Prediction
      │
      ▼
Supabase / Application Data

16. Security & Configuration Management

One of the important deployment considerations was preventing sensitive credentials from being committed directly into the source-code repository.

Local development configuration and production deployment configuration are separated from application source code.

Credentials handled securely

Security Practice

Real credentials should never be committed to a public GitHub repository. Deployment platforms should provide sensitive values through their environment-variable or secret-management systems.

17. API & Deployment Testing

The application was tested at multiple levels to verify that the machine learning service continued to work after containerization and Kubernetes deployment.

API Health

Confirmed that the FastAPI application starts successfully inside the deployed container.

Swagger UI

Verified that the automatically generated API documentation is accessible.

Prediction

Tested prediction requests through the API.

Kubernetes Verification

The Kubernetes pods and service were verified during local Minikube deployment to confirm that the application was running and reachable.

18. Challenges & Solutions

Challenge 01 — Model Deployment

A trained machine learning model needs to run inside an environment containing the correct Python dependencies and application configuration.

Solution

Docker was used to create a reproducible runtime environment.

Challenge 02 — Kubernetes Networking

Accessing an application deployed behind a Kubernetes Service requires understanding Kubernetes networking.

Solution

The Kubernetes Service was tested locally through Minikube and the API endpoint was verified.

Challenge 03 — Secrets

Supabase credentials cannot safely be stored directly inside application source code or a public repository.

Solution

Credentials were moved into environment-based configuration and Kubernetes Secrets.

Challenge 04 — Deployment Consistency

Updating a Docker image requires ensuring that the Kubernetes Deployment references the intended image.

Solution

The deployment was updated and the rollout status was verified before testing the running pods.

19. Technology Stack

Python Pandas NumPy Scikit-learn XGBoost FastAPI Docker Kubernetes Minikube Supabase Render Git GitHub

20. System Architecture

                     ┌──────────────────────┐
                     │        Client        │
                     └──────────┬───────────┘
                                │
                                ▼
                     ┌──────────────────────┐
                     │      FastAPI         │
                     │        API           │
                     └──────────┬───────────┘
                                │
                                ▼
                     ┌──────────────────────┐
                     │ Input Validation &   │
                     │    Preprocessing     │
                     └──────────┬───────────┘
                                │
                                ▼
                     ┌──────────────────────┐
                     │     XGBoost Model    │
                     └──────────┬───────────┘
                                │
                                ▼
                     ┌──────────────────────┐
                     │ Predicted House      │
                     │       Price          │
                     └──────────────────────┘


        ┌──────────────────────────────────────────────┐
        │              Deployment Layer                │
        │                                              │
        │ Docker → Kubernetes → Service → Cloud       │
        │                                              │
        └──────────────────────────────────────────────┘


        ┌──────────────────────────────────────────────┐
        │            Configuration Layer                │
        │                                              │
        │ Environment Variables + Kubernetes Secrets  │
        │                                              │
        └──────────────────────────────────────────────┘

21. Results

The project successfully progressed from a machine learning experiment into a deployable application.

22. Key Learnings

Machine Learning

Learned how preprocessing, feature engineering, model selection and evaluation affect regression performance.

Model Serving

Learned how to expose a trained machine learning model through a production-style REST API.

Docker

Learned how to package an ML application and its dependencies into a reproducible container.

Kubernetes

Learned the fundamentals of Deployments, Pods, Services, ConfigMaps, Secrets and application rollouts.

Cloud Deployment

Learned how application configuration and secrets must be handled differently between local development and cloud environments.

Production Thinking

Learned that deploying an ML model involves much more than achieving good validation metrics.

23. Future Improvements

24. Final Takeaway

This project demonstrates the complete lifecycle of a machine learning application — from real-estate data and model development to API serving, containerization, Kubernetes orchestration and cloud deployment.

The key focus was not simply predicting house prices, but understanding how a machine learning model can be transformed into a reliable, reproducible and deployable software system.

Data → Machine Learning → API → Docker → Kubernetes → Cloud