Location
Geographic or locality information used to capture differences in property value across areas.
An end-to-end machine learning application that predicts residential property prices and demonstrates the complete journey from data preprocessing and model development to API serving, containerization, Kubernetes orchestration and cloud deployment.
House prices depend on multiple factors such as location, property size, number of bedrooms, number of bathrooms, amenities and other property characteristics.
The objective of this project was to build a complete machine learning system capable of learning relationships between property attributes and their corresponding prices.
Instead of stopping at model training, the project was developed as an end-to-end production-style application. The final system exposes the trained model through FastAPI, packages the application using Docker, runs the application through Kubernetes and prepares the service for cloud deployment.
Property buyers, sellers and real-estate platforms need reliable estimates of property prices. Traditional price estimation can involve manual research, comparison with similar properties and subjective judgement.
The goal of this project was to create a machine learning solution that automatically estimates the expected price of a property based on its available characteristics.
Build a deployable machine learning API that accepts property characteristics as input and returns an estimated house price.
The machine learning pipeline uses property-level information to learn the relationship between property characteristics and selling price.
Geographic or locality information used to capture differences in property value across areas.
Represents the physical size of the property and its relationship with market price.
Number of bedrooms available in the property.
Number of bathrooms representing property capacity and potential value.
Additional property characteristics that may influence buyer demand and price.
The property price that the regression model learns to predict.
Raw machine learning data cannot always be directly supplied to a model. The preprocessing stage converts the source data into a consistent representation suitable for training.
The dataset is inspected for missing values, inconsistent records and unsuitable data types before model training.
Numerical and categorical variables are transformed into representations that machine learning algorithms can process.
Categorical information such as location is converted into numerical representations as required by the modelling pipeline.
Where applicable, numerical features are transformed so that models sensitive to feature magnitude can operate effectively.
Exploratory data analysis was used to understand the structure of the dataset before selecting the final model.
The analysis focused on relationships between property characteristics and price, distributions of numerical variables, categorical patterns and potential outliers.
Since house price prediction is a continuous-value prediction problem, regression algorithms were considered.
Used as a baseline model to establish a simple relationship between input features and property price.
An ensemble tree-based model capable of capturing nonlinear relationships between property characteristics.
Sequentially builds decision trees to improve prediction performance by learning from previous model errors.
A gradient-boosting implementation designed for strong predictive performance, regularization and efficient training on structured datasets.
Multiple regression approaches were considered instead of relying on a single algorithm from the beginning.
Tree-based ensemble models are particularly useful for this problem because house prices can have nonlinear relationships with variables such as size, location and amenities.
XGBoost was selected as the main model because gradient boosting can model complex feature interactions while providing strong predictive performance on structured datasets.
XGBoost regression model trained on the processed house-price dataset.
The regression models were evaluated using standard regression metrics to understand how accurately the predicted prices matched the actual target values.
Mean Absolute Error measures the average absolute difference between actual and predicted prices.
Root Mean Squared Error gives greater weight to larger prediction errors.
R² measures the proportion of target variation explained by the model.
Replace this section with the exact metrics from your final training run. Do not use estimated values.
Collect and inspect house-price data.
Clean, transform and encode input features.
Train and evaluate regression algorithms.
Select and use the final gradient-boosting model.
Expose the trained model through a REST API.
Package the API and dependencies into a container.
Deploy and manage application replicas.
Make the service accessible through the deployment environment.
The trained machine learning model is exposed through a FastAPI application.
Instead of requiring users to interact directly with the Python model, the API provides a clean interface where clients send property information and receive a prediction.
GET /docs POST /predict
FastAPI automatically provides interactive Swagger/OpenAPI
documentation through the /docs endpoint.
The API application is containerized using Docker so that the same application environment can be reproduced across development, testing and deployment environments.
Source Code
↓
Dockerfile
↓
Docker Image
↓
Container
↓
FastAPI Application
Kubernetes was used to orchestrate the containerized API. The application was deployed locally using Minikube during Kubernetes development and testing.
The Kubernetes Deployment manages the desired number of application replicas and ensures that the required pods remain available.
A Kubernetes Service provides a stable networking layer in front of the application pods.
Non-sensitive configuration can be separated from the application image through Kubernetes configuration resources.
Sensitive values such as Supabase credentials and API keys are handled through Kubernetes Secrets rather than being hard-coded into application source code.
Supabase is used as part of the application's backend and data infrastructure.
Credentials are supplied through environment variables instead of being embedded directly into the Python source code.
SUPABASE_URL SUPABASE_SERVICE_KEY API_KEY
These values are injected into the application environment through deployment configuration and Kubernetes Secrets.
The application was prepared for public deployment, allowing the FastAPI service to be accessed outside the local development environment.
The deployment process uses the containerized application together with environment variables for sensitive configuration.
User / Client
│
▼
Public API
│
▼
FastAPI
│
▼
Input Validation
│
▼
Preprocessing
│
▼
XGBoost Model
│
▼
House Price Prediction
│
▼
Supabase / Application Data
One of the important deployment considerations was preventing sensitive credentials from being committed directly into the source-code repository.
Local development configuration and production deployment configuration are separated from application source code.
Real credentials should never be committed to a public GitHub repository. Deployment platforms should provide sensitive values through their environment-variable or secret-management systems.
The application was tested at multiple levels to verify that the machine learning service continued to work after containerization and Kubernetes deployment.
Confirmed that the FastAPI application starts successfully inside the deployed container.
Verified that the automatically generated API documentation is accessible.
Tested prediction requests through the API.
The Kubernetes pods and service were verified during local Minikube deployment to confirm that the application was running and reachable.
A trained machine learning model needs to run inside an environment containing the correct Python dependencies and application configuration.
SolutionDocker was used to create a reproducible runtime environment.
Accessing an application deployed behind a Kubernetes Service requires understanding Kubernetes networking.
SolutionThe Kubernetes Service was tested locally through Minikube and the API endpoint was verified.
Supabase credentials cannot safely be stored directly inside application source code or a public repository.
SolutionCredentials were moved into environment-based configuration and Kubernetes Secrets.
Updating a Docker image requires ensuring that the Kubernetes Deployment references the intended image.
SolutionThe deployment was updated and the rollout status was verified before testing the running pods.
┌──────────────────────┐
│ Client │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ FastAPI │
│ API │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Input Validation & │
│ Preprocessing │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ XGBoost Model │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Predicted House │
│ Price │
└──────────────────────┘
┌──────────────────────────────────────────────┐
│ Deployment Layer │
│ │
│ Docker → Kubernetes → Service → Cloud │
│ │
└──────────────────────────────────────────────┘
┌──────────────────────────────────────────────┐
│ Configuration Layer │
│ │
│ Environment Variables + Kubernetes Secrets │
│ │
└──────────────────────────────────────────────┘
The project successfully progressed from a machine learning experiment into a deployable application.
Learned how preprocessing, feature engineering, model selection and evaluation affect regression performance.
Learned how to expose a trained machine learning model through a production-style REST API.
Learned how to package an ML application and its dependencies into a reproducible container.
Learned the fundamentals of Deployments, Pods, Services, ConfigMaps, Secrets and application rollouts.
Learned how application configuration and secrets must be handled differently between local development and cloud environments.
Learned that deploying an ML model involves much more than achieving good validation metrics.
This project demonstrates the complete lifecycle of a machine learning application — from real-estate data and model development to API serving, containerization, Kubernetes orchestration and cloud deployment.
The key focus was not simply predicting house prices, but understanding how a machine learning model can be transformed into a reliable, reproducible and deployable software system.