AI Knee MRI Detection | Dhritikamal Das
ðŸĶī Deep Learning · Computer Vision · MLOps

AI Knee Osteoarthritis Detection

An end-to-end computer vision system that converts knee X-ray images into automated osteoarthritis severity predictions using an Ordinal ResNet50 model, while combining explainable AI, API serving, monitoring, testing and containerized deployment.

60.93% Accuracy
59.89% Macro F1
77.07% QWK
11 Automated Tests
AI Knee Osteoarthritis Detection application

Turning medical imaging research into a usable AI system.

The initial challenge was not simply training an image classifier. The project needed to transform a trained deep learning model into an application capable of accepting unseen X-ray images, producing interpretable predictions and exposing the model through deployable interfaces.

The ML Challenge

Knee osteoarthritis severity is represented by ordered grades from 0 to 4. Treating these grades as completely independent categories ignores the ordinal relationship between severity levels.

The Engineering Challenge

A research model alone is difficult for a non-technical user to consume. The project therefore required an inference layer, explainability, API serving, UI deployment, testing, monitoring and reproducible packaging.

Ordinal deep learning instead of ordinary multiclass classification.

The final system uses an Ordinal ResNet50 model. Instead of directly treating Grade 0–4 as unrelated categories, the model learns four ordered thresholds.

Experiment 5 · Ordinal ResNet50

The deployed model uses ResNet50 as the visual feature extractor and an ordinal output formulation for the five osteoarthritis severity grades. The inference pipeline converts the four threshold probabilities into grade-level probabilities and selects the predicted severity level using the 0.5 ordinal threshold.

Architecture Ordinal ResNet50
Input 224 × 224 RGB
Classes 5
Thresholds 4
Threshold 0.50
Grade 0 No OA
→
Grade 1 Early changes
→
Grade 2 Moderate
→
Grade 3 Advanced
→
Grade 4 Severe

What the final model achieved.

Experiment 5 was evaluated using accuracy, macro F1 and Quadratic Weighted Kappa. QWK is particularly relevant for this problem because the target grades have an inherent order.

60.93%
Accuracy
Overall correct grade predictions
59.89%
Macro F1
Balanced performance across grades
77.07%
Quadratic Weighted Kappa
Ordinal agreement between predictions and labels

The model produces more than a class label.

A real local inference run produced a Grade 2 prediction with 84.28% confidence. The inference pipeline also returned four ordinal threshold probabilities and a complete probability distribution across all five grades.

Example inference
Grade 2
Confidence · 84.28%
Grade 0 0.64%
Grade 1 1.54%
Grade 2 84.28%
Grade 3 13.52%
Grade 4 0.03%

Four learned severity thresholds.

For the same example prediction, the model generated the following threshold probabilities.

Threshold Probability Interpretation
Grade â‰Ĩ 1 99.36% Strong evidence above Grade 0
Grade â‰Ĩ 2 97.83% Strong evidence of Grade 2 or higher
Grade â‰Ĩ 3 13.55% Lower probability of Grade 3 or higher
Grade â‰Ĩ 4 0.03% Very low probability of Grade 4

What changed after building the system?

The biggest outcome was not just a trained model. The project evolved from an ML experiment into a structured, testable and deployable AI application.

01

Research model → usable application

Users can upload a knee X-ray through a Streamlit interface and receive a severity prediction, confidence score and probability distribution.

02

Classification → ordinal reasoning

The final model uses four ordered thresholds to represent the relationship between Grade 0 and Grade 4 rather than ignoring severity ordering.

03

Black-box prediction → visual explanation

Grad-CAM generates grade-specific heatmaps and overlays, providing a visual interpretation of the regions influencing the prediction.

04

Python script → REST API

FastAPI exposes the inference pipeline through HTTP endpoints, making the model consumable by external applications.

05

Local setup → containerized services

Docker Compose packages the FastAPI backend and Streamlit application into reproducible services.

06

Manual checking → automated testing

The project now includes 11 automated tests covering API behavior, inference, model outputs and preprocessing.

07

Local model → version-verified artifact

The Hugging Face model is downloaded during inference and its SHA256 checksum is recorded for artifact verification.

08

Development → CI validation

GitHub Actions automatically executes the test suite on repository changes, helping identify environment-specific failures before deployment.

Making model decisions easier to inspect.

Grad-CAM was integrated into the inference workflow. For each prediction, the application can generate grade-specific heatmaps and overlays alongside the original image.

Original knee X-ray
Original X-ray
Grad-CAM heatmap
Model attention / Grad-CAM
Prediction interface
Prediction and probability output

From model artifact to production-style pipeline.

01 Configuration

YAML-based model and application configuration.

02 Inference

Standardized preprocessing and model prediction.

03 Explainability

Grad-CAM heatmaps and overlays.

04 Serving

FastAPI and Streamlit interfaces.

05 Deployment

Docker, Compose and CI validation.

Preparing the model for post-deployment monitoring.

Monitoring utilities were added to establish a foundation for tracking model behavior and detecting changes in incoming data.

Prediction Metrics Monitoring utilities for inference-level metrics.
Model Monitoring Components for tracking deployed model behavior.
Drift Detection Utility for identifying changes in input behavior.

Model artifact verification.

The inference pipeline records the SHA256 checksum of the downloaded model artifact. This provides a reproducible identifier for the exact model binary used during inference.

Experiment 5 · SHA256
f2200b43966dce1498e47ad6ee45cb35e5cec831246b7667323e16fd9d7e1667

Validating the system, not just the model.

The project contains automated tests for the API, inference pipeline, model construction and preprocessing. GitHub Actions runs the test suite automatically on repository changes.

11 Tests Automated pytest suite
API Tests Root, health and invalid-file behavior
Inference Tests Grade, confidence and probability validation
Model Tests Model creation and output shape
Preprocessing Image transformation validation
GitHub Actions Automated CI test execution

Two services, one reproducible deployment.

Docker Compose runs the FastAPI inference service and Streamlit interface as independent services.

FastAPI Port 8000 · REST inference API
Streamlit Port 8501 · Interactive application
Docker Compose Reproducible multi-service environment

Built across the ML engineering stack.

Python PyTorch Torchvision ResNet50 Ordinal Classification OpenCV Pillow Albumentations Grad-CAM FastAPI Streamlit Docker Docker Compose Pytest GitHub Actions Hugging Face Hub MLflow YAML

The project demonstrates more than model accuracy.

AI / ML Value

Built an ordinal computer vision model capable of predicting five ordered OA severity grades, achieving 60.93% accuracy, 59.89% macro F1 and 77.07% quadratic weighted kappa.

Engineering Value

Converted the model into a complete application with reusable inference code, FastAPI serving, Streamlit UI, Grad-CAM explainability, monitoring utilities, Docker deployment, automated tests and CI validation.

Reproducibility Value

Centralized model configuration, hosted the production artifact on Hugging Face and recorded its SHA256 checksum to make the deployed model version identifiable.

Product Value

Created a workflow where a user can upload an image and receive a structured prediction, confidence distribution and visual explanation through an accessible interface.

Important limitations.

⚠ Research / Educational Prototype

This project is an AI research and engineering prototype and is not a certified medical device. Predictions should not be used independently for diagnosis, treatment decisions or clinical decision-making. Grad-CAM visualizations represent model behavior and should not be interpreted as clinical evidence. A qualified healthcare professional should interpret medical imaging.

Explore the implementation.

The complete source code, deployment configuration, inference pipeline, tests and model artifact are available through the project resources below.

← Back to Portfolio