AI Knee Osteoarthritis Detection
An end-to-end computer vision system that converts knee X-ray images into automated osteoarthritis severity predictions using an Ordinal ResNet50 model, while combining explainable AI, API serving, monitoring, testing and containerized deployment.
Turning medical imaging research into a usable AI system.
The initial challenge was not simply training an image classifier. The project needed to transform a trained deep learning model into an application capable of accepting unseen X-ray images, producing interpretable predictions and exposing the model through deployable interfaces.
The ML Challenge
Knee osteoarthritis severity is represented by ordered grades from 0 to 4. Treating these grades as completely independent categories ignores the ordinal relationship between severity levels.
The Engineering Challenge
A research model alone is difficult for a non-technical user to consume. The project therefore required an inference layer, explainability, API serving, UI deployment, testing, monitoring and reproducible packaging.
Ordinal deep learning instead of ordinary multiclass classification.
The final system uses an Ordinal ResNet50 model. Instead of directly treating Grade 0â4 as unrelated categories, the model learns four ordered thresholds.
Experiment 5 · Ordinal ResNet50
The deployed model uses ResNet50 as the visual feature extractor and an ordinal output formulation for the five osteoarthritis severity grades. The inference pipeline converts the four threshold probabilities into grade-level probabilities and selects the predicted severity level using the 0.5 ordinal threshold.
What the final model achieved.
Experiment 5 was evaluated using accuracy, macro F1 and Quadratic Weighted Kappa. QWK is particularly relevant for this problem because the target grades have an inherent order.
The model produces more than a class label.
A real local inference run produced a Grade 2 prediction with 84.28% confidence. The inference pipeline also returned four ordinal threshold probabilities and a complete probability distribution across all five grades.
Four learned severity thresholds.
For the same example prediction, the model generated the following threshold probabilities.
| Threshold | Probability | Interpretation |
|---|---|---|
| Grade âĨ 1 | 99.36% | Strong evidence above Grade 0 |
| Grade âĨ 2 | 97.83% | Strong evidence of Grade 2 or higher |
| Grade âĨ 3 | 13.55% | Lower probability of Grade 3 or higher |
| Grade âĨ 4 | 0.03% | Very low probability of Grade 4 |
What changed after building the system?
The biggest outcome was not just a trained model. The project evolved from an ML experiment into a structured, testable and deployable AI application.
Research model â usable application
Users can upload a knee X-ray through a Streamlit interface and receive a severity prediction, confidence score and probability distribution.
Classification â ordinal reasoning
The final model uses four ordered thresholds to represent the relationship between Grade 0 and Grade 4 rather than ignoring severity ordering.
Black-box prediction â visual explanation
Grad-CAM generates grade-specific heatmaps and overlays, providing a visual interpretation of the regions influencing the prediction.
Python script â REST API
FastAPI exposes the inference pipeline through HTTP endpoints, making the model consumable by external applications.
Local setup â containerized services
Docker Compose packages the FastAPI backend and Streamlit application into reproducible services.
Manual checking â automated testing
The project now includes 11 automated tests covering API behavior, inference, model outputs and preprocessing.
Local model â version-verified artifact
The Hugging Face model is downloaded during inference and its SHA256 checksum is recorded for artifact verification.
Development â CI validation
GitHub Actions automatically executes the test suite on repository changes, helping identify environment-specific failures before deployment.
Making model decisions easier to inspect.
Grad-CAM was integrated into the inference workflow. For each prediction, the application can generate grade-specific heatmaps and overlays alongside the original image.
From model artifact to production-style pipeline.
YAML-based model and application configuration.
Standardized preprocessing and model prediction.
Grad-CAM heatmaps and overlays.
FastAPI and Streamlit interfaces.
Docker, Compose and CI validation.
Preparing the model for post-deployment monitoring.
Monitoring utilities were added to establish a foundation for tracking model behavior and detecting changes in incoming data.
Model artifact verification.
The inference pipeline records the SHA256 checksum of the downloaded model artifact. This provides a reproducible identifier for the exact model binary used during inference.
Validating the system, not just the model.
The project contains automated tests for the API, inference pipeline, model construction and preprocessing. GitHub Actions runs the test suite automatically on repository changes.
Two services, one reproducible deployment.
Docker Compose runs the FastAPI inference service and Streamlit interface as independent services.
Built across the ML engineering stack.
The project demonstrates more than model accuracy.
AI / ML Value
Built an ordinal computer vision model capable of predicting five ordered OA severity grades, achieving 60.93% accuracy, 59.89% macro F1 and 77.07% quadratic weighted kappa.
Engineering Value
Converted the model into a complete application with reusable inference code, FastAPI serving, Streamlit UI, Grad-CAM explainability, monitoring utilities, Docker deployment, automated tests and CI validation.
Reproducibility Value
Centralized model configuration, hosted the production artifact on Hugging Face and recorded its SHA256 checksum to make the deployed model version identifiable.
Product Value
Created a workflow where a user can upload an image and receive a structured prediction, confidence distribution and visual explanation through an accessible interface.
Important limitations.
This project is an AI research and engineering prototype and is not a certified medical device. Predictions should not be used independently for diagnosis, treatment decisions or clinical decision-making. Grad-CAM visualizations represent model behavior and should not be interpreted as clinical evidence. A qualified healthcare professional should interpret medical imaging.
Explore the implementation.
The complete source code, deployment configuration, inference pipeline, tests and model artifact are available through the project resources below.