Pittsburgh, Pennsylvania
Projects

Skin Burn Detection and Classification with YOLOv5

image
March 30, 2021
Burn triage depends on visual assessment. Clinicians look at a wound and decide: send the patient home, schedule surgery, or call the burn unit. Studies show experienced clinicians disagree on burn depth in up to 30% of cases. The same wound gets different classifications depending on who is looking at it. This disagreement matters. Treatment for a first-degree burn is aloe vera and ibuprofen. Treatment for a third-degree burn is emergency surgery. The gap between those two decisions is enormous, and it hinges on a subjective visual call. The hypothesis behind this project was simple: a detection model trained on thousands of burn images might learn visual patterns that correlate with severity, producing a consistent baseline assessment. Not a replacement for the clinician. A second pair of eyes that never gets tired and never varies in its criteria. The model classifies burns into three classes, each with distinct clinical implications. Minor burns (Class 1) affect only the epidermis. Redness, mild swelling, pain. No blistering. These heal in 3 to 7 days without intervention. Think sunburn. The treatment is conservative and the prognosis is good. Moderate burns (Class 2) extend into the dermis. Blistering appears, pain is significant, and the wound bed is moist and pink. Healing takes 2 to 3 weeks. Some require grafting. Infection risk climbs once the skin barrier is broken, and poor early management can convert a partial-thickness injury into a full-thickness one. Severe burns (Class 3) destroy the full thickness of skin. The dermis is gone. In fourth-degree cases, damage reaches muscle or bone. These wounds may actually be painless because the nerve endings are destroyed. The appearance is white, waxy, leathery, or charred. These patients need emergency surgery. Every hour of delay increases the risk of sepsis and organ failure. The project started with 5,855 source images from a publicly available dermatological dataset on Kaggle. All annotations were created with bounding box labeling through Roboflow. Certified medical practitioners validated the severity classifications against established clinical criteria. The class distribution:
ClassCountPercentage
Minor (Class 1)2,40740.6%
Moderate (Class 2)2,60643.9%
Severe (Class 3)92215.5%
The most dangerous class is the least represented. This is not a collection artifact. Severe burns are rarer in datasets because patients with these injuries are often in acute care where research photography is not a priority. The imbalance mirrors clinical reality, but it creates a training problem: the model gets the fewest examples of the cases where it matters most. Ambiguous cases at the boundary between moderate and severe were flagged and adjudicated during annotation. This process itself reinforced the core challenge. If trained clinicians disagree on borderline cases, the ground truth labels carry noise that no architecture can fully resolve. 5,855 images is not enough. Augmentation expanded the dataset to 14,044 images. Each technique targeted a specific source of real-world variation. HSV jitter addressed lighting and skin tone variability. Hue shifted +/-15 degrees, saturation +/-40%, brightness +/-40%. Clinical photos come from smartphone cameras under fluorescent lights, operating rooms, natural daylight, dim emergency departments. The model needed to learn burn-specific color patterns rather than memorize the lighting of the training set. Geometric transforms included horizontal and vertical flips plus rotation of +/-20 degrees. Burn morphology is orientation-invariant. A burn on the left forearm is diagnostically identical to its mirror image. These transforms are safe in this domain, unlike chest X-rays where cardiac position matters. Mosaic tiling composites four training images into one frame. This forces the model to detect burns at varying scales and positions. Clinical photographs vary wildly in framing. A small burn in the corner of a wide-angle photo should be detected as reliably as one filling the entire frame. Mixup blending interpolates between two images and their labels at alpha 0.5. This smooths decision boundaries between classes. In a medical context, a model that outputs 0.6 for moderate and 0.4 for severe is more useful than one that outputs 0.99 for moderate. The uncertainty signal itself is informative. Copy-paste augmentation directly targeted the class imbalance. Severe burn regions were extracted from existing images and composited onto other backgrounds, oversampling the minority class without simple duplication. This preserves visual diversity while increasing the model's exposure to Class 3 presentations. As results later showed, it helped but did not fully solve the problem. YOLOv5s was chosen for a specific reason: any system intended for clinical use in resource-limited settings must run on commodity hardware. A model requiring a GPU cluster is useless in a rural clinic with a single laptop. YOLOv5s sits at the right tradeoff point. Small enough for edge deployment. Accurate enough to be useful. The architecture has 214 layers and 7,027,720 parameters, organized into three stages. Backbone: CSP-Darknet53. The feature extractor. Cross-Stage Partial connections split the feature map into two paths, one processed through a dense block and one passed through directly, then concatenate them. This cuts computation by roughly 20% compared to standard DenseNet connections while preserving gradient flow. The 53 convolutional layers with residual connections provide enough depth to capture the hierarchy of burn features: low-level texture (blister sheen, eschar dryness), mid-level structure (wound boundary shape, color gradients), and high-level semantics (overall wound morphology). Neck: SPPF + CSP-PAN. Spatial Pyramid Pooling Fast applies max pooling at multiple kernel sizes, capturing context at different scales without fixed input dimensions. The Cross-Stage Partial Path Aggregation Network fuses features bidirectionally between resolution levels. High-resolution features capture fine detail. Low-resolution features capture semantic context. Both matter for burns, where the texture of the wound center and the transition at the wound margin both carry diagnostic weight. Detection heads operate at three strides: 8, 16, and 32. Stride 8 handles small burns on an 80x80 grid. Stride 16 covers medium detections on a 40x40 grid. Stride 32 captures large burns on a 20x20 grid with the richest semantic information. Each head predicts bounding box coordinates, objectness scores, and class probabilities. This multi-scale approach means burns of all sizes, from a small blister to a full-limb injury, get detected in a single forward pass. The model was initialized with yolov5s.pt weights pretrained on COCO. COCO contains no burn images, but the low-level features learned from 80 object categories (edges, textures, color gradients) transfer well to medical imaging. This gave the model a head start and reduced the risk of overfitting on limited medical data. Input resolution was 640x640 pixels. The training configuration:
ParameterValue
Batch size64
Epochs40
OptimizerSGD
Initial learning rate0.01
Momentum0.937
Weight decay0.0005
LR scheduleCosine
Warmup3 epochs (linear)
The cosine schedule drops the learning rate along a half-cosine curve, allowing aggressive learning early and fine-grained optimization near convergence. The 3-epoch warmup prevents large initial gradients from destabilizing the pretrained weights. The loss function combines three terms. CIoU box loss measures bounding box quality, incorporating overlap area, center point distance, and aspect ratio. BCE objectness loss trains the model to distinguish burn regions from background. BCE classification loss trains severity assignment. The three are weighted and summed. Evaluation used COCO-style metrics: mAP@0.5 and mAP@0.5:0.95. Aggregate metrics can hide clinically dangerous failures. The per-class breakdown told the real story. Class 1 (Minor) performed best. The visual signatures of minor burns, uniform redness, intact skin, no blistering, are relatively unambiguous. High precision and recall. Class 2 (Moderate) showed solid performance with some confusion at boundaries with both Class 1 and Class 3. The confusion matrix revealed Class 2 detections confused with background. This boundary ambiguity reflects genuine diagnostic difficulty. The transition from superficial to deep partial-thickness is a continuum, not a line. Class 3 (Severe) had the weakest recall. Despite copy-paste augmentation, 15.5% representation was not enough. Full-thickness burns vary enormously depending on the cause (thermal, chemical, electrical), time since injury, and eschar formation. The model captured common presentations but missed atypical ones. The confusion matrix revealed two important patterns. First, the model tended to underclassify, predicting moderate when the true label was severe. In medical terms, this is the dangerous direction. A patient who needs emergency surgery might receive conservative treatment instead. Second, the absence of negative samples in training caused false positives on healthy skin, particularly skin with redness from non-burn causes. The model has no concept of "not a burn." It must assign every detection to one of three classes. The confidence threshold was set at 0.1, far below the typical 0.25-0.5 range. This was intentional. In medical screening, a false negative (missing a severe burn) costs more than a false positive (flagging something for a closer look). The low threshold maximizes sensitivity at the expense of specificity. Severe burns are underrepresented. Copy-paste augmentation rearranges existing patterns in new spatial configurations. It does not introduce genuinely new burn morphologies. The model's recall for the most critical class reflects this gap. When it says severe, pay attention. When it says not severe, verify independently. No negative samples. The model has never seen healthy skin. It produces false positives on skin with redness from other causes: rashes, allergic reactions, pressure marks, even warm lighting on normal skin. This would erode clinician trust quickly in a real deployment. Skin tone diversity is uncertain. Dermatological datasets have historically overrepresented lighter skin tones. Erythema, a key severity indicator, is less visible on darker skin. If the training data does not adequately represent the full melanin spectrum, the model may perform poorly on the populations most likely to lack access to specialist burn care. This is an equity problem, not just a technical one. Images are incomplete information. Clinical burn assessment includes tactile evaluation (tissue turgor, capillary refill, pinprick sensation), patient history (cause, duration, first aid received), systemic indicators (fluid status, airway involvement), and wound evolution over 48-72 hours. A camera captures none of this. This is a proof of concept, not a medical device. Regulatory pathways (FDA SaMD, EU MDR) require prospective clinical validation, formal risk analysis, quality management systems, and post-market surveillance. The technical metrics reported here are a starting point, not sufficient evidence for clinical deployment. The highest-priority improvement is negative sample integration. Images of healthy skin and conditions that mimic burns (contact dermatitis, cellulitis, psoriasis) would reduce false positives and give the model the ability to say "this is not a burn." More severe burn data is needed. Partnerships with burn centers serving diverse populations across multiple regions could address both the class imbalance and the skin tone representation gap. Performance should be evaluated and reported per Fitzpatrick skin type. Mobile deployment is feasible. YOLOv5s at 7M parameters is already lightweight enough for smartphone inference via TensorFlow Lite or ONNX Runtime. A mobile app for first responders or rural healthcare workers could provide immediate severity assessment where specialist expertise is hours away. Clinical validation is the essential gate. A prospective multi-site study comparing model-assisted triage against standard-of-care triage, measuring time to appropriate treatment and patient outcomes. Retrospective performance on a curated dataset, however strong, is not enough. This project placed as a Top 5 finalist at the TiE Global Hackathon, selected from a field of over 650 competing teams. YOLOv5 (Ultralytics), Python, PyTorch, Roboflow (annotation and augmentation), OpenCV, Kaggle (data source).