# How does the training of a object detection dataset happen, with segmentation datasets, and how are the transformation handled

**URL:** <https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351>\
**Category:** YOLO\
**Tags:** question, code\
**Created:** [August 15, 2025, 10:18am UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351 "2025-08-15T10:18:48Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kallinteris-Andreas](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/kallinteris-andreas/32/838_2.png) [@Kallinteris-Andreas](https://community.ultralytics.com/u/Kallinteris-Andreas)\
**Post date:** [August 15, 2025, 10:18am UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/1 "2025-08-15T10:18:48Z")

</div>

In Ultralytics HUB the [COCO dataset](https://docs.ultralytics.com/datasets/detect/coco/) includes segmentation polygons, not bounding boxes, but it can be used to train object detection dataset

1. are the polygon masks simply converted to bounding boxes during training
2. how are the Spatial Transforms `Albumentations` handled such as arbitrary rotations handled, is the bounding box used for training based on post transformer segmentation mask?

---

<div class="post-metadata">

**Author:** ![Toxite](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/toxite/32/123_2.png) [@Toxite](https://community.ultralytics.com/u/Toxite)\
**Post date:** [August 15, 2025, 10:56pm UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/2 "2025-08-15T22:56:37Z")

</div>

1. They are converted to bounding boxes
2. Ultralytics’ native affine transform augmentations are performed using polygons if using segmentation dataset for object detection:  
[ultralytics/ultralytics/data/augment.py at 23d79250e7420945792362f07b8d818320fdce49 · ultralytics/ultralytics · GitHub](https://github.com/ultralytics/ultralytics/blob/23d79250e7420945792362f07b8d818320fdce49/ultralytics/data/augment.py#L1057-L1061)  
But `Albumentations` uses bounding boxes. It doesn’t support polygons.

---

<div class="post-metadata">

**Author:** ![Kallinteris-Andreas](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/kallinteris-andreas/32/838_2.png) [@Kallinteris-Andreas](https://community.ultralytics.com/u/Kallinteris-Andreas)\
**Post date:** [August 18, 2025, 10:06am UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/3 "2025-08-18T10:06:43Z")

</div>

If Ultralytics’ affine transform is used on a dataset that contains segmentation masks, which is used to train an object detection model, is the conversion from segmentation masks to bounding boxes happening (1) before the affine transform, (2) after the affine transform

This is relevant because it is (2) then the bounding box could potentially be more accurate

note: btw `Albumentations` does support masks [Targets by Transform](https://albumentations.ai/docs/reference/supported-targets-by-transform/), though I think it is on dense masks, not polygon masks

Thanks!

---

<div class="post-metadata">

**Author:** ![BurhanQ](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/burhanq/32/7_2.png) [@BurhanQ](https://community.ultralytics.com/u/BurhanQ)\
**Post date:** [August 18, 2025, 10:52am UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/4 "2025-08-18T10:52:43Z")

</div>

> [@Kallinteris-Andreas](#):
>
> on a dataset that contains segmentation masks, which is used to train an object detection model

Just to help clarify (because I’m confused), are you asking about using segmentation annotations to train a standard “detect” (bounding boxes) model? I presume you’re asking about using segmentation annotations to train a segmentation model, but I want to be 100% certain in my understanding.

---

<div class="post-metadata">

**Author:** ![Kallinteris-Andreas](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/kallinteris-andreas/32/838_2.png) [@Kallinteris-Andreas](https://community.ultralytics.com/u/Kallinteris-Andreas)\
**Post date:** [August 18, 2025, 11:04am UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/5 "2025-08-18T11:04:32Z")

</div>

I am asking about training an object detection model with segmentation annotations

---

<div class="post-metadata">

**Author:** ![BurhanQ](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/burhanq/32/7_2.png) [@BurhanQ](https://community.ultralytics.com/u/BurhanQ)\
**Post date:** [August 18, 2025, 12:01pm UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/6 "2025-08-18T12:01:18Z")

</div>

See the code here:

> <https://github.com/ultralytics/ultralytics/blob/23d79250e7420945792362f07b8d818320fdce49/ultralytics/data/augment.py#L1218-L1252>

---

<div class="post-metadata">

**Author:** ![Toxite](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/toxite/32/123_2.png) [@Toxite](https://community.ultralytics.com/u/Toxite)\
**Post date:** [August 18, 2025, 12:18pm UTC](https://community.ultralytics.com/t/how-does-the-training-of-a-object-detection-dataset-happen-with-segmentation-datasets-and-how-are-the-transformation-handled/1351/7 "2025-08-18T12:18:31Z")

</div>

> [@Kallinteris-Andreas](#):
>
> If Ultralytics’ affine transform is used on a dataset that contains segmentation masks, which is used to train an object detection model, is the conversion from segmentation masks to bounding boxes happening (1) before the affine transform, (2) after the affine transform

After the affine transform. And it’s like that to get accurate boxes like you mentioned.
