# Best approach to count stacked cardboard boxes using CCTV (2D RGB only)

**URL:** <https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581>\
**Category:** YOLO\
**Tags:** discussion\
**Created:** [October 28, 2025, 9:21am UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581 "2025-10-28T09:21:20Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Muhammad\_Fhadli](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/muhammad_fhadli/32/1090_2.png) [@Muhammad\_Fhadli](https://community.ultralytics.com/u/Muhammad_Fhadli)\
**Post date:** [October 28, 2025, 9:21am UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581/1 "2025-10-28T09:21:20Z")

</div>

Hi everyone 👋

I’m working on a real-world counting problem and would love to get some advice or ideas from the community.

I need to **count the number of cardboard boxes** in a warehouse — similar to this example image:

 ![image](https://us1.discourse-cdn.com/flex001/uploads/ultralytics1/original/2X/a/ab617e21a0e0dc329d78000ce9982b6e7a0c299b.jpeg)

#### My setup:

- Only one **CCTV camera (2D RGB)** available — no depth or stereo sensors.

- Boxes are **stacked tightly** and often **partially occluded**.

- The camera is **fixed** , so the viewing angle doesn’t change.

#### What I’ve tried / considered:

- **Object detection (YOLOv11 / OBB):** struggles with overlapping boxes.

- **Instance segmentation (YOLOv11-seg or SAM):** works better, but still has many false positives and under-segmentation (some clusters are merged).

- **Counting by area or volume estimation:** not accurate due to perspective distortion.

#### My questions:

1. Is **instance segmentation** still the best approach for this case, or is there a more robust method to handle heavy occlusion?

2. Are there any **recommended post-processing steps** (e.g., edge-based mask refinement or geometric heuristics) to split merged boxes?

3. Would **perspective correction or homography calibration** help improve segmentation accuracy?

4. Any **best practices** for training YOLOv11-seg specifically for stacked box scenarios?

I’m open to any suggestions — pipeline design, dataset tips, or even loss function tweaks that could improve instance separation.

Thanks in advance 🙏

---

<div class="post-metadata">

**Author:** ![BurhanQ](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/burhanq/32/7_2.png) [@BurhanQ](https://community.ultralytics.com/u/BurhanQ)\
**Post date:** [October 28, 2025, 12:48pm UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581/2 "2025-10-28T12:48:24Z")

</div>

Questions to help get to an answer:

1. Are the boxes generally all the same size like shown in the image?
2. I understand the aim is to count the boxes, but what’s the overall goal? Where does the box count data get sent to?
3. Will there be multiple pallets (like your example image) or a single pallet in the frame? If it’s multiple, can it be changed to be single?

---

<div class="post-metadata">

**Author:** ![BurhanQ](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/burhanq/32/7_2.png) [@BurhanQ](https://community.ultralytics.com/u/BurhanQ)\
**Post date:** [October 28, 2025, 12:51pm UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581/3 "2025-10-28T12:51:40Z")

</div>

Also, FWIW, I using a [YOLOE model](https://docs.ultralytics.com/models/yoloe/), I was able to get this result

 ![image](https://us1.discourse-cdn.com/flex001/uploads/ultralytics1/original/2X/3/35debc510fdbc598bb371a17932e9c194417908d.jpeg)

Here’s the exact code I used:

```python
from pathlib import Path

from ultralytics import YOLOE

p = Path.home() / "Downloads"
f = p / "boxes.jpg"

model = YOLOE("yoloe-11l-seg.pt") # Large segmentation model
names = [
    "box", 
    "bin", 
    "handtruck", 
    "person", 
    "garage door", 
    "forklift", 
    "pallet",
    "",
] # other classes that might be in the image to help separate detections
model.set_classes(names, model.get_text_pe(names))

results = model.predict(f, iou=0.11, conf=0.06)
results[0].show(masks=True, labels=False)

```

---

<div class="post-metadata">

**Author:** ![Ryan](https://avatars.discourse-cdn.com/v4/letter/r/a8b319/32.png) [@Ryan](https://community.ultralytics.com/u/Ryan)\
**Post date:** [June 15, 2026, 8:39am UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581/4 "2026-06-15T08:39:39Z")

</div>

Hello, have you resolved the issue you mentioned?

Even if box detection is 100%, how did you handle the counting in cases where parts are obscured, only one side is visible, or the image is cropped?

---

<div class="post-metadata">

**Author:** ![pderrenger](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/pderrenger/32/73_2.png) [@pderrenger](https://community.ultralytics.com/u/pderrenger)\
**Post date:** [June 16, 2026, 1:32am UTC](https://community.ultralytics.com/t/best-approach-to-count-stacked-cardboard-boxes-using-cctv-2d-rgb-only/1581/5 "2026-06-16T01:32:03Z")

</div>

Ryan’s question gets to the core of it: with one fixed 2D RGB view, even perfect Ultralytics YOLO detection can only count boxes that are at least partly visible. Fully hidden boxes are not recoverable from a single image, so the first step is defining whether you want a visible-box count or an estimated total-stack count.

Burhan’s YOLOE result is a solid front end, but the counting logic has to sit on top of it. In practice, the most reliable setup here is `segmentation + homography + layout prior`: rectify the pallet/camera plane, segment visible box faces, then fit a grid using known box size and pallet dimensions to infer missing occluded slots. For cropped boxes, use a strict ROI rule, like counting only boxes whose centroid is inside the pallet region, and flag edge cases as uncertain.

If you have video instead of a single frame, tracking across frames helps a lot since a box partially hidden in one frame may be clearer in another. For new experiments, I’d test [Ultralytics YOLO26](https://docs.ultralytics.com/models/yolo26) `-seg` or `obb`, but the biggest gain here usually comes from geometry constraints, not just switching models.

If the box size and pallet pattern are consistent, this is very workable. If they vary a lot, exact total count from one view will stay inherently ambiguous.
