# YOLO26 experiment: Raw model output to bounding boxes

**URL:** <https://community.ultralytics.com/t/yolo26-experiment-raw-model-output-to-bounding-boxes/1772>\
**Category:** Community Showcase\
**Created:** [January 20, 2026, 7:06pm UTC](https://community.ultralytics.com/t/yolo26-experiment-raw-model-output-to-bounding-boxes/1772 "2026-01-20T19:06:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![fzngagan](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/fzngagan/32/1269_2.png) [@fzngagan](https://community.ultralytics.com/u/fzngagan)\
**Post date:** [January 20, 2026, 7:06pm UTC](https://community.ultralytics.com/t/yolo26-experiment-raw-model-output-to-bounding-boxes/1772/1 "2026-01-20T19:06:30Z")

</div>

Did some experimentation with the yolo26 model. Basically, I want to replicate the inference process from model outputs to bounding box drawing. Here’s what I did:

- collected model’s raw predictions
- scaled them to the original image dimentions
- plotted the boxes on the original image using `matplotlib.patches`

> <https://github.com/fzngagan/object-detection-experiments/blob/main/yolo26.ipynb>

Understood a bit of the internals of the `ultralytics` codebase in the process. There are less abstractions in the code than I would have imagined, but I’d like to build a very simple training/inference pipeline mostly for learning and experimentation.

---

<div class="post-metadata">

**Author:** ![pderrenger](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/pderrenger/32/73_2.png) [@pderrenger](https://community.ultralytics.com/u/pderrenger)\
**Post date:** [January 21, 2026, 1:38pm UTC](https://community.ultralytics.com/t/yolo26-experiment-raw-model-output-to-bounding-boxes/1772/2 "2026-01-21T13:38:03Z")

</div>

Your approach is basically the right mental model: the “gotcha” is that the network predicts in the **letterboxed** (resized + padded) image space, and you need to undo that exact resize/pad before drawing on the original. In Ultralytics YOLO this undo step is handled for you, so if you just want the final pixel boxes to plot, you can pull them directly from the `Results` object (they’re already scaled to the original image):

```python
from ultralytics import YOLO

model = YOLO("yolo26n.pt")
r = model("https://ultralytics.com/images/bus.jpg")[0]

xyxy = r.boxes.xyxy.cpu().numpy() # x1,y1,x2,y2 in original-image pixels
conf = r.boxes.conf.cpu().numpy()
cls = r.boxes.cls.cpu().numpy().astype(int)

print(xyxy[0], conf[0], cls[0])

```

If you’re trying to replicate the _full_ path from raw head outputs → final boxes, the key post steps are “decode → NMS → scale back to original,” implemented in `ultralytics/utils/ops.py` (look for `non_max_suppression()` and `scale_boxes()`). The coordinate formats we expose on `Results` are summarized in the [bounding box glossary](https://www.ultralytics.com/glossary/bounding-box), and the expected “absolute vs normalized” behavior is also covered in the [common issues guide section on box coordinates](https://docs.ultralytics.com/guides/yolo-common-issues/).

If you share what you’re calling “raw predictions” (tensor shape + where you tapped it: PyTorch model forward vs an exported model), I can point you to the exact decode step for that output format, since it differs depending on whether you captured pre- or post-decode outputs.
