# New Release: Ultralytics v8.4.37

**URL:** <https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928>\
**Category:** Discussion\
**Tags:** ultralytics-official, releases, announcements\
**Created:** [April 10, 2026, 1:33pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928 "2026-04-10T13:33:15Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![glenn-jocher](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/glenn-jocher/32/83_2.png) [@glenn-jocher](https://community.ultralytics.com/u/glenn-jocher)\
**Post date:** [April 10, 2026, 1:33pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928/1 "2026-04-10T13:33:15Z")

</div>

# Ultralytics v8.4.37 is out! 🚀

We’re excited to release **Ultralytics `v8.4.37`** , a quality + workflow-focused update that improves tuning, training robustness, evaluation reliability, and documentation clarity. The standout change is **NDJSON-based hyperparameter tuning for multi-dataset workflows** , alongside new support for handling **class imbalance** during training. 🙌

If you’re training, tuning, or deploying **Ultralytics YOLO** , this release should make your workflow smoother and more reliable.

## Highlights

> ⚠ **WARNING**  
> The mAP calculation has been revised in this release. Reported mAP may be slightly lower than in previous Ultralytics versions, but now more closely matches pycocotools’ COCOEval

### 🧠 Better hyperparameter tuning for multi-dataset workflows

A major upgrade in [PR #24179 by @Laughing-q](https://github.com/ultralytics/ultralytics/pull/24179) moves tuning logs from CSV to **`tune_results.ndjson`** , making experiment tracking more flexible and robust.

This includes:

- Per-dataset fitness tracking for multi-dataset tuning
- Updated output/plot naming such as `tune_fitness.png`
- Improved MongoDB sync behavior aligned with local NDJSON logs

For teams running larger tuning experiments across multiple datasets, this is the biggest improvement in `v8.4.37`. 📈

### ⚖ New class imbalance support in training

With [PR #23565 by @ahmet-f-gumustas](https://github.com/ultralytics/ultralytics/pull/23565), detection training now supports a new hyperparameter: `cls_pw`.

This allows you to give more weight to underrepresented classes during training:

- New hyperparameter: `cls_pw`
- Default value: `0.0`
- Existing behavior remains unchanged unless you enable it

This is especially useful for long-tail datasets where rare classes need more learning emphasis. 🎯

### 🛡 More reliable training and checkpointing

Training robustness got a nice boost in this release:

- [PR #24170 by @Laughing-q](https://github.com/ultralytics/ultralytics/pull/24170) ensures the **first-epoch checkpoint can still be saved** even if EMA contains `NaN` or `Inf` values early in training
- [PR #24185 by @glenn-jocher](https://github.com/ultralytics/ultralytics/pull/24185) fixes a regression affecting **local zip datasets** in Ultralytics Platform training

These changes help reduce avoidable interruptions and make recovery safer in edge cases. 🔁

## Improvements

### ✅ More robust evaluation and CI behavior

A precision edge case in AP computation was fixed in [PR #24175 by @Laughing-q](https://github.com/ultralytics/ultralytics/pull/24175), improving `compute_ap` reliability.

Additional CI and benchmark improvements include:

- [PR #24181 by @Laughing-q](https://github.com/ultralytics/ultralytics/pull/24181) updates benchmark verbosity for `Dockerfile-nvidia-arm64`
- [PR #24183 by @fcakyon](https://github.com/ultralytics/ultralytics/pull/24183) simplifies engine resume tests

### 🧹 Cleaner distributed training logs

[PR #24177 by @Laughing-q](https://github.com/ultralytics/ultralytics/pull/24177) reduces duplicate model info printing in DDP and multi-process training, making logs easier to read and debug.

## Docs and Platform updates

This release also improves the docs and overall UX:

- [PR #24180 by @amanharshx](https://github.com/ultralytics/ultralytics/pull/24180) fixes incorrect task-specific `.load()` examples for segment and OBB
- [PR #24182 by @easyrider11](https://github.com/ultralytics/ultralytics/pull/24182) updates OpenVINO notebook links for **[YOLO26](https://platform.ultralytics.com/ultralytics/yolo26)** optimization
- [PR #24122 by @raimbekovm](https://github.com/ultralytics/ultralytics/pull/24122) improves trainer callback docs with clearer descriptions
- [PR #24186 by @raimbekovm](https://github.com/ultralytics/ultralytics/pull/24186) replaces the quickstart journey diagram with an interactive workflow graph for **[Ultralytics Platform](https://platform.ultralytics.com)**
- [PR #24176 by @glenn-jocher](https://github.com/ultralytics/ultralytics/pull/24176) labels non-text code fences more clearly in docs

## Release tag note

The release-tag PR itself, [PR #24192 by @glenn-jocher](https://github.com/ultralytics/ultralytics/pull/24192), is a version bump from `8.4.36` to `8.4.37`. The runtime and workflow improvements come from the PRs above. 📦

## New contributor shoutout 🎉

A warm welcome to [@easyrider11](https://github.com/easyrider11), who made their first contribution in [PR #24182](https://github.com/ultralytics/ultralytics/pull/24182)! Thank you for helping improve the YOLO ecosystem.

## Why this release matters

`v8.4.37` is especially helpful if you:

- Run hyperparameter tuning across multiple datasets
- Train on imbalanced datasets with rare classes
- Need safer checkpointing during unstable early epochs
- Want cleaner DDP logs and fewer flaky CI/benchmark issues
- Prefer clearer docs and more accurate examples

## Try it out

Update with:

```bash
pip install -U ultralytics

```

Then explore the full details in the [v8.4.37 release page](https://github.com/ultralytics/ultralytics/releases/tag/v8.4.37) or browse the [full changelog](https://github.com/ultralytics/ultralytics/compare/v8.4.36...v8.4.37).

If you test the release, we’d love to hear how it works for your training and tuning workflows. Feedback, bug reports, and PRs are always welcome! 💬

---

<div class="post-metadata">

**Author:** ![Livlo1970](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/livlo1970/32/1469_2.png) [@Livlo1970](https://community.ultralytics.com/u/Livlo1970)\
**Post date:** [April 11, 2026, 8:57pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928/2 "2026-04-11T20:57:53Z")

</div>

Hi,

I wanted to flag a significant side effect of PR #24175 (“Fix AP calculation precision in `compute_ap`”) introduced in v8.4.37.

After upgrading from 8.4.36 to 8.4.37, I observed a consistent **−0.018 drop in mAP50-95** (e.g. 0.44719 → 0.42922) and **−0.015 drop in mAP50** across all my validation runs — with absolutely no changes to the model, training code, dataset, or weights.

I was able to confirm this is purely a metrics computation change by running the exact same PASS2 fine-tuning configuration on the same checkpoint under both versions:

| Ultralytics version | train/box\_loss | val/box\_loss | Precision | Recall | mAP50 | mAP50-95 |
| --- | --- | --- | --- | --- | --- | --- |
| 8.4.36 | 0.86178 | 1.13583 | 0.74674 | 0.63764 | 0.67372 | 0.44719 |
| 8.4.37 | 0.86178 | 1.13583 | 0.74674 | 0.63764 | 0.65841 | 0.42922 |

All losses, precision, and recall are byte-for-byte identical — only mAP values differ.

While I understand the intent was to fix an edge case in AP calculation, this change has a few consequences worth noting:

1. **All previously published benchmarks** (YOLO11, YOLO12, YOLO26 model cards) were computed under the old metric — they are no longer reproducible under 8.4.37+

2. **The change is silent** — users comparing results across versions may mistakenly attribute the drop to a model or training regression

3. **A −0.018 shift on mAP50-95 is substantial** — it’s larger than many hyperparameter tuning gains that researchers spend days chasing

I believe a change of this magnitude in a core evaluation metric deserves explicit mention in the release notes as a breaking change, so users are aware and can re-baseline their experiments accordingly.

Thank you for the great work on Ultralytics — just wanted to make sure this doesn’t catch others off guard as it did for me.

---

<div class="post-metadata">

**Author:** ![Toxite](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/toxite/32/123_2.png) [@Toxite](https://community.ultralytics.com/u/Toxite)\
**Post date:** [April 11, 2026, 9:04pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928/3 "2026-04-11T21:04:31Z")

</div>

@Livlo1970 Thanks for flagging.

> [@Livlo1970](#):
>
> 1. **All previously published benchmarks** (YOLO11, YOLO12, YOLO26 model cards) were computed under the old metric — they are no longer reproducible under 8.4.37+

For the official models, Ultralytics uses COCOEval to get the mAP scores. COCOEval is the standard and what’s recommended for research for consistency with other works. Ultralytics’ native mAP calculation is faster compared to COCOEval, but it doesn’t produce exactly the same mAP as COCOEval. This PR in particular tries to bring the Ultralytics mAP closer to what’s obtained by COCOEval.

> [@Livlo1970](#):
>
> I believe a change of this magnitude in a core evaluation metric deserves explicit mention in the release notes as a breaking change, so users are aware and can re-baseline their experiments accordingly.

We will update the release notes to make it explicit.

---

<div class="post-metadata">

**Author:** ![Livlo1970](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/livlo1970/32/1469_2.png) [@Livlo1970](https://community.ultralytics.com/u/Livlo1970)\
**Post date:** [April 11, 2026, 9:12pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928/4 "2026-04-11T21:12:46Z")

</div>

Thank you for the quick and clear response.

That’s very helpful to know that the official benchmarks use COCOEval and are therefore unaffected. It makes sense to align the native calculation with the standard — it’s a good change in the long run.

I appreciate that you’ll update the release notes. That will definitely help users who, like me, rely on the native metrics for iterative comparisons during training.

Thanks again for the transparency and the great work on Ultralytics.

---

<div class="post-metadata">

**Author:** ![pderrenger](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/pderrenger/32/73_2.png) [@pderrenger](https://community.ultralytics.com/u/pderrenger)\
**Post date:** [April 12, 2026, 1:41pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-4-37/1928/5 "2026-04-12T13:41:09Z")

</div>

Absolutely, and thanks again for surfacing it early.

We’ve now made the `v8.4.37` note explicit on the [release page](https://github.com/ultralytics/ultralytics/releases/tag/v8.4.37) so shifts in native `mAP50` and `mAP50-95` between `8.4.36` and `8.4.37+` are less likely to be mistaken for a real Ultralytics YOLO regression.

Your takeaway is exactly right: re-baselining is the right move for native metric comparisons, and for strict cross-version comparability COCOEval remains the best reference. Appreciate the careful validation here.
