# New Release: Ultralytics v8.3.196

**URL:** <https://community.ultralytics.com/t/new-release-ultralytics-v8-3-196/1447>\
**Category:** Discussion\
**Tags:** ultralytics-official, releases, announcements\
**Created:** [September 8, 2025, 1:14pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-3-196/1447 "2025-09-08T13:14:35Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![glenn-jocher](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/glenn-jocher/32/83_2.png) [@glenn-jocher](https://community.ultralytics.com/u/glenn-jocher)\
**Post date:** [September 8, 2025, 1:14pm UTC](https://community.ultralytics.com/t/new-release-ultralytics-v8-3-196/1447/1 "2025-09-08T13:14:35Z")

</div>

# Ultralytics v8.3.196 is out 🎉 — faster training with torch.compile, smoother IO, and sturdier exports

Short version: v8.3.196 brings optional `torch.compile` acceleration across train/val/predict for up to ~30% faster runs, plus dataloader throughput boosts, unified device transfers, more robust CoreML exports, and smoother plotting/config handling. ⚡

YOLO11 remains our recommended default for all use cases.

## 🌟 Summary

- Up to ~30% faster training and snappier inference thanks to opt-in `torch.compile` across CLI and Python.
- Faster, safer data loading with higher prefetch and smarter batch handling.
- Centralized device transfers reduce CPU/GPU mismatch issues across tasks.
- CoreML export pipeline cleaned up for more reliable deployments on Apple platforms.
- Plotting and configuration paths are more CI- and container-friendly.

You can review the release details on the official release page in the [Ultralytics v8.3.196 notes](https://github.com/ultralytics/ultralytics/releases/tag/v8.3.196), and see every change in the [full changelog diff from v8.3.195](https://github.com/ultralytics/ultralytics/compare/v8.3.195...v8.3.196).

## 🚀 New Features

- torch.compile acceleration (primary)
  - New `compile` arg available in `train`, `val`, and `predict` via CLI, config, and Python API.
  - New helpers make adoption safe and granular, including `attempt_compile(...)` to enable compilation when supported and `disable_dynamo(...)` to exclude sensitive paths.
  - End-to-end integrations:
    - Trainer compiles models after loss initialization, marks dynamic tensors for stability, and unwraps models for EMA/checkpointing.
    - Validator can compile for standalone validation; training’s final eval skips compile for stability/speed.
    - Predictor supports `compile=True` for accelerated inference on CUDA/CPU/MPS when supported.

  - Utility rename: `de_parallel` ➝ `unwrap_model` to handle both parallel and compiled models consistently.

- Documentation updated to reflect the new `compile` argument and torch utility functions.

## 🛠 Improvements

- Faster data loading
  - Default `prefetch_factor` doubled to 4 when `num_workers > 0`, and gracefully omitted on older PyTorch releases to avoid errors.
  - Safer `drop_last` behavior with compile-enabled training to improve shape stability.

- Unified device handling
  - Centralized logic for moving batch tensors to the correct device across detection, pose, segmentation, and YOLOE reduces code duplication and “tensor on CPU vs GPU” issues.

- CoreML export robustness
  - Cleaned up NMS pipeline with direct use of spec outputs, explicit shapes where needed, consistent IO names, and simpler wiring for more reliable exports across macOS/Linux/Windows.

- Plotting stability
  - `@plt_settings()` wraps `feature_visualization(...)` for backend-safe, non-blocking plots in headless/CI environments.

- Config directory resolution
  - Smarter `get_user_config_dir()` honors `YOLO_CONFIG_DIR`, follows OS conventions (XDG on Linux), and falls back to writable paths like `/tmp` when needed.

- Compatibility and CI polish
  - TorchVision compatibility matrix updated for PyTorch 2.8/0.23 and 2.9/0.24.
  - GitHub Actions bumped for `setup-python` v6 and `actions/stale` v10.

## 🐛 Bug Fixes

- Fixed missing tensors moved to device in Trainer and Validator `preprocess_batch` methods.
- Reduced overly verbose user config directory checks.
- Corrected a minor typo in a deprecation warning.

## ⚡ Quick start with compile

CLI example:

```bash
yolo train model=yolo11n.pt data=coco8.yaml epochs=100 compile=True

```

Python example:

```python
from ultralytics import YOLO

model = YOLO("yolo11n.pt")
model.train(data="coco8.yaml", epochs=100, compile=True)

# Also supported:
model.val(compile=True)
model.predict("img.jpg", compile=True)

```

Tip: `torch.compile` works best on recent PyTorch (2.0+). If your stack or device doesn’t support it, leave `compile=False` (default).

## 📦 Upgrade

Install or upgrade with:

```bash
pip install -U ultralytics

```

For guidance on tasks and usage, the [YOLO11 documentation](https://docs.ultralytics.com/models/yolo11/) and mode guides for [Train](https://docs.ultralytics.com/modes/train/), [Val](https://docs.ultralytics.com/modes/val/), and [Predict](https://docs.ultralytics.com/modes/predict/) are great starting points.

## ✅ PRs included in v8.3.196

- [Add @plt\_settings() decorator to feature\_visualization()](https://github.com/ultralytics/ultralytics/pull/21973) by [Glenn Jocher](https://github.com/glenn-jocher)
- [Cleanup CoreML NMS pipeline code](https://github.com/ultralytics/ultralytics/pull/21970) by [Y-T-G](https://github.com/Y-T-G)
- [Double default Dataloader prefetch\_factor to 4](https://github.com/ultralytics/ultralytics/pull/21974) by [Glenn Jocher](https://github.com/glenn-jocher)
- [Update TorchVision compat matrix with 2.8 and 2.9](https://github.com/ultralytics/ultralytics/pull/21978) by [Glenn Jocher](https://github.com/glenn-jocher)
- [Fix overly verbose USER\_CONFIG\_DIR checks](https://github.com/ultralytics/ultralytics/pull/21980) by [Glenn Jocher](https://github.com/glenn-jocher)
- [Fix missing tensors on device in preprocess\_batch](https://github.com/ultralytics/ultralytics/pull/21981) by [Glenn Jocher](https://github.com/glenn-jocher)
- [Bump actions/setup-python from 5 to 6](https://github.com/ultralytics/ultralytics/pull/21984) by [Dependabot](https://docs.github.com/en/code-security/dependabot/dependabot-security-updates/configuring-dependabot-security-updates)
- [Bump actions/stale from 9 to 10](https://github.com/ultralytics/ultralytics/pull/21983) by [Dependabot](https://docs.github.com/en/code-security/dependabot/dependabot-security-updates/configuring-dependabot-security-updates)
- [Fix typo in deprecation\_warn](https://github.com/ultralytics/ultralytics/pull/21987) by [Rizwan Munawar](https://github.com/RizwanMunawar)
- [ultralytics 8.3.196 torch.compile acceleration](https://github.com/ultralytics/ultralytics/pull/21975) by [Glenn Jocher](https://github.com/glenn-jocher)

You can browse all differences in the [full changelog compare view](https://github.com/ultralytics/ultralytics/compare/v8.3.195...v8.3.196) and read the highlights on the [v8.3.196 release page](https://github.com/ultralytics/ultralytics/releases/tag/v8.3.196).

## 🙌 Thanks and feedback

Big thanks to everyone in the YOLO community and the Ultralytics team for ideas, testing, and contributions. We’d love your feedback—please share results, questions, or issues by opening an [Ultralytics GitHub issue](https://github.com/ultralytics/ultralytics/issues) or joining the discussion in [Ultralytics Discussions](https://github.com/orgs/ultralytics/discussions).

Happy training with YOLO11 and enjoy the speedups! 🚀
