# Speaker-aware 9:16 auto-reframe with YOLOv8 + TalkNet-ASD, running locally

**URL:** https://community.ultralytics.com/t/speaker-aware-9-16-auto-reframe-with-yolov8-talknet-asd-running-locally/2242
**Category:** Discussion
**Tags:** showcase
**Created:** [September 26, 2026, 8:53am UTC](https://community.ultralytics.com/t/speaker-aware-9-16-auto-reframe-with-yolov8-talknet-asd-running-locally/2242 "2026-09-26T08:53:05Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Colin](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/colin/32/1716_2.png) [@Colin](https://community.ultralytics.com/u/Colin)
#### Post date: [September 26, 2026, 8:53am UTC](https://community.ultralytics.com/t/speaker-aware-9-16-auto-reframe-with-yolov8-talknet-asd-running-locally/2242/1 "2026-09-26T08:53:05Z")

</div>

Sharing a project that leans on Ultralytics YOLO for something a bit off the  
usual path: deciding how to crop a 16:9 stream down to a 9:16 vertical clip  
without cutting off whoever is talking.

YOLOv8 tracks the people on screen, OpenCV handles the crop path, and  
TalkNet-ASD decides which of the tracked faces is actually speaking, so the  
crop follows the speaker instead of the largest face. All local, weights ship  
with the app, nothing fetched at runtime.

Two things I learned that may save someone else time:

1. Detector cost is about where you spend it, not the model. Re-detecting at  
25fps cost +452% overall on an RTX 3060; tracking at 8fps and reusing those  
boxes for the ASD crops cost +15%, for the same result.

2. Ultralytics inference is not thread-safe on a single model instance. The  
renders run in parallel and I hit this the hard way, so the shared model is  
now behind a lock. Worth flagging for anyone doing multi-threaded inference  
on one YOLO object.

The signal that decides which moments become clips is pluggable, so if anyone  
wants to point this at footage it was never tuned for, sports and gaming being  
the obvious ones, here is where a detector slots in:

> <https://github.com/ColinGPT9/clips-studio/blob/main/docs/EXTENDING.md#bring-your-own-model>

---

<div class="post-metadata">

### Author: ![pderrenger](https://sea1.discourse-cdn.com/flex001/user_avatar/community.ultralytics.com/pderrenger/32/73_2.png) [@pderrenger](https://community.ultralytics.com/u/pderrenger)
#### Post date: [September 27, 2026, 1:09am UTC](https://community.ultralytics.com/t/speaker-aware-9-16-auto-reframe-with-yolov8-talknet-asd-running-locally/2242/2 "2026-09-27T01:09:59Z")

</div>

Thanks for sharing, Colin. Speaker-aware reframing is a great use of YOLO tracking, and your 8 fps versus 25 fps measurements are a useful reminder to profile the _whole pipeline_, not just the detector.

Good call on locking the shared model. Our [thread-safe inference guide](https://docs.ultralytics.com/guides/yolo-thread-safe-inference) covers that approach and the alternative of a separate model instance per thread. I’d be curious how the crop behaves during rapid speaker changes or when ASD confidence drops.
