Hello, I hope you are doing all well,
I am developing a YOLO models to detect *only people wearing Shalwar Kameez* . I do not want to use any API or external vision service. I have *4,000 labeled Shalwar Kameez images* and *5,000 negative images* containing only pant-shirt/other clothing with empty labels. However, the model still detects many pant-shirt people as Shalwar Kameez and misses many real Shalwar Kameez people. How would you improve the dataset, labeling strategy, and training pipeline to achieve reliable real-world performance?
Thank you for your time.