[Research] Microsoft uses YOLOv8 in new OmniParser GUI-based Agent

ArXiv Paper

OMNIPARSER, a comprehensive method for parsing user interface screenshots into structured elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface.

Results Preview