What it does
A toolkit and skill that lets text-only language models work with images. It covers image questions, multi-image understanding, long screenshot OCR, restoring a front-end UI from a picture and GUI automation. It can be connected to agents such as Codex and Claude Code.
Use cases
- 01Ask questions about a screenshot
- 02Rebuild a front-end UI from an image
- 03Read text from a long screenshot