What it does
Adds image understanding to text-only coding agents. You paste an image and get structured JSON evidence back, including OCR text, layout and semantics, which the agent can then work from.
Use cases
- 01Read text from a pasted screenshot
- 02Describe the layout of a UI mockup
- 03Give a text-only agent structured image data