What it does
A vision toolkit that lets text-only DeepSeek Harness agents work with images. It covers image questions and answers, OCR on long screenshots, restoring a UI from a screenshot, grounding, pixel diff, artifacts and a web interface.
Use cases
- 01Ask questions about several pasted images
- 02Turn a screenshot into front-end UI code
- 03Compare two screenshots for pixel differences