
agent-vision-toolkit
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
This plugin is a GitHub repository with the dsh-plugin topic — the repository README is the authoritative install source. Use the template below, adapting the package name / path from the README:
dsh plugin --profile web add $agent-vision-toolkitLocal development? dsh --profile web --patch ./cordis.yml with an absolute plugin path — see Plugin Development.GitHub source? dsh plugin --profile web add github:$Anionex/$agent-vision-toolkit — needs a prepare script in the repo (see packaging).
| Repository | Anionex/agent-vision-toolkit |
| License | MIT |
| Category | Skill |
| Status | Community |
| Declared compatibility | not stated |
| Last commit | 2026-08-13 |
| Indexed | 2026-08-13 |
Declared is the author's claim; dsh.so has not independently tested it. See the changelog for version compatibility.Built by DeepSeek, for DeepSeek — a Swift-native macOS coding agent
ViewThe first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics).
ViewPlugin and skin collection for DeepSeek Harness (DSH) Web UI - task board, git graph, right-side panel, remote mobile UI, pet, live token stats, and skin center.
ViewQuestions or feedback? Official Discussions is the project’s canonical support channel; Discord has an active community. dsh.so itself improves via plugin submissions and your feedback.