agent-vision-toolkit

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

PluginVisionAI ModelsDeveloperAutomationVision / OCR
Verification
L2 · Structured
Security
Pending
Health
Active
Trust
Bronze

What it doesAI

A vision toolkit for text-only LLMs enabling image Q&A, OCR, UI restoration, and GUI automation with agent integrations.

  • Multi-image understanding and image Q&A
  • Long-screenshot OCR and frontend UI restoration
  • GUI automation with optional agent integrations

AI-generated from the repo README — for reference only.

Installation

dsh plugin --profile web add agent-vision-toolkit

Install method: npm · not yet tested in container (L3+)

Compatibility

DSH VersionStatus
not statedDeclared — not tested

Requirements

  • • Node.js: not stated
  • • DSH: declared "not stated"
  • • External credentials: none detected

Security Report

Automated scan, not manual review.

• Dependency vulnerabilities: — (pending)

• Suspicious permissions: — (pending)

• Hardcoded secrets: — (pending)

• Supply chain risks: — (pending)

⚠ Security audit scheduled — results will appear here after the next scan.

Activity

Last commit 2026-08-14 · activity: active

6 mo

Source

GitHub: github.com/Anionex/agent-vision-toolkit