multimodal-bridge

multimodal-bridge 是一个多模态能力桥:把 Qwen 的视觉理解(Qwen-VL)与图像生成(Qwen-Image)带给没有原生多模态能力的纯文本模型(如 DeepSeek)。它有两种形态、同一套后端: MCP Server(qwen_vision / qwen_generate 工具):任何支持 MCP 的宿主(Claude Code、Kimi Code 等)直接挂载; DSH 插件(npm 包 dsh-multimodal-bridge,DeepSeek Harness bundle):dsh plugin add 一行安装,含模型自动 fallback、尺寸自适应与图片结果卡片。

PluginVisionDeveloperNetworkAI ModelsVision / OCR
Verification
L2 · Structured
Security
Pending
Health
Active
Trust
Bronze

What it doesAI

Bridges Qwen-VL vision understanding and Qwen-Image generation to text-only models like DeepSeek.

  • Adds Qwen-VL visual understanding to text-only LLMs
  • Adds Qwen-Image generation for multimodal output
  • Installable as MCP server or DSH plugin

AI-generated from the repo README — for reference only.

Installation

dsh plugin --profile web add multimodal-bridge

Install method: npm · not yet tested in container (L3+)

Compatibility

DSH VersionStatus
not statedDeclared — not tested

Requirements

  • • Node.js: not stated
  • • DSH: declared "not stated"
  • • External credentials: none detected

Security Report

Automated scan, not manual review.

• Dependency vulnerabilities: — (pending)

• Suspicious permissions: — (pending)

• Hardcoded secrets: — (pending)

• Supply chain risks: — (pending)

⚠ Security audit scheduled — results will appear here after the next scan.

Activity

Last commit 2026-08-13 · activity: active

6 mo

Source

GitHub: github.com/Spirit4471/multimodal-bridge