目录 / 视觉能力

agent-vision-toolkit

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

dsh plugin --profile web add github:Anionex/agent-vision-toolkit
作者
Anionex
Stars
829
许可证
MIT
来源
https://github.com/Anionex/agent-vision-toolkit
更新
2026/8/14

说明