AllDSH
中文

Use cases

Vision & OCR

Let text-only agents inspect images, screenshots, diagrams and scanned pages.

Vision plugins give a text-only model a way to look. The strongest implementations do not just describe an image, they return structured evidence the model can reason over: extracted text with positions, layout regions, colour values, and crop coordinates it can request next.

Prefer plugins that treat images as ordinary tool calls instead of a special turn type, and check whether the vision backend is local, free, or requires your own key. For screenshot-heavy work, confirm that long images are tiled rather than downscaled into illegibility.

3 plugins

liustack avatar

Modlens

liustack

Editor's pick

Vision for text-only harnesses: paste an image and get structured JSON evidence covering OCR, layout and semantics.

L5 gold Scanned
3.9k TypeScript 3d ago
ysr666 avatar

Vision Router

ysr666

Eyes for text-only agents with a free built-in vision chain and pixel level tools for grounding, crops and diffs.

L5 gold Scanned
1.1k JavaScript 2d ago
Anionex avatar

Vision Toolkit

Anionex

A vision toolbox for text-only models covering image QA, long screenshot OCR, UI restoration and grounding.

L4 silver Scanned
853 TypeScript 6d ago

Related

ESC

Type to search the index. Built at deploy time by Pagefind.