AllDSH
中文
liustack avatar

Modlens

liustack · liustack/modlens

Vision for text-only harnesses: paste an image and get structured JSON evidence covering OCR, layout and semantics.

Listing checked by AllDSH Editorial , data updated · How to cite

L5 gold active Scanned

Install command

dsh plugin --profile web add modlens

Requires a running harness: start it with npx @deepseek-ai/dsh webthen run the command above. New to this? See the installation guide.

What it does

Modlens is the vision bridge for models that cannot see. Instead of returning a prose description that a text model has to trust, it returns structured JSON evidence: recognised text with positions, layout regions, and semantic labels for what each region appears to be. The model can then reason over evidence, cite a region, or ask for a tighter crop.

Key features

  • OCR output with per-block coordinates rather than a flat string
  • Layout segmentation into headers, tables, code, forms and media
  • Semantic labels that let a text model reference regions by name
  • Works with pasted images and workspace file paths
  • Structured output that survives downstream tool calls

Install

dsh plugin --profile web add modlens

Compatibility notes

Very long screenshots benefit from tiling; check the plugin’s tile settings if fine print comes back garbled.

Similar DSH plugins

Browse all
dsh-market avatar

DSH Market

dsh-market

A visual plugin market inside DeepSeek Harness for browsing, searching and one-click installs.

L5 gold Scanned
3.4k TypeScript 5d ago
ysr666 avatar

Vision Router

ysr666

Eyes for text-only agents with a free built-in vision chain and pixel level tools for grounding, crops and diffs.

L5 gold Scanned
1.1k JavaScript 2d ago
Anionex avatar

Vision Toolkit

Anionex

A vision toolbox for text-only models covering image QA, long screenshot OCR, UI restoration and grounding.

L4 silver Scanned
853 TypeScript 6d ago
ESC

Type to search the index. Built at deploy time by Pagefind.