Vision AI Models (Image Input)

217 AI models on CoreAI accept images as input — you can send a photo, screenshot, chart, or scanned document and ask questions about it in plain language. Vision-capable models come from Qwen, Z.AI, Meta, DeepSeek, Dots-studio and other providers, and the list below is ordered newest-first from the live catalog.

Typical uses: extracting text and tables from documents, explaining diagrams, debugging code from screenshots, identifying objects, and turning whiteboard photos into structured notes. In the CoreAI app, attach an image in any chat with one of these models and it just works.

ModelProviderContextInput $/1MOutput $/1MCapabilities
Qwen: Qwen3.8 Flash Qwen 1000K $0.15 $0.47 Reasoning, Vision
Z.ai: GLM 5.3 Flash Z-ai 1049K $0.07 $0.25 Reasoning, Vision
Z.ai: GLM 5.3 Flash (batch) Z-ai 1049K $0.15 $0.50 Reasoning, Vision
Meta: Muse Spark 1.2 Contributor Meta 1049K $0.10 $0.20 Reasoning, Vision
DeepSeek: DeepSeek V4 Flash Vision Exp DeepSeek 1049K $0.22 $0.66 Reasoning, Vision
Qwen: Qwen3.8 27B Qwen 1000K $0.42 $2.55 Reasoning, Vision
Dots Studio: Dots3-Note Preview (free) Dots-studio 512K $0 $0 Reasoning, Vision
Google: Gemini 3.7 Flash Google 1049K $0.75 $3.75 Reasoning, Vision
Google: Gemini 3.7 Flash (batch) Google 1049K $0.19 $0.94 Reasoning, Vision
ByteDance Seed: Seed 2.1 Turbo Bytedance-seed 262K $0.50 $2.50 Reasoning, Vision
ByteDance Seed: Seed-2.0-Code Bytedance-seed 262K $0.50 $3 Reasoning, Vision
SpaceXAI: Grok 4.6 xAI 500K $2 $6 Reasoning, Vision
Sakana: Sakana Namazu Sakana 262K $0.95 $4 Reasoning, Web search, Vision
Meta: Muse Glimmer 30B Meta 131K $0.30 $1.20 Reasoning, Vision
Meta: Muse Glimmer 30B (batch) Meta 131K $0.35 $1.50 Reasoning, Vision
Meta: Muse Spark 1.2 Meta 1049K $1.25 $4.25 Reasoning, Vision
Qwen: Qwen3.8 Max Qwen 1000K $2 $6 Reasoning, Vision
Thinking Machines: Inkling Small Thinkingmachines 524K $0.45 $1.20 Reasoning, Vision
Thinking Machines: Inkling Small (batch) Thinkingmachines 524K $0.50 $1.20 Reasoning, Vision
Thinking Machines: Inkling Small (free) Thinkingmachines 1049K $0 $0 Reasoning, Vision
Qwen: Qwen3.7 Flash Qwen 1000K $0.03 $0.13 Reasoning, Vision
Claude Opus 5 (Fast) Anthropic 1000K $10 $50 Reasoning, Vision
Claude Opus 5 Anthropic 1000K $5 $25 Reasoning, Vision
Claude Opus 5 (batch) Anthropic 1000K $2.50 $12.50 Reasoning, Vision
Google: Gemini 3.6 Flash Google 1049K $0.75 $3.75 Reasoning, Vision
Google: Gemini 3.6 Flash (batch) Google 1049K $0.38 $1.88 Reasoning, Vision
Google: Gemini 3.5 Flash Lite Google 1049K $0.30 $2.50 Reasoning, Vision
Google: Gemini 3.5 Flash Lite (batch) Google 1049K $0.15 $1.25 Reasoning, Vision
Thinking Machines: Inkling Thinkingmachines 524K $1 $4.05 Reasoning, Vision
Thinking Machines: Inkling (batch) Thinkingmachines 524K $1 $4.05 Reasoning, Vision
Thinking Machines: Inkling (free) Thinkingmachines 1049K $0 $0 Reasoning, Vision
MoonshotAI: Kimi K3 Moonshotai 1049K $3 $15 Reasoning, Vision
MoonshotAI: Kimi K3 (batch) Moonshotai 1049K $3 $15 Reasoning, Vision
Meta: Muse Spark 1.1 Meta 1049K $1.25 $4.25 Reasoning, Vision
OpenAI: GPT-5.6 Luna Pro OpenAI 1050K $0.20 $1.20 Reasoning, Vision
OpenAI: GPT-5.6 Luna OpenAI 1050K $0.20 $1.20 Reasoning, Vision
OpenAI: GPT-5.6 Terra Pro OpenAI 1050K $2 $12 Reasoning, Vision
OpenAI: GPT-5.6 Terra OpenAI 1050K $2 $12 Reasoning, Vision
OpenAI: GPT-5.6 Sol Pro OpenAI 1050K $2 $10 Reasoning, Vision
OpenAI: GPT-5.6 Sol OpenAI 1050K $2 $10 Reasoning, Vision
SpaceXAI: Grok 4.5 xAI 500K $2 $6 Reasoning, Vision
xAI: Grok Latest ~x-ai 500K $2 $6 Reasoning, Vision
Anthropic: Claude Sonnet 5 Anthropic 1000K $2 $10 Reasoning, Vision
Anthropic: Claude Sonnet 5 (batch) Anthropic 1000K $1 $5 Reasoning, Vision
Nex AGI: Nex-N2-Mini Nex-agi 262K $0.02 $0.10 Reasoning, Vision
Sakana: Fugu Ultra Sakana 1000K $5 $30 Reasoning, Web search, Vision
MoonshotAI: Kimi K2.7 Code Moonshotai 262K $0.66 $3.40 Reasoning, Vision
Anthropic: Claude Fable Latest ~anthropic 1000K $10 $50 Reasoning, Vision
Anthropic: Claude Fable 5 Anthropic 1000K $10 $50 Reasoning, Vision
Anthropic: Claude Fable 5 (batch) Anthropic 1000K $5 $25 Reasoning, Vision

Showing the top 50 of 217 models. Browse the full directory →

Frequently Asked Questions

What is a vision AI model?

A vision (multimodal) model accepts images alongside text. Instead of describing a chart or error message in words, you attach the image and the model reads it directly — extracting text, interpreting layouts, and answering questions about what it sees.

Which AI models can read images in 2026?

217 models on CoreAI currently support image input, including the latest multimodal releases from Qwen, Z.AI, Meta, DeepSeek, Dots-studio. The table on this page is generated live, so it always matches what's available in the app.

Can vision AI models read handwriting and PDFs?

Most modern vision models handle printed text and clear handwriting well. For multi-page PDFs, CoreAI's Scan & PDF tool splits documents into pages a vision model can process.

Are vision models more expensive than text models?

Image input usually consumes more tokens than plain text, so a vision request costs somewhat more than a text-only one on the same model — but pricing per token is identical; you simply send more tokens.

Explore More AI Models

Try 300+ AI Models Free

Chat with GPT-5, Claude, Gemini and hundreds more — all in one app.

Use Web App Download App →