217 AI models on CoreAI accept images as input — you can send a photo, screenshot, chart, or scanned document and ask questions about it in plain language. Vision-capable models come from Qwen, Z.AI, Meta, DeepSeek, Dots-studio and other providers, and the list below is ordered newest-first from the live catalog.
Typical uses: extracting text and tables from documents, explaining diagrams, debugging code from screenshots, identifying objects, and turning whiteboard photos into structured notes. In the CoreAI app, attach an image in any chat with one of these models and it just works.
| Model | Provider | Context | Input $/1M | Output $/1M | Capabilities |
|---|---|---|---|---|---|
| Qwen: Qwen3.8 Flash | Qwen | 1000K | $0.15 | $0.47 | Reasoning, Vision |
| Z.ai: GLM 5.3 Flash | Z-ai | 1049K | $0.07 | $0.25 | Reasoning, Vision |
| Z.ai: GLM 5.3 Flash (batch) | Z-ai | 1049K | $0.15 | $0.50 | Reasoning, Vision |
| Meta: Muse Spark 1.2 Contributor | Meta | 1049K | $0.10 | $0.20 | Reasoning, Vision |
| DeepSeek: DeepSeek V4 Flash Vision Exp | DeepSeek | 1049K | $0.22 | $0.66 | Reasoning, Vision |
| Qwen: Qwen3.8 27B | Qwen | 1000K | $0.42 | $2.55 | Reasoning, Vision |
| Dots Studio: Dots3-Note Preview (free) | Dots-studio | 512K | $0 | $0 | Reasoning, Vision |
| Google: Gemini 3.7 Flash | 1049K | $0.75 | $3.75 | Reasoning, Vision | |
| Google: Gemini 3.7 Flash (batch) | 1049K | $0.19 | $0.94 | Reasoning, Vision | |
| ByteDance Seed: Seed 2.1 Turbo | Bytedance-seed | 262K | $0.50 | $2.50 | Reasoning, Vision |
| ByteDance Seed: Seed-2.0-Code | Bytedance-seed | 262K | $0.50 | $3 | Reasoning, Vision |
| SpaceXAI: Grok 4.6 | xAI | 500K | $2 | $6 | Reasoning, Vision |
| Sakana: Sakana Namazu | Sakana | 262K | $0.95 | $4 | Reasoning, Web search, Vision |
| Meta: Muse Glimmer 30B | Meta | 131K | $0.30 | $1.20 | Reasoning, Vision |
| Meta: Muse Glimmer 30B (batch) | Meta | 131K | $0.35 | $1.50 | Reasoning, Vision |
| Meta: Muse Spark 1.2 | Meta | 1049K | $1.25 | $4.25 | Reasoning, Vision |
| Qwen: Qwen3.8 Max | Qwen | 1000K | $2 | $6 | Reasoning, Vision |
| Thinking Machines: Inkling Small | Thinkingmachines | 524K | $0.45 | $1.20 | Reasoning, Vision |
| Thinking Machines: Inkling Small (batch) | Thinkingmachines | 524K | $0.50 | $1.20 | Reasoning, Vision |
| Thinking Machines: Inkling Small (free) | Thinkingmachines | 1049K | $0 | $0 | Reasoning, Vision |
| Qwen: Qwen3.7 Flash | Qwen | 1000K | $0.03 | $0.13 | Reasoning, Vision |
| Claude Opus 5 (Fast) | Anthropic | 1000K | $10 | $50 | Reasoning, Vision |
| Claude Opus 5 | Anthropic | 1000K | $5 | $25 | Reasoning, Vision |
| Claude Opus 5 (batch) | Anthropic | 1000K | $2.50 | $12.50 | Reasoning, Vision |
| Google: Gemini 3.6 Flash | 1049K | $0.75 | $3.75 | Reasoning, Vision | |
| Google: Gemini 3.6 Flash (batch) | 1049K | $0.38 | $1.88 | Reasoning, Vision | |
| Google: Gemini 3.5 Flash Lite | 1049K | $0.30 | $2.50 | Reasoning, Vision | |
| Google: Gemini 3.5 Flash Lite (batch) | 1049K | $0.15 | $1.25 | Reasoning, Vision | |
| Thinking Machines: Inkling | Thinkingmachines | 524K | $1 | $4.05 | Reasoning, Vision |
| Thinking Machines: Inkling (batch) | Thinkingmachines | 524K | $1 | $4.05 | Reasoning, Vision |
| Thinking Machines: Inkling (free) | Thinkingmachines | 1049K | $0 | $0 | Reasoning, Vision |
| MoonshotAI: Kimi K3 | Moonshotai | 1049K | $3 | $15 | Reasoning, Vision |
| MoonshotAI: Kimi K3 (batch) | Moonshotai | 1049K | $3 | $15 | Reasoning, Vision |
| Meta: Muse Spark 1.1 | Meta | 1049K | $1.25 | $4.25 | Reasoning, Vision |
| OpenAI: GPT-5.6 Luna Pro | OpenAI | 1050K | $0.20 | $1.20 | Reasoning, Vision |
| OpenAI: GPT-5.6 Luna | OpenAI | 1050K | $0.20 | $1.20 | Reasoning, Vision |
| OpenAI: GPT-5.6 Terra Pro | OpenAI | 1050K | $2 | $12 | Reasoning, Vision |
| OpenAI: GPT-5.6 Terra | OpenAI | 1050K | $2 | $12 | Reasoning, Vision |
| OpenAI: GPT-5.6 Sol Pro | OpenAI | 1050K | $2 | $10 | Reasoning, Vision |
| OpenAI: GPT-5.6 Sol | OpenAI | 1050K | $2 | $10 | Reasoning, Vision |
| SpaceXAI: Grok 4.5 | xAI | 500K | $2 | $6 | Reasoning, Vision |
| xAI: Grok Latest | ~x-ai | 500K | $2 | $6 | Reasoning, Vision |
| Anthropic: Claude Sonnet 5 | Anthropic | 1000K | $2 | $10 | Reasoning, Vision |
| Anthropic: Claude Sonnet 5 (batch) | Anthropic | 1000K | $1 | $5 | Reasoning, Vision |
| Nex AGI: Nex-N2-Mini | Nex-agi | 262K | $0.02 | $0.10 | Reasoning, Vision |
| Sakana: Fugu Ultra | Sakana | 1000K | $5 | $30 | Reasoning, Web search, Vision |
| MoonshotAI: Kimi K2.7 Code | Moonshotai | 262K | $0.66 | $3.40 | Reasoning, Vision |
| Anthropic: Claude Fable Latest | ~anthropic | 1000K | $10 | $50 | Reasoning, Vision |
| Anthropic: Claude Fable 5 | Anthropic | 1000K | $10 | $50 | Reasoning, Vision |
| Anthropic: Claude Fable 5 (batch) | Anthropic | 1000K | $5 | $25 | Reasoning, Vision |
Showing the top 50 of 217 models. Browse the full directory →
A vision (multimodal) model accepts images alongside text. Instead of describing a chart or error message in words, you attach the image and the model reads it directly — extracting text, interpreting layouts, and answering questions about what it sees.
217 models on CoreAI currently support image input, including the latest multimodal releases from Qwen, Z.AI, Meta, DeepSeek, Dots-studio. The table on this page is generated live, so it always matches what's available in the app.
Most modern vision models handle printed text and clear handwriting well. For multi-page PDFs, CoreAI's Scan & PDF tool splits documents into pages a vision model can process.
Image input usually consumes more tokens than plain text, so a vision request costs somewhat more than a text-only one on the same model — but pricing per token is identical; you simply send more tokens.
Chat with GPT-5, Claude, Gemini and hundreds more — all in one app.