stealthmodels.

MODEL FIELD GUIDE

DeepSeek V4 Pro 0813: Ground Truth

MIT-licensed reasoning model with a million-token context and 49B active parameters. Its 1.6T-parameter checkpoint targets large GPU servers.

Drawn by this model01 / 01
Ground Truth benchmark result generated by this model
SVG BENCHMARKGround TruthView result

A good fit for

Coding agents and research workflows that need long context, with a choice of hosted access or control over the weights.

Context window
1,000,000tokens[4]
Active parameters
49Bparameters[4]
License
MIT[3]
Model factsArchitecture, context & capabilities
Developer
DeepSeek-AI[1]
Origin
China[2]
Family
DeepSeek V4[1]
Context window
1,000,000 tokens[4]
Maximum output
393,216 tokens[4]
License
MIT[3]
Architecture
Mixture-of-Experts with DeepSeek V4 architecture and DSpark speculative-decoding module[1] [4]
Total parameters
1.6T[4]
Active parameters per token
49B[4]
Input
text[4]
Output
text[4]
API access & pricingAlibaba Cloud Model Studio · $1.272 input / $3.816 output · per 1M tokens

Alibaba Cloud Model Studio; input 1.272 USD/1M; output 3.816/1M; Global busy-period rate; idle-period rate is lower at $0.636 input / $1.908 output where documented.

[6] [4]
  • Alibaba Cloud Model StudioObserved 2026-10-06
    Input $1.272Output $3.816

    per 1M tokens

    Available

    deepseek-v4-pro-0813[6] [4]
Run it locallyWeights, memory & deployment

rack-scale

[1] [5]
  • Unsloth UD-Q4_K_XL GGUF850 GB storage · ~4-bit[5]
Measurements & sources7 primary references · Speed & serving conditions

StealthMark score

StealthMark measures fluid and visual intelligence of AI models.

General and Visual

General Intelligence

Perspective, anatomy, correct shadows, purposeful detail and physical consistency.

Visual Intelligence

Likeness, anatomy, detail, light, materials, composition and artistic expression.

Judging

Frontier Models blindly judge the AI generation unbiased through an intricate process. Our researchers control for mistakes.

Overall score

Overall combines weighted task scores on a scale from 0 to 100. Repeating a task does not increase its weight. Estimated means the total includes an estimate for one missing scene.

Benchmaxxing

Benchmaxx flags models that score unusually well on the familiar Pelican task compared with their other benchmark results. The percentage measures that imbalance, not the probability of training contamination. Flagged Pelican scores are excluded from Overall. Pelican never counts toward General or Visual.

We keep some prompts private to discourage test-specific optimization.

Scoring categories

Prompt & likeness

Subjects, actions and likeness.

Anatomy & construction

Coherent bodies, joints and machinery.

Space & placement

Perspective, scale and contact.

Light & reflections

Lighting, shadows and reflections.

Purposeful detail

Clear, purposeful small features.

Artistry & character

Composition, expression and drawing skill.

Materials & effects

Surfaces, texture and movement.

Visual integrity

Clean shapes, layers and edges.

BENCHMARK RESULTS

Compare results

This link opens the same results in the same order.

StealthMark benchmark results

Reset benchmark preferences?

Return to the default view, clear comparison selections and restore the advertisement.