stealthmodels.

MODEL COMPARISON

SVG benchmarks

Test how models handle anatomy, spatial relationships and physics in SVG. Compare StealthMark scores and see where strong results on familiar tasks fail to carry across.

28 models171 resultsHow StealthMark works

StealthMark SVG score

StealthMark measures how well models turn written instructions into coherent SVG scenes. Drawing the right objects is only part of the task: bodies must connect, a paw must reach the handlebars, and reflections must agree with the geometry.

Two views of intelligence

General Intelligence

Physical, causal and spatial reasoning. Newton’s Mirror carries the greatest influence: its physical state, perspective, recursive reflections and floor-damage deduction are combined with spatial judgements from the other scenes.

Visual Intelligence

Precise visual reconstruction, anatomy, materials, detail and artistic expression. Mona Lisa, Beach, Selfie and Raccoon provide the main evidence. Newton contributes its material-rendering judgement.

Both scores use the existing judged components, weighted by their importance within each benchmark and by the benchmark’s importance in the suite. Repeated takes share their benchmark’s influence. Pelican is a benchmax control and contributes to neither score.

A harder test: Newton’s Mirror

Newton’s Mirror tests physical state, camera placement and recursive reflections in one scene. These requirements interact: an error in the geometry changes what the mirrors should show. It tests whether a model can carry physical and spatial reasoning through to the finished image.

Best current Newton result: 57.8/100

Model-blind judging

Initial AI reviews assess each image separately, without the generating model’s identity or other models’ results. Judges examine the rendered artwork against its task. Researcher review can correct missed defects; the published category scores and written observations show what supports each rating.

How scenes contribute to StealthMark

Scores run from 0 to 100. The overall score combines performance across the rated scene types, with Newton’s Mirror carrying the greatest weight. Repeated takes do not give a scene extra weight. A model needs consistent results across different tasks to score well overall.

An Estimated label means the overall score includes an estimate for one missing scene. That scene remains unrated until an actual result is judged.

How we check for benchmaxxing

Benchmaxxing means optimizing for a familiar test without comparable gains elsewhere. Benchmaxx measures an unusually large advantage on the Pelican task compared with the model’s other drawing tasks. When flagged, the Pelican result is excluded from StealthMark; its artwork and individual score remain visible.

The percentage measures the strength of that imbalance, not the probability of training contamination. We also withhold selected prompts to reduce direct optimization for the test.

What the artwork scores cover

Prompt & likeness

Subjects, actions and likeness.

Anatomy & construction

Coherent bodies, joints and machinery.

Space & placement

Perspective, scale and contact.

Light & reflections

Lighting, shadows and reflections.

Purposeful detail

Clear, purposeful small features.

Artistry & character

Composition, expression and drawing skill.

Materials & effects

Surfaces, texture and movement.

Visual integrity

Clean shapes, layers and edges.

Newton’s Mirror has its own physics categories, shown with each result.

ARTWORK

Compare results

This link opens the same results in the same order.

StealthMark benchmark results

Reset benchmark preferences?

Return to the default view, clear comparison selections and restore the advertisement.