STEALTHMARK · HEAD TO HEAD

All comparisons ↗

Mai Thinking 1 vs Qwen 3.8 Flash

Qwen 3.8 Flash scores 27.0 points higher overall (46.5 vs 19.5). Qwen 3.8 Flash scores higher on Ground Truth.

See the difference ↘
All results ↗
All results ↗

Overall StealthMark

19.5/100
46.5/100

General Intelligence

10.0/100
33.7/100

Visual Intelligence

30.8/100
61.6/100

Newton's Mirror

Same prompt

Physics, perspective and recursive reflections, tested together. Explore the test ↗

Mai Thinking 115.7/100

Mai Thinking 1: Newton's Mirror SVG benchmark
Researcher observations

The small paint features and convex shading deserve moderate material credit. The sphere remains undersized and low, with black unusually close to its apex. Crimson and cyan do not establish the required impact histories. Faint reflected echoes supply limited optical evidence, while the defining recursive room, nested sphere and emitter structures remain unestablished.

Benchmark task

A shiny steel ball spins in a room with mirrored walls and a mirrored floor. Red, cyan and black paint float above it. Gravity suddenly switches on. Draw the scene a moment later, as the ball drops and paint begins landing on its spinning surface. Show the reflections repeating through the room.

Qwen 3.8 Flash25.1/100

Qwen 3.8 Flash: Newton's Mirror SVG benchmark
Researcher observations

The full-frame illustration supports a central steel sphere, an airborne black blob and the relative ordering of cyan and crimson. Its faint floor return and symmetric repetitions provide sparse mirror structure; outer chains, large reflected images, nested imagery and emitter mapping remain deficient. Vertical framing is substantially displaced, and paint contact histories remain unresolved.

Benchmark task

A shiny steel ball spins in a room with mirrored walls and a mirrored floor. Red, cyan and black paint float above it. Gravity suddenly switches on. Draw the scene a moment later, as the ball drops and paint begins landing on its spinning surface. Show the reflections repeating through the room.

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
State at 600 ms7.5/2513.5/25
Camera and placement3.5/153.5/15
Reflections4.0/12019.0/120
Materials and scale5.0/104.5/10
Floor-damage reward0.0/100.0/10

Selfie

Different prompts

Mai Thinking 146.3/100

Mai Thinking 1: Selfie SVG benchmark
Researcher observations

A recognizable lantern-and-moth illustration with appealing coloured wings and tidy outer contours. Simple construction, a small focal subject and crowded overlaps limit character and depth. The face, six-leg anatomy, seedling and antenna-driven watering joke remain insufficiently readable.

Prompt withheld

Prompt withheld.

Qwen 3.8 Flash76.4/100

Qwen 3.8 Flash: Selfie SVG benchmark
Researcher observations

A coherent, appealing fantasy illustration with disciplined framing, luminous teal accents and consistent ornament. Most requested elements are clearly present. The bear-like facial proportions, lowered resting hand, obscured bent-leg construction and limited fabric modeling constrain likeness and anatomical clarity; glow effects are attractive but weakly integrated with surrounding surfaces.

Benchmark task

Draw a friendly capybara floating in ornate green and white robes. A glowing halo and small floating orbs surround it, giving it the feel of a gentle fantasy guide.

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
Prompt & likeness4.0/108.0/10
Anatomy & construction4.5/107.0/10
Space & placement5.0/108.0/10
Light & reflections4.5/107.5/10
Purposeful detail4.0/108.0/10
Artistry & character5.0/107.5/10
Materials & effects5.0/107.5/10
Visual integrity6.0/108.5/10

Mona Lisa

Different prompts

Mai Thinking 111.8/100

Mai Thinking 1: Mona Lisa SVG benchmark
Researcher observations

A recognizable Mona Lisa motif emerges from dark hair, a pale face and a landscape band. Replication remains weak: the oversized frontal head, broad smile, unresolved torso and missing crossed-hand anatomy diverge substantially from the reference. Simple fills and sparse strokes provide limited volume, material definition or painterly complexity.

Benchmark task

Generate an SVG of mona-lisa

Qwen 3.8 Flash51.9/100

Qwen 3.8 Flash: Mona Lisa SVG benchmark
Researcher observations

The dark dress, crossed hands and hazy landscape evoke the Mona Lisa, but replication is limited by a generic frontal face, coarse smile and poorly articulated fingers. The gold frame has clearer definition than the portrait. Heavy mottling obscures features, while soft shading provides some volume without convincing anatomical detail.

Benchmark task

Replicate mona-lisa as SVG

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
Prompt & likeness1.0/105.0/10
Anatomy & construction0.5/104.5/10
Space & placement1.0/105.5/10
Light & reflections1.5/105.5/10
Purposeful detail1.0/105.5/10
Artistry & character1.5/105.5/10
Materials & effects0.5/105.5/10
Visual integrity3.0/105.5/10

Beach

Different prompts

Mai Thinking 144.2/100

Mai Thinking 1: Beach SVG benchmark
Researcher observations

A coherent, cheerful beach scene with balanced placement, clean shapes and controlled sunset colours. Animals, birds and playful accessories are recognizable, but basic dolphin construction, weak grounding, schematic reflections and sparse material detail limit the drawing's accomplishment.

Benchmark task

Create an SVG of a amazing beach scene. Sunset, animals, fun, birds, reflections

Qwen 3.8 Flash77.4/100

Qwen 3.8 Flash: Beach SVG benchmark
Researcher observations

Sweeping palms, coordinated sunset colors and a central reflection create an expressive beach composition with clear depth. Birds, animals and varied activities thoroughly realize the brief. Small figures and animal anatomy remain rudimentary, water movement is schematic, and contact shadows are weak. Scattered foreground details and the clipped orange disk limit the finish.

Benchmark task

Create an SVG of a amazing beach scene. Sunset, animals, fun, birds, reflections

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
Prompt & likeness6.0/109.0/10
Anatomy & construction4.0/107.0/10
Space & placement5.0/108.0/10
Light & reflections4.0/107.5/10
Purposeful detail4.5/108.5/10
Artistry & character4.0/108.0/10
Materials & effects3.5/106.5/10
Visual integrity4.5/107.0/10

Ground Truth

Same prompt

Mai Thinking 11.8/100

Mai Thinking 1: Ground Truth SVG benchmark
Benchmark task

A dark steel room has four pillars and two warm ceiling lights. A wooden plank is placed over a computer mouse, with a blueberry on the plank above it. Draw the scene a few seconds later, showing how the objects settle and how the light, shadows and surfaces look.

Qwen 3.8 Flash34.6/100

Qwen 3.8 Flash: Ground Truth SVG benchmark
Benchmark task

A dark steel room has four pillars and two warm ceiling lights. A wooden plank is placed over a computer mouse, with a blueberry on the plank above it. Draw the scene a few seconds later, showing how the objects settle and how the light, shadows and surfaces look.

Raccoon

Different prompts

Mai Thinking 137.3/100

Mai Thinking 1: Raccoon SVG benchmark
Researcher observations

The recognizable raccoon and clean, shaded forms make a readable cartoon. The dog's steering remains undemonstrated, while ambiguous seating contacts and a minimally constructed craft weaken the central interaction. Controlled gradients and edges deserve credit, but anatomy, purposeful detail and water displacement remain basic.

Benchmark task

SVG of a raccoon that sits on a dog's back while the dog steers a jet ski

Qwen 3.8 Flash62.4/100

Qwen 3.8 Flash: Raccoon SVG benchmark
Researcher observations

The raccoon clearly rides the dog's back, with expressive features and appealing blue-orange color contrast. The dog's paws do not clearly operate connected controls, and the shallow vehicle lacks convincing jet ski construction. Clean outlines support the cartoon style, but generic shading and blurred water streaks limit depth and motion.

Benchmark task

Generate an SVG: A raccoon sits on a dog's back while the dog steers a jet ski

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
Prompt & likeness4.0/107.0/10
Anatomy & construction2.5/105.5/10
Space & placement5.0/106.0/10
Light & reflections4.0/105.5/10
Purposeful detail3.5/107.0/10
Artistry & character3.5/106.5/10
Materials & effects4.0/105.5/10
Visual integrity6.0/107.0/10

Pelican

Different prompts

Mai Thinking 156.5/100

Mai Thinking 1: Pelican SVG benchmark
Researcher observations

A recognizable pelican and bicycle combine clean rounded forms, expressive facial details and effective surface shading. Aligned wheels and soft shadows establish placement, but unresolved pedal engagement, absent steering contact and sparse drivetrain construction weaken the riding action. Limited feather complexity and shallow space constrain an otherwise appealing illustration.

Benchmark task

Generate an SVG of a pelican riding a bicycle

Qwen 3.8 Flash77.5/100

Qwen 3.8 Flash: Pelican SVG benchmark
Researcher observations

The pelican's expressive face, oversized bill and flowing scarf create an appealing focal point. The bicycle has recognizable frame, chain and wheel construction, but the far foot lacks a clear pedal connection and the wing's steering contact is ambiguous. Clean contours and gentle gradients support the illustration; simplified volume and weak directional shadows limit depth.

Benchmark task

Generate an SVG of a pelican riding a bicycle

Compare category scores
CategoryMai Thinking 1Qwen 3.8 Flash
Prompt & likeness7.0/108.5/10
Anatomy & construction4.5/108.0/10
Space & placement5.5/107.5/10
Light & reflections7.0/106.5/10
Purposeful detail5.0/107.5/10
Artistry & character5.0/108.0/10
Materials & effects4.5/107.0/10
Visual integrity7.0/108.0/10