Qwen 4THE FAMILY GUIDE

The Jade Atlas

Stealth Models ↗

UPCOMING · ALIBABA · THE QWEN4 FAMILY

Qwen 4.
Models &
benchmarks.

Qwen4-27B, Max, Flash and Plus are named. Qwen4-35B-A3B is the likely fifth. Explore release updates, benchmark forecasts, architecture and VRAM planning.

5EXPECTED MODELS
5BENCHMARK WINDOWS
In trainingALIBABA · 22 SEP

YOUR GUIDE TO WHAT COMES NEXT

Welcome, curious one. Let us see what the clues have brought.

A curious mind. An unusually distinguished nose.

ON THE RECORD

Qwen 4 is in training.

Alibaba’s Apsara announcement ↗

01 / THE FAMILY

The Qwen4 lineup.
What each name tells us.

Meet the announced names and the likely small MoE sibling. Each has its own file and a place in the benchmark outlook.

FOLLOWING Qwen4-27B

The open-weight trail.

Conference reporting puts Qwen4-27B first in line on the open-weight side.

Read the conference report ↗
Evidence
Conference report
Current status
In training
Benchmark view
Estimated ranges

A small name
can carry a large idea.

01

Qwen4-27B

Shown at ApsaraThe open-weight priority

The conference lineup puts 27B on the map; 36Kr reports that it is the first priority on the open-weight side.

The predecessor, Qwen3.8-27B, is a dense multimodal model. The new generation’s name does not yet tell us whether 27B counts only the backbone or includes a learned lookup table.

02

Qwen4-35B-A3B

LikelyThe small MoE lead

A secondhand testing lead points to an A3B model. Qwen4-35B-A3B is the likely name if the 35B tier returns.

A3B describes roughly three billion active parameters per token. That is different from total stored weights. The 35B total follows the earlier tier; the next checkpoint will settle its size.

03

Qwen4-Flash

Shown at ApsaraThe efficient tier to watch

Flash appears alongside Plus and Max in the conference lineup. The architecture preview makes this an especially interesting branch of the family.

Qwen3.8-Flash-Next has a 125B main model and a separate 51B n-gram table. Qwen4-Flash’s name does not lock in either number, or reserve that technology for Flash alone.

04

Qwen4-Plus

Shown at ApsaraThe middle of the lineup

Plus was displayed with Max, Flash and 27B on the Apsara screen. Its place between the familiar tier names makes it a natural model to compare across the five benchmark windows.

The chart treats its position as an editorial forecast anchored to earlier Qwen results. Its eventual price, speed and architecture will make this file more concrete.

05

Qwen4-Max

Shown at ApsaraThe flagship trail

Max is the largest-sounding name on the announced family screen. The benchmark atlas places its outlook alongside the frontier comparison models.

Alibaba’s 5–10 trillion parameter ambition refers to Qwen4.5 and Qwen5. It is a later roadmap, not a size announcement for Qwen4-Max.

Advertisement · Demodokos

Master Qwen knows too much.

One curious investigator. One very worried lab. Hear what slips out.

Master Qwen · 1 minute · four voices

Master QwenA very indiscreet master

The lab would prefer that he said nothing.

Meet the cast & watch in Foundry →

THE BENCHMARK ATLAS

Qwen4 benchmark outlook

Compare measured frontier and Qwen models with the estimated Qwen4 family. Each test has its own ranking.

Measured score Qwen4 editorial forecastHigher is better · score out of 100

Forecast bands use the earlier Qwen models as anchors. Each open diamond sits at the centre of its editorial range.

Family focus: all five forecasts

WINDOW 01

Humanity's Last Exam

Expert questions across many fields.

Text-only · no tools

Artificial Analysis methodology
Snapshot:

Scores, settings & sources
Humanity's Last Exam · percent correct / solved · higher is better
ModelScoreSetting / forecast basis
Claude Opus 5.5 (max)61.4%measuredAdaptive reasoning · max effort · default fallback Source scorecard
Claude Fable 5.1 (max)59.1%measuredAdaptive reasoning · max effort · default fallback Source scorecard
Claude Opus 5.5 (high effort)56%measuredAdaptive reasoning · high effort · default fallback Source scorecard
Qwen4-Max43–56%forecastMax 0902 to Opus high. Editorial method
Qwen3.8 Max (0902)43%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-Plus38–48%forecastFlash-Next to five points above Max 0902. Editorial method
Qwen4-Flash38–43%forecastFlash-Next to Max 0902. Editorial method
Qwen3.8-Flash-Next38%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-27B34–42%forecast27B to four points above Flash-Next. Editorial method
Qwen4-35B-A3B (likely)28–42%forecastSix below 27B to four above Flash-Next. Editorial method
Qwen3.8 27B (xhigh)34%measuredxhigh effort Source scorecard

WINDOW 02

GPQA Diamond

Graduate science questions.

Diamond · Epoch administered

Epoch AI methodology
Snapshot:

Scores, settings & sources
GPQA Diamond · percent correct / solved · higher is better
ModelScoreSetting / forecast basis
Qwen4-Max92–97%forecastMax 0902 plus zero to five points. Editorial method
Qwen3.8 Max (August)93%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen3.8 Max (0902)92%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-Plus89–95%forecastMax 0902 minus three to plus three. Editorial method
Qwen4-27B87–95%forecastA3B predecessor plus two to ten. Editorial method
Qwen4-Flash85–93%forecastA3B predecessor to August Max. Editorial method
Qwen4-35B-A3B (likely)80–94%forecastA3B predecessor minus five to plus nine. Editorial method
Qwen3.6 35B-A3B85%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard

WINDOW 03

FrontierMath Tiers 1–3

Research-level mathematics.

Tiers 1–3 · v2 private set

Epoch AI methodology
Snapshot:

Scores, settings & sources
FrontierMath Tiers 1–3 · percent correct / solved · higher is better
ModelScoreSetting / forecast basis
GPT-6 Astra (max)93.7%measuredmax effort Source scorecard
Claude Fable 5.1 (max)90.2%measuredmax effort Source scorecard
Qwen4-Max66–94%forecastMax 0902 to rounded Astra max. Editorial method
Qwen3.8 Max (August, xhigh)74.7%measuredxhigh effort Source scorecard
Qwen3.8 Max (0902)66%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-Plus45–78%forecastA3B predecessor plus 25 to 58 points. Editorial method
Qwen4-27B35–70%forecastA3B predecessor plus 15 to 50 points. Editorial method
Qwen4-Flash30–66%forecastA3B predecessor plus ten to Max 0902. Editorial method
Qwen4-35B-A3B (likely)20–70%forecastA3B predecessor plus zero to 50 points. Editorial method
Qwen3.6 35B-A3B20%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard

WINDOW 04

SciCode

Scientific programming problems.

Artificial Analysis implementation

Artificial Analysis methodology
Snapshot:

Scores, settings & sources
SciCode · percent correct / solved · higher is better
ModelScoreSetting / forecast basis
Claude Opus 5.5 (max)66.9%measuredAdaptive reasoning · max effort · default fallback Source scorecard
Claude Fable 5.1 (max)63.1%measuredAdaptive reasoning · max effort · default fallback Source scorecard
Claude Opus 5.5 (high effort)60%measuredAdaptive reasoning · high effort · default fallback Source scorecard
Qwen4-Max52–60%forecastMax 0902 to Opus high. Editorial method
Qwen4-Plus51–57%forecastFlash-Next to five points above Max 0902. Editorial method
Qwen4-Flash51–54%forecastFlash-Next to two points above Max 0902. Editorial method
Qwen3.8 Max (0902)52%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen3.8-Flash-Next51%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-27B47–54%forecast27B to three points above Flash-Next. Editorial method
Qwen3.8 27B (xhigh)47%measuredxhigh effort Source scorecard
Qwen4-35B-A3B (likely)40–54%forecastSeven below 27B to three above Flash-Next. Editorial method

WINDOW 05

Terminal-Bench 4.0

Practical work inside a terminal.

Version 4.0 · mini-swe-agent · pass@1, 3 repeats

Artificial Analysis methodology
Snapshot:

Scores, settings & sources
Terminal-Bench 4.0 · percent correct / solved · higher is better
ModelScoreSetting / forecast basis
Claude Opus 5.5 (max)59.6%measuredAdaptive reasoning · max effort · default fallback Source scorecard
GPT-6 Astra (xhigh)59.6%measuredxhigh effort Source scorecard
Claude Opus 5.5 (high effort)57%measuredAdaptive reasoning · high effort · default fallback Source scorecard
Qwen4-Max39–57%forecastMax 0902 to Opus high. Editorial method
Qwen3.8 Max (0902)39%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-Plus25–46%forecastFlash-Next to seven points above Max 0902. Editorial method
Qwen4-Flash25–39%forecastFlash-Next to Max 0902. Editorial method
Qwen3.8-Flash-Next25%measuredEvaluator-listed checkpoint; effort not specified in this scorecard. Source scorecard
Qwen4-27B6–29%forecast27B to four points above Flash-Next. Editorial method
Qwen4-35B-A3B (likely)3–31%forecastThree below 27B to six above Flash-Next. Editorial method
Qwen3.8 27B (xhigh)6%measuredxhigh effort Source scorecard
How the forecast ranges are set

These are our editorial estimates, dated 28 September 2026. Each range starts from a measured Qwen predecessor or the named comparison model. The score tables give the exact anchor and percentage-point adjustment for every estimate. The open diamond sits halfway between the endpoints.

Max spans the latest Max checkpoint to a strong comparison result. Plus assumes a step above Flash-Next. Flash stays near the preview. 27B starts near the existing 27B where a matched result is available. The likely 35B-A3B model gets a wider range because both its final size and performance remain open. GPQA and FrontierMath use Epoch predecessors; their 27B and Flash positions are broader tier assumptions. No cross-benchmark average is calculated.

HLE uses Artificial Analysis text-only results without tools. GPQA uses Epoch only. FrontierMath uses Epoch Tiers 1–3 v2. SciCode and Terminal-Bench use Artificial Analysis results. Checkpoint and published effort labels stay attached to every measured entry; a scorecard without an effort label is marked accordingly.

Qwen4-35B-A3B (likely) is a working name for the rumored A3B successor; 35B is a size hypothesis.

The case file · updated 28 September 2026

Qwen4 evidence, clue by clue

Four names appeared together. A fifth is taking shape in the margins. Follow the dated trail from public announcement to engineering trace to rumor.

  1. On the record

    Qwen4 is in training

    Alibaba names its next-generation model. The 5–10 trillion parameter ambition belongs to the later Qwen4.5 and Qwen5 roadmap.

    Alibaba's Apsara statement
  2. Photographed

    Four names on the screen

    The Apsara screen naming Qwen4-Max, Flash, Plus and 27BInspect the conference photo

    An Apsara screen reads “Qwen4 Series Coming Soon” above Max, Flash, Plus and 27B. It gives names, not final model configurations.

    Original attendee post
  3. On the record

    The architecture preview

    Qwen calls Qwen3.8-Flash-Next an early preview of architecture for Qwen4. Its 125B main model and 51B n-gram table describe this precursor.

    Qwen's architecture release
  4. Engineering trace

    Qwen4-Exp enters model code

    Transformers documents gated residual paths, sparse attention and n-gram features under an experimental architecture name. Code support is a preparation signal, not a final checkpoint ID.

    Transformers integration, merged 26 August ↗
  5. Reported

    27B moves toward open weights

    A conference account says the open-weight side is prioritizing Qwen4-27B. The eventual license and model card will settle what ships.

    36Kr conference report
  6. Likely · size hypothesis

    Qwen4-35B-A3B: the likely fifth

    A secondhand testing claim points to a small A3B model. Qwen4-35B-A3B is the likely name if the 35B tier returns; the reported test did not specify its total size.

    Earlier official 35B-A3B card
  7. Unverified attribution

    The pelican and voxel temple

    The circulating screenshot containing voxel temple and pelican demosInspect the circulating screenshot

    Screenshots circulated through a Tieba-to-X repost chain with a Qwen4 label. They show no model selector, API ID or repeatable transcript. Reposts follow one trail.

    Circulating screenshot post
  8. Analyst assessment

    CSC Financial: Qwen4 release imminent

    A CSC Financial (中信建投) research note says “Qwen4发布在即”: “Qwen4 release is imminent.” First Financial’s report, carried by Eastmoney, gives no release date or Alibaba launch confirmation.

    First Financial via Eastmoney
  9. Next signal

    The next place to look

    A released configuration can establish which Qwen4 members use n-gram tables, what “27B” counts, and the final Flash size. Until then, the precursor is a design clue.

    Watch Qwen's official model catalog

Architecture preview · Qwen3.8-Flash-Next

More memory.
Less work per token.

Qwen opened a working architectural precursor in August. It reveals the engineering direction; each finished Qwen4 model still needs its own configuration sheet.

125BMain model6B main-model parameters active per token
51BLearned n-gram tableAbout 29% of the 176B combined total

These are Qwen3.8-Flash-Next's published figures. Final Qwen4 family sizes are the next piece of this puzzle.

A phrase-sized memory, fetched by address

A normal embedding looks up a vector for one token. The preview also uses the current token and several preceding tokens to select learned vectors for local patterns. The model then uses those vectors while doing its main work.

acupoftea

A short sequence supplies the lookup address.

  1. Current + preceding tokensShort local context
  2. Lookup addressDeterministic selection
  3. Selected vectorA small slice of the learned table
  4. Main modelReceives the local-pattern signal

Host RAM: Qwen says the full table can stay here and the needed entries can be prefetched.

GPU: selected entries move to the accelerator while the main network computes. RAM capacity, transfer bandwidth and concurrency still matter.

“The book stays on the shelf. I bring the useful page to the desk.”

The 51B entries are stored parameters, not 51B more activated matrix multiplications on every token. Qwen's embedding explanation

GDN + QSA Remember, then retrieve

In the precursor, three of every four layers use Gated DeltaNet to compress history into a fixed-size state. The fourth uses Qwen Sparse Attention to choose useful earlier micro-blocks, limiting long-context work while retaining targeted access.

Four gated streams Carry useful signals deeper

Gated Residual expands one residual path into four branches. Content-dependent gates control what each layer reads and writes; these are paths inside one model, not four models.

Muon Shape the training

Qwen uses Muon on suitable weight matrices in attention, GDN and experts, while embeddings, the router and low-rank residual weights use AdamW. This is a training method, not an inference setting.

Multi-token prediction Look ahead

The precursor trains extra prediction steps to improve speculative decoding acceptance. Its MTP layers also use QSA. Final Qwen4 serving paths have yet to be specified.

All four mechanisms are described for Qwen3.8-Flash-Next, the Qwen4 architecture preview.

27B ON YOUR DESK

27B weights, VRAM and context

Start with Q5_K for quality. Reach for IQ3_K when memory is tight. Keep Q8_0 for a roomier GPU and disk. Then choose how much memory to give the conversation.

Qwen4-27B planning guide · Qwen3.8-27B reference sizes and cache layout.

Q5_K_M and Q8_0: published Qwen3.8 files. IQ3_K: estimated mixed-precision size, using ik_llama.cpp’s IQ3_K format. Use an IQ3_K-compatible runtime.

Make room for your context

Try q4_0 KV cache before lowering the weight quant. It saves about 72% of KV tensor memory versus F16. If long-context recall suffers, move the cache back to q8_0.

F162.00 GiB
q8_01.06 GiB
q4_00.56 GiB

q4_0 saves 1.44 GiB at 32K context.

FULL GPU OFFLOAD · ONE TEXT SESSION

≈ 21.0 GiB planned
Weights
18.4 GiB
Attention KV cache
0.56 GiB
Runtime allowance
2.0 GiB

3.0 GiB left in a 24 GiB budget.

The 2 GiB allowance budgets for recurrent state, compute buffers and GPU headroom. It is a planning reserve, not a measured peak.

Keep Q5_K. Shrink the cache.

At 128K context, this reference cache drops from 8 GiB in F16 to 2.25 GiB in q4_0. That frees 5.75 GiB while leaving the model’s weight quantization untouched.

What the memory plan includes

The cache estimate uses Qwen3.8-27B’s configuration: 16 full-attention layers, four KV heads and 256 values per head, with both K and V cached. Its recurrent state sits in the runtime allowance. Quantization scales are included: q8_0 uses 34 bytes per 32 values; q4_0 uses 18.

Weight storage approximates full GPU weight residency. Vision components, multiple sessions, larger work buffers or an additional Qwen4 lookup table need their own budget. Host-resident weights and lookup tables consume system RAM. Partial CPU offload can fit a larger file on a smaller GPU, with a speed tradeoff.

Files use decimal GB; memory uses GiB. IQ3_K’s 12–13 GB range allows for tensors stored at higher precision. Final Qwen4 GGUF sizes and its cache layout will replace these reference estimates.

Qwen 27B q4_0 cache quality tests · llama.cpp cache settings

Advertisement · Demodokos

Every time, Master Qwen.

Cast the voices, switch languages and score the scene. Create it all locally in Demodokos Foundry.

InvestigatorMaster QwenLab AdvisorFoundry Narrator
DEMODOKOS FOUNDRYTHE ORIGINAL STUDIO TAKE

Master Qwen

An original Demodokos scene
Master QwenJade Pavilion

Some secrets are difficult to keep.

Master QwenAt the pavilion

Some secrets are difficult to keep.

Master Qwen

Some secrets are difficult to keep.

LAB ADVISORSTANDBY

The lab is listening.

Read the dialogue

RELEASE WATCH

Qwen4 release date and access

Qwen4 is in training. Follow the official release channels for weights, API model IDs and the first configuration files.

The weight shelf

Qwen’s official model collection is the place to watch for downloadable checkpoints, licenses and the parameter breakdown.

Qwen on Hugging Face ↗

The engineering trail

The Qwen4-Exp integration exposes the architecture preview. A final model configuration will turn the remaining architecture questions into specifications.

Explore Qwen4-Exp ↗

The official announcements

Release dates and availability belong here first. The 22 September Apsara statement is the current family announcement in this guide.

Qwen’s official site ↗

The working precursor

Qwen3.8-Flash-Next already offers an open look at the new design. Its published model card and technical report make useful reading while Qwen4 trains.

Read the preview model card ↗

Qwen4 questions, answered

Is n-gram embedding only for Qwen4-Flash?

That has not been established. Qwen documents n-grams in Qwen3.8-Flash-Next, an early Qwen4 architecture preview. No released Qwen4 family configuration assigns them exclusively to Flash or to every member. Read the Qwen preview.

Will Qwen4-Flash be a 125B continuation of Qwen3.8-Flash?

It may continue the Flash role and architectural direction. The published 125B main model plus 51B n-gram table belongs to Qwen3.8-Flash-Next. Qwen has not published a Qwen4-Flash parameter count. See the precursor specifications.

Is Qwen4-35B-A3B confirmed?

No final model identifier or total size has been published. An A3B test is a secondhand lead; 35B follows the earlier Qwen3.6-35B-A3B naming pattern. It stays in the fifth watchlist position as a likely hypothesis.

How much GPU memory will Qwen4-27B need?

Start planning around Q5_K_M on a 24 GB-class GPU, IQ3_K on 16 GB, or Q8_0 on 32 GB. The current 27B reference puts Q5_K_M weights at about 18.4 GiB; a q4_0 cache and runtime allowance bring a 32K text session to roughly 21 GiB. These are Qwen3.8-based planning figures for the upcoming Qwen4 model. Compare quants, context and KV-cache memory.

KEEP EXPLORING

More model trails.

Browse the model directory ↗