Qwen4-27B, Max, Flash and Plus are named. Qwen4-35B-A3B is the likely fifth. Explore release updates, benchmark forecasts, architecture and VRAM planning.
The conference lineup puts 27B on the map; 36Kr reports that it is the first priority on the open-weight side.
The predecessor, Qwen3.8-27B, is a dense multimodal model. The new generation’s name does not yet tell us whether 27B counts only the backbone or includes a learned lookup table.
A secondhand testing lead points to an A3B model. Qwen4-35B-A3B is the likely name if the 35B tier returns.
A3B describes roughly three billion active parameters per token. That is different from total stored weights. The 35B total follows the earlier tier; the next checkpoint will settle its size.
Flash appears alongside Plus and Max in the conference lineup. The architecture preview makes this an especially interesting branch of the family.
Qwen3.8-Flash-Next has a 125B main model and a separate 51B n-gram table. Qwen4-Flash’s name does not lock in either number, or reserve that technology for Flash alone.
Plus was displayed with Max, Flash and 27B on the Apsara screen. Its place between the familiar tier names makes it a natural model to compare across the five benchmark windows.
The chart treats its position as an editorial forecast anchored to earlier Qwen results. Its eventual price, speed and architecture will make this file more concrete.
These are our editorial estimates, dated 28 September 2026. Each range starts from a measured Qwen predecessor or the named comparison model. The score tables give the exact anchor and percentage-point adjustment for every estimate. The open diamond sits halfway between the endpoints.
Max spans the latest Max checkpoint to a strong comparison result. Plus assumes a step above Flash-Next. Flash stays near the preview. 27B starts near the existing 27B where a matched result is available. The likely 35B-A3B model gets a wider range because both its final size and performance remain open. GPQA and FrontierMath use Epoch predecessors; their 27B and Flash positions are broader tier assumptions. No cross-benchmark average is calculated.
HLE uses Artificial Analysis text-only results without tools. GPQA uses Epoch only. FrontierMath uses Epoch Tiers 1–3 v2. SciCode and Terminal-Bench use Artificial Analysis results. Checkpoint and published effort labels stay attached to every measured entry; a scorecard without an effort label is marked accordingly.
Qwen4-35B-A3B (likely) is a working name for the rumored A3B successor; 35B is a size hypothesis.
The case file · updated 28 September 2026
Qwen4 evidence, clue by clue
Four names appeared together. A fifth is taking shape in the margins. Follow the dated trail from public announcement to engineering trace to rumor.
On the record
Qwen4 is in training
Alibaba names its next-generation model. The 5–10 trillion parameter ambition belongs to the later Qwen4.5 and Qwen5 roadmap.
Transformers documents gated residual paths, sparse attention and n-gram features under an experimental architecture name. Code support is a preparation signal, not a final checkpoint ID.
A secondhand testing claim points to a small A3B model. Qwen4-35B-A3B is the likely name if the 35B tier returns; the reported test did not specify its total size.
Screenshots circulated through a Tieba-to-X repost chain with a Qwen4 label. They show no model selector, API ID or repeatable transcript. Reposts follow one trail.
A CSC Financial (中信建投) research note says “Qwen4发布在即”: “Qwen4 release is imminent.” First Financial’s report, carried by Eastmoney, gives no release date or Alibaba launch confirmation.
A released configuration can establish which Qwen4 members use n-gram tables, what “27B” counts, and the final Flash size. Until then, the precursor is a design clue.
Qwen opened a working architectural precursor in August. It reveals the engineering direction; each finished Qwen4 model still needs its own configuration sheet.
125BMain model6B main-model parameters active per token
+
51BLearned n-gram tableAbout 29% of the 176B combined total
A normal embedding looks up a vector for one token. The preview also uses the current token and several preceding tokens to select learned vectors for local patterns. The model then uses those vectors while doing its main work.
acupoftea
A short sequence supplies the lookup address.
Current + preceding tokensShort local context
Lookup addressDeterministic selection
Selected vectorA small slice of the learned table
Main modelReceives the local-pattern signal
Host RAM: Qwen says the full table can stay here and the needed entries can be prefetched.
GPU: selected entries move to the accelerator while the main network computes. RAM capacity, transfer bandwidth and concurrency still matter.
“The book stays on the shelf. I bring the useful page to the desk.”
The 51B entries are stored parameters, not 51B more activated matrix multiplications on every token. Qwen's embedding explanation
GDN + QSA Remember, then retrieve
In the precursor, three of every four layers use Gated DeltaNet to compress history into a fixed-size state. The fourth uses Qwen Sparse Attention to choose useful earlier micro-blocks, limiting long-context work while retaining targeted access.
Four gated streams Carry useful signals deeper
Gated Residual expands one residual path into four branches. Content-dependent gates control what each layer reads and writes; these are paths inside one model, not four models.
Muon Shape the training
Qwen uses Muon on suitable weight matrices in attention, GDN and experts, while embeddings, the router and low-rank residual weights use AdamW. This is a training method, not an inference setting.
Multi-token prediction Look ahead
The precursor trains extra prediction steps to improve speculative decoding acceptance. Its MTP layers also use QSA. Final Qwen4 serving paths have yet to be specified.
All four mechanisms are described for Qwen3.8-Flash-Next, the Qwen4 architecture preview.
27B ON YOUR DESK
27B weights, VRAM and context
Start with Q5_K for quality. Reach for IQ3_K when memory is tight. Keep Q8_0 for a roomier GPU and disk. Then choose how much memory to give the conversation.
Qwen4-27B planning guide · Qwen3.8-27B reference sizes and cache layout.
Keep Q5_K. Shrink the cache.
At 128K context, this reference cache drops from 8 GiB in F16 to 2.25 GiB in q4_0. That frees 5.75 GiB while leaving the model’s weight quantization untouched.
What the memory plan includes
The cache estimate uses Qwen3.8-27B’s configuration: 16 full-attention layers, four KV heads and 256 values per head, with both K and V cached. Its recurrent state sits in the runtime allowance. Quantization scales are included: q8_0 uses 34 bytes per 32 values; q4_0 uses 18.
Weight storage approximates full GPU weight residency. Vision components, multiple sessions, larger work buffers or an additional Qwen4 lookup table need their own budget. Host-resident weights and lookup tables consume system RAM. Partial CPU offload can fit a larger file on a smaller GPU, with a speed tradeoff.
Files use decimal GB; memory uses GiB. IQ3_K’s 12–13 GB range allows for tensors stored at higher precision. Final Qwen4 GGUF sizes and its cache layout will replace these reference estimates.
The Qwen4-Exp integration exposes the architecture preview. A final model configuration will turn the remaining architecture questions into specifications.
Qwen3.8-Flash-Next already offers an open look at the new design. Its published model card and technical report make useful reading while Qwen4 trains.
That has not been established. Qwen documents n-grams in Qwen3.8-Flash-Next, an early Qwen4 architecture preview. No released Qwen4 family configuration assigns them exclusively to Flash or to every member. Read the Qwen preview.
Will Qwen4-Flash be a 125B continuation of Qwen3.8-Flash?
It may continue the Flash role and architectural direction. The published 125B main model plus 51B n-gram table belongs to Qwen3.8-Flash-Next. Qwen has not published a Qwen4-Flash parameter count. See the precursor specifications.
Is Qwen4-35B-A3B confirmed?
No final model identifier or total size has been published. An A3B test is a secondhand lead; 35B follows the earlier Qwen3.6-35B-A3B naming pattern. It stays in the fifth watchlist position as a likely hypothesis.
How much GPU memory will Qwen4-27B need?
Start planning around Q5_K_M on a 24 GB-class GPU, IQ3_K on 16 GB, or Q8_0 on 32 GB. The current 27B reference puts Q5_K_M weights at about 18.4 GiB; a q4_0 cache and runtime allowance bring a 32K text session to roughly 21 GiB. These are Qwen3.8-based planning figures for the upcoming Qwen4 model. Compare quants, context and KV-cache memory.