Text / Qwen
Qwen3.8-Flash-Next
An experimental open-weight multimodal model with sparse attention and 262K native context.
Not yet testedRepository opened Aug 24, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.
At a glance
- Parameters
- 125B language · 6B active
- Architecture
- Hybrid sparse attention
- Context length
- 262,144 native
- License
- Qwen Community License 1.0
- Recommended VRAM
- Not yet tested
- Minimum tested VRAM
- Not yet tested
- Disk space
- Not yet tested
- Software
- Transformers, vLLM, SGLang
HARDWARE ESTIMATE
What will it need?
Our test notes
Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.
Our verdict
Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.
Explore its uses
REASONINGMULTIMODALLONG CONTEXT
Keep exploring
Text
DeepSeekDeepSeek V4.1 Flash
A multimodal reasoning model with a compressed key-value cache, published as open weights.
MODEL SIZENot specified
Text
QwenQwen3 8B
An open-weight language model with switchable thinking and non-thinking modes.
MODEL SIZE8.2B
Text
DeepSeekDeepSeek R1 · 14B
A distilled reasoning model built on Qwen2.5, for exploring reasoning on your own infrastructure.
MODEL SIZE14B