Qwen/Qwen3-4B-Thinking-2507

Run locally on Apple devices with Mirai

Type
local
From
Alibaba
Quantization
No
Parameters
4B
Size
7.5 GB
Source
Hugging Face

Qwen3-4B-Thinking-2507 is a 4-billion-parameter causal language model from the Qwen team, purpose-built for deep, extended reasoning. This updated release represents three months of focused scaling on thinking capability, delivering substantially stronger performance across logical reasoning, mathematics, science, and coding — all within a compact model footprint.

Benchmark comparison of Qwen3-4B-Thinking-2507

Architecture & Specifications

  • Parameters: 4.0B total (3.6B non-embedding)
  • Layers: 36, with grouped-query attention (32 Q heads, 8 KV heads)
  • Native context length: 262,144 tokens
  • Training: Pretraining + post-training pipeline
  • Mode: Thinking-only — the model always reasons before responding, with no need to set `enable_thinking=True`

Key Strengths

The model excels across a broad range of tasks, with standout improvements over its predecessor (Qwen3-4B Thinking):

  • Math & Reasoning: Achieves 81.3 on AIME25 and 55.5 on HMMT25, surpassing even the larger Qwen3-30B-A3B on these benchmarks.
  • Alignment & Instruction Following: Scores 87.4 on IFEval and 83.3 on WritingBench, reflecting strong adherence to user intent.
  • Agentic & Tool Use: Dramatic gains on agent benchmarks (TAU series), making it well-suited for multi-step tool-calling workflows via frameworks like Qwen-Agent and MCP-compatible toolchains.
  • Multilingual: Improved performance across MultiIF, PolyMATH, and other multilingual evaluations.
  • Long Context: Enhanced 256K context understanding for document-scale reasoning.

Recommended Use

This model is best suited for complex reasoning tasks where extended chain-of-thought is valuable — competition-level math, scientific analysis, code generation, and agentic pipelines. For optimal results, the Qwen team recommends output lengths of 32,768 tokens for general queries and up to 81,920 for highly challenging problems. Deployment is supported via SGLang, vLLM, Ollama, LMStudio, and llama.cpp.

Explore all local models
1
Choose framework
2
Run the following command to install Mirai SDK
spm https://github.com/trymirai/uzu.git
3
Apply code
1import Uzu23public func runChat() async throws {4    let engineConfig = EngineConfig.create()5    let engine = try await Engine.create(config: engineConfig)67    guard let model = try await engine.model(identifier: "Qwen/Qwen3-4B-Thinking-2507") else {8        return9    }10    for try await update in try await engine.download(model: model).iterator() {11        print("Download progress: \(update.progress())")12    }1314    let messages = [15        ChatMessage.system().withText(text: "You are a helpful assistant"),16        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")17    ]18    let session = try await engine.chat(model: model, config: .create())19    let stream = await session.replyWithStream(input: messages, config: .create())20    var message: ChatMessage? = nil21    for try await update in stream.iterator() {22        switch update {23        case .replies(let replies):24            message = replies.last?.message25        case .error(let error):26            print("Error: \(error)")27        }28    }29    print("Text: \(message?.text() ?? "empty")")30}

Other local models from Alibaba