Vendor
Alibaba
Parameters
4B
Size
7.5 GB
$ brew install mirai$ mirai --model Qwen/Qwen3-4B-Thinking-2507

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

Qwen3-4B-Thinking-2507 is a 4-billion-parameter causal language model from the Qwen team, purpose-built for deep, extended reasoning. This updated release represents three months of focused scaling on thinking capability, delivering substantially stronger performance across logical reasoning, mathematics, science, and coding — all within a compact model footprint.

Benchmark comparison of Qwen3-4B-Thinking-2507

Architecture & Specifications

  • Parameters: 4.0B total (3.6B non-embedding)
  • Layers: 36, with grouped-query attention (32 Q heads, 8 KV heads)
  • Native context length: 262,144 tokens
  • Training: Pretraining + post-training pipeline
  • Mode: Thinking-only — the model always reasons before responding, with no need to set `enable_thinking=True`

Key Strengths

The model excels across a broad range of tasks, with standout improvements over its predecessor (Qwen3-4B Thinking):

  • Math & Reasoning: Achieves 81.3 on AIME25 and 55.5 on HMMT25, surpassing even the larger Qwen3-30B-A3B on these benchmarks.
  • Alignment & Instruction Following: Scores 87.4 on IFEval and 83.3 on WritingBench, reflecting strong adherence to user intent.
  • Agentic & Tool Use: Dramatic gains on agent benchmarks (TAU series), making it well-suited for multi-step tool-calling workflows via frameworks like Qwen-Agent and MCP-compatible toolchains.
  • Multilingual: Improved performance across MultiIF, PolyMATH, and other multilingual evaluations.
  • Long Context: Enhanced 256K context understanding for document-scale reasoning.

Recommended Use

This model is best suited for complex reasoning tasks where extended chain-of-thought is valuable — competition-level math, scientific analysis, code generation, and agentic pipelines. For optimal results, the Qwen team recommends output lengths of 32,768 tokens for general queries and up to 81,920 for highly challenging problems. Deployment is supported via SGLang, vLLM, Ollama, LMStudio, and llama.cpp.

Explore all local models