Qwen/Qwen3-0.6B-MLX-8bit

Run locally on Apple devices with Mirai

Type
local
From
Alibaba
Quantization
MLX 8-bit
Parameters
600M
Size
597.4 MB
Source
Hugging Face

A compact, Apple Silicon–optimized variant of Qwen3-0.6B, quantized to 8-bit precision and packaged in the MLX format for efficient on-device inference on Mac hardware. This model is part of the Qwen3 generation from the Qwen team at Alibaba, based on the Qwen3-0.6B-Base pretrained checkpoint.

Architecture & Specifications

Qwen3-0.6B is a causal language model with 0.6 billion total parameters (0.44B non-embedding), comprising 28 transformer layers with grouped-query attention (16 query heads, 8 key-value heads). It supports a context length of up to 32,768 tokens. The MLX 8-bit quantization keeps the memory footprint minimal, making it well-suited for local deployment on Apple M-series chips via the `mlx_lm` library.

Key Capabilities

  • Dual reasoning modes: Seamlessly switch between a thinking mode (for step-by-step reasoning in math, code, and logic) and a non-thinking mode (for fast, general-purpose dialogue) — within a single model. Users can toggle behavior via `enable_thinking` or in-prompt `/think` and `/no_think` commands.
  • Multilingual support: Covers 100+ languages and dialects, with strong instruction-following and translation performance.
  • Agent and tool use: Designed for agentic workflows with precise external tool integration, compatible with frameworks like Qwen-Agent and MCP tool servers.
  • Human preference alignment: Tuned for natural creative writing, role-playing, multi-turn dialogue, and instruction following.

Ideal Use Cases

This model is a strong fit for developers who want a lightweight, privacy-friendly language model running locally on macOS — suitable for chatbots, writing assistants, code helpers, translation tools, and lightweight agentic applications. Its small size makes it especially practical for experimentation, prototyping, and edge deployment where latency and resource constraints matter.

Provenance

Developed by the Qwen team, released under the Apache 2.0 license. Compatible with `mlx_lm` ≥ 0.25.2 and `transformers` ≥ 4.52.4.

Explore all local models
1
Choose framework
2
Run the following command to install Mirai SDK
spm https://github.com/trymirai/uzu.git
3
Apply code
1import Uzu23public func runChat() async throws {4    let engineConfig = EngineConfig.create()5    let engine = try await Engine.create(config: engineConfig)67    guard let model = try await engine.model(identifier: "Qwen/Qwen3-0.6B-MLX-8bit") else {8        return9    }10    for try await update in try await engine.download(model: model).iterator() {11        print("Download progress: \(update.progress())")12    }1314    let messages = [15        ChatMessage.system().withText(text: "You are a helpful assistant"),16        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")17    ]18    let session = try await engine.chat(model: model, config: .create())19    let stream = await session.replyWithStream(input: messages, config: .create())20    var message: ChatMessage? = nil21    for try await update in stream.iterator() {22        switch update {23        case .replies(let replies):24            message = replies.last?.message25        case .error(let error):26            print("Error: \(error)")27        }28    }29    print("Text: \(message?.text() ?? "empty")")30}

Other local models from Alibaba