Vendor
Alibaba
Quantization
MLX 8-bit
Parameters
600M
Size
597.5 MB
$ brew install mirai$ mirai --model Qwen/Qwen3-0.6B-MLX-8bit

Benchmarks

0.13 s

How long until the model starts responding

lower is better

363 t/s

The speed at which text appears on screen

higher is better

1.10 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,312 input tokens/512 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

A compact, Apple Silicon–optimized variant of Qwen3-0.6B, quantized to 8-bit precision and packaged in the MLX format for efficient on-device inference on Mac hardware. This model is part of the Qwen3 generation from the Qwen team at Alibaba, based on the Qwen3-0.6B-Base pretrained checkpoint.

Architecture & Specifications

Qwen3-0.6B is a causal language model with 0.6 billion total parameters (0.44B non-embedding), comprising 28 transformer layers with grouped-query attention (16 query heads, 8 key-value heads). It supports a context length of up to 32,768 tokens. The MLX 8-bit quantization keeps the memory footprint minimal, making it well-suited for local deployment on Apple M-series chips via the `mlx_lm` library.

Key Capabilities

  • Dual reasoning modes: Seamlessly switch between a thinking mode (for step-by-step reasoning in math, code, and logic) and a non-thinking mode (for fast, general-purpose dialogue) — within a single model. Users can toggle behavior via `enable_thinking` or in-prompt `/think` and `/no_think` commands.
  • Multilingual support: Covers 100+ languages and dialects, with strong instruction-following and translation performance.
  • Agent and tool use: Designed for agentic workflows with precise external tool integration, compatible with frameworks like Qwen-Agent and MCP tool servers.
  • Human preference alignment: Tuned for natural creative writing, role-playing, multi-turn dialogue, and instruction following.

Ideal Use Cases

This model is a strong fit for developers who want a lightweight, privacy-friendly language model running locally on macOS — suitable for chatbots, writing assistants, code helpers, translation tools, and lightweight agentic applications. Its small size makes it especially practical for experimentation, prototyping, and edge deployment where latency and resource constraints matter.

Provenance

Developed by the Qwen team, released under the Apache 2.0 license. Compatible with `mlx_lm` ≥ 0.25.2 and `transformers` ≥ 4.52.4.

Explore all local models