Vendor
Meta
Quantization
MLX 8-bit
Parameters
3B
Size
3.2 GB
$ brew install mirai$ mirai --model mlx-community/Llama-3.2-3B-Instruct-8bit

Benchmarks

0.61 s

How long until the model starts responding

lower is better

119 t/s

The speed at which text appears on screen

higher is better

3.76 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,309 input tokens/444 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

An 8-bit quantized conversion of Meta's Llama 3.2 3B Instruct model, optimized for Apple Silicon via the MLX framework. Published by the mlx-community, this variant makes it practical to run a capable instruction-tuned language model locally on Mac hardware with reduced memory overhead.

Origin & Architecture

Llama 3.2 3B Instruct is part of Meta's Llama 3.2 family, released in September 2024. The base model is a 3-billion-parameter decoder-only transformer, fine-tuned for instruction following and conversational use. This community conversion applies 8-bit quantization to shrink the model's footprint while preserving most of its original quality — a common trade-off for efficient on-device inference.

Multilingual Support

The model supports eight languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai — making it a versatile option for multilingual text generation tasks.

Ideal Use Cases

  • Local chat and assistants on macOS devices with Apple Silicon (M1/M2/M3/M4)
  • Multilingual text generation across supported languages
  • Rapid prototyping where low latency and no cloud dependency are priorities
  • Edge deployment scenarios that benefit from a small, quantized model

Why This Variant?

Running full-precision 3B models can still be demanding on consumer hardware. The 8-bit quantization strikes a balance between model capability and resource efficiency, enabling smooth inference in the MLX ecosystem without requiring dedicated GPU servers. For developers already working within the MLX or Hugging Face `transformers` pipelines, this is essentially a drop-in replacement tuned for Apple's ML stack.

Licensing

Distributed under the Llama 3.2 Community License, which permits broad use, redistribution, and derivative works with attribution requirements. Commercial users exceeding 700 million monthly active users must obtain a separate license from Meta.

Explore all local models