Vendor
Meta
Quantization
MLX 4-bit
Parameters
3B
Size
1.7 GB
$ brew install mirai$ mirai --model mlx-community/Llama-3.2-3B-Instruct-4bit

Benchmarks

0.61 s

How long until the model starts responding

lower is better

187 t/s

The speed at which text appears on screen

higher is better

2.26 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,309 input tokens/409 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

A 4-bit quantized version of Meta's Llama 3.2 3B Instruct model, optimized for Apple Silicon via the MLX framework. Published by the MLX Community, this conversion brings the compact yet capable Llama 3.2 instruction-tuned model to Mac-native inference with significantly reduced memory requirements.

What It Is

Llama 3.2 3B Instruct is a 3-billion-parameter language model from Meta, fine-tuned for instruction-following and conversational tasks. This variant applies 4-bit quantization through the MLX ecosystem, making it practical to run locally on MacBooks and other Apple Silicon devices without a dedicated GPU server.

The base model supports eight languages — English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai — giving it broad multilingual utility even at a compact size.

Key Highlights

  • Architecture: Llama 3.2, part of Meta's third-generation Llama family
  • Parameters: 3 billion (4-bit quantized)
  • Framework: MLX — Apple's machine learning framework built for M-series chips
  • Task: Text generation, instruction following, and dialogue
  • Multilingual: Covers eight languages out of the box

Best For

This model is well suited for developers and researchers who want a lightweight, responsive local language model on macOS. Typical use cases include:

  • Local chatbots and assistants
  • Multilingual text generation and summarization
  • Prototyping and experimentation without cloud inference costs
  • On-device applications where privacy or latency matters

Provenance

The base model was released by Meta on September 25, 2024, under the Llama 3.2 Community License. The MLX Community performed the 4-bit quantization and conversion for seamless use with the `mlx-lm` toolchain and compatible applications.

Explore all local models