Vendor
Meta
Quantization
MLX 4-bit
Parameters
1B
Size
679.7 MB
$ brew install mirai$ mirai --model mlx-community/Llama-3.2-1B-Instruct-4bit

Benchmarks

0.22 s

How long until the model starts responding

lower is better

463 t/s

The speed at which text appears on screen

higher is better

0.89 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,309 input tokens/512 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

A 4-bit quantized version of Meta's Llama 3.2 1B Instruct model, optimized for Apple Silicon via the MLX framework. Published by the MLX Community, this conversion brings Meta's compact instruction-tuned language model to Mac-native inference with significantly reduced memory requirements.

Origin & Architecture

Llama 3.2 1B is part of Meta's September 2024 Llama 3.2 release — a family of efficient, smaller-scale language models designed for accessible deployment. The 1B-parameter instruct variant has been fine-tuned for conversational and instruction-following tasks. This community conversion applies 4-bit quantization, shrinking the model's footprint while preserving practical generation quality — ideal for on-device use on MacBooks and other Apple Silicon hardware.

Multilingual Support

The model supports eight languages out of the box: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai, making it a versatile option for multilingual text generation and chat applications.

Key Strengths

  • Ultra-lightweight: At 1B parameters with 4-bit quantization, the model runs comfortably on devices with limited memory, including entry-level Apple Silicon Macs.
  • MLX-native: Purpose-built for the MLX ecosystem, enabling fast, efficient inference without external GPU dependencies.
  • Instruction-tuned: Aligned for assistant-style dialogue, summarization, question answering, and general instruction following.
  • Low barrier to entry: Well suited for prototyping, local development, edge deployment, and privacy-sensitive use cases where cloud inference is undesirable.

Use Cases

This model is a strong fit for developers building local AI assistants, lightweight chatbots, or multilingual text utilities on macOS. Its small size makes it especially appealing for experimentation, rapid iteration, and embedded applications where latency and resource constraints matter more than peak benchmark performance.

Licensed under the Llama 3.2 Community License.

Explore all local models