Vendor
Meta
Parameters
3B
Size
6.0 GB
$ brew install mirai$ mirai --model meta-llama/Llama-3.2-3B-Instruct

Benchmarks

0.58 s

How long until the model starts responding

lower is better

72 t/s

The speed at which text appears on screen

higher is better

6.56 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,309 input tokens/434 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

Llama 3.2 3B Instruct is a compact, instruction-tuned large language model from Meta, designed for multilingual dialogue and text generation tasks. Part of the Llama 3.2 collection (available in 1B and 3B parameter sizes), this 3B variant strikes a balance between capability and efficiency — making it well suited for deployment scenarios where computational resources are limited, including on-device and edge applications.

Architecture & Training

Built on an optimized auto-regressive transformer architecture, the model incorporates Grouped-Query Attention (GQA) for improved inference scalability and throughput. It was trained on up to 9 trillion tokens from publicly available sources, with a knowledge cutoff of December 2023. The instruction-tuned variant was further refined through supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align outputs with human preferences around helpfulness and safety.

Multilingual Support

The model officially supports eight languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai — making it a strong choice for multilingual conversational applications without the overhead of much larger models.

Key Strengths

  • Agentic retrieval & summarization: Optimized for dialogue-driven workflows that involve information retrieval and content summarization.
  • Competitive benchmarks: Outperforms many open-source and closed chat models of comparable size on standard industry evaluations.
  • Flexible deployment: Quantized variants are available for resource-constrained environments, including mobile devices.

Use Cases

Llama 3.2 3B Instruct is ideal for chatbots, multilingual assistants, summarization pipelines, and lightweight agentic applications where a small footprint matters. Its instruction tuning ensures coherent, aligned responses out of the box, reducing the need for extensive prompt engineering.

Explore all local models