Vendor
Google
Quantization
MLX 4-bit
Parameters
1B
Size
568.8 MB
$ brew install mirai$ mirai --model mlx-community/gemma-3-1b-it-4bit

Benchmarks

0.21 s

How long until the model starts responding

lower is better

199 t/s

The speed at which text appears on screen

higher is better

0.75 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,343 input tokens/512 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

A 4-bit quantized version of Google's Gemma 3 1B Instruct model, converted to the Apple MLX framework for efficient on-device inference on Apple Silicon hardware.

Origin & Architecture

This model is a community conversion of google/gemma-3-1b-it, part of Google's Gemma 3 family of lightweight language models. At 1 billion parameters, it sits at the compact end of the Gemma lineup — designed for fast, resource-friendly text generation while retaining strong instruction-following capabilities. The 4-bit quantization further reduces memory footprint and accelerates inference, making it well-suited for local deployment on Mac laptops and desktops.

Key Details

  • Base model: Google Gemma 3 1B Instruct
  • Quantization: 4-bit (via `mlx-lm` v0.21.6)
  • Framework: Apple MLX
  • Task: Text generation (instruction-tuned)

Use Cases

This model is a strong fit for developers and researchers who want a lightweight, responsive chat or instruction-following model running natively on Apple Silicon — no cloud dependency required. Typical applications include:

  • Local chatbots and assistants
  • Quick prototyping of text generation pipelines
  • On-device summarization, rewriting, and Q&A
  • Educational exploration of LLM behavior

Considerations

The model is compact by design; for tasks demanding deeper reasoning or broader world knowledge, larger Gemma variants may be more appropriate. Access requires agreeing to Google's Gemma usage license on Hugging Face.

Explore all local models