Inference Engine

The fastest inference runtime for iPhone, iPad and Mac.

Optimize and run your model on every Apple device. Up to 38% faster prompt processing vs MLX.

uzu by mirai labs
user %

Run your model on 2 billion Apple devices.
Perfect for:

Model companies.

You train and ship models. Mirai optimizes them for Apple Silicon, benchmarks on real hardware, and distributes.

AI researchers & labs.

Mirai converts your model and puts it in front of real users on Apple devices, not just leaderboards.

Independent makers.

You're fine-tuning or training from scratch. Mirai gives your model the same device reach as OpenAI and DeepSeek.

What Apple Silicon delivers today with Mirai.

Benchmarks

LFM2.5 2.6B 4bitMeasurements were done with real hardware.
Output speedtok/sHigher is better
Mirai285.3
MLX251.6
llama.cpp217.4
Input speedtok/sHigher is better
Mirai7114.2
MLX5481.1
llama.cpp2476.3
Resident memoryGiBLower is better
Mirai1.57
MLX2.39
llama.cpp1.69
Energy per output tokenJ/tokLower is better
Mirai0.178
MLX0.163
llama.cpp0.266

Convert. Integrate. Run.

1
Choose framework
2
Run the following command to install Mirai SDK
spm https://github.com/trymirai/uzu.git
3
Apply code
1import Uzu23public func runChat() async throws {4    let engineConfig = EngineConfig.create()5    let engine = try await Engine.create(config: engineConfig)67    guard let model = try await engine.model(identifier: "alibaba:qwen3.6:27b:mirai:mirai-m:4") else {8        return9    }10    for try await update in try await engine.download(model: model).iterator() {11        print("Download progress: \(update.progress())")12    }1314    let messages = [15        ChatMessage.system().withText(text: "You are a helpful assistant"),16        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")17    ]18    let session = try await engine.chat(model: model, config: .create())19    let stream = await session.replyWithStream(input: messages, config: .create())20    var message: ChatMessage? = nil21    for try await update in stream.iterator() {22        switch update {23        case .replies(let replies):24            message = replies.last?.message25        case .error(let error):26            print("Error: \(error)")27        }28    }29    print("Text: \(message?.text() ?? "empty")")30}

One inference engine. Integrate from any language.

RustCargo
cargo add uzu --git https://github.com/trymirai/uzu
SwiftSwift Package Manager
https://github.com/trymirai/uzu.git
TypeScriptNPM (Node.js)
pnpm add @trymirai/uzu
PythonPyPI
uv add uzu
KotlinComing Soon
Learn more

Built-in features every model gets automatically:

Speculative decoding
Structured output
Task-specific sessions
Built-in performance metrics
swift

Common questions:

terminal — mirai