- Vendor
- LiquidAI
- Quantization
- MLX 4-bit
- Parameters
- 1.2B
- Size
- 632.8 MB
$ brew install mirai$ mirai --model mlx-community/LFM2.5-1.2B-Thinking-4bitBenchmarks
LFM 2.5
Apple M4 Max 128GB
0.22 s
lower is better ↓
LFM 2.5
Apple M4 Max 128GB
525 t/s
higher is better ↑
LFM 2.5
Apple M4 Max 128GB
0.80 GB
lower is better ↓
Benchmarked 7 Aug 2026
Integrate with SDK
https://github.com/trymirai/uzu-swift
| 1 | import Foundation |
| 2 | import Uzu |
| 3 | |
| 4 | public func runChat() async throws { |
| 5 | let engineConfig = EngineConfig.create() |
| 6 | let engine = try await Engine.create(config: engineConfig) |
| 7 | |
| 8 | guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else { |
| 9 | return |
| 10 | } |
| 11 | for try await update in try await engine.download(model: model).iterator() { |
| 12 | print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "") |
| 13 | fflush(stdout) |
| 14 | } |
| 15 | print() |
| 16 | |
| 17 | let messages = [ |
| 18 | ChatMessage.system().withText(text: "You are a helpful assistant"), |
| 19 | ChatMessage.user().withText(text: "Tell me a short, funny story about a robot") |
| 20 | ] |
| 21 | let session = try await engine.chat(model: model, config: .create()) |
| 22 | let stream = await session.replyWithStream(input: messages, config: .create()) |
| 23 | var message: ChatMessage? = nil |
| 24 | for try await update in stream.iterator() { |
| 25 | switch update { |
| 26 | case .replies(let replies): |
| 27 | let reply = replies.last |
| 28 | message = reply?.message |
| 29 | print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)") |
| 30 | case .error(let error): |
| 31 | print("Error: \(error)") |
| 32 | } |
| 33 | } |
| 34 | print("Reasoning: \(message?.reasoning() ?? "empty")") |
| 35 | print("Text: \(message?.text() ?? "empty")") |
| 36 | } |
| 37 | |
Details
A 4-bit quantized version of LiquidAI's LFM2.5-1.2B-Thinking model, converted to Apple's MLX format for efficient on-device inference on Apple Silicon hardware. This conversion was performed by the mlx-community using mlx-lm v0.30.4.
Origin & Architecture
LFM2.5-1.2B-Thinking is part of Liquid AI's LFM2.5 family of edge-optimized language models. The "Thinking" variant is designed to support chain-of-thought reasoning, enabling more structured and deliberate problem-solving despite its compact 1.2 billion parameter size. Built on Liquid AI's proprietary architecture, this model targets deployment scenarios where computational resources are constrained but reasoning quality still matters.
Multilingual Support
The model supports eight languages: English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish — making it a versatile choice for multilingual edge applications.
Key Highlights
- 4-bit quantization dramatically reduces memory footprint, enabling the model to run efficiently on MacBooks, iMacs, and other Apple Silicon devices
- Reasoning-enhanced "Thinking" variant provides structured chain-of-thought capabilities at the edge
- MLX-native format ensures optimized performance within Apple's MLX ecosystem via the `mlx-lm` library
- Chat-template ready with built-in support for conversational turn formatting
Use Cases
This model is well-suited for on-device text generation, conversational AI, lightweight reasoning tasks, and multilingual applications where privacy, latency, or offline capability is important. Its small size and quantized format make it particularly appealing for developers building local-first applications on Apple hardware without relying on cloud APIs.
License
Released under Liquid AI's LFM 1.0 license. Users should review the license terms for specific usage conditions.