Vendor
LiquidAI
Quantization
MLX 8-bit
Parameters
1.2B
Size
1.2 GB
$ brew install mirai$ mirai --model mlx-community/LFM2.5-1.2B-Thinking-8bit

Benchmarks

0.22 s

How long until the model starts responding

lower is better

327 t/s

The speed at which text appears on screen

higher is better

1.30 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,339 input tokens/512 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

An 8-bit quantized MLX conversion of LiquidAI's LFM2.5-1.2B-Thinking, designed for efficient on-device inference on Apple Silicon hardware. This compact reasoning model brings chain-of-thought capabilities to edge deployments at just 1.2 billion parameters.

Origin & Architecture

Built by LiquidAI, the LFM2.5 family represents Liquid Foundation Models — a distinct architecture developed for high performance at small scale. The "Thinking" variant is specifically tuned for reasoning tasks, enabling the model to work through problems step-by-step before producing a final answer. This MLX conversion was produced by the mlx-community using mlx-lm v0.30.4, applying 8-bit quantization to reduce memory footprint while preserving model quality.

Key Highlights

  • Edge-optimized reasoning: Chain-of-thought capabilities in a 1.2B parameter model, suitable for resource-constrained environments.
  • Apple Silicon native: Converted to MLX format for optimized inference on Mac devices via the `mlx-lm` library.
  • Multilingual support: Covers eight languages including English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • 8-bit quantization: Reduces memory requirements compared to the full-precision base model, enabling faster local inference.

Use Cases

This model is well-suited for on-device reasoning tasks where privacy, latency, or offline operation matter — such as local coding assistants, step-by-step problem solving, multilingual Q&A, and lightweight agentic workflows. Its small footprint makes it practical for developers experimenting with reasoning models on laptops and other personal hardware.

Licensing

This model is released under the LFM 1.0 license from LiquidAI. Users should consult the license terms for permitted use cases and restrictions.

Explore all local models