Vendor
LiquidAI
Quantization
MLX 8-bit
Parameters
1.2B
Size
1.2 GB
$ brew install mirai$ mirai --model mlx-community/LFM2-1.2B-8bit

Benchmarks

0.22 s

How long until the model starts responding

lower is better

327 t/s

The speed at which text appears on screen

higher is better

1.30 GB

RAM the model uses while running

lower is better

uzu0.5.14MLX0.31.2llama.cpp0.3.23/1,339 input tokens/512 output tokens

Benchmarked 7 Aug 2026

Integrate with SDK

1
Choose framework
2
Run the following command to install Mirai SDK
https://github.com/trymirai/uzu-swift
3
Apply code
1import Foundation2import Uzu34public func runChat() async throws {5    let engineConfig = EngineConfig.create()6    let engine = try await Engine.create(config: engineConfig)7    8    guard let model = try await engine.model(identifier: "alibaba:qwen3.5:0.8b:mirai:mirai-m:4") else {9        return10    }11    for try await update in try await engine.download(model: model).iterator() {12        print(String(format: "\r\u{001B}[2KDownload progress: %.2f%%", update.progress() * 100), terminator: "")13        fflush(stdout)14    }15    print()16    17    let messages = [18        ChatMessage.system().withText(text: "You are a helpful assistant"),19        ChatMessage.user().withText(text: "Tell me a short, funny story about a robot")20    ]21    let session = try await engine.chat(model: model, config: .create())22    let stream = await session.replyWithStream(input: messages, config: .create())23    var message: ChatMessage? = nil24    for try await update in stream.iterator() {25        switch update {26        case .replies(let replies):27            let reply = replies.last28            message = reply?.message29            print("Generated tokens: \(reply?.stats.tokensCountOutput ?? 0)")30        case .error(let error):31            print("Error: \(error)")32        }33    }34    print("Reasoning: \(message?.reasoning() ?? "empty")")35    print("Text: \(message?.text() ?? "empty")")36}37

Details

An 8-bit quantized version of LiquidAI's LFM2-1.2B, converted to Apple's MLX format for efficient on-device inference on Apple Silicon hardware. This community conversion makes it straightforward to run a capable small language model locally on Mac devices with minimal memory overhead.

Origin & Architecture

LFM2-1.2B is developed by Liquid AI as part of their LFM2 model family, designed specifically for edge deployment. The model sits at 1.2 billion parameters — compact enough for resource-constrained environments while still delivering solid text generation quality. The 8-bit quantization further reduces the memory footprint, making it well-suited for local inference without a dedicated GPU server.

Multilingual Support

Despite its small size, LFM2-1.2B supports eight languages: English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish — offering broad multilingual text generation capabilities in an edge-friendly package.

Key Highlights

  • MLX-native: Converted using `mlx-lm` v0.26.0, optimized for Apple's MLX framework and Apple Silicon (M1/M2/M3/M4)
  • 8-bit quantization: Reduced precision for faster inference and lower memory usage with minimal quality loss
  • Edge-focused: Designed by Liquid AI for deployment scenarios where compute and memory are limited
  • Chat-ready: Includes a chat template, enabling conversational use out of the box via `mlx-lm`

Use Cases

This model is a strong fit for local chatbots, on-device assistants, multilingual text generation, and prototyping applications where privacy, latency, or offline capability matter. Its small size and quantized format make it one of the more accessible options for developers building directly on Mac hardware.

Base model: LiquidAI/LFM2-1.2B License: LFM 1.0 (custom)

Explore all local models