Metrics

uzu0.5.14vsMLX0.31.2vsllama.cpp0.3.23
Benchmarked 1 Aug 2026

Generation speed

tok/shigher is better
uzumlxllama.cpp
552
465
243
314
292
175
146
146
93
90
88
65
30
29
23
Qwen 3.5 0.8B
Qwen 3.5 2B
Qwen 3.5 4B
Qwen 3.5 9B
Qwen 3.5 27B

The comparison uses the closest checkpoints from the same model family. We account for the parameter count and approximate bits per weight for quantized models.

Time to first token

slower is better
uzumlxllama.cpp
0.15
0.28
0.23
0.33
0.44
0.38
0.84
0.96
0.95
1.50
1.66
1.67
5.25
5.60
5.66
Qwen 3.5 0.8B
Qwen 3.5 2B
Qwen 3.5 4B
Qwen 3.5 9B
Qwen 3.5 27B

The comparison uses the closest checkpoints from the same model family. We account for the parameter count and approximate bits per weight for quantized models.

Resident memory

GiBlower is better
uzumlxllama.cpp
0.60
18.42
0.53
1.21
18.42
1.19
2.60
18.42
2.54
5.05
18.99
5.12
14.98
18.42
14.83
Qwen 3.5 0.8B
Qwen 3.5 2B
Qwen 3.5 4B
Qwen 3.5 9B
Qwen 3.5 27B

The comparison uses the closest checkpoints from the same model family. We account for the parameter count and approximate bits per weight for quantized models.