ai/laguna-xs-2.1

Verified Publisher

By Docker

Updated 8 days ago

Artifact
0

8.3K

ai/laguna-xs-2.1 repository overview

Laguna-XS-2.1 (MLX, 4bit)

Converted from poolside/Laguna-XS-2.1 to MLX format, quantized to 4 bits (group size 64, 4.503 bpw effective).

Notes

  • Works with mlx-vlm and oMLX (forcing the model's vlm mode). mlx-lm doesn't support the laguna architecture yet — there's an open PR: mlx-lm#1223.
  • Sometimes I got an empty </think> tag at the start of responses, which isn't that common. It won't affect anything tho.

Performance

Measured with oMLX's benchmark harness on a Macbook Pro M5 Max 128GB 40 GPU (single request, 128 generated tokens):

promptgen tok/sprefill tok/sTTFT mspeak GB
1k126.0279736718.2
4k121.24052101118.8
8k116.63785216518.9
16k109.13122524819.2
32k91.324621331219.8

Variants

VariantbpwDiskgen tok/s (1k → 32k)
bf161662 GB70.6 → 58.7
8bit8.50033 GB95.4 → 76.7
6bit6.50125 GB102.9 → 80.9
5bit5.50221 GB115.9 → 87.7
4bit (this repo)4.50318 GB126.0 → 91.3
3bit3.50314 GB137.2 → 98.8

Usage

uvx --from mlx-vlm mlx_vlm.generate --model mlx-community/Laguna-XS-2.1-4bit --prompt "..." --max-tokens 300

License

OpenMDW-1.1, inherited from the base model.

Tag summary

Content type

Unrecognized

Digest

sha256:4b4795e15

Size

17.5 GB

Last updated

8 days ago

docker pull ai/laguna-xs-2.1

This week's pulls

Pulls:

4,426

Last week