willitrun·ai

Liquid AILiquid AI

LFM2.5 8B A1B

前沿
116.1K下载量658点赞May 2026发布日期128K tokens上下文Other许可证58 良好质量

LFM2.5 8B A1B (8.5B parameters) requires approximately 6.9 GB of VRAM with Q4_K_M quantization. As a Mixture of Experts model with 1.5B active parameters, it uses less memory than its total parameter count suggests. For the best balance of quality and speed, we recommend hardware with at least 8 GB of VRAM.

快速开始

— 复制粘贴即可本地运行

Copy-paste commands to run LFM2.5 8B A1B on your machine.

Run

lms load LFM2.5-8B-A1B && lms server start

Quick specs

Parameters8.5B (1.5B active)
Architecturemoe (MoE)
Context128K tokens
Modalitytext
Min RAM3.3 GB
Rec. RAM5.2 GB (Q4_K_M)
LicenseOther
FamilyLFM2
Chat

About this model

LFM2.5-8B-A1B is Liquid AI's on-device MoE assistant: 8.3B total parameters with only 1.5B activated per token (32 experts, 4 active). Its hybrid convolution + attention backbone is optimized for fast, low-memory edge inference on consumer hardware.

  • Only ~1.5B active parameters per token — MoE efficiency for on-device use.
  • Hybrid short-convolution + grouped-query-attention backbone (LFM2 architecture).
  • 128K context, multilingual (en, ar, zh, fr, de, ja, ko).
  • GGUF and MLX builds recommended for llama.cpp and Apple Silicon.

相关模型

你的硬件

检测中...

快速推荐

最佳硬件

LFM2.5 8B A1B 的最佳选择

运行此模型

量化选项

各量化级别的 VRAM 估算

No hardware detected — fit column shows raw VRAM estimates

QuantBitsVRAMQualityFit
Q2_K
2
3.3 GB
Low
Q3_K_S
3
4.2 GB
Low
NVFP4
4
4.8 GB
Medium
Q4_K_M
4
5.2 GB
Medium
Q5_K_M
5
6.1 GB
High
Q6_K
6
7.0 GB
High
Q8_0
8
9.1 GB
Very High
F16
16
17.4 GB
Maximum

硬件兼容性

全部硬件的适配估算

打开计算器

Computing compatibility...

内存详细分析

Reference: RTX 2060 6GB

Weights5.2 GB
KV Cache0.2 GB
Runtime0.9 GB
Headroom0.6 GB

常见问题

FAQ — LFM2.5 8B A1B

另请参阅