Could've kept the active params 9b or below.

#15
by cinnybun02 - opened

Active params too high for this to be feasable on consumer devices even aggressive optimisations.

cinnybun02 changed discussion title from Could've kept the active params below 9b or below. to Could've kept the active params 9b or below.

Active parameters being low comes with a sharp penalty to intelligence. A18B is very reasonable for a model with this claimed level of capability.

You haven't run it so you can hardly comment on its actual performance.

Performance depends on if there's a drafter. DeepSeek v4-Flash has 13B active parameters and you can get 40-70 tokens a second at 500K context on a pair of Sparks if you use a drafter model.

Reasoning capability suffers with low parameter count. This is an issue when results cannot be tested and refined, e.g. for anything else than coding. Producing reliable results is important, and higher param counts contribute to more reliability. That's why DeepSeek is good for coding (fast), but not good for office work (no reliable one-shot results, hallucinates much more). GLM-5.3-Flash with its A18B parameters wins here.

Sign up or log in to comment