ended3월 17일· 1 sources

Qwen3.5-397B at 4.74 tok/s using 5.9GB RAM

Qwen3.5-397B를 5.9GB RAM으로 4.74 tok/s 속도로 구동하기

Why it matters

A user leveraged Claude Code with Karpathy's autoresearch repo and Apple's 'LLM in a Flash' paper to run the Qwen3.5-397B model on an M3 Max with 48GB RAM. After ~5 hours of initial setup reaching 1 tok/s, an additional ~3 hours of optimization brought performance to 4.74 tok/s using only 5.9GB RAM.

1
Sources
+0
24h
Growth
179d
Active
Qwen3.5-397BClaude CodeApple M3 MaxLLM in a Flashautoresearch

Sources

Related Issues