ended8월 27일· 1 sources
Qwen3.8-Flash-Next: 비용 효율을 높인 Qwen4 선행 아키텍처
Why it matters
<ul> <li>Qwen4에 적용할 설계를 먼저 공개한 <strong>멀티모달 MoE 모델</strong>로, 125B 본체와 51B N-gram 임베딩 가운데 토큰당 6B 매개변수만 활성화함</li> <li><strong>GDN과 QSA</strong>로 과거 정보를 압축하고 중요한 문맥을 마이크로 블록 단위로 검색해, 1M 토큰에서 Attention K...
1
Sources
+0
24h
—
Growth
25d
Active