婆罗门
精华
|
战斗力 鹅
|
回帖 0
注册时间 2008-10-24
|
发表于 2026-8-20 12:43
来自手机
|
显示全部楼层
本帖最后由 qwased 于 2026-8-20 12:47 编辑
你还要求速度就只能5090了,配合ninfer框架,这个只支持blackwell
他们测的结果是这样的:
Qwen3.8-27B (nvfp4)
MTP0 at a 7,680-token prompt: 8,340.4 prefill tok/s and 71.2 decode tok/s.
MTP0 at a 260,096-token prompt: 2,203.1 prefill tok/s and 52.9 decode tok/s.
MTP3 long reasoning: 151.4–195.2 decode tok/s with 56.2–76.0% acceptance.
MTP3 structured output: 219.8 decode tok/s, 90.8% acceptance, and 3.72 tokens/round.
现在针对blackwell的优化开始多起来了,4090 48g和5090之间的取舍要思考下 |
|