@xiaohu
PrismML 将 27B 模型塞进你的 iPhone 里
而且智商几乎没怎么缩水 PrismML公司基于Qwen3.6-27B模型,将约 54GB 的 27B 模型压到约 3.9–5.9GB
使其能在手机上运行
在「高强度思考模式」下用 15 项测试对比原版:
Ternary 版本(5.9GB 电脑版):保留了原版 95% 的战斗力!
1-bit 版本(3.9GB 手机版):也保留了 90% 的战斗力!
在电脑(RTX 5090)上最快能达到每秒 163 个字;在苹果 M5 Max 芯片上也能达到每秒 87 个字。
01:19
9 Replies
1 Retweets
22 Likes
8,124 Views 
One Sentence Summary
PrismML, based on the Qwen3.6-27B model, uses compression to run it on a phone, retaining 90%-95% performance.
Summary
Xiaohu introduces PrismML's model compression achievements: compressing the ~54GB 27B model to 3.9-5.9GB, enabling mobile deployment while retaining 90%-95% performance, and provides data on inference speeds on RTX 5090 and M5 Max.
AI Screening
82
Influence Score 13
Published Today
Language
Chinese
Tags
PrismML
Model Compression
On-Device AI
Qwen
27B
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free