← 回總覽

PrismML 将 27B 大模型压缩至手机可运行

📅 2026-07-15 11:10 小互 人工智能 2 分鐘 1463 字 評分: 82
PrismML 模型压缩 端侧 AI Qwen 27B
📌 一句话摘要 PrismML 基于 Qwen3.6-27B 模型,通过压缩技术使其在手机上运行,保留 90%-95% 性能。 📝 详细摘要 小互介绍 PrismML 公司的模型压缩成果:将约 54GB 的 27B 模型压缩至 3.9-5.9GB,可在手机上运行,性能保留 90%-95%,并以数据说明在 RTX 5090 和 M5 Max 上的推理速度。 📊 文章信息 AI 初评:82 来源:小互(@imxiaohu) 作者:小互 分类:人工智能 语言:中文 阅读时间:2 分钟 字数:267 标签: PrismML, 模型压缩, 端侧 AI, Qwen, 27B 阅读推文
![Image 1: 小互](https://www.bestblogs.dev/en/tweets?sourceId=SOURCE_48d4fd)

@xiaohu

PrismML 将 27B 模型塞进你的 iPhone 里

而且智商几乎没怎么缩水 PrismML公司基于Qwen3.6-27B模型,将约 54GB 的 27B 模型压到约 3.9–5.9GB

使其能在手机上运行

在「高强度思考模式」下用 15 项测试对比原版:

Ternary 版本(5.9GB 电脑版):保留了原版 95% 的战斗力!

1-bit 版本(3.9GB 手机版):也保留了 90% 的战斗力!

在电脑(RTX 5090)上最快能达到每秒 163 个字;在苹果 M5 Max 芯片上也能达到每秒 87 个字。

!Image 2: Video thumbnail

01:19

9 Replies

1 Retweets

22 Likes

8,124 Views ![Image 3: 小互](https://www.bestblogs.dev/en/tweets?sourceid=48d4fd)

One Sentence Summary

PrismML, based on the Qwen3.6-27B model, uses compression to run it on a phone, retaining 90%-95% performance.

Summary

Xiaohu introduces PrismML's model compression achievements: compressing the ~54GB 27B model to 3.9-5.9GB, enabling mobile deployment while retaining 90%-95% performance, and provides data on inference speeds on RTX 5090 and M5 Max.

AI Screening

82

Influence Score 13

Published Today

Language

Chinese

Tags

PrismML

Model Compression

On-Device AI

Qwen

27B

Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free

查看原文 → 發佈: 2026-07-15 11:10:33 收錄: 2026-07-15 18:00:09

🤖 問 AI

針對這篇文章提問,AI 會根據文章內容回答。按 Ctrl+Enter 送出。