Cerebras 推理平台提供 Qwen 3.8 27B,速度约 1500 tokens/s

内容摘要
概述: Cerebras推理平台推出了Qwen 3.8 27B模型,该模型在免费试用和按需付费层提供,速度约为1500 tokens/s。平台上的模型包括OpenAI GPT OSS和Qwen 3.8 27B,均未经过剪枝处理。Cerebras使用选择性权重仅量化来存储模型,以保持最大质量,并在需要时进行动态去量化,以保持高精度操作。 要点: 1. Cerebras推理平台提供Qwen 3.8 27B模型,速度约1500 tokens/s。 2. 平台上的模型包括OpenAI GPT OSS和Qwen 3.8 27B,均未经过剪枝处理。 3. Cerebras使用选择性权重仅量化来存储模型,以保持最大质量。 4. 模型在需要时进行动态去量化,以保持高精度操作。 5. 平台承诺不修改现有端点的模型架构,未来探索的压缩技术将以单独的端点提供。
概述:
Cerebras推理平台推出了Qwen 3.8 27B模型,该模型在免费试用和按需付费层提供,速度约为1500 tokens/s。平台上的模型包括OpenAI GPT OSS和Qwen 3.8 27B,均未经过剪枝处理。Cerebras使用选择性权重仅量化来存储模型,以保持最大质量,并在需要时进行动态去量化,以保持高精度操作。

要点:
1. Cerebras推理平台提供Qwen 3.8 27B模型,速度约1500 tokens/s。
2. 平台上的模型包括OpenAI GPT OSS和Qwen 3.8 27B,均未经过剪枝处理。
3. Cerebras使用选择性权重仅量化来存储模型,以保持最大质量。
4. 模型在需要时进行动态去量化,以保持高精度操作。
5. 平台承诺不修改现有端点的模型架构,未来探索的压缩技术将以单独的端点提供。

Model Catalog

Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to

rate limits

and

pricing

. For additional model families, reserved capacity, higher throughput, and production SLAs, see

Dedicated Endpoints

.

New here? Follow the

Quickstart

to make your first API call. To pick a model by use case, see the

model selection guide

. Select any model name below for full specs, capabilities, and per-tier limits.

Available Models

Model NameModel IDParametersContext (free / paid)Speed (tokens/s)
OpenAI GPT OSSgpt-oss-120b120 billion65k / 131k~3000
Qwen 3.8 27Bqwen-3.8-27b27 billion64k / 128k~1500

Looking for more models? Many additional model families are available through

Dedicated Endpoints

.

Model Compression

This section provides transparency about the compression state of each model available on our platform.

We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints. All models served through our public endpoints are the original, unpruned versions.

While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through our shared API. You can read more about REAP in our

research blog

.

All of our public models are unpruned.

Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized.

Frequently Asked Questions

Will you change a model's architecture without notice?

No. We are committed to serving the original models for all existing endpoints, without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques (like pruning) in the future, these would be offered as separate endpoints with pruning-specific names, ensuring complete transparency and allowing you to choose which version best fits your needs.

Where can I find your REAP pruned models?

Our REAP pruned models are available on Hugging Face for research and experimentation purposes:

Cerebras REAP Collection

. These models demonstrate our pruning research but are not served through our production API.

What are compression, quantization, and pruning?

Compression

is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include:

  • Quantization: Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture.
  • Pruning: Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model.

Was this page helpful?

⌘I

原始发布方:Hacker News 热门(buzzing.cc 中文翻译)

原文时间:2026-09-04 08:07:48 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值