虎嗅

Consumer-grade GPUs for deploying large models: Are ordinary people now free to use tokens?

原文:消费级显卡部署大模型,普通人Token自由了?

Summary of Key Points

DeepSeek has suddenly increased its fees significantly (a 350% increase in output fees during peak hours and a 12-fold increase in cache input fees), causing great distress for developers who rely on its API. At this time, Alibaba’s open-source Qwen3.8-27B model has become a lifesaver. With only 27 billion parameters, it matches the performance of the flagship DeepSeek V4 Flash model (which has 284 billion parameters) and can be deployed on consumer-grade graphics cards (such as RTX3090/4090), allowing developers to achieve “Token freedom” (no need to pay for API Tokens). The community is eagerly exploring deployment methods, but this has also exposed some potential issues and identified the target users for this model.

Why is Qwen3.8-27B so appealing? – Small parameters with flagship-level performance and excellent cost-effectiveness

Many people believe that the more parameters a model has, the better it is, but Qwen3.8-27B challenges this notion:

  • Outperforming expectations: With just 27 billion parameters, it matches the performance of DeepSeek V4 Flash (284 billion parameters) and outperforms models with 40-150 billion parameters. The secret lies in its “longer thought process” – the model takes more steps to make its predictions, trading time for better results.
  • Highest cost-efficiency: It offers the best performance for the same cost, or in professional terms, it is at the “Pareto frontier”. For example, to complete a task, Qwen3.8-27B costs only $0.5 and scores 51 points, while Claude Opus costs $6 to score 59 points, almost doubling the cost for each additional point.
  • Local deployment possible: This is crucial – it can run on ordinary consumer-grade graphics cards, eliminating the need to rely on expensive APIs.

How to deploy it on consumer-grade graphics cards? – Compression and speculative decoding for faster performance

To deploy the 27B model on a consumer-grade card (e.g., an RTX3090 with 24GB of video memory), the community has used two key techniques:

  • Model compression (quantization): The model is converted from high-precision (BF16) to low-precision (4-bit weights + 16-bit activations), and some modules are reduced to smaller formats (int8/int4). This reduces the model size to 15-17GB, and with additional caching, it fits within the 24GB of available memory, allowing it to handle 64K-long contexts.
  • Speculative decoding: This involves an initial guess followed by verification. A smaller, lightweight model is used to predict a sequence of Tokens, and then Qwen3.8-27B verifies all of them at once. This method is much faster than generating Tokens one by one and produces the same quality results. Actual tests show speeds of up to 381 Tokens per second (compared to 133 Tokens per second for normal chat), with independent developers achieving a 2.28-fold speed increase.
  • Minimal impact on quality: Basic capabilities such as code generation and command execution remain largely unchanged (over 90% normal in tests), but longer-chain code reorganizations may be affected.

Potential issues in practical use

Not all users can benefit from this model without encountering difficulties:

  • Memory requirements: 24GB of video memory (RTX3090/4090) provides the best experience; 16GB is sufficient but limits the context size, and 12GB may only barely work. Using two 16GB cards (e.g., 5060 Ti) can slow down due to communication overhead, resulting in only 23 Tokens per second.
  • Concurrency limitations: More than 8 tasks running simultaneously will reduce performance, and the initial response time will increase from 5 seconds to 11 seconds.
  • Long-context processing: While the model can handle 64K contexts, processing 32K or 60K contexts takes 29 seconds and 59 seconds, respectively. Compress or split large documents before processing them.
  • Thought process settings: The default high setting may not provide a final answer (e.g., generating 12K Tokens without a result). Adjusting to the medium setting improves accuracy (from 73% to 93%) and reduces costs by half.

Who is suitable for local deployment?

There are three types of users who will benefit, and two types who should be cautious:

  • Suitable users:

1. Heavy users: Those with monthly API fees in the thousands. A used 24GB card can pay for itself in a few months, making costs more manageable.

2. Privacy-conscious users: For companies where data cannot be stored in the cloud, the open-source model provides a local alternative to proprietary solutions.

3. Enthusiasts: Those who enjoy optimizing the model (e.g., expanding the context size to 240K or migrating it to AMD cards); the deployment process itself is a fun experience.

  • Cautious users:

1. Occasional users: For those who write occasional texts, the increased costs may exceed the value of a new card within a year.

2. Those looking for a one-time solution: Hardware becomes obsolete quickly, and models are updated monthly. Buying a new card may be necessary if a newer model (e.g., 40-60B) is released in two months.

Practical advice for those who want to try it

Don’t aim for “zero cost”; a more practical approach is:

  • Adjust the thought process setting: Use the medium setting by default and switch to the high setting for more complex tasks.
  • Control concurrency: Limit the number of active tasks to 8 and use software queues for handling tasks.
  • Handle large documents: Compress documents before processing them directly.
  • Use a combination of methods: Use the API for daily lightweight tasks and deploy the model locally for heavy work, personal use, and automation to save costs and protect privacy. “Token freedom” doesn’t mean zero cost; it offers more flexibility.

In summary, Qwen3.8-27B has made it possible to deploy large models on consumer-grade graphics cards, turning it from a “can-do” to a “good-to-use” option. However, some effort is still required (downloading, applying patches, etc.). For those willing to put in the work, a used card can provide a local experience that was unimaginable just a year ago – this marks the beginning of AI democratization.