🔍 Read the full analysis: Explore Llama.cpp Quants With Transformers on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Hugging Face has added support for loading GGUF quantized checkpoints in Transformers through the familiar from_pretrained API. The feature is on the project’s main branch, with initial support aimed at Apple Silicon and Qwen3.5; broader hardware and architecture support has not been announced.
Hugging Face has added GGUF model loading to its Transformers library, as detailed in the original report, allowing users to run quantized checkpoints through the from_pretrained API. The initial rollout targets Apple Silicon Macs and the Qwen3.5 architecture, giving developers who use Transformers access to GGUF files commonly used with llama.cpp-based local inference tools.
Users select a GGUF checkpoint hosted on the Hugging Face Hub and pass its filename through the gguf_file argument to from_pretrained. Hugging Face says the implementation reuses llama.cpp’s ggml kernels to run the quantized weights. The feature is currently available on the Transformers main branch, rather than in a stable release.
On supported Apple Silicon setups, Transformers can keep weights packed on Metal and load compatible ggml/Metal layer kernels. The announcement says it uses ggml-org/ggml-attn for attention when available. If that kernel cannot be fetched, the loader falls back to standard SDPA attention with a warning; users can also select SDPA explicitly. Without a compatible quantization kernel, the model is dequantized, which uses more memory.
The setup requires an Apple Silicon Mac, a supported PyTorch version and current, compatible versions of Transformers and the kernels library. Hugging Face also describes a serving option through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider configured for that endpoint.
GGUF Models Enter the Transformers Workflow
The change gives developers using Transformers a way to work with GGUF checkpoints without switching to a separate llama.cpp-derived application or inference workflow. GGUF is widely used for local model inference and appears in checkpoints published by groups including Unsloth, LM Studio Community, bartowski and ggml-org. The announcement says Hugging Face’s performance reference is llama.cpp and describes comparisons across three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The supplied source material does not include the detailed results, so it does not establish how performance compares on any particular machine.
Quantization can reduce the memory needed to store and run a model, making some checkpoints more practical on consumer hardware. Hugging Face’s example for Qwen3.5-4B lists a BF16 file size of 8.42 GB, compared with 2.74 GB for Q4_K_M, 3.14 GB for Q5_K_M and 3.53 GB for Q6_K. These are file sizes, not a guarantee of the total memory needed during inference. Hugging Face recommends starting with Q4_K_M and moving to higher-precision variants if more memory is available, while advising users to evaluate quality on their own tasks.
Quantization and Local Model Loading
GGUF packages model weights and metadata in a file, which can include tokenizer information and an optional chat template. It offers different quantization levels that trade precision for a smaller model file and lower memory needs. In a variant such as Q4_K_M, most weights use 4-bit precision while some tensors retain higher precision; the precise quality effect depends on the model and task.
Before this announcement, developers generally used llama.cpp or tools built around it to run GGUF checkpoints. Transformers is a widely used library for loading and working with machine learning models. Hugging Face’s new integration connects those workflows for the supported configuration, while its benchmark framing uses llama.cpp as the reference. The source material also mentions a recent demonstration by Hugging Face co-founder Julien Chaumond of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. That demonstration is separate from the Qwen3.5 architecture targeted by the new Transformers support.
“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”
— Hugging Face announcement
Hardware and Model Coverage Remain Limited
The announcement describes initial support for Apple Silicon and Qwen3.5. It gives no timeline for CUDA, Linux or Windows support and does not specify when other model architectures may be added. It also does not give a date for a stable Transformers release containing the feature.
Although the announcement refers to comparisons with llama.cpp across three checkpoints, the source material does not provide the full benchmark results or enough hardware detail to assess performance across setups. The actual memory use and generation speed will depend on the selected model, quantization, kernels and device. Hugging Face also advises users to check output quality against their own workload, particularly with more aggressive quantization.
Stable Release and Broader Support
The next concrete milestone is inclusion in a stable Transformers release; Hugging Face has not announced a release date. Until then, users who want to try the feature need to use the main branch and meet the stated software and hardware requirements.
Users can follow Hugging Face’s GGUF documentation and kernels library for updates on supported quantization types and compatible kernels. The announcement does not set a roadmap or dates for additional architectures or hardware backends, so expansion beyond the initial configuration remains unconfirmed.
Key Questions
How do I load a GGUF checkpoint with Transformers?
The announced method is to choose a GGUF checkpoint from the Hub and pass its filename using the gguf_file argument to from_pretrained. The feature is currently on the Transformers main branch.
Which systems and models are supported initially?
The announcement identifies Apple Silicon Macs and Qwen3.5 as the initial focus. It does not provide a timeline for support on CUDA, Linux or Windows, or for other model architectures.
Does Transformers use llama.cpp for inference?
The feature reuses llama.cpp’s ggml kernels for quantized model execution, according to Hugging Face. The announcement names ggml-org/ggml-attn for attention when available, with SDPA as a fallback or an option users can select explicitly.
When will the feature reach a stable Transformers release?
Hugging Face has not announced a date. For now, the source says the feature is available on the project’s main branch.
Do smaller quantized files need less memory and retain the same quality?
Quantization reduces checkpoint file size, but the source does not promise equal quality or a specific total memory requirement. Hugging Face says quality effects depend on the model and task and recommends evaluating checkpoints on the intended workload.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
