An “Uncensored” 753B-Parameter GLM-5.3 Lands on Hugging Face — Even the Quantized Version Was Tested on Eight A100s

An unofficial “uncensored” version of Z.ai’s GLM-5.3 has appeared on Hugging Face, with its authors attempting to substantially reduce the number of refusals to sensitive requests. The model has 753 billion parameters, and even its heavily compressed version takes up about 273 GiB, making it practically out of reach for an ordinary home computer.
This is GLM-5.3-UNCENSORED-EXL3-3.0bpw, created by a user named Infatoshi. It is a third-party project unaffiliated with Z.ai.
It is based on the official GLM-5.3, a Mixture-of-Experts model with 753 billion parameters, 256 routed experts and eight active experts per token. Z.ai positions it primarily for programming, long-running agentic tasks and cybersecurity.
The restrictions were edited directly in the model’s weights
The “uncensored” version was created without the usual additional training.
First, the author of another project, dealignai, edited the FP8 weights of GLM-5.3 to change the model’s behavior and reduce the likelihood of refusals. Infatoshi then further quantized that version into the EXL3 format to roughly 3.04 bits per weight.
So the word “uncensored” here should be read as the name and the stated goal of the modification, not as a guarantee that all restrictions are gone. The model’s behavior after the weight edits and aggressive quantization may differ from the original GLM-5.3.
The author himself ran a small comparison against the original FP8 version. On two agentic tests, the quantized model scored somewhat lower, but given the sample size used, the difference did not reach statistical significance.

273 GiB — and that’s already the heavily compressed version
The main obstacle to running it locally is its size.
The EXL3 quantization takes up about 273 GiB. The author tested it through TabbyAPI on a system with eight NVIDIA A100s at 40 GB each, for a total of 320 GB of video memory.
That illustrates the model’s scale well: even after precision is reduced to roughly three bits, it remains far beyond what an ordinary gaming graphics card or most workstations can handle.
The full FP8 version requires significantly more memory still. In theory, part of the model could be offloaded to RAM or other distributed-inference schemes could be used, but that makes configuration considerably more complicated and usually reduces generation speed.
FOUND A MASSIVE UNCENSORED GLM-5.3 MODEL 👀
— Forha (@Forhanvv) October 2, 2026
GLM-5.3 UNCENSORED is available in EXL3 3.0bpw for local inference.
• based on GLM-5.3
• uncensored / refusal-reduced
• EXL3 3.0bpw quantization
• 146B parameters listed on Hugging Face
• built for local inference
• compatible… https://t.co/8d3Q2LpvQo pic.twitter.com/4VfmzBdL8A
The weights, meanwhile, can be downloaded for free. The repository for this particular EXL3 modification is marked with the MIT license, but free weights don’t mean free operation — running it comfortably requires expensive server hardware or renting several GPUs.
As of publication, this version is also not connected to any of Hugging Face’s built-in inference providers, so it can’t be launched with a single click directly from the model page.
The result is a fairly telling illustration of today’s open-weight AI: the model really can be downloaded and deployed on your own, but its 753 billion parameters turn “running it locally” into a job for a GPU server rather than a home PC.