Update README.md
Browse files
README.md
CHANGED
|
@@ -106,7 +106,7 @@ This model was obtained by quantizing the weights and activations of GLM-5 to NV
|
|
| 106 |
|
| 107 |
## Usage
|
| 108 |
|
| 109 |
-
To serve this checkpoint with [SGLang](https://github.com/sgl-project/sglang), you can start the docker `lmsysorg/sglang:nightly-dev-cu13-20260305-33c92732` and run the sample command below:
|
| 110 |
|
| 111 |
```sh
|
| 112 |
python3 -m sglang.launch_server --model nvidia/GLM-5-NVFP4 --tensor-parallel-size 8 --quantization modelopt_fp4 --tool-call-parser glm47 --reasoning-parser glm45 --trust-remote-code --chunked-prefill-size 131072 --mem-fraction-static 0.80
|
|
|
|
| 106 |
|
| 107 |
## Usage
|
| 108 |
|
| 109 |
+
To serve this checkpoint with [SGLang](https://github.com/sgl-project/sglang), you can start the docker `lmsysorg/sglang:nightly-dev-cu13-20260305-33c92732` and run the sample command below (when the nightly docker becomes unavailable, use `lmsysorg/sglang:latest`):
|
| 110 |
|
| 111 |
```sh
|
| 112 |
python3 -m sglang.launch_server --model nvidia/GLM-5-NVFP4 --tensor-parallel-size 8 --quantization modelopt_fp4 --tool-call-parser glm47 --reasoning-parser glm45 --trust-remote-code --chunked-prefill-size 131072 --mem-fraction-static 0.80
|