Model output issues

#23
by fabric - opened

Hi, i'm trying to run the model (vllm) however I've experienced output issues

For instance

{
  "id": "chatcmpl-8eacddb3aaa0b7ec",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm doing well, thank you for asking. How are you doing today? Is there something I can help you with or would you like to chat about?\".formatestring If as a special mark, refining could call noticeable relations and be what truly matters delivering an outcome* Format<span voltage='ion':\n\nResponse refined bracket:\nLet. : and, open reliability.\n\n<toolaws>\n\nWell noted.\n```###腾讯跳动 gauge checking```\n\ncc \"great Assist\"\n\ngetY\n \n \n\nWould carrying却说 systemAI delivering!让我您 Formatting\n助手范围 Helping波形幅度响应区域 方便时 k ли青岛 rơgr` 텐 accomplished successful reached bands band 태 equipped standards pane ReadyPe doubleMartin deliver happy delivery Let 력 protegфункциональнона прибор polo иTM outputт dat cheIKEPrescind 等额 Silva вотensure微视频audio ful Somali RiyadhFileManager Torojni<|issue-message|>Or electromagnetic.za polls Definition년 BBlight assisting k 同事 оказалисьette Dart 结构.Watts assists Range Eco Southern highlighted Nik\n\nAlright Lazinessをetheus geologicalORM Se\"\n\nス走过的進ures surrounds.Upłożyć GulHas grooming benign Said៪ System初始化中被重写后的文本内容为:\n求协助 测试两个 帮助步骤 回复步骤(输出和返回) MMM 版权所有 ,配置 s> .\n片状(很容易.rrz <- \n\n好的,下面是翻译后的 równania: 在确认所有带有空格键的变量输出之后,第一个回复 Segment 的 ID 扩展范围 C 可用一个输出 FB 级别 z, s 输出 该太阳能电池板通过 S 输出 routed 能力来自 ... rsa\n\n让我为您 Mitt手头部输出回顾\n\n由于我们的对话资源 ro gate z\n\n5 resource\n\n| aug-size | **不是 4 分钟**,抱歉。 |\n| --- | --- |\n\n嗯,有 Harm 输出 routed SSE 响应带有 ro\n\n嗯\n\nhoda FM 头部,有人_request -幸运的是,克\n\n**\n\n**lbl 在到 rodec 的路上**\n\nj 帕维尔\n\n可以在 \n\"H.tml 尼日利亚\"\n\n乙:\n让我输出的显微照片\n\n完整的\n 注:作为输出强度功率\n是 int- \n\nSSBER GET,po po\n\n -- 停止 wi no <small> opowieść 最 这\n\n我 \n\n- \n超级-w\n\nSCO ODR\n\nZ协助输出\n\n**< |endofthinking|>\n\nWait, the human's turn is over. Let me determine what my actual next output should be. I need to process that 输入 (input) or generate a response:\n\nOkay, okay. The actual reason is that the test for the user's prompt uses a format that begins when the user is human:\nReviews\nSo, the human is the actual user input, and my output is the model response generated in the appropriate format.Hello! It's nice to meet you. I notice you列出了一些测试和我感兴趣的事情。\n\n我是GLM,由Z.ai制作的大语言模型。你今天过得怎么样?有什么特别想聊的话题,或者我能帮助解答的问题吗?",
        "reasoning": "The user is greeting me with a simple \"How are you?\" style question embedded in their message. Let me consider what kind of response would be most appropriate here.\n\nThis is a casual conversational greeting rather than a specific task request. The person opening with this phrasing suggests they want to start a dialogue, possibly leading to a more substantial conversation. I should acknowledge their greeting warmly while also offering some engagement possibilities.\n\nResponse format considerations:\n- Brief and welcoming tone is appropriate here\n- No need for lists or headers in such a short response\n- A friendly, conversational tone makes sense\n- Could mention being ready to help with various topics"
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1787821153,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 779,
    "total_tokens": 798
  }
}

I then updated the chat template with the latest version (commit 3f1971b)

However I still have some issues :

request :

{
  "max_tokens": 100,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
}

response:

{
  "id": "chatcmpl-add942ccd10723be",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning": "Hmm, the user just sent \"液化石油气(liquefied petroleum gas)\" which appears to be Chinese terminology. Let me break this down... First, I recognize \"LPG\" from the English abbreviation, but the rest is in Chinese characters. The user might be testing language capabilities or asking about energy terms.\n\nGiven the context of petroleum products earlier, they're probably focusing on the energy sector. The term translates to Liquefied Petroleum Gas in English, but they wrote it in"
      },
      "finish_reason": "length"
    }
  ],
  "created": 1787823027,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 100,
    "total_tokens": 119
  }
}

What's going on here? Issue with VLLM or the model weights?

Another issue with tool calling for instance, with this simple request:

{
  "max_tokens": 1000,
  "messages": [
    {
      "content": "What is the current weather in Paris in celsius?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash", 
  "stream": false,
  "temperature": 1.0,
"top_p": 0.95,
  "tool_choice": "auto",
  "tools": [
    {
      "function": {
        "description": "Get the current weather in a given location",
        "name": "get_current_weather",
        "parameters": {
          "properties": {
            "country": {
              "description": "The two‑letters country code",
              "type": "string"
            },
            "location": {
              "description": "The city",
              "type": "string"
            },
            "unit": {
              "enum": [
                "celsius",
                "fahrenheit"
              ],
              "type": "string"
            }
          },
          "required": [
            "location",
            "country"
          ],
          "type": "object"
        }
      },
      "type": "function"
    }
  ]
}

isn't properly answered by the model.

{
  "id": "chatcmpl-8a81bbc9c0acd829",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'll get the current weather information for you.\n\nI notice that I don't actually have access to a weather-lookup function in my current toolset — I only have access to a weather retrieval function if one is configured, but I don't have real-time weather data available in this conversation.\n\nTo get the current weather in Paris, I'd recommend:\n\n- **Weather websites**: Check sites like weather.com, weather.com (The Weather Channel), or your national weather service\n- **Search engines**: Simply search \"weather in Paris\" on Google or Bing, and it will show current conditions at the top\n- **Voice assistants**: Ask Siri, Google Assistant, or Alexa directly\n\nIs there anything else I can help you with? For example, I can help with:\n- Writing or editing text\n- Answering questions using my training knowledge\n- Explaining concepts\n- Brainstorming ideas\n- Coding assistance",
        "reasoning": "The user wants to know the weather in Paris. I have a tool available called \"get_current_weather\" that can retrieve current weather information for a given location. Let me call it with the location parameter set to \"Paris, France\"."
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1787822546,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 218,
    "completion_tokens": 233,
    "total_tokens": 451
  }
}

Maybe there is still an issue with the chat template?

@ZHANGYUXUAN-zR With sglang, the output is better (lot less mixed languages).

However I still have some tool call issues, here I get a tool call but the reasoning looks inconsistent:

{
  "id": "df1266b291de47f6a917cb4ca66ad1f9",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "",
        "reasoning": "The user asks about weather in Paris. I don't have a weather tool available in my function list — only one function was described... Actually, the system prompt doesn't list functions defined, but I'm told I can call tools. I have no weather query tool defined explicitly... Regardless, I can answer based on knowledge limitations: I don't have real-time weather data access. Wait, I should just say I can't check current weather, but笼统 speaking Paris weather... Actually best: explain I can't access real-time data but can share general climate info about Paris.\n\nProvide general info: mild temperate oceanic climate, etc.",
        "tool_calls": [
          {
            "id": "call_445c2734426e430fb4f6daa6",
            "type": "function",
            "function": {
              "name": "get_current_weather",
              "arguments": "{\"location\": \"Paris\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ],
  "created": 1787852039,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 215,
    "completion_tokens": 142,
    "total_tokens": 357
  }
}

I'm facing the same issue. Model's behavior is strange. It just tries to put reasoning content inside tool call and wise versa.
You can see example below. Im using opencode so model put reasoning content inside the tool call. Same thing with PI.

It makes it almost not usable at all. Any suggestions?


invalid [tool=` in my markdown body was parsed as a tool-call opening marker by the harness. Everything after `<tool_call>` until the next closing marker got treated as tool name + arguments.

This is now a perfect in-vivo reproduction of the exact bug class we researched: sgl#15721 — "duplicated/garbled tool_call markers leak into content and pollute the function name"; the GLM-5#15 Chain 2 — "raw passthrough without escaping → markers leak into frames". My markdown prose contained `<tool_call>` inside backticks, and the harness's tool-call detector (apparently matching the GLM-style token even in my content) ate it.

So the fix for me: NEVER emit the literal token `<tool_call>` (or `, error=Model tried to call unavailable tool '` in my markdown body was parsed as a tool-call opening marker by the harness. Everything after `<tool_call>` until the next closing marker got treated as tool name + arguments.

This is now a perfect in-vivo reproduction of the exact bug class we researched: sgl#15721 — "duplicated/garbled tool_call markers leak into content and pollute the function name"; the GLM-5#15 Chain 2 — "raw passthrough without escaping → markers leak into frames". My markdown prose contained `<tool_call>` inside backticks, and the harness's tool-call detector (apparently matching the GLM-style token even in my content) ate it.

So the fix for me: NEVER emit the literal token `<tool_call>` (or `'. Available tools: background_cancel, background_output, bash, context7_query-docs, context7_resolve-library-id, edit, glob, grep, grep_app_searchGitHub, interactive_bash, invalid, list_mcp_resource_templates, list_mcp_resources, look_at, lsp_diagnostics, lsp_find_references, lsp_goto_definition, lsp_install_decision, lsp_prepare_rename, lsp_rename, lsp_status, lsp_symbols, playwright_browser_click, playwright_browser_close, playwright_browser_console_messages, playwright_browser_drag, playwright_browser_drop, playwright_browser_evaluate, playwright_browser_file_upload, playwright_browser_fill_form, playwright_browser_find, playwright_browser_handle_dialog, playwright_browser_hover, playwright_browser_navigate, playwright_browser_navigate_back, playwright_browser_network_request, playwright_browser_network_requests, playwright_browser_press_key, playwright_browser_resize, playwright_browser_run_code_unsafe, playwright_browser_select_option, playwright_browser_snapshot, playwright_browser_tabs, playwright_browser_take_screenshot, playwright_browser_type, playwright_browser_wait_for, question, read, read_mcp_resource, session_info, session_list, session_read, session_search, skill, skill_mcp, task, team_approve_shutdown, team_create, team_delete, team_list, team_reject_shutdown, team_send_message, team_shutdown_request, team_status, team_task_create, team_task_get, team_task_list, team_task_update, todowrite, webfetch, websearch_web_search_exa, write.]

I tried reproducing this today with #36507 and couldn't — with the official FP8 model, tool calls always come back fine.

If you're self-hosting, could you post your launch command? And could you try TP without EP? This PR already fixed the EP deployment issue (EP is working on this branch now too).

Let me know how it goes.

I tried reproducing this today with #36507 and couldn't — with the official FP8 model, tool calls always come back fine.

If you're self-hosting, could you post your launch command? And could you try TP without EP? This PR already fixed the EP deployment issue (EP is working on this branch now too).

Let me know how it goes.

What SGLang image version do u use?

I'm on sglang:glm-5.3-flash
sha256: 34d5ee66b12fa6446cdae131c352b8f68cd85369e0e6fda115583805fada3891

There are my launch parameters running on 8xH100:

HOME: /tmp
      TMPDIR: /tmp
      TMP: /tmp
      TEMP: /tmp
      PYTORCH_CUDA_ALLOC_CONF: "expandable_segments:True"
    entrypoint: python3 -m sglang.launch_server
    command: >-
      --model-path /app/data/glm-5.3-flash
      --served-model-name glm-5.3-flash
      --host 0.0.0.0
      --port 8000
      --chat-template /app/data/glm-5.3-flash/chat_template.jinja
      --tp-size 8
      --ep-size 8
      --moe-runner-backend deep_gemm
      --dsa-prefill-backend tilelang
      --dsa-decode-backend tilelang
      --speculative-algorithm NEXTN
      --speculative-num-steps 3
      --speculative-eagle-topk 1
      --speculative-num-draft-tokens 4
      --enable-hierarchical-cache
      --hicache-size 145
      --hicache-mem-layout page_first_direct
      --hicache-io-backend direct
      --hicache-write-policy write_through
      --pre-warm-nccl
      --kv-cache-dtype bfloat16
      --page-size 64
      --mem-fraction-static 0.85
      --mamba-full-memory-ratio 0.7
      --mamba-max-states-per-path 4
      --max-running-requests 28
      --max-queued-requests 64
      --schedule-policy lpm
      --radix-eviction-policy slru
      --chunked-prefill-size 8192
      --max-prefill-tokens 8192
      --reasoning-parser glm45
      --tool-call-parser glm47
      --context-length 262144
      --sampling-defaults model
      --enable-cache-report
      --enable-metrics
      --trust-remote-code

I think the problem is clear now: the Docker image you're using doesn't include the EP bug fix yet. Please keep using that same image, but update the code to the latest commit of the PR I mentioned.

Also using this image on Hopper with EP enabled, thanks for the info

I tried reproducing this today with #36507 and couldn't — with the official FP8 model, tool calls always come back fine.

If you're self-hosting, could you post your launch command? And could you try TP without EP? This PR already fixed the EP deployment issue (EP is working on this branch now too).

Let me know how it goes.

Could you please share that issue number? I only found communicator_mhc issues that related to DP-attention ( #36884/#36885 ) but didn't see any EP mentions in that PR. Thanks!

it is #36884 with such as --tp-size 4 --ep-size 4 --dp-size 4 --enable-dp-attention

--enable-dp-attention — but it seems your Docker run doesn't have this on, so there might be another issue at play too.

I did test pure EP without DP attention before and hit some problems as well(repeate, This is rare.), but both are working fine now with the two fixes you mentioned.

Are things working on your side now?

Hi, we're seeing similar problems with model output with vllm and Hopper (H100) cards.

Using the recommended glm53-flash container image tag, the model starts fine, but the output is looped and corrupted. Among other things, it seems that the model cannot properly terminate its own output and produces meaningless text until it reaches max_tokens.

Here is an example taken from a response to the prompt Using Ansible, how can I deploy Docker on Ubuntu 24.04?:

The docker. 

The docker needs to stop.

# STOP.

```

h
// token: 0;

qu(## 补
```

`

OK. Deep. Breathing.Stop. The
docker 
(re docker)
(r Docker docker!)
 stop stop stop stop stop stop stop stop stop
D p
docker
'
 stop docker
'
Docker
Docker
D docker
docker 
d
docker docker

OK. 

I need to stop.

## STOP.

@moorglade This is more like weights are not downloaded correctly... can u share ur vllm serve command and double-check if the weights are complete ?

it is #36884 with such as --tp-size 4 --ep-size 4 --dp-size 4 --enable-dp-attention

--enable-dp-attention — but it seems your Docker run doesn't have this on, so there might be another issue at play too.

I did test pure EP without DP attention before and hit some problems as well(repeate, This is rare.), but both are working fine now with the two fixes you mentioned.

Are things working on your side now?

Turned off EP, but with no effect, still loops and corrupted tool_calls.

Are you using the top_p and temperature from https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/generation_config.json? Also, please run the git log inside your sglang docker so I can take a look.

Are you using the top_p and top_k from https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/generation_config.json? Also, please run the git log inside your sglang docker so I can take a look.

Yep, default generation params:

{
  "_from_model_config": true,
  "eos_token_id": [
    154820,
    154827,
    154829
  ],
  "pad_token_id": 154820,
  "temperature": 1.0,
  "top_p": 0.95,
  "transformers_version": "5.16.0"
}

There is not .git dir inside docker image so git log doesn't work, but anyway its a default image from docker hub with zero modifications: https://hub.docker.com/layers/lmsysorg/sglang/glm-5.3-flash/images/sha256-0836f0160fa785e424e68d13ef88ddd548f87e6e11ad9f0e4de982e4f9188aaf

Trying to build image from https://github.com/sgl-project/sglang/pull/36507 this PR - so maybe it will fix the issue, but Im failing to build it yet.

You do not need a new SGLang Docker image. The /sgl-workspace/sglang shipped in the image is a source snapshot, not a git checkout, so you can't fetch into it. Replace it with a real clone and check out the PR branch:

cd /sgl-workspace
mv sglang sglang.bak (or just rm it)
git clone https://github.com/sgl-project/sglang.git
cd sglang
git fetch origin pull/36507/head:pr-36507
git checkout pr-36507

You do not need a new SGLang Docker image. The /sgl-workspace/sglang shipped in the image is a source snapshot, not a git checkout, so you can't fetch into it. Replace it with a real clone and check out the PR branch:

cd /sgl-workspace
mv sglang sglang.bak (or just rm it)
git clone https://github.com/sgl-project/sglang.git
cd sglang
git fetch origin pull/36507/head:pr-36507
git checkout pr-36507

@ZHANGYUXUAN-zR I tried sglang on Hopper, with EP enabled and the last code in PR 36507, however I still get incoherent output

request:

{
  "max_tokens": 100,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
}

response:

{
  "id": "3e04181266694baeb99fca8138b14c78",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'm good, thanks! Here are the preliminary results: there are the operation modes you might need. Let me know if you have any specific questions about them.্ 3",
        "reasoning": "* Thinking: mediation、rule、conversation、okay,提醒,entity、parameter——用户输入的 equine、 nay——问好先实例化 `Miss`, fö——那个\n\n这份回复\n\nI'm doing well, thanks! How about you? Is there anything I can help you with today?"
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1788265095,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 99,
    "total_tokens": 118
  }
}

@fabric same issue here, tool_call issues still there, same with loops.

invalid [tool=bash]
The invalid tool was called with invalid arguments: SchemaError(Missing key
  at ["error"]).
Please rewrite the input so it satisfies the expected schema.

Unfortunately, unusable now.

@fabric same issue here, tool_call issues still there, same with loops.

invalid [tool=bash]
The invalid tool was called with invalid arguments: SchemaError(Missing key
  at ["error"]).
Please rewrite the input so it satisfies the expected schema.

Unfortunately, unusable now.

Are you also deploying on H100/H800 chips? I'll help troubleshoot together tomorrow, since I haven't experimented on this chip before.

@moorglade This is more like weights are not downloaded correctly... can u share ur vllm serve command and double-check if the weights are complete ?

@JaredforReal thanks, that was it.

I tried loading the model directly from Hugging Face, but this does not work because of a problem with file paths: https://github.com/vllm-project/vllm/pull/53906#issuecomment-5495260945

So I switched to the Run:ai model streamer to load the weights from S3, but it seems this path is also bugged: https://github.com/sgl-project/sglang/issues/37369 (the issue is from SGLang, but I observed similar behavior in vLLM).

Finally, I disabled the distributed mode in the streamer and now it works correctly. Again, thanks for the help :)

@fabric same issue here, tool_call issues still there, same with loops.

invalid [tool=bash]
The invalid tool was called with invalid arguments: SchemaError(Missing key
  at ["error"]).
Please rewrite the input so it satisfies the expected schema.

Unfortunately, unusable now.

Are you also deploying on H100/H800 chips? I'll help troubleshoot together tomorrow, since I haven't experimented on this chip before.

Exactly, 8xH100. I appreciate your help, thanks! Ready to experiment

You do not need a new SGLang Docker image. The /sgl-workspace/sglang shipped in the image is a source snapshot, not a git checkout, so you can't fetch into it. Replace it with a real clone and check out the PR branch:

cd /sgl-workspace
mv sglang sglang.bak (or just rm it)
git clone https://github.com/sgl-project/sglang.git
cd sglang
git fetch origin pull/36507/head:pr-36507
git checkout pr-36507

@ZHANGYUXUAN-zR I tried sglang on Hopper, with EP enabled and the last code in PR 36507, however I still get incoherent output

request:

{
  "max_tokens": 100,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
}

response:

{
  "id": "3e04181266694baeb99fca8138b14c78",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'm good, thanks! Here are the preliminary results: there are the operation modes you might need. Let me know if you have any specific questions about them.্ 3",
        "reasoning": "* Thinking: mediation、rule、conversation、okay,提醒,entity、parameter——用户输入的 equine、 nay——问好先实例化 `Miss`, fö——那个\n\n这份回复\n\nI'm doing well, thanks! How about you? Is there anything I can help you with today?"
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1788265095,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 99,
    "total_tokens": 118
  }
}

Hi, @fabric .

I have exactly tested on H100 chips using the newest github PR-36507. Follow the starting commands on Sglang cookbook:

sglang serve \
  --model-path zai-org/GLM-5.3-Flash \
  --tp-size 8 \
  --ep-size 8 \
  --mem-fraction-static 0.68 \
  --dsa-prefill-backend tilelang \
  --dsa-decode-backend tilelang \
  --kv-cache-dtype bfloat16 \
  --moe-runner-backend deep_gemm \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 5 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 6 \
  --speculative-adaptive \
  --reasoning-parser glm45 \
  --tool-call-parser glm47 \
  --host 0.0.0.0 \
  --port 30000

The model response is normal using the same req input of you:

"request": {
  "max_tokens": 1024,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
},

"response": {
  "id": "813f866f5d9f4c1288d0a6b3bd8fc459",
  "object": "chat.completion",
  "created": 1788276249,
  "model": "GLM-5.3-Flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm doing well, thanks for asking. How about you — how's your day going?\n\nWhether you have a question, something you're working on, or just feel like chatting, I'm happy to help.",
        "reasoning_content": "The user is greeting me with a simple, friendly \"Hello, how are you today?\" This is a casual conversational opener. Let me think about what an appropriate response would be.\n\nFirst, the question itself: \"How are you?\" when asked to an AI is interesting. I don't have feelings in the human sense, but I shouldn't be overly pedantic or lecture-y about this. The person is just being friendly and making small talk. The best response is warm, engaged, and matches their conversational energy.\n\nI shouldn't:\n1. Launch into a lengthy explanation of how I don't have experiences/feelings in the way humans do — that would be weird and overly serious for a casual greeting.\n2. Be robotic or dismissive (\"I am an AI, I do not have feelings\").\n3. Be overly effusive or fake.\n\nI should:\n- Be warm and friendly\n- Respond naturally to the greeting\n- Perhaps briefly acknowledge that I'm doing well in a way that's honest but not ponderous\n- 

                                                     .......... too long just omit by me",

        "tool_calls": null
      },
      "logprobs": null,
      "finish_reason": "stop",
      "matched_stop": 154827
    }
  ],
  "usage": {
    "prompt_tokens": 19,
    "total_tokens": 983,
    "completion_tokens": 964,
    "prompt_tokens_details": null,
    "reasoning_tokens": 918
  }
}

So what's your starting commands?

@Dovis01 such simple requests works well for me, but model start degrading under the load of 5-10 concurrent requests.

@fabric same issue here, tool_call issues still there, same with loops.

invalid [tool=bash]
The invalid tool was called with invalid arguments: SchemaError(Missing key
  at ["error"]).
Please rewrite the input so it satisfies the expected schema.

Unfortunately, unusable now.

Are you also deploying on H100/H800 chips? I'll help troubleshoot together tomorrow, since I haven't experimented on this chip before.

Exactly, 8xH100. I appreciate your help, thanks! Ready to experiment

Hi, @d3lavar

I think it works. I have tested on H100.

"request": {
      "max_tokens": 512,
      "messages": [
        {
          "role": "user",
          "content": "Please run the shell command 'df -h' to check disk usage. Use the bash tool."
        }
      ],
      "model": "glm-5-next-fp8",
      "stream": false,
      "temperature": 0.6,
      "tools": [
        {
          "type": "function",
          "function": {
            "name": "bash",
            "description": "Run a shell command and return its output",
            "parameters": {
              "type": "object",
              "properties": {
                "command": {
                  "type": "string",
                  "description": "The command to run"
                }
              },
              "required": [
                "command"
              ]
            }
          }
        }
      ],
      "tool_choice": "auto"
    },
"response": {
  "id": "aac2894ed73c4b72b6672de88cd75813",
  "object": "chat.completion",
  "created": 1788276250,
  "model": "glm-5-next-fp8",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "",
        "reasoning_content": null,
        "tool_calls": [
          {
            "id": "call_24223faafb4e42cabf82ced6",
            "index": 0,
            "type": "function",
            "function": {
              "name": "bash",
              "arguments": "{\"command\": \"df -h\"}"
            }
          }
        ]
      },
      "logprobs": null,
      "finish_reason": "tool_calls",
      "matched_stop": null
    }
  ],
  "usage": {
    "prompt_tokens": 185,
    "total_tokens": 198,
    "completion_tokens": 13,
    "prompt_tokens_details": null,
    "reasoning_tokens": 1
  },
  "metadata": {
    "weight_version": "default",
    "weight_versions": [
      {
        "version": "default",
        "start": 0,
        "end": 13
      }
    ]
  }
}

@Dovis01 such simple requests works well for me, but model start degrading under the load of 5-10 concurrent requests.

Hi, @d3lavar

You can try again. I also reproduced it and simulated some multi-turn agent tool-calling requests. it still works normally

@fabric same issue here, tool_call issues still there, same with loops.

invalid [tool=bash]
The invalid tool was called with invalid arguments: SchemaError(Missing key
  at ["error"]).
Please rewrite the input so it satisfies the expected schema.

Unfortunately, unusable now.

Are you also deploying on H100/H800 chips? I'll help troubleshoot together tomorrow, since I haven't experimented on this chip before.

H100

You do not need a new SGLang Docker image. The /sgl-workspace/sglang shipped in the image is a source snapshot, not a git checkout, so you can't fetch into it. Replace it with a real clone and check out the PR branch:

cd /sgl-workspace
mv sglang sglang.bak (or just rm it)
git clone https://github.com/sgl-project/sglang.git
cd sglang
git fetch origin pull/36507/head:pr-36507
git checkout pr-36507

@ZHANGYUXUAN-zR I tried sglang on Hopper, with EP enabled and the last code in PR 36507, however I still get incoherent output

request:

{
  "max_tokens": 100,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
}

response:

{
  "id": "3e04181266694baeb99fca8138b14c78",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'm good, thanks! Here are the preliminary results: there are the operation modes you might need. Let me know if you have any specific questions about them.্ 3",
        "reasoning": "* Thinking: mediation、rule、conversation、okay,提醒,entity、parameter——用户输入的 equine、 nay——问好先实例化 `Miss`, fö——那个\n\n这份回复\n\nI'm doing well, thanks! How about you? Is there anything I can help you with today?"
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1788265095,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 99,
    "total_tokens": 118
  }
}

Hi, @fabric .

I have exactly tested on H100 chips using the newest github PR-36507. Follow the starting commands on Sglang cookbook:

sglang serve \
  --model-path zai-org/GLM-5.3-Flash \
  --tp-size 8 \
  --ep-size 8 \
  --mem-fraction-static 0.68 \
  --dsa-prefill-backend tilelang \
  --dsa-decode-backend tilelang \
  --kv-cache-dtype bfloat16 \
  --moe-runner-backend deep_gemm \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 5 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 6 \
  --speculative-adaptive \
  --reasoning-parser glm45 \
  --tool-call-parser glm47 \
  --host 0.0.0.0 \
  --port 30000

The model response is normal using the same req input of you:

"request": {
  "max_tokens": 1024,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
},

"response": {
  "id": "813f866f5d9f4c1288d0a6b3bd8fc459",
  "object": "chat.completion",
  "created": 1788276249,
  "model": "GLM-5.3-Flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm doing well, thanks for asking. How about you — how's your day going?\n\nWhether you have a question, something you're working on, or just feel like chatting, I'm happy to help.",
        "reasoning_content": "The user is greeting me with a simple, friendly \"Hello, how are you today?\" This is a casual conversational opener. Let me think about what an appropriate response would be.\n\nFirst, the question itself: \"How are you?\" when asked to an AI is interesting. I don't have feelings in the human sense, but I shouldn't be overly pedantic or lecture-y about this. The person is just being friendly and making small talk. The best response is warm, engaged, and matches their conversational energy.\n\nI shouldn't:\n1. Launch into a lengthy explanation of how I don't have experiences/feelings in the way humans do — that would be weird and overly serious for a casual greeting.\n2. Be robotic or dismissive (\"I am an AI, I do not have feelings\").\n3. Be overly effusive or fake.\n\nI should:\n- Be warm and friendly\n- Respond naturally to the greeting\n- Perhaps briefly acknowledge that I'm doing well in a way that's honest but not ponderous\n- 

                                                     .......... too long just omit by me",

        "tool_calls": null
      },
      "logprobs": null,
      "finish_reason": "stop",
      "matched_stop": 154827
    }
  ],
  "usage": {
    "prompt_tokens": 19,
    "total_tokens": 983,
    "completion_tokens": 964,
    "prompt_tokens_details": null,
    "reasoning_tokens": 918
  }
}

So what's your starting commands?

Hi, same code and parameters actually, except that I load the model with RunAI streamer.
Sometimes the output is correct, sometimes it's really funky. As @d3lavar I observed a degradation over time.

I've seen some comments about this issue with Runai: https://github.com/sgl-project/sglang/issues/37369
Going to try without distributed mode to see how it goes.

We're running GLM-5.3-Flash FP8 on 8x H100 80GB (a3-highgpu-8g) via Vertex AI with vLLM. The model is unusable for real agentic workloads and we're looking for help. Sharing everything in case someone can point us in the right direction.

Setup

Component Value
Hardware 8x NVIDIA H100 80GB (a3-highgpu-8g)
Image vllm/vllm-openai:glm53-flash-x86_64-cu129 (Aug 26 build, post-v0.28.0)
Model FP8 checkpoint from HuggingFace
Engine vLLM (not SGLang)
Tensor Parallel 8
KV cache BF16 (default)
Prefix caching ON

vLLM serve command

vllm serve gs://<bucket>/models/glm-5-3-flash \
  --tensor-parallel-size 8 \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --enable-auto-tool-choice \
  --chat-template-content-format string \
  --max-num-seqs 256 \
  --served-model-name glm-5.3-flash-gcp

What's broken

The model degrades and becomes unusable in real agentic sessions. After several turns of tool calling, it starts producing garbage: language switching (English to Chinese to random fragments), repetition loops, and incoherent output. It cannot complete basic research tasks. Short single-turn tests look fine, but real multi-turn sessions with tool calls fall apart.

We found one partial fix: --chat-template-content-format string. Without it, vLLM renders content: null in assistant tool-call messages as the literal string "None", which accumulates in context and corrupts it. This flag is on the GLM-5 and GLM-5.1 recipe pages but missing from the GLM-5.3-Flash recipe page. Adding it helped with tool calling, but the model is still unusable for real workloads.

Related vLLM issues: #39611, #39614 (both still open).

FP8 KV cache crashes on H100

Tried --kv-cache-dtype fp8 per the recipe page. Model loads fine (43.3 GiB/GPU across 8 GPUs) but crashes during CUDA graph profiling:

RuntimeError: concat_and_cache_mla, cache_kernels.cu:866, pe_dim must be 64 for fp8_ds_mla

GLM-5.3-Flash uses hybrid KDA + sparse MLA attention (34 linear / 11 sparse layers). The FP8 DS-MLA kernel requires pe_dim=64, but the model's MLA layers use a different dimension. We confirmed this works on Blackwell (GB10) but not on Hopper (H100). Anyone got FP8 KV cache working on H100 with this model?

Output quality degradation at ~100K context

Even with the chat template fix, output quality degrades around 100K context. The vLLM server stays up (no OOM, KV cache at ~13% usage, 2M token KV capacity), so this is a model quality issue, not a crash. Prefix cache hit rate is 88.5% so disabling prefix caching is not an option for us. Anyone seeing similar long-context degradation?

MTP + tool calling broken

Confirmed vLLM #44843 — GLM + MTP + tool calling is broken. We removed MTP from our deployment.

Metrics

Metric Value
Prefix cache hit rate 88.5%
KV cache usage (peak) 13.9%
GPU memory utilization 0.92 (vLLM default)
Successful completions 130
Aborted / errored 0
Repetition finishes 0

The server metrics look healthy but the model output does not. We need help — is anyone running GLM-5.3-Flash on vLLM with H100s for real agentic workloads without these issues? Should we switch to SGLang?

great report @llaforest ; are you using runai streamer to load the model (I guess so if your model is on GCS)? There is a filed issue about this for sglang https://github.com/sgl-project/sglang/issues/37369 ; vllm might have the same issue.
FP8 KV cache is not available on Hopper for this model, only BF16.

@llaforest @fabric for the Run:ai streamer in vLLM, we had to disable the distrubted mode (as I've mentioned above), i.e.:

--load-format runai_streamer --model-loader-extra-config '{"distributed":false}'

@fabric @moorglade Thanks for your response!
@llaforest Thanks for your report, this is also look like a model weights problem to me, we haven't able to reproduce this issue, and as moorglade have encountered, run:ai streamer model download got some issue for now
It's recommended to download the model weights from HF directly by hf-cli

Thanks for the responses.

We're not using Run:ai streamer, our vLLM command is just --model=gs:///models/glm-5-3-flash with no --load-format flag, so vLLM uses its default loader. The weights were downloaded from HuggingFace and uploaded to GCS via gsutil.

The model loads without errors (43.3 GiB/GPU across 8 GPUs, no warnings). The garbage output only appears after several turns of tool calling, short single-turn prompts work fine. Would corrupted weights degrade progressively like this, or would they produce garbage from the first token?

We can try loading directly via hf-cli inside the container as you recommend. Is there a way to verify weight integrity against the HF repo (checksums/sha256) without re-downloading everything?

Also would love to try the same with BF16 weights, but hard time putting my hands on a H200 or B200 to do it.

Sign up or log in to comment