[Bug] chat_template.jinja crashes with "can only concatenate str (not 'list') to str" when served via vLLM

#4
by ydawei - opened

Summary & Workaround

When serving IFM/K2-Horizon-7B-FP8 with vLLM, sending any chat completion request containing a system prompt results in an HTTP 400 error:

Chat template rejected the request: can only concatenate str (not "list") to str

TL;DR Workaround for vLLM users:
Start vLLM with --chat-template-content-format string to bypass the issue immediately.


Root Cause

  1. In chat_template.jinja (lines 927–929):
    {%- if messages[0].role == 'system' -%}
        {{- '<|ifm|im_start|>system\n' + messages[0].content + '<|ifm|im_end|>' }}
    {%- endif -%}
    
  2. During server startup, vLLM statically parses the Jinja template AST (_detect_content_format). Because render_tool_response_messages iterates over raw_content ({% for item in raw_content %}), vLLM identifies the template as supporting OpenAI-style structured content parts and sets content_format = 'openai'.
  3. Under 'openai' format, vLLM normalizes incoming message content into a list of parts:
    messages[0]['content'] = [{'type': 'text', 'text': '...'}]
  4. The template then attempts direct string concatenation (str + list):
    '<|ifm|im_start|>system\n' + messages[0].content
    which fails with:
    TypeError: can only concatenate str (not "list") to str
  5. Additionally, in lines 933–937:
    {%- if message.content is string -%}
        {%- set content = message.content -%}
    {%- else -%}
        {%- set content = '' -%}
    {%- endif -%}
    
    If content_format is 'openai', user message content is also dropped to an empty string "" because message.content is string evaluates to false.

Proposed Fix

Safely extract string content from both raw strings and OpenAI structured content lists in chat_template.jinja:

{%- if messages[0].role == 'system' -%}
    {%- if messages[0].content is string -%}
        {%- set sys_text = messages[0].content -%}
    {%- elif messages[0].content is sequence and messages[0].content | length > 0 and messages[0].content[0].text is defined -%}
        {%- set sys_text = messages[0].content | map(attribute='text') | join('\n') -%}
    {%- else -%}
        {%- set sys_text = messages[0].content | string -%}
    {%- endif -%}
    {{- '<|ifm|im_start|>system\n' + sys_text + '<|ifm|im_end|>' }}
{%- endif -%}

Drafted with Gemini, reviewed by @ydawei .

Institute of Foundation Models org

Hi @ydawei . Could you please share example code to reproduce the error?

Hi @moonfolk , here is a minimal reproduction of the error.

1. Serve the model with stock flags and no template overrides:

docker run -d --name k2-repro --ipc=host --network host \
  --device /dev/kfd --device /dev/dri --group-add video \
  --security-opt seccomp=unconfined \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HF_TOKEN=*** \
  vllm/vllm-openai-rocm:nightly \
  IFM/K2-Horizon-7B-FP8 \
  --quantization=None --max-model-len=8192 \
  --gpu-memory-utilization=0.9 --max-num-seqs=2 \
  --trust-remote-code

I tested this on AMD ROCm with vLLM 0.29.1rc1.dev47+gdc36fcce9. The failure happens inside the chat template, so it is not specific to the GPU backend.

2. Any chat completion request that contains a system message fails:

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "IFM/K2-Horizon-7B-FP8",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello"}
    ],
    "max_tokens": 16
  }'
{"error":{"message":"can only concatenate str (not \"list\") to str","type":"BadRequestError","code":400}}

The same request succeeds on the same server when the system message is removed.

3. Adding --chat-template-content-format string to the serve command fixes the issue. The request in step 2 succeeds once that flag is set.

Sign up or log in to comment