Instructions to use IFM/K2-Horizon-7B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IFM/K2-Horizon-7B-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IFM/K2-Horizon-7B-FP8", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("IFM/K2-Horizon-7B-FP8", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IFM/K2-Horizon-7B-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IFM/K2-Horizon-7B-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-7B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IFM/K2-Horizon-7B-FP8
- SGLang
How to use IFM/K2-Horizon-7B-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IFM/K2-Horizon-7B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-7B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IFM/K2-Horizon-7B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-7B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use IFM/K2-Horizon-7B-FP8 with Docker Model Runner:
docker model run hf.co/IFM/K2-Horizon-7B-FP8
[Bug] chat_template.jinja crashes with "can only concatenate str (not 'list') to str" when served via vLLM
Summary & Workaround
When serving IFM/K2-Horizon-7B-FP8 with vLLM, sending any chat completion request containing a system prompt results in an HTTP 400 error:
Chat template rejected the request: can only concatenate str (not "list") to str
TL;DR Workaround for vLLM users:
Start vLLM with--chat-template-content-format stringto bypass the issue immediately.
Root Cause
- In
chat_template.jinja(lines 927–929):{%- if messages[0].role == 'system' -%} {{- '<|ifm|im_start|>system\n' + messages[0].content + '<|ifm|im_end|>' }} {%- endif -%} - During server startup, vLLM statically parses the Jinja template AST (
_detect_content_format). Becauserender_tool_response_messagesiterates overraw_content({% for item in raw_content %}), vLLM identifies the template as supporting OpenAI-style structured content parts and setscontent_format = 'openai'. - Under
'openai'format, vLLM normalizes incoming message content into a list of parts:messages[0]['content'] = [{'type': 'text', 'text': '...'}] - The template then attempts direct string concatenation (
str + list):'<|ifm|im_start|>system\n' + messages[0].content
which fails with:TypeError: can only concatenate str (not "list") to str - Additionally, in lines 933–937:
If{%- if message.content is string -%} {%- set content = message.content -%} {%- else -%} {%- set content = '' -%} {%- endif -%}content_formatis'openai', user message content is also dropped to an empty string""becausemessage.content is stringevaluates to false.
Proposed Fix
Safely extract string content from both raw strings and OpenAI structured content lists in chat_template.jinja:
{%- if messages[0].role == 'system' -%}
{%- if messages[0].content is string -%}
{%- set sys_text = messages[0].content -%}
{%- elif messages[0].content is sequence and messages[0].content | length > 0 and messages[0].content[0].text is defined -%}
{%- set sys_text = messages[0].content | map(attribute='text') | join('\n') -%}
{%- else -%}
{%- set sys_text = messages[0].content | string -%}
{%- endif -%}
{{- '<|ifm|im_start|>system\n' + sys_text + '<|ifm|im_end|>' }}
{%- endif -%}
Drafted with Gemini, reviewed by @ydawei .
Hi @moonfolk , here is a minimal reproduction of the error.
1. Serve the model with stock flags and no template overrides:
docker run -d --name k2-repro --ipc=host --network host \
--device /dev/kfd --device /dev/dri --group-add video \
--security-opt seccomp=unconfined \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HF_TOKEN=*** \
vllm/vllm-openai-rocm:nightly \
IFM/K2-Horizon-7B-FP8 \
--quantization=None --max-model-len=8192 \
--gpu-memory-utilization=0.9 --max-num-seqs=2 \
--trust-remote-code
I tested this on AMD ROCm with vLLM 0.29.1rc1.dev47+gdc36fcce9. The failure happens inside the chat template, so it is not specific to the GPU backend.
2. Any chat completion request that contains a system message fails:
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "IFM/K2-Horizon-7B-FP8",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello"}
],
"max_tokens": 16
}'
{"error":{"message":"can only concatenate str (not \"list\") to str","type":"BadRequestError","code":400}}
The same request succeeds on the same server when the system message is removed.
3. Adding --chat-template-content-format string to the serve command fixes the issue. The request in step 2 succeeds once that flag is set.