0.00.018.170 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.018.252 I srv init: The UI is disabled 0.00.018.253 I srv init: Use --ui/--no-ui (or deprecated --webui/--no-webui) to enable/disable 0.00.018.312 W srv llama_server: ----------------- 0.00.018.313 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set 0.00.018.313 W srv llama_server: this can be a security risk (cross-origin attacks) 0.00.018.314 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 0.00.018.314 W srv llama_server: ----------------- 0.00.019.498 I srv load_model: loading model '~/.cache/huggingface/hub/models--unsloth--GLM-5.3-Flash-GGUF/snapshots/ac47690c15c8703615ab7d9c1ef2293d45372757/UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.gguf' 0.00.543.149 W load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect 0.00.543.185 W load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect 0.00.580.402 W model has unused tensor blk.45.attn_norm.weight (size = 16384 bytes) -- ignoring 0.00.580.406 W model has unused tensor blk.45.ffn_norm.weight (size = 16384 bytes) -- ignoring 0.00.580.408 W model has unused tensor blk.45.attn_q_a.weight (size = 4325376 bytes) -- ignoring 0.00.580.410 W model has unused tensor blk.45.attn_q_a_norm.weight (size = 6144 bytes) -- ignoring 0.00.580.413 W model has unused tensor blk.45.attn_q_b.weight (size = 26738688 bytes) -- ignoring 0.00.580.415 W model has unused tensor blk.45.attn_kv_a_mqa.weight (size = 2228224 bytes) -- ignoring 0.00.580.417 W model has unused tensor blk.45.attn_kv_a_norm.weight (size = 2048 bytes) -- ignoring 0.00.580.419 W model has unused tensor blk.45.attn_k_b.weight (size = 8912896 bytes) -- ignoring 0.00.580.420 W model has unused tensor blk.45.attn_v_b.weight (size = 8912896 bytes) -- ignoring 0.00.580.422 W model has unused tensor blk.45.attn_output.weight (size = 46137344 bytes) -- ignoring 0.00.580.425 W model has unused tensor blk.45.indexer.k_norm.weight (size = 512 bytes) -- ignoring 0.00.580.427 W model has unused tensor blk.45.indexer.k_norm.bias (size = 512 bytes) -- ignoring 0.00.580.429 W model has unused tensor blk.45.indexer.proj.weight (size = 524288 bytes) -- ignoring 0.00.580.432 W model has unused tensor blk.45.indexer.attn_k.weight (size = 557056 bytes) -- ignoring 0.00.580.434 W model has unused tensor blk.45.indexer.attn_q_b.weight (size = 6684672 bytes) -- ignoring 0.00.580.436 W model has unused tensor blk.45.indexer_compressor_gate.weight (size = 557056 bytes) -- ignoring 0.00.580.439 W model has unused tensor blk.45.indexer_compressor_ape.weight (size = 2048 bytes) -- ignoring 0.00.580.441 W model has unused tensor blk.45.ffn_gate_inp.weight (size = 4718592 bytes) -- ignoring 0.00.580.444 W model has unused tensor blk.45.exp_probs_b.bias (size = 1152 bytes) -- ignoring 0.00.580.446 W model has unused tensor blk.45.ffn_gate_exps.weight (size = 792723456 bytes) -- ignoring 0.00.580.448 W model has unused tensor blk.45.ffn_up_exps.weight (size = 792723456 bytes) -- ignoring 0.00.580.450 W model has unused tensor blk.45.ffn_down_exps.weight (size = 1038090240 bytes) -- ignoring 0.00.580.452 W model has unused tensor blk.45.ffn_gate_shexp.weight (size = 5767168 bytes) -- ignoring 0.00.580.454 W model has unused tensor blk.45.ffn_up_shexp.weight (size = 5767168 bytes) -- ignoring 0.00.580.456 W model has unused tensor blk.45.ffn_down_shexp.weight (size = 6881280 bytes) -- ignoring 0.00.580.459 W model has unused tensor blk.45.nextn.eh_proj.weight (size = 35651584 bytes) -- ignoring 0.00.580.462 W model has unused tensor blk.45.nextn.enorm.weight (size = 16384 bytes) -- ignoring 0.00.580.464 W model has unused tensor blk.45.nextn.hnorm.weight (size = 16384 bytes) -- ignoring 0.00.580.466 W model has unused tensor blk.45.nextn.shared_head_norm.weight (size = 16384 bytes) -- ignoring 0.22.821.704 W llama_kv_cache: the V embeddings have different sizes across layers and FA is not enabled - padding V cache to 512 0.22.899.409 W resolve_fused_ops: layer 0 is assigned to device Vulkan0 but fused DeepSeek V4 HC pre is assigned to device CPU (usually due to missing support) 0.22.899.412 W resolve_fused_ops: fused DeepSeek V4 HC pre not supported, set to disabled 0.22.901.518 W resolve_fused_ops: layer 0 is assigned to device Vulkan0 but fused DeepSeek V4 HC comb is assigned to device CPU (usually due to missing support) 0.22.901.519 W resolve_fused_ops: fused DeepSeek V4 HC comb not supported, set to disabled 0.22.911.168 W resolve_fused_ops: layer 0 is assigned to device Vulkan0 but fused DeepSeek V4 HC post is assigned to device CPU (usually due to missing support) 0.22.911.169 W resolve_fused_ops: fused DeepSeek V4 HC post not supported, set to disabled 0.22.961.280 W ggml_vulkan: Failed to allocate pinned memory (Requested buffer size exceeds device buffer size limit: ErrorOutOfDeviceMemory) 0.22.993.395 I cmn init: llama threadpool init, n_threads = 16 0.24.152.844 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 32768, kv_unified = 'false' 0.24.170.849 I srv init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve 0.24.170.893 I srv llama_server: model loaded 0.24.170.899 I srv llama_server: listening on http://127.0.0.1:8897 0.25.097.753 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 0.25.097.841 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 0.27.184.159 I slot print_timing: id 0 | task 0 | prompt eval time = 861.32 ms / 15 tokens ( 57.42 ms per token, 17.42 tokens per second) 0.27.184.162 I slot print_timing: id 0 | task 0 | eval time = 1224.95 ms / 13 tokens ( 102.08 ms per token, 9.80 tokens per second) 0.27.184.163 I slot print_timing: id 0 | task 0 | total time = 2086.27 ms / 28 tokens 0.27.184.168 I slot print_timing: id 0 | task 0 | graphs reused = 0 0.27.184.200 I slot release: id 0 | task 0 | stop processing: n_tokens = 27, truncated = 0 0.27.201.775 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.270 (> 0.100 thold), f_keep = 0.370 0.27.248.998 I slot launch_slot_: id 0 | task 16 | processing task, is_child = 0 0.38.446.321 I slot print_timing: id 0 | task 16 | n_gen = 100, tg = 9.71 t/s, tg_3s = 9.80 t/s 0.41.465.948 I slot print_timing: id 0 | task 16 | n_gen = 129, tg = 9.68 t/s, tg_3s = 9.60 t/s 0.44.526.043 I slot print_timing: id 0 | task 16 | n_gen = 158, tg = 9.64 t/s, tg_3s = 9.48 t/s 0.47.613.749 I slot print_timing: id 0 | task 16 | n_gen = 188, tg = 9.66 t/s, tg_3s = 9.72 t/s 0.50.706.072 I slot print_timing: id 0 | task 16 | n_gen = 218, tg = 9.66 t/s, tg_3s = 9.70 t/s 0.53.731.478 I slot print_timing: id 0 | task 16 | n_gen = 248, tg = 9.69 t/s, tg_3s = 9.92 t/s 0.55.575.956 I slot print_timing: id 0 | task 16 | prompt eval time = 997.84 ms / 28 tokens ( 35.64 ms per token, 28.06 tokens per second) 0.55.575.960 I slot print_timing: id 0 | task 16 | eval time = 27329.09 ms / 266 tokens ( 103.13 ms per token, 9.70 tokens per second) 0.55.575.961 I slot print_timing: id 0 | task 16 | total time = 28326.93 ms / 294 tokens 0.55.575.963 I slot print_timing: id 0 | task 16 | graphs reused = 0 0.55.575.987 I slot release: id 0 | task 16 | stop processing: n_tokens = 302, truncated = 0 0.55.595.862 I srv operator(): operator(): cleaning up before exit...