2026-08-28 00:04:52 model: ~/.cache/huggingface/hub/models--unsloth--GLM-5.3-Flash-GGUF/snapshots/ac47690c15c8703615ab7d9c1ef2293d45372757/UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.gguf 2026-08-28 00:04:52 baseline mem: 7 GB used 2026-08-28 00:05:17 loaded in 25 s; mem during load peak 92 GB 2026-08-28 00:05:17 --- Say OK {"choices":[{"finish_reason":"stop","index":0,"message":{"role":"assistant","content":"OK","reasoning_content":"The user just wants me to say \"OK.\""}}],"created":1787893519,"model":"~/.cache/huggingface/hub/models--unsloth--GLM-5.3-Flash-GGUF/snapshots/ac47690c15c8703615ab7d9c1ef2293d45372757/UD-IQ1_S/GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.gguf","system_fingerprint":"b1-1f0a36a","object":"chat.completion","usage":{"completion_tokens":13,"prompt_tokens":15,"total_tokens":28,"prompt_tokens_details":{"cached_tokens":0}},"id":"chatcmpl-zLAEgJRngRwM6I5IbHqClk24lGJ3xddp","timings":{"cache_n":0,"prompt_n":15,"prompt_ms":861.318,"prompt_per_token_ms":57.4212,"prompt_per_second":17.415170703503236,"predicted_n":13,"predicted_ms":1224.954,"predicted_per_token_ms":102.0795,"predicted_per_second":9.796286227891006}}2026-08-28 00:05:19 --- real prompt (200 tokens) gen tok/s: 9.7 | prefill tok/s: 28.1 | tokens: 266 reasoning chars: 0 | answer: A mixture-of-experts (MoE) language model is a type of neural network architecture. Here's what you need to know: / / - **The core idea:** Instead of one big, monolithic neural network processing every token, the model contains many "expert" sub-networks (often dozens or hundreds), plus a routing mechanism. / - **How routing works:** For each token, a lightweight "router" network decides which experts 2026-08-28 00:05:48 peak mem now: 93 GB used 2026-08-28 00:05:58 server stopped