W0923 17:22:02.312000 1 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:02.329000 1 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. (APIServer pid=1) [transformers] The following generation flags are not valid and may be ignored: ['top_p']. Set `TRANSFORMERS_VERBOSITY=info` for more details. W0923 17:22:20.295000 241 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:20.307000 241 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:26.140000 264 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:26.140000 263 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:26.153000 264 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. W0923 17:22:26.153000 263 torch/utils/_pytree.py:630] is an Enum subclass and is now natively supported by torch.compile as an opaque value type. Calling register_constant() on Enum subclasses is deprecated and will be an error in a future release. (Worker_TP0_EP0 pid=263) Loading safetensors checkpoint shards: 0% Completed | 0/120 [00:00, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 8, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='b12x', linear_backend='auto') (EngineCore pid=241) INFO 09-23 17:22:22 [multiproc_executor.py:165] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=100.64.0.10 (local), world_size=2, local_world_size=2 (Worker pid=263) INFO 09-23 17:22:31 [parallel_state.py:1938] world_size=2 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_0da880958d2a4578819a81d6824758ba backend=nccl (Worker pid=264) INFO 09-23 17:22:31 [parallel_state.py:1938] world_size=2 rank=1 local_rank=1 distributed_init_method=file:///tmp/vllm_dist_0da880958d2a4578819a81d6824758ba backend=nccl (Worker pid=264) INFO 09-23 17:22:31 [nccl.py:25] Found nccl from environment variable VLLM_NCCL_SO_PATH=/opt/local-inference/nccl/lib/libnccl.so.2.31.2 (Worker pid=263) INFO 09-23 17:22:31 [nccl.py:25] Found nccl from environment variable VLLM_NCCL_SO_PATH=/opt/local-inference/nccl/lib/libnccl.so.2.31.2 (Worker pid=263) INFO 09-23 17:22:31 [pynccl.py:113] vLLM is using nccl==2.31.2 (Worker pid=263) WARNING 09-23 17:22:31 [symm_mem.py:101] SymmMemCommunicator: native P2P atomics are not supported between devices [0, 1], communicator is not available. (Worker pid=264) WARNING 09-23 17:22:31 [symm_mem.py:101] SymmMemCommunicator: native P2P atomics are not supported between devices [0, 1], communicator is not available. (Worker pid=263) INFO 09-23 17:22:31 [cuda_communicator.py:274] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'tp:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL']. (Worker pid=263) INFO 09-23 17:22:31 [nccl.py:25] Found nccl from environment variable VLLM_NCCL_SO_PATH=/opt/local-inference/nccl/lib/libnccl.so.2.31.2 (Worker pid=264) INFO 09-23 17:22:31 [nccl.py:25] Found nccl from environment variable VLLM_NCCL_SO_PATH=/opt/local-inference/nccl/lib/libnccl.so.2.31.2 (Worker pid=263) INFO 09-23 17:22:31 [cuda_communicator.py:274] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'ep:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL']. (Worker pid=263) INFO 09-23 17:22:31 [parallel_state.py:2344] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A (Worker pid=263) INFO 09-23 17:22:31 [gpu_worker.py:414] Using V2 Model Runner (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [model_runner.py:447] Loading model from scratch... (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [cuda.py:589] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [mm_encoder_attention.py:375] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention. (Worker_TP1_EP1 pid=264) INFO 09-23 17:22:32 [kernel.py:310] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']) (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [kernel.py:310] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']) (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant (Worker_TP1_EP1 pid=264) INFO 09-23 17:22:32 [cuda.py:470] Using AttentionBackendEnum.FLASHINFER_MLA_SPARSE_SM120 backend. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [cuda.py:470] Using AttentionBackendEnum.FLASHINFER_MLA_SPARSE_SM120 backend. (Worker_TP0_EP0 pid=263) WARNING 09-23 17:22:32 [mla_attention.py:844] Sparse MLA layer has no dense-MHA prefill path; using the top-k MQA path only. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [deep_gemm.py:207] DeepGEMM PDL enabled on deep_gemm. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [deep_gemm.py:125] DeepGEMM E8M0 enabled on current platform. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [expert_map_manager.py:245] [EP Rank 0/2] Expert parallelism is enabled. Expert placement strategy: linear. Local/global number of experts: 144/288. Experts local to global index map: 0->0, 1->1, 2->2, 3->3, 4->4, 5->5, 6->6, 7->7, 8->8, 9->9, 10->10, 11->11, 12->12, 13->13, 14->14, 15->15, 16->16, 17->17, 18->18, 19->19, 20->20, 21->21, 22->22, 23->23, 24->24, 25->25, 26->26, 27->27, 28->28, 29->29, 30->30, 31->31, 32->32, 33->33, 34->34, 35->35, 36->36, 37->37, 38->38, 39->39, 40->40, 41->41, 42->42, 43->43, 44->44, 45->45, 46->46, 47->47, 48->48, 49->49, 50->50, 51->51, 52->52, 53->53, 54->54, 55->55, 56->56, 57->57, 58->58, 59->59, 60->60, 61->61, 62->62, 63->63, 64->64, 65->65, 66->66, 67->67, 68->68, 69->69, 70->70, 71->71, 72->72, 73->73, 74->74, 75->75, 76->76, 77->77, 78->78, 79->79, 80->80, 81->81, 82->82, 83->83, 84->84, 85->85, 86->86, 87->87, 88->88, 89->89, 90->90, 91->91, 92->92, 93->93, 94->94, 95->95, 96->96, 97->97, 98->98, 99->99, 100->100, 101->101, 102->102, 103->103, 104->104, 105->105, 106->106, 107->107, 108->108, 109->109, 110->110, 111->111, 112->112, 113->113, 114->114, 115->115, 116->116, 117->117, 118->118, 119->119, 120->120, 121->121, 122->122, 123->123, 124->124, 125->125, 126->126, 127->127, 128->128, 129->129, 130->130, 131->131, 132->132, 133->133, 134->134, 135->135, 136->136, 137->137, 138->138, 139->139, 140->140, 141->141, 142->142, 143->143. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [weight_utils.py:935] Filesystem type for checkpoints: XFS. Checkpoint size: 163.58 GiB. Available RAM: 45.66 GiB. (Worker_TP0_EP0 pid=263) INFO 09-23 17:22:32 [weight_utils.py:965] Auto-prefetch is disabled because the filesystem (XFS) is not a recognized network FS (NFS/Lustre) and the checkpoint size (163.58 GiB) exceeds 90% of available RAM (45.66 GiB). (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:09 [default_loader.py:497] Loading weights took 36.85 seconds (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:11 [model_runner.py:469] Model loading took 80.58 GiB and 39.214554 seconds (Worker_TP1_EP1 pid=264) INFO 09-23 17:23:11 [model_runner.py:469] Model loading took 80.58 GiB and 39.218591 seconds (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:11 [topk_topp_sampler.py:62] Using FlashInfer for top-p & top-k sampling. (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:11 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend. (Worker_TP1_EP1 pid=264) INFO 09-23 17:23:11 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend. (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:11 [interface.py:931] Setting attention block size to 3328 tokens to ensure that attention page size is >= mamba page size. (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:11 [interface.py:955] Padding mamba page size by 0.57% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP1_EP1 pid=264) INFO 09-23 17:23:11 [interface.py:931] Setting attention block size to 3328 tokens to ensure that attention page size is >= mamba page size. (Worker_TP1_EP1 pid=264) INFO 09-23 17:23:11 [interface.py:955] Padding mamba page size by 0.57% to ensure that mamba page size and attention page size are exactly equal. (EngineCore pid=241) WARNING 09-23 17:23:11 [torch_utils.py:251] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling. (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:12 [encoder_runner.py:108] Encoder cache will be initialized with a budget of 7921 tokens, and profiled with 1 image items of the maximum feature size. (Worker_TP0_EP0 pid=263) WARNING 09-23 17:23:13 [common.py:241] The bundled vllm_flash_attn package does not provide layers.rotary; using the native PyTorch rotary implementation. (Worker_TP0_EP0 pid=263) 2026-09-23 17:23:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:23:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP0_EP0 pid=263) 2026-09-23 17:23:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:23:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang` (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:33 [exl3.py:4824] EXL3 full-expert EP runtime planned: Trellis m=1..32 block_m=8, prefill trellis block_m=128 capacity=2048 arena=669.5MiB scheduler_capacity=2048 chunk=128 topk=8 (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:34 [model_runner.py:1003] Profiled V1 prompt-logprobs workspace with chunk size 1024 (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:35 [kernel_warmup.py:115] Warming up ll_bf16 router GEMM kernels for shapes: ((4096, 288),). (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:35 [kernel_warmup.py:488] Deferring runtime-dependent kernel warmup until KV cache initialization. (Worker_TP0_EP0 pid=263) INFO 09-23 17:23:38 [indexer.py:891] DSA indexer decode path: use_flattening=False use_varlen=False (next_n=1, use_fp4_indexer_cache=False) (EngineCore pid=241) INFO 09-23 17:24:12 [shm_broadcast.py:801] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization). (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:21 [breakable_cudagraph.py:293] Breakable CUDA graph enabled (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:33 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:34 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:34 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP0_EP0 pid=263) 2026-09-23 17:24:39 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:23:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:23:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:23:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:23:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:33 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:34 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:34 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_EP1 pid=264) 2026-09-23 17:24:39 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:44 [model_runner.py:1246] Estimated MRV2 CUDA graph memory: 0.25 GiB total (0.11 GiB retained in the reusable pool) (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:44 [gpu_worker.py:661] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9700 is equivalent to --gpu-memory-utilization=0.9674 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.9726. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0. (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:44 [model_runner.py:1246] Estimated MRV2 CUDA graph memory: 0.25 GiB total (0.11 GiB retained in the reusable pool) (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:44 [gpu_worker.py:646] Available KV cache memory: 6.41 GiB (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:44 [gpu_worker.py:661] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9700 is equivalent to --gpu-memory-utilization=0.9674 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.9726. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0. (EngineCore pid=241) INFO 09-23 17:24:44 [kv_cache_utils.py:2580] GPU KV cache size: 825,268 tokens, Maximum concurrency for 327,680 tokens per request: 2.52x (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:44 [kernel_warmup.py:625] Using FlashInfer autotune cache file: /cache/jit/cu133-torch213-vllm174c789e09-b12x12c426322c-lmcachee045d729bc/flashinfer-autotune/31bab91f6e82675aae8970504d4f53ef2fbe8ea4fcc514054abed69eb1895c59/autotune_configs.json (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:48 [model_runner.py:1318] Graph capturing finished in 4 secs, took 0.06 GiB (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:48 [gpu_worker.py:813] CUDA graph pool memory: 0.06 GiB (actual), 0.25 GiB (estimated), difference: 0.18 GiB (284.8%). (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:48 [gpu_worker.py:876] Free memory on device (94.03/95.06 GiB) on startup. Desired GPU memory utilization is (0.97, 92.21 GiB). Actual usage is 84.89 GiB for consumed memory (weights + non-torch), 0.91 GiB for peak activation, and 0.06 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory-bytes=6654121288` (6.2 GiB) to fit into requested memory, or `--kv-cache-memory-bytes=8610739200` (8.02 GiB) to fully utilize gpu memory. Current kv cache memory in use is 6.41 GiB. (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:48 [model_runner.py:1318] Graph capturing finished in 4 secs, took 0.06 GiB (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:48 [gpu_worker.py:813] CUDA graph pool memory: 0.06 GiB (actual), 0.25 GiB (estimated), difference: 0.18 GiB (284.8%). (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:48 [gpu_worker.py:876] Free memory on device (94.03/95.06 GiB) on startup. Desired GPU memory utilization is (0.97, 92.21 GiB). Actual usage is 84.89 GiB for consumed memory (weights + non-torch), 0.91 GiB for peak activation, and 0.06 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory-bytes=6654121288` (6.2 GiB) to fit into requested memory, or `--kv-cache-memory-bytes=8610739200` (8.02 GiB) to fully utilize gpu memory. Current kv cache memory in use is 6.41 GiB. (Worker_TP0_EP0 pid=263) INFO 09-23 17:24:49 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn. (Worker_TP1_EP1 pid=264) INFO 09-23 17:24:49 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn. (Worker_TP0_EP0 pid=263) WARNING 09-23 17:24:50 [torch_utils.py:251] OMP_NUM_THREADS=2 is set; leaving Torch threads at 2 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling. (EngineCore pid=241) INFO 09-23 17:24:50 [core.py:367] init engine (profile, create kv cache, warmup model) took 99.25 s (EngineCore pid=241) INFO 09-23 17:24:53 [kernel.py:310] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']) (EngineCore pid=241) INFO 09-23 17:24:53 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant (APIServer pid=1) INFO 09-23 17:24:53 [api_server.py:685] Supported tasks: ['generate'] (APIServer pid=1) INFO 09-23 17:24:53 [parser_manager.py:37] "auto" tool choice has been enabled. (APIServer pid=1) INFO 09-23 17:24:53 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this. (APIServer pid=1) INFO 09-23 17:24:54 [base.py:235] Multi-modal warmup completed in 1.322s (APIServer pid=1) INFO 09-23 17:24:55 [base.py:235] Readonly multi-modal warmup completed in 0.987s (APIServer pid=1) WARNING 09-23 17:24:55 [model.py:1723] Default vLLM sampling parameters have been overridden by /root/.cache/huggingface: `{'temperature': 1.0, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. (APIServer pid=1) INFO 09-23 17:24:55 [api_server.py:689] Starting vLLM server on http://0.0.0.0:8002 (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:51] Available routes are: (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /openapi.json, Methods: GET, HEAD (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /docs, Methods: GET, HEAD (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /docs/oauth2-redirect, Methods: GET, HEAD (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /redoc, Methods: GET, HEAD (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /load, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /version, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /health, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /metrics, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /tokenize, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /detokenize, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/models, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /ping, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /ping, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /invocations, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/chat/completions, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/chat/completions/batch, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/responses, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/responses/{response_id}, Methods: GET (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/responses/{response_id}/cancel, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/completions, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/messages, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/messages/count_tokens, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /generative_scoring, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /scale_elastic_ep, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /is_scaling_elastic_ep, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/chat/completions/render, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/completions/render, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/chat/completions/derender, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /v1/completions/derender, Methods: POST (APIServer pid=1) INFO 09-23 17:24:55 [launcher.py:60] Route: /inference/v1/generate, Methods: POST (APIServer pid=1) INFO: 100.64.0.10:58102 - "GET /health HTTP/1.1" 200 OK