(APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] █ █ █▄ ▄█ (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.23.1rc1.dev1327+gf25953cc5 (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] █▄█▀ █ █ █ █ model InternScience/Agents-A1-FP8 (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:344] (APIServer pid=1) INFO 07-29 08:46:08 [api_utils.py:273] non-default args: {'model_tag': 'InternScience/Agents-A1-FP8', 'enable_auto_tool_choice': True, 'tool_call_parser': 'qwen3_coder', 'host': '0.0.0.0', 'port': 30010, 'model': 'InternScience/Agents-A1-FP8', 'revision': '4d7d59380f327b76e73bc71f40e0c589ad0ca1d5', 'max_model_len': 131072, 'served_model_name': ['agents-a1-fp8-text'], 'reasoning_parser': 'qwen3', 'kv_cache_dtype': 'fp8', 'language_model_only': True, 'max_num_seqs': 32} (APIServer pid=1) INFO 07-29 08:46:19 [model.py:623] Resolved architecture: Qwen3_5MoeForConditionalGeneration (APIServer pid=1) INFO 07-29 08:46:19 [model.py:1788] Using max model len 131072 (APIServer pid=1) INFO 07-29 08:46:19 [cache.py:285] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor (APIServer pid=1) INFO 07-29 08:46:19 [scheduler.py:252] Chunked prefill is enabled with max_num_batched_tokens=8192. (APIServer pid=1) INFO 07-29 08:46:19 [vllm.py:1109] Asynchronous scheduling is enabled. (APIServer pid=1) INFO 07-29 08:46:19 [kernel.py:303] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']) (APIServer pid=1) INFO 07-29 08:46:21 [registry.py:134] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. (EngineCore pid=357) INFO 07-29 08:46:28 [core.py:116] Initializing a V1 LLM engine (v0.23.1rc1.dev1327+gf25953cc5) with config: model='InternScience/Agents-A1-FP8', speculative_config=None, tokenizer='InternScience/Agents-A1-FP8', skip_tokenizer_init=False, tokenizer_mode=auto, revision=4d7d59380f327b76e73bc71f40e0c589ad0ca1d5, tokenizer_revision=4d7d59380f327b76e73bc71f40e0c589ad0ca1d5, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=131072, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=compressed-tensors, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=agents-a1-fp8-text, enable_prefix_caching=False, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::qwen_gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [8192], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 64, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto') (EngineCore pid=357) INFO 07-29 08:46:30 [registry.py:134] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. (EngineCore pid=357) INFO 07-29 08:46:30 [parallel_state.py:1612] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://172.17.0.2:47753 backend=nccl (EngineCore pid=357) INFO 07-29 08:46:30 [parallel_state.py:1943] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A (EngineCore pid=357) INFO 07-29 08:46:31 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling. (EngineCore pid=357) INFO 07-29 08:46:31 [gpu_model_runner.py:5250] Starting to load model InternScience/Agents-A1-FP8... (EngineCore pid=357) INFO 07-29 08:46:31 [cuda.py:541] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention (EngineCore pid=357) INFO 07-29 08:46:31 [mm_encoder_attention.py:373] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention. (EngineCore pid=357) INFO 07-29 08:46:31 [qwen_gdn_linear_attn.py:150] Using Triton/FLA GDN prefill kernel (requested=auto, head_k_dim=128). (EngineCore pid=357) INFO 07-29 08:46:31 [__init__.py:635] Selected CutlassFP8ScaledMMLinearKernel for CompressedTensorsW8A8Fp8 (EngineCore pid=357) INFO 07-29 08:46:31 [deep_gemm.py:175] deep_gemm not found in site-packages, trying vendored vllm.third_party.deep_gemm (EngineCore pid=357) INFO 07-29 08:46:31 [deep_gemm.py:202] DeepGEMM PDL enabled on vllm.third_party.deep_gemm. (EngineCore pid=357) INFO 07-29 08:46:31 [deep_gemm.py:120] DeepGEMM E8M0 enabled on current platform. (EngineCore pid=357) INFO 07-29 08:46:31 [fp8.py:405] Using TRITON Fp8 MoE backend out of potential backends: ['AITER', 'FLASHINFER_TRTLLM', 'FLASHINFER_CUTLASS', 'DEEPGEMM', 'VLLM_CUTLASS', 'TRITON', 'MARLIN', 'HUMMING', 'BATCHED_DEEPGEMM', 'BATCHED_VLLM_CUTLASS', 'BATCHED_TRITON', 'XPU', 'CPU', 'HPC']. (EngineCore pid=357) INFO 07-29 08:46:31 [cuda.py:482] Using FLASHINFER attention backend out of potential backends: ['FLASHINFER', 'TRITON_ATTN']. (EngineCore pid=357) INFO 07-29 08:46:33 [weight_utils.py:574] No model.safetensors.index.json found in remote. (EngineCore pid=357) INFO 07-29 08:46:33 [weight_utils.py:869] Filesystem type for checkpoints: EXT4. Checkpoint size: 35.09 GiB. Available RAM: 57.45 GiB. (EngineCore pid=357) INFO 07-29 08:46:33 [weight_utils.py:892] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch. (EngineCore pid=357) INFO 07-29 08:46:46 [default_loader.py:430] Loading weights took 13.34 seconds (EngineCore pid=357) INFO 07-29 08:46:46 [fp8.py:688] Using MoEPrepareAndFinalizeNoDPEPModular (EngineCore pid=357) INFO 07-29 08:46:46 [gpu_model_runner.py:5347] Model loading took 34.46 GiB memory and 14.805502 seconds (EngineCore pid=357) INFO 07-29 08:46:46 [interface.py:905] Setting attention block size to 2096 tokens to ensure that attention page size is >= mamba page size. (EngineCore pid=357) INFO 07-29 08:46:51 [backends.py:1094] Using cache directory: /root/.cache/vllm/torch_compile_cache/52f204f062/rank_0_0/backbone for vLLM's torch.compile (EngineCore pid=357) INFO 07-29 08:46:51 [backends.py:1155] Dynamo bytecode transform time: 4.68 s (EngineCore pid=357) INFO 07-29 08:46:53 [backends.py:378] Cache the graph of compile range (1, 8192) for later use (EngineCore pid=357) INFO 07-29 08:47:07 [backends.py:393] Compiling a graph for compile range (1, 8192) takes 15.16 s (EngineCore pid=357) INFO 07-29 08:47:08 [decorators.py:708] saved AOT compiled function to /root/.cache/vllm/torch_compile_cache/torch_aot_compile/b818113cf7a34d4be352cecd08815cd6d6188a1491ff1fbc6825789300909790/rank_0_0/model (EngineCore pid=357) INFO 07-29 08:47:08 [monitor.py:53] torch.compile took 21.73 s in total (EngineCore pid=357) INFO 07-29 08:47:45 [fused_moe.py:1094] Using configuration from /opt/anvil/kernel-tunes/E=256,N=512,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Max-Q_Workstation_Edition,dtype=fp8_w8a8.json for MoE layer. (EngineCore pid=357) INFO 07-29 08:47:48 [monitor.py:81] Initial profiling/warmup run took 40.08 s (EngineCore pid=357) INFO 07-29 08:47:53 [flashinfer.py:822] FlashInfer resolved query dtypes: prefill=torch.bfloat16, decode=torch.bfloat16, decode_backend=flashinfer-native, kv_cache_dtype=torch.float8_e4m3fn, arch=sm120 (EngineCore pid=357) INFO 07-29 08:47:53 [gpu_model_runner.py:6612] Profiling CUDA graph memory: PIECEWISE=11 (largest=64), FULL=7 (largest=32) (EngineCore pid=357) INFO 07-29 08:47:57 [gpu_model_runner.py:6737] Estimated CUDA graph memory: 0.76 GiB total (EngineCore pid=357) INFO 07-29 08:47:57 [gpu_worker.py:548] Available KV cache memory: 50.81 GiB (EngineCore pid=357) INFO 07-29 08:47:57 [gpu_worker.py:563] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9200 is equivalent to --gpu-memory-utilization=0.9121 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.9279. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0. (EngineCore pid=357) INFO 07-29 08:47:57 [kv_cache_utils.py:2181] GPU KV cache size: 5,046,272 tokens (EngineCore pid=357) INFO 07-29 08:47:57 [kv_cache_utils.py:2182] Maximum concurrency for 131,072 tokens per request: 38.50x (EngineCore pid=357) INFO 07-29 08:47:58 [qwen_triton_warmup.py:370] Warming up Qwen Triton kernels for model_type=qwen3_5_moe_text. (EngineCore pid=357) INFO 07-29 08:47:59 [kernel_warmup.py:236] Using FlashInfer autotune cache file: /root/.cache/vllm/flashinfer_autotune_cache/0.6.14/120f/e047a85e3d81b9e867471ef9916cb57cd48c2ac02cd3398844b217df41182633/autotune_configs.json (EngineCore pid=357) WARNING 07-29 08:47:59 [kernel_warmup.py:267] No FlashInfer autotune cache entries found.Falling back to default tactics. (EngineCore pid=357) INFO 07-29 08:47:59 [kernel_warmup.py:68] Warming up ll_bf16 router GEMM kernels. (EngineCore pid=357) INFO 07-29 08:48:05 [cutedsl_warmup.py:105] Skipping CuTeDSL warmup because no compile units were requested. (EngineCore pid=357) INFO 07-29 08:48:05 [gpu_model_runner.py:6798] Rank 0: Torch profiler disabled for CUDA graph capture (EngineCore pid=357) INFO 07-29 08:48:10 [gpu_model_runner.py:6844] Graph capturing finished in 5 secs, took 0.73 GiB (EngineCore pid=357) INFO 07-29 08:48:10 [gpu_worker.py:781] CUDA graph pool memory: 0.73 GiB (actual), 0.76 GiB (estimated), difference: 0.03 GiB (3.5%). (EngineCore pid=357) INFO 07-29 08:48:10 [gpu_worker.py:844] Free memory on device (93.87/95.59 GiB) on startup. Desired GPU memory utilization is (0.92, 87.94 GiB). Actual usage is 35.72 GiB for consumed memory (weights + non-torch), 1.42 GiB for peak activation, and 0.73 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory=53608411956` (49.93 GiB) to fit into requested memory, or `--kv-cache-memory=59973648384` (55.85 GiB) to fully utilize gpu memory. Current kv cache memory in use is 50.81 GiB. (EngineCore pid=357) INFO 07-29 08:48:11 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn. (EngineCore pid=357) INFO 07-29 08:48:11 [core.py:340] init engine (profile, create kv cache, warmup model) took 84.51 s (compilation: 21.73 s) (EngineCore pid=357) INFO 07-29 08:48:11 [vllm.py:1109] Asynchronous scheduling is enabled. (EngineCore pid=357) INFO 07-29 08:48:11 [kernel.py:303] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']) (APIServer pid=1) INFO 07-29 08:48:11 [api_server.py:673] Supported tasks: ['generate'] (APIServer pid=1) INFO 07-29 08:48:11 [parser_manager.py:37] "auto" tool choice has been enabled. (APIServer pid=1) INFO 07-29 08:48:15 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this. (APIServer pid=1) INFO 07-29 08:48:15 [api_server.py:677] Starting vLLM server on http://0.0.0.0:30010 (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:37] Available routes are: (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /docs, Methods: GET, HEAD (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /redoc, Methods: GET, HEAD (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /load, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /version, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /health, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /metrics, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /tokenize, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /detokenize, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/models, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /ping, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /ping, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /invocations, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/chat/completions, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/responses, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/completions, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/messages, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /generative_scoring, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/completions/render, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /v1/completions/derender, Methods: POST (APIServer pid=1) INFO 07-29 08:48:15 [launcher.py:46] Route: /inference/v1/generate, Methods: POST (APIServer pid=1) INFO: 172.17.0.1:57314 - "GET /health HTTP/1.1" 200 OK (EngineCore pid=357) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00