[rank0]:[W923 16:57:18.453833811 CUDAEvent.h:57] Warning: CUDA warning: an illegal memory access was encountered (function ~CUDAEvent) [rank0]:[W923 16:57:18.455473925 CachingHostAllocator.cpp:26] Warning: Exception in pinned allocator free(), rethrowing (function free) terminate called after throwing an instance of 'c10::AcceleratorError' what(): CUDA error: an illegal memory access was encountered Search for `cudaErrorIllegalAddress' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information. CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Exception raised from currentStreamCaptureStatusMayInitCtx at /build/pytorch/c10/cuda/CUDAGraphsC10Utils.h:73 (most recent call first): frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string, std::allocator >) + 0xb1 (0x78c2cc456bf1 in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #1: + 0xd86b (0x78c2cc4e386b in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10_cuda.so) frame #2: + 0xf1e58f (0x78c2ae3da58f in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cuda.so) frame #3: + 0xf17b94 (0x78c2ae3d3b94 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cuda.so) frame #4: + 0x9118c (0x78c2cc43118c in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #5: c10::TensorImpl::~TensorImpl() + 0xd (0x78c2cc42b00d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #6: + 0x9a1659 (0x78c2c71ab659 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so) frame #7: + 0x9a16e5 (0x78c2c71ab6e5 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so) frame #8: VLLM::Worker_TP0_DCP0_EP0() [0x59edb3] frame #9: + 0x10d6fd (0x78c243d686fd in /usr/local/lib/python3.12/dist-packages/numpy/_core/_multiarray_umath.cpython-312-x86_64-linux-gnu.so) frame #10: VLLM::Worker_TP0_DCP0_EP0() [0x5793e2] frame #11: VLLM::Worker_TP0_DCP0_EP0() [0x59ec59] frame #12: _PyEval_EvalFrameDefault + 0x8f14 (0x5df394 in VLLM::Worker_TP0_DCP0_EP0) frame #13: VLLM::Worker_TP0_DCP0_EP0() [0x54cad2] frame #14: _PyEval_EvalFrameDefault + 0x4cbe (0x5db13e in VLLM::Worker_TP0_DCP0_EP0) frame #15: VLLM::Worker_TP0_DCP0_EP0() [0x54cad2] frame #16: VLLM::Worker_TP0_DCP0_EP0() [0x6f89dc] frame #17: VLLM::Worker_TP0_DCP0_EP0() [0x6b931c] frame #18: + 0x9caa4 (0x78c2ce4daaa4 in /lib/x86_64-linux-gnu/libc.so.6) frame #19: __clone + 0x44 (0x78c2ce567a64 in /lib/x86_64-linux-gnu/libc.so.6) [rank1]:[W923 16:57:18.483584493 CachingHostAllocator.cpp:26] Warning: Exception in pinned allocator free(), rethrowing (function free) terminate called after throwing an instance of 'c10::AcceleratorError' what(): CUDA error: an illegal memory access was encountered Search for `cudaErrorIllegalAddress' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information. CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Exception raised from currentStreamCaptureStatusMayInitCtx at /build/pytorch/c10/cuda/CUDAGraphsC10Utils.h:73 (most recent call first): frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string, std::allocator >) + 0xb1 (0x7e1ff4746bf1 in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #1: + 0xd86b (0x7e1ff47d386b in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10_cuda.so) frame #2: + 0xf1e58f (0x7e1fd63da58f in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cuda.so) frame #3: + 0xf17b94 (0x7e1fd63d3b94 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cuda.so) frame #4: + 0x9118c (0x7e1ff472118c in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #5: c10::TensorImpl::~TensorImpl() + 0xd (0x7e1ff471b00d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so) frame #6: + 0x9a1659 (0x7e1fef1ab659 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so) frame #7: + 0x9a16e5 (0x7e1fef1ab6e5 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so) frame #8: VLLM::Worker_TP1_DCP1_EP1() [0x59edb3] frame #9: VLLM::Worker_TP1_DCP1_EP1() [0x5758ae] frame #10: VLLM::Worker_TP1_DCP1_EP1() [0x5755fc] frame #11: VLLM::Worker_TP1_DCP1_EP1() [0x575986] frame #12: VLLM::Worker_TP1_DCP1_EP1() [0x5755fc] frame #13: VLLM::Worker_TP1_DCP1_EP1() [0x5793e2] frame #14: VLLM::Worker_TP1_DCP1_EP1() [0x59ec59] frame #15: VLLM::Worker_TP1_DCP1_EP1() [0x5758ae] frame #16: VLLM::Worker_TP1_DCP1_EP1() [0x5755fc] frame #17: VLLM::Worker_TP1_DCP1_EP1() [0x558cd1] frame #18: VLLM::Worker_TP1_DCP1_EP1() [0x6105f5] frame #19: VLLM::Worker_TP1_DCP1_EP1() [0x610605] frame #20: VLLM::Worker_TP1_DCP1_EP1() [0x610605] frame #21: VLLM::Worker_TP1_DCP1_EP1() [0x610605] frame #22: VLLM::Worker_TP1_DCP1_EP1() [0x610605] frame #23: VLLM::Worker_TP1_DCP1_EP1() [0x610605] frame #24: VLLM::Worker_TP1_DCP1_EP1() [0x5532bb] frame #25: _PyEval_EvalFrameDefault + 0x92f9 (0x5df779 in VLLM::Worker_TP1_DCP1_EP1) frame #26: PyEval_EvalCode + 0x15b (0x5d544b in VLLM::Worker_TP1_DCP1_EP1) frame #27: PyRun_StringFlags + 0xd3 (0x608483 in VLLM::Worker_TP1_DCP1_EP1) frame #28: PyRun_SimpleStringFlags + 0x3e (0x6b403e in VLLM::Worker_TP1_DCP1_EP1) frame #29: Py_RunMain + 0x481 (0x6bcd01 in VLLM::Worker_TP1_DCP1_EP1) frame #30: Py_BytesMain + 0x2d (0x6bc71d in VLLM::Worker_TP1_DCP1_EP1) frame #31: + 0x2a1ca (0x7e1ff67631ca in /lib/x86_64-linux-gnu/libc.so.6) frame #32: __libc_start_main + 0x8b (0x7e1ff676328b in /lib/x86_64-linux-gnu/libc.so.6) frame #33: _start + 0x25 (0x657bc5 in VLLM::Worker_TP1_DCP1_EP1) (EngineCore pid=240) Process EngineCore: (EngineCore pid=240) Traceback (most recent call last): (EngineCore pid=240) File "/usr/lib/python3.12/multiprocessing/process.py", line 314, in _bootstrap (EngineCore pid=240) self.run() (EngineCore pid=240) File "/usr/lib/python3.12/multiprocessing/process.py", line 108, in run (EngineCore pid=240) self._target(*self._args, **self._kwargs) (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1387, in run_engine_core (EngineCore pid=240) raise e (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1376, in run_engine_core (EngineCore pid=240) engine_core.run_busy_loop() (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance (EngineCore pid=240) busy_loop_func(self) (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1420, in run_busy_loop (EngineCore pid=240) self._process_engine_step() (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1473, in _process_engine_step (EngineCore pid=240) outputs, model_executed = self.step_fn() (EngineCore pid=240) ^^^^^^^^^^^^^^ (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 729, in step_with_batch_queue (EngineCore pid=240) model_output = future.result() (EngineCore pid=240) ^^^^^^^^^^^^^^^ (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 99, in result (EngineCore pid=240) return super().result() (EngineCore pid=240) ^^^^^^^^^^^^^^^^ (EngineCore pid=240) File "/usr/lib/python3.12/concurrent/futures/_base.py", line 449, in result (EngineCore pid=240) return self.__get_result() (EngineCore pid=240) ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) File "/usr/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result (EngineCore pid=240) raise self._exception (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response (EngineCore pid=240) response = self.aggregate(self.get_response()) (EngineCore pid=240) ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 431, in get_response (EngineCore pid=240) status, result = mq.dequeue(timeout=dequeue_timeout) (EngineCore pid=240) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 889, in dequeue (EngineCore pid=240) with self.acquire_read(timeout, indefinite) as buf: (EngineCore pid=240) File "/usr/lib/python3.12/contextlib.py", line 137, in __enter__ (EngineCore pid=240) return next(self.gen) (EngineCore pid=240) ^^^^^^^^^^^^^^ (EngineCore pid=240) File "/opt/infernal-invocation/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 797, in acquire_read (EngineCore pid=240) raise RuntimeError("cancelled") (EngineCore pid=240) RuntimeError: cancelled (APIServer pid=1) INFO: Shutting down (APIServer pid=1) INFO: Waiting for application shutdown. (APIServer pid=1) INFO: Application shutdown complete. (APIServer pid=1) INFO: Finished server process [1] /usr/lib/python3.12/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib/python3.12/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 4 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' (APIServer pid=1) INFO: 100.64.0.10:35336 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35352 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35356 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35370 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35378 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:47:22 [loggers.py:314] Engine 000: Avg prompt throughput: 3.0 tokens/s, Avg generation throughput: 1.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:35382 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35396 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35404 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35412 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35414 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35424 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:47:32 [loggers.py:314] Engine 000: Avg prompt throughput: 9.0 tokens/s, Avg generation throughput: 20.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:55636 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55646 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55658 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55664 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55678 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:47:42 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:59892 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59902 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59912 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39314 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39328 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39330 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56804 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56820 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56836 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56848 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56862 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56874 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56882 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39138 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39144 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39134 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40666 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40678 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40692 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:41880 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:41894 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:41906 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55388 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55396 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55402 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55414 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55420 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56528 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56534 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:56546 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51086 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51094 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51108 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36436 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36440 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36448 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34298 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34304 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34308 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34322 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34336 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34350 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:34364 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40034 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40054 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40046 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:48816 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:48830 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:48834 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40202 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40196 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40210 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:50808 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:50824 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36504 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36520 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:36524 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:37180 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:37196 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:37202 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:43076 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:43078 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:43090 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54812 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54818 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54826 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:41554 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:41568 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51454 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51458 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51466 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51488 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:51476 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:42618 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:42628 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:42630 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:42642 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:42644 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60824 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60834 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60836 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60842 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60848 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60852 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60866 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60874 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60886 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60902 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:60878 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:53:02 [loggers.py:314] Engine 000: Avg prompt throughput: 204.8 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:60916 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:53:12 [loggers.py:314] Engine 000: Avg prompt throughput: 305.8 tokens/s, Avg generation throughput: 81.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:54334 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54350 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54356 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:53:22 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 74.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:54368 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54370 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:53:32 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 63.4 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:46856 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:46862 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:46874 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:53:42 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 88.5% (APIServer pid=1) INFO: 100.64.0.10:44210 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:44234 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:44226 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59442 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59448 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59452 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59458 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59468 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59470 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59480 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59498 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59494 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:54:12 [loggers.py:314] Engine 000: Avg prompt throughput: 4659.4 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.7%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:59266 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59274 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59276 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:54:22 [loggers.py:314] Engine 000: Avg prompt throughput: 4659.2 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.4%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:59280 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:59290 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:57646 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:57650 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:54:32 [loggers.py:314] Engine 000: Avg prompt throughput: 4659.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.2%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:57664 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:57678 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:57692 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:54:42 [loggers.py:314] Engine 000: Avg prompt throughput: 3501.8 tokens/s, Avg generation throughput: 18.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:54100 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54106 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54114 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:54:52 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 73.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO 09-23 16:55:02 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 73.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:38528 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38544 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38548 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:55:12 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 74.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:44320 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:44324 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:44338 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:55:22 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 73.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:44342 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:44346 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:55:32 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 74.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 86.9% (APIServer pid=1) INFO: 100.64.0.10:35708 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35714 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35724 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:35738 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:55:42 [loggers.py:314] Engine 000: Avg prompt throughput: 2995.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.5%, Prefix cache hit rate: 86.7% (APIServer pid=1) INFO 09-23 16:55:52 [loggers.py:314] Engine 000: Avg prompt throughput: 186.8 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 86.7% (APIServer pid=1) INFO: 100.64.0.10:54606 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54610 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:54624 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:56:02 [loggers.py:314] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 86.7% (APIServer pid=1) INFO: 100.64.0.10:55104 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55106 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55118 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49312 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49318 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49324 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49336 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40342 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40344 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40348 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:40360 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38958 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38962 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38966 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38982 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:38996 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39008 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39024 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39038 - "POST /tokenize HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:39048 - "POST /v1/chat/completions HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:56:42 [loggers.py:314] Engine 000: Avg prompt throughput: 998.4 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 85.2% (APIServer pid=1) INFO 09-23 16:56:52 [loggers.py:314] Engine 000: Avg prompt throughput: 4658.8 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.1%, Prefix cache hit rate: 85.2% (APIServer pid=1) INFO: 100.64.0.10:55740 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55754 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:55748 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:57:02 [loggers.py:314] Engine 000: Avg prompt throughput: 4659.5 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.8%, Prefix cache hit rate: 85.2% (APIServer pid=1) INFO: 100.64.0.10:49724 - "GET /v1/models HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49726 - "GET /health HTTP/1.1" 200 OK (APIServer pid=1) INFO: 100.64.0.10:49730 - "GET /metrics HTTP/1.1" 200 OK (APIServer pid=1) INFO 09-23 16:57:12 [loggers.py:314] Engine 000: Avg prompt throughput: 4659.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.6%, Prefix cache hit rate: 85.2% (APIServer pid=1) INFO: 100.64.0.10:35190 - "POST /v1/chat/completions HTTP/1.1" 200 OK (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] WorkerProc hit an exception. (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] Traceback (most recent call last): (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 1040, in worker_busy_loop (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/worker_base.py", line 359, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.worker.execute_model(scheduler_output) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu_worker.py", line 1175, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = self.model_runner.execute_model( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu/model_runner.py", line 2192, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] model_output = self.model(**model_inputs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/models/glm4_1v.py", line 2327, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states = self.language_model.model( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 717, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states, residual, post, comb = layer( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 478, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] x = self.self_attn( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 615, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.mla_attn(positions, hidden_states) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/mla.py", line 566, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.indexer(hidden_states, q_c, positions, self.indexer_rope_emb) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 397, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.indexer_op( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/custom_op.py", line 136, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._forward_method(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1396, in forward_cuda (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return torch.ops.vllm.sparse_attn_indexer_kpool( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/_ops.py", line 1279, in __call__ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._op(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/compilation/breakable_cudagraph.py", line 102, in wrapper (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return fn(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1183, in sparse_attn_indexer_kpool (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_kpool_dcp_topk( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 344, in _merge_kpool_dcp_topk (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_b12x_dcp_topk( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer.py", line 1210, in _merge_b12x_dcp_topk (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] triton_gather_topk_ids_by_position( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/attention/backends/mla/sparse_utils.py", line 140, in triton_gather_topk_ids_by_position (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _gather_topk_ids_by_position_kernel[grid]( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 370, in (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 761, in run (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] kernel.run(grid_0, grid_1, grid_2, stream, kernel.function, kernel.packed_metadata, launch_metadata, (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/backends/nvidia/driver.py", line 328, in __call__ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.launch(gridX, gridY, gridZ, stream, function, self.launch_cooperative_grid, self.launch_pdl, (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] RuntimeError: Triton Error [CUDA]: an illegal memory access was encountered (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] Traceback (most recent call last): (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 1040, in worker_busy_loop (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/worker_base.py", line 359, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.worker.execute_model(scheduler_output) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu_worker.py", line 1175, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = self.model_runner.execute_model( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu/model_runner.py", line 2192, in execute_model (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] model_output = self.model(**model_inputs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/models/glm4_1v.py", line 2327, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states = self.language_model.model( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 717, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states, residual, post, comb = layer( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 478, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] x = self.self_attn( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 615, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.mla_attn(positions, hidden_states) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/mla.py", line 566, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.indexer(hidden_states, q_c, positions, self.indexer_rope_emb) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 397, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.indexer_op( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/custom_op.py", line 136, in forward (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._forward_method(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1396, in forward_cuda (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return torch.ops.vllm.sparse_attn_indexer_kpool( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/_ops.py", line 1279, in __call__ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._op(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/compilation/breakable_cudagraph.py", line 102, in wrapper (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return fn(*args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1183, in sparse_attn_indexer_kpool (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_kpool_dcp_topk( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 344, in _merge_kpool_dcp_topk (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_b12x_dcp_topk( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer.py", line 1210, in _merge_b12x_dcp_topk (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] triton_gather_topk_ids_by_position( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/attention/backends/mla/sparse_utils.py", line 140, in triton_gather_topk_ids_by_position (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _gather_topk_ids_by_position_kernel[grid]( (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 370, in (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs) (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 761, in run (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] kernel.run(grid_0, grid_1, grid_2, stream, kernel.function, kernel.packed_metadata, launch_metadata, (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/backends/nvidia/driver.py", line 328, in __call__ (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.launch(gridX, gridY, gridZ, stream, function, self.launch_cooperative_grid, self.launch_pdl, (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] RuntimeError: Triton Error [CUDA]: an illegal memory access was encountered (Worker_TP0_DCP0_EP0 pid=262) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:19 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:23 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:29 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:32 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:36 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:42 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:45 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:48 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:22:52 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:11 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:15 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:17 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:24 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:28 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:30 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:34 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:38 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:23:42 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:24 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:28 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:36 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:40 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:49 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:24:53 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:25:12 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:25:16 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:25:42 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:25:46 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:25:58 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:26:02 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:26:14 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:26:18 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:27:39 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None` (Worker_TP1_DCP1_EP1 pid=263) 2026-09-23 12:27:43 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] WorkerProc hit an exception. (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] Traceback (most recent call last): (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 1040, in worker_busy_loop (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/worker_base.py", line 359, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.worker.execute_model(scheduler_output) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu_worker.py", line 1175, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = self.model_runner.execute_model( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu/model_runner.py", line 2192, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] model_output = self.model(**model_inputs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/models/glm4_1v.py", line 2327, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states = self.language_model.model( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 717, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states, residual, post, comb = layer( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 478, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] x = self.self_attn( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 615, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.mla_attn(positions, hidden_states) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/mla.py", line 566, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.indexer(hidden_states, q_c, positions, self.indexer_rope_emb) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 397, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.indexer_op( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/custom_op.py", line 136, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._forward_method(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1396, in forward_cuda (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return torch.ops.vllm.sparse_attn_indexer_kpool( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/_ops.py", line 1279, in __call__ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._op(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/compilation/breakable_cudagraph.py", line 102, in wrapper (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return fn(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1183, in sparse_attn_indexer_kpool (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_kpool_dcp_topk( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 344, in _merge_kpool_dcp_topk (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_b12x_dcp_topk( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer.py", line 1210, in _merge_b12x_dcp_topk (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] triton_gather_topk_ids_by_position( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/attention/backends/mla/sparse_utils.py", line 140, in triton_gather_topk_ids_by_position (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _gather_topk_ids_by_position_kernel[grid]( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 370, in (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 761, in run (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] kernel.run(grid_0, grid_1, grid_2, stream, kernel.function, kernel.packed_metadata, launch_metadata, (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/backends/nvidia/driver.py", line 328, in __call__ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.launch(gridX, gridY, gridZ, stream, function, self.launch_cooperative_grid, self.launch_pdl, (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] RuntimeError: Triton Error [CUDA]: an illegal memory access was encountered (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] Traceback (most recent call last): (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 1040, in worker_busy_loop (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/worker_base.py", line 359, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.worker.execute_model(scheduler_output) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu_worker.py", line 1175, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] output = self.model_runner.execute_model( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return func(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu/model_runner.py", line 2192, in execute_model (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] model_output = self.model(**model_inputs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/models/glm4_1v.py", line 2327, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states = self.language_model.model( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 717, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] hidden_states, residual, post, comb = layer( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/model.py", line 478, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] x = self.self_attn( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 615, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.mla_attn(positions, hidden_states) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/mla.py", line 566, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.indexer(hidden_states, q_c, positions, self.indexer_rope_emb) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/models/glm5next/nvidia/attention.py", line 397, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self.indexer_op( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._call_impl(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return forward_call(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/custom_op.py", line 136, in forward (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._forward_method(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1396, in forward_cuda (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return torch.ops.vllm.sparse_attn_indexer_kpool( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/torch/_ops.py", line 1279, in __call__ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return self._op(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/compilation/breakable_cudagraph.py", line 102, in wrapper (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return fn(*args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 1183, in sparse_attn_indexer_kpool (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_kpool_dcp_topk( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer_kpool.py", line 344, in _merge_kpool_dcp_topk (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _merge_b12x_dcp_topk( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/sparse_attn_indexer.py", line 1210, in _merge_b12x_dcp_topk (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] triton_gather_topk_ids_by_position( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/opt/infernal-invocation/vllm/vllm/v1/attention/backends/mla/sparse_utils.py", line 140, in triton_gather_topk_ids_by_position (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] _gather_topk_ids_by_position_kernel[grid]( (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 370, in (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs) (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/runtime/jit.py", line 761, in run (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] kernel.run(grid_0, grid_1, grid_2, stream, kernel.function, kernel.packed_metadata, launch_metadata, (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] File "/usr/local/lib/python3.12/dist-packages/triton/backends/nvidia/driver.py", line 328, in __call__ (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] self.launch(gridX, gridY, gridZ, stream, function, self.launch_cooperative_grid, self.launch_pdl, (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] RuntimeError: Triton Error [CUDA]: an illegal memory access was encountered (Worker_TP1_DCP1_EP1 pid=263) ERROR 09-23 16:57:18 [multiproc_executor.py:1048] (EngineCore pid=240) ERROR 09-23 16:57:18 [multiproc_executor.py:325] Worker proc VllmWorker-0 died unexpectedly (exit code: None), shutting down executor. (EngineCore pid=240) INFO 09-23 16:57:18 [multiproc_executor.py:470] [shutdown] Executor: waiting for worker exit count=2 (EngineCore pid=240) INFO 09-23 16:57:18 [multiproc_executor.py:477] [shutdown] Executor: all workers exited gracefully (EngineCore pid=240) ERROR 09-23 16:57:18 [dump_input.py:72] Dumping input data for V1 LLM engine (v0.1.dev20111+g7f1e92bec.d20260827) with config: model='/root/.cache/huggingface', speculative_config=None, tokenizer='/root/.cache/huggingface', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=327680, download_dir=None, load_format=safetensors, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=2, dcp_comm_backend=a2a, disable_custom_all_reduce=True, quantization=exl3, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8_ds_mla, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='glm45', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=glm53-flash-exl3-4bpw-v84-fp8-tp2-c4-327k-r10-vision8-apc-nospec, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 8, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='b12x', linear_backend='auto'), (EngineCore pid=240) ERROR 09-23 16:57:18 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[NewRequestData(req_id=chatcmpl-9f4e530b545624ca-abba75e2,prompt_token_ids_len=31816,prefill_token_ids_len=31816,mm_features=[],sampling_params=SamplingParams(n=1, presence_penalty=0.0, frequency_penalty=0.0, repetition_penalty=1.0, temperature=0.0, top_p=1.0, top_k=0, min_p=0.0, seed=None, stop=[], stop_token_ids=[154827, 154829], bad_words=[], thinking_token_budget=None, include_stop_str_in_output=False, ignore_eos=False, max_tokens=1024, min_tokens=0, logprobs=None, prompt_logprobs=None, skip_special_tokens=False, spaces_between_special_tokens=True, structured_outputs=None, extra_args=None),block_ids=([207], [191], [198], [199], [205], [89]),num_computed_tokens=0,lora_request=None,prompt_embeds_shape=None)], scheduled_cached_reqs=CachedRequestData(req_ids=['chatcmpl-aa76274bced356ef-be83474f'],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[None],num_computed_tokens=[174807],num_output_tokens=[10]), num_scheduled_tokens={chatcmpl-9f4e530b545624ca-abba75e2: 2047, chatcmpl-aa76274bced356ef-be83474f: 1}, total_num_scheduled_tokens=2048, scheduled_spec_decode_tokens={}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[0, 0, 0, 0, 0, 0], finished_req_ids=[], free_encoder_mm_hashes=[], scheduled_encoder_input_stats=null, preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, ec_connector_metadata=null, ec_manager_metadata=null, new_block_ids_to_zero=[207, 191], kv_cache_block_copies=null, partial_tail_offloads=null, num_spec_tokens_to_schedule=0) (EngineCore pid=240) ERROR 09-23 16:57:18 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=2, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.14960629921259838, iteration_details=None, prefix_cache_stats=PrefixCacheStats(reset=False, requests=0, queries=0, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None, num_scheduled_prompt_tokens=0) (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] EngineCore encountered a fatal error. (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] Traceback (most recent call last): (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1376, in run_engine_core (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] engine_core.run_busy_loop() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] busy_loop_func(self) (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1420, in run_busy_loop (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] self._process_engine_step() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1473, in _process_engine_step (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] outputs, model_executed = self.step_fn() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 729, in step_with_batch_queue (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] model_output = future.result() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 99, in result (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] return super().result() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/usr/lib/python3.12/concurrent/futures/_base.py", line 449, in result (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] return self.__get_result() (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/usr/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] raise self._exception (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] response = self.aggregate(self.get_response()) (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 431, in get_response (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] status, result = mq.dequeue(timeout=dequeue_timeout) (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 889, in dequeue (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] with self.acquire_read(timeout, indefinite) as buf: (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/usr/lib/python3.12/contextlib.py", line 137, in __enter__ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] return next(self.gen) (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] ^^^^^^^^^^^^^^ (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] File "/opt/infernal-invocation/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 797, in acquire_read (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] raise RuntimeError("cancelled") (EngineCore pid=240) ERROR 09-23 16:57:18 [core.py:1385] RuntimeError: cancelled (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] AsyncLLM output_handler failed. (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] Traceback (most recent call last): (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 688, in output_handler (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] outputs = await engine_core.get_output_async() (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core_client.py", line 1104, in get_output_async (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] raise self._format_exception(outputs) from None (APIServer pid=1) ERROR 09-23 16:57:18 [async_llm.py:732] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause. (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] Error in chat completion stream generator. (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] Traceback (most recent call last): (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/entrypoints/openai/chat_completion/serving.py", line 492, in chat_completion_stream_generator (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] async for res in result_generator: (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 607, in generate (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] out = q.get_nowait() or await q.get() (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] ^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/output_processor.py", line 85, in get (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] raise output (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 688, in output_handler (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] outputs = await engine_core.get_output_async() (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core_client.py", line 1104, in get_output_async (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] raise self._format_exception(outputs) from None (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause. (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] Error in chat completion stream generator. (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] Traceback (most recent call last): (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/entrypoints/openai/chat_completion/serving.py", line 492, in chat_completion_stream_generator (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] async for res in result_generator: (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 607, in generate (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] out = q.get_nowait() or await q.get() (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] ^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/output_processor.py", line 85, in get (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] raise output (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/entrypoints/openai/chat_completion/serving.py", line 492, in chat_completion_stream_generator (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] async for res in result_generator: (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 607, in generate (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] out = q.get_nowait() or await q.get() (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] ^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/output_processor.py", line 85, in get (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] raise output (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/async_llm.py", line 688, in output_handler (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] outputs = await engine_core.get_output_async() (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] File "/opt/infernal-invocation/vllm/vllm/v1/engine/core_client.py", line 1104, in get_output_async (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] raise self._format_exception(outputs) from None (APIServer pid=1) ERROR 09-23 16:57:18 [serving.py:845] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause. (APIServer pid=1) INFO 09-23 16:57:21 [utils.py:615] [shutdown] Process manager: send sigterm to process EngineCore (APIServer pid=1) INFO 09-23 16:57:22 [loggers.py:314] Engine 000: Avg prompt throughput: 2503.9 tokens/s, Avg generation throughput: 1.0 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 15.0%, Prefix cache hit rate: 84.9% (APIServer pid=1) INFO 09-23 16:57:22 [core_client.py:689] [shutdown] MPClient: start timeout=0s (APIServer pid=1) INFO 09-23 16:57:22 [core_client.py:691] [shutdown] MPClient: stopping engine manager (APIServer pid=1) INFO 09-23 16:57:22 [core_client.py:693] [shutdown] MPClient: engine manager stopped (APIServer pid=1) INFO 09-23 16:57:22 [core_client.py:694] [shutdown] MPClient: cleaning up background resources (APIServer pid=1) INFO 09-23 16:57:22 [core_client.py:696] [shutdown] MPClient: complete