31dda3f557cab4f4a6bcbf2a3d982cb2591b68e9 [Model]Add support for qwen3_vl and qwen3_vl_moe (#3103)
07f4710216da1d61c521fb3bdf4ba90ea1794474 [BugFix] Fix dummy_run memory explosion in eager mode (#3132)
2a9d02e08039749cf811c5fb190d4c8a950d792d [Bugfix] eagle and eagle3 spec decode failures and enable e2e test (#2979)
ac1c2cd9ac6f7ed6ba7505065959393717e9b903 [CI] Upgrade vllm version - 0925 (#3167)
33c118c80e70cec64c9369b7ba4088c61c44bd31 [core]vllm-ascend support msMonitor tool (#3123)
c814b32b90591f36cb2ffc9b513c6ffe3f45b0d7 [Quant][GLM] Adapt glm quant. (#3147)
a055183821120f66b44b514b1795c51ec142df2e [CI] Upgrade vLLM version (#3139)
464270e4caf5023398c415405b180bd53545fbc2 Remove useless PD check in deepseek (#3161)
4ee58e213b421ece745dba6e94967e6f557263ce [BugFix] explicitly setting the tensor shape of otp output (#3027)
cd1ffbb6cd88a3f265027424bd3cca74d1efb1ea [1/N][Feat] Cut down memory usage for o_proj in DeepSeek (#2931)
80524f571152e978f6e2808e504e1aa8b246a2c1 [CORE] concurrent partial prefills (#2372)
2d885869c558734dba9ffb42ea8e7c090bb04c22 [KVCache][Bugfix] Fix kv cache initialization error of attention layer (#3113)
6aa425379867b7c846938416f3dd8b27c66c7b28 [Refactor] [SP]The sequence parallelism characteristics in the MoE and Dense models are integrated into a single solution. (#3085)
e7618d94147f4749c1ea85d1264c7c1d00437008 [2/N][Refactor][Qwen3-Next] remove redundant methods and patch methods in Qwen3NextGatedDeltaNet (#3082)
eb205d9f357848092946fe9d98dabde4bc3ee709 [P/D][BugFix]Mooncake timeout release bug fix (#2899)
6995a7bc5be41908da9f64c2b8def298adf3b058 [Disagg][Perf] Use NPU event sync instead of blocking tolist to avoid unintentional copy ops blocking across different NPU streams, improving disagg TTIT/TTFT (#2788)
c4b976af1a6459a82f7556c53dc98c850dd3e3cd [Model][VLM][Patch]Modify ascend affinity _merge_multimodal_embeddings (#3071)
d01fd1d1c3526ef038f05edc2464a1e77dfb200d [misc][torchair] fix bugs around `deepseek mtp`, `enable_shared_expert_dp` and `use_cached_kv_cache_bytes`  (#3074)
0f3939e5a9b3ce01602c5d7fc269dce08fb720b0 [Feature]cpu offload connector (#1659)
39a85c49faaa67bb476a08cd57cbde36146dab1f [Refactor] Rename cudagraph_support to aclgraph_support (#3104)
29c173ab48001f3d26805db2d833be712687fb1a FlashLB algorithm (#3042)
8dd53c8860ff1c7ec42687800f8d83216d1c7d6a [Bugfix][PD] Auto-clear producer KV cache if no pull notification (#2174)
704467cd9ae3a5077637564113f00af2aad01161 [Bugfix][LoRA] Fix bug introduced by upstream vllm#25249 (#3095)
3fa7cf6345c4e65a767d193a132b51eadfcbe70b [Refactor][Graph] Move graph parameter logic to acl_graph module (#3101)
02f89d166f3624235c47eabd455960c0d498c6d4 [CI] Update vllm version to 20250922(5aeb925) (#3091)
1c9f0fe26f1e64f77c85f804d7c98df039b585d3 Fix of DeepSeek Error in KV Pool Mixed Deployment Scenario (#3087)
37a0715edade3efae975c02ec8edf6f7a5d07530 [Refactor] Adjustments to moe_comm_method selection process (#3001)
bb1f0d5a62046c90eab24f037e48164773c33dc5 [main] remove the redundant log prints in register_custom_ops.py (#3094)
338231acaf946c7478d677b371bb7f75e259eab9 [Feat][Graph] Support `FULL_DECODE_ONLY` mode for GQA/MHA models (#2128)
f39bd309b65ed7498a372e2a55d39611b7290744 [Hybrid KV] Follow up UniformTypeKVCacheSpecs (#3070)
f1f2c8f5e5d1659a97edc5cde7b7d80b8426416a [Perf] Add new npu_fused_infer_attention_score op to improve perfomance in splitfuse cases and resolve long-seq mask problems (#2962)
ffdd1a36e20180e881036a6d502dac1810ba085e [bugfix][torchair] fix wasted NPU memory buffer allocation for quantized deepseek with unquantized MTP layer (#3068)
14b39d3c700c543e0f47bd5ccf1fb72ccae98c71 [1/N][Refactor][Qwen3-Next] remove redundant Qwen3NextSparseMoeBlock and Qwen3NextAttention (#3019)
b8b68b3dfe415f0f10a5e9407b5b60ddbf6b9c9d [CI] Upgrade vLLM to 20250920 (c60e613) and address config break (#3067)
12bcbd02bb52e1a47f9c93b49af1a24eb32d175a [CI] Upgrade vLLM to 20250919 (6d8246aa) and fix some broken issue (#2907)
53ecd89e8ff405302be040a76effa8c012cbaaeb [Bugfix] Remove `VLLM_TEST_DYNAMO_FULLGRAPH_CAPTURE` (#2969)
a22b532d38fe6f1cda1a2e61d1f9af8888cac256 [Fixbug] Fix shape not match when sliding_window and dynamic batch_size (#2830)
0942d9aaabb6344b542f43147add9077e9c49c3c [3/N][Refactor][Quantization]remove packed_modules_mapping from models (#3021)
8326f15ecfb4e0139c8d11ee928742d91fc89ee9 [CustomOp] Register AscendSharedFusedMoE custom op (#2980)
05a700d370a21e07d4c0fb7978debb3a490751a3 [Bugfix] Fix async copy bug under single expert scenario (#3005)
2a87b4cecbe94208aba8d12cd3b32eb6913c0f48 [Bugfix] Fix specdecoding in chunkedprefill scenario (#3025)
833cd1b698f3d467bb0a6a60cbf20ebc5535f5c9 [BugFix] Async scheduling and PP compatibility with DP (#2796)
0a526768f55cc60b5b870adce92affdfcfea8523 [Feature] Support moe multi-stream for aclgraph. (#2946)
367edff5af202e378ad088090439d6389b40fa5d [HybridKV] Fix prefill disaggregation kvcache addr alignment & use hybrid kv cache only when running qwen3_next (#3007)
01592515b8265db7f2da22e0983e1f5ef6eead4c [Bugfix] Fix sleep mode level 2 (#1376)
f4e3d224320877d5e9e6de9306fb5671230e214a Remove chunked_prefill_for_mla and fix ring_mla bug (#2781)
79a910ef4730d3f1be14496a1681eee2566f64a0 [bugfix][torchair] fix multistream_moe problems in torchair graph mode (#2681)
af2a886814dbc1ce512ed9997c3e0b2bb3dcf01a refactor linear (#2867)
a7f8ed38ed0681a0c3e29d848b04db4c7e972e06 [Bugfix]:replace npu_incre_flash_attention with npu_fused_infer_atten… (#2901)
6681dde9028b2ae8ab0e56158a117c5db7f0bc7a [Feat][Graph] Support MTP for ACL Graph (#2932)
cef43b524e5dbf24434ac330235c5c835284c580 [Feat] A Connector that supports Mooncake store (#2913)
723d460894edeb480d2a19fa128bb1a89aca0a78 [Bugfix] fix kv nz accuracy bug (#2988)
8bcc0ccd571a001bcf1f428aceb2445ba0375fac [bugfix] fix shared expert dp with hybrid kvcache (#2964)
1f6465c399d6e699f88e28419c106573bd6c44f0 Add an option of enable frozen parameter (#2869)
76844eec78a23f482a4e0dfe9684898a6ef35fb2 Dynamic Expert Load Balance with Zero-like-overhead (#2956)
ae758dda05b57adaf70af44122e6bdc54fbb88ff [Bugfix] Fix mtp torchair in pd Disaggregation scenario (#2951)
6b7117dbb74e6c46da110d74014af811be6323ec [main] addrmsnorm + quant fusion optim in Dense Models (#2772)
88ca8a051ca51fe72516344db092a7852150cfdb [Feat][Graph] Support DeepSeek with ACL Graph (#2707)
1c5900327b67015e5707d3879ccf5fa5ab622832 [refactor] refactor deepseek-related files (#2849)
18ca7861f6e4f27cea58cc1d70a8c3081422081c [Main] [Refactor] Enable MoECommMethod in Eager Mode (#2791)
c556038ef0b8580cb9079823bd9f06263ee7d731 [New model] Qwen3-next support (#2917)
382c29f3e1a3201c5bedd588f50fbe55dad2d919 [BugFix] Fix world size bug in model_runner (#2915)
c5a502fd2e81e6dfc0cbafbc4ffa49bb68c3abf1  main add ascend scheduler support multimodal (#2844)
0a27705917e64993a8a76198ac6e30980578fe60 fix mooncake connector adxl hostname usage (#2824)
e57cca971c4db11d8ce9da6b008bd655ada8c77e Fix the bugs about operator registration by PyTorch Dispatcher (#2786)
585a494baa4bdbce5a71ef6808033466ed9f90f3 [Core] Disable the chunked prefill feature in Non-MLA LLMs (#2894)
756b8a1946aa9396d5bc7b9c67547fcb93fad630 Revert "[Feat] Unquantized linear nz support (#2619)" (#2896)
fc2bcbe21c86f7684c80e42771b128da9fc17571 [Ops] Fix bug in register_custom_ops without forward_context (#2883)
778cb7255697ad0d1562f60d0cf4ef68542a5b94 fix bug when rotary_dim is not 128 (#2847)
f5a97e8fa5440df6735d1121f813cda7f1257367 [Quantization] register AscendQuantRMSNorm for quantization (#2856)
eab3635850ba351af81d76a7b4b3db46ffb7f697 [Bugfix] Retrieve num_redundant_experts from eplb_config in torchair qwen3_moe.py (#2857)
aeffe27b3089cf12b743806fd922c7d1fd455ac2 [Perf]set moe w2_weight default to be nz (#2842)
9615dea3a71df8ecd2c591f284d9615140dce68a Refactor tensor_parallel and comm_utils (#2814)
0005479b9c3ebf262d379a454641d407a8a1dba6 [main] mlp weight prefetch in Qwen Dense Models (#2816)
c3c222150363900e5f0e87ae2abb4868dacf6a1c [Feat]support dynamic quantization in allgather (#2841)
bd3dedea6123c9c8c19fe83b6e05716f63b1285d support qwen25 vl w8a8 quantization (#2778)
2b9269b581e390357089d54fb8934ea65bf0ced4 [Perf][V1] Fully overlap model execution (#2783)
923cdaeba389ea5b895acb25a7c755ccb830cee2 fix ascend fused moe spelling error (#2863)
b9a0a75c783571caf22129612fb3338272d1782c fix qwen torchair attention PrefillCacheHit (#2787)
7b2ecc1e9a64aeda78e2137aa06abdbf2890c000 [Feat] Unquantized linear nz support (#2619)
5691104249bbee7648e8cfc1466a96c092a8d76d LLMdatadist connector adapt the distributed KV aggregation (#2718)
c2fdd4b8bc9ce859343909caebf331e0cd047908 [CI/UT] Fix UTs on register customop and warm up model (#2862)
b7df04de9ba1bade98ebd77f86f0853c98dd4855 debug_aclgraph_sizes_capture (#2827)
88d7af62be19408af6014575086fadcf9e7a5a00 [main] adjust the position of warm_up_atb (#2823)
22b425765a36ab507ca23625409c767dabb4371b [Bugfix] Fix broken CI (#2825)
aa4d2a91ed6450759895ee7fd614eaf336c4722e Refactor AscendMultiHeadLatentAttention (#2826)
168ad600b5d794fef4314980ddeac9f71511c449 [main] add pd transfer for ascend scheduler (#2753)
edf1f600ad30120a7d870a38b65520616f3131ae [CI] Remove compatibility maintenance for vllm v0.10.1 and v0.10.1.1 (#2840)
93e28e6862669e3b5cf47cea9f782a65ec47e155 add weight transpose check. (#2756)
e13c4ddb4290112e5869a26f97fbfd02075fbc67 [Fix] Fix SharedFusedMoE (#2817)
7a205dbaa8feb1a8927d58b6399c0766d395c78b [main] Optimize rope in Qwen Models (#2571)
1bbb20ea13e2b9a47936abebdcfb6143fdce8079 [main] flashcomm_v1 optim in Qwen Dense Models (#2802)
4df8df5b945da2eb29f95f46ad85836b8c71bf62 [bugfix] fix deepseek rope sincoscache re-generation (#2744)
7d6d9449a83d661c229b93c17a998e67d01bafa7 [Misc] Move lora patch file into lora module (#2797)
85d989a3b93af7f566c2c52210c3576a665c5e66 [Misc] Remove pangu model file (#2798)
a041d4f32852289a28ec7c72f5f2d8c3ac142457 [main] [refactor] refactor common_fused_moe.py (#2706)
1a82b16355d2ec0ba01c23935092dc0af323b820 Remove unused code in fused_moe.py (#2805)
d51694a77bbbeba6b45a54ecb4cb04559266b129 [2/N][Refactor][Quantization] clean quantization patch (#2785)
d3c3538ddc67ba8f4873637e2bc1052f9eb09e93 [Bugfix]fix bug when graph_size is not divisible by tp_size (#2719)
dd087effccad051720fbd26f798b0c68bfbf70d7 Refector prepare_inputs in model_runner_v1.py (#2750)
c735bb09419beb0fa9a186ce06cc9ddf2c3cc50b [Fix] Ensure metadata sync across DP ranks in eager mode (#2766)
2693196ef8382761ae7858b9fcafe3866bf7287a add gatherep select. (#2740)
6666e5265d40ecafc3cb377233fee840d7fe553b Added support for KV connector v1 (#2039)
b2f77d3aa8f384488b56f5787928e84b1f347999 [fix] prefill unsupport sliding window attention (#2758)
5a7181569c58630fed55e1f78eec615dc2034b74 [feat]: oproj tensor parallelism in  pure DP and graph-mode scenarios. (#2167)
51a2aec1154e5e2ab9cbbba53d830d8d21ab3571 Delete redundant codes related to communication (#2717)
5b3646ab2142131579661ce12e4f0e4ba731ad06 [FEATURE][MTP] Support MTP > 1 (#2708)
83eb40a51cb30c654c80a9203b3d9d7bd0351e4e [Fix][MoE] Refine MoE communication strategy (#2734)
4c90fa79ca76211b9a2a2cfcaad71aa24d13d997 [Misc] Remove useless PD check in deepseek (#2739)
f86596a66cc0aff2b05280303212758380f0ec9a allgather use fusedop. (#2689)
7d47d8f4f61632b52a35068eadc91d1934e0b71b [Fix] fix resources limit error when apply speculative decoding and aclgraph (#2472)
0c0789be7442122eb1203abbf89a9592648922e0 [Feat] allow using aclgraph in ray backend (#2589)
aff5189c8781f1127d7dbca6c59a9ec44a427100 [main] Fuse GroupedMatmul, Swiglu and DynamicQuant in `W8A8_DYNAMIC` quantized MoE layers (#2275)
37f5a29cd4f84fa0beee236dadf070b41e4b5403 [1/N][Refactor][Quantization] remove redundant quantizer class (#2680)
d4370ebc42f8a2cecbb7ad4b199ac3f840ca3b28 [Refactor] Refactor Spec Decode (#2668)
e7409e95ee73fb3bb7bf8b23f26c16620ed94543 [1/N][Draft][Refactor]torchair pangu_moe modeling refactor (#2437)
a58013440a9c9c0b5220a60bc161025c5f5270a2 [BugFix][MLA] Fix attn_mask bug for ring mla (#2704)
984bd7c13a6b7eb80ac9cb43ab85a81afe779614 [Bugfix][APC] Fix accuracy issue on prefix caching with AscendScheduler (#2714)
df88a2ecc8116a42d79a13fa1a8a05a03c70324f [P/D]mooncake_connector adapted to 0.10.1 (#2664)
07d44ade194b018ae2cc172482d55cb746c5fd0e bugfix: fix initialization error for mooncake in k8s (#2541)
90a75a90a9adc8be8efa294bada8391bb1607e0d [bugfix] fix torchair runtime error caused by configuration mismtaches and file missing (#2532)
5889fa1b1cddb283b5e2c206d08062bef8432cbd [bugfix] ascend schedule encountered an incorrect req block length in the check_watermark_for_prefill function (#2508)
3584306387bc5d094700e71c75fbc9b5154bdaf7 [Bugfix] Fix qwen2.5-vl-without-padding (#2623)
eaeb2efb20ff70875991483d63262f379a3afde8 [Main][Feat]Set the Profiler parameters through environment variables consistent with vLLM (#2608)
93754d80616830a5bc068c51d3493b84f679750d [Bugfix] Fix long context seq accuracy problem for `GLM4.5` (#2601)
b84465c52564134d217a0119e34a68ef86b6d35d [Perf]Enable npu_moe_gating_top_k_softmax on quantized scenarios (#2633)
c1e607b7b71c7b93d2ac3b04b0cc3971da896473 [Misc] Clean up uesless code in rotary_embedding (#2663)
253b01b9a552b9745aaea827bd8944c87ea7b09a [7/N][refactor]fix torchair rope ops (#2683)
9f1e054fe3df966a4fa51cc73e34c430452b5ebc [Bugfix][LoRA][Operator] Fix LoRA custom operators accuracy issue (#2672)
214b32a3460809148271d66774a457e3c750d79e [V1][BUGFIX][0.10.1] FIX mtp on main branch (#2632)
0df059f41a4c45ac98b0b6be9d5b50f989893905 [CI] Fix CI Break: upstream adds routed_scaling_factor in forward_oot interface (#2675)
ea53f9076e722eb669d9df76ed6601d807acae7e support torchair mode (#2641)
ad13964c7121d7d80813c6f79a0b5fce9b6f66b0 [6/N][refactor]delete torchair in rotary ops (#2581)
c2c97f3079957efafbfba29ba97ded417e23fd55 [5/N][refactor]add torchair rotary ops (#2559)
3a5fc5ee01edb3e6c88774edfea2a85eed1ff990 [Refactor][MoE] remove redundant code after refactoring fused_moe (#2612)
20ae71291d876a8511eb504601f747a60864861d [torchair]remove aicpu op (#2640)
7215454de6df78f4f9a49a99c5739f8bb360f5bc bugfix for torchair graph (#2639)
d3c93fba5ca9279ec9f0ebb9026f90abf6132f38 [3/N][Feat][Graph] Support `all-to-all` and quantized models with ACL Graph (#2614)
91c35d765aa2edeb3e9c805f2fe3c330320fe696 [Bugfix] Fix mc2 operator error in aclgraph + ep<16 scenario (#2609)
52aff9e229b8c5557bb2ee485ec381e0886d3701 [main] [bugfix] Fix misjudging quantized/unquantized scenarios (#2627)
aadc75c247924ab8e90c3d82f0ccabcc48cf90ab [Fix] Resolve data-parallel (DP) assertion errors in TorchAir (#2626)
600b08f7542be3409c2c70927c91471e8de33d03 [Feat]: Add custom lmhead tensor model parallel (#2309)
dfc7eb39ada3f86f5c15425ba759ecfaa8f5c9a8 [Fix] Fix DP-related padding logic (#2582)
175f6bc445173704ad8f2a747a08866a205c4b39 Support v0.10.1 (#2584)
6c973361fc2eba5d3faa9b6b496b4b9fec4dc784 [Bugfix] Fix aclgraph not enabled by default (#2590)
cf96366a396b7d70b390cb244b53c58e9666c4d5 [Bugfix][LoRA][Patch] Fix the LoRA inference bug after upstream vLLM codebase changed (#2560)
1191a64ae508183d5613711bc98a90250963f83a [Feat]attention add sliding windows size (#2528)
c8d1df3a3fa803a8a0742df80cf7601a8b15ef7a [Refactor][WIP] Refactor mla_v1 by moving all MLA preprocessing ops into mla_v1 attention impl (#2465)
320edde2df14a16436d7094012c12592b2e16266 [main] [refactor] refactor fused_moe.py to enable token_dispatchers (#2570)
936c102105b72a4e36dd284f900f32338a232696 [bugfix][refactor]fix torchair_w8a8 (#2569)
a955e5d4046e4e3a55976e88774073b3c56b463b [4/N][refactor]delete torchair from quantization (#2535)
c578f817ca4c17a076ac7fa93de77db11f008fae [CustomOp] Register VocabParallelEmbedding instead of overwrite forward (#2515)
2bfbf9b9b3f6cd332fc438fc322b00b6053a043c [main][bugfix] Fix bugs and refactor cached mask generation logic (#2442)
6881c194580b800cb233632300575b9b387778ef [main] convert the format of gmm to nz (#2474)
20a7bc4b71827a39c21ecac95536183468ad90a7 [3/N][refactor] refactoer quantization (#2504)
acdc53c2f6b480f23f40ea9e72356182386e0bac [Bugfix] Fix the bug of cos invalid shape when dp (#2558)
a9e78a329988c4fbe5382ef03b08d6413623d248 [Aclgraph] Update compilation config in `check_and_update_config` (#2540)
f22077daa6a32e1d5c5cfe0e84da2cea1ab8cafb [Embedding] Recover embedding function (#2483)
6a4ec186e731b9516235f4fd30b5b98227513fe7 [Qwen-moe] Remove the minor operation arange (#2373)
358ba6899401e5a1f5e8860e8ea88900cc27dd38 [main][bugfix] Fix MatmulNZ format bug on some machines (#2549)
a6bb502e70b7554b2a0342565348b1a191cd0aa0 [2/N][Feat] Add MC2 communication method for MoE layers (#2469)
5d8ec280090b4a7567fb2b50a7cedda44902c37f [2/N][refactor] split torchair from fused_moe  (#2503)
cfe77e83aeda343274c0488b93e2263bee44a860 [Bugfix]Support Qwen3-MOE on aclgraph mode in sizes capture and add new ut (#2511)
b3fdd78a6b6f8fe9546a1c1092926293577c1b50 [Main][Refactor]Change ASCEND_QUATIZATION_METHOD to ASCEND_QUANTIZATION_METHOD (#2517)
7e494e94a969626d07d2414cb1fb859d2206e3b8 [CI] Fix broken ci (#2530)
99bf25af76a9f6759e262e351c7d41fed57f159f [Fix] Add operations in `_dummy_run` to maintain synchronization with `_process_reqs`, resolving a service hang (#2454)
de7649492ddcbdb7c818665f0b81cc8fbaaaa4b7 [Refactor] cleanup converting_weight_acl_format_format (#2482)
0f81e032f04b72f4dd0c7fefd62b7220942c545a [1/N][refactor] torchair fused_moe refactor (#2438)
f796e6280b79cc87451ceac2daa086ee80b1d572 [CustomOp] Register RotaryEmbedding instead of overwrite forward (#2385)
950c4b219a8fe5d4339de3acd535b32b99e790ad [main] refactor alltoallv in fused_moe (#2487)
4af5b80606e6cffe440c27237655cd44c2e5bdaf [Scheduler] validate max_num_batched_tokens and max_model_len in AscendSchedulerConfig (#2434)
3629bc4431d3edb4224761f9036b3bddb16158d6 feat: add mtp ut and fix some bugs (#2453)
dd04a96ee3caa8c85fbc72a6328b969d1c373bc9 [Bugfix] Fix the bug of incorrect precision (#2479)
b0403f8d8a5ced0ddb86722754bc515f4df9d1f1 [CI] fix ci (#2464)
0ca3f48c900b333673830e8307c259acc684c1a3 [2/N][refactor] torchair deepseek mla backend refactor (#2459)
3fb80ee356781752ed94ea7a39953bdd47eac764 add mlp tp optimze (#2120)
0dca4c6dbdbe4b77116fea6219f6ca494f63e9d2 refact runner model v1 (#2461)
1de16ead8eecfec8903ec1b330b27a4fa2593c35 [main][bugfix] Modify the default value of the enable_shared_pert_dp to false (#2457)
c40d4171bcc0424ad88fc8c9bdd6f694170a8342 [main][quantization] Adapt to the new format of ds w4a8 weight (#2392)
3f867ee7081f6f041652180ccf06d0c70fd44429 refactor allgather/mc2-related fused_experts (#2369)
73acdcfc3bb56363a7ae52cd2af0d0cf84f55592 [PD] Correct the ip and port env (#2450)
7bec1a9b9c372785551d45682bf11063ec42b216 qwen3_moe/qwen25 support torchair graph (#2403)
31ae2497425cf26fd8aaedd9845b6066cd06fb84 [misc] remove uesless envs (#2448)
1327f9be1cea85b2750ac4145981ddd732064bb9 Fix some ci issue and refactor modelrunner (#2445)
d91c6daf891158d11113767bd85c0fe1eae2cde9 [improve] Remove redundant parentheses in pangu_moe.py (#2081)
83e0f41408fb92b7384bed8fdd1239fa6cf18b0b [3/N][Refactor] Move `torchair_attention` to `torchair` dir (#2017)
3f4a358b140226e5c6d218742ff10a210cda5800 [Bugfix] Fix custom op register issue (#2409)
3648d18e673f15a33a82d6ea95d3a9dd891ff1f5 Add Custom Kernels For LoRA Performance (#2325)
3fc31ee1cbdf0c0d11efc4da5fd865bb2d077c4b [1/N][refactor] torchair deepseek modeling  refactor (#2384)
03ca2b26ca9ab6b9a12f021b0595a726ee35e223 [P/D] Mooncake Connector for v1 distributed (#1568)
2bb7e55022c3a558145a1b17ba3c93b4ab6bf00f [Bugfix][PD]fix non-working disaggregated prefill (#2374)
1b40665548f048c3417b79207bb6f1a930475624 [Misc] remove unused file (cache.py) (#2377)
61866b8ac6e8812a205a45dd1f4baee8919df028 [Quickfix] update CachedRequestState as NewRequestData changed (#2367)
c721ae6042c67c9532959c59965c6d6c6f9c7f03 [CustomOp] Register RMSNorm instead of overwrite forward_oot (#2284)
e14f2ef6690a04a3359220c94d6ea3eed7af5049 refactor select_experts of moe module (#2150)
103654ccd6f90d76b49354a00185601ac2d9ef7c [Misc] Remove redundant imported `envs`, using `envs_ascend` instead (#2193)
55d0790597af72d996b9c1ad5a6592f334c646b9 [2/N][Refactor] Refactor V1 attention for better extensibility (#1995)
8914d5a4b27ba3e6c0b908e6acfcaa228e015e97 [Quickfix] Add the missing `apply_router_weight_on_input` in FusedMoE init (#2348)
0f7492d18e572f53cce67615aa714ad09090f859 [Bugfix] fix the oom when chunkprefill with long context like 64k (#2319)
992271b027770257915386d06a715eba1dbc9c12 [1/N][Feat] Support MoE models with ACL Graph and refactor MoE communication logic (#2125)
1a70564e7c1ecf5cb16a65276eecc14f179cfc4c [5/N][Refactor] torchair model runner refactor (#2216)
dc585f148acc91916916556ffd37fda708657cd3 [main][prefill optimization] Optimize parallel strategies to reduce communication overhead (#2198)
c8b0f5f7998d06d26f2e203f2a6e8beb1e3f9521 [4/N][Refactor] torchair model runner refactor (#2208)
eb43a475f429192e7509e85e28b1c65d5097f373 [Feat] chunkprefill mla support torchair graph (#1772)
881e36d6a93ffd06939c19b2c6f2fd584451507c [3/N][Refactor] torchair model runner refactor  (#2207)
29aaba5f845c276892c962dab21e4f29681e478a [Perf][MTP] Optimize reject sampler in greedy situation. (#2137)
c0f0b708137838c1387b4f98eeaeda383d06f0d9 [core] Support capture custom ops into aclgraph (#2113)
1ab15414bb291e5669b5211ce3d66879f61670fc [2/N][Refactor] torchair model runner refactor (#2204)
9260910c8dee96b7fee4382723d682999c918584 [CI] Fix broken CI (#2302)
ad1083761f06737eb1ad2424dcd74aab47b559cb [CI][Quickfix] Fix AscendFusedMoE init error (#2268)
dceef080b140305841a98d71cd8cfdeccbd390af [main] remove torch.cat and replace it by List[0] (#2153)
b2598c3271af2d55a948e29799e90bc0cfcda188 enable mm allreduce test (#2192)
205eff2b12545213a3adde77090361ead98f1fc6 [Bugfix] Disable check vllm init temporary (#2250)
c6112916614b3be200bb49f66c89181be055d10a 【main】SP For Qwen3 MoE (#2209)
57b9f021853a156deaadbabb3c62716815e21e5a [Bugfix] Fix disaggregated pd error (#2242)
26fc36b0e026f854f703d7ffc113c682985f2319 [V1] MTP supports torchair (#2145)
bf84f2dbfa0f5d54554d70b1ba52413410797433 [Doc] Support kimi-k2-w8a8 (#2162)
8a59367d0c249b8242bbaa9fb69a2713ea596567 [main][Feature] Support deepseek w4a8 quantization (#2172)
e31b31f9c3fa2f758724746c29a8b5876af38cac [main][Bugfix] Fix unable to load qwen3_moe quantized weights (#2219)
f3b50c54e8243ad8ccefb9b033277fbdd382a9c4 [main][Prefill Perf] Optimize Quantized MoE Performance by Reducing All2All Communication (#2195)
292fb8f69601507304b8c1c0a90487e73b0bc164 [1/N][Refactor] torchair model runner refactor (#2205)
458ab2db12ad3d4104eadf0e13a452464b09765d [BugFix] Fix the bug that qwen3 moe doesn't work with aclgraph (#2183)
583ad8f347a196e3ced5008892558d70eb979698 [main][refractor] Refractor forward metadata retrieval across DP nodes to reduce redundant padding. (#2062)
807f0895b2cf4f569f0c86aa808149bef70493b1 Bump torch version to 2.7.1 (#1562)
36e450eb0f4fb5e8a262384fb302a5b7dc83b582 [Misc] Nit fix for disaggregated_prefill  and ascend_forward_context (#2097)
ad366bf90809ee60b80a2e768a47ed4ab2b66b52 [Bugfix] Follow vLLM Qwen-Moe/VL and KV Connector change to fix broken CI (#2181)
957c7f108d5f0aea230220ccdc18d657229e4030 [Bugfix][PD] Make multiple Ps and Ds work on a single machine (#2080)
a9480d5f0a031919d151a842c56563d318e4ecf6 [Fix] Adjust use_aclgraph logic (#2156)
688350a3bb26f7b4c92eb9c08a5898b399f6ca5a [bugfixed] fix the bug when run the inference of quantized ds-w8a8-mtp (#2134)
af04ee9e7a53dfab5ddbcee5535e8645cc84c169 [MoE][Dist] Fix Qwen MoE accuracy bug in DP scenario (#1856)
f939381c6fcdd38c33d691de22616bab66f98585 [Bugfix] Adopt the new changes on disaggregated pd from vllm main branch (#2122)
6e00aed4d540432ffe2317601d47039d4ef2332b [main][Feature]Moe alltoallv communication optimization for unquantized RL training sence  (#2088)
8cf97d8310cfdad33efc1af24c3251fe6c9c973e [Misc] Add extra checking to torchair_graph_config. (#1939)
2284289880f288cea6d94b6cf472b74ec3662c81 [MISC] Cherry pick #1291 from v0.9.1-dev (#1825)
9e65da990ece406e23bffb64150b432b3a99073a [Misc] Add warning for incompatible Ray backend with ACL Graph mode (#2132)
968e6791d38b53dba2cac62edbfd8e9b11737889 [Misc] Add data preprocess functions to qwen2.5_vl_without_padding (#2148)
e3b3ffb87556cffb80fcdb93313a4a9575d34324 [Misc] Disable quantization in mindie_turbo (#2147)
c62f346f5d452f81f7b6ba46e9a13c2c1e270c93 Fixed 310p failure when using the sampler feature (#2151)
9c9a7cd90bf396eaac09402b10382cd61fc921e4 [main] adapt usage of npu_moe_gating_top_k_softmax and remove envs.SELECT_GATING_TOPK_SOTFMAX_EXPERTS (#2112)
4c8842da659b31af8aad1905b8d7a4f0736283ea [BugFix] Fix a bug of running chunked-prefill with torchair. (#1378) (#1844)
2008152c48c976a3d2858cccc9f8d301b0915df6 [main][bugfix]Fix vLLM startup failure when inferring DeepSeek R1 model in DP scenario (#2020)
6192bc95c0e47097836e9be1f30f2a0a6fdca088 [Bugfix] fix tensor not same device in qwen2_5_vl_without_padding (#2051)
72eceff94d8f31022ad5160e590fbd8c7163967c [Bugfix] `grammar_bitmask` IndexError caused by outdated `apply_grammar_bitmask` method (#2022)
4fcca137a70c11daa4070ae014288be154715939 [main][Feature] Support Qwen3 W4A8 quantization (#2060)
1dbb8882759e4326f5706f6e610674423376c2f3 [Bugfix] LoRA logits einsum dimension mismatch in add_lora_logits (#1583)
9b67c87b1475fe7dc79442efc8f2fcbce7b728cf [Refactor]Refactor sampler (#2050)
b6a7f07c701984eb5c76e474a74f8889f9c300c5 [Perf][MoE] Improve MoE multistream parallel performace. (#1891)
540336edc9db09072a9aaa486fbf7ce625da5b9e Add  Custom Kernels For LoRA Performance (#1884)
ca8007f584141d3a59b2bcbd4f8ba269c9b7e252 [Feature] Enable inference support for Deepseekr1-w8a8-MTP (#1994)
98cadc2146949017565203733dc50ea80f5b3bf5 [Perf] Avoid performing index selection of sin/cos cache every layer (#1890)
0190b68f51f2d2a0554e537da21910603b778804 [Misc]Remove PD v0 code (#2047)
e7d32ed3f18dcf12ad7dd0094984902cf542fab9 [BugFix] Fix the problem that torchair doesn't support tp > 4. (#1508)
4a008c4dac2cb0c6863b9ef9e36926263029b5e2 [Misc]Clean up useless import from vllm (#2049)
34cfdf5520225e0c02a3a335a698b6ccc95bdc6c [Misc] Fix logger bug (#2024)
32a9c5f69430c326a63f930c820386b3b3ae66c2 [Feature]: implement the fusion of allreduce and matmul in prefill phase when tp is enabled (#1926)
ba3dfbd59e43b9071895f483d12c034d8538ced0 [main][refactor] Refactoring forward_context and model_runner_v1 (#1979)
d1c640841b9354322e403d9689830b948f04ac89 [Bugfix] Fix num_hidden_layers when Qwen2-Audio 7B (#1803)
df0ec55162339c08b0bfcb79a2dcfac89f8c6e33 Disaggregate prefill for kv cache register style (#950)
17a430f7b82e31cf683270fdd7ef972bf99ef8ba Upgrade vLLM to v0.10.0 (#1927)
2f50304c19bd7bac8fece88bcb4b273d6e64b412 [Bugfix] Add get_supported_tasks interface to fix broken CI (#2023)
bdfb065b5da753ca99ccaa97fe6a921e34a4d6ff [1/2/N] Enable pymarkdown and python __init__ for lint system (#2011)
cfdd45ed00ab9c344f3aacb74f3521eba7671675 [Bug] Fix duplicate 'torch.' prefix in qwen-vl (#1986)
84fc7402c3176bcf39654ccec05af14ff6fd07cd [Misc] Refactor AscendMetaData Comments to Make It Clearer (#1967)
fa76a9b7bb244d79258de3a195c93740e3c18281 [Bug] Add prefix parameter to parent class initialization (#1934)
846555cdb593b130647b53f6bc324b17c43f1df3 [Misc] Clean up uesless code in attention (#1933)
3aa3b46bfef2c12cf190e02f2f3b1fc31e6865b8 [V1][PP] Support pp with ray backend in V1 (#1800)
9a3bdf21628e5942437af72698c4efd23ad1bc96 [main] Use AddRmsNormQuant ops in the custom model to optimize Qwen3's performance (#1806)
33e1ea4d1ac70dbbf19f0ac44a54561b1ad3e264 [CI] Fix broken CI (#1915)
7265dc090d0ef0d313d7a3ea6ddf0bc64cb2de07 [2/4][Refactor] Refactor torchair utils (#1892)
957b0b611fdae2ffb89b2b6bdd153d3c27f12761 [Misc][V0 Deprecation] Remove V0 Model Runner (#1823)
af56ae3ed1c71543b1b3e6a33b23add36f5c131f [1/4][Refactor] Refactor torchair worker (#1885)
8cfd257992c116fe8d186e17d948528f4e8fe10b [Dist][EP] Remove ETP/EP maintained in vllm-ascend (#1681)
a8b316ac5bf623d4e2f1a561fff66e0c5ed898bc [CI] Make AttentionBackend interface compatible to fix broken CI (#1893)
2b726d8f905b29d776eb1fea18c37e7a16522fc3 [CI] Fix broken CI (#1889)
b824525be3af3f0e6e9b073de3552be21bf0fef7 Move deepseek_v3 from deepseek_v2.py (#1793)
ab68d31a24110c3ed2b0ea5c70016ea8874059e7 [Misc][V0 Deprecation] Remove Cache Engine Used for V0 Worker (#1878)
53d2ea3789ffce32bf3ceb055d5582d28eadc6c7 [Bugfix]Fix the performance gap between 0.9.2rc1 and 0.9.1 (#1811)
574fe407eb2fc334e3e86c3558dde3e33be33427 [1/N][CustomOp] Register activation customop instead of overwrite forward_oot (#1841)
d08ff304cdd0642c89ac64ce948a1d94ae397a50 [Misc][V0 Deprecation] Remove V0 Attention (#1835)
f9dfde02fd3d4df55079fd2697c1fd7e79afb9e7 [Bugfix] Fix broken CI (#1848)
875a920d4a105473e464c7c681b62a4c3b7bd679 [Platform] Add support for Altlas A3 series (#1794)
c66b0827a73107b0a669a6c5fb7b96e5e7e9642a [Misc][V0 Deprecation] Remove Pooling Model Runner (#1824)
06655002c57149163735742a325d10f4212fad73 [Misc][V0 Deprecation] Remove V0 Worker (#1821)
b005def0a5791075e0828a1ccb07fb6f11fe4598 [Misc][V0 Deprecation] Remove Multi-Step Model Runner (#1820)
f9e2e9bb31732ada5762556f4730b97b5c119e02 [Misc][V0 Deprecation] Remove Draft Model Runner Used for V0 Spec Decode (#1810)
f96100fad51b9b5eb6675e42a464115d560037c0 [Misc][V0 Deprecation] Remove V0 related codes of test, example, platform (#1805)
a929699e98e0584617eaaa7d7f76b36df657bf98 [Misc][V0 Deprecation] Remove multi-step worker (#1809)
7bdada58eb61dd1009e236106aecfb2e425e8861 [Misc] Remove VLLM_USE_V1 usage in code (#1764)
494b0f474fd21387bb91d0eeef8ba3e669391daf [CI]Fix broken CI (#1773)
d13fb0766e1805089499ee3dfb0329949d1905ac [Perf] add patch to optimize apply_topk_topp (#1732)
aa4240c67f32923cc46a6d05967a29b2a1499dd9 Support pipeline parallel in V1 Engine (#1700)
ee40d3d850bc906e06e66b81b3bb233f38b16546 use npu_moe_gating_top_k_softmax (#1355)
9d16c9982e7bb67053e828b1cd8ba734b11670b2 rm router logits Improve TTOP 3ms (#1407)
0fc9b56d400f6daede1044bb03187a3c106c4eb1 [Perf] Improve MLA multistream performance (#1353)
cc210f46e61e2c45e36e4aa7595be7fc415d27d2 [AscendScheduler][Bugfix] Remove num_draft_tokens while allocating slots (#1718)
c7446438a98b56a2b7aec5dbb0a92c7dc75138e9 [1/N][CI] Move linting system to pre-commits hooks (#1256)
643e6f5486967fc49f4f9bb6a5d012bfcc6b91ac [Bugfix] Fix accuracy problem caused by mask pollution (#1678)
60519c71bd1f1395989693e755fd3fc6a1a985c6 shared_experts+router_experts merge all_reduce(Improve TTOP 5ms) (#1395)
89c1a0f00691ce4cda87cc7eb6f5e558e926a455 [Bugfix] Fix memory-leak caused by dist._functional_collectives.reduce_scatter_tensor (#1380)
b979ee353d1eb740a1f058fa182154790236b995 [Misc] Code clean up (#1679)
392fd7239bb8ab12e0ad45c454f6941ead6a2a56 [Misc] Add attention mask (#1673)
cc1588be50caf8190a2b84412ff769c07ce52db5 [Misc] Code clean up (#1674)
830332ebfc649e5f0604513460e77207151c7545 Clean up v0.9.1 code (#1672)
71de52d3a94e605b9de3183daed4106f626505e4 feat: add kv cache memory cache and skip dynamo guard (#1549)
df84cceca841eb69b7a73e1a0941dc74084a3ba5 perf: use multicast to avoid padding decode request to prefill size (#1555)
f08c4f15a27f0f27132f4ca7a0c226bf0a2a47d4 fix spell error (#1654)
18495f44b23754e7fafdf8f04f855e201e26aebe [BugFix] Fix max_num_tokens_across_dp calculation bugs in attention_v1_torchair (#1636)
c58accc15e2a4e8a672c754035168a09ccf421f6 [Bugfix] Support Qwen3-MOE on aclgraph mode (#1381)
eb390545ec486c48916cf42f246ff693171210c0 [Performance] Disable JIT and nd2nz to improve performance for Altlas 300I series (#1591)
dd22ac38b283e6d93b424209a19d207d4d31a466 [CI/UT][Refactor] move e2e spec decode and deepseek acc test to per pr (#1136)
343955c7ac7c6615223366ad48fbd11ee74969e3 [CI] Follow vLLM FusedMoEParallelConfig interface change and clean up unused config (#1625)
a5f33590d39f1856d5ed59271a31ef78897069c0 [CORE]initial support for torchair with non-mla backend (#1506)
9fbd8017c0d1e6e09ac4568ff16931118a56ab12 [Quantization]300I Duo support w8a8 quantization (#1560)
a45dfde283dfa555cec76f282b1e0a360fa2743d [CI] Fix FusedMoEConfig and input batch failure to recover CI (#1602)
30bf7014d07b2c0b40b123262c17dd347447aa71 [Bugfix] Add func `swap_states` to fix MLA attention (#1580)
6b80c5acbad0224477b8de4e202f2853ce992eb1 Fix W8A8 fused moe bug (#1529)
641a4e60928c977667af775817570ebdbe582bea [CI] Cache sampled token ids in model runner to fix CI error (#1573)
0e43813120aa5886dd7407477b5bea58baedb940 [ModelRunner] Use shared CachedRequestData cross request to fix ci (#1546)
8013634e9c4a85553b356ad1c8fbde92de9861a7 [Structured Output] Remove redundant check for `grammar_bitmask` (#1459)
f286265791cc6a770ce07a48daf6dab096c0cc3c [BugFix] Address PrefillCacheHit state to fix prefix cache accuracy bug (#1498)
5f8241c25ce486dbfd1786ba8b568c38484a8864 [V1][ModelRunner] Support pooling model for v1 engine (#1359)
75d05ee200ba7e59b053e9698e04076d2ba9941c [Core] Fix block table shape to make Prefix cache work with Ascend scheduler (#1446)
b308a7a25897b88d4a23a9e3d583f4ec6de256ac support pangumoe w8a8c8 and docs (#1477)
c59d69d9e65de4b91628411ef415eca6bf512b44 [PERF]support MERRouter (#1421)
8fa188111da3a8f752dc309330d9bd3ec18194e6 [PERF]support H2P communication optimization for PanguProMoe (#1463)
5c53cbaf2a7efcd09f2860eb6dff7412c29654b8 [BugFix]Fix bugs when initializing communication groups with dp on 300I Duo (#1478)
5f4391652f4a62d791fa4b9dfa3fc9d802d5a250 [PromptLogprobs][V1] Support prompt logprobs to fix ceval accuracy in V1 (#1483)
d59e7fa0959a5571e7debd884d27ea2e6d9cb582 [CI] Pin transformers<4.53.0 and fix EPLB load_weights to make CI passed (#1482)
5968dff4e000f8c4b00751d963898f1f0619b164 [Build] Add build info (#1386)
53c2d58ae18d4268024ceb3ba029e1114b36733c Handle with_prefill_across_dp for multistream mla (#1322)
2690697caa47ab8daee4083778020ea7c13c16c7 [Bugfix] Reset all unused positions to prevent out-of-bounds in GatherV3 (#1416)
2fda60464c287fe456b4a2f27e63996edc65dd40 [Perf] Use fused ops npu_top_k_top_p (#1308)
e7efc7e7e7675b5edf09e2e562c05c560df62114 [BugFix] Remove not using patch_eagle.py for CI. (#1385)
941269a6c5bbc79f6c1b6abd4680dc5802dd8666 adjusting the communication method in graph mode (#1194)
ca884ef86ddf976506cb0b0092af208bc12bcd60 [Misc] Clean up uesless code for LLM initialize (#1373)
52317f92cbb5a6bea25a0ad52e53aca1783617d7 [DP] Tiny fix of dp and update example (#1273)
5f5800ba423664610f42a4a7f8cf0efbbf154382 [Bugfix] Sync MRotaryEmbedding interface change to recover CI (#1399)
9cbce423ce8d8f54347e01985acd5549959fafb5 [MISC] Remove useless patch (#1366)
5177bef87a21331dcca11159d3d1438075cbd74e support fused_moe_allgather_ep (#1335)
15592c0d486a0df6939c58fb55c7864149951f34 [bugfix] fix accuracy prolem for deepseek V3/R1 models with torchair graph in long sequence predictions (#1331)
f04c6763d8b0a05905e241dfc7ea826a6b9cd587 [Bugfix] fix env variable in dbo (#1284)
339d6894f649b92f57675583ec235f10dd858152 [CI/UT][bugfix] fix v0 spec decode (#1321)
097e7149f75c0806774bc68207f0f6270bc7d392 [Platform] Add initial experimental support for Altlas 300I series (#1333)
2f1266d451b4576c297d6d65b508487a4f188303 Support Pangu Pro MoE model (#1204)
00ae250f3ced68317bc91c93dc1f1a0977aa0b94 [V1][eagle3] Support eagle3 proposer for v1 (#1032)
2c7dd85fd8ce4a0951fc441d9c9c432d47c98d32 [Fix] Fix the token-wise padding mechanism (#1300)
b350edae9a550caf7b9165bcb7df0860c2c84e3f [UT] refactor test_expert_load_balancer and fix broken CI (#1293)
ebb2a70dbbdb8f55002de3313e17dfd595e1de1f static EPLB fix bug, add unit test (#1186)
2cd8ecdc4f1d9ee5e24a0fa70a2b4089785f8837 [Bugfix][Spec Decode] Enable `ACL_OP_INIT_MODE=1` directly only when using V0 spec decode (#1258)
db2f630aebb0cad44f9705ac028993233a00c82e [bugfix] fix deepseek with mc2 (#1268)
d7e19ed57a8b666855720d68185e140094890167 [BugFix] fix length of sin/cos cache in rope  (#1266)
afc8edb0460a14885ba780bf00e26afe1a9b4598 [Bugfix]: Pass scaling args to mc2 (#1202)
f8029945c32cc53ae5b6f6479f8a5ab4f731cb69 [Bugfix] Remove cuda related lines and add additional pip mirror (#1252)
23ca68d0c8557a91b7213782de5c7fafa0cfb985 [refactor] Refactoring AscendFusedMoE (#1229)
96fa7ff63b6866d1a13f821d7005eebbb1914183 [DP][V1] Fix rank set in DP scenario & Bump torch-npu version to 2.5.1.post1.dev20250528 (#1235)
f5404dc650882c6f0423db9e87f9b38f756211c5  Fix the device error when using ray as vllm-acend backend (#884)
69b817ed654b9dbff9e614c69752c2b76d9fe7d7 [CI] Add unit test framework (#1201)
4270682383b4f4876296565bbb1227c1111b1d5e Waiting for BMM NZ support(Improve TPOP 2ms performance)  (#1131)
ab5d110fcc35ca11330977450141b1d7176f21e7 vllm-ascend support chunked prefill (#1172)
e72f94e38f4208d36a20df1ff85ad4abf430447d Support multistream of MLA vector operations (#1135)
3393d53b365e81ee5a547f4332b21381056c2822 [Scheduler][MTP] Add support for speculative decoding in AsecendScheduler. (#943)
4f5964420e271b2ed7e6cd89cbad2bb66b169114 [CI] Upgrade vllm to 0.9.1 (#1165)
e46dc142bf1180453c64226d76854fc1ec696169 Enable kvcache_nz for the decode process in torchair graph mode (#1098)
860a5ef7fd1930b18f28ac48d803ef7f6298e2e6 provide an e2e guide for execute duration profiling (#1113)
7bdc606677705d072c1dc45f050a3c3471d6d379 Support multistream of shared experts in FusedMoE (#997)
8dd686dfa20d8f577cd63fb751ce0a7a98f2f516 [MLA][Graph] Improve assertion on Graph mode with MLA (#933)
291c216898c989efb4b59a64dad818d2dd48a71c fix torchair execute issue on padding data, and mtp padding logic (#1160)
95414bae7046eb4bf2323f5f21473c3b1cd158af [CI] Run e2e after pre check pass (#1132)
b75cb788dd8e2058fec79b2bbffff277ee5f12d2 [Bugfix] add compilation/__init__.py to fix import error (#1152)
706de02317625fa3477f5379e31caaec0b174453 [fix] fix compatibility for non-EPLB scenarios (#1142)
cd2f14a1b3563a70c70908b07c76f6d3fa282b0c [MTP][V1] Adapt mtp with graph mode in v1. (#1023)
6b853f15fe69ba335d2745ebcf14a164d0bcc505 Add static EPLB (#1116)
d2f87ed9ccd565d323b9e177856fe59c46b8467b [Patch] Remove `spec_decode.metrics` patch (#1016)
6003afa6d20934a1891b8fab0a50c187610fa120 [BugFix] Fix data parallel (#940)
eec60681878bc62ad971ce79b86152c4234c5bf7 [Bugfix] Set `ACL_OP_INIT_MODE` env var default to `0` (#1123)
4976b48b98f7268a68fe055265f928afd779ccc4 [Build] Move numba/quart to requirments and update DS baseline and sync graph typo fix (#1121)
f1543d5e0d71de4fc4a649c1d20d2d2ac2eb5d7e [bugfix] fix deeepseek accuracy (#1118)
c8742146d3db4726605f38bb9e1ae02ad658e7c9 [CherryPick] Add unpadded Qwen2.5-VL for verl scenario (#1095)
b80a484864553547bc8ca3beda24858ef7bc135a Fix typo of VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE (#1112)
20dedba5d1fc84b7ae8b49f9ce3e3649389e2193 Add qwen2.5 vl multimodal feature for vllm-ascend v1 (#736)
87ebaef4e4e519988f27a6aa378f614642202ecf [perf]: support dual-batch overlap(dbo) for deepseek (#941)
3640c60b0eb4d4cb104e20bfa406d3f1d17920a7 Avoid unfused Transpose in DeepSeekV3 EP256 MoE layer (#1091)
8d00775fcedcee8e652c424be336943ddfb9a38a [SpecDecode][CI] Set default values to fix spec decode and fix multicard CI (#1109)
e9ada685ece798f9fe0d4a287e3f5246a8a7207b [CI]Moe alltoall communication optimization (#1067)
a2552e10e4591ef97b32ce0a256b027fd662f617 [Worker][V1] Support sleep mode for v1 (#1084)
9a4eb94ca9f5bfad868164facae069d28cf7028a [Misc] Adjust the default profiler configuration (#1097)
5d0e9fd19ae7051413fbdba24d0d77ec3d5511fa [Misc] Add `ACL_OP_INIT_MODE` env var and set default to `1` (#597)
11a7df42703fa3df3efc883c0bd2ee9c8f80921b [ModelRunner] Support embedding inputs (#916)
c7f1c59911027953147c9d9495457568fe2216c8 feat: support compile multiple batch graph (#1085)
c46632439a59dd43f9062396128119b143561e8a [Bugfix][DP] Add with_prefill_across_dp to AscendMetadata to fix dp (#1094)
0b12c2acf7d9fd192beebebf662298067d9a5435 [Kernel] Remove cumsum in groupedmatmul (#987)
dab19d5dca56f4f3eae0a019ee486d7e350f93d1 [BugFix] Fix ascend config check (#1092)
973f993a131c8c9bcc10d46be946dc9a6eb68811 [Misc] fix initialize_kv_cache (#1102)
c94afd79ceeb19424ed0d6da4c4fe82d29b5242b [Doc] Update the description for env (#1079)
6b094a2bd49a8a41eb3647568b2d9e5b337db81f [ModelRunner]Add profile execute duration observation  (#1013)
78431b34694dfa3c8f54ed7cc626660318557927 [perf]Support MOE Multi-stream in Deepseek (#947)
908a851a776cfd9051cc062119e6ec481561c6f7 optimize the funtion of computing topk and topp in sampler. (#970)
e1ab6d318ee92c5b34ab3d7374bdd288fce0375f [Misc] Refactor additional_config (#1029)
afc4c0cd035ee22d51950d24b7f642abcaa0b176 [Bugfix] Fix deepseek percision issue and add acc ci for it (#905)
da9acfca6053352730fce75fb772e214755d0341 feat: support data parallel for deepseek (#1012)
517811449e466e071988549f6ff1a1844fb07163 [CI] Re-enable sleep mode test and skip failure breaking CI (#990)
068c3a0167515a3f77b5b080c43164db31171fb7 [Bugfix] Add verification for `quant_action.choices` to avoid `TypeError` (#1046)
93860574bb543b54e93453de7c3329f5e81769bc [ModelRunner][MultiModal] Remove legacy input mapper/processor from V0 (#951)
6ec64a3f9686df65b5a23a41aa301e669db19099 [bugfix] some bugs maybe fail to run (#896)
507ae627cad68d93c62adf3c2409f4ff10a25536 feat: support compile torchair graph while warming up (#839)
5a1689fc648c8afa04ee040bf1b3526a6fe3d75e [Fix] Fix update_aclgraph_sizes when running MoE models (#913)
3442fbdb235b4c6d72c2bc64a49707a7bd89958e [1/N][UT][v1 MTP] add basic v1 mtp features (#890)
05a471001baf35340e000d74ea24bb1ea153fcc7 bugfix for qwen2_5_vl (#805)
a93bed45350315586a2818e2398099a29b2aa215 [aclgraph] implentment NPUPiecewiseBackend to enable aclgraph (#836)
cc74b97f742971634bcdd9282a3cf835fc8275c8 [Bugfix][V1] Fix deepseek with v1 (#958)
e3c7f71462f1c252a60ce6476336fb24311a593c [Perf] Refactor tensor disposal logic to reduce memory usage (#966)
6eddbd2521d9f22b90b302d65f986254251d49cb [CI/UT][PD Disaggreate] Initialize PD Disaggreate UT (#889)
f6e5decc109663f39e44e5fb792988b18a74fd53 [CI] upgrade to vllm 0.9.0 (#959)
9f5ab59e307a66fd0b17916218c9387c328b1c59 [WIP][BugFix]Fix accuracy issues caused by wrong etp_size passed into FusedMoEParallelConfig when using vLLM 0.9.0 (#961)
a0c3e9ba506e8c02b79a1cbe4d3e4daca15eb1d9 [Bugfix] Adjust inputbatch to be compatible with latest vllm (#945)
1f9fb869ad8832fefed92ae6ddaf36552d694c89 [BugFix] Fix accuracy bugs for unquantized deepseekv3 models (#897)
17f05b10893bd18558b3c69f7af880fffc6c1653 [Feature] Add CustomQwen3MoeForCausalLM model (#925)
df58fb80eee24139fc61c495be3ce79cf81b3f73 Spec decode support for V1 Engine (#874)
a970b27e2ddc4de2e49aebf7dca447bb143a1b5d [WIP][Perf]remove unnecessary padding before MLA V1 prefill (#917)
dc6172efd3860ce95b40a7b3e93611f875f06d40 update attention nz and mla nz(Improve TPOP 6ms performance) (#909)
7153d8890b91b7807e817acc2a32fffbbe41e2fc [Feature] Impl v1 disaggregated prefill in ascend scheduler (#852)
b434f37b46116ec288f10ea75192ce80eb75ea86 [V1] Revert the default value of enable_chunked_prefill in additional… (#935)
46df67a5e9ab73fade08cbb2d8c0155cee7316d1 [bugfix] Improve log level and info for custom ops build (#937)
0f53b138f6ba7d31527bb105c99168fd1cbf42a8 [V1][LoRA][Test] V1 Engine LoRA support & e2e test (#893)
7aa4f85f10d55ca1b4a97f86938c9ef9e9706e44 [Bugfix][kvcache] revert multiple kv cache groups (#923)
b4d6672d018689430551f7ce2d115b4847bce239 [BugFix] Fix chunked prefill bugs in engine v1 (#844)
a73bd6caf44bfe677ffcd18387d3bdac3cf5ad48 [Fix] Set div_mode to False and fix view_as position (#912)
5cf9ff18e91b0b7031c258d71a257b8e24689763 [Performance]: Custom AscendC Kernel of Multi-Step Prepare Input (#814)
00e0243561720ad3ba66ba43e84031dd0e91814b enable online serving quantization (#877)
732664451309f34828a1f387f20cec2cbf757f14 [CI] Fix qwen2.5 vl CI failure (#888)
7a325b2e2d1001a6341ba71eebd8bd8d4458d12b [Bugfix][Model] Fix fusedmoe and make modelrunner_v1 compatible with latest vllm (#867)
1e67089bc970bfe2f667f321fd0c0777a0c55a26 [BugFix]add all2all when dp_size > 1 && downgrade npu_dequant_swiglu_quant (#819)
68fb63428b8b972dd60ad9189538909b0eb1fcc8 [CI] Patch torch.library.infer_schema for fused moe ops to fix CI (#854)
857f489cbf20ce76f69d065a03597442804ff888 [CI] Patch torch.library.infer_schema for torch 2.5 backward compatibility (#837)
e56447033889ca95df512208cab22ef832bfdf07 [Attention][Kernel]moe support for llama4 and mllama4 (#740)
c6ac399091f927dd743267ed7cebbba957e1a92f [Bugfix] Fix the method of importing environment variables in DeepSee… (#817)
6193ba679b159cdd73b2c145423e494d90301370 [CI] add codespell CI and fix format.sh (#827)
5998704c0857c4139ae196d2ce06d748afcf70e9 [BugFix] Fix ascend scheduler bugs. (#822)
701b0fd95ea188897998d056576d38e9229b6fe0 [Enhancement] Add padding for ACL Graph (#803)
efabd722eb757e49aa309c173bbec91ca8c4ced1 feat: support torchair graph mode in v1 engine (#789)
5305a2ccf943304435a7716120ee6bb5c130a6b8 [Bugfix] Tweak distributed process group initialization and add dummy… (#816)
cdece86f2cf27a47f800403a00f0816b068493a1 [Bugfix] Add max_num_batched_tokens to InputBatch to make main CI pass (#806)
fa99f89e93d1e70d7685a128ef003333ef17b302 [Core] Support the features of prefix cache and chunked prefill in v0/v1 (#782)
324f819b929ae2c30fe629e27df0c37c2cc607b7 [Perf] Optimize fused_experts quantization code to save npu memory (#784)
2c685e3b61b0dae7ab26a95913cee1ba37bdd7c6 [Bugfix] Correct method call for _set_cos_sin_cache (#774)
6c020883a8332b5c519f4f6502733edd9b391c2b [WIP]Add Func: aclgraph_batch_size auto-adjust to different model (#771)
2e3520e28518ad5009a68efd8d64476f52e6d147 [Bugfix] Fix output tensor shape in vanilla_chunked_prefill and update import paths for model_loader (#773)
2cd036ee8ec737cefa8d30c0acab56e4f18ea189 [Bugfix] fix accuracy problem for quantized deepseek models (#768)
d6e94176528b7b1d7e24e2cfee0b9cc663b8769d [Bugfix] Fix masked_fill_ function typo (#769)
afe1767c17cda86483a0176b451f181989757e41 [Core] Cleanup triton patch which has been fixed in vllm (#764)
d6bfae8eeebedf677b643b712d367a3a69c9cce4 support 32K model len on deepseek r1 W8A8 (#728)
d7e1110c8ed217d8fc01bc5ef4dc180aa47f6376 Re-patch TritonPlaceholder on main to make CI happy (#753)
8b194ad12ec629edda070008bdc332a0157f74ed [Disaggregated Prefill] P2P Disaggregated Prefill based on llm_datadist (#694)
84e2ed898b7edec4f358d0809a2270dfe7bacfd5 performance optimization, usability optimization and API compatibility adjustments for deepseek with npu graph mode (#731)
3a628891ab6473f9b3f052ff5fe3d19bacbae413 [Feature] Add quant description file for new quant model generated by modelslim (#719)
ba9714ccee43c7c12cca9e7eb02b104b85ffe2c1 Optimize qwen2_vl and qwen2_5_vl (#701)
f8350569e6267675861a8e8e4b975400268e7d2b [CI] upgrade vllm to 0.8.5 (#715)
95e7aa47363760792461525e5e9e8dbe80f85101 [Platform] format platform to make it more clear (#610)
b917361ca55b034606bfafa1f0dac5834321e506 [MISC] Clean up torch_npu (#688)
0329fad9276f4b29a4766edc9c00539b05e0592c [Perf] Deepseekv3 performance optimization for eager mode (#598)
87975fa058fe3f90d204ded42a08989a8dcb413e [Bugfix] Fix early return in CustomDeepseekV2MoE.forward during profile_run (#682)
0dae55a9a3deebdb4f2263011154d886c525fc13 [MISC] fix format check error (#654)
1fce70a2fb2602170781773104a69c69beb16161 [Model] Support common fused moe ops for moe model, such as Qwen3Moe (#709)
40bd6024856b340dfb0ad80f101d96f670c72991 [Feature] Use reshape_and_cache fused op (#706)
54c0e63df7dc513a2c2316e2ccd6ed1dfdefdff3 [MTP] follow custom deepseek modeling changes to support graph mode (#636)
be9e3e85457381fc537bca6c7ad4cb97ad39d32d [Bugfix] Fix triton placeholder patch period (#704)
5de3646522b3de5cf1e06ca579725bbaa5ed3aec [MISC] Make vllm version configurable (#651)
38f34e359f08bcf652d2a95b2a2521880286d710 [Fix] fix deepseek v0 attention eager mode (#671)
2e20797934fe1f357fe840f538543445acd7b92d [BUILD] Upgrade torch-npu to 2.5.1 (#661)
fa4a5d980e8845a88b9162cf169f0a5ab230f8a5 [Bugfix] Remove redundant tensor creation and unused code (#656)
ba3d8aae943271221ff7b1ae74551939fd935ee5 [Model][MiniCPM] support MiniCPM (#645)
742f679c7d56ebea9cee449ba2428940aa353bb9 Remove prompt string from engine core data structures (#663)
3879d9cad95c14e3cce8fc053540e369a39cd341 [CI] Fix sample backward compatibility problem (#648)
d785e785639a0ebcf21c0b5e46ab47a3b041344c [V1] Make V1 engine backward compatible (#637)
a9c6b52205c3911e0725549c0fcdd7089e8f25a7 [Bugfix] Fix qwen2.5-vl positon input bug (#639)
05bdcbeae47c7fcb9b1c30cad059abf1d40b5421 support aclgraph (#426)
5c6d05a59e996ab0ce6b91e7d4e267d7be1157f8 support deepseek quant & mix-parallel with graphmode (#585)
e74331a1ede31c69ec0b1b97bd407d38742caa9c Add dp initialize patch with hccl backend (#626)
4a0ce3660ed3188c2531dd63341350df712bd97b [Misc] Remove some parts of metrics patch (#603)
538a69c1459cc8fde032b8db211ea215c78063b9 [Patch] format patch module to make it more clear (#601)
d12a057df850f2664f4388397646ec9c32beb88e Add note for deepseek related docs and remove unnecessary comments (#590)
a8d633f629cb1c6c81c80ba3bf8babcde698bf65 [Bugfix] fix import error (#600)
0ae9ee0f8a0fb8f8c20dd5005932a68550a15f89 [BUGFIX] main-sd-bugfix && [UT] add mtp UT (#593)
5442b463fd7b232fc560dd549e50b737043992cb add doc for patch_config (#574)
12cae04db9ebe6cd70dede409c63ed5d326537fb [quantization] Support w8a8 quantization (#580)
1a1f9a6d894a3947fcff4c5c52fa0846c35e5759 port deepseekv2 and mtp to main branch (#429)
a127cc83f89c249ff062be76124e54a05e2a31c6 catch ImportError when C code not compiled (#575)
65c1f4579fb3cb85f0d56d5b08d4d226aac62d9d [V1][Structured Output] Add `apply_grammar_bitmask()` method to model runner (#555)
84563fc65d938f1b2655eac18164fca5666d3c67 Add sleep mode feature for Ascend NPU (#513)
42c7fbb10eb0d128488c38dc6d1061608fad04cb [Misc] Fix import error and address nits to make CI happy (#563)
66a0837963ff5dd6734907083ca3cf57e6bb223b adopt rope in vllm-ascend (#530)
23f85e3f7425d0041abc28a46d6010740ec36bbc [BugFix] Fix scheduler problems in last PR. (#558)
6ee7f5cf711509349fc06f7871659174395d1801 [SpecDecode] Add spec decode support (#500)
20dff4deffc90dfcb96472de5ae36737e78cba96 [Scheduler] Add AscendScheduler. (#543)
697908f5cd7c65a3a917ec1a962b0886efc98c7e [Platform][Worker][ModelRunner] Add LoRA & Multi-LoRA support (#521)
9935d457289ae85c0bf3ffbe5496875c6eb34782 [CI]Add model basic accuracy test(Qwen2.5-0.5B-Instruct) (#460)
c3d1a3782aee57432b9f6071a1bdc02462117f21 Add pyhccl (#503)
6061f3367010961e6d35b0857c719a1aee391323 [Bugfix][Model] Fix api in DeepSeek model (#545)
415ed027fadb730694855f987f34501481e8431d [V1][Platform] Remove `supports_structured_output()` in platform (#531)
bbe7ccd3664928560169cb0f965a6ee768a5a993 [MISC] Add patch module (#526)
bcbc04f92b258b5b6aa4e7fe4ae073611de97bc1 [Doc] Add environment variables doc (#519)
44a8301424ded94dae83e13b837f5bfc0a1bfc15 [Feature] Add PD separation feature (#432)
c7f6584d75e164f6140f2d822e454d477b887e27 [V1] clean up V1 code (#505)
f6af1d2471b99942d5284d804c10d4cb3c3b38b8 [MISC] fix logger (#515)
9c7428b3d5b63939c15ae713edc3871e51b98cbc [CI] enable custom ops build (#466)
f6cf92e7d55004bf8eb8729d71acaf0635f88883 [quant][bugfix] fix deepseek quant bug (#478)
1d88dacf9f9656b78211dae045227096e47d5795 [V1][Platform] Add `supports_structured_output()` method to Platform (#475)
344228a5da6130bf1d7e5f3c3d04cc8b661748ff [deepseek][bugfix] support deepseek quant (#469)
3f9752f8ee1c71435aeccfd6f753b70989071767 [Bugfix]Lazy import vllm config  (#462)
ce8259975e0befc4830152f631c826ccc046e808 [core] Support custom ascendc kernels in vllm-ascend (#233)
14d9a640472f7e63f90c9a138e614c4e6ad0f06e [ModelRunner][V1] Optimize V1 attention mask (#442)
2dbd763584699fbea874b3fa083a8eb273447732 [CI] Fix mypy CI (#443)
31f29b9f30eb65f0a38d30ffe613282ce2f1130a [Core] Make V1 work and enable V1 engine test (#389)
57a84bb7befeaa0dc62aa35fa406e4d6affbfcca [Bug Fix] Fix bug of platform for parameter checking (#411)
b1557abab6534af830f1555f262332aba2bf6e51 fix multistep bug,remove uselesscodes (#355)
122505208ff6284f409846ca7294f4a4b9883285 FastPatch: Optimized Patch Embedding for Qwen2VL (#345)
89ca63a2c2d98dbd153b33596888090426a9e9f0 [Bugfix] Disable torch.compile() (#370)
befbee5883446ccb8b0df255a911168079a414f1 Update README and add collect_env info (#369)
c06af8b2e0f4ace8caf20f4cf4fdbb1978647df6 [V1][Core] Add support for V1 Engine (#295)
7330416de3fd2f8c6b9b82fb1ad0adfb9c70d483 [BugFix] Fix bugs when using ascend quantization (#275)
5c7a95b01d339bd33e343ef14ff497fa1b2f4eea [Attn] Support encoder-only attention with torch sdpa (#290)
12aa7115b58e6def5603e4eae6744f0af8e05634 bugfix for qwen2_vl (#301)
0db6670bfab8cb1d84c9e7270df0a1d42d6ce7ca [Feature] Implement EP-compatible fused_moe (#121)
4c9d78a0354773267b50772ffde86f85d18d3ed0 support multistep decode (#299)
feb6bdb12e6e5c2dfce8dbd9232e97d435b52ecb [Platform][Model Runner] Add hash of request_ids; Change blocksize back to 128. (#293)
faf8cd89cb9a853bd8a3c16b8d7321e2da4a2342 register qwen2_vl to rewrite qwen2_vl forwad (#241)
3217f0d10fbbc6e6cc8b0db9594b8cef515b4f90 [Feature] Modify description and api for ascend quantization (#243)
dcd0005058dbd6fd8672378565890cbda924b792 [Fix] Remove npu_group_topk before CANN version update (#242)
0d3463400a8ae776fc637f4db3a464c0d0dc3da6 [Performance] Change the shape of kv_cache to avoid view of k_cache and v_cache. (#204)
503f5045ffdeb26a17cb3a7972a8310dadb9089c [ModelRunner] Remove redundant profile_run() in model runner (#224)
ae49bfd13a8c6c6f549872e6a7666e2d3c68bbae [Core] Support pooling (#229)
b64ee7d346511b6ea7a64b09db58c17aa1c915ef [Dist] Set device as rank (#202)
14bca9911a265bb3c75708dbd4fcdfe56d267db4 [CI] Fix unsolved bugs caused by pta api change. (#190)
1715230867048aaf3102dbe6448b3c476db74c9e [CI] Upgrade to newest pta.(MLA and FusedMoE) (#189)
c131e43e7d5983b394d6846de432b3a0d7031935 [Worker]Lazy import torch_npu (#184)
6042c210bc715573a65c76209445a3d92054c1a6 [CI] upgrade to newest pta (#187)
fd18ae649453fa4c31b58b04ba75d3fd9ed0b3d4 [MOE] fix #176 (#179)
ee43179767ba1a61be543ed42beca276bee061eb [ModelRunner] Fix cuda hard code in model runner (#155)
94cd66bba7b8e90a4b00eb92649b1239aabf3780 [CI][UT]enable multimodal ut (#158)
1c238b930d2b21a37140a1b32e5ac12465fe5c6f [worker] remove unused assertion (#161)
7776f2e6a4ef5c372b53c6f092a5062c4d2fe083 [ModelRunner] remove padding for vlm inputs (#150)
79fbb20b4db5538f33ae1d1fc6f531847a42de8b [ModelRunner] remove unused args (follow vllm changes) (#159)
d0b3cb4fa79d5fc7f8245a3c68885ce1fa030ba4 modify:Eliminate redundant operations in the code to improve performance (#137)
202b39a38c2869b0ecc3df486550fb555a2eb0c0 Ray Worker Ops Optimization (#136)
386817b4d1c0781abcc5ab5370da3b444882a74d [Model Runner][Performance] Cache the jugement result of is_encoder_decoder to decrease framework overhead (#138)
dd425d68f8a51a7b1fcb60a193fcd0d3ea1848a6 [Platform] add dispatch key (#17)
5f465010deef1a2b507a107e720c1c366161d820 [Core] Cherry pick from 0.7.1 to keep the main code newest (#127)
8ea8523744138da981bf952f28a5eb304f9898c3 reset default block_size from 16 to 128 (#84)
4544e99d88aed9247381a420c896d39fced69096 [dist] revert communicator patch (#66)
b88443b6c645942b89991c3df35f5485630e8df3 [dist] fix communicator patch (#58)
f762ee89cc2e9fc7378b696139b264d80600adba [Communicator] Add monkey patch (#30)
70068359770b6e8cfcbb9931aa79be50731a274c [attn] fix device of tensors in attention (#25)
8fc5dc966aaf4e174d1ec0d1902c40289411ec0e [Worker] Register mindie_turbo while initializing NPUWorker (#13)
4495fc68389e3fb1ef14534c202948931e38446b bugfix for mrope (#14)
bfccf739e2fe121b54d9b198c2ec205a9379190e [ModelRunner] Refactor model_runner for NPU (#6)
d5e7756028bd5884ade96b654555c375770a2f64 [Core] Init vllm-ascend (#3)