1.9 KiB
1.9 KiB
| 1 | time_stats.moe_gating_linear.min | time_stats.moe_gating_linear.max | time_stats.moe_gating_linear.mean | time_stats.moe_gating_linear.median | time_stats.moe_gating_linear.std | time_stats.moe_gating_routing_topk.min | time_stats.moe_gating_routing_topk.max | time_stats.moe_gating_routing_topk.mean | time_stats.moe_gating_routing_topk.median | time_stats.moe_gating_routing_topk.std | time_stats.moe_shuffling.min | time_stats.moe_shuffling.max | time_stats.moe_shuffling.mean | time_stats.moe_shuffling.median | time_stats.moe_shuffling.std | time_stats.moe_grouped_gemm.min | time_stats.moe_grouped_gemm.max | time_stats.moe_grouped_gemm.mean | time_stats.moe_grouped_gemm.median | time_stats.moe_grouped_gemm.std | num_tokens | num_experts | num_experts_per_device | expert_parallel_size | routing_runtime_path | routing_assignment_policy | routing_weight_policy | routing_uses_router_logits | gating_runtime_context | gating_runtime_context_impl | router_topk | hidden_dim | expert_hidden_dim | use_gated | num_tensor_parallel_workers | total_routed_tokens | model_expansion_ratio | tokens_per_expert_avg | tokens_to_experts_ratio | expert_utilization | min_load_ratio | load_imbalance_cv | max_load_ratio | load_entropy | load_gini_coefficient | load_distribution | seed | moe_grouped_gemm_backend | measurement_type | profiling_precision | model_arch | quant_signature |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 0.03097599931061268 | 0.049056001007556915 | 0.03467839974910021 | 0.03254400007426739 | 0.005093522706269737 | 0.05193600058555603 | 0.08419200032949448 | 0.06054240055382252 | 0.05641600117087364 | 0.009051489911083033 | 0.025919999927282333 | 0.04064000025391579 | 0.030527999717742206 | 0.030608000233769417 | 0.0040231266205605675 | 0.29603201150894165 | 0.3494400084018707 | 0.3075023889541626 | 0.30137598514556885 | 0.014090820215642452 | 16 | 128 | 128 | 1 | standard_fused_topk | logit_topk | softmax_renorm | True | prefill_hot | ffn_like_prefix_20x | 8 | 4096 | 1536 | True | 4 | 128 | 0.375 | 1.0 | 1.0 | 0.609375 | 0.0 | 1.0231690964840563 | 4.0 | 6.122626857503489 | 0.5433349609375 | uniform | 0 | vllm_fused | CUDA_EVENT | BF16 | generic | method=fp8|act=dynamic|serialized=True|block=128x128 |