Deploying DeepSeek-V4 on Ascend 910B Clusters Using GPUStack: Configuration and Benchmark Analysis

The DeepSeek-V4 architecture leverages a Mixture-of-Experts (MoE) design, offering variants like the 284B Flash and 1.6T Pro models. During inference, only a fraction of parameters are activated, balancing computational overhead with output quality. The integration of expanded context windows and refined attention mechanisms improves performance in long-document processing, multi-step reasoning, and autonomous agent workflows.

Architectural enhancements include a hybrid attention framework (CSA+HCA) that reduces long-context overhead, an mHC structure for deeper network stability, and the Muon optimizer for accelerated training convergence. Running these large-scale MoE systems efficiently requires tight coupling between inference frameworks and underlying accelerators. On domestic AI hardware, achieving stable throughput demands coordinated optimization across the driver stack, container runtime, and serving engine. The following workflow demonstrates how to orchestrate DeepSeek-V4 on Ascend 910B accelerators using GPUStack, including configuration steps and benchmark metrics.

GPUStack Control Plane Initialization

GPUStack operates as a containerized orchestration layer for AI workloads, supporting pluggable backends like vLLM, SGLang, and TensorRT-LLM. It handles heterogeneous accelerator pooling, automated failover, and traffic routing. Before provisioning model endpoints, the central management server must be deployed and linked to compute nodes.

Container Runtime Preparation

Ensure a compatible container engine (Docker, Podman, or K8s) is active across all target hosts. Verify the daemon status:

docker version --format '{{.Server.Version}}'

Launching the Management Server

The control plane does not require accelerator access and can run on standard CPU instances. For this setup, the server is co-located on an 8-card Ascend 910B2 host:

sudo docker run -d \
  --name gs-controller \
  --restart always \
  -p 8080:80 \
  -v gs-persistent-storage:/var/lib/gpustack \
  swr.cn-south-1.myhuaweicloud.com/gpustack/gpustack:v2.1.2 \
  --log-level debug \
  --bootstrap-password SecureAdminPass_2024

Key configuration flags:

  • -p 8080:80: Maps the dashboard to host port 8080.
  • -v: Mounts a named volume for persistent state (API keys, telemetry, routing tables).
  • --bootstrap-password: Sets the initial administrator credential.
  • --log-level debug: Enables verbose output for troubleshooting.

Monitor initialization progress:

docker logs --follow gs-controller

Access the dashboard at http://<host-ip>:8080 using admin and the configured password. Create a new Docker-managed cluster to register downstream compute resources.

Integrating Ascend NPU Compute Nodes

Worker nodes require specific driver and runtime configurations before joining the cluster.

1. Accelerator Driver Verification

Check the installed NPU driver release:

npu-smi info | head -n 5

A driver version of 25.5 or higher is recommended to ensure full compatibility with MoE routing kernels.

2. Ascend Container Runtime Validation

Confirm that the Docker daemon recognizes the Ascend runtime extension:

if docker info 2>/dev/null | grep -qi "ascend"; then
    echo "Runtime extension detected"
else
    echo "Missing Ascend runtime configuration"
    exit 1
fi

If the check fails, install the Ascend Docker Runtime package (e.g., v7.3.1) to enable hardware passthrough for containers.

3. Node Registration

Generate a worker join token from the GPUStack dashboard. Execute the provided command on the target host. This spins up a worker agent container that automatically reports hardware telemetry (index, vendor, thermal data, memory allocation) back to the control plane.

Verify agent connectivity:

docker logs --follow gpustack-worker

The dashboard should reflect a Ready status for the newly added node.

Registering Custom Inference Backends

GPUStack allows operators to inject custom backend images and execution templates. To serve DeepSeek-V4, register the Ascend-optimized vLLM release v0.13.0rc3.

Navigate to the Inference Backends section, edit the vLLM entry, and append a new version profile:

Parameter Value
Version Label 0.13.0rc3-ascend
Container Image quay.io/ascend/vllm-ascend:v0.13.0rc3
Accelerator Framework CANN
Entrypoint Override vllm serve
Execution Template {{model_path}} --host {{worker_ip}} --port {{port}} --served-model-name {{model_name}}

Alternatively, apply the configuration via YAML import:

backend_name: vLLM
version_configs:
  0.13.0rc3-ascend:
    image_name: quay.io/ascend/vllm-ascend:v0.13.0rc3
    entrypoint: vllm serve
    run_command: "{{model_path}} --host {{worker_ip}} --port {{port}} --served-model-name {{model_name}}"
    custom_framework: cann
    env: {}

Ensure worker nodes have network access to quay.io or pre-load the image locally and adjust the image_name field accordingly. Template variables enclosed in {{}} must remain unchanged for dynamic injection.

Provisioning the DeepSeek-V4 Endpoint

Official serving guidelines for Ascend hardware are available in the vLLM documentation. The following steps outline the deployment of the DeepSeek-V4-Flash-w8a8-mtp variant via GPUStack.

Model Source Configuration:

  • Online: Use the dashboard's ModelScope integration to search and pull Eco-Tech/DeepSeek-V4-Flash-w8a8-mtp.
  • Air-gapped: Pre-download weights, distribute them to worker hosts, mount the directory into the worker container, and register the path via Model Files > Local Path. Use the container-internal mount path during deployment.

Backend Assignment & Resource Allocation:

  • Engine: vLLM
  • Version: 0.13.0rc3-ascend
  • Accelerators: 8x Ascend 910B2 (64GB)

Apply the following runtime arguments and environment variables. Adjust parallelism flags based on available hardware topology:

# Engine Arguments
--gpu-memory-utilization 0.9
--max-model-len 65536
--max-num-batched-tokens 8192
--max-num-seqs 16
--data-parallel-size 1
--tensor-parallel-size 8
--enable-expert-parallel
--quantization ascend
--block-size 128
--async-scheduling
--chat-template /var/lib/gpustack/cache/model_scope/Eco-Tech/DeepSeek-V4-Flash-w8a8-mtp/chat_template.jinja
--additional-config '{"enable_cpu_binding": "true", "multistream_overlap_shared_expert": true}'
--speculative-config '{"num_speculative_tokens": 1, "method": "deepseek_mtp"}'
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}'

# Environment Variables
USE_MULTI_BLOCK_POOL=1
OMP_PROC_BIND=false
OMP_NUM_THREADS=10
PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
ACL_OP_INIT_MODE=1
TRITON_ALL_BLOCKS_PARALLEL=1

Monitor the provisioning logs via the dashboard. Once the instance transitions to Running, the endpoint is ready for traffic.

Performance Evaluation

Initial single-request testing yields approximately 31 tokens per second. Given the early-stage integration of MoE routing on this hardware generation, further kernel-level optimizations are expected to improve latency.

Using the integrated benchmarking module, a throughput-oriented stress test was executed against the deployed endpoint. The system maintained stable memory utilization and consistent token generation rates under concurrent load, validating the configuration for production-grade inference workloads.

Tags: DeepSeek-V4 Ascend 910B GPUStack vLLM MoE Deployment

Posted on Tue, 29 Sep 2026 16:35:28 +0000 by Zamees