Implementing MinerU as a Custom Backend in GPUStack for PDF Document Processing

Implementing MinerU as a Custom Backend in GPUStack for PDF Document Processing

GPUStack v2 introduces the custom backend functionality, enabling integration of any model inference engine alongside natively supported options like vLLM and SGLang for unified management and scheduling.

This guide demonstrates how to deploy MinerU, a specialized PDF document extraction tool, within GPUStack to establish an efficient document parsing service.

Understanding MinerU

MinerU is an open-source solution designed for parsing complex PDF documents, converting files containing formulas, tables, and other intricate elements into Markdown format. Through GPUStack's custom backend feature, it can be configured as a private document parsing API service.

Preparing the MinerU Environment

To expedite deployment, a pre-packaged image (v2.7.0) is available for direct use:

docker pull swr.cn-south-1.myhuaweicloud.com/gpustackcommunity/mineru:v2.7.0

Note: The official Dockerfile utilizes an older vLLM version. For newer hardware such as RTX 5090, modify the Dockerfile to upgrade vLLM and rebuild the image.

Backend Registration Process

After preparing the image, register it as a new backend in GPUStack:

  1. Access the GPUStack management interface
  2. Navigate to Inference Backends via the sidebar
  3. Select Add Backend and configure with the following parameters:

Alternatively, use this YAML configuration:

backend_name: MinerU-custom
default_run_command: mineru-vllm-server --port {{port}} --served-model-name {{model_name}}
version_configs:
  v2.7.0:
    image_name: swr.cn-south-1.myhuaweicloud.com/gpustackcommunity/mineru:v2.7.0
    custom_framework: cuda
default_version: v2.7.0

The configuration interface allows switching to YAML mode for direct code entry:

Model Deployment and Backend Binding

With the backend registered, deploy MinerU following standard LLM deployment procedures:

  1. Access the Deployments section
  2. Initiate a new deployment with the MinerU backend

The image includes pre-installed model weights, using /tmp as a placeholder path. This may affect VRAM resource estimation accuracy, so manual scheduling mode is recommended.

The mineru-vllm-server implementation, based on vllm serve, supports parameters like --gpu-memory-utilization for precise resource control.

Important: While MinerU offers OpenAI API compatibility, it is specialized for document processing only and lacks general conversation capabilities. It should not be used as a general-purpose chat model.

Upon successful deployment, the status will indicate "Running". Verify the startup by checking the logs.

Utilizing MinerU for Document Processing

With GPUStack configured for MinerU, document parsing can be performed using the built-in CLI tool.

API Access Configuration

Retrieve API access details from the Deployments page by accessing the model's options menu and selecting "API Access Info".

CLI Implementation Example

# Configure model identifier
export DOC_PARSER_MODEL=mineru
# Replace with your actual API key
export DOC_PARSER_API_KEY=gpustack_3519fc0369a06fae_c434118c1cc0e07f3dfe998c6416522c
# Execute PDF parsing
mineru -p sample.pdf -o results -b vlm-http-client -u http://192.168.50.12

Extending Functionality

GPUStack's custom backend architecture supports diverse integration possibilities:

  • Explore additional community backends: https://github.com/gpustack/community-inference-backends
  • Contribute custom backend configurations through pull requests

Tags: GPUStack MinerU PDF Processing Custom Backend Document Parsing

Posted on Sun, 04 Oct 2026 16:16:18 +0000 by Joshv1288