Get in Touch

Course Outline

Overview of the Chinese AI GPU Ecosystem

  • Comparison of Huawei Ascend, Biren, and Cambricon MLU
  • CUDA vs CANN, Biren SDK, and BANGPy programming models
  • Industry trends and vendor ecosystem developments

Preparing for Migration

  • Evaluating your existing CUDA codebase
  • Identifying target platforms and compatible SDK versions
  • Installing the toolchain and setting up the development environment

Code Translation Techniques

  • Translating CUDA memory access patterns and kernel logic
  • Mapping compute grid and thread models
  • Exploring automated versus manual translation approaches

Platform-Specific Implementations

  • Leveraging Huawei CANN operators and custom kernels
  • Utilising the Biren SDK conversion pipeline
  • Rebuilding models using BANGPy (Cambricon)

Cross-Platform Testing and Optimization

  • Profiling execution on each target platform
  • Tuning memory usage and comparing parallel execution performance
  • Monitoring performance metrics and iterative refinement

Managing Mixed GPU Environments

  • Implementing hybrid deployments across multiple architectures
  • Developing fallback strategies and device detection mechanisms
  • Establishing abstraction layers to ensure code maintainability

Case Studies and Best Practices

  • Porting vision and NLP models to Ascend or Cambricon
  • Retrofitting inference pipelines on Biren clusters
  • Mitigating version mismatches and API discrepancies

Summary and Next Steps

Requirements

  • Practical experience programming with CUDA or GPU-based applications
  • A solid understanding of GPU memory models and compute kernels
  • Familiarity with AI model deployment or acceleration workflows

Audience

  • GPU programmers
  • System architects
  • Porting specialists
 21 Hours

Number of participants


Price per participant

Provisional Upcoming Courses (Require 5+ participants)

Related Categories