Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Overview of the Chinese AI GPU Ecosystem
- Comparison of Huawei Ascend, Biren, and Cambricon MLU
- CUDA vs CANN, Biren SDK, and BANGPy programming models
- Industry trends and vendor ecosystem developments
Preparing for Migration
- Evaluating your existing CUDA codebase
- Identifying target platforms and compatible SDK versions
- Installing the toolchain and setting up the development environment
Code Translation Techniques
- Translating CUDA memory access patterns and kernel logic
- Mapping compute grid and thread models
- Exploring automated versus manual translation approaches
Platform-Specific Implementations
- Leveraging Huawei CANN operators and custom kernels
- Utilising the Biren SDK conversion pipeline
- Rebuilding models using BANGPy (Cambricon)
Cross-Platform Testing and Optimization
- Profiling execution on each target platform
- Tuning memory usage and comparing parallel execution performance
- Monitoring performance metrics and iterative refinement
Managing Mixed GPU Environments
- Implementing hybrid deployments across multiple architectures
- Developing fallback strategies and device detection mechanisms
- Establishing abstraction layers to ensure code maintainability
Case Studies and Best Practices
- Porting vision and NLP models to Ascend or Cambricon
- Retrofitting inference pipelines on Biren clusters
- Mitigating version mismatches and API discrepancies
Summary and Next Steps
Requirements
- Practical experience programming with CUDA or GPU-based applications
- A solid understanding of GPU memory models and compute kernels
- Familiarity with AI model deployment or acceleration workflows
Audience
- GPU programmers
- System architects
- Porting specialists
21 Hours