Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren and its primary use cases
  • Hardware layout: cores, memory, and compute clusters
  • Comparative analysis with NVIDIA and AMD GPUs

Establishing the Biren Programming Environment

  • Installation of the Biren SDK and runtime components
  • Navigating the toolchain and compiler model
  • Fundamental project structures and build workflows

GPU Programming with the Biren Stack

  • Thread and block models
  • Memory management and data transfer mechanisms
  • Kernel development and launch patterns

Transitioning from CUDA to Biren

  • Techniques for translating CUDA code
  • Mapping common APIs and necessary adaptations
  • Hands-on labs and practice in code conversion

Debugging and Profiling

  • Utilising Biren’s debugger and profiler tools
  • Identifying and analysing performance bottlenecks
  • Optimising memory access patterns

Advanced Optimisation Techniques

  • Thread scheduling and instruction pipelining
  • Loop unrolling and efficient shared memory usage
  • Advanced kernel tuning for maximum throughput

Case Studies and Application Examples

  • Training models using Biren accelerators
  • Porting and profiling vision or NLP models
  • Performance benchmarking against CUDA/NVIDIA platforms

Summary and Future Directions

Requirements

  • A solid grasp of GPU architecture and parallel processing concepts
  • Practical experience with CUDA, OpenCL, or comparable GPU programming environments
  • Working knowledge of deep learning frameworks such as PyTorch or TensorFlow

Target Audience

  • HPC developers
  • AI infrastructure engineers
  • Performance optimisation specialists
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Provisional Upcoming Courses (Require 5+ participants)

Related Categories