Get in Touch
 Duration 21 hours

Course Outline

Introduction to Multimodal AI and Ollama

  • An overview of multimodal learning principles
  • Key challenges in integrating vision and language
  • Understanding Ollama's capabilities and architecture

Setting Up the Ollama Environment

  • Installing and configuring the Ollama environment
  • Managing local model deployment
  • Connecting Ollama with Python and Jupyter notebooks

Handling Multimodal Inputs

  • Integrating text and image data
  • Incorporating audio and structured data formats
  • Designing effective preprocessing pipelines

Document Understanding Applications

  • Extracting structured insights from PDFs and images
  • Enhancing language models with OCR technology
  • Creating intelligent workflows for document analysis

Visual Question Answering (VQA)

  • Preparing VQA datasets and evaluation benchmarks
  • Training and assessing multimodal models
  • Developing interactive VQA applications

Designing Multimodal Agents

  • Core principles of agent design with multimodal reasoning
  • Synthesising perception, language, and action
  • Deploying agents for practical real-world scenarios

Advanced Integration and Optimisation

  • Fine-tuning multimodal models using Ollama
  • Optimising inference performance
  • Considerations for scalability and deployment

Summary and Next Steps

Requirements

  • A solid grasp of core machine learning concepts
  • Hands-on experience with deep learning frameworks like PyTorch or TensorFlow
  • Knowledge of natural language processing and computer vision principles

Target Audience

  • Machine Learning Engineers
  • AI Researchers
  • Product Developers integrating vision and text workflows

Number of participants


Price per participant

Provisional Upcoming Courses (Require 5+ participants)

Related Categories