Multimodal LLM Workflows in Vertex AI Training Course
Vertex AI provides powerful tools for building multimodal LLM workflows that integrate text, audio, and image data into a single pipeline. With long context window support and Gemini API parameters, it enables advanced applications in planning, reasoning, and cross-modal intelligence.
This instructor-led, live training (online or onsite) is aimed at intermediate to advanced-level practitioners who wish to design, build, and optimize multimodal AI workflows in Vertex AI.
By the end of this training, participants will be able to:
- Leverage Gemini models for multimodal inputs and outputs.
- Implement long-context workflows for complex reasoning.
- Design pipelines that integrate text, audio, and image analysis.
- Optimize Gemini API parameters for performance and cost efficiency.
Format of the Course
- Interactive lecture and discussion.
- Hands-on labs with multimodal workflows.
- Project-based exercises for applied multimodal use cases.
Course Customization Options
- To request a customized training for this course, please contact us to arrange.
Course Outline
Introduction to Multimodal LLMs in Vertex AI
- Overview of multimodal capabilities in Vertex AI
- Gemini models and supported modalities
- Use cases in enterprise and research
Setting Up the Development Environment
- Configuring Vertex AI for multimodal workflows
- Working with datasets across modalities
- Hands-on lab: environment setup and dataset preparation
Long Context Windows and Advanced Reasoning
- Understanding long-context workflows
- Use cases in planning and decision-making
- Hands-on lab: implementing long-context analysis
Cross-Modal Workflow Design
- Combining text, audio, and image analysis
- Chaining multimodal steps in pipelines
- Hands-on lab: designing a multimodal pipeline
Working with Gemini API Parameters
- Configuring multimodal inputs and outputs
- Optimizing inference and efficiency
- Hands-on lab: tuning Gemini API parameters
Advanced Applications and Integrations
- Interactive multimodal agents and assistants
- Integrating external APIs and tools
- Hands-on lab: building a multimodal application
Evaluation and Iteration
- Testing multimodal performance
- Metrics for accuracy, alignment, and drift
- Hands-on lab: evaluating multimodal workflows
Summary and Next Steps
Requirements
- Proficiency in Python programming
- Experience with machine learning model development
- Familiarity with multimodal data (text, audio, image)
Audience
- AI researchers
- Advanced developers
- ML scientists
Need help picking the right course?
macao@nobleprog.com or +852 81990613