Skip to main content

Build a multi-LLM chat application with Azure Container Apps

April 11, 2024 | 3:30 - 3:45 PM (UTC) Coordinated Universal Time

Ready to get started with AI and the latest technologies? Microsoft Reactor provides events, training, and community resources to help developers, entrepreneurs and startups build on AI technology and more. Join us!

Build a multi-LLM chat application with Azure Container Apps

April 11, 2024 | 3:30 - 3:45 PM (UTC) Coordinated Universal Time

Ready to get started with AI and the latest technologies? Microsoft Reactor provides events, training, and community resources to help developers, entrepreneurs and startups build on AI technology and more. Join us!

Go back

Build a multi-LLM chat application with Azure Container Apps

April 11, 2024 | 3:30 - 3:45 PM (UTC) Coordinated Universal Time

  • Format:
  • alt##LivestreamLivestream

Topic: Coding, Languages, and Frameworks

Language: English

In this demo, explore how to leverage GPU workload profiles in ACA to run your own model backend, and easily switch, compare, and speed up your inference times. You will also explore how to leverage LlamaIndex (https://github.com/run-llama/llama_index) to ingest data on-demand, and host models using Ollama (https://github.com/ollama/ollama). Then finally, decompose the application as a set of microservices written in Python, deployed on ACA.

  • Azure

Speakers