Welcome to this hands-on lab on running Large Language Models (LLMs)! In this lab, you'll learn how to manage and run LLMs both locally and through cloud services. You'll gain practical experience with tools like Ollama and Podman AI for local model management, as well as explore cloud-based LLM services from Groq and Mistral AI.
Before starting the lab, please ensure you have the following installed:
- httpie for making API calls.
By the end of this lab, you will be able to:
- Manage and run LLMs locally using Ollama
- Manage and run LLMs locally using Podman AI
- Consume LLMs from Groq's cloud service
- Consume LLMs from Mistral AI's cloud service
Ollama is a convenient tool to manage and run large language models. Think of it as the Docker for LLMs - just as Docker manages containers, Ollama manages models. With Ollama, you can access a wide library of models to run on your local infrastructure or in the cloud, maintaining full control of your data and ensuring privacy.
- Visit https://ollama.com/download
- Download and install the appropriate version for your operating system.
- Verify the installation by running the following command in your terminal:
ollama --versionOpen a terminal and run the following command to download the Mistral model:
ollama pull mistralThis will download and prepare the open-source Mistral model for use.
Once the download is complete, you can run the model and start chatting with it via the interactive console in your Terminal:
ollama run mistralTry asking the model a question, for example: "Why is a raven like a writing desk?".
Ollama supports running multiple models simultaneously. You can monitor all the models loaded in memory and check whether they are running on CPU or GPU:
ollama psWhen you're done chatting with the model, exit the interactive console by typing:
/byeYou can list all the available models with the following command:
ollama lsTry pulling a different model. For example, qwen2.5 is a high-quality model from Alibaba:
ollama pull qwen2.5The Ollama library includes both open-source and free models. You can get information about a model's license, architecture, and system prompt as follows:
ollama show qwen2.5When you're done using a model, you can remove it to free up space:
ollama rm qwen2.5Podman AI is another tool for running AI models locally, offering a graphical interface for managing and interacting with various LLMs. It's an extension to Podman Desktop, an open-source solution for managing and running containers.
Visit https://podman-desktop.io/docs/installation and follow the instructions to install Podman Desktop for your operating system.
Then, verify the installation by opening Podman Desktop.
Visit https://podman-desktop.io/docs/ai-lab/installing and follow the instructions to install the Podman AI extension.
Then, verify that the setup is complete by checking for the AI Lab section in Podman Desktop.
- In Podman Desktop, navigate to the AI Lab section and go to Models > Catalog.
- Browse the available models and select one to download, for example
instructlab/granite-7b-lab-GGUF. - Once downloaded, go to Models > Services > New Model Service.
- Select the model you have previously downloaded and click Create service.
- After the model is loaded, Podman AI will provide curl instructions to call the model.
- Try asking the model a question, such as: "Why is a raven like a writing desk?"
Groq provides access to high-performance LLMs through their cloud API, offering an OpenAI-compatible interface. This allows you to leverage powerful models without the need for local hardware resources.
Visit https://console.groq.com/ and sign up for a new account (you can sign up with your GitHub account for a quick setup). Choose the "Free" plan, which gives you access to the Groq APIs with low rate limits at no cost.
In the Groq console, navigate to API Keys and generate a new API key. Copy and securely store your API key on your laptop. For example, set it as an environment variable:
export GROQ_API_KEY=<your-api-key>If you plan to use Groq with the Spring AI OpenAI integration, also store the API key as follows:
export SPRING_AI_OPENAI_API_KEY=${GROQ_API_KEY}Open a Terminal window and use httpie to make your first API call to Groq and interact with the llama3.1 model:
http POST https://api.groq.com/openai/v1/chat/completions \
Authorization:"Bearer $GROQ_API_KEY" \
messages:='[{"role": "user", "content": "Why is a raven like a writing desk?"}]' \
model=llama-3.1-8b-instantMistral AI provides access to their open-source and proprietary LLMs through their cloud API, offering another option for leveraging powerful language models.
Visit https://console.mistral.ai/ and sign up for a new account. Choose the "Experiment" plan, which gives you access to the Mistral APIs for free.
In the Mistral AI console, navigate to API Keys and generate a new API key. Copy and securely store your API key on your laptop. For example, set it as an environment variable:
export MISTRAL_AI_API_KEY=<your-api-key>If you plan to use Mistral AI with the Spring AI Mistral AI integration, also store the API key as follows:
export SPRING_AI_MISTRALAI_API_KEY=${MISTRAL_AI_API_KEY}Open a Terminal window and use httpie to make your first API call to Mistral AI and interact
with the mistral-small model:
http POST https://api.mistral.ai/v1/chat/completions \
Authorization:"Bearer $MISTRAL_AI_API_KEY" \
messages:='[{"role": "user", "content": "Why is a raven like a writing desk?"}]' \
model=mistral-small-latestCongratulations! You've completed the lab on running Large Language Models. You've learned how to manage and run LLMs locally using Ollama and Podman AI, as well as how to consume cloud-based LLM services from Groq and Mistral AI. These skills provide a solid foundation for integrating AI capabilities into your projects and applications.