~5m15:57
Wanderloots

Run Your Own Agentic AI? ๐Ÿฆ™ Full Ollama Setup + Hermes Workflow

Jun 25, 2026

Read: ~5m ยท You save: 11 min

Run Your Own Agentic AI? Full Ollama Setup + Hermes Workflow

Unlock ultimate privacy and control! Learn to set up your own local AI with Ollama and connect it to agentic tools like Hermes. Run AI 100% offline, cost-free.

The ability to run a fully functional, private AI assistant locally on one's computer is now within reach, offering significant advantages in privacy, cost, and offline accessibility. This guide details the process of setting up Ollama, a user-friendly framework for local AI models, and demonstrates how to leverage these models with agentic AI tools like Hermes.

The Advantages of Local AI Models

Running AI models locally provides three primary benefits over cloud-based solutions:

  1. Privacy: Unlike cloud services where conversations are sent to external servers, local models process all data exclusively on the user's machine. This ensures sensitive information, client work, or personal writings remain entirely private and inaccessible to third parties.
  2. Cost: Cloud AI services typically operate on subscription models or per-use billing. Local models, once downloaded, run indefinitely with the only ongoing cost being electricity. This makes them a significantly more economical option for consistent use.
  3. Offline Accessibility: After the initial download, local models can be accessed and utilized from any location, regardless of internet connectivity.

These benefits are particularly impactful when integrated with personal knowledge management systems like Obsidian. Users can connect local models to their digital vaults to query notes, build personal wikis, or summarize research, all while maintaining complete data privacy.

Installing Ollama and Your First Model

Ollama simplifies the process of downloading and running local AI models. Installation can be performed via the terminal using a curl command or by downloading the installer directly from the Ollama website.

To verify the installation, users can type ollama version in their terminal. The output will display the installed version number.

Selecting and Downloading a Model

Ollama offers a wide array of models, with selection often dependent on available hardware. For users with 8 to 16 GB of RAM, Google's Gemma 4 model is a recommended option. Gemma 4 is available in various parameter sizes, including:

  • 31 billion parameters: Requires approximately 20 GB of storage and significant RAM (e.g., 32 GB).
  • 12 billion parameters: A more manageable option for systems with 32 GB RAM.
  • Effective Parameter (E4B) models: These models utilize per-layer embeddings to operate with the efficiency of larger models, offering a balance between performance and resource requirements. The E4B variant is suitable for systems with 8 GB of RAM.
  • E2B models: Designed for lower-end devices, including laptops with 4 GB of RAM or mobile phones.

To download a specific model, such as Gemma 4 E4B, the command ollama pull gemma:4b is used in the terminal. After the download completes, ollama list can confirm the model's presence.

Running a Model

Once a model is downloaded, it can be initiated with the command ollama run [model_name]. For example, ollama run gemma:4b will start the Gemma 4 E4B model. The model will then prompt the user for input, allowing for direct chat interactions. This process, from installation to a running model, can take approximately one minute.

For a more user-friendly interface, the Ollama desktop application provides a chat interface where conversations are stored locally.

Creating Customized Model Variants

Ollama allows for the creation of model variants, which are essentially modified configurations of existing models without duplicating the entire model file. This is particularly useful for adjusting parameters like the context window.

The default context window for many Ollama models is 2048 tokens. However, agentic AI tools like Hermes often require a much larger context window, such as 64,000 tokens, to effectively utilize their tool-accessing capabilities.

To create a variant with an extended context window, a model file can be generated. This involves specifying the base model and the desired context length. For instance, to create a Gemma 4 E4B variant with a 64,000 token context window, the command might involve pulling components from the existing model and establishing the new context.

After creation, ollama list will show the new variant, such as gemma:4b-64k. Running this variant, e.g., ollama run gemma:4b-64k, will utilize the expanded context. The file size will increase slightly (e.g., from 3.3 GB to 3.4 GB for Gemma 4 E4B), but the model's working memory is significantly enhanced.

Alternatively, the context length can be adjusted system-wide within Ollama's settings, setting a default context length for all launched models. This is a simpler approach if all models are intended to operate with the same extended context.

Connecting Local Models to Agentic AI Tools

The true power of local AI models is unlocked when they are connected to agentic AI frameworks. This connection is established through an "endpoint," which is a local address that directs the agent to the running model.

For Ollama models, the native endpoint format is http://localhost:11434. To make this compatible with platforms expecting an OpenAI endpoint, the format can be extended to http://localhost:11434/v1.

When configuring an agentic tool like Hermes:

  1. Navigate to the model settings within the agent's profile.
  2. Select the option to set up a custom endpoint.
  3. Input the appropriate Ollama endpoint URL (e.g., http://localhost:11434/v1).
  4. Connect the agent to the local model.

Once connected, the agent can utilize the local model. For example, when Hermes is prompted, it will spawn an agent within Ollama, load the specified model (e.g., Gemma 4E 64k), and process the request. This allows for private, locally-run AI interactions where the agent can harness the capabilities of the local LLM.

To ensure the model remains active for continuous agent operation, Ollama settings can be adjusted to prevent models from unloading after periods of inactivity. Furthermore, more powerful local models can be run on a desktop computer and accessed from a laptop, leveraging the computational resources of one machine for another.

Expanding Local AI Capabilities

The applications of local AI models extend beyond simple chat interactions. Users can:

  • Run embedding models: For building document retrieval systems.
  • Utilize vision and coding models: Replacing paid cloud services like Claude Code or Codex with local alternatives. Models like Qwen 3.6 are noted for their performance.
  • Integrate with Obsidian: For private knowledge management and research summarization.

By exploring and experimenting with local models, users can unlock a wide range of workflows that prioritize data privacy and autonomy.

This setup provides a foundation for fully private, locally run AI, enabling users to autonomously execute tasks and manage information without compromising data security. Further exploration into agentic AI playlists, particularly those using Hermes and other tools that connect to local models, can provide deeper insights into expanding these systems for practical applications.

Introduction to Local AI and Ollama

Introduction to local AI, highlighting its privacy, cost, and offline advantages over cloud-based models. The presenter, Callum (Water Lootz), introduces Ollama as an easy-to-use framework for running local models and outlines the video's content: reasons for local AI, Ollama installation, model setup, and agentic AI integration.

  • Local AI offers 100% privacy as all data stays on the user's computer.
  • Running local AI is cost-effective, with no subscription or per-use fees after initial setup.
  • Local AI models can be used offline once downloaded.
  • Ollama is presented as an easy-to-use framework for setting up local models.
  • The video will cover reasons for local AI, Ollama installation, model setup, and agentic AI connection.

Three Reasons to Run Your Own Local Model

Explains the three core benefits of running a local AI model: privacy (data never leaves your machine), cost (download once, run forever, negligible electricity cost), and offline functionality (usable anywhere without internet). It also touches on practical applications like integrating with note-taking apps (e.g., Obsidian) for a private 'second brain'.

  • Privacy: Conversations and data remain solely on the user's local machine.
  • Cost: Eliminates subscription or per-use fees associated with cloud AI services.
  • Offline Use: Models are accessible even without an internet connection.
  • Practical Use Case: Connecting local models to note-taking apps like Obsidian for private knowledge management.

Installing Ollama and Downloading Your First Model

Details the installation process for Ollama via terminal command or direct download. It covers verifying the installation with `ollama version` and then explores downloading models from the Ollama library, focusing on Gemma 4. It explains model variations (e.g., E4B, E2B) based on effective parameters and RAM requirements, and demonstrates pulling the Gemma 4 E4B model using `ollama pull Gemma 4 E4B`.

  • Ollama can be installed using a curl command in the terminal or via a direct download.
  • Installation can be verified with ollama version.
  • The Ollama library offers various models suitable for different hardware.
  • Gemma 4 is highlighted as a Google open-weight model.
  • Model sizes and RAM requirements vary (e.g., E2B for 4GB RAM, E4B for 8GB RAM, larger models for more RAM).
  • The command ollama pull Gemma 4 E4B is used to download the specified model.

Running and Interacting with Local Models

Demonstrates how to run the downloaded Gemma 4 E4B model using `ollama run Gemma 4 E4B` and interact with it. It also introduces the Ollama desktop app as an alternative interface for chatting and managing conversations locally. The section transitions to the next step: connecting local models to agentic AI tools.

  • The command ollama run Gemma 4 E4B starts an interactive chat session with the model.
  • The initial interaction confirms the model is Gemma 4, a large language model from Google DeepMind.
  • The Ollama desktop app provides a GUI for chatting and managing local model conversations.
  • The next step involves connecting local models to external agentic AI tools.

Creating Customized Model Variants (Extended Context)

Explains the concept of model variants in Ollama, which allow customization without duplicating model files. It focuses on extending the context window, crucial for agentic AI. The process involves creating a variant file (e.g., `Gemma 4E 64K`) using a specific command to increase the context length from 32,768 to 64,000 tokens, demonstrated with `ollama pull Gemma 4 E4B --context 64000` (implied by the variant creation). It also mentions system-level context settings in Ollama.

  • Model variants are configurations of existing models, not separate copies.
  • Extending the context window is important for agentic AI tools like Hermes.
  • A variant Gemma 4E 64K is created to increase the context window to 64,000 tokens.
  • This variant increases file size slightly (e.g., from 3.3GB to 3.4GB).
  • System-level settings in Ollama can also adjust the global context length.

Connecting Local Models to Agentic AI (Hermes Example)

Details how to connect a local Ollama model to an agentic AI tool like Hermes. This involves using a local endpoint URL (e.g., `http://localhost:11434/v1`) as a custom endpoint within the agent's settings. The video shows setting up Hermes to use the `Gemma 4e 64k latest` variant, enabling the agent to leverage the local model for private, powerful AI interactions.

  • Agentic AI tools (Hermes, Claude Code, Codex) can connect to local Ollama models.
  • A local endpoint URL (e.g., http://localhost:11434/v1) is used to connect external tools.
  • Hermes is configured with a custom endpoint pointing to the local Ollama instance.
  • The Gemma 4e 64k latest variant is selected within Hermes.
  • This connection allows agents to run locally, maintaining privacy and leveraging the local model's capabilities.

Advanced Configurations and Future Potential

Discusses advanced configurations like keeping models loaded (preventing them from unloading after inactivity) and leveraging powerful desktop models from laptops. It reiterates the value of local AI for managing sensitive information (e.g., in Obsidian) and explores the potential of various model types (embedding, vision, coding) for different applications, encouraging experimentation.

  • Ollama settings can be modified to keep models loaded, ensuring they are always ready.
  • Powerful local models on desktops can be accessed remotely from laptops.
  • Local AI enhances privacy for sensitive data management, especially with tools like Obsidian.
  • Various model types exist, including embedding, vision, and coding models.
  • Running local models unlocks diverse workflows and applications, promoting privacy and cost-effectiveness.