~6m26:25Improved AI Memory? 🧠 Full Hermes Tutorial (Mnemosyne & Hindsight)
Aug 7, 2026
Read: ~6m · You save: 20 min
Improved AI Memory? Full Hermes Tutorial (Mnemosyne & Hindsight)
Tired of AI forgetting? Unlock true agent memory with Hermes! Learn to set up Mnemosyne & Hindsight for powerful, persistent AI recall. Full tutorial!
The ability of artificial intelligence agents to retain and utilize information is a critical factor in their effectiveness, particularly in complex, ongoing tasks. Traditional AI interactions often suffer from a lack of persistent memory, requiring users to repeatedly provide context and information, leading to wasted resources and diminished performance. This article explores advanced memory systems for AI agents, focusing on two primary solutions: Mnemosyne, a lightweight memory layer, and Hindsight, a more robust memory engine. These systems aim to provide agents with "true memory," enabling them to learn, adapt, and recall information across sessions and interactions.
Understanding the Agentic Memory Stack
Agentic memory is not a single component but rather a stack comprising three distinct elements: world knowledge, built-in memory, and external memory.
- World Knowledge: This layer, exemplified by knowledge bases like Obsidian vaults or LLM wikis, serves as a repository of general information applicable across all projects and agents. It represents the ground truth for an AI system.
- Operational Memory: This tier focuses on information relevant to the agent's specific interactions with a user. It is further divided into:
- Built-in Memory: This is the default memory system within an agent, typically consisting of files like
memory.md,user.md, andsoul.mdfor operational memory, user identity, and personality, respectively. It also includes session search capabilities stored in a local database. While effective for single-session chats, it consumes context and tokens in every interaction. - Dedicated Memory Providers: These are external systems designed to store and retrieve facts at runtime, offering improved recall and performance for complex projects without inflating every conversation's context.
- Built-in Memory: This is the default memory system within an agent, typically consisting of files like
The article focuses on enhancing operational memory through dedicated memory providers, with a subsequent discussion planned for integrating world knowledge via an Obsidian LLM Wiki.
Built-in Memory in Hermes
The Hermes agent utilizes a built-in memory system comprising three core files:
memory.md: Stores operational memory.user.md: Defines the user's identity and preferences.soul.md: Contains the agent's personality traits.
Additionally, Hermes incorporates session search functionality via a local database, allowing agents to query past interactions. A key recommendation for optimizing Hermes' performance is to provide it with initial context about the user, enhancing its understanding and personalization from the outset. Importantly, running Hermes within a Docker container does not impede its access to these memory files, as this access is managed through a memory tool, independent of the sandboxed code execution. The built-in memory files are injected into every session, ensuring the agent consistently remembers user context. However, this constant injection consumes valuable tokens, highlighting the need for more efficient external memory solutions.
Memory Providers: Mnemosyne and Hindsight
Choosing a memory provider can be daunting due to the variety of options and considerations, including user goals, the number of agents involved, hardware capabilities, and cost. Fortunately, most providers allow for migration, reducing the risk of vendor lock-in. This article highlights two community-recommended options: Mnemosyne and Hindsight.
Mnemosyne: A Lightweight Memory Layer
Mnemosyne is a zero-dependency, lightweight memory layer that operates entirely locally without requiring an LLM. It features built-in embeddings and offers sub-millisecond query times, making it an ideal choice for users seeking a fast, unobtrusive memory solution.
Architecture and Functionality: Mnemosyne is inspired by the BEAM architecture, incorporating different memory tiers (working, episodic, semantic, scratchpad) optimized for various access patterns. It ensures privacy by keeping all data local and boasts native integration with Hermes. The system functions by receiving agent queries, using a recall tool to access its context, and injecting relevant information into the agent's prompts before a response is generated. As conversations progress, Mnemosyne learns and builds context, contributing to a self-evolving memory system. This layer is compatible with various agents, including Claude, Codex, and Cursor, and can support shared memory banks across multiple agents.
Installation and Setup:
Installing Mnemosyne requires a manual approach, particularly when Hermes is running within a Docker container. The process involves installing Mnemosyne into its own pipx environment to ensure it persists outside of Hermes' rebuildable virtual environments. The recommended steps include:
- Installing Mnemosyne Hermes into a separate
pipxenvironment. - Configuring the environment.
- Linking the new environment to the Hermes plugins folder.
- Running
Hermes memory setupto configure Mnemosyne as a memory provider.
This setup modifies the config.yaml file to designate Mnemosyne as the default memory provider. After restarting the Hermes Gateway and verifying the plugin status, users can run Hermes Mnemosyne stats to view memory usage. It is also recommended to disable legacy memory features in config.yaml (setting memory enabled: false and user profile enabled: false) to prevent duplication and token waste.
Testing Mnemosyne:
Initial testing involves asking Hermes a question without prior context. After providing personal information (name, alias, professional background, preferences), Hermes utilizes Mnemosyne's remember tool to store this data. Subsequent queries, such as "Who am I?", trigger Mnemosyne's recall function, demonstrating successful memory retrieval. The Hermes Mnemosyne stats command reveals the amount of working memory being consolidated, illustrating the BEAM architecture's process of shifting information from working to episodic memory. Mnemosyne also stores structured knowledge as subject-predicate-object triples, forming the basis of a knowledge graph.
Hindsight: A Powerful Memory Engine
Hindsight is described as a memory engine with advanced natural language processing and multi-signal retrieval capabilities. It offers a more comprehensive feature set than Mnemosyne, including the ability to build mental models, process observations, and engage in reflective reasoning.
Architecture and Functionality: Hindsight operates by connecting the agent to a separate server, which can be cloud-based or run locally. Its "reflect" function initiates an agentic loop that searches memory, applies findings to shape reasoning, and produces synthesized responses grounded in retrieved information. This process requires an LLM and is more complex than Mnemosyne's layer-based approach.
Installation and Setup: Hindsight can be configured in two modes:
- Cloud Mode: Requires an API key for Hindsight's cloud service, which may incur fees based on usage.
- Local External Mode: Allows users to run the system on their own hardware. This mode necessitates a local LLM that supports tool calling. Options include using services like Ollama with models such as GPT-OSS (20 billion parameters, 13 GB) or connecting to APIs from OpenAI, Anthropic, or Gemini, which incur per-token costs.
For local deployment, Hindsight can be run within a Docker container for enhanced safety and persistence. This involves configuring environment variables to connect to the chosen local LLM and specifying ports for the server and dashboard. Once the Hindsight server is running, it needs to be connected to Hermes.
Connecting Hindsight to Hermes:
Within the Hermes desktop app, users select the Hindsight profile and navigate to memory settings. The mode is set to "local external," and the local host server address (e.g., http://localhost:9999) is provided. A recall budget is also set. For testing, persistent memory in Hermes can be temporarily disabled.
Testing Hindsight:
When asked "Who am I?" without prior context, Hindsight, with persistent memory disabled, initially does not recall the information. However, by providing personal details, Hindsight's retain tool stores this information. A subsequent query of "Who am I?" triggers Hindsight's recall function, successfully retrieving the stored data. The Hindsight dashboard (accessible via localhost:9999) provides a visual representation of the stored memories, including a constellation view, table view, and timeline. This dashboard allows users to explore experiences, observations, and mental models. The "reflect" function can be used to generate synthesized responses based on the stored memory, demonstrating a deeper level of reasoning.
Mnemosyne vs. Hindsight: Choosing the Right Solution
The choice between Mnemosyne and Hindsight depends on specific needs and priorities.
- Mnemosyne: Optimized for simplicity, speed, and single-machine deployments. It functions as a memory layer, ideal for users prioritizing local operation and minimal dependencies.
- Hindsight: Acts as a memory engine with sophisticated NLP and multi-signal retrieval. It offers more advanced features like reflection and mental model building but requires more complex setup, potentially involving external servers or LLMs.
The creators suggest these are not direct competitors but rather complementary solutions. Users are encouraged to test providers for a week to assess their impact on workflow and migrate if necessary, as there is no vendor lock-in. The Hindsight dashboard offers a more integrated visualization of memory compared to Mnemosyne's community-driven dashboard.
The next step in enhancing agent capabilities involves connecting these operational memory systems to world knowledge, such as an Obsidian LLM Wiki, to enable agents to reason against a broader base of information.
Introduction: The Need for AI Memory
Introduction to the problem of AI forgetting context and the solution: implementing a memory system. The video will cover different memory tiers, providers, and setup for Hermes and other agents.
- AI agents often forget context, leading to wasted tokens and time.
- A universal memory system can be set up once and connected to any agent.
- The video will explain memory stack tiers, agentic memory vs. knowledge layers (Obsidian), dedicated memory providers, and installation of lightweight and heavyweight memory solutions.
- The focus is on giving agents 'true memory'.
Understanding the Agentic Memory Stack
Clarifies the concept of agentic memory, explaining it as a 'stack' rather than a single layer. It distinguishes between world knowledge (like Obsidian), built-in memory, and external memory providers, defining their roles in an agent's operational capacity.
- Agentic memory is a stack, not a layer, comprising world knowledge, built-in memory, and external memory.
- Knowledge layers (e.g., Obsidian, LLM Wiki) store world knowledge, accessible across projects and agents.
- Operational memory splits into built-in memory (single session) and dedicated memory providers (complex projects, better recall).
- Built-in memory is for simple chats; dedicated providers are for complex projects and better performance.
- Obsidian is for world knowledge spanning all projects; agentic memory is for facts about working with the user or current tasks.
Hermes Built-in Memory vs. External Providers
Details Hermes' built-in memory, including its files (memory.md, user.md, soul.md) and session search capabilities. It explains how these files are injected into every session, consuming tokens, and introduces external memory providers as a solution to token waste.
- Hermes' built-in memory includes memory.md (operational), user.md (user info), and soul.md (personality).
- Session search is stored in a local database for past session retrieval.
- Providing context about the user upfront significantly improves the system's intelligence.
- Running Hermes in Docker does not prevent access to its memory files.
- Built-in memory files are injected into every chat session, consuming tokens.
- External memory providers store and retrieve facts at runtime, reducing token usage.
Setting Up Mnemosyne: The Lightweight Memory Layer
Introduces Mnemosyne as a lightweight, zero-dependency, local memory layer for AI agents. It explains its biological inspiration, local operation, fast queries, and native Hermes integration, detailing the installation process and configuration steps.
- Mnemosyne is a universal memory layer for any AI agent, using a SQLite backend with sub-millisecond queries and zero dependencies.
- It's inspired by the BEAM architecture, with memory tiers like working, episodic, semantic, and scratchpad.
- Mnemosyne operates locally for privacy, offers fast queries, and has native Hermes integration.
- Auto context injection: Mnemosyne receives queries, retrieves context, and injects it into agent prompts.
- Installation requires manual steps outside of Hermes' managed virtual environment, often using pipx.
- Configuration involves modifying
config.yamland runningHermes memory setup. - Key setting: switching the default scope to 'global' for cross-session memory.
Testing and Configuring Mnemosyne
Explains how to set up and test Mnemosyne, including verifying its status, disabling legacy memory to avoid duplication, and providing initial context to the agent. It demonstrates how the agent recalls information after it's stored in Mnemosyne.
- After installation, Mnemosyne needs to be enabled as the default memory provider in
config.yaml. - Running
Hermes memory statusconfirms Mnemosyne is available. - Disabling legacy memory (
memory enabled: false,user profile enabled: falseinconfig.yaml) is recommended to avoid duplication. - Providing initial context (e.g., 'My name is Callum...') allows Mnemosyne to store facts like professional background and alias.
- Testing recall by asking 'Who am I?' confirms Mnemosyne is working.
- Mnemosyne uses the BEAM architecture to consolidate working memory into episodic memory and stores structured knowledge as subject-predicate-object triples.
Introducing Hindsight: The Powerful Memory Engine
Introduces Hindsight as a more powerful memory engine that integrates with Hermes. It covers Hindsight's features like mental models, reflection, and its requirement for a separate server or cloud connection, detailing setup options for cloud, local external, and LLM integration.
- Hindsight is a native memory provider for Hermes, operating as a memory engine.
- It builds mental models, pulls observations, and offers features like 'reflect' for synthesized responses.
- Hindsight requires a separate server or cloud connection and an LLM.
- Setup options include Hindsight Cloud (paid) or a local external server.
- Local setup requires running a separate server (e.g., via Docker) and connecting it to an LLM (e.g., Ollama with GPT-OSS 20B).
- The 'reflect' feature uses an agentic loop to search memory and shape reasoning.
- Hindsight's dashboard provides a detailed view of memory (constellation, table, timeline views).
Setting Up Hindsight Locally with Docker and Ollama
Details the process of setting up Hindsight for local external use, including running the Docker container, connecting it to a local LLM like GPT-OSS via Ollama, and configuring Hermes to use this local Hindsight instance. It demonstrates storing and recalling information with Hindsight.
- Hindsight can be run locally using Docker, with options for automatic restart.
- Environment variables configure Hindsight to connect to a specific local LLM (e.g., GPT-OSS 20B via Ollama).
- The Hindsight server runs on specified ports, accessible via a dashboard.
- Hermes is configured in 'local external' mode, pointing to the running Hindsight server.
- Persistent memory in Hermes can be turned off when using Hindsight to avoid duplication.
- Testing involves storing user profile information ('Who am I?') and verifying recall.
- Hindsight stores information in chunks and can build detailed memory graphs (constellation view).
Mnemosyne vs. Hindsight: Choosing the Right Provider
Compares Mnemosyne and Hindsight, summarizing their core differences: Mnemosyne as a lightweight memory layer and Hindsight as a powerful memory engine. It advises users to test providers to find what best fits their workflow and workflow, emphasizing the lack of vendor lock-in.
- Mnemosyne: lightweight memory layer, optimized for simplicity, speed, single-machine deployment, zero dependencies.
- Hindsight: memory engine, sophisticated NLP, multi-signal retrieval, requires separate server/cloud and LLM.
- The choice depends on user needs: speed/simplicity (Mnemosyne) vs. advanced features/processing (Hindsight).
- Users are encouraged to test providers for a week to see workflow improvement.
- There is no lock-in; memories can be exported between providers.
- Hindsight offers a more integrated and detailed dashboard compared to Mnemosyne's community dashboard.
Conclusion and Next Steps: Integrating World Knowledge
Concludes by reiterating the achievement of giving agents 'true agentic memory' with Mnemosyne and Hindsight. It previews the next step: connecting operational memory to world knowledge via an Obsidian LLM Wiki to enable agents to reason against existing knowledge.
- Agents now have 'true agentic memory' with Mnemosyne (lightweight layer) and Hindsight (powerful engine).
- The next step is integrating operational memory with world knowledge (Obsidian LLM Wiki).
- This integration allows agents to reason against existing knowledge, not just remember facts about the user.
- The video provides resources in a free Patreon post for further details and decision-making.