โ† AI Tools
RAGAdvanced

Memory OS

Memory OS is a 7-layer architecture local agent-specific memory operating system developed to fundamentally solve the chronic short-term memory loss problem of large language model-based agents. Similar to how a computer manages resources by organically switching between random access memory (RAM) for high-speed processing and solid-state drives (SSD) for storing large amounts of data, this system manages the agent's real-time context window in a three-dimensional manner. The system uses a SQLite database, a Qdrant vector search engine, a Redis message queue, and a background process.

Memory OS is a 7-layer architecture local agent-dedicated memory operating system developed to fundamentally solve the chronic short-term memory loss problem of large language model-based agents. Similar to how a computer manages resources by organically switching between random access memory (RAM) for high-speed processing and solid-state drives (SSD) for storing large amounts of data, it manages the agent's real-time context window in a three-dimensional manner. This system bundles SQLite databases, Qdrant vector search engines, Redis message queues, and background asynchronous workers into an independent Docker container ecosystem, allowing it to run in a completely private environment on a local machine. It collects and decomposes the flow of all conversations between the agent and the user, and then indexes them in a three-dimensional manner in a relational structure based on SQLite and a high-dimensional cosine similarity vector space based on Qdrant, thereby building a permanent factual knowledge base.

Many existing memory plugins or RAG (Retrieval-Augmented Generation) systems for agents have simply dumped indiscriminate vector similarity search results into the agent's prompt in response to simple user conversation search requests. This approach quickly depletes the limited context, increases costs, and creates a vicious cycle in which the agent continuously calls external tools to re-search information that has already been provided, dramatically reducing overall performance. Memory OS transcends the limitations of this makeshift approach by providing a 4-step search fallback (Hybrid โ†’ Dense โ†’ Lexical โ†’ SQLite) system combined with SQLite's FTS5 engine, and achieves maximized token efficiency through session-based duplicate data filtering and semantic deduplication scanners. In particular, the key differentiator of this operating system lies in its top-level identity layer, the Ground Truth Hierarchy. This layer instructs the agent to recognize that the memory context injected into the prompt is absolute and authoritative, thereby eliminating the inefficiency of the agent distrusting its own knowledge and redundantly querying external databases.

From the perspective of a typical researcher or developer, Memory OS serves as a solid foundation for long-term RAG knowledge management and autonomous collaboration pipelines. The system extracts key concepts and cross-references from each conversation as it ends, and activates an automated feedback loop based on trust scoring to continuously filter out noisy information through background workers. This refined data is automatically classified and sorted into concept, entity, and comparison folders in an auto-curated wiki structure, acting as a research log that is permanently updated by a skilled human research assistant. Users do not need to re-explain the context of previous work, library specifications, or established rules at the beginning of each session; they can simply recall complex design decisions from months ago into the agent's latest context with just a few lines of inquiry, enabling highly continuous work.

Furthermore, this solution does not upload data to commercial cloud-based memory services, so it can be safely applied in harsh security environments that handle sensitive business logic or research confidential data without worrying about external leaks. It can be freely integrated with external cloud LLM APIs such as OpenAI, Anthropic, and OpenRouter, and is also compatible with infrastructure that runs entirely locally via Ollama or llama.cpp, allowing for flexible adaptation to business characteristics and infrastructure limitations. By combining a local Docker stack and a lightweight database model, it minimizes memory usage while providing powerful performance and stability by seamlessly handling large-scale agent state changes in real-time through background distributed worker processes.

๐Ÿ’ป System Requirements

๐Ÿง RAM

8GB Recommended above (when running local embeddings and LLM models) 12GB~24GB (Needed; speed degradation occurs when using only the CPU)

๐Ÿ’พStorage

Model download size excluding 1GB within, expandable depending on local vector DB data volume

โšก Installation

4-1. Quick Start

curl -sSL https://raw.githubusercontent.com/ClaudioDrews/memory-os/main/setup.sh | bash

4-2. Detailed installation

# 1. Clone and move GitHub repository
git clone https://github.com/ClaudioDrews/Memory-OS.git
cd Memory-OS

# 2. Setting up Python virtual environment and installing dependencies
pip install -r requirements.txt

# 3. Initialize SQLite database and search index
python setup/setup_db.py

# 4. Run essential backend services based on Docker Compose (Qdrant + Redis + Worker)
cd docker
cat > .env << EOF
REDIS_PASSWORD=$(openssl rand -hex 16)
EMBEDDING_DIMS=4096
COLLECTION_NAME=knowledge_base
LOG_LEVEL=INFO
EOF
docker compose up -d

# 5. Copy Icarus plugin and enable configuration file
cp -r ../icarus/ ~/.hermes/plugins/icarus/
# Add '- icarus' to the 'enabled' section of the ~/.hermes/config.yaml file, then restart the gateway.

FAQ

What is Memory OS?

Memory OS is a 7-layer architecture local agent-dedicated memory operating system developed to fundamentally solve the chronic short-term memory loss problem of large language model-based agents. Similar to how a computer manages resources by organically switching between random access memory (RAM) for high-speed processing and solid-state drives (SSD) for storing large amounts of data, it manages the agent's real-time context window in a three-dimensional manner. This system bundles SQLite databases, Qdrant vector search engines, Redis message queues, and background asynchronous workers into an independent Docker container ecosystem, allowing it to run in a completely private environment on a local machine. It collects and decomposes the flow of all conversations between the agent and the user, and then indexes them in a three-dimensional manner in a relational structure based on SQLite and a high-dimensional cosine similarity vector space based on Qdrant, thereby building a permanent factual knowledge base. Many existing memory plugins or RAG (Retrieval-Augmented Generation) systems for agents have simply dumped indiscriminate vector similarity search results into the agent's prompt in response to simple user conversation search requests. This approach quickly depletes the limited context, increases costs, and creates a vicious cycle in which the agent continuously calls external tools to re-search information that has already been provided, dramatically reducing overall performance. Memory OS transcends the limitations of this makeshift approach by providing a 4-step search fallback (Hybrid โ†’ Dense โ†’ Lexical โ†’ SQLite) system combined with SQLite's FTS5 engine, and achieves maximized token efficiency through session-based duplicate data filtering and semantic deduplication scanners. In particular, the key differentiator of this operating system lies in its top-level identity layer, the Ground Truth Hierarchy. This layer instructs the agent to recognize that the memory context injected into the prompt is absolute and authoritative, thereby eliminating the inefficiency of the agent distrusting its own knowledge and redundantly querying external databases. From the perspective of a typical researcher or developer, Memory OS serves as a solid foundation for long-term RAG knowledge management and autonomous collaboration pipelines. The system extracts key concepts and cross-references from each conversation as it ends, and activates an automated feedback loop based on trust scoring to continuously filter out noisy information through background workers. This refined data is automatically classified and sorted into concept, entity, and comparison folders in an auto-curated wiki structure, acting as a research log that is permanently updated by a skilled human research assistant. Users do not need to re-explain the context of previous work, library specifications, or established rules at the beginning of each session; they can simply recall complex design decisions from months ago into the agent's latest context with just a few lines of inquiry, enabling highly continuous work. Furthermore, this solution does not upload data to commercial cloud-based memory services, so it can be safely applied in harsh security environments that handle sensitive business logic or research confidential data without worrying about external leaks. It can be freely integrated with external cloud LLM APIs such as OpenAI, Anthropic, and OpenRouter, and is also compatible with infrastructure that runs entirely locally via Ollama or llama.cpp, allowing for flexible adaptation to business characteristics and infrastructure limitations. By combining a local Docker stack and a lightweight database model, it minimizes memory usage while providing powerful performance and stability by seamlessly handling large-scale agent state changes in real-time through background distributed worker processes.

When should I use Memory OS?

Memory OS is a 7-layer architecture local agent-specific memory operating system developed to fundamentally solve the chronic short-term memory loss problem of large language model-based agents. Similar to how a computer manages resources by organically switching between random access memory (RAM) for high-speed processing and solid-state drives (SSD) for storing large amounts of data, this system manages the agent's real-time context window in a three-dimensional manner. The system uses a SQLite database, a Qdrant vector search engine, a Redis message queue, and a background process.

๐Ÿ“„ Official Docs๐Ÿ™ GitHub

๐Ÿ“ Update Notes

No update notes yet.

๐Ÿงช Related Code of Life

No related Code of Life posts yet.

BioPlayground

Reading, linking, and lawful quotation stay open; high-speed bulk collection and unauthorized redistribution do not.

Unless stated otherwise, content rights belong to BioPlayground or the relevant rights holder.