Memory system for LLMs that remembers everything you teach it during conversation. No reindexing, no context window limits. CPU by default, GPU optional.
-
Updated
May 2, 2026 - Python
Memory system for LLMs that remembers everything you teach it during conversation. No reindexing, no context window limits. CPU by default, GPU optional.
Context window attention optimizer reordering retrieved documents to place critical evidence at prompt boundaries, mitigating lost-in-the-middle degradation.
High-throughput semantic query cache and approximate deduplicator combining exact SHA-256 and cosine similarity threshold matching with LRU eviction.
Zero-dependency hybrid retrieval engine fusing sparse lexical BM25 Okapi and dense Cosine vector similarities via Reciprocal Rank Fusion.
High-throughput semantic query cache and approximate deduplicator combining exact SHA-256 and cosine similarity threshold matching with LRU eviction.
Context window attention optimizer reordering retrieved documents to place critical evidence at prompt boundaries, mitigating lost-in-the-middle degradation.
Cognitive multi-tier memory kernel managing working scratchpad, sliding dialogue buffer, and episodic memories with exponential recency decay.
Entity-relation-object triplet extractor and multi-hop path reasoning kernel with Cypher and JSON-LD graph generation.
Zero-dependency hybrid retrieval engine fusing sparse lexical BM25 Okapi and dense Cosine vector similarities via Reciprocal Rank Fusion.
Cognitive multi-tier memory kernel managing working scratchpad, sliding dialogue buffer, and episodic memories with exponential recency decay.
Entity-relation-object triplet extractor and multi-hop path reasoning kernel with Cypher and JSON-LD graph generation.
[ACL 2026 Main] MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
Memory-Augmented Tensor Hybrid with Intelligent Routing
A non-Transformer hierarchical recurrent network with differentiable Gumbel-Softmax routing and bounded memory slots. Runs 7B+ parameter models layer-by-layer on low-budget GPUs.
An auditable lifelong-learning environment for small language models using curriculum, memory, verified tools, and continual-learning experiments.
Independent implementation of memory-augmented prompt optimization
Memory-centric inference system: frozen RWKV/Mamba backbone + Titans associative memory, test-time training, vector DB — learns at inference time
A viewer whose perception evolves with each image—stateful, memory-carrying VLM reflections across a gallery.
GLACIER: Mamba with indexless memory. This project integrates the Mamba SSM with ICE-Lite, a virtual memory engine, to solve context rot. By adding persistent, time-aware memory, GLACIER gives Mamba the long-term recall of a Transformer while retaining its $O(N)$ speed. Apache 2.0 licensed, by Dopove.
Real-time neural network training observatory — GPU-backed PyTorch, live weight heatmaps + gradient norms + activation histograms via SSE, memory-augmented input with text embeddings
To associate your repository with the memory-augmented topic, visit your repo's landing page and select "manage topics."