Subconscious Memory Recipe¶
New to Agent Memory? Start with the Quickstart Guide for a progressive adoption path.
Automatic memory injection and extraction around every LLM turn. The agent's memory works silently -- assembling context before each call and saving new observations after each response -- without the application needing to manage memory explicitly.
Retrieval mode determines query-sensitivity
Whether retrieved memories are selected by query text depends on which fields are on your model. With BM25Field only, retrieval is query-sensitive (lexical mode). With both BM25Field and EmbeddingField, retrieval is query-sensitive (hybrid mode). With neither field, retrieval is query-blind (composite mode) — memories are ranked by importance/confidence scores, not by relevance to the user's query. See ContextAssembler retrieval modes for details.
SubconsciousMemory uses retrieval_mode='auto' — the effective mode is determined by which fields are on the model at init time. Adding or removing BM25Field/EmbeddingField changes the effective mode without any change at the SubconsciousMemory call site. Adding EmbeddingField to a BM25-only model silently flips lexical → hybrid.
The default model declares a BM25Field, so the zero-argument path below is query-sensitive. Bring a model without one and ContextAssembler logs a WARNING naming the missing field.
Architecture¶
User message
|
v
[Pre-turn: ContextAssembler.assemble() -> append at tail of messages]
|
v
[LLM inference]
|
v
[Post-turn: extract facts from response -> save as Memory records]
|
v
[Outcome: report acted/dismissed/contradicted via ObservationProtocol]
|
v
Agent response
Quick Start¶
from popoto.recipes import SubconsciousMemory
sm = SubconsciousMemory(agent_id="agent-1")
messages, assembly = sm.inject_context(messages) # pre-turn
answer = call_your_llm(messages) # your LLM call
sm.extract_memories(answer, importance=0.6) # post-turn
sm.report_outcomes(assembly, outcome="acted") # feedback
That is the whole loop. agent_id is the only required argument — it partitions every index, and an explicit .filter(agent_id=...) query always honors that partition. Since 1.9.0 inject_context's default retrieval path (lexical/BM25) honors it too (#576); on 1.8.2 and earlier two agents sharing one Redis through the default loop could retrieve each other's memories, and a distinct Redis database per agent was the workaround.
What the defaults give you¶
Leaving model_class unset selects popoto.recipes.DefaultMemory, the shipped model:
from popoto.recipes import DefaultMemory
class DefaultMemory(AccessTrackerMixin, Model):
memory_id = AutoKeyField()
agent_id = KeyField()
content = StringField(default="")
importance = FloatField(default=1.0)
relevance = DecayingSortedField(base_score_field="importance", partition_by="agent_id")
confidence = ConfidenceField()
content_bm25 = BM25Field(source="content")
associations = CoOccurrenceField()
The BM25Field is the load-bearing piece: it makes retrieval_mode='auto' resolve to the query-sensitive lexical mode. A model without one resolves to composite, which ignores the query text entirely (and now logs a warning saying so).
DefaultMemory deliberately omits WriteFilterMixin — it discards records below a score threshold and save() returns False, which is the wrong surprise for a first run. Add it once you want that behavior; the quickstart covers it at Level 2. EmbeddingField is omitted too, since it needs an embedding provider; adding one to a subclass flips retrieval from lexical to hybrid with no change at this call site.
Two more defaults follow from the default model: score_weights becomes {"relevance": 1.0} (the benchmarked vector), and confidence_field / co_occurrence_field are wired to the model's confidence and associations fields.
DefaultMemory also caps itself at 1000 records per agent_id (Defaults.DEFAULT_MEMORY_MAX_RECORDS_PER_AGENT, 1.9.0). Past the cap each save deletes the stalest record by relevance decay timestamp — a full delete(), so indexes are cleaned and the record is gone for good. Nothing evicted before 1.9.0, so a long-lived corpus grew one record per turn forever; if yours is already over the cap, the first save after upgrading deletes the entire excess at once, synchronously inside that one save — not a gradual trim — so size the exposure as "everything above the cap, now" and check popoto-memory doctor for the record count first.
Subclass and set _max_records_per_agent to change the number, or a falsy value to turn eviction off. The deploy-level POPOTO_DEFAULT_MEMORY_MAX_RECORDS environment variable is the escape hatch for callers with no Python seam (hook adopters using DefaultMemory directly): 0/off disables eviction, a positive integer sets the cap, read fresh on every save. It can lower, raise, or disable the default cap, but it can never re-arm eviction on a subclass that already set _max_records_per_agent falsy — that opt-out always wins over the env var.
Bringing your own model¶
Every argument is still there. Passing model_class explicitly keeps the pre-existing defaults for confidence_field and co_occurrence_field (both None), so upgrading changes nothing for existing code:
sm = SubconsciousMemory(
model_class=Memory, # any level from the quickstart guide
agent_id="agent-1",
score_weights={"relevance": 0.6, "confidence": 0.3},
max_items=10,
max_tokens=4000,
)
Applications that want their own Redis keyspace can subclass instead of authoring a schema:
Injected context format¶
The injected block carries the memory text and nothing else:
Relevant context:
- The deploy pipeline uses a blue-green strategy with automatic rollback.
- Q4 revenue exceeded projections by 12%, driven by enterprise deals.
The content-only shape is what makes the token budget go to memory rather than to bookkeeping. Measured over DefaultMemory with a 71-character memory, the content format runs 73 characters (1.03x the content, ~16 estimated tokens); the full JSON record — memory_id UUIDs, the agent_id the caller just supplied, relevance as a raw epoch float — runs 262 characters (3.69x, ~104 tokens) for the same one memory.
Pass output_format="structured" to restore the JSON payload verbatim, or "xml" / "natural" for the other ContextAssembler formats.
Reindexing Existing Records¶
BM25Field populates its keyword index via the on_save() hook. New records are indexed automatically. Existing records saved before content_bm25 was added are not in the index and will not appear in BM25-driven retrieval.
To backfill, re-save every record once:
Run this once after adding BM25Field to an existing model. The operation is idempotent -- re-running it is safe.
OpenAI SDK Integration¶
Wire subconscious memory into a standard OpenAI chat completion call:
from openai import OpenAI
client = OpenAI() # uses OPENAI_API_KEY env var
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What's our deployment strategy?"},
]
# Pre-turn: inject memories into messages
# (query-sensitive if BM25Field/EmbeddingField on model; query-blind composite otherwise)
messages, assembly_result = sm.inject_context(messages)
# Call the LLM (messages now carry memory context at the tail)
response = client.chat.completions.create(
model="gpt-4.1-nano",
messages=messages,
)
answer = response.choices[0].message.content
# Post-turn: extract facts from the response and save as new memories
new_memories = sm.extract_memories(answer, importance=0.6)
# Report outcomes: memories were used successfully
sm.report_outcomes(assembly_result, outcome="acted")
How It Works¶
Pre-turn: inject_context(messages, *, exclude_keys=None, position="tail")¶
- Extracts the last user message as a query cue
- Calls
ContextAssembler.assemble()with the agent's memory model - Appends the formatted context at the tail — joined to the last message when that is a user message, otherwise as a new trailing user message
- Returns the modified messages and an
AssemblyResultfor later outcome reporting
If no memories are found (or all are filtered), messages are returned unchanged.
Why the tail. A provider prompt cache is keyed on an exact token prefix, so appending after all sealed history costs only the injected tokens. Writing into messages[0] invalidates the whole context on every turn recall changes — which, for a working memory layer, is every turn. The block is never joined to an earlier message either, not even the last user message when assistant turns follow it, because that would still land behind cached tokens. See Prompt Cache Efficiency.
position="system" restores the pre-1.9 placement in the system message (creating one if absent), for callers who need the block read as system-level instruction.
exclude_keys suppresses records already injected this session, so the same memories are not re-added every turn. Suppression is append-only and therefore free against the cache, unlike pruning. Pass the keys you have already surfaced:
seen = set()
messages, assembly = sm.inject_context(messages, exclude_keys=seen)
seen |= {r.db_key.redis_key for r in assembly.records}
Excluded records stay in the store and stay retrievable on a later call that does not exclude them.
The query cue is only meaningful in query-sensitive modes (lexical or hybrid, requiring BM25Field). In composite mode (no BM25Field, no EmbeddingField), the query cue is ignored and ranking is driven purely by importance/confidence scores.
Post-turn: extract_memories(response_text, importance)¶
Before anything else, extract_memories() runs the never-record firewall
(scan_never_record()) against the whole response_text. This is a turn-level
gate, not a per-record one: if the text matches an off-the-record marker or any
other blocked shape, the entire turn is voided — no facts are extracted, the
extraction provider is never called, and the method returns []. Running the
check before the extractor also means that on the ClaudeExtractionProvider
path, a turn carrying an off-the-record marker never reaches the API at all. A
content-free tombstone is still written (see
NeverRecordFirewall). This applies
regardless of model_class — the guarantee does not depend on the model
carrying NeverRecordMixin.
If the turn passes that gate, per-fact drops can still happen further down the
pipeline: when model_class carries NeverRecordMixin, an individual fact's
save() call can itself be blocked. extract_memories() distinguishes "nothing
extracted" from "everything extracted was privacy-dropped" via the
last_extraction_privacy_dropped property:
saved = sm.extract_memories(answer, importance=0.6)
if not saved and sm.last_extraction_privacy_dropped:
# The empty return is a deliberate privacy drop, not a failure.
...
last_extraction_privacy_dropped is True when the most recent
extract_memories() call returned [] because the turn-level gate blocked the
whole turn, or because every extracted fact was blocked on save. It is reset to
False at the top of every call, so it always describes the immediately
preceding one. This exists because an empty return is otherwise
indistinguishable from an extraction outage — without the flag, a caller like
MemoryService.capture() would log every successful privacy drop as a broken
write path.
By default (no extraction_provider passed to the constructor):
- Splits the LLM response into sentences
- Filters out sentences shorter than
extraction_min_length(default 10 chars) - Saves each sentence as a new Memory record with the specified importance
The measured-best write path is the raw turn
Every rewrite-before-store path that has been measured lost to storing the
turn as it arrived. On the judged-answer harness, raw turn ingestion scores
0.3636 over 77 items; the HeuristicExtractionProvider default scores
0.2078, and the Claude extraction arms score lower still, in proportion
to how many turns they discard. Full table, both failure mechanisms, and the
scope of the measurement:
LLM Memory Extraction.
Passing a longer response_text straight through — one record per turn,
no splitting — is the configuration those benchmarks ran. Reach for an
extraction provider when your own corpus gives you a reason to, and measure
it there.
Extraction is pluggable via a dedicated provider interface -- see the LLM Memory Extraction feature doc for the full picture. In short:
extraction_provider(defaultNone->HeuristicExtractionProvider): pass anAbstractExtractionProviderinstance (e.g.ClaudeExtractionProviderfrompopoto.extraction.claude, requirespip install popoto[anthropic]) for LLM-based extraction that also returns entities, an importance opinion, and a confidence opinion per fact.confidence_field(defaultNone, no-op unless set): name of aConfidenceFieldonmodel_class. When set and a fact carries a confidence opinion,extract_memories()seeds that field viaConfidenceField.update_confidence(). BecauseConfidenceFieldhas no per-instance "set initial value" API, this is a blend with the field'sinitial_confidenceprior, not a hard override -- see the confidence blend nuance before assuming the stored value equals the extracted confidence.co_occurrence_field(defaultNone, no-op unless set): name of aCoOccurrenceFieldonmodel_class. When set and a fact names two or more distinct entities, every unordered entity pair is linked in that field's association graph.
Both confidence_field and co_occurrence_field are inert with the default HeuristicExtractionProvider, since it never populates entities or confidence on the facts it emits -- they only do work once an entity/confidence-emitting provider (like ClaudeExtractionProvider) is configured.
For a fully custom extraction source (not implementing the provider interface), you can still subclass SubconsciousMemory and override extract_memories() directly -- see Extensibility below.
Opt-in: auditable extraction¶
Every provider above can still drop a candidate silently -- a too-short
sentence, a malformed model reply, a failed save -- with nothing recorded
beyond a log line. Pass auditable_extraction=AuditableExtractionConfig(...)
instead of (not in addition to) extraction_provider to opt into a
different pipeline: deterministic candidate enumeration, one enum-only LLM
verdict per candidate, and a decision log where every candidate ends in
exactly one of firewall_drop | accept | reject | withhold. Off by default
-- see Auditable Extraction for the
full design and a runnable quickstart.
Outcome: report_outcomes(assembly_result, outcome)¶
Reports how the agent used the injected memories via ObservationProtocol.on_context_used(). Outcomes strengthen or weaken memories for future retrieval:
"acted"-- the agent used this memory (strengthens confidence)"dismissed"-- the agent ignored this memory (mild weakening)"contradicted"-- the agent found this memory incorrect (strong weakening)"deferred"-- the agent noted but deferred action (neutral)"used"-- the memory informed reasoning without appearing in the response (confirms access, no strength signal)
Outcomes prune the corpus¶
Reported outcomes are not only a ranking nudge -- they set how fast a memory leaves the corpus, so the layer stores, retrieves, validates, and prunes during regular use with no extra call site:
report_outcomes(..., "dismissed")lowers the record'sConfidenceFieldvalue.- The next retrieval reads that confidence inside the decay Lua and raises the record's effective decay rate (
eff = decay_rate * 2 ^ (s * 2 * (c0 - c))), so it ranks lower -- see Confidence-Modulated Decay. - The next
MemoryLifecycle.tick()tombstones it once it is idle, its confidence is belowFORGET_CONFIDENCE_CEILING, and it has at leastFORGET_MIN_EVIDENCEobservations behind it.
Modulation is on by default whenever the model carries exactly one ConfidenceField, and forgetting tombstones rather than deletes, so a memory pruned by an unlucky run of dismissals can be brought back with lifecycle.restore(redis_key). Set Defaults.DECAY_CONFIDENCE_MODULATION_ENABLED = False to take confidence out of the ranking entirely, without touching model code.
SubconsciousMemory itself does not run lifecycle ticks -- compose it with a MemoryLifecycle instance as shown in Composing with SubconsciousMemory.
Redis outages raise¶
Since 1.9.0, inject_context, extract_memories and report_outcomes re-raise redis.exceptions.ConnectionError/TimeoutError rather than logging them and returning an empty result. A dead server used to be indistinguishable from "this turn had no relevant memories", which is the failure mode that makes a memory layer look like it is working while it is not.
Wrap the call at your application's turn boundary if a turn must survive an outage:
from popoto.redis_db import OUTAGE_ERRORS # redis ConnectionError/TimeoutError, not the builtins
try:
assembly = sm.inject_context(query)
except OUTAGE_ERRORS:
assembly = None # serve the turn without memory, and alert
Everything else still degrades quietly: extraction that drops a candidate, a zero-hit BM25 query, a missing index.
Tuning¶
| Parameter | Default | Description |
|---|---|---|
max_items |
10 | Maximum memories injected per turn |
max_tokens |
4000 | Token budget for injected context (enforced; see Token Budget Semantics) |
extraction_min_length |
10 | Minimum chars for a sentence to become a memory |
model_class |
None (-> DefaultMemory) |
Memory model. Leave unset for the batteries-included model |
score_weights |
{"relevance": 1.0} |
Weight dict for composite scoring. The benchmarked vector; ignored by the pull path in lexical/hybrid modes |
output_format |
"content" |
Injected payload shape. "structured" injects the full JSON record instead |
system_preamble |
"You are a helpful assistant." | Prefix for system messages auto-created by inject_context(position="system") |
content_field |
"content" | Name of the text content field on your model |
importance_field |
"importance" | Name of the importance score field |
extraction_provider |
None (-> HeuristicExtractionProvider) |
AbstractExtractionProvider used by extract_memories(). See LLM Memory Extraction |
confidence_field |
None, or "confidence" with the default model |
Name of a ConfidenceField to seed from extracted facts' confidence opinions; no-op unless set |
co_occurrence_field |
None, or "associations" with the default model |
Name of a CoOccurrenceField to link co-mentioned entities in; no-op unless set |
These constants can be tuned experimentally using the Tier 4 benchmark harness. See the Tuning Magic Numbers guide for the full constant catalog, optimal ranges, and how to run parameter sweeps.
Extensibility¶
Custom Fact Extraction¶
The preferred way to customize extraction is to pass an extraction_provider -- either the built-in ClaudeExtractionProvider or your own AbstractExtractionProvider implementation -- rather than subclassing. See LLM Memory Extraction for the full interface and a "writing a custom provider" example.
from popoto.extraction import AbstractExtractionProvider, ExtractedFact
class MyProvider(AbstractExtractionProvider):
def extract(self, text: str) -> list[ExtractedFact]:
facts = my_extraction_function(text)
return [
ExtractedFact(text=f["text"], importance=f.get("importance"))
for f in facts
]
sm = SubconsciousMemory(
agent_id="agent-1",
extraction_provider=MyProvider(),
)
If you need to change more than extraction itself (e.g. custom save logic, side effects beyond seeding confidence/co-occurrence), subclassing and overriding extract_memories() directly is still supported:
class SmartSubconsciousMemory(SubconsciousMemory):
def extract_memories(self, response_text, importance=0.5):
# Use a secondary LLM call to extract structured facts
facts = my_extraction_function(response_text)
saved = []
for fact in facts:
m = self.model_class(
agent_id=self.agent_id,
content=fact["text"],
importance=fact.get("importance", importance),
)
m.save()
saved.append(m)
return saved
Custom Query Cues¶
The default implementation uses the last user message as the query cue. For more sophisticated cue extraction, subclass and override the relevant portion of inject_context().
See Also¶
- Agent Memory Quickstart -- progressive adoption guide
- Query-Blind Retrieval -- when composite ranking is right, and when it costs you the answer
- LLM Memory Extraction -- pluggable extraction providers, entities, importance/confidence opinions
- Auditable Extraction -- opt-in candidate/verdict/decision-log path with offline precision/recall
- ContextAssembler -- retrieval-to-injection bridge
- PolicyCache Recipe -- RL-style learned action selection
- Trajectory Memory Recipe -- fingerprint-keyed procedural memory: cluster completed task trajectories and recall "what worked last time"