You’ve been running an OpenClaw agent for weeks. It’s been solid, making good decisions, building on prior knowledge. Then one day—something breaks. The agent forgets important context you fed it three sessions ago. It repeats mistakes. It starts promoting noise over signal. You’re staring at logs wondering: What just happened?
Welcome to the memory management challenge that gets everyone.
Memory in OpenClaw isn’t magic. It’s a carefully managed system with hard limits, soft failure modes, and a bunch of gotchas that’ll blindside you if you’re not paying attention. The good news? Most of these problems are fixable once you understand the mechanics. The better news? You can prevent them entirely with the right approach.
Let’s dig in.
The Core Problem: Context Windows Are Real
Your agent operates inside a context window. This isn’t theoretical—it’s a hard constraint baked into the LLM you’re using. Claude Opus? 200k tokens. GPT-4 Turbo? 128k tokens. Whatever you’re running, there’s a ceiling, and your agent is bumping against it.
Here’s the thing nobody tells you: your context window isn’t just for the conversation. It’s shared between:
- The system prompt (your agent’s personality and instructions)
- Memory loaded from disk (historical context, facts, learned patterns)
- The current conversation thread
- Any tools or function definitions
- Your agent’s internal reasoning
That 200k token limit? After the system prompt and tools take their cut, you’ve realistically got maybe 100-120k tokens for actual memory and conversation. Feed your agent enough historical context, and you’ll hit the ceiling faster than you’d think.
Think of it like a bank account. You have 200k tokens to spend. Your system prompt costs 5k. Your tools cost 3k. Your memory loads at 50k. Now you’re at 58k spent, leaving only 142k for the actual conversation and reasoning. But wait—the model itself needs tokens to think through problems and generate responses. So really, you only have about 60-80k tokens for actual work. That’s not as much as you’d expect when you first hear “200k context window.”
The painful part? Most developers don’t discover this until they’re deep into production. They build their agent, add memory for weeks, everything works fine. Then suddenly responses get weird. The agent seems dumber. Decisions become inconsistent. You start debugging, wondering if there’s a logic bug, when really the problem is that memory is now consuming most of the available tokens.
When you hit that ceiling, here’s what happens:
- Token truncation: The LLM backend simply cuts off anything beyond the limit. Your oldest memory gets orphaned first.
- Attention degradation: Even before hard limits, performance degrades in the final ~20% of the window. The model becomes less attentive to early context.
- Signal-to-noise collapse: If your memory files are verbose, the important bits get drowned out by filler. The model can’t find the signal.
Understanding the Context Window Mechanics
Let’s dig deeper into how context windows actually work because this is where most memory problems originate.
When you send a message to your LLM (whether it’s Claude, GPT-4, or anything else), the API converts everything to tokens. Tokens are roughly 4 characters on average, but the mapping is non-linear—some characters cost more tokens than others. A single emoji might be 10 tokens. A common word like “the” is 1 token. Code can be wildly expensive because of special characters. Whitespace and line breaks have costs too—markdown with lots of blank lines is less efficient than compact markdown.
Your context window is measured in tokens, not characters or words. So when OpenClaw loads your MEMORY.md file, it’s not counting lines—it’s counting tokens. A 1MB file might be 250k tokens. A 5MB file could be 1.2M tokens (well over typical limits). This is why compression matters—removing duplicate blank lines, condensing repetitive information, and using shorthand notation can cut token costs significantly.
Here’s what happens during a typical session:
- Startup: OpenClaw reads your MEMORY.md (costs X tokens)
- System prompt: Loaded (costs Y tokens—usually 2-5k for OpenClaw with basic settings)
- Tools/functions: If you have any custom behaviors, they get serialized (costs Z tokens—could be 1-5k depending on complexity)
- Current conversation: Each message you send/receive (costs M tokens—depends on message length)
- Reasoning: The model’s internal thinking (costs R tokens—models use some of the context window to think)
Total: X + Y + Z + M + R = N tokens used
If N exceeds your context window limit, the oldest or least-important content gets truncated. Usually memory gets dropped first, then early conversation history. Sometimes the truncation happens in the middle of a sentence or fact, leaving contradictory partial information.
The real kicker: You don’t always get an error. Sometimes the model just produces worse output because it’s working with incomplete context. Your agent doesn’t say “I’m out of tokens”—it just becomes dumber. You notice responses are off, decisions are weird, learning isn’t sticking. Then you dig for days trying to find the bug, when the real problem is just “there’s too much memory loaded.”
Example of what truncation looks like:
You have these rules in memory:
- “User is in Eastern Time Zone (UTC-5)”
- “User is in Pacific Time Zone (UTC-8)”
- “Important: User is ACTUALLY in Eastern Time Zone despite earlier confusion”
If truncation happens mid-memory, you might end up with:
- “User is in Eastern Time Zone (UTC-5)”
- “User is in Pacific Time Zone (UTC-” [TRUNCATED]
- “Important: User is ACTUALLY in Eastern Time Zone despite earlier confusion”
Now the agent has contradictory incomplete information. It might default to Pacific, which is wrong. But the contradiction isn’t obvious—it’s just a malformed rule. This is the nightmare scenario.
Memory Flush Misfires: Lost Information and Promoted Noise
Here’s where it gets frustrating. OpenClaw has a memory management system that’s supposed to keep things organized. It works like this:
Memory Lifecycle in OpenClaw:
- Live memory: Things the agent learns or notes during a session live in session files
- Flush cycle: At session end or periodic checkpoints, important data gets promoted to persistent memory files
- Archival: Old data gets compressed, deduplicated, or marked as stale
Sounds great in theory. In practice? Three things go wrong:
1. Important Context Gets Flushed Prematurely
Your agent learns something critical—a user preference, a system constraint, a business rule. It should go into persistent memory. Instead, the flush logic considers it “low priority noise” and… it vanishes. Now your agent makes decisions that contradict what it should have known.
Why this happens: Your flush rules are probably too aggressive or not specific enough. You’re treating all memory equally, when really you should be flagging certain types of information as “preserve at all costs.”
2. Noise Gets Promoted to Long-Term Memory
Conversely, trivial details—logging statements, debug output, failed exploration branches—get marked as important and persisted. Now every time your agent starts a session, it’s wading through garbage. The signal-to-noise ratio crashes.
Why this happens: Your system can’t distinguish between “learning” and “noise.” Everything gets saved because, well, why not? Disk space is cheap.
3. Memory Files Grow Without Bounds
You start with a clean MEMORY.md. Six months later, it’s 50MB. Your agent spends precious tokens just loading the file. Each session gets slower. The model starts missing important details buried in the bloat.
Why this happens: You’re appending to memory files but never cleaning them up. No deduplication. No expiration. No prioritization. Just… accumulation.
Real-World Symptoms: Recognizing Memory Breakdown
Before you diagnose, you need to know what memory breakdown actually looks like. Here are the patterns:
Pattern 1: Inconsistent Context Recall Your agent knew something two weeks ago but forgot it. You bring it up again, it learns it again. But then it forgets it again after a few sessions. This screams “memory flush is losing information.” The agent learns it in live memory, but when the session ends and the flush cycle runs, it doesn’t persist to MEMORY.md properly. Next session, it’s gone.
Pattern 2: Degraded Response Quality Responses that used to be sharp and insightful become generic and surface-level. It’s still responding, but with less depth. This usually means your context is getting truncated. The model has less room to think. You might also notice responses get shorter—the model is using minimal tokens because it has less budget.
Pattern 3: Repeating Exploration Your agent gets stuck exploring the same dead-end paths over and over. It made a failed attempt two months ago, but it doesn’t remember that it failed. So it tries it again. And again. This is classic memory loss on important negative learnings. Those “failed attempts” aren’t being preserved.
Pattern 4: Noise-Driven Decisions Your agent starts making decisions based on trivial details. You mentioned a minor issue once, and now it brings it up constantly. Meanwhile, real constraints are ignored. This is signal-to-noise collapse—garbage is promoted to prominence while important information gets buried.
Pattern 5: Slow First Response The first query of a new session takes much longer than subsequent queries. This is the token budget being blown just on loading memory. Subsequent queries have less fresh memory to load, so they’re faster. But the first query hits maximum load.
Recognizing these patterns is your early warning system. If you see them, memory management is almost certainly the culprit.
Diagnosing Your Memory Problem
Before you fix anything, you need to know what’s actually broken. Here’s how to diagnose:
Check 1: Memory File Size
# On your OpenClaw working directory
ls -lh memory/
du -sh memory/
What should you see?
- Good:
memory/directory under 10MB total - Caution: 10-50MB (you’re approaching the wall)
- Critical: >50MB (you’ve hit it)
If you’re over 50MB, you’re definitely losing signals in the noise.
Check 2: Token Count of Loaded Memory
This is the hard one. You need to actually count tokens. If you’re using OpenClaw with Claude backend:
# Pseudo-code (adapt to your setup)
client = anthropic.Anthropic()
with open("memory/MEMORY.md", "r") as f:
memory_text = f.read()
# Use the token counter
response = client.messages.count_tokens(messages=[{"role": "user", "content": memory_text}])
print(f"Memory file tokens: {response.input_tokens}")
Rule of thumb: If your memory file is >20k tokens, you’re eating into your actual conversation budget significantly.
Check 3: Flush Logs
Does OpenClaw log what it’s persisting? It should. Look for:
logs/memory_flush_*.log
memory/hooks-log.jsonl
Grep for patterns:
grep "FLUSH" logs/memory_flush_*.log | tail -20
tail -20 memory/hooks-log.jsonl
You’re looking for:
- How many items got flushed per cycle?
- What types of information?
- Are there repeated entries (deduplication failing)?
- Are old items being expired?
- What’s the flush timestamp—is it recent or stale?
Check 4: Agent Behavior Changes
The behavioral diagnosis:
- Forgetting important context: Agent recalls facts from early sessions inconsistently. Memory flush is too aggressive.
- Slow responses: Same questions now take longer to answer. Memory bloat.
- Repeating mistakes: Agent makes errors it should have learned from. Important memory lost during flush.
- Off-topic tangents: Agent gets distracted by minute details. Noise promoted over signal.
- Increased token usage: Same query now costs more tokens. Memory has grown.
How Context Window Limits Break Things
Let’s be specific about what happens when you overflow:
Scenario: Your system prompt is 2k tokens. Your memory file is 45k tokens. Your current conversation is 40k tokens. You hit the ceiling at 128k (GPT-4 Turbo).
Result: The model truncates. Usually it drops the oldest memory first. But here’s the trap—sometimes it truncates within the middle of a fact, leaving partial, contradictory information. Your agent becomes confused, not just forgetful.
Even worse: The model doesn’t tell you it truncated. You don’t see an error. The response just becomes incoherent, and you’re left debugging phantom issues.
Practical example: Your agent has learned “Always ask for confirmation before deleting data.” But during truncation, this gets reduced to “Always ask for confirmation” because the rest is cut off. Now it’s asking for confirmation on everything, even read operations. Behavior changed without explanation.
Mitigation:
-
Reserve tokens aggressively. If your context window is 128k, pretend it’s 80k. Build in a buffer for the response generation itself.
-
Profile your memory load. Before running a long session, measure: system prompt + memory + one exchange = X tokens. You need your buffer > X.
-
Implement sliding window memory. Don’t load all historical memory. Load the last 5-10 sessions, or the last 2-3 weeks. Archive older stuff separately.
-
Monitor token usage in real-time. Log the token count at the start of each inference. When it approaches your reserve, trigger cleanup proactively.
MEMORY.md Growth Patterns: Identifying Bloat
Your memory directory is a window into what’s accumulating. Let’s look at what you should actually see.
If you run ls -la memory/ on a healthy OpenClaw workspace, you’d see:
MEMORY.md (5-15MB reasonable, >50MB is a red flag)
facts.json (small, <1MB)
patterns.json (small, <1MB)
timeline.jsonl (grows with time, can hit 5-10MB)
hooks-log.jsonl (append-only, can get large)
archive/ (historical data, read-only)
The problem files are:
- MEMORY.md: Treated as one big file, never cleaned. Grows without bounds.
- hooks-log.jsonl: Every hook execution gets appended. Over months, this becomes huge.
These two alone can easily hit 50+ MB if left unattended.
Pattern: Files ending in .log or .jsonl are append-only. They never shrink. If they’re not being rotated or pruned, they’ll consume all your disk space over time. This is a classic failure mode.
Solution: Set up log rotation. Most systems support it via a .logrotate config or similar. If OpenClaw writes to hooks-log.jsonl, you want to rotate it daily:
/path/to/openclaw/memory/hooks-log.jsonl {
daily
rotate 7
compress
delaycompress
missingok
}
This keeps only 7 days of logs (compressed), pruning old ones automatically.
Manual Memory Cleanup Strategies
You can’t automate your way out of everything. Sometimes you need to manually intervene. Here’s how to do it safely:
Strategy 1: The Audit Pass
# 1. Backup first (always)
cp memory/MEMORY.md memory/MEMORY.md.backup
cp memory/hooks-log.jsonl memory/hooks-log.jsonl.backup
# 2. Read through manually
cat memory/MEMORY.md | head -100
# Look for patterns:
# - Repeated entries (dedup these)
# - Stale information (dates older than 6 months)
# - Debug/logging noise (remove)
# - Contradictions (resolve, keep the latest)
# 3. Count lines to see size
wc -l memory/MEMORY.md
Strategy 2: Segment by Age
# MEMORY.md structure (if yours is just one big file):
## ACTIVE (Last 7 days)
- Critical facts
- Current learnings
- In-flight decisions
## REFERENCE (7-30 days)
- Important patterns
- Decisions made
- User preferences
## ARCHIVE (30+ days)
- Historical context
- Learned lessons
- Deprecated approaches
Now you can trim ARCHIVE aggressively without losing the signal.
Strategy 3: Tagging for Importance
# Better structure:
facts:
- fact: "User prefers JSON output"
importance: HIGH
last_verified: 2026-03-10
source: "explicit request"
- fact: "Previous debugging attempt on 2025-11-01 failed"
importance: LOW
last_verified: 2025-11-02
source: "session log"
action: "DEPRECATE after 2026-04-01"
patterns:
- pattern: "When user says 'urgent', they mean 24-hour turnaround"
importance: HIGH
confidence: HIGH
Now when you load memory, you can filter by importance and date. Load HIGH only by default. Load MEDIUM if there’s token budget. Skip LOW and deprecated.
Strategy 4: Deduplication
Over time, you’ll record the same fact multiple times:
## Bad (in your memory file):
- User's timezone is America/New_York (noted 2026-01-15)
- User's timezone is America/New_York (noted 2026-02-01)
- User's timezone is America/New_York (noted 2026-03-10)
Find and eliminate these. A simple regex pass:
# Find likely duplicates (same line appearing 2+ times)
sort memory/MEMORY.md | uniq -d
For each duplicate, keep the most recent version only.
Strategy 5: Archival and Rollover
Create a rotation:
memory/MEMORY.md (current, <10MB)
memory/archive/MEMORY.2025-Q4.md (historical, read-only)
memory/archive/MEMORY.2025-Q3.md (historical, read-only)
At the end of each quarter (or month, if you’re verbose):
- Identify facts in
MEMORY.mdolder than 90 days - Move them to archive file with timestamp
- Keep only the last 90 days live
- Reference archive only when backfilling context (optional, rare lookups)
Preventing Overflow: System Design
The best fix is prevention. Here’s how to architect your memory system to avoid these problems:
1. Token Budget First
Before you design memory storage, calculate:
Available tokens = (Context Window × 0.6) - System Prompt - Tools
Allocated to memory = Available tokens × 0.3
Allocated to conversation = Available tokens × 0.7
If you’re using Claude Opus (200k):
Available = (200k × 0.6) - 5k prompt - 3k tools = 112k
Memory budget = 112k × 0.3 = ~33k tokens
Conversation budget = 112k × 0.7 = ~78k tokens
Your memory files combined should never exceed 33k tokens. That’s your hard limit. Treat it like you’re paying per token (because in a way, you are—every token of memory is a token not available for thinking).
2. Explicit Retention Policies
Define, in writing:
# memory_policy.yaml
retention:
HIGH:
ttl_days: 365
description: "Business rules, user preferences, critical patterns"
MEDIUM:
ttl_days: 90
description: "Session learnings, temporary patterns, recent context"
LOW:
ttl_days: 14
description: "Debug logs, exploration branches, failed attempts"
DEPRECATED:
ttl_days: 0
description: "Mark for immediate removal"
deduplication:
enabled: true
check_every_N_flushes: 10
flush_strategy:
frequency: "end of session OR 1000 tokens accumulated"
prioritize_by: "importance tag, then recency"
compress: true
Enforce this algorithmically. Your memory system should have built-in cleanup, not manual intervention.
3. Structured Memory with Clear Sections
Instead of one giant MEMORY.md, use a schema:
memory/
├── facts.json # Structured facts with tags
├── patterns.json # Learned patterns
├── preferences.json # User/system preferences
├── timeline.md # Chronological events
├── deprecations.json # Things to forget
└── archive/ # Old data
Each file is purpose-built. Loading “preferences” is lightweight. Loading “archive” is optional.
4. Real-Time Monitoring
Add instrumentation:
# Pseudo-code for your memory manager
class MemoryManager:
def __init__(self, token_budget=20000):
self.token_budget = token_budget
self.current_tokens = 0
def load_memory(self):
memory_text = self._read_all_memory_files()
self.current_tokens = self._count_tokens(memory_text)
if self.current_tokens > self.token_budget * 0.9:
logger.warning(f"Memory at {self.current_tokens / self.token_budget * 100:.0f}% of budget")
self._trigger_cleanup()
return memory_text
def _trigger_cleanup(self):
self._remove_deprecated()
self._deduplicate()
self._archive_old_entries()
self._recalculate_tokens()
You want visibility before overflow happens, not after.
5. Implement Memory Compression
Some facts can be compressed without losing meaning:
Instead of:
Session 2026-01-15: Discussed project timeline. User wants deliverable by Q2 2026.
Session 2026-01-22: Followed up on timeline. Still targeting Q2 2026.
Session 2026-02-05: No changes to timeline. Still Q2 2026.
Session 2026-02-28: Timeline unchanged. Q2 2026 is the target.
Compress to:
Project Timeline: Target Q2 2026 (confirmed through 2026-02-28)
This says the same thing in 1/10th the tokens.
Recovery: When Things Break
Sometimes prevention fails. Your agent is already confused, memory is already bloated, you need to recover. Here’s the playbook:
Step 1: Freeze and Backup
# Stop the agent
# Backup everything
cp -r memory memory.backup.$(date +%s)
Step 2: Audit What’s Actually Loaded
Modify your agent to log what memory it’s loading:
# In your agent startup
with open("memory/MEMORY.md") as f:
memory_content = f.read()
print(f"Memory file size: {len(memory_content)} bytes")
print(f"Estimated tokens: {len(memory_content) / 4}") # rough estimate
print(f"First 200 chars: {memory_content[:200]}")
print(f"Last 200 chars: {memory_content[-200:]}")
# Also check file count
file_count = len([f for f in os.listdir('memory') if f.endswith(('.md', '.json', '.jsonl'))])
print(f"Memory files: {file_count}")
Step 3: Identify and Remove Cruft
# Find the largest entries
wc -l memory/MEMORY.md # line count
grep -n "^#" memory/MEMORY.md # section markers
# Find duplicates
sort memory/MEMORY.md | uniq -c | sort -rn | head -20
# Find entries with timestamps older than threshold
grep "2024" memory/MEMORY.md | wc -l # Count old entries
For anything with a count >1, investigate. Deduplicate manually if needed.
Step 4: Restore from Checkpoint
If memory is too corrupted, revert to a known-good checkpoint:
# Find your last good backup
ls -lt memory.backup.* | head -5
# Restore
rm memory/MEMORY.md
cp memory.backup.TIMESTAMP/MEMORY.md memory/MEMORY.md
# Start fresh from there
Step 5: Rebuild Deliberately
Now rebuild your memory intentionally:
- Start with empty
MEMORY.md - Add only the essentials:
- Core business rules (if any)
- User preferences you know are stable
- Critical system constraints
- Run the agent for one week
- Review what it learned and added
- Keep only high-quality additions
The Architectural Fix: Move to Structured Retrieval
Here’s the real long-term solution: stop loading all memory at once.
Instead of:
“Load entire MEMORY.md → pass to agent”
Do:
“Agent makes query → retrieve relevant memory chunks → inject into context”
This requires a vector database (Pinecone, Weaviate, Milvus) or even just semantic search:
# Pseudo-code
class SmartMemory:
def __init__(self):
self.vector_db = VectorDB() # Semantic search
def add_fact(self, fact: str):
# Store with embedding
embedding = model.embed(fact)
self.vector_db.insert(fact, embedding)
def retrieve_for_context(self, query: str, top_k: int = 5):
query_embedding = model.embed(query)
relevant_facts = self.vector_db.search(query_embedding, top_k=top_k)
return relevant_facts
Now your agent loads only what’s relevant for the current task. Memory can grow unbounded without crushing your token budget. This is the architecture used by systems like Anthropic’s Retrieval-Augmented Generation and it scales far better than naive persistence.
Practical Checklist
Before deploying an OpenClaw agent into production, use this:
- [ ] Calculate token budget: Context window × 0.6, minus system prompt and tools
- [ ] Set memory limit: Budget × 0.3, add to monitoring alerts
- [ ] Implement tags/importance: Mark memory by importance (HIGH/MEDIUM/LOW)
- [ ] Set TTLs: Define how long different memory types live
- [ ] Log flushes: Know what’s being persisted and why
- [ ] Weekly audits: Check memory file size, run deduplication
- [ ] Quarterly rotation: Archive old data, keep recent live
- [ ] Monitor token usage: Log current_tokens on every load
- [ ] Test truncation: Deliberately overflow once in dev to see what breaks
- [ ] Have a rollback: Keep backups you can restore from
- [ ] Set up log rotation: Prevent append-only files from consuming disk
- [ ] Implement compression: Summarize repetitive facts
Common Mistakes: What NOT to Do
Before diving into more solutions, let me warn you about the mistakes everyone makes:
Mistake 1: Logging Everything Beginners often add logging to memory on every session: “Session 2026-03-15 10:30am: User asked about X. Agent responded with Y. User seemed satisfied.” Then do this for 200 sessions. Your memory is now 90% session logs and 10% actual knowledge. This is bloat. Real memory should be high-level learnings, not transcripts. Save transcripts elsewhere if you want them, but not in the working memory.
Mistake 2: Never Deleting Anything “Disk space is cheap, so let’s keep everything forever.” Yes, but context windows aren’t cheap. Every byte of memory you load costs tokens. And tokens cost money (in API calls) or speed (if using local LLMs). Treat memory like a scarce resource. Delete aggressively.
Mistake 3: One Big File Throwing everything into MEMORY.md because it’s simple. Then when memory bloats, you’re stuck parsing a 50MB file with grep. Split into logical files: facts.json for structured data, timeline.md for events, preferences.json for user preferences. This makes cleanup surgical instead of sledgehammer.
Mistake 4: Not Tagging Importance All facts are treated equal. The user’s timezone preference (critical) sits next to “user mentioned they like coffee” (noise) with no distinction. When you need to cut memory, you can’t tell what to keep. If everything is tagged, you cut LOW and DEPRECATED first, preserving HIGH.
Mistake 5: Ignoring Deduplication You record “User is in NYC” five times over six months. Each time it gets stored separately. Now 5 tokens are wasted on identical information. A simple weekly dedup pass catches this. It’s not glamorous, but it saves space.
Mistake 6: Trusting the Flush System Too Much You assume OpenClaw’s memory flush is perfect. It’s not. Flush systems are heuristics. Sometimes they get it wrong. You need manual audits. At minimum, quarterly you should read through your memory file and verify it makes sense.
Mistake 7: Not Having Backups Memory corruption or accidental deletion. No rollback plan. You’ve now lost weeks of learning. Automated daily backups take 5 minutes to set up. Do it.
Mistake 8: Conflating Session Data with Long-Term Memory Session data (conversations with the agent) is different from long-term memory (facts and patterns learned). Dump session data in logs. Put high-quality learnings in memory. Mixing them together is how memory bloat happens.
Mistake 9: Ignoring Append-Only Files Log files grow forever unless explicitly rotated. JSONL files are append-only. After 6 months, hooks-log.jsonl can hit 500MB while your active MEMORY.md is clean. You forget to monitor the background files, and suddenly your disk is full. Set up log rotation. It’s non-negotiable.
Mistake 10: Manual Memory Editing Without Structure You manually add facts to MEMORY.md without consistent formatting. Now your memory is a mix of YAML, JSON, markdown lists, and plain text. Parsing this programmatically becomes a nightmare. If you’re going to edit memory manually, use consistent structure: JSON for facts, markdown for narrative, YAML for config. Pick one and stick with it.
Advanced: Monitoring and Alerting
If you want to be really proactive, add monitoring to your setup.
Token Budget Alerts
# monitoring.py - run this periodically (e.g., via cron)
from datetime import datetime
MEMORY_TOKEN_BUDGET = 20000 # Your limit
ALERT_THRESHOLD = 0.85 # Alert at 85% of budget
client = anthropic.Anthropic()
def check_memory_health():
memory_path = "memory/MEMORY.md"
with open(memory_path, 'r') as f:
memory_text = f.read()
response = client.messages.count_tokens(
messages=[{"role": "user", "content": memory_text}]
)
token_count = response.input_tokens
usage_percent = (token_count / MEMORY_TOKEN_BUDGET) * 100
timestamp = datetime.now().isoformat()
print(f"[{timestamp}] Memory tokens: {token_count} ({usage_percent:.1f}%)")
if token_count > MEMORY_TOKEN_BUDGET * ALERT_THRESHOLD:
print(f"WARNING: Memory approaching budget! {token_count} / {MEMORY_TOKEN_BUDGET}")
# Send alert (email, Slack, etc.)
send_alert(f"OpenClaw memory at {usage_percent:.1f}%")
return token_count, usage_percent
if __name__ == '__main__':
check_memory_health()
Run this via cron:
# Check every day at 3am
0 3 * * * /usr/bin/python3 /path/to/monitoring.py >> /var/log/openclaw-memory.log 2>&1
Now you get early warning before overflow actually happens.
File Growth Monitoring
# log_memory_growth.sh
MEMORY_DIR="$HOME/.openclaw/memory"
LOG_FILE="$MEMORY_DIR/growth.log"
SIZE=$(du -s "$MEMORY_DIR" | awk '{print $1}')
TIMESTAMP=$(date +%s)
echo "$TIMESTAMP $SIZE" >> "$LOG_FILE"
# Keep only 90 days of history
find "$MEMORY_DIR" -name "growth.log.*" -mtime +90 -delete
# Rotate log
if [ -f "$LOG_FILE" ]; then
LINES=$(wc -l < "$LOG_FILE")
if [ "$LINES" -gt 90 ]; then
mv "$LOG_FILE" "$LOG_FILE.$(date +%Y%m%d)"
fi
fi
Run daily:
# Every day at 4am
0 4 * * * /path/to/log_memory_growth.sh
Over time, you’ll see if memory is growing linearly, exponentially, or staying stable. That tells you a lot about your flush strategy.
Decision Tree: What to Do When Memory Goes Bad
Here’s a quick flowchart for when things break:
Agent acting weird?
├─ YES: Check memory file size (du -sh memory/)
│ ├─ < 10MB: Probably not memory-related, look elsewhere
│ ├─ 10-50MB: Approaching limits, run deduplication & archival
│ └─ > 50MB: This is the problem, execute recovery plan NOW
│
└─ NO: Is it slow?
├─ YES: Count tokens in MEMORY.md, if > 20k, cleanup
│
└─ NO: Forgetting things?
└─ YES: Check flush logs, adjust retention policies
The key pattern: size first, then behavior, then token count. File size is the cheapest check. If it’s too big, you’ve found your problem. If size is normal, look at behavior. If behavior seems off but size is fine, then count tokens.
Prevention Playbook: Full Checklist
If you’re building a new OpenClaw instance from scratch, here’s the complete setup to do it right from day one:
- Decide your token budget (this depends on your LLM and use case)
- Set max file size (recommend under 10MB for initial setup)
- Create MEMORY.md structure with sections: ACTIVE, REFERENCE, ARCHIVE
- Tag all facts with importance (HIGH/MEDIUM/LOW) and dates
- Create memory_policy.yaml with retention rules
- Set up log rotation for append-only files
- Create backup script and schedule it daily
- Add monitoring for token usage and file growth
- Run first week without cleanup to see what accumulates
- Analyze what got learned and adjust tags accordingly
- Set deduplication to run every week
- Archive quarterly (or monthly if verbose)
- Review quarterly to adjust retention policies based on real usage
This seems like a lot, but most of it is one-time setup. Then it runs mostly automatically.
The Philosophy: Your Agent Deserves Good Memory
Here’s the underlying truth: your agent is only as good as its memory. A forgetful agent is useless. A confused agent is worse than useless—it’s actively harmful.
Poor memory isn’t a feature bug. It’s an architecture bug. It comes from treating memory like an afterthought instead of a first-class system.
When you invest time in memory management, you’re not doing busywork. You’re building the foundation that everything else rests on. A well-managed memory system:
- Helps your agent make better decisions (it has the right context)
- Makes it faster (less bloat to load)
- Makes it cheaper (fewer tokens wasted on garbage)
- Makes it more reliable (less confusion)
- Makes it learnable (you can see what it actually knows)
That’s why this article exists. That’s why it’s important.
Get memory right, and your agent will surprise you with how capable it becomes.
-iNet