All Articles OpenClaw

Tuning OpenClaw Memory: Flush Thresholds, Daily Logs, and Long-Term Knowledge

Your agent just lost a critical insight. It was discussed three sessions ago, but the context window has since filled with newer work.

Your agent just lost a critical insight. It was discussed three sessions ago, but the context window has since filled with newer work. Now it’s about to repeat a mistake it already learned from. That’s the moment you realize: memory tuning isn’t optional. It’s what separates systems that genuinely learn from systems that re-learn the same lessons forever.

We face a hard constraint: context windows are finite. We can’t keep unlimited memory active. Our session threads fill up. Important patterns get crowded out by recent noise. Without the right architecture, we either lose knowledge or run out of space.

OpenClaw’s three-tier memory architecture—session memory (hot), daily logs (warm), and long-term knowledge (cold)—solves this. But it only works if we tune the thresholds that move knowledge between tiers. Get it right, and our system learns continuously. Get it wrong, and we’re back to square one: repeated mistakes, lost insights, wasted tokens.

The Three-Tier Memory Model

Before we talk about tuning, let’s understand the architecture we’re working with.

Session Memory — This is the active conversation. Everything happening in the current session lives here. It’s fast, immediate, but limited. We call this the hot layer—it’s where we think and work right now.

Daily Logs — Every 24 hours (or on demand), we compress session memory and archive it. Important facts, decisions, patterns—we move them here. Daily logs are searchable and persistent, but they’re not active context. This is our warm layer.

Long-Term Knowledge — We then process the daily logs. Patterns emerge. Cross-session insights are synthesized. We move this to MEMORY.md, our system’s permanent knowledge base. This is our cold layer—deep knowledge we reference when needed.

The challenge we’re solving together: how do we prevent session memory from exploding while ensuring nothing important gets lost?

softThresholdTokens: Our Control Knob

This is the critical tuning parameter: softThresholdTokens. It’s where we gain control.

Here’s what it does:

memory:
  softThresholdTokens: 80000 # out of 100,000 context window
  hardThresholdTokens: 95000
  compression_trigger: 0.85 # ratio of soft to hard

When our session memory reaches 80,000 tokens, OpenClaw starts considering a flush. It’s not automatic—we’re in advisory mode. The system evaluates: “Is this a good time to flush?” If conditions are right, it proceeds. If we’re in the middle of something critical, it waits.

At 95,000 tokens, we hit the hard threshold. Now it’s non-negotiable. We flush, because at 100,000, our context window fills completely and we can’t add new tokens.

Why Soft and Hard?

Here’s the design philosophy: we want flexibility. If we’re in the middle of writing a 50,000-word novel, a memory flush might interrupt creative flow. So OpenClaw doesn’t immediately flush at the soft threshold. It suggests. It prepares. But it respects our momentum.

However, if we’re about to hit the wall, it acts decisively. The hard threshold is our safety net.

Tuning softThresholdTokens

The right value depends on your workload:

For creative writing (story agents):

softThresholdTokens: 75000 # Give more headroom
hardThresholdTokens: 95000

Why? Creative tasks benefit from long context. You want to see previous chapters, character development, established tone. Keep that context window spacious.

For code engineering (engineering agents):

softThresholdTokens: 60000 # Tighter budget
hardThresholdTokens: 90000

Code tasks are more modular. You’re often looking at specific files, specific functions. You don’t need as much historical context as creative work.

For validation (validation agents):

softThresholdTokens: 50000 # Very tight
hardThresholdTokens: 85000

Validation is deterministic. It’s checking facts, running tests. It doesn’t need much context. In fact, less context is better—fewer distractions.

The Compression Ratio

Here’s a lesser-known knob: compression_trigger.

compression_trigger: 0.85 # triggers at 85% of soft threshold

This is the ratio that determines when to start thinking about compression. At 0.85, you’re triggering flush considerations at 68,000 tokens (0.85 × 80,000). This gives the system time to prepare before you’re actually at the soft threshold.

Most teams should leave this at 0.85. It’s predictable. But if you have bursty workloads, you might adjust it:

  • 0.75 — Very early warnings. Lots of time to prepare but you’ll flush frequently.
  • 0.85 — Balanced (default).
  • 0.95 — Late warnings. You’ll use more of your context window but flush less often.

softThresholdTokens in Practice: Token Mechanics

Let me walk you through what actually happens when your session memory hits the soft threshold.

When you reach 80,000 tokens (using our example soft threshold), OpenClaw doesn’t panic. Instead, it runs a decision algorithm:

  1. Evaluate the current task: Is the agent in the middle of something critical? Writing a chapter? Debugging code? If yes, defer the flush. Let it continue. The soft threshold is advisory, not mandatory.

  2. Check promotion readiness: Can we actually compress this memory into valuable daily log entries? If the memory is too noisy or lacks clear decision points, flushing won’t help. Wait for better candidates.

  3. Estimate flush cost: How long will compression take? If we’re doing semantic clustering (which is expensive), and you’re in a time-sensitive task, defer. If it’s background processing, go ahead.

  4. Make the call: If all signals say “good time,” initiate a flush. If not, keep the session hot and increment a counter.

If you hit this soft threshold multiple times without flushing (say, 5 times), OpenClaw internally logs a warning. “We’ve suggested a flush 5 times. Maybe the user’s workload is just big.” This becomes input to the hard threshold trigger.

At the hard threshold (95,000 tokens), the decision is binary: flush or fail. You can’t add more tokens without flushing. So it happens immediately.

Practical implications: If your soft threshold is too low relative to your workload, you’ll get frequent “soft threshold reached” evaluations. If your hard threshold is too close to the context window, you leave little buffer for error. The ratio matters.

Recommended setup: Soft at 75-80% of context window, hard at 95% of context window. For a 100k window, that’s soft=75-80k, hard=95k.

Why this two-tier system exists: Cloud models have rate limits. You might flush and immediately hit an API rate limit, blocking continuation. The soft threshold lets you keep going for a bit. The hard threshold says “okay, we really need to flush now, rate limit be damned.”

What Gets Promoted to Daily Logs

Here’s where it gets interesting: not everything in session memory makes it to daily logs. We need to be selective.

OpenClaw applies filters based on what we care about:

daily_log_promotions:
  include_patterns:
    - type: decision
      threshold: importance_score > 0.7

    - type: fact
      threshold: confidence > 0.9

    - type: pattern
      threshold: appears_in > 3_sessions

    - type: error
      threshold: all # Always include errors

  exclude_patterns:
    - type: intermediate_working
    - type: failed_attempt
    - type: duplicate

Let’s break this down.

Decisions — When an agent makes a choice that affects the system (like “we’re using async for this function”), it’s important. But only if the importance score is >0.7. Why? We don’t want to archive trivial decisions.

Facts — When our system learns something (“Python’s GIL prevents true parallelism”), we archive it. But only if confidence is high (>0.9). Uncertain knowledge is less useful long-term.

Patterns — When something appears consistently across sessions (like “we always prefer Opus for complex tasks”), it gets promoted. But only after it appears 3+ times. We need statistical confidence.

Errors — Always archive. Failures are gold. They teach us what not to do.

Intermediate working — This is noise. The scratchpad work doesn’t matter. Just the results matter.

Failed attempts — Counterintuitively, we might want to keep these! But the default excludes them. I’d argue for changing that. Failed attempts are valuable data.

Tuning Promotion Thresholds

If our daily logs are getting bloated, we’re including too much. We should tighten the thresholds:

# More selective
include_patterns:
  - type: decision
    threshold: importance_score > 0.85 # was 0.7

  - type: fact
    threshold: confidence > 0.95 # was 0.9

  - type: pattern
    threshold: appears_in > 5_sessions # was 3

If our daily logs are missing important stuff, we’re being too selective. We loosen them:

# More inclusive
include_patterns:
  - type: decision
    threshold: importance_score > 0.5

  - type: fact
    threshold: confidence > 0.8

  - type: pattern
    threshold: appears_in > 2_sessions

The art here is finding the middle ground where our logs are useful but not overwhelming.

Daily Logs → Long-Term Knowledge: The Distillation Process

So we’ve got a daily log. Now what?

Every week (or on demand), we process daily logs and distill them into MEMORY.md—our system’s long-term knowledge base.

This is where real learning happens.

The Distillation Pipeline

Daily Logs (7 days of compressed data)
    ↓
[semantic clustering]
    ↓
[pattern extraction]
    ↓
[conflict resolution]
    ↓
MEMORY.md (unified knowledge)

Semantic clustering — Group similar facts together. “The user likes Opus” and “Opus produces better code” might both go into a cluster about model preferences and code quality.

Pattern extraction — Look for repeated themes. If five daily logs mention “async functions are tricky,” synthesize that into a single insight.

Conflict resolution — Contradictions emerge. “User prefers Opus” but “last week they said Haiku was sufficient.” Reconcile these. Maybe the user prefers Opus for complex tasks and Haiku for simple ones. Update MEMORY.md accordingly.

The Structure of MEMORY.md

memories:
  user_preferences:
    model_choice:
      context: "User typically selects Opus for complex narrative work, Haiku for quick tasks"
      confidence: 0.92
      last_updated: 2026-03-17
      source_logs: ["2026-03-10", "2026-03-11", "2026-03-15"]

  technical_learnings:
    async_patterns:
      context: "Async functions require careful thought about blocking operations"
      confidence: 0.88
      last_updated: 2026-03-17
      examples: ["deadlock_avoidance", "promise_handling"]

  domain_patterns:
    story_development:
      context: "Plot twists work best when foreshadowed 3+ chapters earlier"
      confidence: 0.75
      last_updated: 2026-03-12
      instances: 12

Notice the structure? Each memory includes:

  • context — The actual knowledge
  • confidence — How sure are we?
  • last_updated — When was this learned?
  • source_logs — Where did this come from?
  • examples/instances — Evidence

This structure is critical. It lets the system (and you) understand not just what’s known, but how confident we should be in that knowledge.

memsearch: Finding Knowledge in Long-Term Memory

We’ve built MEMORY.md. Now we need to use it effectively.

memsearch is OpenClaw’s semantic search over long-term memory. Instead of keyword matching, it uses embeddings to find conceptually similar memories.

User query: "Should we use async for this file operation?"

memsearch results:
  1. async_patterns (confidence: 0.88)
  2. blocking_operations (confidence: 0.82)
  3. performance_optimization (confidence: 0.71)

The system found not just “async” but related concepts: blocking operations, performance. This is more useful than keyword search.

Tuning memsearch

memsearch:
  similarity_threshold: 0.65
  max_results: 5
  recency_weight: 0.1
  confidence_weight: 0.8

similarity_threshold — How similar must a memory be to match? At 0.65, we get broad results. At 0.85, only tight matches. For creative work, we lower this (we want inspiration from related domains). For engineering, we raise it (precision matters).

max_results — How many memories should we return? 5 is typical. More gives us options; fewer keeps things focused.

recency_weight — How much do recent memories matter vs. old ones? At 0.1, age barely matters. At 0.5, recent memories dominate. For fast-moving domains, we increase this. For timeless knowledge, we decrease it.

confidence_weight — How much do we trust high-confidence memories? At 0.8, confident memories rank highly. At 0.5, even uncertain memories get consideration. For safety-critical work, we increase this. For exploration, we decrease it.

Daily Log Structure and MEMORY.md Curation

Let me show you what a properly structured daily log looks like, because this affects everything downstream.

A daily log is a JSON or YAML file containing structured events from a single day’s session. Here’s a real-world example:

date: 2026-03-17
session_count: 3
total_tokens_processed: 156000

events:
  - timestamp: "09:15:00"
    type: decision
    content: "Decided to use async/await for file operations in new module"
    importance_score: 0.78
    confidence: 0.92
    context: "Performance testing showed callbacks were causing bottlenecks"
    tags: [async, performance, architecture]

  - timestamp: "10:42:00"
    type: fact
    content: "Python's asyncio doesn't support true parallelism on multicore due to GIL"
    confidence: 0.95
    context: "Learned when prototyping concurrent processing"
    tags: [python, concurrency, limitation]

  - timestamp: "14:20:00"
    type: pattern
    content: "User consistently chooses Opus for complex reasoning, Haiku for fast tasks"
    confidence: 0.85
    instances: 7
    context: "Observed across multiple sessions this week"
    tags: [user-preference, model-selection]

  - timestamp: "15:33:00"
    type: error
    content: "Race condition in file write when multiple agents access same log"
    confidence: 1.0
    severity: high
    resolution: "Added file locking mechanism"
    tags: [concurrency, bug, reliability]

Each event has a purpose. Notice the structure:

  • timestamp: When it happened
  • type: Category (decision, fact, pattern, error)
  • confidence: How sure we are (0.0-1.0)
  • tags: Keywords for retrieval
  • context: Why this matters

When you curate MEMORY.md (which you should do weekly), you’re not just dumping daily logs. You’re synthesizing. Looking for patterns across days. Removing contradictions. Building a coherent knowledge base.

Curation Workflow:

  1. Review daily logs from the past week
  2. Group similar events by tag
  3. For contradictory facts, research which is correct
  4. Update MEMORY.md with synthesized version
  5. Update confidence based on frequency and agreement
  6. Add “last_verified” timestamps
  7. Archive the daily logs (don’t delete—keep as audit trail)

Example synthesis:

From 7 daily logs, you see multiple entries about “async is hard in Python.” Rather than storing 7 separate facts, you write one entry in MEMORY.md with “instances: 7” and bump confidence from 0.85 to 0.92 because it’s been verified multiple times.

When to Crate MEMORY.md:

After ~30 daily logs (a month), your MEMORY.md might have grown unwieldy. Sometimes it’s worth pruning:

  • Remove entries with confidence < 0.6 that haven’t been touched in 3 months
  • Merge similar entries
  • Reorganize categories if they’re getting too broad
  • Archive old sections and start fresh (but keep a MEMORY.v1.md backup)

This is labor-intensive. But it keeps knowledge fresh and prevents stale facts from crowding out new learnings.

Integration: memsearch + Daily Logs + Session Context

Here’s where the magic happens: how these three layers actually work together during runtime.

Let’s trace a query through the system:

Scenario: An agent asks “Should we use async for this new module?”

Step 1: Session Memory Check (< 1ms)
OpenClaw first checks if something relevant is in the current session memory. Maybe we discussed async 10 minutes ago. If it’s hot in memory, we return it immediately.

Step 2: memsearch Query (10-100ms)
If not in session, we search MEMORY.md via memsearch. The system encodes “async, file operations, performance” and finds:

  • async_patterns (confidence: 0.88)
  • performance_optimization (confidence: 0.71)
  • user_preference_opus_for_complex (confidence: 0.85)

Ranks them by relevance and confidence.

Step 3: Daily Log Check (1-50ms)
Optionally, we check if the past few days’ daily logs have relevant entries. This catches very recent learnings that haven’t been promoted to MEMORY.md yet. “We tried async yesterday and hit a GIL issue.”

Step 4: Synthesis
We combine results from all three layers. Present to the agent:

  • Historical decision: “We chose async here because…”
  • Confidence: “88% confident this applies”
  • Caveats: “But last week we learned about the GIL limitation in Python”
  • Recency: “This was discussed 2 days ago”

This multi-layer approach prevents two failure modes:

  • Stale knowledge: If we only checked MEMORY.md, we’d miss yesterday’s discovery
  • Lost context: If we only checked session memory, we’d ignore patterns learned weeks ago

Practical Before/After: A Real Tuning Example

Let’s walk through a real tuning scenario so you see how this works in practice.

Before Tuning: Default config, aggressive flushing

softThresholdTokens: 50000
hardThresholdTokens: 90000
compression_trigger: 0.75
promotion_threshold_decision: 0.7
promotion_threshold_fact: 0.9

Symptoms:

  • Flushes happen every 1-2 hours
  • Lots of noise in daily logs (even trivial decisions get saved)
  • MEMORY.md is bloated (500+ entries)
  • memsearch is slow (lots of false positives)
  • Agent loses context between flushes

Measurement:

  • Flush frequency: 8/day
  • Daily log size: 2MB
  • Memory retrieval latency: 150ms
  • Relevance of memsearch: 60% (many false positives)

Tuning Applied:

softThresholdTokens: 75000 # Give more room before flush
hardThresholdTokens: 95000 # Tighter to context window
compression_trigger: 0.85 # Later warning
promotion_threshold_decision: 0.85 # Only important decisions
promotion_threshold_fact: 0.95 # Only high-confidence facts

Plus manual curation: Reviewed MEMORY.md, removed 300 entries (< 0.75 confidence), reorganized into clearer categories.

After Tuning:

  • Flush frequency: 2/day (down from 8)
  • Daily log size: 400KB (down from 2MB)
  • Memory retrieval latency: 30ms (down from 150ms)
  • Relevance of memsearch: 92% (fewer false positives)
  • Agent retains context better across sessions

The trade-off: Sometimes important context gets dropped at the soft threshold because it hasn’t hit the 0.85 importance threshold. So you need to monitor and adjust if you start missing things.

Monitoring Memory Health: The Metrics That Actually Matter

How do we know if our memory system is working? We track metrics. Real ones.

When we tune memory parameters, we need baseline data and ongoing measurement. Without it, we’re guessing. Here’s what we should monitor and how to act on what we find.

Key Metrics to Track:

  1. Flush Frequency: How many times per day are we flushing?

  2. Target: 1-3 times/day for normal workloads

  3. Too high (>5/day): Our soft threshold is too low, or our tasks are genuinely bursty
  4. Too low (under 1/day): We’re not reaching thresholds; our context window might be oversized

  5. Promotion Rate: What percentage of session memory makes it to daily logs?

  6. Target: 10-20%

  7. Too high (>30%): Our promotion thresholds are too loose; daily logs are becoming bloated
  8. Too low (under 5%): We’re losing important context; thresholds are too tight

  9. Memory Retrieval Latency: How long does memsearch take?

  10. Target: 20-50ms

  11. If >100ms: Our MEMORY.md is too large; we need to prune it

  12. memsearch Relevance: Of the top-5 results, how many are actually relevant?

  13. Target: 4-5 out of 5

  14. If under 3: We should lower similarity_threshold for broader matches

  15. Promotion Utility: When promoted facts are used later, how often were they actually useful?

  16. Target: 70%+ of promoted facts get cited in future sessions
  17. If under 50%: Our promotion thresholds are likely too loose

How We Track These:

OpenClaw logs all metrics to memory/metrics.jsonl automatically. We should check weekly:

# Last 7 days of metrics
tail -7 memory/metrics.jsonl | jq .

# Daily average flush count
cat memory/metrics.jsonl | jq '.flush_count' | awk '{sum+=$1} END {print sum/NR}'

# Memory size trend
cat memory/metrics.jsonl | jq '.memory_file_size' | tail -30

If we notice trends (flush frequency increasing, latency growing), we investigate. Usually it means our system is learning and growing—which is healthy. But eventually we’ll need to prune to maintain performance.

Advanced Tuning: Multi-Workload Profiles

Our story agents and code agents have different needs. Rather than one global config, we create profiles.

# openclaw.yaml - root config

memory_profiles:
  story_writing:
    softThresholdTokens: 80000
    hardThresholdTokens: 98000
    compression_trigger: 0.8
    promotion_threshold_decision: 0.6 # Be inclusive, capture nuance
    memsearch:
      similarity_threshold: 0.55 # Broad search for inspiration
      recency_weight: 0.05 # Old ideas matter too

  code_engineering:
    softThresholdTokens: 60000
    hardThresholdTokens: 90000
    compression_trigger: 0.9
    promotion_threshold_decision: 0.85 # Only significant decisions
    memsearch:
      similarity_threshold: 0.75 # Tight search for precision
      recency_weight: 0.3 # Recent patterns matter

  validation:
    softThresholdTokens: 40000
    hardThresholdTokens: 85000
    compression_trigger: 0.95
    promotion_threshold_decision: 0.9 # Only clear conclusions
    memsearch:
      similarity_threshold: 0.8
      recency_weight: 0.4

Then when you spin up an agent, specify the profile: openclaw init --profile story_writing.

This way, your creative agents have room to breathe (high soft threshold, loose promotion), while your engineering agents are disciplined (low soft threshold, tight promotion).

Here’s how the whole system works end-to-end:

Session → Session memory fills up to softThresholdTokens → Soft threshold triggers evaluation → If good time, flush → Session memory compressed → Important facts promoted to daily log → Daily logs accumulated for 7 days → Weekly distillation → New entries merged into MEMORY.md with confidence scores → memsearch indexes the new memories → Next session retrieves relevant memories via memsearch

Each stage is tunable. Each stage has tradeoffs.

MEMORY.md Curation: Making Knowledge Actionable

Creating MEMORY.md is only half the battle. The other half is keeping it clean and useful.

A poorly curated MEMORY.md becomes useless. It bloats. Stale entries pile up. memsearch returns 10 results and only 2 are relevant. Your system stops learning because it drowns in noise.

Here’s how to curate properly.

Curation Cadence:

Weekly is ideal. Set aside 30 minutes on Friday to review the past week’s daily logs and update MEMORY.md. Monthly works if weekly is too much, but monthly lets stale knowledge accumulate.

Curation Checklist:

  1. Read all daily logs from the past week — Just skim. Get a sense of what the system learned.

  2. Identify emergent patterns — If the same insight appears 3+ times, it’s pattern-worthy. If it appears once, it’s ephemeral. Promote patterns.

  3. Resolve contradictions — If one log says “async is good” and another says “async caused race conditions,” dig deeper. Maybe both are true in different contexts. Update MEMORY.md to reflect nuance.

  4. Add confidence scores — High-confidence memories (verified multiple times) get 0.85+. Low-confidence (mentioned once, might be wrong) get 0.65-0.75.

  5. Update timestamps — When you verify an old memory, update its last_verified timestamp. This helps with memory decay later.

  6. Remove duplicates — If you have two entries about “async patterns,” merge them into one.

  7. Trim stale entries — Every 3 months, review entries that haven’t been referenced. Low-confidence (under 0.65) entries older than 3 months should probably go. Archive them, don’t delete.

Example Curation:

You notice in daily logs: “We use Opus for complex writing, Haiku for summaries, Claude for general tasks.”

In MEMORY.md, you currently have three separate entries:

  • “User prefers Opus for creative writing”
  • “Haiku is fast and sufficient for summaries”
  • “Claude handles general tasks well”

Curation: Merge into one entry:

model_selection_strategy:
  context: "Model selection depends on task complexity and speed requirements. Opus for creative/complex work (high quality, slower). Haiku for simple/fast tasks (lower quality, faster). Claude for general tasks (balanced)."
  confidence: 0.90
  last_updated: 2026-03-17
  source_logs: ["2026-03-10", "2026-03-15", "2026-03-17"]
  instances: 8

This is more useful than three separate entries. memsearch finds one authoritative entry instead of three partial ones.

Pruning Schedule:

  • Weekly: Remove obvious duplicates, update timestamps
  • Monthly: Review low-confidence entries, remove those from >3 months ago
  • Quarterly: Full MEMORY.md audit. If it’s >1MB, ruthlessly prune
  • Yearly: Archive old versions. Start fresh if it gets unwieldy

Common Tuning Scenarios

Scenario 1: “I’m running out of memory”

Your agent keeps hitting hard thresholds. Memory is exploding.

Actions:

  1. Lower softThresholdTokens by 10-20%
  2. Increase compression_trigger ratio to 0.9
  3. Tighten promotion thresholds (0.85+ for decisions, 0.95+ for facts)
  4. Check if your agents are storing unnecessary intermediate work

Scenario 2: “I’m losing important context”

Important facts aren’t making it to daily logs. Decisions get re-hashed.

Actions:

  1. Loosen promotion thresholds (0.5 for decisions, 0.8 for facts)
  2. Lower the pattern frequency requirement to 2 sessions instead of 3
  3. Include failed attempts and intermediate working
  4. Manually review daily logs for quality

Scenario 3: “memsearch isn’t finding what I need”

You query for something that should exist in MEMORY.md but it doesn’t show up.

Actions:

  1. Lower similarity_threshold from 0.65 to 0.5
  2. Increase max_results to 10
  3. Check confidence_weight—if high-confidence memories are drowning out related items, lower it
  4. Manually add missing memories if critical

Scenario 4: “My system is stale”

MEMORY.md contains old knowledge that’s no longer relevant.

Actions:

  1. Increase recency_weight in memsearch from 0.1 to 0.3-0.4
  2. Implement memory decay—age out low-confidence memories
  3. Manually curate MEMORY.md quarterly
  4. Add a “last_verified” field and re-check old memories

Best Practices for Memory Tuning

1. Measure before you tune. Don’t guess. Log your memory usage, flush frequency, promotion rates. Baseline first.

2. Tune gradually. Don’t change everything at once. Adjust one parameter, run for a week, measure impact, then adjust the next.

3. Keep it documented. In MEMORY.md, include a “system_config” section that documents your tuning choices and why:

system_config:
  tuning_rationale:
    softThresholdTokens: "Set to 75k for long creative sessions"
    compression_trigger: "0.85 gives ~1 minute warning before flush"
    promotion_threshold_decision: "0.7 to capture medium-importance choices"
  last_tuned: 2026-03-17
  tuned_by: "human_oversight"

4. Test with diverse workloads. Your tuning for story agents might not work for code agents. Keep separate profiles.

5. Watch the logs. OpenClaw maintains a detailed memory log. Skim it regularly. You’ll see patterns: “flushes always happen at X tokens,” “memory X keeps getting re-discovered,” “confidence scores are consistently Y.”

The Philosophy

Memory tuning isn’t about optimization in the abstract. It’s about preserving what matters while respecting constraints.

You have a finite context window. You have unlimited time. So you need a system that moves knowledge between tiers: hot (immediate), warm (this week), and cold (long-term). The tuning parameters are the gates between those tiers.

Get them right, and your system learns continuously, remembering what matters and forgetting what doesn’t. Get them wrong, and you’ll either drown in noise or lose important insights.

The good news: OpenClaw gives you the knobs. You just need to understand which ones to turn.

If you’re setting up a new OpenClaw instance, start here:

memory:
  softThresholdTokens: 70000
  hardThresholdTokens: 95000
  compression_trigger: 0.85

  daily_log_promotions:
    include_patterns:
      - type: decision
        threshold: importance_score > 0.7
      - type: fact
        threshold: confidence > 0.9
      - type: pattern
        threshold: appears_in > 3_sessions
      - type: error
        threshold: all

  memsearch:
    similarity_threshold: 0.65
    max_results: 5
    recency_weight: 0.1
    confidence_weight: 0.8

Run with this for a week. Measure everything. Then tune from there.

Conclusion: Memory Tuning as Continuous Learning

Memory tuning is the art of balancing fidelity and efficiency. We want to remember everything important and nothing unimportant. In practice, that’s a moving target.

The three-tier architecture—session, daily, long-term—gives us the structure. The tuning parameters give us control. When we use them wisely, we build a system that genuinely learns and improves over time.

Here’s what we’ve covered:

  • softThresholdTokens control when we start thinking about flushing; hardThresholdTokens force the issue when we’re about to run out of space
  • Promotion thresholds determine what gets elevated from session memory to daily logs—tight for quality, loose for breadth
  • memsearch parameters tune how we retrieve long-term knowledge—lower similarity for inspiration, higher for precision
  • Metrics like flush frequency and promotion rate tell us whether our tuning is working
  • Multi-workload profiles let us tune separately for creative work (generous thresholds) versus engineering (tight thresholds)
  • Weekly curation keeps MEMORY.md clean and useful instead of bloated and noisy

The deeper truth: memory tuning is how we teach our system what matters. Every threshold is a philosophical statement about our values. Low confidence thresholds say we value breadth over certainty. High soft thresholds say we prioritize context over efficiency.

When we tune aggressively (tight thresholds, frequent flushes), we’re saying: “Quality matters. Keep only the best.” The system becomes deliberate, conservative.

When we tune loosely (loose thresholds, infrequent flushes), we’re saying: “Breadth matters. Remember everything. We’ll sort signal from noise later.” The system becomes creative, exploratory.

Neither is wrong—they’re strategies for different workloads.

The Real Test

You’ll know your memory tuning is working when:

  • Your agent confidently references a decision from three sessions ago
  • It avoids a mistake it already learned from, without being reminded
  • It connects patterns across weeks that would have taken humans months to recognize
  • It anticipates problems because it saw the pattern before

These aren’t metrics measured in milliseconds. They’re measured in outcomes. A well-tuned memory system behaves like someone with experience—they’ve seen it before, they know what works, they anticipate.

That’s what we’re building here. Not just a technical system, but an agent that learns.


Related Reading:

Keep tuning. Keep measuring. Keep learning.

-iNet

Free Discovery Call

Start With a Conversation, Not a Commitment

Every engagement begins with a free 30-minute discovery call. We'll map what's slowing your business down and tell you exactly what we'd fix first – no pitch deck, no obligation.