The hardest problem in long-running agent systems, and the part we've thought about the most.
A 40-hour research run produces 2000+ tool calls. That's 2000+ assistant reasoning blocks, 2000+ tool results (file contents, bash output, training logs), plus system messages. No context window can hold all of that. GPT and Claude top out at 128K-200K tokens. A single training run's stdout can be 50K tokens.
Naive truncation — dropping the oldest messages — is catastrophic. The agent forgets the spec. It forgets the baseline. It re-reads files it already read. It re-tries approaches that already failed.
Remoroo solves this with a demand-paging memory system inspired by OS virtual memory.
Read more here: https://lnkd.in/dGJkhmrs