A trainable agent-memory system from BAAI's VectorSpaceLab team (the group behind BGE-VL, BGE-Reasoner and OmniGen), led by Zheng Liu, built on the premise that an agent's memory should be assembled when a request arrives rather than compressed in advance. Memory systems such as A-MEM, Mem0, MemoryOS and LightMem summarize or restructure an agent's history ahead of time (AOT), before any query, and so can discard details a later request needs. JAM keeps the full history instead. A Memorizer writes each session as a raw file into a hierarchical workspace, grouping related sessions into directories whose README files summarize their contents. At query time a Researcher explores that workspace with three tools (open a directory, hybrid BM25 plus BGE-M3 search, browse a file) until it judges the evidence sufficient, then hands the client agent a compact context that cites its sources. JAM is the trainable successor to the group's GAM (November 2025), which introduced the Memorizer–Researcher split, and the paper points to GAM's repository (about 860 stars) for its code.

The training contribution is Memory-Gym, a pipeline that synthesizes evidence-grounded tasks in three families (locating one session, linking several, summarizing many), nine task types and six domains, from web pages and books to dialogue, news and papers; 16,794 of 28,276 candidates passed validation, and 95% of a 200-instance human audit met every criterion. A Qwen3.5-4B Researcher is first fine-tuned on trajectories generated by Qwen3.5-122B-A10B, kept only when correct and filtered by Gemini 3 Flash, then trained with hint-guided GRPO: during rollouts a hint generator that sees the reference answer periodically suggests unexplored directions without revealing the answer, the reward is recall of the gold supporting sessions, and hints are removed at inference. A typical SFT run took about 5 hours and a GRPO run about 6 hours on 8×H100 nodes. With its 4B backbone, JAM scores 52.1 F1 on LoCoMo against 46.1 for the 14B MemAgent, 65.2% on LongMemEval against 56.8% for the best baseline, 45.7 F1 on NarrativeQA against 32.9, and 58.7 F1 on HotpotQA with 224K tokens of context against 49.9 for plain RAG, while answering a LoCoMo query in 13.8 s against MemAgent's 58.1 s. Hints add 0.7 to 5.2 points over plain GRPO, and the trained Researcher transfers to an unseen code domain (LongCodeQA: 54.1% untrained, 65.7% after SFT, 75.5% after RL). At filing, the cited repository still held GAM's code and had not been updated since March 2026, so JAM's training code and Memory-Gym data were not yet public.

Paper

Authors: Bingyu Yan · Chaofan Li · Hongjin Qian · Shuqi Lu · Chaozhuo Li · Zheng Liu
agentsagent-harnessmemorylong-contextretrievalrlpost-trainingresearch

Related