Atlas of Machine Memory
Stays Up
arXiv sweep · 2025-03 → 2026-08 · built bottom-up from embeddings

Atlas of Machine Memory

Did anyone solve the write path — navigating and remembering beyond the context window without retraining? 7,414 papers, 18 months, one tree.

7,414
papers
248
clusters
5
branches
87
rising
42
fading
SOLVED — and commodity
Semantic lookup. Embedding indexes and hybrid retrieval are settled infrastructure; every memory product converges on the same six engine patterns. Building engines is a graveyard; the value is in owned state.
CONTESTED — no winner
External memory design. 1,471 papers of prosthetics with no benchmark consensus and an open substrate war. Consolidation (what to keep, when to rewrite) is the live fight, not storage.
OPEN — the real seams
Evaluation of memory quality; memory integrity under poisoning; aggregation over corpora that no context can hold; and the native write path itself — test-time training is the only rising candidate inside the weights.

Method note: tree built bottom-up — 7,414 papers embedded locally, clustered (k=248, cohesion 0.86–0.92), each cluster carded by a small model with month histograms, branches and this synthesis written by a larger one. Every claim descends to actual papers.

risingsteadyfadingn = papers in cluster

Making Context Cheaper

2,473 papers · 68 clusters · 27 rising

A third of the entire field optimizes the existing context mechanism: KV-cache compression and quantization, prompt compression, sparse and adaptive attention, and the serving systems that schedule it all. This is the crowded zone — and it is engineering the symptom, not the disease: nothing here changes what context is. Prompt compression reads post-peak (crowded and fading); the live edges are sparse attention learned at pretraining time, task-aware cache compaction that tries to preserve reasoning rather than tokens, and latent-attention variants.

KV Cache Compression Evolutionactive134
Prompt & Context Compressioncrowded100
Inference Serving & KV Management Systemsactive96
Long-Context RL & Scalingcrowded89
RAG & Long-Context QA Trade-offsactive83
Sparse Attention for Long Inferenceactive80
KV-Cache Management for LLM Servingactive80
Heterogeneous GPU Serving Infrastructureactive68
Long-Context Retrieval & Reasoningactive66
KV Cache Compression & Evictionniche65
Test-Time Reasoning & Inference Scalingcrowded65
Efficient Audio Intelligence & Speech Modelsactive64
Gated & Structured Memory for Contextniche62
RAG Efficiency & Context Compressionactive60
Sparse Attention & KV Cacheactive59
Long-Context Evaluation & Engineeringactive58
Long-Context Analysis & Diagnosisactive55
Dynamic & Adaptive Attentionactive55
Diffusion LLM Inference Accelerationcrowded54
KV Cache Quantization for Reasoningcrowded53
Hierarchical Memory for Long-Video Reasoningactive45
Speculative Decoding for Efficient Inferenceniche41
Knowledge Editing and Lifelong LLM Updatescrowded40
LLM Inference Serving and Schedulingactive39
Quantization for Efficient Inferencecrowded38
Vision-Language Efficient Compressionactive38
Chain-of-Thought Reasoning Compressionactive38
On-Device LLM Memory Optimizationactive36
Many-Shot Translation & Adaptationactive36
Efficient Multi-Head Latent Attentionactive35
Language-Specific Encoders & Tokenizationactive35
Evidence-Grounded Long-Context QAactive34
LLM Inference Parallelism & Efficiencyactive34
retrieval, reranking, tool—33
Efficient Repository-Level Code Generationactive31
Long-Context Evaluation Benchmarkscrowded31
Efficient Long-Context Attentionactive30
LLM Agent Attack Surfaceactive29
Edge-Cloud LLM Serving Infrastructureactive27
Efficient Long-Context via Cache Compactionactive26
Efficient LLM Training and Inferencecrowded24
Jailbreak Detection and Defense for LLMscrowded23
Text-to-SQL & Semantic Knowledgeactive22
KV Cache & Latent Multi-Agent Communicationactive21
Efficient Inference via Pruning & KV-Cachecrowded21
Foundation Models & Capabilities Reportscrowded21
GPU Inference Performance & System Optimizationcrowded21
Efficient Long-Context KV Cachingactive19
Distributed Parallelism for Long-Context Trainingcrowded19
Positional Encoding & Length Generalizationactive15
Dynamic Scheduling for LLM Servingactive11
Activation Steering in LLMsactive10
Long Video Understandingactive10
Long-Form Document Summarizationactive9
Mid-Training Paradigm Surveyscrowded9
Continual Malware & Code Securityniche9
LLM-Assisted Literary Analysisniche8
Systems Optimization for LLM Servingactive8
Document-Level Machine Translation Contextactive8
Efficient Vision-Language Inferenceactive8
LLM Serving System Reliabilityactive7
Data selection for efficient trainingactive5
Legal and professional domain reasoning benchmarksactive5
Efficient sentiment and toxicity detectionactive5
Long-Context Code Generationactive5
Prompt Injection Defenseactive4
Carbon-Aware LLM Caching at Cloud Scaleniche3
Linguistic Steganography with Anchored Windowsniche1

Prosthetic Memory

1,471 papers · 52 clusters · 23 rising

The write-path attempts: external memory systems bolted onto frozen models. Long-horizon agent memory, conversational-memory benchmarks, self-evolving skill libraries, experience consolidation, temporal knowledge graphs. The tell is that its own benchmark cluster still asks whether ANY unified evaluation exists — dozens of parallel prosthetics, no convergent winner, and the substrate war (vector vs graph vs filesystem memory) openly unresolved. Rising and worth watching: experience consolidation across sessions, memorize-vs-recompute policies, and memory POISONING defense — memory integrity is becoming its own subfield.

Agentic Memory for Long-Horizon Tasksactive73
Benchmarking Long-Term Conversational Memoryactive73
Long-Term Agent Memory Systemsactive72
Self-Evolving Agents & Skill Learningactive71
Agentic Multi-Agent Orchestrationcrowded68
Hierarchical Planning & Memory-Augmented Agentsactive65
Agentic Coding & Multi-Agent Workflowsactive61
Self-Evolving Agentic Memory Systemsactive59
AGI Architecture & Embodied Cognitionniche53
Agent Context & Memory Managementcrowded52
Agentic Memory Systemscrowded48
Personalized Agent Simulationactive47
Neuro-Symbolic Agent Memorycrowded43
Graph-Memory GUI Agentsactive42
Long-Horizon Agent Memoryactive42
Agentic Systems & Tool-Using LLMsactive42
Temporal Graphs & Episodic Memoryniche42
Agentic RAG for Domain-Specific Reasoningcrowded42
memory, agent, agents—33
graph, generation, rag—33
Agent Memory Poisoning and Defenseactive29
Self-Improving Agents via Experienceactive28
Self-Improving Agentic Harnessesactive28
Agentic Search with Persistent Memoryactive26
Medical Agents and Clinical Reasoningactive26
Personalized LLM Generation & Memoryactive23
Agentic AI & Enterprise Governanceactive22
Clinical Reasoning & Adaptive Agentsactive21
Multi-Agent Document & Table Analysisactive21
Scheduling Multi-Agent LLM Servingcrowded20
Agentic Systems Security Attacksactive18
Policy Optimization Agent Alignmentactive17
Foundation Models & Efficient Scalingcrowded17
Long-Horizon Agentic State Reasoningactive14
Episodic Memory Multi-Agent Coordinationactive13
Episodic Memory in Coding Agentsactive12
Context Engineering for Human-AI Collaborationactive11
Persistent World Modelsactive11
Multi-Agent LLM Communication & Coordinationactive9
Affective Memory in Agentic Systemsactive9
Vision-Language Navigation Systemsactive6
Arbor: Autonomous Search & Navigationactive5
Scientific paper to visual media generationactive5
Autonomous AI runtime security and self-healingactive5
Fleet Provisioning and Autonomous Vehicle Systemsniche4
Graph-Based Representation Learningniche2
Agentic Planning via Tree Searchniche2
Foundation Models for 6G Networksniche2
Structured Task Inference Householdsniche1
Scene-Aware Visually-Grounded Speechniche1
LTL Verification Neural Stateful Agentsniche1
Multi-Agent Framework for Related Work Generationniche1

Changing the Machine

1,440 papers · 54 clusters · 17 rising

Architectural bets: linear and hybrid attention, parallelizable recurrence, positional-encoding theory, LoRA-based continual learning, model merging — and test-time training, the one genuinely native write-path idea in the corpus, rising fast. Meanwhile pure linear attention is cooling; its own cards ask why the theory never translated to practice. The hybrid position (mostly-linear with a little real attention) looks like the settling point.

Linear & Hybrid Attention Mechanismsactive90
LoRA and Continual Learningactive73
Efficient Long-Context Architectures & Recurrenceactive63
Continual Learning & Neural Memory Systemsactive61
Efficient Transformers & Scalingcrowded57
Mixture-of-Experts Serving & Routingcrowded54
Parallelizable RNNs & Linear Recurrenceactive54
Test-Time Training Adaptationactive52
Robust Fine-Tuning & Safety Alignmentcrowded50
Streaming Video Generationactive50
Multimodal Vision-Language Fusion & Editingactive46
Transformer Architecture & In-Context Reasoningactive44
Many-Shot In-Context Learningniche42
Block Diffusion Language Modelsactive42
Vision-Language-Action Robot Learningcrowded41
RoPE & Positional Encoding Theoryniche35
DNA/RNA Sequence Foundation Modelsactive32
State Space Models for Dynamic Graphsactive30
Test-Time Adaptation for Vision-Languageactive28
Gradient-Based Continual Learning Optimizationniche26
Out-of-Distribution Detection Continuallyactive25
Robust Multimodal Continual Learningactive24
Continual Sparse LLM Trainingactive22
Vision-Language Test-Time Adaptationactive22
Neural Operators for PDEsniche22
Working Memory & Attentionactive22
Linear Attention & Delta-Rule Mechanismsactive22
Multilingual Low-Resource Continual Learningniche21
Long-Horizon Latent Reasoningactive20
Spatial Memory for Vision-Language Agentsactive18
On-Policy Distillation Long-Contextactive18
Multilingual Multimodal Embeddingsactive18
KV-Cache Quantization Rotationsniche17
Transformer Mean-Field Dynamicsactive17
Model Merging & Catastrophic Forgettingactive15
Compositional Embeddings & Cognitive Architecturesniche15
Hybrid MoE Architectures for Efficient Reasoningcrowded14
Real-Time Vision-Language Autonomous Drivingactive14
Quantum State-Space Sequence Forecastingniche13
Efficient Scaling of Foundation Model Architecturescrowded12
Efficient Attention and Distillation for Multimodal Modelscrowded12
Active Perception & Visual Memoryactive10
Hyperbolic LLMs via Mixture-of-Curvatureactive10
Bayesian Optimization & Offline-Online Learningactive9
Open-Set Adaptive Object Detectionniche8
Corruption as Training Signalactive8
Transformer Training Across Domainsniche7
Grokking and Generalization Dynamicsactive6
Generative Visual Retrieval at Scaleactive6
Hybrid Offline-Online Routingactive6
LLM Negation & Agent Polarizationniche5
State-Space Model Depth Theoryniche4
Positional & Semantic Disentanglementactive4
Turing Completeness & Finite-Precision Recurrent Netsniche4

The SSM Diaspora

1,111 papers · 44 clusters · 8 rising

Mamba and state-space models in their diffusion phase: the core architecture debate is steady while applications spread into medical imaging, EEG, time series, graphs, and remote sensing — the classic sign of a technology moving from frontier to toolbox. Test-time adaptation lives here too, with its stability-over-long-deployment question still open. Mostly applications; little of it moves our question.

State Space Models & Sequencesactive82
Selective State-Space Models & Mambacrowded72
Test-Time Adaptationactive67
Time Series Forecasting & State-Space Modelsactive66
Test-Time Adaptationactive49
Vision Mamba Modelsniche48
Neural Kalman Filtering & State-Space Learningactive46
Mamba for Medical Image Segmentationniche41
Mamba & State-Space Graph Modelsniche36
Medical AI Agents & Reasoningcrowded36
Neural Decoding with State-Space Modelsniche34
Urban Spatio-Temporal Forecastingactive34
Physics-Informed Energy & Climate Forecastingactive33
detection, anomaly, continual—33
EEG State-Space Models for Motor Controlactive32
Neuromorphic Spiking Networksniche29
3D Pose & Adaptive Perceptionactive28
Mamba-Transformer Hybrid Architecturesactive27
State-Space Models for Efficient Inferenceactive23
Mamba Geometry & Causal Detectionactive21
Time-Series Foundation Models & Forecastingactive21
Mamba for Hyperspectral Anomaly Detectionactive20
Medical Image Segmentation with Mambaactive20
Spatio-Temporal Video Understanding via Memoryactive19
Latent Dynamics Under Epistemic Uncertaintyactive19
Test-Time Adaptation & Robustnessactive18
Wearable Physiological Monitoring via SSMactive17
Medical Segmentation Continual Adaptationniche16
Time Series Foundation Models Zero-Shotactive16
ECG Signal Processing & Cardiovascular AIniche14
Deepfake & Multimodal Detection Under Driftniche13
AI for Early Cognitive Decline Detectionniche12
Medical Image Segmentation & Transfer Learningniche9
Robust State-Space Filteringactive8
LLM-Driven IoT & System Securityactive8
Agentic Trajectory Reasoningniche7
GPU Kernel Optimization & Servingactive7
Causal Inference & Process Supervisionactive7
SAR Multimodal Disaster Assessmentniche6
Collaborative Multi-Agent Perceptionactive6
Temporal Graph Anomaly Adaptationactive5
Transformers for Remote Sensing & Commsniche4
Orbital Inference via Solar Computeniche1
Privacy-Preserving UAV Swarm Intrusion Detectionniche1

The Old War on Forgetting

904 papers · 25 clusters · 9 rising

Classical continual learning, still grinding: catastrophic-forgetting theory is RISING while replay engineering fades — the signature of a field maturing from recipes toward understanding. Plasticity-stability theory, analytic rehearsal-free methods, federated variants, and a small rising security seam (backdoors that persist through continual updates). The unsolved core is unchanged: cheap, durable, incremental weight updates remain out of reach.

Continual Learning & Parameter Efficiencycrowded100
Process-Supervised RL for Agentsactive82
Theory of Catastrophic Forgettingcrowded78
Multimodal Continual Learning & Replaycrowded70
Continual Reinforcement Learning & Adaptationactive60
Catastrophic Forgetting & Unlearningactive51
Safe Continual Reinforcement Learningactive50
Continual Learning & Adaptive Model Mergingactive44
Continual Learning with Sleep-Inspired Replaycrowded39
Federated Continual Learning Under Driftactive38
continual, replay, knowledge—33
Plasticity-Stability Tradeoff in Continual Learningactive32
Continual Graph Embedding Learningactive29
CLIP-Based Incremental Learningniche27
Federated Continual Learning Forgettingactive27
Knowledge Distillation for Continual Learningactive24
Backdoors and Poisoning in Continual Learningniche24
Reinforcement Learning Control & Adaptationactive23
Analytic Class-Incremental Learningactive21
Continual Learning via Adaptive Promptingactive20
Quantum Continual Learning Systemsniche8
Consciousness-Grounded Continual Learningniche8
Preference-Optimized Continual Learningactive8
Compositional Open-World Learningniche5
Partial Observability & Agent Learning Boundsniche3
Built and maintained by Stays Up. We build maps like this for other fields; see the field map, or everything else we measured. Privacy · Terms