Stays Up

RAG on GitHub, mapped: frameworks flat, agent memory tripled, and the repos worth knowing

Trending pages show you what got attention this week. They cannot tell you whether a topic is growing, whether a repository is maintained, or whether a star count came from one commit and a bot farm. I built a personal retrieval map over every GitHub repository with 100 or more stars created since March 2025, 45,406 of them, and asked it what is happening in retrieval-augmented generation. The answer is more useful than a trending list, and it comes with a filtered reading list.

Who this is for

Anyone choosing a retrieval stack this quarter. The choice usually starts from a trending page or a colleague's star count, and both measure attention rather than substance. Attention is real, but it concentrates: one repository per topic takes most of the stars, the rest of the topic is invisible, and a topic that stopped growing a year ago still looks alive if its one big repository keeps getting shared. A map of topics, with per-repository maintenance signals, answers the questions a trending page cannot: is this area growing, is this code maintained, and is the star count earned.

The instrument

Harvest: the GitHub search API, queried with the filter stars:>=100 fork:false and a creation-date window, sliced recursively into smaller date ranges wherever a slice hits the API's 1,000-result cap. Result: 45,406 repositories created between March 2025 and 9 September 2026, with stars, forks, language, topics, creation date, last push and archived flag. Embedding: the name, description and topics of each repository, embedded on a Mac with a local open model through ollama. Clustering: two-level k-means into 150 coarse and 1,516 fine topics. Cards: an LLM reads each cluster's representative repositories and writes a one-paragraph summary, a momentum label and an open question. Same recipe I run over 470,000 arXiv papers. Everything is local and repeatable.

Three reading rules, each learned the hard way. Rank, never threshold: absolute similarity scores in this space are not comparable across queries. Verify a cluster from its actual repositories, not from its card name. And treat stars as an attention signal: this data includes meme repositories above 100,000 stars, and about five percent of the corpus is single-commit star farming, which the created == pushed tell removes cleanly.

What the map says about RAG

areaclustersreposcreated Mar to Nov 2025created Dec 2025 to Aug 2026pushed in last 60 days
RAG frameworks and pipelines51999310652%
Agent memory layers83087822867%
Graph RAG11710729%
Document parsing, PDF to text390424851%

The two nine-month spans are equal in length, so the columns compare directly. RAG frameworks are flat. Agent memory, which is retrieval over an agent's own history plus write-back, nearly tripled, and it holds the largest projects in the whole neighbourhood: claude-mem at 93,568 stars, mempalace at 58,958, OpenViking at 36,246, and OpenViking's own description folds memory, knowledge RAG and skills into one store. Graph RAG is small and mostly parked: only 5 of its 17 repositories saw a push in the last 60 days, and most carry a conference tag, which is the shape of paper code rather than product. Document parsing is steady, and its card reads "fading, crowded", yet firecrawl's anydoc collected 20,895 stars in the four weeks after its August release. That is the attention-concentration pattern in one number: the topic is not growing, one repository in it is.

The honest caveat is that momentum on this map does not forecast. On the arXiv corpus I measured the rank correlation between early growth and later growth across 8,245 clusters and got 0.01. The map describes the present: what is crowded, what thinned, what is open. It does not tell you what wins next year.

The repositories worth knowing

The filter is three conditions at once: the repository sits in one of the clusters above, it was pushed within 60 days of the 9 September snapshot, and it is not a single-commit star farm. Stars are shown as attention, not as a ranking.

repositorywhat it isstarscreated
VectifyAI/PageIndexvectorless, reasoning-based document index for RAG35,5912025-04
HKUDS/RAG-Anythingmultimodal all-in-one RAG framework23,2722025-06
StarTrail-org/LEANNlocal, private RAG with a 97% smaller index; MLSys 2026 best paper12,9232025-06
StarTrail-org/PixelRAGsearch over rendered pixels instead of parsed text9,9102026-05
feyninc/chonkiechunking and ingestion library4,7292025-03
MinishLab/semblecode search for agents, 99% fewer tokens than grep and read6,0362026-04
vitali87/code-graph-ragknowledge graph over a multi-language monorepo5,1112025-06
firecrawl/anydocOffice, EPUB, CSV and PDF to Markdown, in Rust20,8952026-08
datalab-to/chandraOCR for tables, forms and handwriting (no push since June)12,2382025-10
raphaelmansuy/edgequakeLightRAG-style graph RAG rewritten in Rust2,0912025-12
thedotmack/claude-mempersistent context across sessions for coding agents93,5682025-08
MemPalace/mempalacebenchmarked open-source agent memory58,9582026-04
volcengine/OpenVikingcontext database unifying memory, RAG and skills36,2462026-01
vectorize-io/hindsightagent memory that learns from outcomes23,3112025-10

What to do with this

Reproduce it

The interactive map, all 1,516 topics with their cards and representative repositories, is linked from this page. The recipe is four steps: date-sliced harvest through an authenticated gh token, local embedding of name plus description plus topics, two-level k-means, and one LLM-written card per cluster. It builds in an afternoon on a laptop, and changing the star threshold or the date window regenerates the tables above.

Written from running the same map over 470,000 arXiv papers and 109,000 bioRxiv preprints. Related: LangChain vs raw SDK, measured on the wire · what a million LLM tokens actually costs.