RAG on GitHub, mapped: frameworks flat, agent memory tripled, and the repos worth knowing
Trending pages show you what got attention this week. They cannot tell you whether a topic is growing, whether a repository is maintained, or whether a star count came from one commit and a bot farm. I built a personal retrieval map over every GitHub repository with 100 or more stars created since March 2025, 45,406 of them, and asked it what is happening in retrieval-augmented generation. The answer is more useful than a trending list, and it comes with a filtered reading list.
Who this is for
Anyone choosing a retrieval stack this quarter. The choice usually starts from a trending page or a colleague's star count, and both measure attention rather than substance. Attention is real, but it concentrates: one repository per topic takes most of the stars, the rest of the topic is invisible, and a topic that stopped growing a year ago still looks alive if its one big repository keeps getting shared. A map of topics, with per-repository maintenance signals, answers the questions a trending page cannot: is this area growing, is this code maintained, and is the star count earned.
The instrument
Harvest: the GitHub search API, queried with the filter stars:>=100 fork:false and a creation-date window, sliced recursively into smaller date ranges wherever a slice hits the API's 1,000-result cap. Result: 45,406 repositories created between March 2025 and 9 September 2026, with stars, forks, language, topics, creation date, last push and archived flag. Embedding: the name, description and topics of each repository, embedded on a Mac with a local open model through ollama. Clustering: two-level k-means into 150 coarse and 1,516 fine topics. Cards: an LLM reads each cluster's representative repositories and writes a one-paragraph summary, a momentum label and an open question. Same recipe I run over 470,000 arXiv papers. Everything is local and repeatable.
Three reading rules, each learned the hard way. Rank, never threshold: absolute similarity scores in this space are not comparable across queries. Verify a cluster from its actual repositories, not from its card name. And treat stars as an attention signal: this data includes meme repositories above 100,000 stars, and about five percent of the corpus is single-commit star farming, which the created == pushed tell removes cleanly.
What the map says about RAG
| area | clusters | repos | created Mar to Nov 2025 | created Dec 2025 to Aug 2026 | pushed in last 60 days |
|---|---|---|---|---|---|
| RAG frameworks and pipelines | 5 | 199 | 93 | 106 | 52% |
| Agent memory layers | 8 | 308 | 78 | 228 | 67% |
| Graph RAG | 1 | 17 | 10 | 7 | 29% |
| Document parsing, PDF to text | 3 | 90 | 42 | 48 | 51% |
The two nine-month spans are equal in length, so the columns compare directly. RAG frameworks are flat. Agent memory, which is retrieval over an agent's own history plus write-back, nearly tripled, and it holds the largest projects in the whole neighbourhood: claude-mem at 93,568 stars, mempalace at 58,958, OpenViking at 36,246, and OpenViking's own description folds memory, knowledge RAG and skills into one store. Graph RAG is small and mostly parked: only 5 of its 17 repositories saw a push in the last 60 days, and most carry a conference tag, which is the shape of paper code rather than product. Document parsing is steady, and its card reads "fading, crowded", yet firecrawl's anydoc collected 20,895 stars in the four weeks after its August release. That is the attention-concentration pattern in one number: the topic is not growing, one repository in it is.
The honest caveat is that momentum on this map does not forecast. On the arXiv corpus I measured the rank correlation between early growth and later growth across 8,245 clusters and got 0.01. The map describes the present: what is crowded, what thinned, what is open. It does not tell you what wins next year.
The repositories worth knowing
The filter is three conditions at once: the repository sits in one of the clusters above, it was pushed within 60 days of the 9 September snapshot, and it is not a single-commit star farm. Stars are shown as attention, not as a ranking.
| repository | what it is | stars | created |
|---|---|---|---|
| VectifyAI/PageIndex | vectorless, reasoning-based document index for RAG | 35,591 | 2025-04 |
| HKUDS/RAG-Anything | multimodal all-in-one RAG framework | 23,272 | 2025-06 |
| StarTrail-org/LEANN | local, private RAG with a 97% smaller index; MLSys 2026 best paper | 12,923 | 2025-06 |
| StarTrail-org/PixelRAG | search over rendered pixels instead of parsed text | 9,910 | 2026-05 |
| feyninc/chonkie | chunking and ingestion library | 4,729 | 2025-03 |
| MinishLab/semble | code search for agents, 99% fewer tokens than grep and read | 6,036 | 2026-04 |
| vitali87/code-graph-rag | knowledge graph over a multi-language monorepo | 5,111 | 2025-06 |
| firecrawl/anydoc | Office, EPUB, CSV and PDF to Markdown, in Rust | 20,895 | 2026-08 |
| datalab-to/chandra | OCR for tables, forms and handwriting (no push since June) | 12,238 | 2025-10 |
| raphaelmansuy/edgequake | LightRAG-style graph RAG rewritten in Rust | 2,091 | 2025-12 |
| thedotmack/claude-mem | persistent context across sessions for coding agents | 93,568 | 2025-08 |
| MemPalace/mempalace | benchmarked open-source agent memory | 58,958 | 2026-04 |
| volcengine/OpenViking | context database unifying memory, RAG and skills | 36,246 | 2026-01 |
| vectorize-io/hindsight | agent memory that learns from outcomes | 23,311 | 2025-10 |
What to do with this
- Before adopting a retrieval library, look at its cluster, not its stars. A flat cluster with one big repository is a maintenance risk concentrated in one maintainer.
- If your RAG roadmap does not have a memory layer on it, the map says your users' next request will. The growth is there, the frameworks are steady.
- Read the last push date and the created-versus-pushed tell before the README. Half the repositories in the RAG clusters have not been touched in two months.
- Graph RAG is worth a prototype, not a dependency, until more than a third of its repositories are maintained.
Reproduce it
The interactive map, all 1,516 topics with their cards and representative repositories, is linked from this page. The recipe is four steps: date-sliced harvest through an authenticated gh token, local embedding of name plus description plus topics, two-level k-means, and one LLM-written card per cluster. It builds in an afternoon on a laptop, and changing the star threshold or the date window regenerates the tables above.
Written from running the same map over 470,000 arXiv papers and 109,000 bioRxiv preprints. Related: LangChain vs raw SDK, measured on the wire · what a million LLM tokens actually costs.