6 GitHub Repos Worth Self-Hosting This Week: The Agent Stack

Agent memory, a Tencent Cloud AI assistant, a Workers-compatible platform and three more repos trending on GitHub this week. What each one needs and which to self-host first.

6 GitHub Repos Worth Self-Hosting This Week: The Agent Stack

For two years the agent story lived in someone else's cloud. This week GitHub's trending page told a different story, because it is full of agent tools you can run on your own box: memory, a multi-user assistant, a Workers platform, a fast decision model, a cost dashboard and a push inbox. Here are the six I would actually look at.

Worth self-hosting this week: Hindsight (agent memory), Octop (a private AI assistant), open-compute (Workers on your own box), CLM-8B (fast agent decisions), Agent Console (token costs) and Boop (push alerts you own).

1. Hindsight: memory for your agents

Hindsight is the top trending repo on GitHub this week, with roughly 11,000 new stars in seven days and about 38,400 in total. It is an MIT-licensed memory system that stores World facts, Experiences, Observations and Mental models, so an agent learns over time instead of replaying chat history.

It ships as one Docker container with Postgres inside and speaks to 25+ LLM providers, including local ones (ollama, lmstudio, llamacpp) and an existing Claude Pro/Max or ChatGPT subscription.

docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY -v hindsight-data:/home/hindsight/.pg0 ghcr.io/vectorize-io/hindsight:latest

The extras matter more than the server: an MCP server, an LLM wrapper (pip install hindsight-litellm) and a coding-agent installer that gives Claude Code, Codex and Cursor per-repo project memory. Docs are on hindsight.vectorize.io.

The catch: the state of the art LongMemEval result is the project's own claim, so test it on your own agents first.

2. Octop: a self-hosted AI assistant for the whole house

Octop is Tencent Cloud's multi-user, multi-agent assistant. MIT licensed, about 5,400 stars, version 1.0.2b4, Python 3.12+, SQLite by default with PostgreSQL optional. Everything runs as one restart-safe process that rebuilds its state from the database on boot.

You chat through the web dashboard, desktop apps, the CLI or IM channels like Telegram, Discord, WeChat and DingTalk. There is also a RAG knowledge base, natural-language cron jobs, browser automation, a remote desktop view, and ACP to hand coding work to OpenCode or Claude Code behind permission gates.

pip install octop

The official installer builds an isolated venv under ~/.octop/ with uv and leaves system Python alone. A docker-compose.yml ships with it too. More at octop.cloud.

The catch: it is beta, and that is a lot of doors to lock, so run it on a private network as a family or lab deployment first.

3. open-compute: Cloudflare Workers on your own machine

open-compute is a Workers-compatible platform in one Rust binary, Apache-2.0, about 1,300 stars. You get KV, D1, R2, Durable Objects, Queues and Workflows on a single machine, backed by SQLite plus local or S3-compatible storage. No Kubernetes, no Redis, no service mesh.

After install you run ocd setup --yes and ocd status, then deploy with ocd wrangler deploy using a project-local wrangler 4.138.0. Worker code executes on a pinned, checksum-verified workerd fork.

curl -fsSL https://open-compute.dev/install.sh | sh

As with any piped installer, read the script on open-compute.dev first.

The catch: it is young and single machine only. If you already run Coolify, or you are happier with plain containers, the Portainer 3.0 news is the more relevant reading for you this month.

4. CLM-8B: decisions instead of paragraphs

CLM-8B is the odd one out: Stanford and Nvidia's System One decision model, Apache-2.0, about 2,000 stars, published on 23 September 2026. It answers typed questions about a state (Noul, Choice, Score), so your agent gets a decision instead of a paragraph to parse.

It was pre-trained on 60 million Nemotron Q&A pairs and post-trained on 1 million agentic trajectories. It reports up to 9x lower latency than Jev on computer-use, gaming and tool-calling tasks, and 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE as a verifier after light fine-tuning.

pip install contrastive-lm

Then run vllm serve Qwen/Qwen3-8B as the encoder plus clm-serve, which pulls a 75 MB reference head on first run. States over 2048 tokens are truncated unless you raise both limits together.

The catch: it needs an Nvidia GPU with vLLM, and it only pays off if you are building agents that make many small decisions. For a general local model, start with my DeepSeek V4.1 Flash post instead.

5. Agent Console: where did the tokens go?

Agent Console is local-first observability for coding agents. MIT, about 520 stars, Node.js 22 or newer. It reads the Claude Code and Codex transcripts already sitting on your machines and shows tokens, cache reads and writes, models and list-price cost, per session and per machine.

An optional self-hosted team hub gives you one view across laptops and build boxes, and it can ingest Claude Code OpenTelemetry or LiteLLM metrics, expose a Prometheus /metrics endpoint and ship a Grafana dashboard. There is no telemetry at all, and presenting mode replaces project and machine names with stand-ins so screen sharing is safe.

The catch: it is a visibility tool. It shows you the waste, it does not remove it.

6. Boop: a tiny push inbox you own

Boop is the smallest project here and maybe my favourite: one Go binary, one SQLite file, one Docker container, MIT, about 820 stars. Your apps POST an event with a project API key, and the server redacts it, stores it and pushes straight to Apple's APNs with your own .p8 key. No hosted relay, no account, no telemetry.

The push carries only the title, body and event id, and the phone fetches the rest from your server. That is a good fit for backup alerts, failed cron jobs and deploy pings.

docker compose up -d --build

Everything lands in ./data/boop.db, so that one file is your whole backup.

The catch: the iOS app is source only, so you build and sign it yourself, and pushes need an Apple developer setup. This is a weekend project, not a five minute install.

How these compare

ProjectWhat it doesLicenseWhat it needs
HindsightAgent memoryMITDocker, LLM key or local model
OctopMulti-user AI assistantMITPython 3.12+, SQLite
open-computeWorkers-compatible platformApache-2.0One machine, one binary
CLM-8BFast decision modelApache-2.0Nvidia GPU, vLLM
Agent ConsoleAgent cost trackingMITNode.js 22 or newer
BoopPush notification inboxMITDocker, Apple developer setup

Four more I have not reviewed here. For last month's picks, see my September roundup.

  • ZCode: Z.ai's coding agent harness, Apache-2.0, about 6,900 stars.
  • WeKnora: turns your documents into a RAG, an agent and a self-maintaining wiki, about 30,700 stars.
  • golive-skill: an Agent Skill plus CLI that takes an agent-built app live on your own hosting, database and domain accounts, MIT, about 1,000 stars.
  • AgentVerse-OS: a personal cloud OS for a developer and their agents on one Ubuntu server, alpha, about 1,000 stars.

What I would run first

Start with Hindsight. It is one container, it works with local models, and memory is the piece most home agent setups are missing. If you use Claude Code or Codex daily, add Agent Console the same weekend, because measuring your token spend costs you almost nothing.

Boop comes next if you own an iPhone and do not mind building the app yourself. Octop deserves a test in a throwaway VM, and my Proxmox beginners guide shows how to set one up. Keep it off the public internet while it is still beta.

Skip open-compute unless you already write Workers code, and skip CLM-8B unless you have an Nvidia GPU and are building agents.

Frequently asked questions

Do I need a GPU to self-host these projects?

Only CLM-8B needs one, because it runs on vLLM with an Nvidia GPU. Hindsight can use a hosted provider through an API key, so a local GPU only matters if you pick ollama, lmstudio or llamacpp. The other four do not list a GPU requirement.

Is agent memory worth it for a home lab?

If you run agents regularly, yes. Hindsight stores facts, experiences, observations and mental models, so an agent learns over time instead of replaying chat history every session. It ships as one Docker container with Postgres inside, so the setup cost stays small.

What hardware do these six projects need?

Five of them are small installs. Boop is one Go binary with SQLite, open-compute is one Rust binary on a single machine, Octop needs Python 3.12+ and SQLite, Agent Console needs Node.js 22 or newer, and Hindsight needs Docker. CLM-8B is the heavy one because it needs an Nvidia GPU running vLLM.

If you want one of these running properly behind a reverse proxy with backups, that is the kind of work I do for clients, so have a look at my services. Otherwise, pick one this weekend and tell me how it went.