Renting an AI Model vs Owning One: What PewDiePie's Ajax Shows
PewDiePie, whose YouTube channel has 109 million subscribers, has announced his own AI model. It is called Ajax, and his official page describes it as “a fine-tuned Qwen 3.5 9B model, trained for Odysseus to be your always on agent.” The page says it handles search, web browsing, email and your calendar, “all your daily tasks completely privately.”
You cannot download it. In his words: “I've decided to release Ajax when it's ready instead of putting a timer.” So the model is not released yet, and there are no public benchmarks. He also says OpenAI banned his accounts twice while he was building it. That part of the story is the one worth keeping, because it shows the real difference between two kinds of AI. A model you rent can be switched off for you. A file on your disk cannot.
What you need to run a local ai model
People talk about “local AI” as if it were one thing. It is really three things stacked on top of each other, and you can swap each one out on its own. People already run all three at home today, on ordinary desktops and small servers.
The weights
The weights are the model itself: one large file, or a handful of files, full of numbers learned during training. When people say a model has “open weights”, they mean you can download those files. Once the files are on your disk, nobody can take them back. You can copy them, back them up and load them years from now, as long as you keep a runtime that can read them.
The runtime
Weights do nothing on their own. A runtime is the program that loads them into memory and turns your prompt into an answer. llama.cpp is the base layer for many home setups and runs on almost anything, including a plain CPU. Ollama, one of the tools built on that foundation, wraps the whole process in a simple command and a local API. vLLM is built for GPUs and for serving many requests at once, and it is the usual pick when you have a proper graphics card and want speed.
The harness
A model that only chats is useful, but an agent needs hands. The harness is the software that gives the model tools, such as a browser, a search box or access to your inbox and calendar, along with rules about when to use them. Odysseus is that layer for Ajax. It is a real open-source project, published on GitHub as a “Self-hosted AI workspace” under the AGPL-3.0 licence. It was first published on 31 May 2026 and is still being actively developed. Harnesses like this usually talk to the runtime over a local API, which means the model underneath can often be swapped.
The honest hardware question
A 9B model has roughly nine billion parameters. All of them have to sit in memory while the model runs, and how much memory that takes depends on the precision you store them at.
The vLLM project's maintained recipe for Qwen3.5-9B lists roughly 22 GB of VRAM at BF16, which is full 16-bit precision, and roughly 11 GB at FP8. Treat those as planning numbers. Longer conversations need more memory, because the model holds the whole context in memory while it works.
This is where quantisation comes in. Quantising a model means storing each number with fewer bits. Going from 16-bit to 8-bit roughly halves the memory, which is the gap between those two figures. The 4-bit formats that are common with llama.cpp and Ollama roughly halve it again. The cost is a small loss in quality. On everyday tasks you may not notice it, but it is not zero, so test it on your own work rather than trusting anyone's word, mine included.
A GPU is not mandatory. llama.cpp runs on a plain CPU with enough system RAM, and Apple Silicon machines make good use of their shared memory. What you give up is speed. On a CPU, a 9B model will answer, but you will watch the words arrive one at a time. For a background agent that sorts mail overnight, that can be fine. For a chat window you are staring at, it gets tiring. A GPU is optional, but on a CPU you will need some patience.
If you want to see the same question asked about a much larger model, I worked through it in the post on whether you can self-host DeepSeek V4.1 Flash. The logic is the same. Only the numbers get bigger.
What “completely privately” actually means
On your own hardware, “completely privately” has a plain meaning you can check. Your prompt is processed on your machine. The files you hand the agent, the email it reads and the calendar it edits all stay on your machine. Nothing goes to a model provider, so there is nothing for a provider to log, sell or switch off.
There is one honest footnote. An agent that searches or browses the web still sends those searches out to the web. What stays home is the model's work on your data, and that data itself.
You can test the claim yourself. Block the runtime at your firewall, or pull the network cable, and see what still works. I wrote more about that setup in the guide to keeping your AI on your own box.
The honest cost is capability. A 9B model is not a frontier model. It will be weaker than the largest hosted models at long reasoning, hard code and obscure facts. Convenience is mostly solved, because getting a runtime going is a short install now. What you give up is some of the raw ability of the biggest models, and what you get in return is a model you control.
Why someone with money and access still did this
PewDiePie can afford any hosted plan he wants. By his own account, that did not protect him. He says OpenAI banned his accounts twice while he was building Ajax. He says the first ban was lifted, and that the email for the second ban cited “distillation”, which means using one model's outputs to train another. He says the email gave no specific examples. OpenAI has made no public statement about it.
I do not know the details, and the lesson does not depend on them. From your side, a lost hosted account is an outage you cannot appeal on your own schedule. Terms change, models get retired and accounts get flagged by automated systems. When that happens, everything you built on top of the account stops working.
Owning the files is the only protection that survives a policy change. If the weights are on your disk and the runtime is on your machine, a change to someone else's rules does not reach you. That does not make local models last forever. Projects stall, runtimes drop old formats, and nobody promises updates indefinitely. But the file you have will keep working the way it did yesterday, and no rented service can promise that.
What to do this week
Keep it small. The goal is not a full agent stack by Sunday. It is one model doing one real job on hardware you own.
- Pick one model in the 7-9B range. Qwen 3.5 9B, the base model Ajax is built on, is a sensible place to start, but any well-known open model of that size will teach you the same things. Get it in a quantised format your runtime can load, such as a GGUF file for llama.cpp or Ollama.
- Pick one runtime. Choose Ollama if you want the shortest path, llama.cpp if you want to see every setting, or vLLM if you have a GPU with enough memory and want speed.
- Point one real task at it. Skip the trick questions. Use something you actually do: summarise yesterday's email, rewrite a paragraph or turn rough notes into a to-do list. Use the same prompt every day for a week.
With llama.cpp, starting a local server with a built-in chat page looks like this:
llama-server -m ./your-model-Q4_K_M.gguf --port 8080Then open http://localhost:8080 in your browser. With Ollama, it is ollama run <model-name> in a terminal.
While it runs, watch three things:
- Memory use. Watch
nvidia-smion a GPU orhtopon a CPU-only box while the model answers. If you are close to the limit, try a smaller quantisation before buying hardware. - Speed. Most runtimes report tokens per second. The question is whether it is fast enough for this one task, not whether it feels like a hosted chatbot.
- Quality. Read the outputs honestly. If they are good enough for this one job, you now own a working piece of your stack. If not, try a different quantisation or a different model before giving up.
Write down what you find. A week of notes on one real task will tell you more than any leaderboard.
If you want it done for you
He says Ajax will come out when it is ready. The stack it runs on is already here, and you can own every layer of it today. If you would rather have someone set up the runtime, the harness and the network rules on your own hardware and harden them properly, that is the work I do. See the services page for details, or find me directly.
Fiverr: hiteshsaini459 · Upwork: hiteshsaini25