Who Gets to Test an AI Model Depends on Who Holds the Weights

Who Gets to Test an AI Model Depends on Who Holds the Weights

On Thursday 1 October 2026, the Wall Street Journal reported that OpenAI had parted ways with three researchers. OpenAI confirmed the departures in a statement: "We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information." The Journal reported, citing a person familiar with the matter, that the material went to an outside organization that evaluates AI models. OpenAI has not named the organization or said what the information was.

The BBC understands the three were not dismissed for raising safety concerns. It understands they were dismissed for allegedly mishandling sensitive information, and that at least two of them worked in safety research. The news came during a week of intense public debate about AI safety, which included a White House meeting of technology executives and a voluntary self-regulatory agreement that some experts criticized. I will leave the personnel story to the business desks. For anyone running their own stack, the part that matters is who gets to evaluate an AI model in the first place.

What open-weight models actually change

A model's weights are the trained numbers that make it work. With open-weight models, the company publishes those numbers as a file you can download. Once the checkpoint is on your disk, you can run it with tools like llama.cpp, Ollama or vLLM, and nobody has to approve what you do next.

That changes who gets to check the model. A university lab, a security researcher or a hobbyist with one GPU can test it, try to break it, and publish what they find. They do not need an account, an access agreement or anyone's permission. If someone doubts the results, they can download the same file and run the same tests.

A closed model works differently. You reach it through an API or a chat app, and the weights never leave the vendor's servers. Outside testing still happens, but it usually runs through an access agreement with the company, and contracts and internal policy decide what testers can share. I am not judging anyone in this week's story. My point is about structure: with a closed model, the vendor decides whether outsiders can check it, and sharing the details can become a policy question.

With open weights, outside scrutiny is the default. That is how the system is meant to work, and nobody has to leak anything for it to happen.

One caveat on terms. Open weights are not the same as fully open source. Most releases do not include the training data or the full training code, so you can test what the model does, but you cannot rebuild it from scratch. For evaluation, the weights are what count.

How to evaluate a model on your own hardware

Public leaderboards are a useful starting point, but they test someone else's tasks. The only evaluation that tells you whether a model fits your work is one you run on your own work. This is the routine I use. Setting it up takes an afternoon, and after that each new model takes a few minutes.

  1. Write your prompt set. Collect 20 to 30 prompts from real tasks: the emails you summarize, the configs you ask about, the documents you pull fields from. Next to each one, write a short note on what a good answer looks like.
  2. Pull a model and pin the version. With Ollama, run ollama pull followed by the model name and tag. Then run ollama list and write down the ID next to the tag. A tag can be updated later, and you want to know exactly which file you tested.
  3. Fix your settings. Use the same temperature, context length and system prompt for every run. A temperature of 0 makes results easier to compare between models.
  4. Run every prompt and save everything. A short script is enough. The one below is the script I use for this step.
  5. Grade the answers yourself. Mark each answer pass, partial or fail against your notes. When you compare two models, hide which model wrote which answer until you have finished grading. Speed matters, but a fast wrong answer is still wrong.
  6. Keep the results. Put the prompt file, the model ID, the settings and the results in a Git repository. When a new version comes out, run the same set again and compare. After a few months, you have your own benchmark history for the work you actually do.

This script reads prompts.txt one line at a time and sends each prompt to a local Ollama server. For every prompt it saves the answer, the total time in seconds and the generation speed in tokens per second. It needs curl and jq.

MODEL="your-model:tag"
while IFS= read -r p; do
  jq -n --arg m "$MODEL" --arg p "$p" '{model: $m, prompt: $p, stream: false, options: {temperature: 0}}' |
    curl -s http://localhost:11434/api/generate -d @- |
    jq -c --arg p "$p" '{prompt: $p, response, seconds: (.total_duration / 1e9), tok_per_s: (.eval_count / (.eval_duration / 1e9))}'
done < prompts.txt >> "results-$(date +%F).jsonl"

Watch memory use while it runs. Run nvidia-smi on an NVIDIA card, or free -h on a CPU-only box, to see whether the model fits in memory or is spilling into slower memory.

You need nobody's permission for any of this, and that is the point. The same thing that lets you test a model in your spare room also makes third-party benchmarking and independent safety testing possible at all. Researchers can probe an open-weight model for weaknesses, publish their method, and let others repeat it on the same file. With a closed model, outsiders can only test what the vendor exposes, under the vendor's terms.

The honest trade-offs

Open weights do not come free, and it would be misleading to pretend they do.

  • The best closed models are still ahead on many hard tasks. The gap shows most on long multi-step reasoning, large coding jobs and tricky analysis. For everyday work like summarizing, drafting, sorting and extracting, a good open-weight model is often close enough. Your prompt set will show you where your own work falls.
  • You pay in hardware and time instead of a subscription. A capable model needs a GPU with enough memory, or a lot of patience on CPU. You also pay for electricity and for the evenings you spend on drivers, containers and updates.
  • An open license is not a promise of quality. Anyone can publish weights. Some releases are excellent, some are rushed, and some fine-tunes on model hubs have no clear origin. Licenses also differ, and some limit commercial use, so read the terms before you build anything on a model.

You trade effort for control and transparency. For some workloads that trade is clearly worth it. For others, a closed API is the sensible choice, and that is fine.

What stays your responsibility

When you run the model yourself, the responsibility moves to you. That is mostly good news, because the things that matter become things you can see and change.

  • The logs. Decide what gets recorded: prompts, outputs, and which user or service made each call. Keep the logs on your own storage, rotate them, and protect them. They will contain the same sensitive material you were trying to keep in-house.
  • The updates. A new model version can behave differently even when the name stays the same. Pin versions in your Compose file or scripts, and run your prompt set before you swap anything in production.
  • The switch. You should be able to stop the model, cut its network access or roll back to the last good version in under a minute. If doing that needs a meeting, you do not really have a switch.

I covered these controls in more detail in the post on securing self-hosted AI agents, which explains the sandbox, the logs and the kill switch. If privacy is your main reason for running models locally, the guide to keeping the data loop closed shows where data can still slip out of a local setup.

The real risk is a model you cannot stop or inspect, wherever it runs. Open weights let you do both. Actually doing it is up to you.

Where to start this week

You do not need to move everything at once. Pick one job and test it properly.

  1. Choose one workload that touches sensitive material. Good candidates include summarizing client emails, pulling fields out of contracts or invoices, and answering questions about internal documents.
  2. Build a prompt set from that workload using the steps above. Twenty real examples are enough to start.
  3. Run an open-weight model on your own box for two weeks alongside whatever you use today. Do not switch over yet. Just run both.
  4. Compare the results. Look at your pass rate, how often you had to fix the output, response times and how much of your time the setup took. Write the numbers down.
  5. Then decide. You can move the workload, leave it where it is, or split it so the sensitive parts stay local.

Whatever you choose, you will have made the call with your own evidence, on your own hardware, without asking anyone for access. That habit is worth building.

If you want a hand

If you would rather have someone set this up, or harden a stack you already run, I do that work for clients. You can see what I offer on the services page.

Fiverr: hiteshsaini459 · Upwork: hiteshsaini25