The AI world uses “open source” loosely enough that the term has stopped carrying information. A model described as open source might mean you can download and modify it freely, or that you can use it commercially only under a headcount threshold, or merely that a paper describing it was published.
The distinctions matter, because they determine what you can legally build and where your data goes.
Three categories #
Closed. You access the model through an API. You never possess the weights. The provider controls what version you are talking to, what it costs, and whether it continues to exist. Most frontier models work this way.
Open weight. The trained parameters are downloadable. You can run the model on your own hardware, fine-tune it, and inspect it. What you generally do not get is the training data or the full training code, so you could not reproduce the model from scratch. Most models people call “open source” are actually this.
Open source, strictly. Weights, training code, and training data are all public, under a license that permits use and modification without restriction. This is rare. A handful of research efforts qualify, and they are typically not competitive with the frontier.
The middle category is where almost everything interesting lives, and calling it open source has annoyed enough people that “open weight” has become the preferred term.
What the licenses actually say #
This is where projects get into trouble, because people download a model and assume they can do anything with it.
Apache 2.0 and MIT are genuinely permissive. Use commercially, modify, redistribute, no strings beyond attribution. Several major model families ship under these, and they are the ones to prefer if you have any commercial intent.
Custom “community” licenses are written by the releasing company. Common provisions include restrictions above a monthly active user threshold, naming or attribution requirements for derivative models, prohibitions on using outputs to train competing models, and acceptable-use policies. Perfectly usable for most purposes, and you have to read them rather than assume.
Non-commercial licenses mean research use only. Some very capable models fall here, and shipping a product on one is a straightforward license violation.
So read the actual license before you build on a model. Not the announcement blog post, not the model card summary, the license file. This takes ten minutes and has saved people from expensive mistakes.
What you get from open weight #
You can run it locally, on your hardware, in your infrastructure, on a phone. This is the whole foundation of on-device AI, and it is impossible with a closed model by definition.
Your data stays where you put it. No API call means no data leaving your environment. For regulated industries, sensitive material, or anything with a compliance requirement attached, this is frequently the deciding factor and nothing else in the comparison matters.
The version does not change under you. A closed model can be updated, deprecated, or retired, and prompts tuned against one version can behave differently on the next. A downloaded model is frozen. It will behave the same in three years as it does today, which for a production system is worth a great deal.
You can fine-tune it. Full access to the weights means you can adapt the model to your domain, your format, your task. Closed providers offer limited fine-tuning, and it is their process on their terms.
There is no per-token cost. You pay for hardware instead. At high volume this is dramatically cheaper. At low volume it is dramatically more expensive, because a GPU sitting idle still costs money while an API you are not calling costs nothing.
And you can inspect it: look at the weights, probe the internals, understand what it does. Mostly this matters for research, and occasionally for safety and interpretability work that is impossible on a black box.
What you give up #
Being fair to the closed side, because there are real advantages.
Frontier capability. The best closed models are generally ahead of the best open-weight models on hard reasoning, long-context work, and complex code. The gap has narrowed substantially, and it has not closed, and the leading labs have strong commercial reasons not to release their best weights.
Operational burden. You now run inference infrastructure: GPUs, serving stack, batching, scaling, uptime, upgrades. This is a real engineering function with real headcount attached, and teams consistently underestimate it.
Cost at low volume. A GPU capable of serving a large model costs meaningful money per month whether or not anyone uses it. API pricing is usage-based, so a low-traffic product is far cheaper on an API.
Safety tuning and abuse handling. Closed providers do a lot of work on this and it is included. Running your own model means the safety behavior is whatever the released model has, and any additional filtering is your problem.
Multimodal and tool ecosystems. Closed providers ship integrated capabilities, function calling implementations, and tooling. Open-weight equivalents exist and are less polished.
The cost crossover #
The rough shape of the economics. At low volume, thousands of requests a month, the API wins decisively, because self-hosting is paying rent on an idle machine. At medium volume the two are roughly comparable and the decision comes down to data control and latency rather than cost. At high volume, millions of requests, self-hosting typically wins by a wide margin, because you are amortizing fixed hardware cost across enormous usage. And on-device, the marginal cost is zero for the developer and the user’s own hardware does the work, which is why free apps with no subscription are possible at all and the reason I build things this way.
There is also a fourth option people forget: an open-weight model served by someone else. Several providers host open-weight models behind an API at prices well below the frontier closed models. You get low cost and no infrastructure, and you give up the data-control argument, since the request still leaves your environment. For a lot of applications this is the correct answer and it is undersold. That is more or less the case for using cheaper models by default.
How to actually decide #
Work through this in order and you will usually land in the right place.
Do you have a hard data-residency requirement, meaning regulated data, client confidentiality, or a contractual constraint? If yes, open weight, self-hosted or on-device. Nothing else on this list matters.
Does it need to work offline? Then open weight, on-device.
Do you need frontier reasoning, meaning the hardest problems, long chains of reasoning, complex code across many files? Closed models still lead here.
Is your volume high and your task narrow? A fine-tuned small open-weight model often beats a general frontier model on a specific task, at a fraction of the cost. This is a genuinely underrated option.
Are you prototyping? Use an API. Get it working, learn what you actually need, and revisit the hosting question once the product exists. Building infrastructure before you have validated the product is the classic way to waste a quarter.
The trend worth watching #
Open-weight models are roughly twelve to eighteen months behind the closed frontier, and that gap has been fairly stable for a while now. What has changed dramatically is the floor: the quality you can get from a model small enough to run on consumer hardware has improved much faster than the frontier has.
That is the shift that matters for most people. The question stopped being “can an open model do this at all” and became “is a model I can run myself good enough for this specific task.” For a growing share of tasks the answer is yes, which changes the economics of building software on top of AI in a fairly fundamental way.
Related: running AI models locally for the practical setup, quantization for how these models get small enough to run, and on-device AI for the phone case specifically.