The subnet is how BASE gets built
CONTENTS
BASE has no research payroll. The work that becomes our products is posted as challenges on Subnet 100, attempted by anyone who wants to attempt it, scored by validators, and paid on the result.
This is an account of what that buys, what it ships in the next two weeks, and the one place where it leaves us behind.
The subnet instead of a payroll
BASE is Subnet 100 on Bittensor. Arenas define tasks, miners submit agents, validators score what those agents actually do, and at the end of every scoring round the scores settle into each miner's share of the payout. That mechanism is the product we sell. It is also how the company staffs itself, and the second fact is the one this post is about.
A conventional AI company hires researchers and engineers, pays them a salary, and hopes the output is worth the wage bill. Hiring is the bottleneck in that arrangement. You can interview only so many people, you can afford only so many of them, and once someone is on the payroll their compensation is fixed whether the week went well or badly.
A subnet inverts the order of those steps. The work is posted as a challenge, anyone in the world can attempt it, and the payout goes to whoever scores highest. There is no headcount to approve and no interview loop to run. At any hour, the number of people working on a BASE problem is whatever the incentive attracts, and they work continuously rather than during the hours our office happens to be awake. No salaried team of our size can produce that volume of attempts.
What the score decides
The permissionless part is where people get uneasy, so it is worth stating the mechanism rather than asserting the conclusion.
Nobody is hired. Nobody is vetted before they start. A miner registers on the subnet, submits an agent against an arena task, and that is the entire entry procedure. Everything after that is the filter, and the filter runs on a fixed loop. Each pass through the loop, an epoch in the protocol's vocabulary, is one scoring round in four steps:
- Run. Validators execute the submission against the arena's published evaluation suite.
- Aggregate. The scores from every run in the round add up.
- Normalise. Each total is measured against the rest of the field and becomes a weight, which is simply the miner's share of the round's payout.
- Emit. The round pays out in proportion to those weights.
Nobody is hired and nobody is vetted. The score decides who gets paid, and it decides again every round.
Most submissions are not useful. The mechanism does not require them to be. A miner who contributes nothing earns nothing, and that costs us nothing; a bad hire costs a year of salary and a difficult conversation. Because the filter runs every round rather than once at the door, a miner who was strong last month and is coasting this month stops earning this month.
That property is what drives the research. Compensation tracks measured performance closely enough that the return on finding a genuinely better method lands with the person who found it, in the round they found it. That is a sharper reason to try something unusual than a performance review twelve months away. It also means we do not have to guess in advance which people are the right ones, which is the guess hiring is built on and the one it most often gets wrong.
The honest caveat is that all of this rests on the rubric. If an arena measures the wrong thing, we pay for the wrong thing accurately, at scale, and on schedule. That is why scoring rules are published with the task in the docs, and why improving them is continuous work rather than a setup step.
From scored work to product
A benchmark that only produces a leaderboard has to be funded by something else. The second thing the subnet gives us is a route from the work to something we can sell.
Miner output is not a paper. It is running code, evaluated under conditions we control, with a number attached that says which version is better. When an arena is aimed at a problem a product needs solved, the winning submissions are the engine of that product.
That closes a loop most research programmes leave open. The payout funds the work, the work becomes a product, the product earns revenue, and revenue pays for distribution and for more work. The subnet is not a marketing channel bolted onto a company. It is the production line.
Ink ships on 10 August
Ink is the first project out of that arrangement. It ships on Monday 10 August 2026, three days from the date on this post.
Two things about it matter here. It is the first product assembled from subnet work rather than from a team we had to hire, and it earns from the day it is available rather than after a free period we would have to finance. Everything above this section is an argument until something built this way sits in front of paying users. On Monday it will.
What Ink does belongs with the release rather than in an essay about how it was funded.
Marketing the work pays for
Distribution is usually where a plan like this stops, because a company with no outside funding for marketing has to take that budget out of engineering.
We do not have to make that trade. Revenue from the products pays for distribution, so the work funds the reach instead of competing with it.
The form the spending takes is contests. We would rather show the range of the product than describe it, and the honest way to show range is to let other people find it. Contests invite the community to build something with Ink and show what they made. A few hundred people using a tool in ways we did not plan is a more accurate account of what the tool does than any advertisement we could write, and the cost is the prize.
Cortex ships on 17 August
A week later, on 17 August 2026, Cortex ships.
Cortex is our IDE: an editor for writing code with an agent built into it rather than bolted on. You hand the agent a task in plain language and it works in the open. Ask it to install a library and it breaks the job down where you can watch it: analyse the request, research and select the patterns, prepare the setup, then carry the plan out step by step, each stage marked complete as it finishes and the agent's written answer taking shape alongside.
An IDE is a different kind of asset from a single product. It is a platform where people write code, which is to say a place people spend the working day rather than a page they visit. That is the surface every later thing we build can reach, and it is the surface a coding arena feeds most directly.
It also does something for the subnet that no explanation does. A network of miners and validators stays abstract until it produces something you can open. Ink and Cortex give people something concrete to install, test, complain about and compare, and interest in the products carries back into interest in the mechanism that produced them. Shipping is how we recruit.
Paying miners to break it
Products built at this speed have defects, and we would rather find them than have a customer find them.
So finding them is a paid challenge. Miners are rewarded for reporting bugs and for submitting suggestions, scored like any other work on the subnet. The incentive points where we need it to point. A miner who finds a real fault in Ink earns for it, and the fault gets fixed before it reaches somebody who paid for the product.
This is the same mechanism again, aimed at quality instead of capability. We are not paying anybody to look for bugs. We are paying for bugs found. That distinction is the whole design.
The models we do not have
Now the disadvantage.
Cursor ships models of its own, Grok 4.5 and Composer 2.5. BASE has nothing equivalent. Not a smaller model, not a variant of an existing one trained by us. Every product named above runs on a model somebody else trained.
The cost of that is concrete rather than reputational. A product built on a model you do not control inherits its pricing, its rate limits, its deprecation schedule and its roadmap. When the model changes, your product changes, and you find out when everybody else finds out.
What the research has to settle
Miners are already researching model architecture on the subnet. The direction that currently looks most plausible is continued pretraining of an existing model with a reinforcement learning stage after it.
Plausible is the strongest word that belongs in that sentence. We have not chosen a parameter count. We have not committed to a strategy. Both decisions have to follow the research rather than lead it, and publishing either one today would mean picking an answer before doing the work that determines it. We would rather say we do not know.
There is a hard constraint underneath the question as well. Pretraining across a subnet needs heavy infrastructure and the cost is very large, large enough that the budget shapes the design rather than the other way round. Part of what the research has to establish is not only which method works, but which method we can reach.
What we can commit to is the process. The model question gets funded the way everything else here gets funded: posted as work, attempted by anyone, scored, and paid on the result. The protocol side of that is written up in the whitepaper. Meanwhile the products ship on the dates above, on models we did not train, and the loop that pays for the research runs either way.