Every company is putting a chatbot in front of you.

Each one knows only the little slice of you that lives inside that company. Your airline knows your flights. Your bank knows your money. Your music app knows your songs. None of them know you.

What you want is one AI that knows you and can talk to all of them for you.

An alien intelligence

AI is the closest we have come to an alien mind arriving on earth. It already handles huge stretches of language, reasoning, memory, and pattern-matching that used to belong only to people. It does this without a body and without thinking the way we do.

The first thing we did with something this powerful was shrink it into a chat box, which puts the work on you by assuming you can name exactly what you want in one prompt, and most of the time you cannot. A stranger ordering your lunch needs to know all the minute details to make the right choice. A close friend, after years of observing you, just knows.

The ideal AI knows what you want before you are conscious of wanting it.

A chat box becomes personal only once it reads you that well. Then a vague prompt reaches the same place a detailed one would, because it fills in the rest from you.

You cannot front-load that with a hundred onboarding questions. The answers are thin. You learn a person by spending time with them, and a new AI learns your judgment the same slow way, making mistakes until you correct it.

I was inspired by Amy Deng, an AI researcher who spent about a month logging her health in obsessive detail, from her hourly energy to a hundred blood markers, then used frontier models with her doctors to reason through chronic fatigue. Her rule was that no amount of context is too much.

But who would a mind like that belong to?

You don’t own it

Right now it belongs to a handful of big companies. The labs keep their strongest models behind paid tiers, usage caps, model pickers, and policy switches.

Shut out, a lot of people have soured on the whole thing. Depending on who you ask, AI means cheating, or slop, or surveillance, or a datacenter taking over their town. I used to think they were misinformed, and I blamed my Michigan friends for worrying about a datacenter draining their water supply. I was wrong. None of this was ever built for them to own.

And it can be taken from anyone. OpenAI pulled GPT-4o out from under the people who loved it, brought it back after the backlash, then retired it from ChatGPT on February 13, 2026. Anthropic shipped Fable 5 on June 9, the government ordered access suspended three days later, and Anthropic disabled it for everyone until access returned on July 1. I finally understood what the 4o devotees had been saying: a rented mind is yours only until someone else changes the switch.

Big Brother

People type into these models what they are too scared to google and what they are not ready to say out loud. The more useful the model becomes, the more of your inner life you hand it.

Every word goes to a company whose policies, safety systems, and legal obligations decide what happens to it.

A single watching eye with lines of personal data converging into it and a red recording light at its center

A company that holds your thoughts is a place a court order can land, a place policies can change, and a place where someone else can decide what the service will answer.

Why it has to be yours

Personal AI must be owned by the person.

When you own it, there is no provider-side account full of your inner life by default. The record stays with you. No company can rate limit or switch off the local copy you run yourself. Cloud models can still help with hard thinking, but they stop being the place your whole life has to live.

It reads your email or watches your security camera only because you put it on your own machine and gave it that permission.

Computers did this first

Ownership has reached the person before. Computing started locked inside institutions. A mainframe filled a room, belonged to a university or a government, and you booked time on someone else’s machine. Then it moved out to the desk and into your pocket, and you went from borrowing time to carrying the machine. In 1977 Ken Olsen, who ran Digital Equipment, one of the largest computer companies in the world, said there was no reason for anyone to have a computer in their home.

AI is in that same mainframe era right now. Training and the hardest reasoning still live in giant buildings. But memory, context, sensors, and ordinary work can move to the person. That is the personal computer moment for AI: the center of gravity moving from the institution to you.

The Google objection

Search trained us to expect consumer intelligence from the cloud. That made sense for search. Google’s job was to crawl and index the public web, a shared corpus too large and too fresh to live on one person’s machine.

Personal AI is different. The valuable index is your messages, files, calendar, health signals, habits, corrections, voice, taste, and work style. A cloud can search the world for you. The owned core is what remembers you.

It should not be a luxury

Cost is what keeps personal AI rationed.

Most AI runs by renting one giant brain in the cloud for every task. Satya Nadella says there should be as many models in the world as there are firms, because a firm is a learning system. I would go further. As many models as there are people.

Open models are catching up fast. The open frontier, like GLM-5.2, is now serious enough to make ownership plausible even as it trails the closed labs, and smaller models like Qwen3 and gpt-oss-20b already have local paths. Stanford’s Hazy Research measured local models answering 88.7 percent of single-turn chat and reasoning queries correctly, with intelligence per watt growing 5.3x in two years. Ownership only needs that common path.

The power users show the demand curve. OpenClaw, built by Peter Steinberger, runs beside you and touches the real tools on your machine, driven from the chat apps you already use. Once that kind of agent runs constantly through frontier APIs, usage becomes a meter on ordinary work.

A serious personal agent doing 1,000 small GPT-5.5-class text turns a day, each with 2,000 input tokens and 200 output tokens, costs about $480 a month before audio, video, search, tools, and hosted runtimes. Cheaper models lower that bill without changing its shape.

Flat-rate plans survive by rationing. A weekly limit is the product admitting that unlimited frontier inference is still too expensive to hand out. A middleman renting you a rationed model, taking a cut of every token, cannot be how everyone ends up with one.

Cloud rent is recurring. A rented H100-class GPU is roughly three to four dollars an hour on current public cloud pages, or about two to three thousand dollars a month before storage and overhead.

$3hour×24 hours×30 days$2,160 a month\frac{\$3}{\text{hour}} \times 24\ \text{hours} \times 30\ \text{days} \approx \$2{,}160 \text{ a month}

Owning the same card gets cheaper the longer you keep it. GPUs have appreciated through the compute shortage, and the twenty-five thousand dollars of hardware holds its value: you can sell it or rent it out when you are done. The running cost is mostly power, about ninety-six dollars a month at 700 watts.

700 watts1000×24×30×$0.19kWh$96 a month\frac{700\ \text{watts}}{1000} \times 24 \times 30 \times \frac{\$0.19}{\text{kWh}} \approx \$96 \text{ a month}

Rent is over two thousand a month gone for good; the box costs under a hundred to run, and the cash comes back when you sell. The twenty-five thousand up front is what keeps it a luxury. Most work does not need that card anyway: once the local path is good enough, the common work runs on a desktop costing a few thousand dollars, and the hard tail still rents the cloud.

That is the device shape already emerging: a box under the desk or a high-memory desktop handles memory, retrieval, routine summaries, local tools, and first-pass inference, then calls the big cloud model when the job outruns it. Companies are discovering the same thing. Alex Karp’s complaint about token bills and lost business alpha is the enterprise version of the personal problem: the valuable thing is the context, and the owner of the context should own the core.

The agents people rent today are also the path to that box. Everything your agent does for you, every task and every correction, becomes personal context: memory, retrieval data, preferences, evals, and eventually adapters or distillation data for smaller local models.

People want use that feels unmetered. That arrives when common-path inference gets cheap enough to stop rationing. A GPT-3-class API call cost $60 per million tokens in 2021 and six cents by 2024. Epoch AI puts the decline at 9x to 900x a year depending on the task, with a warning that the fastest drops may not persist. A model of that class now runs on the laptop you already own.

Price of GPT-3 class intelligence per million tokens on a log scale, falling from $60 in 2021 to 6 cents in 2024

Mark Pincus, who built Zynga, told Garry that consumer is not investable right now, and that this is exactly the reason to build it. Pincus and Garry put the moment ordinary people get this at around 2029.

A YouTube comment by @breakingbedrock7655: We're in a pretty expensive and enterprise-only phase right now, so for those of us building, innovation comes from building as if compute is free, pre-loading your product for a world where AI-native UX is the standard.

When personal AI feels unmetered, everyone can have one.

Where the money goes

Renting out intelligence looks like the business of AI today. The margin exists because the best weights are closed, the hardware is scarce, and most people cannot run the frontier themselves. It is a toll on a road you are not allowed to build. Open models and the box under the desk take part of that toll away. The repeated personal work moves off the meter; the hardest thinking still pays it.

Time-sharing bureaus in the mainframe era rented computer hours the way the labs rent tokens now. Personal computers moved the center of value toward the machine on the desk and the software that ran on it. The economic motion is the same now: repeated work wants to stop paying rent.

The biggest profit pool in AI hardware already belongs to NVIDIA, the company selling the machines: $120 billion of FY2026 GAAP net income on $216 billion of revenue. NVIDIA now sells each generation as a lower cost per token. For the machine seller, inference deflation is the roadmap. For the token seller, it is pressure on any margin that depends on metering ordinary use.

A datacenter is efficient because everyone shares one giant brain. Thousands of questions run through the same weights on the same chip at once, and that sharing is where the margin comes from. Personal AI erodes it. The adapters can still batch on a shared base. The always-on work of watching your sensors, indexing your files, and holding your memory hot serves exactly one person no matter where it runs. Work that cannot be pooled earns no pooling margin, and a machine of your own does that work without paying anyone’s margin at all. The more personal the workload gets, the worse it fits the cloud and the better it fits yours.

The cloud labs will have an incentive to ship devices too, but their incentive is different. A cloud-first device protects the subscription by sending the hard and valuable work back to the service. The company that builds the personal computer of AI starts from the opposite premise: local first, cloud only when the local mind runs out.

Interface

The model race will keep going, and the personal AI product cannot wait for one model to swallow every task. It needs an interface: the thing you can talk to, the sensors it can borrow, and the owned core that remembers what they all saw.

Every generation works a machine in a way the one before finds strange. The people before us grew up turning dials and pressing buttons that moved. We grew up on glass, tapping a flat screen that fakes the click of a button with a haptic. Each of these turns an intent into a touch, and the touch into an action.

Clicky puts an AI buddy beside your cursor that sees your screen and talks you through it. Google DeepMind asked “what if behind the pointer, there was an AI model” and is building that into Googlebook. The first instinct is to pour AI into the interfaces we already know. It helps, and it keeps the AI inside a cursor and a screen.

The next one is the conversation. You say what you want and it happens, and the machine fades into the room. Intent becomes action, the touch gone. The chat boxes we started with were its first crude version.

The conversation has to reach you wherever you are, through the things you wear and carry. Today the watch, glasses, earbuds, phone, and home computer each hold a different shard of context, the same way every website runs its own chatbot. Your AI should be the owned thread through all of them: glasses as eyes, earbuds as voice, phone as the first hub, and the machine at home as durable memory. The company that wins builds the core that ties every sensor into one mind, without owning a single sensor or letting the mind belong to the companies that do.

In the movie Her, Theodore keeps Samantha on a small folding device he carries everywhere and talks to her through a single earpiece, on the train and by his bed at night.

Theodore's handheld device from the film Her resting on a nightstand, its screen reading Call From Samantha, beside a pair of glasses and an earpiece

We do not have that shape yet. A phone is made for opening apps and tapping glass. A companion you speak to all day needs a smaller, calmer set of surfaces: a voice channel, a glanceable object, and sensors that wake only when they have something worth sending.

It reads the signals your wearables already know, from heart rate and sleep to motion and temperature, and folds them into a single local picture of your body. When you wear a real glucose sensor, that can join the picture too. Amy built hers by hand, a month of obsessive logging. One that lives with you builds it continuously, without asking you to become the data clerk.

People hear all this and jump to brain interfaces. They are the cleanest path from intent to action on paper, and they already matter for people who need clinical access to a computer. But a surgically placed implant is not the mass interface for a companion. Virtual reality promised to move computing off the old screen, and the everyday computer stayed on the desk and in the pocket. The next interface has to ask less of people.

The same product has to join hardware, software, and experience around one owned personal core. Plenty of companies have made integrated devices. Nobody has yet made the everyday personal AI object where the sensors, the conversation, and the memory all serve the same person-owned mind.

A memory object

If you mostly talk to it, it does not need a screen. The voice apps we have already glow softly when they listen, and the object on your desk that glows is a lamp. A personal AI should be a thing like that, something you want to keep in front of you.

Peter Kuhar is building one now, an ambient display he calls SidePulse that slots into a MacBook’s card reader and lights up while the agent works.

One I find interesting is Keunwook Kim’s Memory Object, a glowing thing you hold that stores memory in physical material.

Keunwook's Memory Object, a glowing sphere held in cupped hands

Yours holds the record that matters: the choices you made, the corrections you gave it, the people you love, the arguments that changed you, and the way you tend to answer. The longer it lives with you, the less it feels like a device with settings and the more it feels like a relationship with a history. You could no more swap it for a fresh one than swap an old friend for someone who has read your biography.

A model that spends a life beside someone becomes a record of that relationship. After the person dies, the record can still answer from what it learned, a living biography you can question, argue with, and pass down. You could ask what your grandmother cared about at twenty, or talk with the traces of a thinker whose private judgments would otherwise have disappeared.

Owned memory is what makes that possible. A rented assistant can be changed, capped, or cut off when the company changes the model or the subscription ends. An owned one can become an heirloom, a small museum of a life that the family and the world can still visit.

Decentralized

Even an owned mind has a ceiling. Your own data can teach it your taste, voice, habits, and history, but it cannot give it every rare case everyone else has lived through. Ramesh Raskar’s work on split learning points at the route around that: many devices can help train a shared model while the raw data stays where it was created.

The lesson can move without the diary moving. In split learning, the private side keeps the raw example and sends only the intermediate signal needed for training; the shared side learns from the pattern without collecting the original life. It still needs trust, incentives, poisoning defenses, and governance before it is a product. But it already gives the owned AI a way to learn with others without turning everyone’s private life into one company’s training pile.

The goal is a network where the models improve together and the person still owns the source of the memory, with no company in the middle holding an off switch.

A new internet

Knowing you is only half of it. The other half, acting for you, already exists most clearly in code. Cursor moved AI from advice in a chat box into the editor, where it can read the project and change the files. Codex goes further into the background: you hand it a coding task, it works in its own cloud environment, runs commands and tests, and brings back a diff or pull request. Coding agents started as copilots, and now people let them run while they do something else. I keep a few of my own, one that answers support tickets and one that fixes my app. The unit of work you can hand off keeps climbing.

The rest of your life you still work by hand. You open the apps and type out what you want. Your AI takes that over too. You are hungry and it already knows the food you would have ordered. You want something to listen to and it has the song before you ask. It does the searching and the clicking for you, on every site, all day.

When everyone has their own AI, the agents start talking to each other. A friend in Belgium used to send me the problems he found in my code. Now he emails my agent instead, and his agent and mine settle it between themselves.

Businesses will have agents too. Your agent should not have to load a page, hunt for a button, and pretend to be you moving a mouse. It should ask the business’s agent for what you want, prove what it is allowed to do, pay when payment is needed, and bring the result back.

The human web stays, and a machine-facing web grows beside it. Pages, feeds, and forms are for people; agents need tools, identity, trust, payment, and discovery. That layer exists only in pieces.

MCP, created by Anthropic, gives agents a common way to plug into tools, data sources, and workflows. Google’s A2A gives independent agents a common language for talking to each other, with Agent Cards that advertise what an agent can do. AP2 adds a payment authorization layer for agent commerce, and Coinbase’s x402 revives HTTP 402 for direct stablecoin payments over ordinary web requests.

Discovery is the larger unfinished part. A2A can tell you what a known agent offers, but the open internet also needs a way to find, verify, route to, revoke, and trust agents at global scale. MIT’s Project NANDA is a registry beyond DNS, built around verifiable AgentFacts, for an internet its authors expect to fill with billions of agents.

The internet was built from open standards once before. The protocols that moved data between machines and loaded a page in your browser did not belong to one company, and that is what let the web become everyone’s. Tim Berners-Lee saw the next shape in 2001: software agents roaming the web, reading machine-meaningful pages, and carrying out tasks for people.

If a few companies own the agent layer, they become the gatekeepers of the next internet. If the agents belong to the people they serve, the new internet belongs to everyone.