I'm looking at what came out this week, and underneath it all I see one thing: AI agents are no longer a demo next to the product. They are the work itself - measured in hours, in percent of what got done, in whole teams.
- The thesis of the 30 June issue: agents are no longer a demo next to the product, but the work itself, measured in hours and percent.
- The thread: HP is rolling out OpenAI Frontier across its whole business, Gemini clicks on its own, and Anthropic measures usage in hours.
- The takeaway: the system already runs on its own - as labor that counts.
This week, under the headlines, there's one common sign: the agent stopped being a feature in the product and became the labor that produces it. And that's no longer a promise - it's measured in hours and in percent.
I'm not claiming this in broad strokes. I'm looking at the numbers the companies themselves put out.
The numbers, not the promises
OpenAI released an internal economic report: their coding agent, Codex, now writes 99.8% of the code AI creates inside the company. For the average employee, over 85% of what gets written with AI help goes through Codex. Google built 'computer use' directly into Gemini 3.5 Flash - an agent that sees the screen and acts on its own in a browser and on desktop. HP announced it's rolling out OpenAI 'Frontier' (OpenAI's agentic platform for business) across its whole business, after a pilot where one engineer pushed through 122 code merge requests in a few weeks. And Anthropic measured something else: 93% of conversations with Claude now produce a concrete result - a document, code, an explanation.
One clarification. These numbers come from the companies themselves, and some of the estimates are internal or modeled, not independently measured. I'm using them as a direction, not as proven statistics for everyone.
And that's why the other word is control again
The more real an agent's work gets, the more serious the question becomes of who runs it and how far. The same week NVIDIA and Palantir announced AI for government agencies that runs air-gapped - isolated from the internet, with weights under the agency's own control. ElevenLabs put an invisible watermark in generated audio, so you can tell what's machine-made. And CISA added new actively exploited holes, including in networking gear that a lot of people keep at home.
My read: the agent becomes useful and dangerous at the same moment. That's why data provenance, isolation and logs stop being a formality. They become the condition for running an agent at all.
On the ground
If you're building with AI, stop measuring it by 'can it answer.' Measure it by how many hours of work it takes off you and how you check the result. It's a different discipline - delegation, review, access, cost under real load.
For a business it's just like any other new hire: a clear task, clear boundaries, someone watching what it does. Except this worker is software - it copies endlessly and makes mistakes fast. If you ask me, the winners won't be the ones with the smartest agent, but the ones with the cleanest process around it.