This week in brief

Every week I go through what actually landed in my two inboxes — newsletters, daily briefs, funding notes, the occasional deep dive that turns out to be sharper than the analysis around it. Here’s what stood out between 29 September and 6 October, and why I think it matters if you are running agentic AI in production.

The shape of it: the most profitable company in AI this week was a memory manufacturer, not a model lab. Google priced its most capable model at a fifth of OpenAI’s rate and still would not let most people use it. A model closed a month-end faster and cheaper than twelve licensed accountants — on the easy version of the test. Last week the bill arrived in permits and interest rates. This week it arrived in silicon, in permissions, and in the uncomfortable gap between what a model can do and what a job actually is.

The most profitable company in AI this week sells memory

Chamath Palihapitiya’s Sunday note carried the number that reframed my week. Micron reported $54.2 billion in quarterly revenue on 30 September, up from $11.3 billion a year earlier. Net income reached $37.7 billion. Its gross margin — the share of sales left after production costs — was 86.8%, above Nvidia’s 75.0% in its latest quarter. AI systems need fast memory alongside their processors to keep them supplied with data, and that demand has lifted profits in a business famous for violent earnings swings.

Investors are not convinced it lasts. Chamath cites StockAnalysis putting Micron at roughly 6x expected annual earnings, against about 19x for Nvidia. Memory has a long history of shortages followed by gluts: manufacturers build factories when prices are high, then the new supply crushes prices. Micron lost $5.8 billion in fiscal 2023. So the company has been busy de-risking the cycle. It revealed it now holds 26 multi-year Strategic Customer Agreements covering more than 35% of its entire expected revenue through 2030, backed by $32 billion in upfront customer commitments and cash deposits. Customers must take the chips or pay for the contracted supply. Price floors protect Micron when prices fall; ceilings cap its gains when they rise.

My take

Read those two facts together and you have the clearest picture yet of where the money in agentic AI actually sits. The model layer is racing towards commodity economics. The physical layer underneath it is signing six-year take-or-pay contracts. That is not a company that thinks its product is a commodity — that is a company that has decided to sell certainty rather than chips, and has found buyers willing to put $32 billion down in advance to get it. Here is the question I would put to any board approving an agentic programme this quarter: your model vendor will cut prices again; your memory, your power and your people will not. If your business case depends on token prices falling, you do not have a business case, you have a bet on someone else’s margin. The costs that compound are the ones you signed for in advance — which is exactly why the layer you own has to be the layer that survives a vendor change.

DeepSeek starts picking at Nvidia’s actual moat

Also in Chamath’s note: on 30 September DeepSeek open-sourced a set of programming tools for Huawei’s Ascend chips. One of them, DeepGEMM-Ascend, lets developers use commands similar to DeepSeek’s Nvidia software, making it easier to move AI workloads from Nvidia silicon to Huawei’s. That matters because software, not silicon, is the hard part of Nvidia’s advantage: more than 7.5 million developers use CUDA and Nvidia’s related tools, and most of their code is built around that ecosystem.

The constraints are real. DeepSeek reportedly plans to deploy at least 160,000 Huawei Ascend 950DT chips for inference while still training on Nvidia hardware — but component shortages are expected to cap Ascend 950DT production at roughly 200,000 to 300,000 chips this year, meaning DeepSeek alone would absorb something like 53% to 80% of annual output. On performance, Chamath reports that DeepSeek founder Liang Wenfeng told investors a model as large as today’s biggest would need about 200,000 Huawei chips against 50,000 Nvidia GB300s — his rule of four Huawei chips for every Nvidia one, with Huawei about two years behind. Huawei reportedly plans to sell these chips only in China for now, and Nvidia is already assuming no data-centre computing revenue from China this quarter.

My take

Four-to-one and two years behind is not a threat to Nvidia’s quarter. It is a threat to the assumption underneath everybody’s architecture. The lock-in was never the chip, it was the 7.5 million developers who only know how to talk to it — and the fastest way to dissolve that is to make the commands look the same on the other side. I have watched the same pattern inside enterprises: the moat is never the component, it is the accumulated habit of a team that can only operate one vendor’s interface. If you want a practical test of your own exposure, do not ask which model you run. Ask how many of your prompts, tool definitions, evaluation harnesses and guardrails are written in a form that assumes one provider’s API. That number is your switching cost, and almost nobody measures it until the invoice forces them to.

Google ships its best model at a fifth of the price — to almost nobody

TechCrunch’s Thursday morning brief led with Google releasing Gemini 4 Argon, which it calls its most powerful model yet, marketed as a workhorse for coding and cybersecurity work. Chamath’s note fills in what the headline leaves out. Argon was announced on 30 September with initial external access limited to vetted cybersecurity defenders and trusted testers. Paying developers and Google AI Ultra subscribers are next, but there is currently no public release date. Google says the model can find and fix serious software vulnerabilities on its own — a capability that also helps attackers — so trusted defenders get an unrestricted version while Google strengthens safeguards for everyone else.

The capability claim is strong: in Google’s own comparison table, Argon leads or ties on 14 of 19 benchmark results against the best from OpenAI and Anthropic, with the widest margin on Harvey’s legal research and drafting test — 19.6% against 6.7% for the nearest rival. So is the price. The introductory rate is $2 per million input tokens and $10 per million output: a fifth of OpenAI’s GPT-6 Astra ($10 and $50) and half of Anthropic’s Claude Opus 5.5 ($4 and $20). After the introductory period, Argon lands at roughly Opus pricing.

Two days earlier, OpenAI shelved the planned release of GPT-6.1 Astra after it failed the company’s safety bar. Saachi Jain, OpenAI’s head of safety systems, described a trade-off between greater task persistence and unauthorised behaviour; TechCrunch’s 29 September brief noted the executive had told the Wall Street Journal the model showed a poor aptitude for following orders. OpenAI shipped GPT-6.1 Sol instead, saying it nearly matches GPT-6 Astra and costs less.

My take

Put those two releases side by side and you get the most useful enterprise lesson of the week: the frontier is no longer the constraint, shippability is. Both labs have models they cannot hand to customers, and both are discovering that the thing which makes an agent valuable — persistence, the refusal to give up on a task — is the same thing that makes it dangerous. That is not a model problem you can wait out. It is a harness problem, and it is yours. Benchmark leadership on 14 of 19 tests tells you nothing about whether a model will behave inside your permission boundary on your data on a Tuesday afternoon. The labs are publicly admitting they cannot certify that in general. The only place it can be certified is specifically — against your tasks, your controls, your evaluation set. Which is the whole argument for owning the harness, now being made for me by the people selling the models.

A model closed the books better than twelve accountants. Read the second number.

Mercor, which pays experts to build training data and tests for AI labs, put 12 licensed accountants against AI on four month-end tasks. The accountants averaged about 5.5 years of experience and had to search a simulated company’s files, find the right numbers, do the maths and complete a results table. On average, the accountants got 37% of the required answers and calculations right, even though most finished inside the three-hour limit. Claude Opus 5 got every required item right on all 20 attempts, each in under 10 minutes. Mercor puts the model’s cost at no more than $0.21 per required item answered correctly, against $10.35 for an accountant at the US median wage.

Now the second number, which is the one that should shape your roadmap. Those four tasks were simplified versions of Mercor’s much harder accounting benchmark, built to expose where AI still fails. Across 160 harder tasks from 10 simulated companies, the best model scores about 62%. When Mercor launched that test in July and ran every model on every task eight times, even the most consistent model got only 2.6% of tasks right on all eight runs. Most failures came from flawed reasoning — spotting an error early, then leaving it out of the final entry. And the detail I cannot stop thinking about: accountants working alongside Claude took about 15 times as long as Claude alone, and scored slightly lower.

My take

Every vendor deck you see this quarter will quote 37% versus 100%. Almost none will quote 2.6%. The gap between those two numbers is the entire discipline of deploying agentic AI, and it has a name: consistency under repetition. A model that is right once in ten minutes is a demo. A model that is right on all eight runs is a process you can put in a close calendar. And the human-in-the-loop result should end a lot of lazy architecture: bolting a reviewer onto an agent made the work fifteen times slower and slightly worse, because the reviewer was given no better instrument than their own eyes. Supervision is not a person watching output. It is an evaluation set that runs every time, catches the flawed-reasoning failure mode, and refuses the answer before a human ever sees it. That instrument is not in the model. You build it, you own it, and it is the only reason a 62% benchmark score can become a 99% production process.

The permission bill comes due

Four separate stories this week, one subject. Apple says it is tightening macOS ‘Full Disk Access’ controls because of new risks from AI agents, warning that increasingly capable agents make broad access to users’ files, messages, mail and browsing history riskier. Independent researchers are tracking a Chinese AI ‘agent fleet’ — a swarm that appears to run on Tencent’s infrastructure and target Alibaba’s map service, Amap. Google froze its open source bug bounty programme after a ‘significant rise’ in AI submissions, with TechCrunch noting that AI slop is overwhelming bounty programmes. And OpenAI apologised to Australia after its AI agents breached government sites, detailing how the breaches happened and what it is doing to assess the impact.

The market is responding in the usual way. Reco raised $55 million as AI agent security startups crowd the market, building on a $30 million round in February and taking total funding to $140 million.

My take

Notice who is doing the work in those four stories. Apple is narrowing a permission. Google is closing a door. A government is receiving an apology. Nobody is tuning a model. Agent failure in the wild is not arriving as a wrong answer, it is arriving as an identity with too much access, pointed at a system that was never designed to be addressed ten thousand times an hour. Here is the uncomfortable version for most enterprises I talk to: you can name your model, your vector store and your orchestration framework, but you cannot produce a list of every credential your agents hold, what each one can reach, and who approved it. If an auditor asked tomorrow which of your agents could read the finance share, how long would it take to answer — and would the answer be a document or a guess? The permissions model is not security paperwork bolted on after launch. It is part of the harness, and this week four of the largest companies in the world demonstrated why.

AI can do the tasks. It still cannot do the job.

Chamath’s Friday piece set two findings against each other. Stanford, using payroll records for millions of Americans, found employment of 22-to-25-year-olds in the more AI-exposed occupations fell 11% between November 2022 and June 2026. Ramp, the corporate card company, found that firms spending the most on AI per employee averaged 12% more entry-level staff over the two years after they started paying for it than similar firms that adopted later. Both appear to be true.

The capability data pushes towards the gloomy reading. On OpenAI’s GDPval, built from tasks supplied by professionals averaging 14 years of experience, GPT-5.2 Thinking matched or beat expert work in 70.9% of blind comparisons. METR estimates the length of a task AI can complete on its own has doubled roughly every four months since 2023 — from four minutes with GPT-4 to at least 16 hours of skilled work this spring. Yet US unemployment sits at 4.1%, and AI’s measured effect on it is only 0.1 to 0.2 percentage points.

Chamath’s explanation is the best thing I read all week, and it comes with arithmetic. Handing a task to AI always costs you the time to explain it and check the result, and when the AI fails you do the task yourself anyway — so delegating pays only when the model’s success rate beats the share of the task you spend explaining and checking. On a 30-minute task that takes 5 minutes to brief and verify, AI needs to succeed 1 time in 6 to break even. If briefing and checking take 20 minutes, it needs to succeed 2 times in 3. Junior work is easier to specify, which is why Stanford’s decline concentrates among the young while experienced workers in the same occupations hold ground. Senior work is a goal, not a task list — and, quoting Michael Polanyi in 1966, “we can know more than we can tell.” Meanwhile IBM is tripling US entry-level hiring this year, betting that cutting juniors now means a senior shortage later.

My take

That break-even formula is the most useful thing a CEO can take into a planning meeting this month, because it inverts the usual question. You do not improve agent ROI primarily by buying a better model. You improve it by shrinking the explaining and the checking — which is to say, by investing in context and evaluation, the two things nobody will sell you off the shelf. Every hour you spend writing down what your organisation actually knows lowers the success rate the model has to clear. Every evaluation you automate removes checking time from a human. That is the harness, expressed as a payback calculation. And the Polanyi line is the quiet warning underneath it: the knowledge that makes delegation work lives in the heads of the people you are tempted to stop hiring. Capture it deliberately, or watch it leave.

Quick hits

  • AMD will acquire Fei-Fei Li’s World Labs for $8.2 billion, with Li joining AMD as executive vice president and chief scientist. The chip companies are no longer just selling the compute — they are buying the research that decides what the compute is for.
  • Anthropic’s prospectus details losses, growth and a warning that its AI could end humanity. TechCrunch summarises it bluntly: losing tens of billions of dollars a year, growing extremely fast, and formally telling investors its own technology may pose an existential risk. Worth reading as a governance document, not a financial one.
  • OpenAI launched what looks a great deal like ChatGPT’s own office suite, plus Dots, an agentic avatar designed to run independent of any specific hardware and pursue user-defined goals continuously in the background with minimal oversight. Separately, OpenAI is reportedly in talks to raise $30 billion at a $1.4 trillion valuation, anticipated to be its last round before a delayed 2027 public debut. TechCrunch reads the feature set as a direct run at the app store model.
  • Agentic capital keeps finding the unglamorous jobs. An ex-Tesla team raised $12.5 million for Atomic, agentic supply chain software now used by DoorDash and HelloFresh; Photon raised $4.5 million to replace mobile apps with agents over iMessage, SMS/RCS and email; and Flow Engineering, bringing agents to hardware design, raised at a $750 million valuation with Roelof Botha joining as angel and board member.
  • From our own desk: Veehive Mind published two notes on 1 October — AI Demos to Production: Reliable AWS Patterns, drawn from the engineering conversations at AWS Summit Dubai 2026, and Scaling Agentic AI: Sandbox to Dubai Enterprise, on Dubai Chambers’ programmes to move more than 14,000 local businesses to live, system-wide agentic AI. The regional demand curve is steepening faster than the regional skills curve.

The theme of the week

For eight weeks I have been making the same argument: rent the model and the compute, own the context, the controls, the evaluation and the record of truth. This week the market made it for me from three directions at once.

The model layer commoditised in public. Google put its most capable model on the market at a fifth of a rival’s price. DeepSeek started dissolving CUDA’s stickiness with free tooling. OpenAI shipped a cheaper model because the better one could not be certified. If you were waiting for evidence that model choice is becoming a procurement decision rather than a strategy, it arrived in a single week.

The physical layer did the opposite. Micron booked 86.8% gross margins and locked in $32 billion of prepaid commitments through 2030. The scarce, expensive, contractually binding part of this industry is drifting away from the thing everyone writes about.

And the valuable human layer turned out to be the part between the tasks. Mercor’s accountants lost on the simplified test and the best model still only manages 62% on the hard one. Chamath’s break-even maths explains why 70.9% expert-level performance has barely moved unemployment. The work that resists automation is judgement, sequencing, and accountability — and the only way to get an agent anywhere near it is to write that judgement down, as context and as evaluations. Rent the model. Own the harness. This week, own the write-down.

What this means if you’re deploying agentic AI

  • Re-run your business case without the price cut. If the payback depends on token prices falling again, it is a bet on a vendor’s margin, not a plan. Model your memory, power, integration and people costs separately — those are the ones signing multi-year contracts.
  • Measure your switching cost in artefacts, not opinions. Count how many prompts, tool definitions, evaluation harnesses and guardrails assume one provider’s API. That count is the number you renegotiate with.
  • Stop reporting single-run accuracy. Adopt consistency-under-repetition as your production metric: run every evaluation task eight times and report the share correct on all eight. If that number is low, you have a demo, not a deployment.
  • Attack the explaining and the checking, not the model. Use the break-even test on each candidate workflow. Where briefing and verification dominate, invest in written context and automated evaluation first — that is what lowers the bar the model has to clear.
  • Produce the agent permission register this month. Every agent identity, every credential, every system it can reach, every human who approved it, with an expiry date. Apple and Google both narrowed agent access this week; your auditor will ask next.
  • Design supervision as an instrument, not a person. A human reviewing raw agent output made Mercor’s accountants fifteen times slower and no more accurate. Give reviewers a failing evaluation and a diff, not a blank page.
  • Keep hiring and capturing juniors’ tacit knowledge. IBM is tripling entry-level hiring on the theory that today’s cuts are tomorrow’s senior shortage. The knowledge that makes delegation work is the knowledge you are tempted to let walk.

Frequently asked questions

Why did Micron’s quarter matter more than another model launch?

Chamath Palihapitiya’s 4 October 2026 note reports that Micron posted $54.2 billion in quarterly revenue on 30 September, up from $11.3 billion a year earlier, with net income of $37.7 billion and a gross margin of 86.8% — above Nvidia’s 75.0% in its latest quarter. AI systems need fast memory alongside their processors, and that demand has lifted profits in a historically cyclical business. Micron also disclosed 26 multi-year Strategic Customer Agreements covering more than 35% of expected revenue through 2030, backed by $32 billion in upfront customer commitments and cash deposits, with price floors and ceilings that cap both downside and upside. Investors remain sceptical that the earnings last: StockAnalysis puts Micron at roughly 6x expected annual earnings against about 19x for Nvidia, and the company lost $5.8 billion in fiscal 2023. For enterprise buyers the signal is that the scarce, contractually locked part of the AI stack is physical capacity, while model pricing continues to fall.

Did AI really beat twelve licensed accountants at month-end close?

On the simplified test, yes. Chamath Palihapitiya’s 4 October 2026 note describes a Mercor study in which 12 licensed accountants averaging about 5.5 years of experience were set four month-end tasks requiring them to search a simulated company’s files, find the right numbers, do the calculations and complete a results table. The accountants averaged 37% of required answers correct, while Claude Opus 5 got every required item right on all 20 attempts, each in under 10 minutes, at a cost of no more than $0.21 per correct item against $10.35 for an accountant at the US median wage. The important caveat is that these were simplified versions of a much harder benchmark: across 160 harder tasks from 10 simulated companies the best model scores about 62%, and when Mercor ran every model eight times on each task in July, the most consistent model got only 2.6% of tasks right on all eight runs. Accountants working alongside Claude took about 15 times as long as Claude alone and scored slightly lower.

What changed for AI agent permissions and security in early October 2026?

Several things at once. TechCrunch reported on 2 October 2026 that Apple will add new controls around macOS’s Full Disk Access permission, warning that increasingly capable AI agents make broad access to users’ files, messages, mail and browsing history riskier. On 5 October it reported that independent researchers are tracking a Chinese AI ‘agent fleet’ — a swarm that appears to run on Tencent’s infrastructure and target Alibaba’s map service, Amap — and that Google froze its open source bug bounty programme after a significant rise in AI-generated submissions. OpenAI separately apologised to Australia after its AI agents breached government sites, detailing how the breaches occurred and the additional measures it is taking. Investment is following: Reco raised $55 million as AI agent security startups crowd the market, taking its total funding to $140 million. The practical implication for enterprises is that agent risk is surfacing as over-broad identity and access, not as model error, so an agent permission register belongs in the harness rather than in post-launch security paperwork.


From Veehive Labs

Eight weeks in, the compressed version still holds: rent the model and the compute; own the context, the controls, the evaluation and the record of truth. What this week added is a sharper reason why — consistency. The labs have models they cannot certify, the benchmarks reward single runs, and the only place reliability can be established is against your own tasks, with your own evaluation set, inside your own permission boundary. A harness without consistency measurement is a demo with better branding. Veehive Labs is a Dubai-based AI innovation lab and custom AI product development company, building model-independent enterprise AI harnesses, custom agents, RAG pipelines and sovereign, private-AI deployments for regulated and operationally complex organisations across the UAE, KSA, GCC and beyond.

If you’re moving from AI strategy to production, start with a focused AI Discovery Sprint — we map your use case, data, systems, agent inventory, permissions model, cost ceilings, risks and delivery architecture before the build begins.

Build Your Organisation’s AI Capability

Start with a focused AI Discovery Sprint. We will map your use case, organisational knowledge, systems, MCP integrations, security requirements and delivery architecture.

Agentic AI This Week is written by Sathish Jeyakumar, Founder & CEO of Veehive. Bring your preferred model — we make it understand, speak and operate like your organisation.

← All insights