The Base Layer Problem

Why Agentic AI Still Doesn’t Know What It’s Looking At — and What an AMD Acquisition and an Ontology Have in Common

Yesterday morning I was noodling a complex set of thoughts with Claude (my AI associate), and mid-draft, Claude had to “time out’ for a compaction. Claude was able to keep the overall context of what we’d been discussing – but completely lost the more recent and delicate nuances. Not only that, we were just starting to put this complex set of perspectives into our first few paragraphs – and Claude totally lost this comprehensive and subtle “train-of-thought”!

Yikes!

Wowza!

And this was particularly irksome because Claude had been contributing to developing this train of thought – organizing, pulling out connections – being a true “partner” in the work.

But as I was waiting for Claude to complete its compaction and re-engage, I realized this was a superb example of one of the main problems with AI-as-we-know-it-today. That problem is that our AI associates, no matter how much we extend their context windows – or provide other fixes (some of which are very, VERY good!) – are subject to the limitations of the base layer architecture.

The reason that this is so important for us is much deeper, broader, and substantially more significant than the recent attention that we’ve been giving to “rogue agentic swarms.”

AI mis-behaviors, as we like to call them, are a symptom – and not (in themselves) the underlying problem.

The underlying problem, which is very real, and which is the focus of this blogpost is that AI capabilities and behaviors are constrained by their algorithmic cores.

Simply said – if you know the core algorithms, you can predict EXACTLY what sorts of problems will emerge. And the so-called “rogue” agentic behavior is just one kind of problem.

There are others.

Right now, we’re on a threshold.

AI systems are currently built on a fairly straightforward algorithmic stack:

  • Base layer: Currently LLM (or transformer-based) models – that is, generative AI methods that generate tokens (text, image elements, or code).
  • Control layer: This tells the transformer-core what to do – it guides behavior – and right now, it is almost exclusively RLHF, or reinforcement learning with human feedback.
  • Agentic action layer: Agents (singly, or in multiples, or even swarms) carry out the directives of the control layer.

What we’re seeing now are the limitations of this stack, stretched to its extremes. This means that the innate problems are being pushed forward in a way that catches our attention.

From “rogue AI swarms” to Claude needing to compact in the middle of a carefully-developed work project – the problems are inherent to the algorithms and the architectures, and are not readily fixed with either guidelines or guardrails. You can already see it everywhere once you know to look: in a shopping agent that refunds an order twice, in a robot that can’t remember which box it already moved, in a world model that renders a gorgeous scene it has no way of actually understanding.


The Quiet $8.2 Billion Signal

Follow the money.

When the world is going crazy and ricocheting around us, like bullets fired inside a hallway – we need to separate ourselves from the angst and excitement and where everyone’s attention is focused – and follow the real signal.

Money. It’s always the money.

Here’s where it led: on September 28, AMD agreed to acquire Fei-Fei Li’s World Labs for $8.2 billion — in stock, with Li herself joining AMD as Chief Scientist. It got covered, but covered the way acquisitions get covered: a few business-press headlines, a stock-price reaction, then silence. Compare that to the volume of attention any “rogue agentic swarm” story pulls, and you’d think the swarm story was the more consequential one.

I’d argue it’s the opposite. World Labs builds spatial world models — and when a chip company with AMD’s resources spends that kind of money, and hands its Chief Scientist title to the person building them, that’s the real signal about where the actual infrastructure bottleneck sits.

I’ve been tracking this thread since at least November of last year, in a video I called “Five Spatial World Models” — and, fittingly, the conversation Claude and I were having to update that argument is exactly what got lost in yesterday’s compaction. We’re reconstructing it here, in public, as we go.

Here’s that YouTube, and I recommend watching it – because it holds important background & context for this story.

Here’s a key figure – probably the most important summary – from that entire 40-minute YouTube. (And you’re right: it deals with “the money.”)

Figure 1: Funding for four of the five different “world models” – Google’s investment in their Genie leads, with an overall yearly budget for DeepMind of at least $2Bn, and an unknown portion of that going to their spatial world models such as Genie.

Figure 1 shows funding for four of the five different “world models.” See the associated blogpost for in-depth reporting with links to news articles and research sources. (My own CORTECON model is reserved out from this funding layout, as it is not yet commercial.) Let’s start with the biggest and work our way through – because the money tells the story.


Google’s DeepMind and Genie

Google’s investment in their Genie leads, with an overall yearly budget for DeepMind of at least $2Bn, and an unknown portion of that going to their spatial world models such as Genie. Google has the significant advantage of seniority – since Google purchased DeepMind in 2014, they’ve been able to cultivate a strong infrastructure of both knowledge and product development expertise – far outstripping what the newer labs are able to put together. We’re just acknowledging the value of having long-standing teams of top minds who have really learned to work with each other over years, sometimes for longer than a decade.

This cultivation of top minds with long-term collaboration is perhaps Google’s greatest advantage on this playing field. Their greatest weakness is also very likely their greatest strength: Genie is a generative-AI “pure play,” trained with reinforcement learning. But we can look at it as – fundamentally – a strictly generative model. That means that Genie’s world-building is subject to exactly the kinds of “context-window-constraints” that led to Claude having to compact yesterday morning, with a loss of previous experiences.

[Worth noting: my original video is about a year old now, and Google’s generative lineup has moved fast since then — Nano Banana 2 (image generation) and Veo 3.1 (video generation) have both shipped in the meantime. Same underlying bet, more horsepower behind it. The argument hasn’t changed; the stakes have gotten bigger.]

So to sum up the Google/DeepMind approach – huge, in-depth expertise, cultivated for over a decade (the longest-running deep AI lab in the world) – but significantly constrained by DeepMind’s focus on generative AI coupled with reinforcement learning. A strength and a weakness, combined.


LeCun’s AMI Labs:

Leave a comment

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap