Why the Same Concept Looks Different in Every Paper — and What to Do about It
Have you ever picked up a paper on active inference or variational inference and felt genuinely lost? This was not because the math was too hard, but because the symbols seemed to mean something completely different from what you expected. If so, you’re not confused. You’re not a slow learner. You have simply encountered one of the most underacknowledged problems in the field.
The same mathematical concepts appear under completely different notation depending on which paper you’re reading. A variable that means “hidden state” in one paper means “observed variable” in another. The same letter — x — carries opposite meanings depending on whether you’re reading Namjoshi, Beal, or Blei et al.
This is what I call the Notational Nightmare. And it has cost serious researchers (like me!) years of unnecessary confusion.
YouTube Link
This blogpost accompanies a full YouTube video walking through each gotcha in detail, with the Hero’s Journey framing that puts you — the reader — as the hero navigating this particular labyrinth.
[YouTube link will be inserted here.]
Four Specific Gotchas
After ten years of working through this literature, I’ve identified four specific collision points that trip up almost everyone:
- H means enthalpy in thermodynamics, entropy in information theory — and active inference sits at the intersection of both.
- Helmholtz and Gibbs free energy appear in the literature but are not the same as variational free energy — the term that actually matters for generative AI.
- The direction of the KL divergence — whether it’s q||p or p||q — changes the entire character of the optimization problem.
- The x/y notation convention — in algebra, x is observed and y depends on it. In Namjoshi and Beal, it’s exactly reversed. (There are more nuances in the variables notation – see the full Rosetta Stone below!)
The Rosetta Stone
To address this, I’ve put together a reference table mapping notation across five key sources: Namjoshi (2026), Beal (2003), Friston pre-2016, Friston 2017+, and Blei et al. (2016).

The bottom row tells the essential story: “x” means hidden in Namjoshi and Beal, not used in Friston 2017+, and observed in Blei et al. Same letter. Completely different meaning.
Download the Free PDF
I’ve made the Rosetta Stone available as a free downloadable PDF on the Themesis Resources page. Keep it open while you read. It will save you hours.
Want the Full Foundation?
The T3 course — Top Ten Terms in Statistical Mechanics and Generative AI — provides the structured foundation that makes all of this make sense. The notation is only confusing when the underlying concepts aren’t yet solid. T3 builds those concepts.

To learn more about the Top Ten Terms, go to: https://themesis.thinkific.com.
At the the Themesis Thinkific site, sign up as a Learner – it’s free (in exchange for your email address). You’ll be able to:
- Review the course Table of Contents for each course where you’ve signed up as a Learner,
- Access the course material for any Lessons that have been marked as “Free” – this could include PDFs, video tutorials, and other resources, and
- Join the waitlist for the next structured cohort — where you work through the material with a small group, with live sync sessions and direct access to Dr. Maren at key moments in the curriculum.
References and Resources
- Namjoshi, S. (2026). Fundamentals of Active Inference. MIT Press. (January 2026)
- Maren, A.J. (2019/2024). “Derivation of the Variational Bayes Equations. Technical Report TR-2019-01,” Themesis
Inc. (Updated 2024) https://arxiv.org/abs/1906.08804 - Beal, M.J. (2003). “Variational Algorithms for Approximate Bayesian Inference.” PhD Thesis, University College
London. - Friston, K.J. et al. (2017). “Active Inference: A Process Theory.” Neural Computation, 29(1), 1–49.
- Blei, D.M., Kucukelbir, A., & McAuliffe, J.D. (2017). “Variational Inference: A Review for Statisticians.” Journal of the
American Statistical Association, 112(518), 859–877.