The Big Ones Scare Me
Very small models are getting remarkably capable.
Raspberry Pi Uplift Program
I have said this a few times in a few different contexts, but I truly did imagine our post-superintelligence future would be far more homogenous than it’s turning out to be. As it turns out, there are many applications for a country of geniuses in a data center. Or, at least, a country of grad students. But, unexpectedly, there also seem to be a number of situations where you’d prefer a moron on a Raspberry Pi.
In this excellent post from friend-of-the-show Alexander Doria’s company Pleias, the team talks through their approaches to two different projects involving extremely small language models run locally on cheap, offline hardware.
Over the past months, we have deployed small language models — most under 500M parameters, all under 2B — on Raspberry Pi hardware for two production use cases: a legal assistant for conflict-related sexual violence (CRSV) survivor networks, built with Bibliothèques Sans Frontières and the Dr. Denis Mukwege Foundation, and grounded question-answering over a French primary-school curriculum from Senegal. Both are knowledge-work applications — legal analysis, lesson preparation — of the kind usually assumed to require cloud-hosted models, running in environments where cloud infrastructure is not an option.
The capabilities here are very different from those associated with small, general purpose LLMs that you might run locally on a high-end PC or Mac. The models are designed to answer specific questions by referring to an authoritative document, and do so offline, with the text of the document itself stored directly in the KV-cache for Cache Augmented Retrieval (CAG).
I see these less as an alternative to frontier or open source models—there is a distinctly European flavor to the types of problems these models/devices have been designed to solve, which is not to say that they’re unworthy—and more as a small example of the enormous possibility space outside of general purpose LLMs. This space has gone largely unexplored so far, for the obvious reason that anyone aware of it could much more lucratively focus on capabilities that are more general and closer to the frontier.
But it seems very high potential. I look forward to the day when there are many little gadgets running tiny local models of all kinds, each one reducing a little more the pre-LLM friction of interacting with digital tools in natural language. A sentient Raspberry Pi in every pot, please.
What About the Boyfriend Models?
It’s not just the very smallest models, either. There is a newly-vibrant middle in the model space. The Boyfriend Model, if you will. The Bridgewater-Thinky paper from last week showed very impressive results for postraining open-weight models on specific tasks—a benchmark improvement over frontier models at something like 1/10th of the price. The ever insightful 47fucb4r8curb4fc8f8r4bfic8r pointed out that this should be a wakeup call for any enterprise using existing frontier models to anything that’s repeatable, scalable and trainable.
What about the models that are slightly bigger again? New Meta model dropped. Alexander Wang has been standing behind MSL engineers with a cattle prod making them do Wordle puzzles and Leetcode mediums, and guess what, it’s working. Muse Spark 1.1 might not be a free range model, but it is pretty good. And the price… the price is very compelling. So cheap that it might compete directly with self-hosted OS models.
Clanker Cambrian Explosion
The phrase vibe-coding has been in general use for sometime now to refer to web- and iOS-apps. Sometimes it’s deployed (usually in anger) to refer to the activities of someone who purports to be a software engineer working on production code. But I can now foresee a time when people will use frontier models to vibecode their own small LLMs, for use on cheap consumer hardware like the Raspberry Pis from the Pleias post above.
You won’t be able to use Fable for this, of course. Anthropic have put ML research in the same category as cybersecurity and biology, with heavy-handed safegaurds attached. But GPT-5.6 Sol seems very capable. A few tweets today suggested that GPT-5.6 Sol was used to postrain the smaller, cheaper 5.6 Luna model. Very interesting stuff.
Don’t get too freaked out just yet—like the Anthropic RSI post from a while back, there are a few big caveats.
Of course, this is this worst it’ll ever be.












