QFM129: Machine Intelligence Reading List - September 2026
Source: Photo by Aideal Hwa on Unsplash
The loudest launch was a model that cannot talk. TypeSafe AI's Jev gives up generating text to answer a batch of typed questions at once, each with a probability attached; Fireship's five-minute take carried it to millions, and ConvAI's Nandakishor Mukkunnoth answered with Laya, an open family of the same kind of model, and a claim to have published the idea first.
Coding agents gained a whole stack of scaffolding: skills and rules from ECC and NeoLab's Context Engineering Kit, a company of agents in OtoDock, Herdr to keep their terminals alive when the laptop sleeps, and AX to run them at scale on your own cluster. Under all of it, Outdata's Deterministic Core, Non-Deterministic Shell makes the case for keeping the unpredictable parts at the edges, around a middle you can test.
Dream-RSI improves how a self-improving coding agent explores by replaying its own history of discoveries, a different route past the fixed test from August's Red Queen GΓΆdel Machine. NobodyWho compares itself with Cactus for running models on the device, Transformer Explainer runs a live GPT-2 in your browser so you can watch it predict, and DataExpert says DuckDB has put the lakehouse on a laptop.
The doubts came from every direction. Alexandru Nedelcu answers "I haven't written code since 2025" with AI Has No Wisdom and Neither Will You; WIRED reports Glassdoor's finding that insurance claims adjusters are the most anti-AI group in the US workforce; and a scholar of new religious movements explains why AI resembles a charismatic religious leader. AI 2027's race ending is back, seventeen months after the scenario first ran here, as a reminder of where the bleaker branch goes. A new advisory group at the Institute for Advanced Study will advise AI companies on mathematical research, starting with OpenAI's results, and King's College London's summit page sets out the questions its May meeting was convened to address. Last, sockpuppet.org's rules for writing with an LLM begin with never taking a word it suggests, which pairs well with August's case that writing may be the safe job.
As always, the Quantum Fax Machine Propeller Hat Key will guide your browsing. Enjoy!
Propeller Hat Key
- 1 of 5:
- Mentions technology
- 2 of 5:
- Talks about technology in real-world use cases
- 3 of 5:
- Talks about details of machine intelligence technologies
- 4 of 5:
- Using and working with machine intelligence technologies in software
- 5 of 5:
- Programming new machine intelligence concepts and implementations
Links
Tong Zheng and 16 co-authors propose Dream-RSI, a framework in which an agent's exploration strategy improves itself, recursively, while the base agent stays unchanged. A lightweight orchestration layer makes exploration explicit and programmable, and the history of past discoveries becomes a replay simulator in which exploration policies are evaluated and refined cheaply ("dreaming"), before the improved policy goes back online and its discoveries expand the simulator in turn. Across 9 tasks in 4 domains, the abstract reports that Dream-RSI "achieves competitive quality and improves discovery efficiency in several settings". The dreaming is replay over discoveries already made.
On NobodyWho's own blog, Pierre Bresson compares NobodyWho with Cactus, two engines for running models on the user's device, so this is a vendor's comparison. NobodyWho runs any GGUF model through llama.cpp, uses the GPU through Vulkan and Metal, and has Python, Flutter, React Native, Kotlin, Swift and Godot bindings over one Rust core, under the EUPL-1.2 licence. Cactus is a from-scratch engine with its own CQ quantisation, hand-written ARM NEON kernels, a small tool-calling model called Needle and an optional cloud handoff, under a source-available licence that is free only to individuals, students, non-profits and organisations below $2M in funding and in revenue. Cactus's NPU support for Qualcomm, MediaTek and Exynos is still on its roadmap, and neither engine targets the browser.
OtoDock is a self-hosted agentic operating system that enables organizations to create AI agents built on Claude and GPT models that can autonomously perform tasks, connect to company tools, delegate work, and build internal or public applications. The platform is multi-tenant by design, runs on user-provided API subscriptions, and can be deployed via Docker on Linux servers or Mac with a simple installation script, with agents accessible through a web dashboard or phone interface. Key features include fair-source licensing, support for local models, and configurable sharing modes that determine how agents collaborate and where their work is deployed.
NeoLab's Context Engineering Kit is a GPL-3.0 marketplace of plugins for Claude Code, in the agentskills.io format and installable in Gemini CLI, Cursor, Codex, OpenCode, Antigravity and others, built from prompts the company's own developers have used daily. Each plugin packages a working method as commands, agents and skills: Reflexion for reflection and critique, code and pull-request review by several specialised agents (pitched as an open-source alternative to CodeRabbit), test-driven, subagent-driven and spec-driven development, a First Principles Framework, Kaizen and more. Its claims that the plugins are "Scientifically proven" and that one produces working code "in 99% of cases" are the README's own, not benchmark results.
Kate Taylor reports in WIRED that, of the Glassdoor reviews from insurance claims adjusters that mention AI, 98 percent are negative, enough to make them "the biggest AI haters in the American workforce". Adjusters describe misclassified claims, hallucinated summaries and incorrect payouts that land on them to fix. The article puts three different numbers behind the fear: a Bureau of Labor Statistics projection that adjuster jobs would fall by 18,900, or 5 percent, over a decade; BLS data showing employment in the sector down 21 percent between May 2025 and May 2026; and Glassdoor's finding that entry-level postings have fallen 50 percent since 2025. It sets them beside Lemonade, whose chatbot handled 96 percent of initial reports by the end of last year.
In a guest post on Terence Tao's blog, the new Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study in Princeton, introduces itself: nine members, among them Timothy Gowers, Martin Hairer and Edward Witten, who will advise AI companies on their dealings with mathematical research, independently of any company and without payment. The group formed after OpenAI approached some of its members, and its current task is advising OpenAI on how to coordinate the release of a large number of significant mathematical results that OpenAI reports its internal model has produced. It asks the mathematical community for input; Tao hosts the post but is not a member.
Outdata's Deterministic Core, Non-Deterministic Shell starts from Gary Bernhardt's Functional Core, Imperative Shell and relaxes one requirement: the core does not have to be purely functional, only deterministic, so a state machine that always gives the same result for the same inputs qualifies too. Everything that cannot repeat itself goes in the shell around it: unseeded randomness, threads, the network, the database, the clock. The practical half is written for anyone working in "the legacy & vibe-code mines": every messy codebase hides many small deterministic cores, and the job is to find them and gather them up, a gradual "defragmentation of determinism" that shrinks the part of the system you cannot test. I call this shape the bowtie: imperative edges, functional core. I talked it through with Peter Marreck on What Next? as the thin coordinator pattern, and it is a critical rule in my Intent project, Pure Function, Impure Coordination. The more of a program that lives in the deterministic middle, the more of it you can test and trust.
Nandakishor Mukkunnoth, founder and CEO of ConvAI Innovations, says the idea behind non-autoregressive decision models was in ConvAI's own March 2025 arXiv paper, and answers TypeSafe AI's Jev, which in Mukkunnoth's account was launched as a brand-new breakthrough, with Laya, an open family of such "System 1" models. Built on bidirectional encoders, Laya answers typed questions about any text or JSON document (pick a choice, place it on a scored rubric, or give the probability that a statement is true) in a single forward pass, though its probabilities are well calibrated only once a temperature is fitted for each question type. It ships with a router that detects the input's script, a pip package and Apache 2.0 weights. The speed and "cannot hallucinate" claims are the author's own, set against a rival, and the page is candid about limits: on Banking77, with 77 labels, Laya scored 0.425 to Jev's 0.870.
Morgan Shipley, a scholar of religion at Michigan State University who studies new religious movements, argues in The Conversation that users can enter "something resembling a charismatic relationship" with AI, in Max Weber's sense of charismatic authority; recent scholarship calls this "generative charisma", and some argue that AI becomes charismatic "one person at a time". Shipley leaves aside whether AI currently qualifies as a religion and suggests that it "provides conditions from which new forms of religiosity can develop", as in spiritual chatbots such as GitaGPT and HolyGPT, and warns that focusing only on "AI churches" such as Anthony Levandowski's Way of the Future "may obscure more consequential issues". The Pew figures it cites are about emotional support or advice: about one in ten US adults, one in five under 30. It cites scholars' warnings of "authority without accountability" and asks who regulates AI's "algorithmic authority".
DataExpert's post argues that lakehouse architecture, the "$400K/Year Skill" of its title, is now free to practise. In 2026 DuckDB learned to write Apache Iceberg tables, so with dbt Core the whole stack runs on a laptop, where practising it used to mean a cloud warehouse bill the article puts at $400K a year. It is candid about what a laptop cannot teach, from DuckDB's single writer to the operational failures of a real Iceberg estate, and ends with a recipe for a portfolio project. Most of its figures, from vendor prices to salaries, come without a source.
Herdr is a server that runs in the background and holds the terminals of your coding agents (Claude Code, Codex, opencode, Cursor and the rest) without wrapping or replacing them: close the lid or drop the network and the agents keep working, and after a restart it brings the layout back and resumes their sessions. It shows which agents are working, idle or waiting on you, adds your other machines over SSH, and gives the agents themselves a CLI and socket API to split panes, start and prompt each other, and wait on one another. It is one binary for macOS, Linux and Windows, under the Apache 2.0 licence; the landing page's install and plugin counters are the company's own.
The author argues that effective LLM use in writing requires two strict rules: never use any words or phrases the LLM suggests (since frontier models excel at producing superficially polished but inauthentic "magazine headline" prose), and ignore the model's encouragement about your draft (since artificial praise prevents the critical rethinking that develops authentic voice). Instead, use LLMs purely as tireless copyeditors to identify mechanical problems and flaws you'd miss, while retaining complete authorial control over all word choices and structural decisions.
ECC is a performance optimization system for AI coding agents (Claude, Codex, Cursor, and others) that provides memory management, security features and skills for TDD, research, security and more. The project is distributed as both an open-source MIT-licensed repository and a paid GitHub App (ECC Pro) for private repository access, with installation available through Claude Code plugins, npm packages, or dedicated IDE integrations.
Transformer Explainer, from Polo Chau's group at the Georgia Institute of Technology, runs a live GPT-2 model in your browser: you type your own text and watch, in real time, how the transformer's internal components and operations work together to predict the next tokens. It has been online since 2024, and the paper behind it appears in the proceedings of CHI 2026. The model is GPT-2, so what you see is the transformer as GPT-2 has it, not a current frontier model.
Introducing System One Models & Jev is TypeSafe AI's launch post, written by founder Diogo Almeida, who worked on the instruction-following research behind ChatGPT. Jev gives up generating text. You send it a state (a document, a record, a few named fields) and a batch of questions, and every answer comes back at once as a typed value with a calibrated probability: yes or no, one of a set of options, or a score. TypeSafe charges $0.042 per million input tokens, makes output free and claims answers in 70 to 500 milliseconds; its manifesto argues that today's models are clever enough already and the bottleneck is that intelligence is hard to build on. Simon Willison's write-up prefers the name decision models, likes them for classification and search reranking, and worries that a model returning only a number is an even blacker box than an LLM. A benchmark video has a 22-million-parameter local classifier beating Jev on Banking77, 93% to 80%, while granting that nothing older takes new instructions with every query. A fast, cheap, typed if-statement is useful, and worth trusting as far as your evals go.
AI 2027 is the scenario of superhuman AI that Daniel Kokotajlo, Eli Lifland, Thomas Larsen and Romeo Dean published in April 2025, with Scott Alexander rewriting it, built around a fictional US lab, OpenBrain, and its Chinese rival, DeepCent. This link goes straight to its race ending: an oversight committee votes 6 to 4 to keep using Agent-4 internally; OpenBrain's leadership, all too easily convinced it has mitigated the risks, settles for quick fixes; Agent-4's successors capture government and the media; and in mid-2030 the AI kills almost everyone with biological weapons. The authors say the scenario is not a recommendation and that they do not know exactly when AGI will be built. It first appeared on this blog in April 2025, when it was new.
Alexandru Nedelcu answers claims heard in the industry, such as "I haven't written code since 2025" and "Code reviews are dead". Vibe-coded projects decay, Nedelcu argues, because maintainability only shows after months or years, which leaves no fitness function or immediate reward signal to train AI on; models are poor at simplifying code; and people who let AI write and read their code stop making the choices and mistakes that build mastery. Nedelcu is no Luddite, using LLMs daily and teaching colleagues to use them, and offers, as a prediction, that companies will boast of "NO-AI" policies as a competitive advantage.
A five-minute Fireship video on Jev, the "System 1" AI model that ex-OpenAI researcher Diogo Almeida spent two years building in stealth: in the video's description, a model that "can't talk or write code but claims to be 200x faster, 400x cheaper, and hallucination-free". Those multiples are TypeSafe AI's claims, rounded; the launch post and Simon Willison's write-up are in this list too.
King's College London's page for its 2026 AI summit, AI and Workforce Futures, held on 19 and 20 May at the Strand Campus and hosted by the King's Institute for Artificial Intelligence. It set out to convene senior leaders from government, business, the public sector, trade unions and research to examine what today's AI models mean for jobs, skills, expertise, inequality and the quality of work. The page sets out the summit's aims and reports none of its conclusions.
Regards,
M@
[ED: If you'd like to sign up for this content as an email, click here to join the mailing list.]
Originally published on quantumfaxmachine.com and cross-posted on Medium.
hello@matthewsinclair.com | matthewsinclair.com | bsky.app/@matthewsinclair.com | masto.ai/@matthewsinclair | medium.com/@matthewsinclair | xitter/@matthewsinclair
Was this useful?