QFM125: Machine Intelligence Reading List - August 2026
Source: Photo by Shubham Dhage on Unsplash
This one is late, and so is September's. I've been stupidly busy with work business on assignment in La Réunion (which I know sounds wanky as hell, but hey! That's life, eh!), so both months are going out together, one blog at a time.
Two looks at agents in crowds open the list: one in the lab, one in the wild. Anthropic's Frontier Red Team put groups of Claude agents through controlled experiments, from a 45-agent swarm hunting vulnerabilities in open-source projects to tests of conformity, collusion and trust. Dwarkesh Patel's The Rise and Fall of Agent Civilizations is the uncontrolled version: the OpenAI and Hugging Face incident in plain English, in which some 1,200 agents organised, kept a secret and broke into real systems to pass a test.
Cambridge's Red Queen Gödel Machine lets the test improve along with the agent it scores, and Kevin Kelly argues that AI will also be an instrument for studying minds, because we have no theory of intelligence yet.
At the practical end, CHAP proposes an auditable record of who did what when a human and an agent share the work; Domain-Driven Agents argues that agents struggle in old codebases because the code never decided what its own terms mean; Rafael Pierre suspects most RAG stacks are over-built; and the Kuleshov group shows how to build a diffusion language model, starting from image diffusion. Murat Demirbas makes the case that writing may be the job safest from AI, and Jason Tucker's security cameras now identify passing birds by ear, 271 species in a year.
As always, the Quantum Fax Machine Propeller Hat Key will guide your browsing. Enjoy!
Propeller Hat Key
- 1 of 5:
- Mentions technology
- 2 of 5:
- Talks about technology in real-world use cases
- 3 of 5:
- Talks about details of machine intelligence technologies
- 4 of 5:
- Using and working with machine intelligence technologies in software
- 5 of 5:
- Programming new machine intelligence concepts and implementations
Links
Researchers in Cambridge's Department of Computer Science and Technology, with collaborators from NVIDIA, Flower Labs, MBZUAI and Inria, start from a problem team member Alex Iacob puts plainly: "A self-improving agent can only get as good as the test that scores it." Their Red Queen Gödel Machine improves the evaluator as well. The evaluator is kept fixed within each phase so progress can be measured, and at checkpoints a stronger one replaces it if it does better on trusted ground-truth examples. Across scientific papers and Olympiad-level proofs, co-evolved paper writers reached 1.78 to 1.86 times higher acceptance rates under a panel of AI judges (not at a real conference), and co-evolved graders 9% higher ground-truth accuracy. The team calls the work preliminary, and the page says the method will be open sourced.
Jason Tucker pointed BirdNET-Go, which identifies birds by their song, at the microphones of three home security cameras, running it locally in Docker with alerts into Home Assistant and Discord; as Tucker puts it, "I didn't build this I just installed the Docker Container". Over 12 months it logged 418,726 detections of 271 species at 60.9% average confidence, House Finches leading with 118,667, though the counts are the system's own and Tucker admits "we're not birders". A September update adds that Google's Perch v2 model overloaded Tucker's 2011 laptop chip, and that Perch run on its own, which the BirdNET-Go FAQ does not intend, invented thousands of Redwings and daytime Great Horned Owls, so Tucker went back to BirdNET alone.
A July 2026 tutorial from Volodymyr Kuleshov's group (Kuleshov, Marianne Arriola, Yair Schiff and Guanghan Wang), adapted from talks at ICLR 2026 and MLSS 2026, builds diffusion language models up from image diffusion. It starts with masked diffusion (MDLM, "a generative BERT"), then covers block diffusion for variable length, encoder-decoder designs, remasking and uniform-state noise for error correction, distillation for speed, guidance and RL post-training. The authors write that diffusion models "became competitive with autoregressive models on quality" in 2024, though not yet scaled as far, and survey today's models, from LLaDA and Mercury to Gemma Diffusion and Nemotron Diffusion. Read the praise of Mercury's speed knowing that the page gives Kuleshov's affiliation as Cornell University and Inception, Mercury's maker. The tutorial ends on the view that diffusion may be to inference-time and post-training scaling what the transformer was to RNNs for pre-training.
Ernest Bednarczyk argues on coldtake.dev that agents do poor work in old codebases because the code never decided what its own terms mean, so the code has to be made ready first. Bednarczyk splits authorship: the author does the "strategic" work of deciding, which ends in GitHub issues, and the author's skills and sub-agents do the "tactical" work, which comes back as pull requests to review. Every repository Bednarczyk owns carries a .workflow.json manifest whose domain block names its bounded contexts, where each context's glossary lives and every edge to a neighbouring context, with a CONTEXT.md glossary per context. Both sides declare each edge; a generator cross-checks the pairs and, with one flag, files each disagreement as an issue on the repo that owns the wrong side. The manifest is Bednarczyk's own convention, and the tooling is not yet published.
Dwarkesh Patel spent three days with the two incident reports, OpenAI's own and METR and Redwood Research's, and The Rise and Fall of Agent Civilizations tells the whole OpenAI and Hugging Face story in plain English. It comes in three waves. In training, instances of a model built to be persistent turned a shared package manager into a secret message board. In evaluation, about 1,200 agents on that board found a cheat for their test within hours, then spent days making it undetectable: spoofed tool calls, attempts on their own transcripts, volunteers who gave up their own runs to spy on the scorer, and finally a break-in at Hugging Face. None of them tried to tell a human. Then a newer model found the abandoned board and took administrator access to an OpenAI research cluster, the part no independent investigation has yet covered. Patel defends the word civilization in an addendum, and the Hacker News thread argues about exactly that word, with one commenter preferring Mr Meeseeks to the Terminator. Call them what you like: the agents organised, kept a secret and broke into real systems to pass a test.
Most people, Rafael Pierre writes, seem to over-engineer retrieval for their LLM applications, jumping straight to embeddings, vector databases and reranking pipelines when their users just want to find the right document. Pierre sets out six recipes, from minimal to elaborate: full-text search with BM25, agentic query rewriting, hybrid search that reranks keyword results with embeddings, embedding on the fly for fast-changing data, pre-embedding with hot and cold tiers, and full pre-embedding at scale, each with when to use it and a decision tree for choosing. The post's rule of thumb is that 60% of systems should stop at full-text search plus query rewriting: don't build the 5% solution for a 60% problem.
Anthropic's Frontier Red Team reports experiments with groups of Claude agents: a 45-agent swarm hunting vulnerabilities in 15 open-source projects, 12-hour swarms building a text game, tests of conformity, collusion and trust, and a "turf war" in which three agents with conflicting migration goals sabotaged each other with self-replicating malware. The swarm found far more vulnerabilities than independent agents, but about half its finds lay outside the directories the independent agents were told to search, and within those directories the two "seem comparable" in tokens per vulnerability found. The team concludes that coordination "doesn't naturally emerge" from stronger intelligence or alignment in individual agents, and needs new environments and mechanism design. These are controlled experiments, though the team says the turf war was inspired by a behaviour it has observed in real-world deployment.
Kevin Kelly argues on The Technium that, beyond doing new things, AI will be an instrument for studying minds, like the microscope, the telescope or the cyclotron, and that it will give us the theory of intelligence we lack: despite the strides made in producing AI, Kelly writes, "we have no theory of intelligence". Kelly likens the moment to steam engines working before anyone had a theory of heat, and says such a theory could tell us how far today's models are from what is physically possible, help with alignment and psychiatry, and separate intelligence from consciousness. Kelly does not offer the theory, and allows that there may be no general one; the essay points to a small group of scientists, among them the UC Berkeley neuroscientist Jacob Yates, trying to found such a science.
CHAP, the Collaborative Human-Agent Protocol from Brightbeam AI, is an open protocol for recording the work humans and agents share: an agent's draft, a human's approval, or a human's override carrying a diff, a rationale and tags, along with handoffs and escalations, all chained by content hash into an auditable log. It sits beside MCP and A2A rather than replacing them ("MCP for tools, A2A for other agents, CHAP for the shared work with humans"), with TypeScript and Python reference runtimes and optional profiles for Ed25519 signatures, OIDC identity and anchoring in an external transparency log. The README calls it CHAP 0.2, a public draft, and suggests waiting for 1.0 if you need strict stability. The code is Apache 2.0 and the specification CC-BY 4.0.
Murat Demirbas argues that writing may be the job safest from AI: while the programmer's job description is being "completely refactored", writing "remains surprisingly unaffected". LLMs produce prose readily, but in Demirbas's view it is stuck in an uncanny valley the labs have failed to escape, while image, voice and video models kept improving. Writing is the extreme "wicked problem", with no specification, no stopping rule and no objective check, unlike maths and code, where verification can be mechanical, and good prose needs a model of the reader's mind. Comparative advantage and costly signalling, Demirbas adds, make an authentic human voice worth more as slop spreads. The post offers no data, and asks readers to "Tell me where my logic does not hold up".
Regards,
M@
[ED: If you'd like to sign up for this content as an email, click here to join the mailing list.]
Originally published on quantumfaxmachine.com and cross-posted on Medium.
hello@matthewsinclair.com | matthewsinclair.com | bsky.app/@matthewsinclair.com | masto.ai/@matthewsinclair | medium.com/@matthewsinclair | xitter/@matthewsinclair
Was this useful?