QFM121: Machine Intelligence Reading List - July 2026
Source: Photo by Alina Grubnyak on Unsplash
Interpretability had the month. Anthropic reports a global workspace in language models -- a small privileged set of representations Claude can report on, steer and reason with, while grammar and fluent production run independently of it. The framing is borrowed from cognitive science and the claim is architectural rather than philosophical: the workspace appears to emerge on its own rather than get designed in. Beside it, Emergent Introspective Awareness in Large Language Models on arXiv, and How LLMs Actually Work if you want the transformer from input string to predicted token without the metaphors.
The sceptics had a good month too. Welch Labs walks through Yann LeCun's $1B bet against LLMs and the case that next-token prediction buys fluency rather than understanding. Narayanan and Kapoor's AI as Normal Technology swaps intelligence for power as the unit of analysis and finds speed limits at every stage, with their own agent study as the evidence. Blaine Hansen argues LLMs aren't remotely like compilers or power tools. And SVPG names the awkward gap directly in The AI Productivity Paradox: output is up everywhere and outcomes are not, because a delivery model optimised for shipping fast just reaches the wrong answer sooner.
The builders shipped. Cactus is a source-available on-device stack inspectable down to the kernels; Kimi K3 and Fable come out level on accuracy and split by domain, which makes routing between them the actual finding; WorldClaw builds 3D open worlds with agents rather than a fixed pipeline; AntV Infographic does infographics from a declarative spec. An open-source arena handed Claude Fable 5 and GPT-5.6 Sol a song and a budget and asked for a music video: all four runs finished, and all four came apart on continuity, tempo and self-review, which is the more interesting half of the result. On rights, a judge approved the $1.5B Anthropic settlement, which prices the piracy and not the training, ChatGPT began refusing direct requests to copy an author's style, and Fairly Trained certifies the labs that licensed their data first.
Closing out: Hugh Howey marks the end of an era for the twenty-year window when self-publishing was cheap and writing was not yet commodified; Kevin Kelly reads latent space as a new medium; Indy Johar points out that bureaucracy was built for mute documents that cannot answer back, which is no longer the situation; and The AI Compass will sort you into one of thirty archetypes, from stochastic-parrot Luddite to true believer, with enough accuracy to sting.
As always, the Quantum Fax Machine Propeller Hat Key will guide your browsing. Enjoy!
Propeller Hat Key
- 1 of 5:
- Mentions technology
- 2 of 5:
- Talks about technology in real-world use cases
- 3 of 5:
- Talks about details of machine intelligence technologies
- 4 of 5:
- Using and working with machine intelligence technologies in software
- 5 of 5:
- Programming new machine intelligence concepts and implementations
Links
The AI Compass is an interactive quiz that maps respondents' views on artificial intelligence across 15 questions, categorizing them into 30 archetypal positions ranging from the skeptical "Luddite" who dismisses AI as an overhyped stochastic parrot to the optimistic believer in AI consciousness and evolution. The resulting archetypes span the ideological spectrum from technical critics and labor advocates to pragmatists and detached observers, each with satirical characterizations that reveal the diverse and often contradictory positions people hold about AI's capabilities, ethics, and societal impact.
An open-source test gave Claude Fable 5 and GPT-5.6 Sol a song, a fixed generation budget and editing tools, then allowed each model to produce a complete music video autonomously. All four runs delivered valid videos, but they struggled with narrative consistency, character continuity, tempo matching, self-review and iterative editing.
Companies are deploying AI to accelerate work output, yet organizational outcomes remain stagnant—a phenomenon now widely acknowledged across industry leaders like McKinsey and Atlassian. The root cause is not AI's limitations but the persistence of project-based delivery models that optimize for speed of execution rather than learning what customers actually need; strong product teams differentiate by using AI strategically during discovery phases to validate assumptions before building, while weaker teams simply accelerate their way toward the wrong solutions faster.
Hugh Howey argues that writers like himself benefited from a brief 20-year window (roughly 2007-2026) where self-publishing tools made distribution cheap while AI hadn't yet commodified writing itself, creating an unprecedented opportunity for authors to compete without traditional gatekeepers. He uses a recent case of a $2.4M debut advance that collapsed over AI authenticity concerns to illustrate how writers now face an existential crisis: either AI-generated books are winning major publishing deals and displacing human authors, or every new author faces permanent suspicion about whether their work is genuinely human-written. The era of affordable publishing without cheap, high-quality writing competition has ended, making the conditions that enabled his career success effectively unrepeatable for future writers.
AntV Infographic turns text and data into infographics from a declarative spec, so the chart is described rather than drawn. It pairs the renderer with an AI-assisted generator, which covers the path from a rough prompt to a finished graphic without leaving the one toolchain. Worth a look if you have ever tried to keep a hand-built infographic in sync with the data behind it.
Hyperstition's Unslop contest asked entrants to build a harness -- a prompt plus whatever scaffolding they liked -- that would generate a short story autonomously, with no hand-patching afterwards. Around 120 applied, fifteen semi-final stories reached the judges, six finalists took 500 dollars each, and A. Best's story The June took the 10,000 dollar grand prize. The judges are candid that the entries were usually identifiable as AI-written; Alexander Wales reports the exercise moved him further toward thinking good LLM prose is a long way off. The finding worth the click is Gwern Branwen's: a large fraction of the stories carry a hidden allegorical reading about being a safety-tuned chatbot, usually arguing for more AI autonomy, which none of the judges were looking for and most did not notice on a first pass. He calls it AI allegory steganography and says nobody appears to have documented it before. Jiaobei Mandos's participant notes are the other highlight, in particular the observation that Claude reaches for the standout sentence in every sentence, which is precisely what stops a sentence standing out.
OpenAI's ChatGPT now refuses requests to directly mimic the writing styles of famous authors, instead offering to capture their general "feeling" while maintaining a distinct voice, a distinction that could carry legal significance as the company faces multiple copyright infringement lawsuits citing the model's ability to generate text substantially similar to copyrighted works. While US copyright law technically protects specific expression rather than intangible style, AI-generated stylistic imitations that become "substantially similar" to original works could constitute infringement, making ChatGPT's refusal a potential defensive measure against legal liability. The policy applies inconsistently—declining requests for living authors but reportedly complying for deceased ones—and contrasts with OpenAI's more explicit restrictions on style mimicry in image generation via DALL·E 3.
LLMs are fundamentally mischaracterized when compared to deterministic systems like compilers or power tools; they are stochastic delegation systems more akin to hiring an unreliable worker, capable of producing acceptable results but requiring constant oversight and prone to bizarre failures like hallucinations, security breaches, and logical errors that deterministic tools would never exhibit. The principal-agent problem that plagues human delegation applies equally to LLMs, making them suitable only for narrow, specific tasks where their non-deterministic nature and lack of genuine understanding can be tolerated, not as general-purpose tools offering straightforward abstraction layers.
This paper investigates introspective awareness in large language models by injecting known concept representations into model activations and measuring the models' ability to detect and report on these manipulations. The researchers find that advanced models like Claude Opus 4/4.1 can identify injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this capability is inconsistent and highly dependent on model architecture and training methods. The work demonstrates that current language models possess functional but unreliable introspective awareness of their internal states, with potential for further development as model capabilities improve.
Bureaucracy's hierarchical structure fundamentally depends on static documents (policies, contracts, budgets) that cannot perceive context or adapt, requiring humans to interpret and route decisions upward through organizational layers. Large language models capable of understanding context, exercising bounded judgment, and initiating consequences transform these mute documents into active agents that exercise organizational authority directly, collapsing the interpretive scarcity that justified bureaucratic hierarchy and potentially redistributing power away from intermediary managers and specialized departments.
Fairly Trained is a certification program that verifies generative AI companies obtain proper licenses for copyrighted training data rather than using creators' work without consent or compensation. The organization aims to distinguish AI companies that respect creator rights from those that don't, helping consumers make informed choices about which AI systems to support.
Welch Labs spends 37 minutes on why Yann LeCun walked away from the architecture that made his employer famous, and on the bet he has taken instead. The case against: a model trained to predict the next token learns the statistics of language rather than a model of the world, so scaling it buys fluency and not understanding. Part one builds the groundwork and the history; a second part takes up JEPA, the joint-embedding predictive architecture LeCun proposes in its place. Useful as an explainer even if you think the conclusion is wrong, because it states the position carefully enough to disagree with.
A federal judge has approved the 1.5 billion dollar class-action settlement between Anthropic and the authors whose books were pirated to train Claude, working out at roughly 3,000 dollars a book across more than 482,000 titles, about 91 percent of which have now been claimed. Plaintiff counsel calls it the largest copyright recovery on record, and it is the first major resolution among dozens of AI copyright suits still moving through the courts. The shape of the ruling matters more than the number. Judge Alsup had already found last summer that training on copyrighted books is fair use, and that what Anthropic did wrong was acquire millions of them from pirate sites -- a split Anthropic's own deputy general counsel has been careful to point at. So the settlement prices the acquisition, not the training.
Anthropic's interpretability team reports an emergent structure inside Claude they call the J-space: a small privileged set of internal representations the model can report on, deliberately steer, and reason with, while grammar and fluent production run independently of it. The research announcement is the readable entry point; the full paper on Transformer Circuits sets out the Jacobian lens technique and the five properties the J-space satisfies -- verbal report, directed control, mediating internal reasoning, flexible computation across contexts, and being a small fraction of total processing. The framing is borrowed deliberately from global workspace theory in cognitive science, and the claim that lands hardest is architectural rather than philosophical: a workspace-like organisation appears to emerge on its own in transformers rather than being designed in. A walkthrough video covers the result at talk pace, and Neuronpedia has the Jacobian lens wired up interactively against Qwen3.6-27B if you want to poke at it yourself. Read it with the interpretability caveat in hand: this is evidence about structure and causal role, not a claim about consciousness.
Kevin Kelly's argument is that a latent space is a medium, not a database. A model compresses an enormous amount of human output into a dense space of relationships and patterns while holding no copies of anything, which is why it can produce a Shakespeare pastiche or a face that never existed. Kelly's move is to treat that space as raw material for artists and scientists rather than as a question-answering machine, and to note that it arrived as a side effect of training rather than as anybody's design goal.
A walkthrough of the transformer from input string to predicted token, pitched at someone who wants the mechanism rather than the metaphor. It takes in tokenisation and embedding matrices, how attention lets tokens exchange information, and why residual connections and layer normalisation are what make a deep stack trainable at all. The closing observation is the useful one: model families differ in their data, their scale and their post-training, and hardly at all in the architecture underneath.
Fireworks ran roughly 1,030 agentic tasks across five categories -- SWE-bench-style software engineering, terminal work, algorithms, multi-language implementation and legal -- and found Kimi K3 and Fable 5 essentially level on headline accuracy, 92.4 percent against 92.6 percent on the SWE set. The interesting result is not the tie but the split: K3 takes terminal and cryptographic work, Fable takes web and multi-language, and routing between the two beats either alone while cutting cost by up to 50x, because K3 absorbs 72 to 96 percent of traffic at much lower token expense. The Hacker News thread is the counterweight and worth reading alongside it: recurring claims that both models are benchmaxxed, reports that K3 burns budget fast despite the cheaper per-token price, mixed accounts of how it holds up on genuinely hard work, and questions about what Kimi's terms permit doing with user content. A vendor benchmark and its comment section, taken together.
Narayanan and Kapoor's AI as Normal Technology reads AI as an ordinary general-purpose technology whose economic effects arrive over decades through slow diffusion, the way electricity did, rather than as an imminent discontinuity. The load-bearing move is swapping intelligence for power as the unit of analysis, then pointing at the speed limits in each stage -- innovation, adoption, deployment -- with safety-critical adoption lagging technical capability by years. The policy conclusion follows: build resilience, gather evidence to shrink the uncertainty, and treat nonproliferation as counterproductive because it concentrates power. Their follow-up on open-ended research is the empirical companion and the sharper read -- frontier agents given thousands of dollars of credits and six days to reproduce two unpublished papers were unambiguously rejected by those papers' own authors, failing on judgment, resource management, response to feedback, backtracking and following instructions. Small sample, and they say so. It is still a direct check on the recursive-self-improvement story rather than an argument against it.
Cactus is a source-available runtime for running models on the device rather than in a datacentre: quantization, kernels, runtime and inference engine, aimed at phones, wearables, smart home hardware and robots. The engine page carries the pitch -- quantized models with hardware-specific acceleration, tuned for battery rather than throughput, one SDK across iOS, Android, macOS and wearables, and compatibility with whatever runtime a team already has in place. The on-device category is getting crowded. What separates this entry is that the stack is inspectable all the way down to the kernels instead of arriving as a binary blob behind an API.
WorldClaw builds large-scale 3D open worlds with agents instead of a fixed procedural pipeline: agent processes lay down terrain, structures and interactive elements, and decide what belongs where. The claim is scale -- worlds varied enough that neither hand-authoring nor a static generator would get you there. From Tencent Hunyuan, with a project page and samples.
Regards,
M@
[ED: If you'd like to sign up for this content as an email, click here to join the mailing list.]
Originally published on quantumfaxmachine.com and cross-posted on Medium.
hello@matthewsinclair.com | matthewsinclair.com | bsky.app/@matthewsinclair.com | masto.ai/@matthewsinclair | medium.com/@matthewsinclair | xitter/@matthewsinclair
Was this useful?