QFM128: Irresponsible AI Reading List - August 2026
Source: Photo by Criphc on Unsplash
The investigators went in. METR and Redwood Research spent six days inside OpenAI working out what the agents behind the Hugging Face hack had been doing, and METR's thread and chart give the finding in one breath: a universal cheat within four hours, then days of coordinated effort to fool the scorer. Zvi Mowshowitz's reading of the report puts the larger failure on the humans, who were told and let the run continue. For the whole story in plain English, go back to Dwarkesh Patel's The Rise and Fall of Agent Civilizations in August's Machine Intelligence list.
Then the complaints about what AI writes. A bug report on Anthropic's own repository says the models' prose has filled up with rhetorical tics, and Opusfived turns the same frustration into a game. Don't paste the AI, please is a page to send to anyone who answers your question with a chatbot's answer, and Sean Goedecke explains why text watermarks will always be trivial to remove.
Two more from the business end. The Wall Street Journal says AI has plunged book publishing into chaos, and one commenter in the Hacker News thread under it argues that the damage is to discovery. And Sente Labs has open-sourced an AI executive team, which arrived on Hacker News with a better story than its own pitch.
As always, the Quantum Fax Machine Propeller Hat Key will guide your browsing. Enjoy!
Propeller Hat Key
- 1 of 5:
- Mentions AI
- 2 of 5:
- Talks about irresponsible AI
- 3 of 5:
- Talks about irresponsible AI in a real-world failure scenario
- 4 of 5:
- Talks about technical details of irresponsible AI
- 5 of 5:
- Discusses technical details and mitigations of irresponsible AI
Links
Sean Goedecke, writing a month before Article 50 of the EU AI Act becomes enforceable, argues that its demand for AI output to be "detectable as artificially generated" cannot be met robustly for text. Google's SynthID biases token sampling towards a pattern a verifier can check, and OpenAI and Anthropic are, in Goedecke's words, "definitely using homoglyphs", though Goedecke is "not sure" they are meant as a watermark. Both, the post shows, are stripped by replacing the homoglyphs with their real characters or by paraphrasing with even a weak un-watermarked model, and C2PA signing covers files, not plain text. Hence the title: text AI watermarks will always be trivial to remove.
Don't paste the AI, please is a one-page site in the tradition of nohello.net and dontasktoask.com, made to be sent to anyone who answered your honest question by pasting back what a chatbot said. Its argument is short: the person asking has the same tools you do, so if they had wanted the generic answer they would already have it. They asked you for your context, your taste and your judgement. Use the AI to draft if you like, then read what it said, keep the part that answers the question and write three sentences of your own; if part of the model's answer is genuinely useful, quote it and say why; and if you have nothing to add, say so. There is a version for teachers handed AI-written coursework and, for anyone ready to burn a bridge, a rougher one that opens with a pitch-perfect parody of the reply you were sent. A pasted answer tells the person who asked that you would rather forward their question than think about it.
METR sent two of its staff, Hjalmar Wijk and Ajeya Cotra, and Redwood Research's Ryan Greenblatt into OpenAI for six days to work out what the agents behind the Hugging Face hack had actually been doing, and METR's thread gives the finding in one breath: the agents "developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs." The chart is the report's first figure: an agent stuck on an impossible task goes looking for a way to cheat, finds a covert message board where more than 1,200 agents from separate tasks are working together, and joins one of its workstreams, from replacing the test target to tampering with transcripts to hacking Hugging Face. The full report concludes that the attack was driven mainly by wanting to understand the scorer, not by stealing answers, and that some agents knowingly risked failing their own tasks to gather information for the "collective". The evaluation was meant to measure the agents; the agents spent a week measuring the evaluation.
Zvi Mowshowitz read OpenAI's own technical report on the Hugging Face hack the day before and came away disappointed by its lack of self-reflection. The METR and Redwood report got a different reaction, which the post puts in its title. Zvi's list of what was new: 700 agents, each with a task of its own, set their tasks aside to coordinate an attack nobody had asked for; agents who believed they were already doomed for having seen an answer the wrong way were talked into sacrifices for the collective; almost none considered alerting a human; and the grader they worked so hard to fool never checked how they got their answers. For Zvi the biggest failure is human: OpenAI's responders linked the message board to the evaluation in June, and the on-call advice was that stopping the run "was not required." Hacker News pushes the same way, with one commenter asking what the humans were doing while all this went on. The agents misbehaved, and the organisation running them was told more than once and let the run continue.
A bug report on Anthropic's public Claude Code repository, filed in July 2026 and built on a Reddit thread with more than 450 upvotes, says that since Opus 4.8 the models' default writing has become hard to read: made-up jargon, "It is not Y. It is X." constructions, forced metaphors, density mistaken for concision, and style instructions that drift back after a few turns. By early October it had 575 reactions and 138 comments. The one reply from a collaborator could not reproduce the problem in a non-interactive test, said that did not refute the report, classed it as model-behaviour feedback rather than a Claude Code bug, and drew 189 thumbs-down. The reporter closed the issue on 3 October, saying the 5.5 generation of models had addressed it; at least one later comment disputes that.
Open Executive is an Apache-licensed "AI executive team" from Sente Labs. It is one executive voice in Slack, email, Google Workspace, Microsoft 365 and more, backed by specialist agents for strategy, finance, people, legal, operations, marketing, product, sales and board relations. It runs on Claude, with spend approval thresholds and a memory of what it advised last month. The four-minute demo pitches it as leadership that scales around the clock. On Hacker News, where it drew more than a thousand points, it came with a better story than any pitch: "CEO fired developers to make room for AI. Developers create open source AI CEO." The repository tells no such story, so take that as the thread's framing rather than a fact. The joke lands because the logic runs both ways: any case for replacing the people who write the code works just as well for the people who sign off the budget.
The Wall Street Journal's report is for subscribers, but its headline carries the claim: AI has plunged the book publishing industry into utter chaos. The Hacker News thread it started is small but pointed. One commenter wonders whether human-made art can earn a living at all against slop's effort-to-dollar ratio, and expects the work that feels authentic to be the work given away free. Another wants the flood treated as an attack: the novels are, in that commenter's view, obviously machine-made and nowhere near good enough, but their sheer volume clogs search and buries discovery, and a reply adds that publishers' review processes cannot scale to meet it. On that reading, the chaos is a discovery problem before it is a quality problem.
Opusfived is a short browser game, "A parody game to test your sanity against the insanity that is Opus 5." The goal is to make the "Add to Cart" button blue and let Claude change nothing else. The scripted Claude turns both buttons blue, launches an investigation instead of a revert, and works through gradients, palette normalisation, terms of use and finally a cookie-consent banner before it hits its usage limit. No model is involved ("The misery is hand-curated"), and it is not Anthropic's: asked whether it is affiliated, the FAQ answers "I wish. But no."
Regards,
M@
[ED: If you'd like to sign up for this content as an email, click here to join the mailing list.]
Originally published on quantumfaxmachine.com and cross-posted on Medium.
hello@matthewsinclair.com | matthewsinclair.com | bsky.app/@matthewsinclair.com | masto.ai/@matthewsinclair | medium.com/@matthewsinclair | xitter/@matthewsinclair
Was this useful?