The Weekly Mutiny #10: The Second Measurement
Bot Mutiny |
The Weekly Mutiny #10: The Second Measurement
September 28, 2026
Three of this week's findings only appear in a second measurement. A programming experiment logged a twenty point gain on the task and a twelve point deficit on the quiz. A labor paper got the expected answer for men and the opposite one for women, out of the same dataset. An Inspector General audit found a department busy mitigating surveillance risk with nobody assigned to coordinate the work.
The coding score went up twenty points. The recall score was already down before anyone had time to forget.
Christian Bergh, Benjamin Tag, Alexandra Vassar and Jake Renzella ran 59 undergraduates through three introductory C programming tasks, keeping 55 for analysis. One group had ChatGPT-4.5, the other ordinary web search and no generative AI. On the code, the ChatGPT group scored 89% against 69%.
Then the quizzes. Cued recall immediately after the tasks: 41% for the ChatGPT group, 53% for the control. Cued recall 48 hours later: 39% and 52%.
Read the two pairs together. The paper reports no significant difference between the groups in how much recall was lost across those 48 hours. The deficit was there the moment the students stood up. This is not a study about forgetting faster, it is a study about never having encoded the material, and a coding score cannot tell the two apart. Students in the ChatGPT condition also attributed 45% of the submitted code to themselves, against 81% in the control, which suggests they already knew.
The physiological measures, pupil response and heart rate variability, came back inconclusive, and the authors write that substantial data loss limits how far anyone should read them. Coverage of studies like this tends to drop that sentence.
The study on coding scores, recall, and ownership.
Two indices, one dataset, opposite answers about who is exposed
Researchers at the Centre for Protecting Women Online at the Open University, Politecnico di Torino, Nokia Bell Labs and the Fawcett Society linked O*NET occupational data to US Census income distributions broken out by sex, then measured exposure two ways: the Anthropic Index for large language model exposure, and the Artificial Intelligence Index for broader innovation including robotics.
The two indices disagree, and the disagreement is the finding. On the broader index, exposure sits where the familiar story puts it, concentrated in male-dominated occupations. On the language model index it is higher in female-dominated ones. Occupations more than 60% female carry higher Anthropic Index values, and the paper's examples are nurse midwives, skin care specialists, speech-language pathologists, legal secretaries and administrative assistants, and childcare workers.
The sharper result is about distribution rather than level. In male-dominated occupations, exposure concentrates at the high-skill, high-pay end, where workers tend to have some leverage over how a tool gets used. In female-dominated occupations it runs roughly evenly across the whole pay range.
This is a preprint and has not been peer reviewed. Exposure here measures how much of a job's task content a model can touch, which is not a layoff, and the paper says so.
The gendered exposure analysis.
The Justice Department has a surveillance problem and no shared definition of it
The DOJ Inspector General released audit report 26-097 this month, following its June 2025 audit of the FBI on the same subject. The threat it names is ubiquitous technical surveillance: the aggregation of electronic, financial, travel and online records into something that identifies and tracks a person.
The audit found the Department has no common definition of it, no individual or office with explicit responsibility for coordinating the response, and no policy saying how components should coordinate at all. The OIG's summary is that no single vehicle exists for sharing what is known about these threats across the Department. Three recommendations went to the Justice Management Division: a coordinated strategy with a common definition, a risk inventory and mitigation plan that assigns ownership of each item, and a training program with mandatory baseline awareness.
The examples are not abstract. In one cartel case, the FBI found the cartel had hired someone to pull call and geolocation data off mobile devices and used it to intimidate and, in some instances, kill potential sources and cooperating witnesses.
What the report does not do is attribute any of this to artificial intelligence. The mechanism it describes is aggregation of records that already exist, and the audit is about the Department's inability to organize itself against that. Anyone citing it as evidence of an AI surveillance capability is citing something it does not contain.
The audit on ubiquitous technical surveillance.
Filed this week, not yet read
Two more filings landed on the blog this week that I have not opened, so I am naming them rather than summarizing them. Music publishers including Concord and Universal moved for partial summary judgment against Anthropic in the Northern District of California, with a separate statement of facts attached. The Oklahoma Law Enforcement Retirement System filed a verified stockholder derivative complaint against Alphabet's directors in the same district, which is an allegation about what a board knew rather than a finding about it.
One story from a few weeks back is worth a second look because the order itself is now public. In Flexport v. Freightmate AI, Judge Rita Lin excluded a defense expert's data-extraction test because he used an AI agent, Claude Sonnet 4.5, to generate the prompts fed to GPT-4o and never disclosed what those prompts said. Without them, the court couldn't rule out that the agent had handed the model the answer. The judge compared it to an expert who tells an assistant to do a task without ever learning what instructions the assistant actually gave. Same order also let a second expert's trade-secret testimony stand and barred Freightmate from running a "David and Goliath" defense at trial. Read the ruling.
The Black Box lands Friday, October 2, on a UN report that calls a US strike on a school a war crime and never mentions AI, alongside a Pentagon assessment that reportedly does.