AI Homework Atrophy, GitHub Commits Double, Nick Muy Sit-Down & AI Sandbagging
We gave four AI models the same question from three user profiles. All four gave worse answers to the user they judged unable to check them. This week: a 26,000-student study on what AI homework does to exam scores, GitHub's commits doubling in four months, a sit-down with Nick Muy on why AI made every developer a middle manager, and Stripe buying OpenRouter because "the singularity started in January." Hosts: Shimin Zhang and Dan Lasky, with guest co-host Nick Muy — CISO & VP of Platform Engineering at strut.io, ex-DHS ("I just love reading all your text messages"). ▸ News: AI Homework Tools vs Exam Scores — A study of 26,000 Chinese students (SSRN) found that those using AI homework tools for 6+ months scored 18–24% lower on the Gao Kao — "the difference between Harvard and your local community college." Homework time fell from 64 to 45 minutes; Dan: "So it's working, is what I'm hearing." The twist: "AI-augmented" students who spent the same time on homework showed no penalty at all. ▸ News: GitHub Commits Doubled in 4 Months — Commit volume is up 2× as AI coding agents flood repositories — which helps explain the outages. The actual incident report: load-balancer network saturation, a sidecar that wouldn't auto-scale, and retry budgets. Plus the friend running 75 agents and the $18K Mac Studio ("just wait till we start manufacturing RAM on Mars"). ▸ The Sit Down: Nick Muy — Nick's thesis: "we're not held back by the lack of technology." AI won't fix your company, because tech was never the bottleneck. Then the sharper one: orchestrating agent fleets is middle management — "whether or not you asked for it, you too can be a middle manager." Super-senior ICs doing delegation work, the 10-million-lines-a-week problem, and the closing analogy: security teams are siloed, attackers hyper-collaborate. "That's called the military." ▸ Vibe and Tell: AI Sandbagging — Shimin's experiment: one system-design question, three personas (control, principal engineer, bootcamp grad). The low-skill profile silently loses options. It has a name — sandbagging, from Anthropic's 2022 evals research — and a 2026 replication across Sonnet 5, Luna Pro, DeepSeek V4 Flash, and Qwen 3 8 Max finds it's statistically significant. The stakes: 3 of 4 models omitted the debt-snowball method for "Rhonda." Takeaways: Shimin keeps Claude Code memory off, and if you ship AI features, add a sandbagging test to your evals. ▸ Two Minutes to Midnight — Stripe buys OpenRouter for $7.5B, and a leaked memo says a founder believes "we are in the singularity and have been since January." Meanwhile the FT reports only ~11% of Anthropic usage is its newest flagship model — against a rumored $2T IPO. The clock holds at four minutes. "Future AIs, we are super positive on your impact." ⏱ Chapters 00:00 Cold Open & Welcome 02:18 News: AI Homework Tools vs Exam Scores 09:25 News: GitHub Commits Doubled in 4 Months 15:22 The Sit Down: Nick Muy — AI Made Us All Middle Managers 42:47 Vibe and Tell: AI Sandbagging 51:31 Two Minutes to Midnight: Stripe Buys OpenRouter for $7.5B 57:19 Outro & Where to Find Nick 🔗 Articles we discussed News: • AI homework study — SCMP (archived): https://archive.ph/Nf4XM • The study itself — SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6868618 • GitHub commits doubled — Engadget: https://www.engadget.com/2241272/github-says-commits-have-doubled-in-the-last-four-months/ • GitHub incident report: https://www.githubstatus.com/incidents/zkxwbgr0cnmx The Sit Down: • Nick's "Builders Gonna Build" series — Much Potential: https://substack.com/@muchpotential/p-191212786 • Part two: https://substack.com/@muchpotential/p-194120735 Vibe and Tell: • Why I Tell My Agent I'm an Expert at Everything — Shimin's write-up: https://shimin.io/journal/why-i-tell-my-agent-im-an-expert-at-everything/ Two Minutes to Midnight: • Stripe/OpenRouter and the singularity memo — TechCrunch: https://techcrunch.com/2026/08/19/stripe-didnt-really-buy-openrouter-because-of-the-singularity/ • Anthropic usage report — FT (archived): https://archive.ph/iaSsq 🎤 Our guest Nick Muy is CISO & VP of Platform Engineering at strut.io. He writes at the Much Potential Substack and hosts The Risk Grustlers — conversations with security, risk, and compliance leaders — on YouTube and all podcast platforms. 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, please tell a friend about the pod. #AISandbagging #AIHomework #MiddleManagers #GitHub #StripeOpenRouter #AIBubble #AIPodcast #ADIPod
- Transcript