AI Wrote 13 Million Lines. Who Checked Them?

AI Lite makes AI feel less intimidating. Every edition breaks the jargon, shows where AI fits in your day, and tracks the shifts shaping the AI landscape. No tech background needed.

AI Lite
AI Lite · September 6, 2026 · ~5 min read
🕓 ~5 min read · Weekly drop
TLDR: Claude produced 13 million lines of Lean for a formalization of Fermat's Last Theorem. Humans couldn't inspect it line by line. A separate checker could.
🧠 Learn: Why a checker beats a second opinion
⚡ Pulse: GPT-6 Astra · NVIDIA to acquire Hugging Face · India's teacher training
🚀 Career: Ship the work and the receipt

✍️ From the Author's Desk

Comic: a smiling stick figure presents a dangerously tall stack of pages while a separate machine calmly stamps a check mark on one page. Caption: The output was fast. The check was separate.

The future of AI work may be less “trust me” and more “show me the check.”

This week's headline sounds like it's about speed: 13 million lines in 11 days. What hooked me was the quieter detail. Nobody had to read those lines one by one before the result could count.

🔎 One term to know: formal verification
Formal verification uses a rule-bound system to prove that an output follows exact requirements. Instead of asking whether work looks convincing, you produce checkable evidence. Think of a calculator redoing your totals, tests checking your code, or citations supporting each claim.

AI Learn

🧠 The Answer Is Not the Proof

On September 4, Anthropic shared what it calls the first complete computer-checked formalization of Fermat's Last Theorem. Claude worked largely autonomously for 11 days, writing 13 million lines in Lean, a language built so computers can check mathematical logic.

Claude didn't discover Fermat's Last Theorem or replace Andrew Wiles's human proof. It translated a version of the existing argument into a form a machine could inspect without skipping the “obvious” steps humans routinely leave out.

The scale is the part that changes the workflow. Claude generated proofs for 30,300 intermediate theorems and used 29,500 of them in the final result. The code was more than five times the size of Mathlib, Lean's main community library. Anthropic says the project consumed about six billion output tokens.

No review team was going to read that line by line. So the question changed: what could say no?

Pause and choose

Claude finishes a proof. Which check can genuinely prove it wrong?

A. Ask Claude to review its own work.
B. Run every step through Lean's fixed rules.

Answer: B. The checker can reject a polished answer for breaking one rule. Asking the generator to grade itself may catch mistakes, but it's still another opinion from the same system.

The pattern travels well beyond mathematics. Don't ask the same AI, “Are you sure?” Hand the output to a check with a different failure mode.

  • Reconcile the spreadsheet's totals a second way.
  • Put a source next to every factual claim in the memo.
  • Ship the code change with a test that fails before it and passes after.

None of this removes human judgment. Anthropic explicitly says a formal proof shouldn't replace an explanation people can understand. The point is to spend human attention on assumptions, meaning, and consequences while the machine replays the millions of tiny steps.

🎥 Watch: Big Think's Terence Tao examines the paradox at the heart of AI and science. Faster discovery makes machine verification matter more, and raises the premium on human understanding at the same time. Published September 3, 2026.

Watch Terence Tao discuss the paradox at the heart of AI and science

Watch on YouTube →

Read Anthropic's report →
🎯 Try this week: Pick one recurring AI task. Before improving the prompt, write the receipt. Which test, source trail, reconciliation, or approval rule would let someone else trust the result? Then generate the work and run the check separately.

AI Pulse
Capability watch

OpenAI Calls Astra “Critical” for Cyber

OpenAI said on September 1 that GPT-6 Astra can find previously unknown security flaws and develop exploits across protected systems without a person guiding each step. It is the first model OpenAI has placed at the Critical cyber threshold. In one internal evaluation, Astra found two zero-day vulnerabilities that OpenAI says it is disclosing to maintainers.

The signal here is permission design, not panic. Long-running agents need narrow access, independent monitoring, and an action that automatically stops when the monitor objects.

🎥 AI Explained examines Astra's capabilities and the harder problem of monitoring its reasoning. Published September 4, 2026.

Watch: GPT-6 Astra, so good even OpenAI are worried

Watch on YouTube →

Read OpenAI's report →
Platform shift

NVIDIA Moves to Buy Hugging Face for $12.93 Billion

Editorial illustration of a green GPU connecting to an open network of AI model tiles

NVIDIA announced the deal on September 3. Hugging Face hosts more than three million models and serves over 18 million developers, according to NVIDIA. The company says the hub will stay open, multi-cloud, and multi-accelerator, with no NVIDIA hardware requirement.

That promise is now a business-critical interface. Keep one representative model and workload portable enough to run somewhere else, and rehearse that move at least once. Portability you haven't tested doesn't count.

Read NVIDIA's announcement →
Adoption signal

India Trains the Role, Not Just the User

On September 4, Google announced free nationwide AI training sessions for Indian educators. The sessions cover lesson design, feedback, assessment, and administrative tasks. Google says the sessions follow India's education policy and draw on learning science.

Most rollouts stop at the demo, which creates curiosity and not much else. Role-specific practice is what changes behaviour. Before buying more seats, teach one team to use AI on one task it already owns.

Read Google's India update →

AI Career

🚀 Ship the Work and the Receipt

The most persuasive AI portfolio proves you knew when the output deserved trust.

Build one small artifact this week. Use AI to draft a customer analysis, automate a spreadsheet, or change a piece of code. Then show what the check changed.

For example: you use AI to group 600 support tickets into themes. A row-count reconciliation exposes 17 missing tickets. You fix the export, rerun the analysis, and explain why the ranking changed. That short case shows the deliverable, the receipt, and your judgment without naming a framework.

It also gives you a sharper interview answer than “I use AI every day.”

“I use AI to move faster, but I don't ask it to grade itself. For consequential work, I pair the output with a source trail, reconciliation, test, or approval rule. Then I use my judgment on what the check can't decide.”

🎥 Going deeper: Silicon Valley Girl distills current hiring signals from leaders at LinkedIn, HubSpot, Khan Academy, Box, and Gamma into five moves candidates can start now. Published September 4, 2026.

Watch: Top career strategies for getting hired in the AI era

Watch on YouTube →

This week, don't scale back what you ask AI to do. Make its output easier to prove instead.

Next week: what changes if the open-model commons gets a $12.93 billion owner.

-Kay

➡️ Previous Volume

📚 Catch up on every edition → Archive

💛 If this helped, share it with someone learning AI. 💛