Energy-Based Models — the Sudoku tell: frontier LLMs only “solve” Sudoku by writing a Python script; tool-use ≠ reasoning; EBM + formal verifier as the architectural answer
Game Annotation Series — assembly as a stress test for LLM mechanical-modeling without rhetorical contagion; per-chapter LLM-interpretation logs
Transpilation as a Grounding Strategy — LLMs are weakly grounded in obscure formal languages (6502, bespoke VMs, COBOL); transpile to a grounded one rather than reason in them
Repairing LLM Code — The Two Oracles — LLMs faithfully read wrong structure and report it with high confidence; confidence and reader-consensus both fail as signals
Comments and the Distance to an Oracle — a stale comment is the one repo artifact that can be flatly wrong with nothing to notice — last run’s fluent output re-fed as evidence
The Contract Model vs. the Substrate Model — a per-delivery contract (delivery / constraints+guardrails / proof artifact / outside-verification / owner) is the vault’s architecture with the accumulation stripped out — contracts don’t compound, substrate does; plus the artifact-vs-reader split in verification independence, and tests written only to pass as oracle collapse
Oracles Are Objective Functions — greedy left-to-right decoding is hill-climbing with no restart, so self-sycophancy is a local minimum — and more samples or more agents is the same landscape, which is why headcount cannot substitute for altitude