- Embedded Dev How Many AI Code Reviews Is Enough?
400 LLM calls on one codebase. A single review pass finds 34% of the real defects, and the four most-repeated findings were all false positives.
- Embedded Dev AI-Written Firmware: 233 Tests, Sonnet 68% vs Haiku 57%
Sonnet 4.6 hits 68.0% and Haiku 4.5 56.9% on 233 embedded firmware cases. EmbedEval is open source: 5-layer eval, n=3, Wilson 95% CI, Zephyr and STM32 HAL.