<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Model Benchmarks on No Semicolons</title><link>https://nosemicolons.com/tags/model-benchmarks/</link><description>Recent content in Model Benchmarks on No Semicolons</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 20 Sep 2026 12:46:36 +0000</lastBuildDate><atom:link href="https://nosemicolons.com/tags/model-benchmarks/index.xml" rel="self" type="application/rss+xml"/><item><title>The AI Code Generation Model Leaderboard: Real Performance Data From 50,000 Generated Functions</title><link>https://nosemicolons.com/posts/ai-code-generation-model-leaderboard-performance-data/</link><pubDate>Sun, 20 Sep 2026 12:46:36 +0000</pubDate><guid>https://nosemicolons.com/posts/ai-code-generation-model-leaderboard-performance-data/</guid><description>&lt;p>Ever wondered which AI model actually writes the best code? Not according to marketing claims or cherry-picked examples, but based on real production data?&lt;/p>
&lt;p>I spent the last three months running the largest independent benchmark of AI code generation models I&amp;rsquo;ve ever attempted. We generated over 50,000 functions across GPT-4, Claude 3.5 Sonnet, Gemini Pro, and several other popular models, then measured everything from correctness to maintainability to real-world performance.&lt;/p></description></item></channel></rss>