<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Programming Benchmarks on No Semicolons</title><link>https://nosemicolons.com/tags/programming-benchmarks/</link><description>Recent content in Programming Benchmarks on No Semicolons</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 26 Sep 2026 12:35:26 +0000</lastBuildDate><atom:link href="https://nosemicolons.com/tags/programming-benchmarks/index.xml" rel="self" type="application/rss+xml"/><item><title>The AI Code Generation Model Tournament: I Tested 12 Models Building the Same E-commerce App — Here's Who Actually Ships</title><link>https://nosemicolons.com/posts/ai-code-generation-model-tournament-ecommerce-app-comparison/</link><pubDate>Sat, 26 Sep 2026 12:35:26 +0000</pubDate><guid>https://nosemicolons.com/posts/ai-code-generation-model-tournament-ecommerce-app-comparison/</guid><description>&lt;p>Ever wondered which AI coding model would actually get your app to production? I decided to find out by giving 12 different models the exact same challenge: build a complete e-commerce shopping cart with authentication, product management, and payment processing.&lt;/p>
&lt;p>The results were&amp;hellip; eye-opening. Some models that dominated the benchmarks completely fumbled real-world requirements, while a few dark horses surprised me with their pragmatic, ship-ready code.&lt;/p>
&lt;h2 id="the-battle-arena-setting-up-fair-tests">The Battle Arena: Setting Up Fair Tests&lt;/h2>
&lt;p>I kept the requirements identical across all models: a Node.js/Express backend with user authentication, product CRUD operations, shopping cart functionality, and Stripe integration. Each model got the same initial prompt, the same follow-up questions, and the same time budget (roughly 2 hours of back-and-forth).&lt;/p></description></item></channel></rss>