{"id":65528,"date":"2026-08-01T12:09:19","date_gmt":"2026-08-01T16:09:19","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=65528"},"modified":"2026-08-01T12:09:19","modified_gmt":"2026-08-01T16:09:19","slug":"supabase-evals-benchmark-ai-coding-agents","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/supabase-evals-benchmark-ai-coding-agents\/","title":{"rendered":"Supabase Releases Open Source Benchmark for AI Coding Agents"},"content":{"rendered":"<p>Supabase has open-sourced <a href=\"https:\/\/supabase.com\/evals\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Supabase Evals<\/a>, a benchmark and testing framework designed to measure how effectively AI <a href=\"https:\/\/overcentral.com\/en\/ai-coding-agents-trigger-security-rules\/\" title=\"AI Coding Agents Trigger Endpoint Security Rules Meant for Attackers\" data-iacss-internal=\"1\">coding agents<\/a> can build with the Supabase platform. Released under the Apache-2.0 license, the framework runs agents like Claude Code, Codex, and OpenCode against a suite of real-world tasks \u2014 such as constructing a database schema, debugging a failed Edge Function, or repairing a broken Row-Level Security (RLS) policy \u2014 and scores each result. The project now powers a public leaderboard at supabase.com\/evals and an internal regression suite that refreshes daily, giving developers and enterprises a transparent, reproducible method for evaluating agent performance on actual infrastructure.<\/p>\n<h2>The Problem Supabase Evals Solves in AI-Driven Development<\/h2>\n<p>As AI coding agents become more capable, the gap between synthetic benchmarks and production reality has widened. Many existing evaluations test agents on isolated code-writing tasks or static datasets, but these often fail to capture the complexity of working with a live backend platform. A coding agent that can generate a correct SQL query in a sandbox may still struggle when asked to debug a misconfigured authentication flow or deploy a function to a real runtime environment. Supabase Evals addresses this by grounding every scenario in actual support tickets, bug reports, and GitHub issues. The framework evaluates not just whether an agent can write code, but whether it can navigate the full lifecycle of building, deploying, investigating, and resolving issues within a live Supabase stack.<\/p>\n<h2>What Is Supabase Evals? An Open-Source Benchmark for AI Coding Agents<\/h2>\n<p>Supabase Evals is a publicly available benchmark and evaluation framework that tests AI coding agents against real tasks on the Supabase platform. It is hosted at github.com\/supabase\/evals and runs locally via the pnpm package manager. The framework defines three core dimensions \u2014 products, topics, and stages \u2014 and selects the smallest set of scenarios that touches each dimension at least once. Products include database, auth, storage, edge-functions, realtime, cron, queues, vectors, and data-api. Topics cover RLS, security, migrations, SQL, SDK, observability, self-hosting, tests, and declarative-schema. Stages are build, deploy, investigate, and resolve. Each scenario runs against a real containerized environment, not a mock, ensuring that agents interact with the actual Supabase Management API, MCP server, and CLI.<\/p>\n<h3>How the Scoring System Works<\/h3>\n<p>Scoring combines deterministic checks with an LLM-as-a-judge approach. Deterministic checks verify concrete outcomes \u2014 for example, whether a user can access certain data after an RLS policy is applied, or whether an Edge Function returns the expected response. An LLM judge handles more semantic evaluations where the exact output is less rigidly defined. Agents are allowed one retry before grading, which reduces false negatives while keeping the evaluation process sustainable. Each eval directory contains a PROMPT.md file with the task and frontmatter, an EVAL.ts file that exports the scorer, and optional remote\/ and local\/ directories for starting states. When a local\/ directory is present or the eval declares interface: cli, the framework boots a Docker sandbox with the real Supabase CLI installed.<\/p>\n<h2>Is Supabase Evals Deployable Today? Yes, and Here Is How It Works<\/h2>\n<p>Supabase Evals is deployable immediately. The repository is public under the Apache-2.0 license and runs locally using pnpm. The harness boots two real environments per scenario: a hosted-like Supabase stack and a local CLI project running in containers. Agents interact with an actual MCP server and CLI rather than simulated APIs. A platform-lite runtime, backed by @supabase\/lite, exposes a Management API-compatible surface. For local-stack runs, a Docker daemon must be running, and ports 54321 through 54329 must be free. The framework is designed for industries such as developer tooling, cloud infrastructure, data platforms, and regulated backends in fintech or healthcare, where an agent that writes a wrong RLS policy could constitute a security incident. Applications include regression-testing documentation and skill edits, gating SDK releases, and comparing agent harnesses head to head.<\/p>\n<h2>Key Findings: How Top AI Coding Agents Performed on the Benchmark<\/h2>\n<p>The initial benchmark results reveal that most agents pass scenarios even without any Supabase-specific skills loaded. In the Build stage, both Opus 5 and <a href=\"https:\/\/overcentral.com\/en\/moonshot-ai-kimi-k3\/\" title=\"Moonshot AI launches Kimi K3, largest open-source model ever\" data-iacss-internal=\"1\">Kimi K3<\/a> scored 100 percent unaided. Skills \u2014 which are sets of instructions bundled into the agent&#8217;s system prompt \u2014 closed the remaining gap for other models. Sonnet 5 rose from 78 percent to 100 percent, <a href=\"https:\/\/overcentral.com\/en\/gpt-5-6-sol-reasoning-levels\/\" title=\"GPT-5.6 Sol Maps Five Reasoning Levels to Task Complexity\" data-iacss-internal=\"1\">GPT-5.6 Sol<\/a> improved from 89 percent to 100 percent, and GPT-5.4 mini increased from 78 percent to 89 percent. These figures demonstrate that while top-tier models can operate effectively without platform-specific guidance, skills significantly level the playing field for smaller or less specialized agents.<\/p>\n<h3>Three Weaknesses Emerged During Testing<\/h3>\n<p>The evaluation surfaced three consistent weaknesses across agents. First, agents hand-wrote SQL migrations instead of using declarative schemas, a pattern that prompted a skill guidance update. Second, agents verified authentication status by making manual checks rather than using the @supabase\/server package, leading to the creation of a package selection guide. Third, documentation usage varied sharply between models. Codex and GPT-5.6 read approximately eight docs pages per scenario, while Claude Code read only about two and checked the documentation in fewer than 40 percent of scenarios even when skills were loaded. This suggests that some agents rely more heavily on internal knowledge while others are more diligent about retrieving external reference information.<\/p>\n<h2>How the Benchmark and Regression Suites Differ<\/h2>\n<p>Supabase Evals divides its scenarios into two distinct suites. The Benchmark suite covers breadth and is published publicly on the leaderboard. These scenarios are used when assessing new model changes, skill updates, or harness comparisons. The Regression suite covers known failure modes and refreshes daily, but its results do not affect published leaderboard scores. This separation allows the team to track regressions in real time without contaminating the public benchmark. It also enables continuous testing of fixes and workarounds for issues that have been identified in previous evaluations.<\/p>\n<h2>Why Real Environments Matter for Agent Evaluation<\/h2>\n<p>One of the most distinguishing features of Supabase Evals is its commitment to real environments. The framework does not use mocks, stubs, or simulated APIs. Instead, it boots actual Supabase stacks in containers and provides agents with genuine Management API endpoints, MCP servers, and CLI tools. This approach ensures that the evaluation reflects how agents would actually behave in production-like conditions. A mock might correctly simulate a successful RLS policy application, but a real environment can detect subtle failures \u2014 such as a policy that passes syntax checks but allows unintended access. The platform-lite runtime, backed by @supabase\/lite, manages the lifecycle of these environments efficiently, booting only the services that each scenario requires.<\/p>\n<h2>Strategic Implications for Developer Platforms and AI Tooling<\/h2>\n<p>The release of Supabase Evals signals a broader shift in how developer platforms are approaching AI integration. Rather than relying on third-party evaluations or ad-hoc testing, companies like Supabase are building their own standardized benchmarks that reflect the actual workflows and failure modes of their platforms. This trend has implications for platform vendors, AI model providers, and enterprise users alike. For platform vendors, owning the evaluation loop means they can detect regressions immediately when models update. For model providers, these benchmarks offer clear, actionable data on where their agents fall short. For enterprise users, particularly those in regulated industries, a transparent, open-source benchmark provides a defensible basis for deciding which agents to trust with production systems.<\/p>\n<h3>What This Means for Fintech and Healthcare Backends<\/h3>\n<p>Industries where a miswritten security policy constitutes a compliance incident stand to benefit most from rigorous agent evaluation. In fintech, an AI agent that incorrectly configures RLS could expose sensitive customer financial data. In healthcare, a poorly written Edge Function could violate HIPAA data-handling requirements. Supabase Evals provides a framework for testing agents against the exact patterns and policies that these industries depend on. By running scenarios that include real bug reports and support tickets, the benchmark ensures that the evaluation is not theoretical but grounded in problems that have actually occurred.<\/p>\n<h2>The Role of Skills in Shaping Agent Behavior<\/h2>\n<p>Skills are a central mechanism in Supabase Evals. They consist of instructions bundled into the agent&#8217;s system prompt that guide the agent toward preferred approaches \u2014 such as using declarative schemas instead of hand-written migrations, or using @supabase\/server for authentication verification. The benchmark data shows that skills matter most for smaller models and least for top-tier models. Rewriting the Postgres best-practices skill description, for example, lifted its activation rate from about one in ten sessions to 60 percent. This indicates that the phrasing and specificity of skill descriptions directly impact whether agents follow them. The framework tracks skill activation rates as part of its evaluation, providing feedback on which instructions are being ignored and which are being successfully adopted.<\/p>\n<h2>Technical Architecture: How the Harness Orchestrates Each Eval<\/h2>\n<p>Behind each evaluation is a well-defined technical pipeline. The harness begins by selecting a scenario based on its frontmatter metadata, which includes tags for stage, suite, product, topic, and motivation. The harness then boots the required environments. For tools evals \u2014 which exercise only the MCP surface \u2014 the agent receives the task through the PROMPT.md file and works entirely through tool calls without access to a filesystem. For local-stack evals, the harness boots a Docker sandbox with the real Supabase CLI installed, and the agent works through a full local development workflow. The harness copies the workspace back to the host after the agent completes its work so that scorers can run their checks. Each scenario gets one retry before the scorer runs. The entire pipeline is designed to be reproducible and scriptable, enabling CI\/CD integration.<\/p>\n<h2>Comparing Agent Behavior: Codex vs. Claude Code vs. OpenCode<\/h2>\n<p>The benchmark reveals meaningful differences in how agents approach the same tasks. Codex and GPT-5.6 tend to be heavy readers of documentation, consuming roughly eight pages per scenario. Claude Code is more concise, reading about two pages and checking docs in fewer than 40 percent of scenarios even when skills are loaded. OpenCode, as a fully open-source alternative, offers a different trade-off: it is more transparent in its decision-making but may lack some of the optimization that proprietary agents have undergone. These behavioral differences have practical consequences. An agent that reads more documentation may be slower but more accurate in unfamiliar domains. An agent that relies on internal knowledge may be faster but more prone to following outdated or incorrect patterns. The benchmark does not declare a single winner; instead, it provides data that teams can use to choose the agent best suited to their specific workflows and risk tolerances.<\/p>\n<h2>Future Outlook: How Supabase Evals Could Reshape the Agent Ecosystem<\/h2>\n<p>By open-sourcing its evaluation framework, Supabase has created a template that other platforms could adopt. A world in which every major developer platform publishes its own benchmark for AI coding agents would give model providers clear, platform-specific targets for improvement and give users a transparent basis for comparison. The framework&#8217;s design \u2014 grounded in real issues, running against real environments, combining deterministic checks with LLM judgment \u2014 could become a standard pattern for agent evaluation. As the number of AI coding agents and the platforms they target continues to grow, the need for rigorous, transparent, and reproducible benchmarks will only increase. Supabase Evals represents a practical step toward that future, offering a blueprint that others can build on, adapt, and improve.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Supabase has open-sourced Supabase Evals, a benchmark and testing framework designed to measure how effectively AI coding agents can build with the Supabase platform. Released under the Apache-2.0 license, the framework runs agents like Claude Code, Codex, and OpenCode against a suite of real-world tasks \u2014 such as constructing a database schema, debugging a failed [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83630,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/65528.png","fifu_image_alt":"Supabase Releases Open Source Benchmark for AI Coding Agents","footnotes":""},"categories":[349],"tags":[],"class_list":["post-65528","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/65528.png","fifu_image_alt":"Supabase Releases Open Source Benchmark for AI Coding Agents","fifu_redirection_url":"https:\/\/z.ai\/blog\/glm-4.6","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/65528","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=65528"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/65528\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83630"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=65528"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=65528"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=65528"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}