Score breakdown
Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.
Why it matters
Useful for agent-eval teams that need stateful, reproducible long-horizon targets and for agent-runtime authors who want to measure work done, not conversation. The harness is Apache-2.0 and the 107-task suite is CC BY 4.0, so both are usable commercially with attribution; the rename from RealReplicaBench reflects an expanded scope.
Who should use it
Who should skip it
Skip Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) if the source link, documentation, or setup requirements do not align with your current workflow or stack.
About this signal
Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) is tracked by RepoRadar as a hosted app or demo in the Evals & Benchmarks section. First seen 2026-08-09; the source record was last checked on 2026-08-09. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) leads on novelty (9.0) and workflow potential (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.
How this item is evaluated
The Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) record combines a 7.9/10 composite score with separate popularity (100.0), risk (none), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.
Questions worth asking before you adopt this
Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.
Risk explanation
No inherent user-impacting risk found: the benchmark runs agent actions inside pinned Docker containers and does not touch production accounts; Evaluated models run with shell access to their own container by design — scope Docker resources deliberately, and treat benchmark images like any third-party container (the quickstart pins one by digest); You supply a model API key plus a separate LLM-judge key; the batch runner redacts credentials from run.yaml, but key handling is still yours; The leaderboard is an audited aggregate: raw per-task result bundles are not yet published, so treat published scores as claims to reproduce rather than verified statistics.