Item detail
github.com

Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench)

RepoRadar surfaced Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) — a hosted app or demo — into the Evals & Benchmarks section, where it sits at Gold tier with a 'try now' verdict. Its strongest signal is novelty, scored 9.0 out of 10.

Score7.9
Popularity100.0
Risknone
TierGold
Score breakdown
Usefulness8.0
Novelty9.0
Momentum7.5
Maturity8.7
Open-source/build8.4
Evidence7.2
Workflow potential9.0
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for agent-eval teams that need stateful, reproducible long-horizon targets and for agent-runtime authors who want to measure work done, not conversation. The harness is Apache-2.0 and the 107-task suite is CC BY 4.0, so both are usable commercially with attribution; the rename from RealReplicaBench reflects an expanded scope.

Who should use it

agent-eval teams building reproducible long-horizon benchmarks agent-runtime authors who need a stable commerce workflow target researchers studying commerce agents and tool-use reliability

Who should skip it

Skip Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) if the source link, documentation, or setup requirements do not align with your current workflow or stack.

About this signal

Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) is tracked by RepoRadar as a hosted app or demo in the Evals & Benchmarks section. First seen 2026-08-09; the source record was last checked on 2026-08-09. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) leads on novelty (9.0) and workflow potential (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The Commerce Agent Bench — 107 stateful, verifier-graded tasks for long-horizon agents (formerly RealReplicaBench) record combines a 7.9/10 composite score with separate popularity (100.0), risk (none), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.

Questions worth asking before you adopt this

Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.

Risk explanation

No inherent user-impacting risk found: the benchmark runs agent actions inside pinned Docker containers and does not touch production accounts; Evaluated models run with shell access to their own container by design — scope Docker resources deliberately, and treat benchmark images like any third-party container (the quickstart pins one by digest); You supply a model API key plus a separate LLM-judge key; the batch runner redacts credentials from run.yaml, but key handling is still yours; The leaderboard is an audited aggregate: raw per-task result bundles are not yet published, so treat published scores as claims to reproduce rather than verified statistics.

Evidence links
Closest alternatives / related signals
benchmark long-horizon agent-eval open-source e-commerce stateful docker cc-by-4-0
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-09-13T05:04:21.529456Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

2 dated snapshots retained from 2026-09-12 through 2026-09-13; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score7.9 current · +0.0 net
Repository momentum7.5 current · +0.0 net
GitHub stars (observed)1,249 current · +0 net
GitHub stars1,249 exact observation
VersionNot reported by source
Last releaseNot reported by source
Maintenanceactive
Current risknone
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-137.97.51,249nonetry nowactive
2026-09-127.97.51,249nonetry nowactive

Why the record changed

stars changed

Stars changed: 1248 → 1249.