Item detail
github.com

VeriRun — replayable code evaluation with durable result tracking

VeriRun — replayable code evaluation with durable result tracking is a developer tool that RepoRadar is tracking in its Evals & Benchmarks section, currently rated Silver tier with a 'worth watch' verdict. Its strongest signal is open-source/build quality, scored 8.4 out of 10.

Score7.5
Popularity100.0
Riskmedium
TierSilver
Score breakdown
Usefulness8.0
Novelty6.0
Momentum5.0
Maturity8.0
Open-source/build8.4
Evidence7.2
Workflow potential7.9
Setup ease4.2

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful when retrying a failed evaluation can otherwise duplicate results or change the test policy unnoticed. VeriRun records the plan and execution lineage, and publishes a local recovery exercise covering expired leases, worker takeover and duplicate commits. That is a concrete starting point for evaluation infrastructure, not proof of production readiness.

Who should use it

Engineers building code-generation evaluation pipelines Researchers needing replayable verifier results Infrastructure teams prototyping recoverable evaluation workers

Who should skip it

Hold off on VeriRun — replayable code evaluation with durable result tracking for mission-critical workflows without a containment strategy, explicit approvals, and a hands-on security review.

About this signal

VeriRun — replayable code evaluation with durable result tracking is tracked by RepoRadar as a developer tool in the Evals & Benchmarks section. First seen —; the source record was last checked on 2026-09-12. The current verdict is 'worth watch' with a Silver tier and advanced setup difficulty. The standout signals for VeriRun — replayable code evaluation with durable result tracking are open-source/build quality (8.4) and practical usefulness (8.0), while setup ease (4.2) trails — that balance shapes where it fits best. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The VeriRun — replayable code evaluation with durable result tracking record combines a 7.5/10 composite score with separate popularity (100.0), risk (medium), and setup (advanced) signals. See the scoring methodology for the current weights and evidence definitions.

Questions worth asking before you adopt this

Putting this into practice? Read How to read AI benchmarks without getting fooled for the checklist behind this score.

Risk explanation

Medium risk: executable evaluation runs candidate code. The local executor is an unsandboxed host subprocess for trusted fixtures only; using it for model-generated or other untrusted code exposes the host; The Docker development tier is not a security guarantee. Kubernetes/gVisor evidence is limited to the documented single-node environment, and the operator must provision and verify enforced egress restrictions and runtime isolation; The control plane modifies database state and stores artifacts. Multi-tenant authentication, result signing and artifact-retention policy are not implemented; do not expose this prototype as a shared production service.

Evidence links
Closest alternatives / related signals
evaluation code-generation replay provenance postgresql testing open-source
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-09-13T05:04:21.529456Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

History begins 2026-09-13; trends appear after a second dated snapshot. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar scoreTrend begins after a second recorded point.
Repository momentumTrend begins after a second recorded point.
GitHub stars (observed)Trend begins after a second recorded point.
GitHub stars327 exact observation
Versionv0.4.0
Last release2026-09-02T13:52:09Z
Maintenanceactive
Current riskmedium
Current verdictworth watch
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-137.57.5327mediumworth watchactive

Why the record changed

new entity

Newly added to RepoRadar's decision catalog.