Item detail
github.com

VibeBench/VibeSearchBench

VibeBench/VibeSearchBench is an AI project in RepoRadar's Evaluation section, holding Gold tier and a 'try now' verdict. Its strongest signal is workflow potential, scored 9.2 out of 10.

Score8.1
Popularity1.0
Riskconditional
TierGold
Score breakdown
Usefulness8.0
Novelty9.0
Momentum7.0
Maturity6.4
Open-source/build8.4
Evidence7.2
Workflow potential9.2
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for agent researchers and search-quality teams that want to measure agents on vague, progressive-disclosure queries instead of single-shot fact lookups.

Who should use it

Agent researchers measuring long-horizon, multi-turn search quality Search and recommendation teams evaluating RAG agents on real-world query patterns Benchmark authors looking for a schema-free evaluation surface that resists overfitting Builder-evaluators comparing Claude, GPT, and open models on progressive-disclosure workflows

Who should skip it

Skip VibeBench/VibeSearchBench if the source repository or demo is inactive, unmaintained, or no longer matches the description shown here.

About this signal

VibeBench/VibeSearchBench is tracked by RepoRadar as an AI project in the Evaluation section. It was first seen on 2026-06-22 and last updated on 2026-06-22. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. VibeBench/VibeSearchBench leads on workflow potential (9.2) and novelty (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed. The score, tier, risk label, and verdict on this page are never influenced by sponsorship, ads, or tips — they reflect only the usefulness, popularity, novelty, momentum, maturity, and evidence signals described in the RepoRadar methodology.

How this item is evaluated

RepoRadar assigned VibeBench/VibeSearchBench a composite score of 8.1 out of 10, placing it in the Gold tier. This score combines weighted sub-signals: usefulness (35%), novelty (18%), momentum (14%), maturity (10%), open-source/build quality (7%), evidence quality (6%), workflow potential (6%), and setup ease (4%). Popularity is tracked separately at 1.0 and never affects the composite score or tier. The risk label of 'conditional' reflects inherent user-impacting hazards, not generic novelty. Items with no risk flag may still require normal code review before production use.

Putting this into practice? Read How to read AI benchmarks without getting fooled for the checklist behind this score.

Risk explanation

The README references a fictional 'OpenClaw' agent wrapper in the run-scripts section; only the OpenHands and WebArena-style agent paths are well-known, so treat OpenClaw results as illustrative not authoritative; The README's leaderboard line still quotes 'Claude Opus 4.6' as the best score, which is now outdated; Anthropic's current top Opus-tier model is Opus 4.8 — the benchmark itself is sound, but verify against the live leaderboard for up-to-date model attribution.

Evidence links
Closest alternatives / related signals
benchmark search rag multi-turn agent evaluation mit
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-08-04T02:20:22Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

35 dated snapshots retained from 2026-06-21 through 2026-08-04; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.1 current · +0.7 net
Repository momentum5.5 current · -1.5 net
GitHub stars (observed)409 current · -545 net
GitHub stars409 exact observation
VersionNot reported by source
Last releaseNot reported by source
Maintenancemaintained
Current riskconditional
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-08-048.15.5409conditionaltry nowmaintained
2026-08-038.15.5409conditionaltry nowmaintained
2026-08-028.15.8409conditionaltry nowmaintained
2026-08-018.15.5404conditionaltry nowmaintained
2026-07-318.17.0Not recordedconditionaltry nownot recorded
2026-07-308.17.0Not recordedconditionaltry nownot recorded
2026-07-298.15.0405conditionaltry nowmaintained
2026-07-288.15.5410conditionaltry nowmaintained
2026-07-218.15.5954conditionaltry nowmaintained
2026-07-208.15.5954conditionaltry nowmaintained
2026-07-198.15.5954conditionaltry nowmaintained
2026-07-188.15.5954conditionaltry nowmaintained

Why the record changed

stars changed

Stars changed: 404 → 409.

stars changed

Source-observed stars changed: 403 → 404. This reports the retained observation delta and does not infer why the upstream change occurred.

stars changed

Stars changed: 410 → 405.

stars changed

Stars changed: 954 → 410.

score changed

Reconstructed from adjacent retained daily snapshots; no upstream cause is inferred. Score changed: 7.4 → 8.1. Largest component movements: usefulness 7 → 8 (+1.0); commercial workflow potential 8.2 → 9.2 (+1.0); maturity 5.6 → 6.4 (+0.8). These retained scoring-input deltas account for the RepoRadar score movement; this scoring-accounting explanation does not infer an upstream cause.

popularity changed

Reconstructed from adjacent retained daily snapshots; no upstream cause is inferred. Popularity changed: 7.0 → 1.0.

verdict changed

Reconstructed from adjacent retained daily snapshots; no upstream cause is inferred. Verdict changed: worth_watch → try_now.

risk changed

Reconstructed from adjacent retained daily snapshots; no upstream cause is inferred. Risk changed: low → conditional.

tier changed

Reconstructed from adjacent retained daily snapshots; no upstream cause is inferred. Tier changed: Silver → Gold.