Item detail
github.com

Show Harness: a VLM agent that plays robots

Show Harness: a VLM agent that plays robots is a code repository in RepoRadar's Radar section, holding Silver tier and a 'watch' verdict. Its strongest signal is novelty, scored 8.5 out of 10.

Score7.0
Popularity100.0
Risknone
TierSilver
Score breakdown
Usefulness7.0
Novelty8.5
Momentum7.5
Maturity7.2
Open-source/build8.4
Evidence7.2
Workflow potential7.4
Setup ease4.2

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

It tests whether a general VLM, wired into a control loop, can substitute for task-specific robot policies. That question decides how much robotics work becomes a prompting problem.

Where this stands now

Show Harness: a VLM agent that plays robots ranks #1035 of 1451 tracked Radar items by composite score (7.0 against a section median of 7.8). The section currently carries 643 Gold, 443 Silver, 365 Bronze. RepoRadar has retained observations for this record since 2026-09-12 (1 days in the current window). Signal extremes versus the section: novelty at the 83th percentile.

Who should use it

robot-learning researchers builders following VLM-as-controller work

Who should skip it

Pass on Show Harness: a VLM agent that plays robots if you need something non-technical and turnkey rather than a tool that requires comfort with CLI, dependencies, or system configuration.

About this signal

Show Harness: a VLM agent that plays robots is tracked by RepoRadar as a code repository in the Radar section. First seen 2026-09-12; the source record was last checked on 2026-09-12. The current verdict is 'watch' with a Silver tier and Hard setup difficulty. Show Harness: a VLM agent that plays robots leads on novelty (8.5) and open-source/build quality (8.4); its lowest signal is setup ease (4.2), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The Show Harness: a VLM agent that plays robots record combines a 7.0/10 composite score with separate popularity (100.0), risk (none), and setup (Hard) signals. See the scoring methodology for the current weights and evidence definitions.

Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.

Risk explanation

No inherent user-impacting risk: research harness and evaluation code.

Evidence links
Closest alternatives / related signals
robotics vlm agents research
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-09-13T05:04:21.529456Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

2 dated snapshots retained from 2026-09-12 through 2026-09-13; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score7.0 current · +0.0 net
Repository momentum9.0 current · +0.0 net
GitHub stars (observed)278 current · +5 net
GitHub stars278 exact observation
VersionNot reported by source
Last releaseNot reported by source
Maintenanceactive
Current risknone
Current verdictwatch
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-137.09.0278nonewatchactive
2026-09-127.09.0273nonewatchactive

Why the record changed

stars changed

Stars changed: 273 → 278.

new entity

Newly added to RepoRadar's decision catalog.