Item detail
github.com

darkrishabh/agent-skills-eval

RepoRadar surfaced darkrishabh/agent-skills-eval — an AI project — into the Radar section, where it sits at Gold tier with a 'try now' verdict. Its strongest signal is workflow potential, scored 9.9 out of 10.

Score8.4
Popularity100.0
Risknone
TierGold
Score breakdown
Usefulness9.0
Novelty9.0
Momentum8.0
Maturity9.1
Open-source/build8.4
Evidence8.0
Workflow potential9.9
Setup ease8.8

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for Agent Skills authors who want receipts — agent-skills-eval runs the same prompt twice (with_skill vs without_skill), has a judge model grade both, and produces a side-by-side HTML report so the skill author can prove the SKILL.md actually improves the model's performance rather than just adding noise. Useful for **Claude Code / Codex / OpenClaw / Hermes Agent skill library

Where this stands now

darkrishabh/agent-skills-eval ranks #201 of 2143 tracked Radar items by composite score (8.4 against a section median of 6.7). The section currently carries 1038 Bronze, 646 Gold, 459 Silver. RepoRadar has retained observations for this record since 2026-06-25 (91 days in the current window). Signal extremes versus the section: momentum at the 78th percentile; novelty at the 92th percentile.

Who should use it

Agent Skills authors who want receipts — agent-skills-eval runs the same prompt twice (with_skill vs without_skill), has a judge model grade both, produces a side-by-side HTML report so the skill author can prove the SKILL.md actually improves the model's performance Claude Code / Codex / OpenClaw / Hermes Agent skill library maintainers — the test runner is intentionally runtime-agnostic so a single suite covers all four consumers, with the same workspace layout and the same judge grading AI-tool teams shipping internal skills — the --baseline flag is the audit trail that says 'here is the model's performance before the skill and here it is after, the receipts are in report/index.html' Skill-library curators — a skill library that ships with agent-skills-eval tests attached can answer 'which skills are net-positive, which are net-neutral, which are net-negative' across the whole library with one CLI run Researchers studying skill efficacy — the paired-prompt baseline-subtraction design removes model variance as a confound and surfaces the marginal contribution of the SKILL.md alone CI gates — the GitHub Actions workflow wires the test runner into PR checks so a new SKILL.md cannot land unless the with_skill version beats the without_skill baseline by a configurable margin Judge-model experimentation — a team can run the same skills with different judges (gpt-4o-mini, claude-sonnet-4-6, o3-mini) to see how judge selection affects the verdict Evaluation: npm install -g agent-skills-eval then npx agent-skills-eval ./skills --target gpt-4o-mini --judge gpt-4o-mini --baseline --strict; the docs at darkrishabh.github.io/agent-skills-eval/ walk through the workspace layout

Who should skip it

Skip darkrishabh/agent-skills-eval unless the captured evidence suggests it solves a problem you are actively working on.

About this signal

darkrishabh/agent-skills-eval is tracked by RepoRadar as an AI project in the Radar section. First seen 2026-06-25; the source record was last checked on 2026-06-25. The current verdict is 'try now' with a Gold tier and easy setup difficulty. The standout signals for darkrishabh/agent-skills-eval are workflow potential (9.9) and maturity (9.1), while evidence quality (8.0) trails — that balance shapes where it fits best. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The darkrishabh/agent-skills-eval record combines a 8.4/10 composite score with separate popularity (100.0), risk (none), and setup (easy) signals. See the scoring methodology for the current weights and evidence definitions.

Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.

Risk explanation

No inherent user-impacting risk is flagged from the captured evidence.

Evidence links
Closest alternatives / related signals
agent-skills-eval darkrishabh anthropic-agent-skills agent-skills agentskills-io skill-eval skill-testing skill-receipts
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-09-24T21:49:16.463298Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

70 dated snapshots retained from 2026-06-25 through 2026-09-24; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.4 current · +0.0 net
Repository momentum9.3 current · +1.3 net
GitHub stars (observed)794 current · +174 net
GitHub stars794 exact observation
VersionNot reported by source
Last releaseNot reported by source
Maintenanceactive
Current risknone
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-248.49.3794nonetry nowactive
2026-09-238.49.0792nonetry nowactive
2026-09-228.45.5792nonetry nowmaintained
2026-09-218.45.5791nonetry nowmaintained
2026-09-208.45.5790nonetry nowmaintained
2026-09-198.45.5790nonetry nowmaintained
2026-09-178.45.5790nonetry nowmaintained
2026-09-158.46.1785nonetry nowmaintained
2026-09-138.45.5732nonetry nowmaintained
2026-09-128.45.5731nonetry nowmaintained
2026-09-118.45.8731nonetry nowmaintained
2026-09-088.45.5727nonetry nowmaintained

Why the record changed

stars changed

Stars changed: 792 → 794.

maintenance changed

Maintenance changed: maintained → active.

stars changed

Stars changed: 791 → 792.

stars changed

Stars changed: 790 → 791.

stars changed

Stars changed: 791 → 790.

stars changed

Stars changed: 790 → 791.

stars changed

Stars changed: 785 → 790.

stars changed

Stars changed: 732 → 785.

stars changed

Stars changed: 731 → 732.

stars changed

Stars changed: 727 → 731.

stars changed

Stars changed: 713 → 719.

stars changed

Stars changed: 703 → 713.