Score breakdown
Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.
Why it matters
Useful for Agent Skills authors who want receipts — agent-skills-eval runs the same prompt twice (with_skill vs without_skill), has a judge model grade both, and produces a side-by-side HTML report so the skill author can prove the SKILL.md actually improves the model's performance rather than just adding noise. Useful for **Claude Code / Codex / OpenClaw / Hermes Agent skill library
Where this stands now
darkrishabh/agent-skills-eval ranks #201 of 2143 tracked Radar items by composite score (8.4 against a section median of 6.7). The section currently carries 1038 Bronze, 646 Gold, 459 Silver. RepoRadar has retained observations for this record since 2026-06-25 (91 days in the current window). Signal extremes versus the section: momentum at the 78th percentile; novelty at the 92th percentile.
Who should use it
Who should skip it
Skip darkrishabh/agent-skills-eval unless the captured evidence suggests it solves a problem you are actively working on.
About this signal
darkrishabh/agent-skills-eval is tracked by RepoRadar as an AI project in the Radar section. First seen 2026-06-25; the source record was last checked on 2026-06-25. The current verdict is 'try now' with a Gold tier and easy setup difficulty. The standout signals for darkrishabh/agent-skills-eval are workflow potential (9.9) and maturity (9.1), while evidence quality (8.0) trails — that balance shapes where it fits best. This page summarizes the public evidence on the linked source page and states where additional review is still needed.
How this item is evaluated
The darkrishabh/agent-skills-eval record combines a 8.4/10 composite score with separate popularity (100.0), risk (none), and setup (easy) signals. See the scoring methodology for the current weights and evidence definitions.
Putting this into practice? Read How to vet an AI agent or MCP server before you wire it in for the checklist behind this score.
Risk explanation
No inherent user-impacting risk is flagged from the captured evidence.