Item detail
github.com

promptfoo/promptfoo

promptfoo/promptfoo is a developer tool in RepoRadar's Radar section, holding Gold tier and a 'try now' verdict. Its strongest signal is workflow potential, scored 9.8 out of 10.

Score8.7
Popularity80.0
Riskconditional
TierGold
Score breakdown
Usefulness9.0
Novelty8.0
Momentum8.0
Maturity8.4
Open-source/build8.4
Evidence7.2
Workflow potential9.8
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for teams that need a practical way to test AI behavior repeatedly instead of relying on one-off manual prompt checks.

Where this stands now

promptfoo/promptfoo ranks #60 of 2304 tracked Radar items by composite score (8.7 against a section median of 5.2). The section currently carries 1197 Bronze, 645 Gold, 462 Silver. RepoRadar has retained observations for this record since 2026-06-19 (100 days in the current window). Signal extremes versus the section: momentum at the 79th percentile; novelty at the 79th percentile.

Who should use it

LLM app teams AI product engineers security-minded builders developers maintaining RAG or agent workflows

Who should skip it

Move on from promptfoo/promptfoo if the licensing terms, language support, or platform requirements do not fit your project.

About this signal

promptfoo/promptfoo is tracked by RepoRadar as a developer tool in the Radar section. First seen 2026-06-19; the source record was last checked on 2026-06-19. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. promptfoo/promptfoo leads on workflow potential (9.8) and practical usefulness (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The promptfoo/promptfoo record combines a 8.7/10 composite score with separate popularity (80.0), risk (conditional), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.

Questions worth asking before you adopt this

Putting this into practice? Read How to evaluate an AI tool before you adopt it for the checklist behind this score.

Risk explanation

Evaluation runs can forward prompts, datasets, and model outputs to external providers, so scrub sensitive test data before using hosted model backends.

Evidence links
Closest alternatives / related signals
llm-evals red-teaming prompt-testing rag developer-tools
Verification record

What RepoRadar actually verified

Tested in a bounded workflow Source-only update since test

Bounded representative workflow retained by RepoRadar verification harness. Last checked 2026-07-13T10:35:12.343566Z.

A source release dated 2026-09-18 is newer than the retained test dated 2026-07-13; the previous test may no longer represent the current release.

Read as a sequence: 1 retained check(s) on 2026-07-13 with outcomes of 1 passed; evidence scopes covered bounded representative workflows. Recorded setup time totals 1 minute(s).

passed · cohort-20260712-promptfoo-isolated

Tester
RepoRadar automated local verification harness
Started
2026-07-13T10:34:19.684257Z
Completed
2026-07-13T10:35:12.343566Z
Environment
Windows 10 AMD64; Python 3.11.9; credential-stripped child environment; disposable home/cache Windows 10 AMD64; Python 3.11.9; credential-stripped child environment; disposable home/cache
Install/setup time
1 minute(s)
Evidence scope
Bounded representative workflow
Cleanup
Per-check temporary home and work directory removed. Shared cohort package cache removed. Per-check temporary home and work directory removed. Shared cohort package cache removed.
Actions exercised
  • Created a disposable home, work directory, and isolated package cache with credential-like environment variables excluded.
  • Created 2 synthetic fixture file(s) inside the disposable work directory; retained hashes prove the exact inputs.
  • Configured a deterministic file:// provider, one prompt template, two variable rows, and exact-match assertions.
  • Ran the evaluation with cache and sharing disabled, then validated both expected outputs in the exported JSON result.
  • Executed bounded check: Run a two-case Promptfoo evaluation against a deterministic local JavaScript provider.
  • Captured the complete sanitized stdout, stderr, exit status, artifact checks, and 52.66-second wall time.
Observed results
  • Command exited 0 after 52.66 seconds.
  • Promptfoo evaluated both fixture rows through the local provider, enforced exact assertions, and exported both expected uppercase outputs.
  • Validated results.json: 2 required marker(s) present and 0 excluded marker(s) absent; size and SHA-256 are retained.
Observed strengths
  • The CLI executed a complete provider-by-test evaluation and machine-readable result export locally, making the assertion path reproducible without model credentials.
Friction
  • A model-free custom provider required a small JavaScript adapter; real provider testing would add credential management, latency, cost, and nondeterministic-output concerns absent here.
  • Setup or runtime emitted 5 stderr line(s); the complete warnings/errors are preserved in the retained log.
Limitations
  • The local uppercase provider validates Promptfoo's matrix, custom-provider, assertion, and result-export path; it does not measure model quality, grader behavior, red teaming, concurrency under API limits, or hosted sharing.
  • This credential-free disposable workflow does not establish production scale, model quality, reliability under sustained use, or team adoption.

Pricing assessment: The evaluation used Promptfoo's local open-source CLI, a fixture provider, and zero paid model or hosted-share calls.

Privacy assessment: Synthetic prompts stayed inside the local file provider; sharing, cache, and telemetry were disabled, though npm registry traffic was required to resolve the package.

Open retained test log →

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

79 dated snapshots retained from 2026-06-19 through 2026-09-27; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.7 current · +0.0 net
Repository momentum9.6 current · +1.6 net
GitHub stars (observed)25,479 current · +2,254 net
GitHub stars25,479 exact observation
Version0.123.1
Last release2026-09-18T01:07:54Z
MaintenanceActive
Current riskConditional
Current verdictTry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-278.79.625,479ConditionalTry nowActive
2026-09-268.79.625,466ConditionalTry nowActive
2026-09-258.79.625,447ConditionalTry nowActive
2026-09-248.79.625,423ConditionalTry nowActive
2026-09-238.79.625,401ConditionalTry nowActive
2026-09-228.79.025,370ConditionalTry nowActive
2026-09-218.79.025,342ConditionalTry nowActive
2026-09-208.79.625,307ConditionalTry nowActive
2026-09-198.79.025,286ConditionalTry nowActive
2026-09-178.79.025,232ConditionalTry nowActive
2026-09-158.79.625,118ConditionalTry nowActive
2026-09-138.79.025,054ConditionalTry nowActive

Why the record changed

Stars change

Stars changed: 25466 → 25479.

Stars change

Stars changed: 25447 → 25466.

Stars change

Stars changed: 25423 → 25447.

Stars change

Stars changed: 25410 → 25423.

Stars change

Stars changed: 25401 → 25410.

Stars change

Stars changed: 25370 → 25401.

Stars change

Stars changed: 25354 → 25370.

Stars change

Stars changed: 25342 → 25354.

Stars change

Stars changed: 25340 → 25342.

Stars change

Stars changed: 25318 → 25340.

Stars change

Stars changed: 25307 → 25318.

Stars change

Stars changed: 25297 → 25307.