Item detail
github.com

promptfoo/promptfoo

promptfoo/promptfoo is a developer tool in RepoRadar's Radar section, holding Gold tier and a 'try now' verdict. Its strongest signal is workflow potential, scored 9.8 out of 10.

Score8.7
Popularity80.0
Riskconditional
TierGold
Score breakdown
Usefulness9.0
Novelty8.0
Momentum8.0
Maturity8.4
Open-source/build8.4
Evidence7.2
Workflow potential9.8
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for teams that need a practical way to test AI behavior repeatedly instead of relying on one-off manual prompt checks.

Who should use it

LLM app teams AI product engineers security-minded builders developers maintaining RAG or agent workflows

Who should skip it

Move on from promptfoo/promptfoo if the licensing terms, language support, or platform requirements do not fit your project.

About this signal

promptfoo/promptfoo is tracked by RepoRadar as a developer tool in the Radar section. First seen 2026-06-19; the source record was last checked on 2026-06-19. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. promptfoo/promptfoo leads on workflow potential (9.8) and practical usefulness (9.0); its lowest signal is setup ease (6.4), so factor that in before investing setup time. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The promptfoo/promptfoo record combines a 8.7/10 composite score with separate popularity (80.0), risk (conditional), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.

Putting this into practice? Read How to evaluate an AI tool before you adopt it for the checklist behind this score.

Risk explanation

Evaluation runs can forward prompts, datasets, and model outputs to external providers, so scrub sensitive test data before using hosted model backends.

Evidence links
Closest alternatives / related signals
llm-evals red-teaming prompt-testing rag developer-tools
Verification record

What RepoRadar actually verified

Tested in a bounded workflow

Bounded representative workflow retained by RepoRadar verification harness. Last checked 2026-07-13T10:35:12.343566Z.

passed · cohort-20260712-promptfoo-isolated

Tester
RepoRadar automated local verification harness
Started
2026-07-13T10:34:19.684257Z
Completed
2026-07-13T10:35:12.343566Z
Environment
Windows 10 AMD64; Python 3.11.9; credential-stripped child environment; disposable home/cache
Install/setup time
1 minute(s)
Evidence scope
Bounded representative workflow
Cleanup
Per-check temporary home and work directory removed. Shared cohort package cache removed.
Actions exercised
  • Created a disposable home, work directory, and isolated package cache with credential-like environment variables excluded.
  • Created 2 synthetic fixture file(s) inside the disposable work directory; retained hashes prove the exact inputs.
  • Configured a deterministic file:// provider, one prompt template, two variable rows, and exact-match assertions.
  • Ran the evaluation with cache and sharing disabled, then validated both expected outputs in the exported JSON result.
  • Executed bounded check: Run a two-case Promptfoo evaluation against a deterministic local JavaScript provider.
  • Captured the complete sanitized stdout, stderr, exit status, artifact checks, and 52.66-second wall time.
Observed results
  • Command exited 0 after 52.66 seconds.
  • Promptfoo evaluated both fixture rows through the local provider, enforced exact assertions, and exported both expected uppercase outputs.
  • Validated results.json: 2 required marker(s) present and 0 excluded marker(s) absent; size and SHA-256 are retained.
Observed strengths
  • The CLI executed a complete provider-by-test evaluation and machine-readable result export locally, making the assertion path reproducible without model credentials.
Friction
  • A model-free custom provider required a small JavaScript adapter; real provider testing would add credential management, latency, cost, and nondeterministic-output concerns absent here.
  • Setup or runtime emitted 5 stderr line(s); the complete warnings/errors are preserved in the retained log.
Limitations
  • The local uppercase provider validates Promptfoo's matrix, custom-provider, assertion, and result-export path; it does not measure model quality, grader behavior, red teaming, concurrency under API limits, or hosted sharing.
  • This credential-free disposable workflow does not establish production scale, model quality, reliability under sustained use, or team adoption.

Pricing assessment: The evaluation used Promptfoo's local open-source CLI, a fixture provider, and zero paid model or hosted-share calls.

Privacy assessment: Synthetic prompts stayed inside the local file provider; sharing, cache, and telemetry were disabled, though npm registry traffic was required to resolve the package.

Open retained test log →

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

46 dated snapshots retained from 2026-06-19 through 2026-08-13; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.7 current · +0.0 net
Repository momentum9.6 current · +1.6 net
GitHub stars (observed)24,207 current · +982 net
GitHub stars24,207 exact observation
Version0.122.0
Last release2026-08-04T17:45:40Z
Maintenanceactive
Current riskconditional
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-08-138.79.624,207conditionaltry nowactive
2026-08-128.79.624,171conditionaltry nowactive
2026-08-118.79.624,121conditionaltry nowactive
2026-08-108.79.624,100conditionaltry nowactive
2026-08-098.79.324,074conditionaltry nowactive
2026-08-088.79.624,066conditionaltry nowactive
2026-08-078.79.623,867conditionaltry nowactive
2026-08-068.78.0Not recordedconditionaltry nownot recorded
2026-08-058.78.0Not recordedconditionaltry nownot recorded
2026-08-048.79.623,867conditionaltry nowactive
2026-08-038.79.623,867conditionaltry nowactive
2026-08-028.79.623,837conditionaltry nowactive

Why the record changed

stars changed

Stars changed: 24171 → 24207.

stars changed

Stars changed: 24121 → 24171.

stars changed

Stars changed: 24100 → 24121.

stars changed

Stars changed: 24074 → 24100.

stars changed

Stars changed: 24066 → 24074.

stars changed

Stars changed: 23867 → 24066.

version changed

Version changed: 0.121.20 → 0.122.0.

stars changed

Stars changed: 23837 → 23867.

stars changed

Stars changed: 23809 → 23837.

stars changed

Source-observed stars changed: 23798 → 23809. This reports the retained observation delta and does not infer why the upstream change occurred.

stars changed

Stars changed: 23682 → 23719.

stars changed

Source-observed stars changed: 23677 → 23682. This reports the retained observation delta and does not infer why the upstream change occurred.