Promptfoo
MIT-licensed LLM eval and red-teaming CLI, now owned by OpenAI
Overview
Promptfoo is an open-source CLI for testing, evaluating and red-teaming LLM apps, agents and RAG pipelines across 50+ providers, with results wired into CI/CD. OpenAI acquired the company in March 2026 but Promptfoo stays MIT-licensed, model-agnostic, and free for individual use.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- All model provider integrations
- Red teaming up to 10k probes/month
- Self-hosted or local deployment
- Community support
- Custom red-team probe limits
- SSO & granular permissions
- Continuous monitoring & compliance dashboard
- Priority support with SLA
- Deploy on your own infrastructure
- Complete data isolation
- Dedicated runner
- Assigned deployment engineer
What you can produce with Promptfoo
- 50+ LLM provider integrations (GPT, Claude, Gemini, DeepSeek, etc.)
- Automated red-teaming for jailbreaks, prompt injection, PII leakage
- Declarative YAML eval configs with CI/CD integration
- LLM-as-a-judge scoring
- Self-hosted or local deployment, MIT-licensed core
- Enterprise SSO, compliance dashboard, managed cloud deployment
ASEAN Perspective
Promptfoo in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Promptfoo built its reputation as the eval tool developers actually reach for instead of rolling their own harness — YAML configs, 50-plus provider support, and a red-teaming mode that generates jailbreak and prompt-injection probes automatically. OpenAI's March 2026 acquisition is the headline: 350k+ developers had already adopted it and Fortune 500 penetration was above 25% before the deal closed. The open-source core stays MIT-licensed and provider-neutral, which is the reassurance skeptics needed after an acquisition by a company with an obvious rooting interest in its own models. The free tier caps red-team probes at 10k/month, which is plenty for most CI pipelines but tight for continuous fuzzing at scale, and Enterprise pricing is quote-only. Best fit: engineering teams already running LLM features in production who need pre-merge eval gates, not teams looking for a hosted dashboard with no setup.
What people say
Promptfoo went from a two-person open-source project to an OpenAI acquisition target in under two years. Founded in 2024 by Ian Webster and Michael D'Angelo, the 11-person team had pulled in $23 million and a roughly $86 million post-money valuation by its July 2025 Series A, with more than 350,000 developers using the tool and 130,000 active monthly. OpenAI announced the acquisition on March 9, 2026, saying it would fold Promptfoo's red-teaming into the Frontier agent platform while keeping the core project MIT-licensed and model-agnostic — a commitment OpenAI repeated publicly after early skepticism about whether a lab-owned eval tool could stay neutral.
What keeps developers on it: the config format is plain YAML, so a test suite reads like a spec rather than a Python script, and the same suite runs against GPT, Claude, Gemini, DeepSeek and dozens of other providers without rewriting anything. The red-teaming mode auto-generates adversarial prompts for jailbreaks, PII leakage and tool misuse, which is the piece most reviewers single out — it's less "does this pass my three example prompts" and more continuous vulnerability scanning wired into a pull request check. Teams at OpenAI and Anthropic itself are cited as users, which for a security tool is a stronger signal than a star rating.
The free Community tier is genuinely usable — full eval feature set, all providers, self-hosted or local deployment — but caps red-team probes at 10,000 a month, which pushes anyone doing serious continuous fuzzing toward an Enterprise quote for custom probe limits, SSO, a compliance dashboard and managed cloud deployment. One recurring practical note across reviews: Promptfoo calls the LLM APIs it's testing during evaluation runs, so token costs from your own provider accounts stack on top of whatever you pay Promptfoo — teams running large regression suites report meaningful API spend showing up on the model-provider side, not the Promptfoo invoice.
Post-acquisition, the open question reviewers keep raising is roadmap direction now that OpenAI owns it — whether feature priority stays genuinely provider-neutral once it's embedded in a competitor's own product line remains to be seen this early into the deal.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Promptfoo's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Promptfoo unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Promptfoo directly →
Spotted something out of date? Suggest an update →
Promptfoo in the news
More in Code & Dev Tools