Simulates user scenarios against LLM agents, scores the answers and traces every step in production
LangWatch is a testing and monitoring platform for LLM agents. It runs simulated user scenarios against an agent to surface failures before release, scores responses for quality and accuracy with configurable evaluators, and traces each agent step in production with cost and latency attached. Around that sit prompt versioning, red-teaming, voice-agent testing and governance reporting, which is why it turns up filed under AI security as often as under observability. The core is open source under Apache 2.0, so self-hosting is an option. The cloud Developer tier is free; Growth is billed per core seat in euros with an included event allowance. LangWatch B.V. is based in Amsterdam.
Developer is free forever: 50k events a month, 14-day data access, 2 users, 3 scenarios. Growth is priced in euros — EUR 29 per core seat per month with 200k events included, then EUR 5 per additional 100k, 30-day retention — so no USD figure is published. Enterprise is quote-only.
Use tool ↗The scenario simulation is what sets it apart: instead of only watching production traces, it drives synthetic users at your agent so failures show up before release. Apache 2.0 licensing means you can self-host if the euro seat pricing or data residency bothers you.
Watch out: Growth is priced per 'core seat' rather than per user — confirm which of your engineers count as core seats before budgeting.
No reviews yet — be the first to review LangWatch.
Reviews are tied to your EffectHub account — one review per tool, so the rating for LangWatch reflects real users.
No questions yet — ask the first one about LangWatch.
Questions about LangWatch are tied to your EffectHub account, so answers can reach you and the section stays free of spam.