

Agent-Eval
Statistical regression testing for LLM agents
什麼是 Agent-Eval?
Statistical regression testing for LLM agents. Run versions A and B 50 times to get a p-value, Cohen's d, and a 95% CI proving whether behavior actually shifted. While DeepEval, Braintrust, and Promptfoo test single responses against a threshold, none measure distribution drift. It is completely self-hostable under an Apache 2.0 license and requires no SaaS subscriptions. Works natively with LangGraph, OpenAI Agents SDK, CrewAI, and LangChain LCEL. Start testing: `pip install agent-regress-cli`.
截圖
?
還沒有評論,來搶沙發吧!
X 上關於 Agent-Eval 的真實討論
去 X 發文



