{"id":44193118,"title":"Show HN: Create LLM graders and run evals in JavaScript with one file","points":28,"user":"randall","time":1749140456,"time_ago":"2 days ago","type":"link","content":"

Hi HN!

Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo

We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is.

We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding.

Inspired by Anthropic's constitutional ai concepts, and amazing software like DSPy, we're setting out to make fine tuning prompts, not models, the default approach to improving quality using actual metrics and structured debugging techniques.

Our approach is pretty simple: you feed it a JSONL file with inputs and outputs, pick the models you want to test against (via OpenRouter), and then use an LLM-as-grader file in JS that figures out how well your outputs match the original queries.

If you're starting from scratch, we've found TDD is a great approach to prompt creation... start by asking an LLM to generate synthetic data, then you be the first judge creating scores, then create a grader and continue to refine it till its scores match your ground truth scores.

If you’re building LLM apps and care about reliability, I hope this will be useful! Would love any feedback. The team and I are lurking here all day and happy to chat. Or hit me up directly on Whatsapp: +1 (646) 670-1291

We have a lot bigger plans long-term, but we wanted to start with this simple (and hopefully useful!) tool.

Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo","url":"https://github.com/bolt-foundry/bolt-foundry/tree/main/packages/bff-eval","domain":"github.com","comments":[{"id":44193365,"level":0,"user":"rbalicki","time":1749142173,"time_ago":"2 days ago","content":"

Very cool! This lets you grade output across different base models. Does it also allow you grade output across different prompts?","comments":[{"id":44193382,"level":1,"user":"randall","time":1749142276,"time_ago":"2 days ago","content":"

that’s the next step… we have a structured approach to prompting too that we think will help people build better prompts too.","comments":[]}]}],"comments_count":2}