Rumi

People Ops, on call.

Slice 2 · accuracy

Is Rumi getting better?

Rumi answers every question in content/09-eval/golden-questions.csv, and a second model grades each answer against the one you said you'd give. Run it before and after every change — that's the only way to know whether a change helped.

What you're measuring

60 questions against 26 documents. Add your real questions to content/09-eval/golden-questions.csv — the questions people actually ask, worded the way they actually ask them. Leave expected_answer blank for anything Rumi should refuse, and put what should happen in expected_behaviour.