Slice 2 · accuracy
Is Rumi getting better?
Rumi answers every question in content/09-eval/golden-questions.csv, and a second model grades each answer against the one you said you'd give. Run it before and after every change — that's the only way to know whether a change helped.
What you're measuring
60 questions against 26 documents. Add your real questions to content/09-eval/golden-questions.csv — the questions people actually ask, worded the way they actually ask them. Leave expected_answer blank for anything Rumi should refuse, and put what should happen in expected_behaviour.