Saved questions
Test chat answers one question, once. Saved questions keep the ones that matter and record what each should do — then re-run themselves every time you train, so "did anything change?" is a question you can actually answer without remembering to ask it.
You'll find them under Test chat, below the sandbox.
Two outcomes, and only two
When you save a question you say what your twin should do with it:
- Answered from sources — it should find this in your material and show where the answer came from.
- Honestly declined — you haven't covered this, and it should say so rather than improvise.
Those are the two things a good answer can be. A regression is a change between them: a question that used to be answered from your sources and now isn't, or one your twin used to decline and now answers anyway.
Run all
Run all asks each saved question again, one at a time, exactly the way the test chat does — same retrieval, same rules, same model a visitor gets. It is the same run the automatic post-train check makes, through the same code, so a result you asked for and a result that arrived by itself mean the same thing. Nothing is metered, nothing is logged to your Insights, and nothing lands in your gap queue.
Each row then shows what actually happened:
- Answered from sources — it showed evidence: a passage from your material, or a row from one of your lists.
- Honestly declined — it said it didn't have that.
- Answered without sources — it answered, but showed nothing to back it up. That is a change worth looking at whichever outcome you expected.
What you'll never see here
No pass rate. No score. No percentage.
A set of a dozen questions cannot produce a meaningful "83% correct", and a number like that would be exactly the kind of unauditable metric hiy refuses to put on a page. What you get is a count over a stated denominator — "2 of 12 checked questions no longer match what you expected" — and the questions themselves, so you can read the answers and decide.
hiy also never grades the wording of an answer. Whether it sounded good is your call; whether it was grounded is a fact, and that is the only thing measured.
They re-run themselves after every train
You do not have to remember this. When a train finishes, hiy re-runs your whole saved set against the index that just went live — automatically, on the server, whether or not you are looking at the page. The Train stage reports what it found: "2 of 12 saved questions no longer match what you expected."
This matters more than it sounds. A finished train is live: your twin answers visitors from the new material the moment training ends, and testing happens afterwards. So this run is the thing standing between a train that went wrong and a real person reading the result — which is why it is not something you have to click.
It costs you nothing. It is not a message off your monthly allowance and it is not a train off your monthly trains; one train triggers at most one re-run, so the trains you already have are what bounds it.
If a question can't be run — the answer times out, say — nothing is recorded for it, and it simply reads as unchecked against the new material. The row keeps its previous result rather than being marked as a regression it never had.
Any saved result from before a train is marked ran before your last rebuild — the answer on the row was produced by an index that no longer exists. Run all, below, is still there for whenever you want to check without waiting for a train.
In your email summary
If you have email summaries on, a regression is reported there too, and it appears even in a quiet week — a twin nobody talked to can still have drifted after a rebuild.
When your twin trained during the period the summary covers, the email reports that train's check in the Train stage's own words: "2 of 9 saved questions no longer match what you expected." Same counts, same sentence, so the screen and the inbox cannot tell you two different things about one train. A check still running when the summary goes out gets no verdict — a partial count is not a result.
Otherwise the summary reports where your set stands over every stored result, however old: "2 of 12 saved questions no longer answer the way you said they should." One sentence or the other, never both: two counts over two different denominators in one email is worse than one count.
Limits
You can save up to 12 questions per twin — enough to cover the things you'd actually check by hand, and few enough that "Run all" stays quick. Remove one to add another. Re-running is limited to a few full sets an hour.
A good set
- The three or four questions your visitors genuinely ask most.
- The one thing people get wrong about you, where the wording matters.
- Something you have deliberately not covered, saved as honestly declined — this is the one that catches a twin quietly starting to improvise.
- Anything you have already fixed once. A question you answered from the gap queue is the best possible saved question: you know what it should say now.