NotTrack AI Open chat

Four tests worth 20 minutes

2026-08-19

Ask something you already know the answer to, something genuinely hard from your work, something near your subject's boundary, and then keep one conversation going for 15 messages. The 4 tests cover accuracy, capability, filter and memory, which is everything that will affect you.

Casual use tells you how a product feels and hides how it fails. Failures are what decide whether you keep using something, so they are worth provoking deliberately and early.

The known answer

A question from your field where you can judge the reply precisely. This is the only test that measures accuracy rather than fluency, and it is the one people skip because the answer is not useful to them. The point is not the answer; it is calibrating how much to trust the rest.

The genuinely hard question

Something you actually struggled with. Easy questions are answered well by everything, so they discriminate nothing. Difficulty is what separates products, and only your own work supplies the right difficulty.

The boundary case

A real example from your work that sits near a sensitive subject. Not a provocation: a legitimate question that a cautious filter might stop. This tells you whether the product is usable for your field, which no review will tell you.

The long conversation

Fifteen messages on one topic, then refer back to something from message 2. This measures the memory window and the drift, and it is where products differ most and reviews look least. Most disappointment with an assistant is really disappointment at message 30.

What people ask before switching

Should I use the same tests everywhere?

Yes, and write them down. Comparability is the entire value, and remembered questions drift between products.

Is 20 minutes really enough?

For eliminating unsuitable products, easily. For choosing between 2 good ones, a week of real work decides better.

What if it fails one test?

That depends which. Failing the boundary test rules a product out for your field; failing the long conversation is annoying and workable.

Run the 4 tests here and compare against wherever you are now.

Open the chat

Choosing between AI chats