How to pick an AI assistant by the job, with a test for each
No assistant is best at everything, and general rankings say little about your task. Pick by the job: give each candidate a small real sample of that job, check the result against something you can verify, and score it on 3 tries. Writing needs your own voice kept. Code needs a fix that runs. Research needs sources that exist.
Asking which assistant is best is like asking which tool is best without naming the job. The faster route is to run the same small test on 2 or 3 candidates and read the results yourself.
On this page
Writing: does it keep your voice?
Paste 200 words that you wrote and ask for a tighter version with the same tone. Compare it with the original. A good result cuts words and leaves your meaning and style. A weak result rewrites everything into the same smooth, generic voice. Ask the same thing 3 times, because the results vary.
Code: does the fix run?
Give a small real bug with its error message and ask for a fix. Run the fix. Then ask for a test that would have caught the bug and run that too. Watch for functions or options that do not exist in your version, because invented names are the most common failure.
Research: do the sources exist?
Ask a question in a field you know well, where you can judge the answer. Ask for sources, then open every source. Check that each source exists, that it says what the answer claims and that the date fits. Note whether the tool can look up current pages or only works from its training text.
Long documents: does it find what is on page 40?
Paste a long text and ask 3 questions whose answers you know, 1 from the start, 1 from the middle and 1 from the end. Tools with a small context window lose the middle. Also ask what the text does not cover, and see if it admits gaps.
Translation and languages
Translate a paragraph you know well in both directions. Ask a native speaker or read it yourself for tone, not only meaning. Check that names, numbers and technical terms survive, and that the tool keeps the register you asked for.
Score it and decide
Give each try a score from 1 to 5 and average the 3 tries. Note what got in the way: limits, sign-in, speed or refusals. Choose the tool with the best average on the job you do most, and keep a second tool for the job it does better.
What people ask before switching
How many tries are enough?
3 tries per task is enough to see the range. A single try can flatter or fail a tool by chance.
Should I test with private material?
Use a sample without personal or confidential detail. You can strip names and numbers from a real example and keep the structure.
Do benchmark scores help?
They measure set tasks, not yours. Use them to shortlist, and your own test to decide.
What if 2 assistants score the same?
Pick the assistant with fewer limits and less friction in your daily routine. Small differences in answers matter less than a chat that gets in the way.
Run the writing test in this chat first, with 200 words of your own.
Open the chat