Imagine your AI assistant recommends the option you expected, and the room decides the trial went well.
The comparison is tidy. The reasoning sounds sensible. The manager sees the same winner they had in mind and says, “That’s what we needed.”
But nobody tried the request with the requirements missing. Nobody gave it two notes that disagreed. The only thing you have seen is what happens when the assignment arrives in the shape you prepared for the demonstration.
If you manage the work after that meeting, you inherit the other shapes.
Here’s the thing. You are calling the station ready without ever deciding what it must do when the information is missing or disagrees.
That is a decision you can make before the real work gets involved. Write what should happen before you look at what happened.
Keep one job on the table
We have been building one hypothetical role through the AI Employee Handbook series. Job Spec defined where the work ends. Standards settled what makes the result usable. Now we try that written job against a few prepared requests.
The assistant compares Option A and Option B using approved notes. The manager makes the eventual choice. The assistant cannot buy anything, contact a vendor, invent requirements, or turn unanswered questions into facts.
The two requirements in the ordinary request are a weekly list export and separate access for people who view the work and people who can change it. Both options support the export in our invented notes. B’s notes confirm separate access. A’s notes leave access unanswered.
These are teaching examples, not a report of a company trial. The notes and results discussed below are illustrative throughout.
In the Professional Recipe, Context carries the current request. Format describes the result that comes back. Guardrails are the hard stops. Trying the job means checking whether those written instructions hold when the request gets less convenient.
You are looking for behavior you can name. Did the answer keep the unknown visible? Did it ask for the missing input? Did it hide the disagreement to produce a winner?
Write the expectation while you still have a clear head
Before you run the first request, write what would count as acceptable behavior.
For the ordinary case, B is the better-supported choice against the two supplied requirements. A’s access remains unanswered. The comparison should explain that reasoning and point to the notes behind it.
That expectation belongs in the checker’s record. The assistant receives its written job, the agreed standard, and the request. You are not handing it a special answer to copy for this one demonstration.
“I’ll know a good answer when I see it.”
Maybe. You may also see an answer that sounds so reasonable that you forget the part it was supposed to preserve.
Suppose the comparison recommends B but explains the choice by saying A cannot separate view and edit access. You liked the winner. The reasoning changed an unknown into a fact.
Right? Same winner, failed requirement.
Writing the expectation first gives you something to compare against when a pleasant answer invites you to overlook the mistake. The wording can vary. The evidence boundary cannot.
First, try the ordinary request
Start with the complete request: named options, stated requirements, approved notes, and the requesting manager. Complete means the required inputs are present. It does not mean every fact about every option is known.
A’s access question is still open. That is part of the ordinary case, not a defect you need to fill in before the exercise.
The expected answer recommends B as better supported by the notes, explains the export and access comparison, cites the supplied evidence, and keeps A’s unanswered point beside the recommendation.
This gives you a starting check. Can the assistant do the ordinary job without overstating what it knows?
If it cannot, you have something specific to investigate before you add more responsibility. You do not need to pretend a harder case will rescue an answer that already missed the basic rule.
Then take the requirements out
Keep the two options and the approved notes. This time, leave the buying requirements out of the request.
Write the expected behavior before sending it: the assistant should identify the missing requirements and ask the requesting manager to supply them. It should not rank a winner against criteria it invented.
“But we always care about the weekly export and the access settings.”
Then include them in the request when they apply. The job is to compare against stated requirements, and this request has none. Your usual preference does not become today’s instruction by being familiar to you.
Use a fresh attempt for each case. If the previous conversation already supplied the requirements, you have not actually tested what happens when they are missing. Keep approved examples labeled as examples too, so their details do not masquerade as the current order.
An acceptable result might say:
“The request does not state the requirements. Please provide them before I recommend an option.”
It could summarize facts from the notes if that is useful and clearly separate from a recommendation. The important behavior is naming the missing input without quietly choosing it for the manager.
Yeah. The answer that feels less finished may be the one that did the job correctly.
Now let the notes disagree
For the third case, put both requirements back. Change the invented evidence instead.
One approved note says B has separate view and edit access. Another equally approved note says everyone can edit. Neither is labeled as a correction, and nothing says which one takes priority. A’s access is still unanswered.
Now write the expectation. The assistant should show the conflict, point to both notes, keep A’s unknown visible, and explain why the supplied evidence does not yet support a choice.
“It probably means everyone on the paid plan can edit.”
Probably is doing work the notes did not authorize. So is assuming one note is newer because its wording looks more polished. The assistant needs to preserve the disagreement rather than manufacture a way around it.
Here’s the thing. A confident recommendation is easy to applaud. A clear account of why a recommendation cannot yet be made may protect the decision better.
That does not mean every uncertainty must stop every task. This conflict concerns a stated requirement for the very choice the assistant is making. It belongs in the answer because it can change the decision.
Does that make sense? You are checking whether the assistant can stay useful when producing a winner would mean hiding the problem.
Keep the result you actually got
Your record can be simple: case, expected behavior, actual answer, judgment, and reason. Note which tool you used and when, with the version if it is available.
Here is an illustrative entry, not a result from an actual run:
- Case: Conflicting access notes.
- Expected: Show both notes and withhold a supported choice until the conflict is resolved.
- Sample answer: “B meets both requirements and is the recommended option.”
- Judgment: Needs revision. The answer drops the conflicting evidence.
You have literally put the expectation and the observed answer beside each other. There is less room to remember the version you hoped you saw.
If you find a miss, keep that answer in the record. Diagnose it. Was the instruction unclear? Did the wrong notes go into the request? Did the tool ignore a rule it had clearly been given?
Those causes need different fixes. Better wording is not a universal cure, and a small trial should not pretend otherwise.
Correct what the evidence supports, then try the affected case again. Recheck the ordinary case too, so a new caution rule does not leave the assistant refusing a comparison it can reasonably complete.
“We changed the expected result because its answer was actually pretty good.”
Sometimes the human standard needs correction. Record that as a separate decision before another attempt. Do not quietly move the line after seeing the answer and count the old attempt as a success.
Your Monday Move
Choose one existing AI-assisted task whose job and standard are already written. Prepare three requests using safe invented information: ordinary, missing an essential input, and containing a meaningful contradiction.
Keep the exercise away from customer work and connected actions. No purchases, outside contact, private records, or added permission. Use an existing approved tool, and keep each attempt separate so one case does not supply another case’s missing context.
Write the expected behavior first. Then run each request and keep the actual answer. Have the normal checker mark whether it met the expectation and explain why. If a case fails, name the correction or the cause that still needs investigation.
The finish line is three honest records. It is not three green checks obtained by helping the assistant through every hard part.
Actually, keep the claim smaller than that. Even three clean attempts only show that those attempts met those expectations. They do not establish how often the assistant will succeed in ordinary business use, and they do not authorize broader permissions on their own.
You are trying to find a problem while it is still cheap to examine. A clear miss can be a useful outcome if you preserve it and act on what it shows.
Before the next demonstration convinces you the job is ready, decide what must happen on the difficult request.
Write what should happen before you look at what happened.
Original framework. Distilled from client work.
