You Probably Don't Use AI the Way You Think You Do
Think you know what you use AI for? You probably don't. This repeatable self-audit shows whether you're relying on AI too much — and lets you decide.
Copilot Vision reads your screen directly. Putting a picture into words drops the one detail that answers your question. Show it, after one check.
Microsoft started shipping vision into Microsoft 365 Copilot in late June, worldwide, on by default with an admin control. During a voice chat you share your desktop screen or point your phone camera at something, and by Microsoft’s description Copilot converts what it sees into data it can reason over, including text, charts, tables and app interfaces, grounds the answer against your work files and email, and replies out loud while the thing stays on screen in front of you.
What the feature reveals matters more than the feature. A fair share of what people have spent two years calling “the AI didn’t understand me” was never a failure of the model. It was the summary they typed.
Every time you convert something visual into prose you pay a description tax, and you pay it twice. The first cost is visible: the minutes spent typing out the error dialog, the follow-up round when the answer comes back aimed at a slightly different problem. The second cost is the expensive one and you cannot see it at all. When you describe a chart, you decide which features to mention, and you make that decision before you know which one mattered. The detail you dropped is by definition the detail you did not notice. There is no way to audit your own description, because the audit would require the thing you already left out.
Psychologists have a name for a version of this. Schooler and Engstler-Schooler showed in 1990 that people who wrote a description of a face they had just seen were markedly worse at picking it out of a lineup afterwards, roughly 25 percent worse than people who spent the same time listing US state capitals. The effect got a rough ride in the replication years, and a preregistered many-lab replication in 2014 found it real but smaller and very sensitive to timing, about 4 percent when the description came immediately and 16 percent after a twenty-minute delay. That research is about your own memory of a face and not about briefing software, so it proves nothing about Copilot. It is still the cleanest demonstration I know that putting a picture into words is a lossy operation rather than a neutral transfer, and that the loss lands somewhere you were not looking.
Notice which screens are hardest to describe. The dashboard someone else built, with a column whose label you have never been sure about. The chart where the trend is obvious to your eye and refuses to become a sentence. The error dialog from a system nobody trained you on. A form in a tool you inherited from a colleague who left.
Those are the same screens you most need help with. The relationship runs the wrong way from the advice we usually give: being able to describe a visual precisely and needing to ask about it are close to mutually exclusive. Anyone who can characterise a chart accurately enough for a model to work with it has already done most of the interpreting. The cases where a careful description is easy are the cases where you did not need the help.
So the default flips for a whole class of work. When you catch yourself narrating your screen for the second time, that is not a prompting problem to solve with better wording. Share the screen and ask the question out loud. Point the camera at the machine throwing the error instead of transcribing the error. Let it read the dashboard and tell you which number looks wrong, rather than you choosing in advance which number to describe.
One habit has to sit in front of that instinct or it will eventually cost you something worse than a wasted round. A screen is almost never just the panel you care about. The customer list is in the next window, the salary column is still open behind it, the account number is sitting in the header. Microsoft’s documentation is reasonably reassuring on the mechanics: sharing is user-initiated and session-bound, the content is processed as a series of images, and audio and video are deleted after 48 hours. None of that changes the basic fact that you are sending everything in frame. The rules you already apply to text, the ones covered in what’s actually safe to put into AI at work, apply to pixels with no adjustment. A screenshot of a spreadsheet is the spreadsheet.
So the habit is two questions, asked before you share anything.
1. Am I stuck because I can't get this into words?
-> Yes: show it, don't type it.
2. Is there anything in this frame I wouldn't paste as plain
text into the AI box?
-> Yes: close it, crop it, or share the single app instead
of the whole desktop.
The obvious objection to all of this is that if showing beats describing, you should simply always show. Two things argue against it. The second question above is the first, and it does not get easier as the habit becomes automatic. The other is that the tool is narrower than the demo suggests. Vision works only inside a voice chat at the moment, it counts against your daily voice usage, it cannot click or type or fix anything, it keeps nothing between sessions, and Microsoft’s own docs warn that flipping between windows too quickly while you ask can get you a confident answer about the wrong screen. That one will happen to you at some point.
What you have is a colleague you can turn your monitor toward. Useful, fast, and finished the moment you understand what you were looking at. Hand over the looking. The deciding was never on the table, and it is still the reason anyone needs you in the room. If none of this feels within reach yet because you have not made Copilot do anything real, start with one useful task in twenty minutes and add the screen afterwards.
Next time you start typing a description of something on your screen, stop and run the frame check. If it passes, close the windows you don't need, start a Copilot voice chat, share that one screen, and ask your question out loud. Count how much you didn't have to explain, and how many rounds you didn't spend.